<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>A Comparison of Representation Models in a Non-Conventional Semantic Similarity Scenario</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Andrea Amelio Ravelli</string-name>
          <email>andreaamelio.ravelli@unifi.it</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Oier Lopez de Lacalle and Eneko Agirre</string-name>
          <email>e.agirre@ehu.eus</email>
          <email>e.agirre@ehu.eus oier.lopezdelacalle@ehu.eus</email>
          <email>oier.lopezdelacalle@ehu.eus</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>University of Florence</institution>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>University of the Basque Country</institution>
        </aff>
      </contrib-group>
      <abstract>
        <p>Representation models have shown very promising results in solving semantic similarity problems. Normally, their performances are benchmarked on well-tailored experimental settings, but what happens with unusual data? In this paper, we present a comparison between popular representation models tested in a nonconventional scenario: assessing action reference similarity between sentences from different domains. The action reference problem is not a trivial task, given that verbs are generally ambiguous and complex to treat in NLP. We set four variants of the same tests to check if different pre-processing may improve models performances. We also compared our results with those obtained in a common benchmark dataset for a similar task.1</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1 Introduction</title>
      <p>
        Verbs are the standard linguistic tool that
humans use to refer to actions, and action verbs are
very frequent in spoken language ( 50% of total
verbs occurrences)
        <xref ref-type="bibr" rid="ref10">(Moneglia and Panunzi, 2007)</xref>
        .
These verbs are generally ambiguous and
complex to treat in NLP tasks, because the relation
between verbs and action concepts is not one-to-one:
e.g. (a) pushing a button is cognitively separated
from (b) pushing a table to the corner; action (a)
can also be predicated through press, while move
can be used for (b) and not vice-versa
        <xref ref-type="bibr" rid="ref11 ref12">(Moneglia,
2014)</xref>
        . These represent two different pragmatic
actions, despite of the verb used to describe it, and
all the possible objects that can undergo the
action. Another example could be the ambiguity
behind a sentence like John pushes the bottle: is the
1Copyright c 2019 for this paper by its authors. Use
permitted under Creative Commons License Attribution 4.0
International (CC BY 4.0).
agent applying a continuous and controlled force
to move the object from position A to position B,
or is he carelessly shoving an object away from its
location? These are just two of the possible
interpretation of this sentence as is, without any other
lexical information or pragmatic reference.
      </p>
      <p>
        Given these premises, it is clear that the task
of automatically classifying sentences referring to
actions in a fine-grained way (e.g. push/move vs.
push/press) is not trivial at all, and even humans
may need extra information (e.g. images, videos)
to precisely identify the exact action. One way
could be to consider action reference similarity
as a Semantic Textual Similarity (STS) problem
        <xref ref-type="bibr" rid="ref1">(Agirre et al., 2012)</xref>
        , assessing that lexical
semantic information encodes, at a certain level, the
action those words are referring to. The simplest
way is to make use of pre-computed word
embeddings, which are ready to use for computing
similarity between words, sentences and documents.
Various models have been presented in the past
years that make use of well-known static word
embeddings, like word2vec, GloVe and FastText
        <xref ref-type="bibr" rid="ref17 ref2 ref9">(Mikolov et al., 2013; Pennington et al., 2014;
Bojanowski et al., 2017)</xref>
        . Recently, the best STS
models rely on representations obtained from
contextual embeddings, such as ELMO, BERT and
XLNet
        <xref ref-type="bibr" rid="ref18 ref6">(Peters et al., 2018; Devlin et al., 2018;
Yang et al., 2019)</xref>
        .
      </p>
      <p>In this paper, we are testing the effectiveness of
representation models in a non-conventional
scenario, in which we do not have labeled data to
train STS systems. Normally, STS is performed
on sentence pairs that, on one hand, can have very
close or distinct meaning, i.e. the assertion of
similarity is easy to formulate; on the other hand, all
sentences derive from the same domain, thus they
share some syntactic regularities and vocabulary.
In our scenario, we are computing STS between
textual data from two different resources,
IMAGACT and LSMDC16 (described respectively in
5.1 and 5.2), in which the language used is highly
different: from the first, synthetic and short
captions; from the latter, audio descriptions. The
objective is to benchmark word embedding models in
the task of estimating the action concept expressed
by a sentence.
2</p>
    </sec>
    <sec id="sec-2">
      <title>Related Works</title>
      <p>Word embeddings are abstract representations of
words in the form of dense vectors, specifically
tailored to encode semantic information. They
represent an example of the so called transfer
learning, as the vectors are built to minimize
certain objective function (i.e., guessing the next
word in a sentence), but successfully applied on
different unrelated tasks, such as searching for
words that are semantically related. In fact,
embeddings are typically tested on semantic
similarity/relatedness datasets, where a comparison of the
vectors of two words is meant to mimic a human
score that assesses the grade of semantic similarity
between them.</p>
      <p>
        The success of word embeddings on
similarity tasks has motivated methods to learn
representations of longer pieces of text such as
sentences
        <xref ref-type="bibr" rid="ref14">(Pagliardini et al., 2017)</xref>
        , as representing
their meaning is a fundamental step on any task
requiring some level of text understanding.
However, sentence representation is a challenging task
that has to consider aspects such as
compositionality, phrase similarity, negation, etc. The
Semantic Textual Similarity (STS) task
        <xref ref-type="bibr" rid="ref3">(Cer et al., 2017)</xref>
        aims at extending traditional semantic
similarity/relatedness measures between pair of words in
isolation to full sentences, and is a natural dataset
to evaluate sentence representations. Through a
set of campaigns, STS has distributed set of
manually annotated datasets where annotators measure
the similarity among sentences with a score that
ranges between 0 (no similarity) to 5 (full
equivalence).
      </p>
      <p>
        In the recent years, evaluation campaigns that
agglutinate many semantic tasks have been set
up, with the objective to measure the
performance of many natural language understanding
systems. The most well-known benchmarks are
SentEval2
        <xref ref-type="bibr" rid="ref15 ref16 ref19 ref5">(Conneau and Kiela, 2018)</xref>
        and GLUE3
        <xref ref-type="bibr" rid="ref26">(Wang et al., 2019)</xref>
        . They share many of existing
2https://github.com/facebookresearch/
SentEval
3https://gluebenchmark.com/
tasks and datasets, such as sentence similarity.
3
      </p>
    </sec>
    <sec id="sec-3">
      <title>Problem Formulation</title>
      <p>We cast the problem as a fine-grained action
concept classification for verbs in LSMDC16 captions
(e.g. push as move vs push as press, see
Figure 1). Given a caption and the target verb from
LSMDC16, our aim is to detect the most
similar caption in IMAGACT that describe the action.
The inputs to our model are the target caption and
an inventory of captions that categorize the
possible action concepts of the target verb. The model
ranks the captions in the inventory according to
the textual similarity with the target caption, and,
similar to a kNN classifier, the model assigns the
action label of k most similar captions.
4</p>
    </sec>
    <sec id="sec-4">
      <title>Representation Models</title>
      <p>In this section we describe the pretrained
embeddings used to represent the contexts. Once we get
the representation of each caption, the final
similarity is computed based on cosine of the two
representation vectors.
4.1</p>
      <sec id="sec-4-1">
        <title>One-hot Encoding</title>
        <p>
          This is the most basic textual representation, in
which text is represented as binary vector
indicating the words occurring in the context
          <xref ref-type="bibr" rid="ref8">(Manning
et al., 2008)</xref>
          . This way of representing text creates
long and sparse vectors, but it has been
successfully used in many NLP tasks.
4.2
        </p>
      </sec>
      <sec id="sec-4-2">
        <title>GloVe</title>
        <p>
          The Global Vector model (GloVe)4
          <xref ref-type="bibr" rid="ref17">(Pennington
et al., 2014)</xref>
          is a log-linear model trained to
encode semantic relationships between words as
vector offsets in the learned vector space, combining
global matrix factorization and local context
window methods.
        </p>
        <p>Since GloVe is a word-level vector model, we
compute the mean of the vectors of all words
composing the sentence, in order to obtain the
sentence-level representation. The pre-trained
model from GloVe considered in this paper is the
6B-300d, counting a vocabulary of 400k words
with 300 dimensions vectors and trained on a
dataset of 6 billion tokens.</p>
        <p>
          4https://nlp.stanford.edu/projects/
glove/
The Bidirectional Encoder Representations from
Transformer (BERT)5
          <xref ref-type="bibr" rid="ref6">(Devlin et al., 2018)</xref>
          implements a novel methodology based on the so called
masked language model, which randomly masks
some of the tokens from the input, and predicts the
original vocabulary id of the masked word based
only on its context.
        </p>
        <p>Similarly with GloVe, we extract the token
embeddings of the last layer, and compute the mean
vector to obtain the sentence-level representation.
The BERT model used in our test is the
BERTLarge Uncased (24-layer, 1024-hidden, 16-heads,
340M parameters).
4.4</p>
        <p>
          USE
The Universal Sentence Encoder (USE)
          <xref ref-type="bibr" rid="ref4">(Cer et al.,
2018)</xref>
          is a model for encoding sentences into
embedding vectors, specifically designed for
transfer learning in NLP. Based on a deep averaging
network encoder, the model is trained for a
variety text length, such as sentences, phrases or short
paragraphs, and in a variety of semantic task
including the STS. The encoder returns the
corresponding vector of the sentence, and we compute
similarity using cosine formula.
5
        </p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>Datasets</title>
      <p>In this section, we briefly introduce the resources
used to collect sentence pairs for our similarity
test. Figure 1 shows some examples of data,
aligned by action concepts.</p>
      <sec id="sec-5-1">
        <title>5.1 IMAGACT</title>
        <p>
          IMAGACT6
          <xref ref-type="bibr" rid="ref11 ref12">(Moneglia et al., 2014)</xref>
          is a
multilingual and multimodal ontology of action that
provides a video-based translation and
disambiguation framework for action verbs. The resource
is built on an ontology containing a fine-grained
categorization of action concepts (acs), each
represented by one or more visual prototypes in the
form of recorded videos and 3D animations.
IMAGACT currently contains 1,010 scenes, which
encompass the actions most commonly referred to in
everyday language usage.
        </p>
        <p>
          Verbs from different languages are linked to
acs, on the basis of competence-based annotation
from mother tongue informants. All the verbs
5https://github.com/google-research/
bert
6http://www.imagact.it
that productively predicates the action depicted in
an ac video are in local equivalence relation
          <xref ref-type="bibr" rid="ref15 ref16">(Panunzi et al., 2018b)</xref>
          , i.e the property that
different verbs (even with different meanings) can
refer to the same action concept. Moreover, each
ac is linked to a short synthetic caption (e.g. John
pushes the button) for each locally equivalent verb
in every language. These captions are formally
defined, thus they only contain the minimum
arguments needed to express an action.
        </p>
        <p>
          We exploited IMAGACT conceptualization due
to its action-centric approach. In fact, compared
to other linguistic resources, e.g. WordNet
          <xref ref-type="bibr" rid="ref7">(Fellbaum, 1998)</xref>
          , BabelNet
          <xref ref-type="bibr" rid="ref1 ref13">(Navigli and Ponzetto,
2012)</xref>
          , VerbNet
          <xref ref-type="bibr" rid="ref23">(Schuler, 2006)</xref>
          , IMAGACT
focuses on actions and represents them as visual
concepts. Even if IMAGACT is a smaller
resource, its action conceptualization is more
finegrained. Other resources have more broad scopes,
and for this reason senses referred to actions
are often vague and overlapping
          <xref ref-type="bibr" rid="ref15 ref16">(Panunzi et al.,
2018a)</xref>
          , i.e. all possible actions can be gathered
under one synset. For instance, if we look at the
senses of push in Wordnet, we find that only 4 out
of 10 synsets refer to concrete actions, and some
of the glosses are not really exhaustive and can be
applied to a wide set of different actions:
push, force (move with force);
push (press against forcefully without
moving);
push (move strenuously and with effort);
press, push (make strenuous pushing
movements during birth to expel the baby).
        </p>
        <p>In such framework of categorization, all
possible actions referred by push can be gathered under
the first synset, except from those specifically
described by the other three.</p>
        <p>For the experiments proposed in this paper, only
the English captions have been used, in order to
test our method in a monolingual scenario.
5.2</p>
      </sec>
      <sec id="sec-5-2">
        <title>LSMDC16</title>
        <p>
          The Large Scale Movie Description Challenge
Dataset7 (LSMDC16)
          <xref ref-type="bibr" rid="ref21">(Rohrbach et al., 2017)</xref>
          consists in a parallel corpus of 128,118 sentences
obtained from audio descriptions for visually
impaired people and scripts, aligned to video clips
7https://sites.google.com/site/
describingmovies/home
from 200 movies. This dataset derives from the
merging of two previously independent datasets,
MPII-MD
          <xref ref-type="bibr" rid="ref20">(Rohrbach et al., 2015)</xref>
          and M-VAD
          <xref ref-type="bibr" rid="ref25">(Torabi et al., 2015)</xref>
          . The language used in
audio descriptions is particularly rich of references
to physical action, with respect to reference
corpora (e.g. BNC corpus)
          <xref ref-type="bibr" rid="ref22">(Salway, 2007)</xref>
          .
        </p>
        <p>For this reason, LSMDC16 dataset could be
considered a good source of video-caption pairs of
action examples, comparable to data from
IMAGACT resource.
6</p>
      </sec>
    </sec>
    <sec id="sec-6">
      <title>Experiments</title>
      <p>Given that the objective is not to discriminate
distant actions (e.g. opening a door vs. taking a
cup) but rather to distinguish actions referred to
by the same verb or set of verbs, the experiments
herein described have been conducted on a sub-set
of the LSMDC16 dataset, that have been manually
annotated with the corresponding acs from
IMAGACT. The annotation has been carried on by one
expert annotator, trained on IMAGACT
conceptualization framework, and revised by a supervisor.
In this way, we created a Gold Standard for the
evaluation of the compared systems.
6.1</p>
      <sec id="sec-6-1">
        <title>Gold Standard</title>
        <p>The Gold Standard test set (GS) has been created
by selecting one starting verb: push. This verb has
been chosen according to the fact that, as a general
action verb, it is highly frequent in the use, it
applies to a high number of acs in the IMAGACT
Ontology (25 acs) and it has a high occurrence
both in IMAGACT and LSMDC16.</p>
        <p>From the IMAGACT Ontology, all the verbs in
relation of local equivalence with push in each of
its acs have been queried8, i.e all the verbs that
predicate at least one of the acs linked to push.
Then, all the captions in LSMDC16 containing
one of those verbs have been manually annotated
with the corresponding ac’s id. In total, 377
videocaption pairs have been correctly annotated9 with
18 acs, and they have been paired with 38
captions for the verbs linked to the same acs in
IMAGACT, consisting in a total of 14,440 similarity
8The verbs collected for this experiment are: push, insert,
press, ram, nudge, compress, squeeze, wheel, throw, shove,
flatten, put, move. Move and put have been excluded from
this list, due to the fact that this verbs are too general and
apply to a wide set of acs, with the risk of introducing more
noise in the computation of the similarity; flatten is connected
to an ac that found no examples in LSMDC16, so it has been
excluded too.</p>
        <p>9Pairs with no action in the video, or pairs with a novel or
difficult to assign ac have been excluded from the test.
judjements.</p>
        <p>It is important to highlight that the manual
annotation took into account the visual information
conveyed with the captions (i.e. videos from both
resources), that made possible to precisely assign
the most applicable ac to the LSMDC16 captions.
6.2</p>
      </sec>
      <sec id="sec-6-2">
        <title>Pre-processing of the data</title>
        <p>As stated in the introduction, STS methods are
normally tested on data within the same domain.
In attempt to leverage some differences between
IMAGACT and LSMDC16, basic pre-processing
have been applied.</p>
        <p>
          Length of caption in the two resources vary:
captions in IMAGACT are artificial, and they only
contain minimum syntactic/semantic elements to
describe the ac; captions in LSMDC16 are
transcription of more natural spoken language, and
usually convey information on more than one
action at the same time. For this reason, LSMDC16
captions have been splitted in shorter and
simpler sentences. To do that, we parsed the
original caption with StanforNLP
          <xref ref-type="bibr" rid="ref19">(Qi et al., 2018)</xref>
          , and
rewrote simplified sentences by collecting all the
words in a dependency relation with the targeted
verbs. Table 1 shows an example of the splitting
process.
        </p>
        <p>FULL As he crashes onto the platform,
someone hauls him to his feet
and pushes him back towards
someone.</p>
        <p>SPLIT he crashes onto the platform and
As someone hauls him to his feet
pushes him back towards
someone
3</p>
        <p>LSMDC16 dataset is anonymised, i.e. the
pronoun someone is used in place of all proper names;
on the contrary, captions in IMAGACT always
have a proper name (e.g. John, Mary). We
automatically substituted IMAGACT proper names
with someone, to match with LSMDC16.</p>
        <p>Finally, we also removed stop-words, which
are often the first lexical elements to be pruned
out from texts, prior of any computation, because
they do not convey semantic information, and they
sometimes introduce noise in the process.
Stopwords removal has been executed in the moment
of calculating the similarity between caption pairs,
i.e. tokens corresponding to stop-words have been
used for the representation by contextual models,
but then discharged when computing sentence
representation.</p>
        <p>With these pre-processing operations, we
obtained 4 variants of testing data:
plain (LSMDC16 splitting only);
anonIM (anonymisation of IMAGACT
captions by substitution of proper names with
someone);
noSW (stop-words removing from both
resources);
anonIM+noSW (combination of the two
previous ones).
7</p>
      </sec>
    </sec>
    <sec id="sec-7">
      <title>Results</title>
      <p>To benchmark the performances of the four
models, we also defined a baseline that, following a
binomial distribution, randomly assigns an ac of
the GS test set (actually, baseline is calculated
analytically without simulations). Parameters of the
binomial are calculated from the GS test set. Table
2 shows the results at different recall@k (i.e. ratio
of examples containing the correct label in the top
k answers) of the three models tested.</p>
      <p>All models show slightly better results
compared to the baseline, but they are not much
higher. Regarding the pre-processing, any
strategy (noSW, anonIM, anonIM+noSW) seems not
to make difference. We were expecting low
results, given the difficulty of the task: without
taking into account visual information, also for a
human annotator most of those caption pairs are
ambiguous.</p>
      <p>Surprisingly, GloVe model, the only one with
static pre-trained embeddings based on statistical
distribution, outperforms the baseline and other
contextual models by 0.2 in recall@10. It is
not an exciting result, but it shows that STS with
pre-trained word embedding might be effective to
speed up manual annotation tasks, without any
computational cost. Probably, one reason to
explain the lower trend in results obtained by
contextual models (BERT, USE) could be that these
systems have been penalized by the splitting
process of LSMDC16 captions. Example in Table</p>
      <sec id="sec-7-1">
        <title>Model</title>
        <p>ONE-HOT ENCODING</p>
        <sec id="sec-7-1-1">
          <title>GLOVE</title>
        </sec>
        <sec id="sec-7-1-2">
          <title>BERT USE</title>
          <p>Pre-processing
plain
noSW
anonIM
anonIM+noSW
plain
noSW
anonIM
anonIM+noSW
plain
noSW
anonIM
anonIM+noSW
plain
noSW
anonIM
anonIM+noSW
Random baseline
1 shows a good splitting result, while processing
some other captions leads to less-natural sentence
splitting, and this might influence the global result.</p>
        </sec>
      </sec>
      <sec id="sec-7-2">
        <title>Model</title>
        <p>GLOVE
BERT
USE</p>
      </sec>
      <sec id="sec-7-3">
        <title>Pre-processing</title>
        <p>plain
plain
plain</p>
        <p>
          We run similar experiments on the publicly
available STS-benchmark dataset10
          <xref ref-type="bibr" rid="ref3">(Cer et al.,
2017)</xref>
          , in order to see if the models show similar
behaviour when benchmarked on a more
conventional scenario. The task is similar to the one
presented herein: it consists in the assessment of pairs
of sentences according to their degree of
semantic similarity. In this task, models are evaluated
by the Pearson correlation of machine scores with
human judgments. Table 3 shows the expected
results: Contextual models outperform GloVe based
model in a consisted way, and USE outperform
the rest by large margin (about 20-30 points better
overall). It confirms that model performances are
task-dependent, and that results obtained in
nonconventional scenarios can be counter-intuitive if
compared to results obtained in conventional ones.
10http://ixa2.si.ehu.es/stswiki/index.php/STSbenchmark
8
        </p>
      </sec>
    </sec>
    <sec id="sec-8">
      <title>Conclusions and Future Work</title>
      <p>In this paper we presented a comparison of four
popular representation models (one-hot encoding,
GloVe, BERT, USE) in the task of semantic
textual similarity on a non-conventional scenario:
action reference similarity between sentences from
different domains.</p>
      <p>
        In the future, we would like to extend our Gold
Standard dataset, not only in terms of dimension
(i.e. more LSMDC16 video-caption pairs
annotated with acs from IMAGACT), but also in
terms of annotators. It would be interesting to
observe to what extend the visual stimuli offered
by video prototypes can be interpreted clearly by
more than one annotator, and thus calculate the
inter-annotator agreement. Moreover, we plan to
extend the evaluation to other representation
models as well as state-of-the-art supervised models,
and see if their performances in canonical tests
are confirmed on our scenario. We would also try
to augment data used for this test, by exploiting
dense video captioning models, i.e. videoBERT
        <xref ref-type="bibr" rid="ref24">(Sun et al., 2019)</xref>
        .
      </p>
    </sec>
    <sec id="sec-9">
      <title>Acknowledgements</title>
      <p>This research was partially supported by the
Spanish MINECO (DeepReading RTI2018-096846-B-C21
(MCIU/AEI/FEDER, UE)), ERA-Net CHISTERA LIHLITH
Project funded by Agencia Esatatal de Investigacin (AEI,
Spain) projects PCIN-2017-118/AEI and
PCIN-2017085/AEI, the Basque Government (excellence research
group, IT1343-19), and the NVIDIA GPU grant program.</p>
      <p>Zhilin Yang, Zihang Dai, Yiming Yang, Jaime G. Carbonell,
Ruslan Salakhutdinov, and Quoc V. Le. 2019. Xlnet:
Generalized autoregressive pretraining for language
understanding. CoRR.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          <string-name>
            <surname>Eneko</surname>
            <given-names>Agirre</given-names>
          </string-name>
          , Daniel Cer, Mona Diab, and Aitor GonzalezAgirre.
          <year>2012</year>
          . SemEval
          <article-title>-2012 Task 6: A pilot on semantic textual similarity</article-title>
          .
          <source>In *SEM 2012 - 1st Joint Conference on Lexical and Computational Semantics</source>
          , pages
          <fpage>385</fpage>
          -
          <lpage>393</lpage>
          .
          <source>Universidad del Pais Vasco</source>
          , Leioa, Spain, January.
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          <string-name>
            <given-names>Piotr</given-names>
            <surname>Bojanowski</surname>
          </string-name>
          , Edouard Grave, Armand Joulin, and
          <string-name>
            <given-names>Tomas</given-names>
            <surname>Mikolov</surname>
          </string-name>
          .
          <year>2017</year>
          .
          <article-title>Enriching Word Vectors with Subword Information. Transactions of the Association for Computational Linguistics (TACL</article-title>
          ),
          <volume>5</volume>
          (
          <issue>1</issue>
          ):
          <fpage>135</fpage>
          -
          <lpage>146</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          <string-name>
            <given-names>Daniel</given-names>
            <surname>Cer</surname>
          </string-name>
          , Mona Diab, Eneko Agirre, In˜ igo Lopez-Gazpio, and
          <string-name>
            <given-names>Lucia</given-names>
            <surname>Specia</surname>
          </string-name>
          .
          <year>2017</year>
          . SemEval
          <article-title>-2017 task 1: Semantic textual similarity multilingual and crosslingual focused evaluation</article-title>
          .
          <source>In Proceedings of the 11th International Workshop on Semantic Evaluation (SemEval-2017)</source>
          , pages
          <fpage>1</fpage>
          -
          <lpage>14</lpage>
          , Vancouver, Canada, August. Association for Computational Linguistics.
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          <string-name>
            <given-names>Daniel</given-names>
            <surname>Cer</surname>
          </string-name>
          , Yinfei Yang,
          <string-name>
            <surname>Sheng-yi Kong</surname>
            , Nan Hua, Nicole Limtiaco,
            <given-names>Rhomni</given-names>
          </string-name>
          <string-name>
            <surname>St</surname>
          </string-name>
          . John, Noah Constant, Mario Guajardo-Cespedes, Steve Yuan, Chris Tar,
          <string-name>
            <surname>Yun-Hsuan</surname>
            <given-names>Sung</given-names>
          </string-name>
          , Brian Strope, and
          <string-name>
            <given-names>Ray</given-names>
            <surname>Kurzweil</surname>
          </string-name>
          .
          <year>2018</year>
          .
          <article-title>Universal sentence encoder</article-title>
          .
          <source>CoRR.</source>
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          <string-name>
            <given-names>Alexis</given-names>
            <surname>Conneau</surname>
          </string-name>
          and
          <string-name>
            <given-names>Douwe</given-names>
            <surname>Kiela</surname>
          </string-name>
          .
          <year>2018</year>
          .
          <article-title>Senteval: An evaluation toolkit for universal sentence representations</article-title>
          .
          <source>arXiv preprint arXiv:1803</source>
          .05449.
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          <string-name>
            <given-names>Jacob</given-names>
            <surname>Devlin</surname>
          </string-name>
          ,
          <string-name>
            <surname>Ming-Wei</surname>
            <given-names>Chang</given-names>
          </string-name>
          ,
          <string-name>
            <given-names>Kenton</given-names>
            <surname>Lee</surname>
          </string-name>
          ,
          <string-name>
            <given-names>and Kristina</given-names>
            <surname>Toutanova</surname>
          </string-name>
          .
          <year>2018</year>
          .
          <article-title>BERT - Pre-training of Deep Bidirectional Transformers for Language Understanding</article-title>
          .
          <source>CoRR</source>
          ,
          <year>1810</year>
          :arXiv:
          <year>1810</year>
          .04805.
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          <string-name>
            <given-names>Christiane</given-names>
            <surname>Fellbaum</surname>
          </string-name>
          .
          <year>1998</year>
          .
          <article-title>WordNet: an electronic lexical database</article-title>
          . MIT Press.
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          <string-name>
            <surname>Christopher D. Manning</surname>
          </string-name>
          , Prabhakar Raghavan, and Hinrich Schu¨ tze.
          <year>2008</year>
          . Introduction to Information Retrieval. Cambridge University Press, New York, NY, USA.
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          <string-name>
            <given-names>Tomas</given-names>
            <surname>Mikolov</surname>
          </string-name>
          , Ilya Sutskever, Kai Chen, Greg S Corrado, and
          <string-name>
            <given-names>Jeff</given-names>
            <surname>Dean</surname>
          </string-name>
          .
          <year>2013</year>
          .
          <article-title>Distributed Representations of Words and Phrases and their Compositionality</article-title>
          . In
          <string-name>
            <surname>C J C Burges</surname>
            ,
            <given-names>L</given-names>
          </string-name>
          <string-name>
            <surname>Bottou</surname>
            ,
            <given-names>M</given-names>
          </string-name>
          <string-name>
            <surname>Welling</surname>
            ,
            <given-names>Z</given-names>
          </string-name>
          <string-name>
            <surname>Ghahramani</surname>
          </string-name>
          ,
          <article-title>and</article-title>
          K Q Weinberger, editors,
          <source>Advances in Neural Information Processing Systems</source>
          <volume>26</volume>
          , pages
          <fpage>3111</fpage>
          -
          <lpage>3119</lpage>
          . Curran Associates, Inc.
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          <string-name>
            <given-names>Massimo</given-names>
            <surname>Moneglia</surname>
          </string-name>
          and
          <string-name>
            <given-names>Alessandro</given-names>
            <surname>Panunzi</surname>
          </string-name>
          .
          <year>2007</year>
          .
          <article-title>Action Predicates and the Ontology of Action across Spoken Language Corpora</article-title>
          . In M Alca´
          <article-title>ntara Pla´</article-title>
          and Th Declerk, editors,
          <source>Proceedings of the International Workshop on the Semantic Representation of Spoken Language (SRSL</source>
          <year>2007</year>
          ), pages
          <fpage>51</fpage>
          -
          <lpage>58</lpage>
          , Salamanca.
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          <string-name>
            <given-names>Massimo</given-names>
            <surname>Moneglia</surname>
          </string-name>
          , Susan Brown, Francesca Frontini, Gloria Gagliardi, Fahad Khan, Monica Monachini, and
          <string-name>
            <given-names>Alessandro</given-names>
            <surname>Panunzi</surname>
          </string-name>
          .
          <year>2014</year>
          .
          <article-title>The IMAGACT Visual Ontology. An Extendable Multilingual Infrastructure for the representation of lexical encoding of Action</article-title>
          . LREC, pages
          <fpage>3425</fpage>
          -
          <lpage>3432</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          <string-name>
            <given-names>Massimo</given-names>
            <surname>Moneglia</surname>
          </string-name>
          .
          <year>2014</year>
          .
          <article-title>The variation of Action verbs in multilingual spontaneous speech corpora</article-title>
          .
          <source>Spoken Corpora and Linguistic Studies</source>
          ,
          <volume>61</volume>
          :
          <fpage>152</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          <string-name>
            <given-names>Roberto</given-names>
            <surname>Navigli</surname>
          </string-name>
          and Simone Paolo Ponzetto.
          <year>2012</year>
          .
          <article-title>Babelnet: The automatic construction, evaluation and application of a wide-coverage multilingual semantic network</article-title>
          .
          <source>Artificial Intelligence</source>
          ,
          <volume>193</volume>
          (
          <issue>0</issue>
          ):
          <fpage>217</fpage>
          -
          <lpage>250</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          <string-name>
            <given-names>Matteo</given-names>
            <surname>Pagliardini</surname>
          </string-name>
          , Prakhar Gupta, and
          <string-name>
            <given-names>Martin</given-names>
            <surname>Jaggi</surname>
          </string-name>
          .
          <year>2017</year>
          .
          <article-title>Unsupervised learning of sentence embeddings using compositional n-gram features</article-title>
          .
          <source>CoRR, abs/1703</source>
          .02507.
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          <string-name>
            <given-names>Alessandro</given-names>
            <surname>Panunzi</surname>
          </string-name>
          , Lorenzo Gregori, and Andrea Amelio Ravelli. 2018a.
          <article-title>One event, many representations. mapping action concepts through visual features</article-title>
          .
          <source>In James Pustejovsky and Ielka van der Sluis</source>
          , editors,
          <source>Proceedings of the Eleventh International Conference on Language Resources and Evaluation (LREC</source>
          <year>2018</year>
          ), Miyazaki,
          <string-name>
            <given-names>Japan. European</given-names>
            <surname>Language Resources Association (ELRA).</surname>
          </string-name>
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          <string-name>
            <given-names>Alessandro</given-names>
            <surname>Panunzi</surname>
          </string-name>
          , Massimo Moneglia, and
          <string-name>
            <given-names>Lorenzo</given-names>
            <surname>Gregori</surname>
          </string-name>
          . 2018b.
          <article-title>Action identification and local equivalence of action verbs: the annotation framework of the imagact ontology</article-title>
          .
          <source>In James Pustejovsky and Ielka van der Sluis</source>
          , editors,
          <source>Proceedings of the Eleventh International Conference on Language Resources and Evaluation (LREC</source>
          <year>2018</year>
          ), Miyazaki,
          <string-name>
            <given-names>Japan. European</given-names>
            <surname>Language Resources Association (ELRA).</surname>
          </string-name>
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          <string-name>
            <surname>Jeffrey</surname>
            <given-names>Pennington</given-names>
          </string-name>
          , Richard Socher, and
          <string-name>
            <given-names>Christopher D</given-names>
            <surname>Manning</surname>
          </string-name>
          .
          <year>2014</year>
          .
          <article-title>GloVe: Global vectors for word representation</article-title>
          .
          <source>In Conference on Empirical Methods in Natural Language Processing</source>
          , pages
          <fpage>1532</fpage>
          -
          <lpage>1543</lpage>
          . Stanford University, Palo Alto, United States, January.
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          <string-name>
            <surname>Matthew E Peters</surname>
            , Mark Neumann, Mohit Iyyer, Matt Gardner, Christopher Clark,
            <given-names>Kenton</given-names>
          </string-name>
          <string-name>
            <surname>Lee</surname>
            ,
            <given-names>and Luke</given-names>
          </string-name>
          <string-name>
            <surname>Zettlemoyer</surname>
          </string-name>
          .
          <year>2018</year>
          .
          <article-title>Deep contextualized word representations</article-title>
          .
          <source>In Proceedings of NAACL-HLT</source>
          , pages
          <fpage>2227</fpage>
          -
          <lpage>2237</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          <string-name>
            <given-names>Peng</given-names>
            <surname>Qi</surname>
          </string-name>
          , Timothy Dozat, Yuhao Zhang, and
          <string-name>
            <given-names>Christopher D</given-names>
            <surname>Manning</surname>
          </string-name>
          .
          <year>2018</year>
          .
          <article-title>Universal Dependency Parsing from Scratch. CoNLL Shared Task</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          <string-name>
            <given-names>Anna</given-names>
            <surname>Rohrbach</surname>
          </string-name>
          , Marcus Rohrbach, Niket Tandon, and
          <string-name>
            <given-names>Bernt</given-names>
            <surname>Schiele</surname>
          </string-name>
          .
          <year>2015</year>
          .
          <article-title>A dataset for Movie Description</article-title>
          .
          <source>In 2015 IEEE Conference on Computer Vision and Pattern Recognition (CVPR).</source>
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          <string-name>
            <given-names>Anna</given-names>
            <surname>Rohrbach</surname>
          </string-name>
          , Atousa Torabi, Marcus Rohrbach, Niket Tandon, Christopher Pal, Hugo Larochelle, Aaron Courville, and
          <string-name>
            <given-names>Bernt</given-names>
            <surname>Schiele</surname>
          </string-name>
          .
          <year>2017</year>
          . Movie Description.
          <source>International Journal of Computer Vision</source>
          ,
          <volume>123</volume>
          (
          <issue>1</issue>
          ):
          <fpage>94</fpage>
          -
          <lpage>120</lpage>
          , January.
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          <string-name>
            <given-names>Andrew</given-names>
            <surname>Salway</surname>
          </string-name>
          .
          <year>2007</year>
          .
          <article-title>A corpus-based analysis of audio description</article-title>
          .
          <source>In Jorge D´ıaz Cintas</source>
          , Pilar Orero, and Aline Remael, editors,
          <source>Media for All</source>
          , pages
          <fpage>151</fpage>
          -
          <lpage>174</lpage>
          . Leiden.
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          <string-name>
            <given-names>Karin</given-names>
            <surname>Kipper Schuler</surname>
          </string-name>
          .
          <year>2006</year>
          .
          <article-title>VerbNet: A Broad-Coverage, Comprehensive Verb Lexicon</article-title>
          .
          <source>Ph.D. thesis</source>
          , University of Pennsylvania.
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          <string-name>
            <given-names>Chen</given-names>
            <surname>Sun</surname>
          </string-name>
          , Austin Myers, Carl Vondrick, Kevin Murphy, and
          <string-name>
            <given-names>Cordelia</given-names>
            <surname>Schmid</surname>
          </string-name>
          .
          <year>2019</year>
          .
          <article-title>Videobert: A joint model for video and language representation learning</article-title>
          .
          <source>CoRR</source>
          , abs/
          <year>1904</year>
          .01766.
        </mixed-citation>
      </ref>
      <ref id="ref25">
        <mixed-citation>
          <string-name>
            <given-names>Atousa</given-names>
            <surname>Torabi</surname>
          </string-name>
          , Christopher J Pal, Hugo Larochelle, and Aaron C Courville.
          <year>2015</year>
          .
          <article-title>Using Descriptive Video Services to Create a Large Data Source for Video Annotation Research</article-title>
          . cs.CV:arXiv:
          <fpage>1503</fpage>
          .
          <fpage>01070</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref26">
        <mixed-citation>
          <string-name>
            <given-names>Alex</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Amanpreet</given-names>
            <surname>Singh</surname>
          </string-name>
          ,
          <string-name>
            <surname>Julian Michael</surname>
          </string-name>
          , Felix Hill,
          <string-name>
            <given-names>Omer Levy</given-names>
            , and
            <surname>Samuel</surname>
          </string-name>
          <string-name>
            <given-names>R.</given-names>
            <surname>Bowman</surname>
          </string-name>
          .
          <year>2019</year>
          .
          <article-title>GLUE: A multi-task benchmark and analysis platform for natural language understanding</article-title>
          .
          <source>In International Conference on Learning Representations.</source>
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>