<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>In Search for Linear Relations in Sentence Embedding Spaces</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Petra Barancˇíková</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Ondrˇej Bojar</string-name>
          <email>bojar@ufal.mff.cuni.cz</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Charles University Faculty of Mathematics and Physics Institute of Formal and Applied Linguistics</institution>
        </aff>
      </contrib-group>
      <abstract>
        <p>We present an introductory investigation into continuous-space vector representations of sentences. We acquire pairs of very similar sentences differing only by a small alterations (such as change of a noun, adding an adjective, noun or punctuation) from datasets for natural language inference using a simple pattern method. We look into how such a small change within the sentence text affects its representation in the continuous space and how such alterations are reflected by some of the popular sentence embedding models. We found that vector differences of some embeddings actually reflect small changes within a sentence.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1 Introduction</title>
      <p>Continuous-space representations of sentences, so-called
sentence embeddings, are becoming an interesting object
of study, consider e.g. the BlackBox workshop.1
Representing sentences in a continuous space, i.e. commonly
with a long vector of real numbers, can be useful in
multiple ways, analogous to continuous word representations
(word embeddings). Word embeddings have provably
made downstream processing robust to unimportant input
variations or minor errors (sometimes incl. typos), they
have greatly boosted the performance of many tasks in
low data conditions and can form the basis of
empiricallydriven lexicographic explanations of word meanings.</p>
      <p>
        One notable observation was made in [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ], showing that
several interesting relations between words have their
immediate geometric counterpart in the continuous vector
space.
      </p>
      <p>
        Our aim is to examine existing continuous
representations of whole sentences, looking for an analogous
behaviour. The idea of what we are hoping for is illustrated
in Figure 1. As with words, we would like to learn if and to
what extent some simple geometric operations in the
continuous space correspond to simple semantic operations
on the sentence strings. Similarly to [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ], we are
deliberately not including this aspect in the training objective
of the sentence presentations but instead search for
properties that are learned in an unsupervised way, as a
sideeffect of the original training objective, data and setup.
      </p>
      <p>A sad boy is walking.</p>
      <p>Look how sad my cat is.</p>
      <p>A little boy is walking.</p>
      <p>Look at my little cat!
A little boy is running.</p>
      <p>A dog is walking past a field.
s
e
c
n
e
t
n
e
S
f
o
e
c
a
p
S</p>
      <p>A man is walking</p>
      <p>in the field.
itrsoan .o.f.aamdoagn..in.stead
e
p
O
f
o
e
c
a
p
S</p>
      <p>There is a dog running past the field.</p>
      <p>...being sad, not little...
...running instead of walking...</p>
      <p>...a grown-up instead of a child...</p>
      <p>This approach has the potential of explaining the good or
bad performance of the examined types of representations
in various tasks.</p>
      <p>The paper is structured as follows: Section 2 reviews
the closest related work. Sections 3 and 4, respectively,
describe the dataset of sentences and the sentence
embeddings methods we use. Section 5 presents the selection of
operations on the sentence vectors. Section 6 provides the
main experimental results of our work. We conclude in
Section 7.
2</p>
    </sec>
    <sec id="sec-2">
      <title>Related Work</title>
      <p>
        Series of tests to measure how well their word
embeddings capture semantic and syntactic information is
defined in [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ]. These tests include for example
declination of adjectives (“easy"!“easier"!“easiest"),
changing the tense of a verb (“walking"!“walk") or getting
the capital (“Athens"!“Greece") or currency of a state
(“Angola"!“kwanza"). References [2; 13] have further
refined the support of sub-word units, leading to
considerable improvements in representing morpho-syntactic
properties of words. Vylomova, Rimmel, Cohn and
Baldwin [
        <xref ref-type="bibr" rid="ref26">26</xref>
        ] largely extended the set of considered semantic
relations of words.
      </p>
      <p>
        Sentence embeddings are most commonly evaluated
extrinsically in so called ‘transfer tasks’, i.e. comparing
the evaluated representations based on their performance
in sentence sentiment analysis, question type prediction,
natural language inference and other assignments.
Reference [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ] introduce ‘probing tasks’ for intrinsic evaluation
of sentence embeddings. They measure to what extent
linguistic features like sentence length, word order, or the
depth of the syntactic tree are available in a sentence
embedding. This work was extended to SentEval [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ], a toolkit
for evaluating the quality of sentence embedding both
intrinsically and extrinsically. It contains 17 transfer tasks
and 10 probing tasks. SentEval is applied to many recent
sentence embedding techniques showing that no method
had a consistently good performance across all tasks [
        <xref ref-type="bibr" rid="ref18">18</xref>
        ].
      </p>
      <p>
        Voleti, Liss and Berisha [
        <xref ref-type="bibr" rid="ref25">25</xref>
        ] examine how errors (such
as incorrect word substitution caused by automatic speech
recognition) in a sentence affect its embedding. The
embeddings of corrupted sentences are then used in textual
similarity tasks and the performance is compared with
original embedding. The results suggest that pretrained
neural sentence encoders are much more robust to
introduced errors contrary to bag-of-words embeddings.
3
      </p>
    </sec>
    <sec id="sec-3">
      <title>Examined Sentences</title>
      <p>
        Because manual creation of sentence variations is costly,
we reuse existing data from SNLI [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ] and MultiNLI [
        <xref ref-type="bibr" rid="ref27">27</xref>
        ].
Both these collections consist of pairs of sentences—a
premise and a hypothesis—and their relationship
(entailment/contradiction/neutral). The two datasets together
contain 982k unique sentence pairs. All sentences were
lowercased and tokenized using NLTK [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ].
      </p>
      <p>From all the available sentence pairs, we select only a
subset where the difference between the sentences in the
pair can be described with a simple pattern. Our method
goes as follows: given two sentences, a premise p and the
corresponding hypothesis h, we find the longest common
substring consisting of whole words and replace it with
a variable. This is repeated once more, so our sentence
patterns can have up to two variables. In the last step, we
make sure the pattern is in a canonical form by switching
the variables to ensure they are alphabetically sorted in p.
The process is illustrated in Figure 2.</p>
      <p>Ten most common patterns for each NLI relation are
shown in Figure 3. Many of the obtained patterns clearly
match the sentence pair label. For instance the pattern no.
2 (“X man Y ! X person Y”) can be expected to lead to
a sentence pair illustrating entailment. If a man appears in
a story, we can infer that a person appeared in the story.
The contradictions illustrate typical oppositions like man–
woman, dog–cat. Neutrals are various refinements of the
content described by the sentences, probably in part due to
the original instruction in SNLI that hypothesis “might be
a true” given the premise in neutral relation.</p>
      <p>We kept only patterns appearing with at least 20
different sentence pairs in order to have large and variable sets
of sentence pairs in subsequent experiments. We also
ignored the overall most common pattern, namely the
identity, because it actually does not alter the sentence at all.
Strangely enough, identity was observed not just among
entailment pairs (693 cases), but also in neutral (41 cases)
and contradiction (22) pairs.</p>
      <p>Altogether, we collected 4,2k unique sentence pairs in
60 patterns. Only 10% of this data comes from MultiNLI,
the majority is from SNLI.
4</p>
    </sec>
    <sec id="sec-4">
      <title>Sentence Embeddings</title>
      <p>We experiment with several popular pretrained sentence
embeddings.</p>
      <p>
        InferSent2 [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ] is the first embedding model that used a
supervised learning to compute sentence representations.
It was trained to predict inference labels on the SNLI
dataset. The authors tested 7 different architectures and
BiLSTM encoder with max pooling achieved the best
results. InferSent comes in two versions: InferSent_1
is trained with Glove embeddings [
        <xref ref-type="bibr" rid="ref17">17</xref>
        ] and InferSent_2
with fastText [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]. InferSent representations are by far the
largest, with the dimensionality of 4096 in both versions.
      </p>
      <p>
        Similarly to InferSent, Universal Sentence Encoder [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ]
uses unsupervised learning augmented with training on
supervised data from SNLI. There are two models available.
USE_T3 is a transformer-network [
        <xref ref-type="bibr" rid="ref23">23</xref>
        ] designed for higher
accuracy at the cost of larger memory use and
computational time. USE_D4 is a deep averaging network [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ],
where words and bi-grams embeddings are averaged and
used as input to a deep neural network that computes the
final sentence embeddings. This second model is faster
and more efficient but its accuracy is lower. Both models
output representation with 512 dimensions.
      </p>
      <p>
        Unlike the previous models, BERT5 (Bidirectional
Encoder Representations from Transformers) [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ] is a deep
unsupervised language representation, pre-trained using
only unlabeled text. It has two self-supervised training
objectives - masked language modelling and next
sentence classification. It is considered bidirectional as the
Transformer encoder reads the entire sequence of words
at once. We use a pre-trained BERT-Large model with
2https://github:com/facebookresearch/InferSent
3https://tfhub:dev/google/universal-sentenceencoder-large/3
      </p>
      <p>4https://tfhub:dev/google/universal-sentenceencoder/2
5https://github:com/google-research/bert
Whole Word Masking. BERT gives embeddings for every
(sub)word unit, we take as a sentence embedding a [CLS]
token, which is inserted at the beginning of every sentence.
BERT embeddings have 1,024-dimensions.</p>
      <p>
        ELMo6 (Embedding from Language Models) [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ] uses
representations from a biLSTM that is trained with the
language model objective on a large text dataset. Its
embeddings are a function of the internal layers of the
bidirectional Language Model (biLM), which should
capture not only semantics and syntax, but also different
meanings a word can represent in different contexts
(polysemy). Similarly to BERT, each token representation of
ELMo is a function of the entire input sentence - one word
gets different embeddings in different contexts. ELMo
computes an embedding for every token and we compute
the final sentence embedding as the average over all
tokens. It has dimensionality 1024.
      </p>
      <p>
        LASER7 (Language-Agnostic SEntence
Representations) [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ] is a five-layer bi-directional LSTM (BiLSTM)
network. The 1,024-dimension vectors are obtained by
max-pooling over its last states. It was trained to
translate from more than 90 languages to English or Spanish at
the same time, the source language was selected randomly
in each batch.
5
      </p>
    </sec>
    <sec id="sec-5">
      <title>Choosing Vector Operations</title>
      <p>
        Mikolov, Chen, Corrado and Dean [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ] used a simple
vector difference as the operation that relates two word
embeddings. For sentences embeddings, we experiment a
little and consider four simple operations: addition,
subtraction, multiplication and division, all applied elementwise.
More operations could be also considered as long as they
are reversible, so that we can isolate the vector change for
a particular sentence alternation and apply it to the
embedding of any other sentence. Hopefully, we would then land
in the area where the correspondingly altered sentence is
embedded.
      </p>
      <p>The underlying idea of our analysis was already
sketched in Figure 1. From every sentence pair in our
dataset, we extract the pattern, i.e. the string edit of the
sentences. The arithmetic operation needed to move from
the embedding of the first sentence to the embedding of
the second sentence (in the continuous space of sentences)
can be represented as a point in what we call the space of
operations. Considering all sentence pairs that share the
same edit pattern, we obtain many points in the space of
operations. If the space of sentences reflects the
particular edit pattern in an accessible way, all the corresponding
points in the space of operations will be close together,
forming a cluster.</p>
      <p>
        To select which of the arithmetic operations best suits
the data, we test pattern clustering with three common
clustering performance evaluation methods:
6https://github:com/HIT-SCIR/ELMoForManyLangs
7https://github:com/facebookresearch/LASER
Adjusted Rand index [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ] is measure of the
similarity between two cluster assignments adjusted with
chance normalization. The score ranges from 1 to
+1 with 1 being the perfect match score and values
around 0 meaning random label assignment.
Negative numbers show worse agreement than what is
expected from a random result.
      </p>
      <p>
        V-measure [
        <xref ref-type="bibr" rid="ref19">19</xref>
        ] is harmonic mean of homogeneity
(each cluster should contain only members of one
class) and completeness (all members of one class
should be assigned to the same cluster). The score
ranges from 0 (the worst situation) to 1 (perfect
score).
      </p>
      <p>
        Adjusted Mutual Information [
        <xref ref-type="bibr" rid="ref21">21</xref>
        ] measures the
agreement of the two clusterings with the correction
of agreement by chance. The random label
assignment gets a score close to 0, while two identical
clusterings get the score of 1.
      </p>
      <p>
        As the detailed description of these measures is out of
scope of this article, we refer readers to related literature
(e.g. [
        <xref ref-type="bibr" rid="ref24">24</xref>
        ]). We use these scores to compare patterns with
labels predicted by k-Means (best result of 100 random
initialisations). The results are presented in Table 1. It is
apparent that the best distribution by far is achieved using
the most intuitive operation, vector subtraction.
      </p>
      <p>There seems to be a weak correlation between the size
of embeddings and the scores. The smallest embeddings
USE_D and USE_T are getting the worst scores, while the
largest embeddings InferSent_1 are the best scoring
embeddings. However, InferSent_2 with dimensionality 4096
is performing poorly. The fact that several of the
embeddings were trained on SNLI does not to seem benefit those
embeddings. Between the three top scored embeddings,
only InferSent_1 was trained on the data that we use for
evaluation of embeddings.
6</p>
    </sec>
    <sec id="sec-6">
      <title>Experiments</title>
      <p>For the following exploration of the continuous space of
operations, we focus only on the ELMo embeddings. They
scored second best in all scores but unlike the best scoring
Infersent_1, ELMo was not trained on SNLI, which is the
major source of our sentence pairs.</p>
      <p>
        The t-SNE [
        <xref ref-type="bibr" rid="ref22">22</xref>
        ] visualisation of subtractions of ELMo
vectors is presented in Figure 4. The visualisation is
constructed automatically and, of course, without the
knowledge of the pattern label. It shows that the patterns are
generally grouped together into compact clusters with the
exception of a ‘chaos cloud’ in the middle and several
outliers. Also there are several patterns that seem inseparable,
e.g. “two X ! X" and “three X ! X", or “X white Y -&gt;
X Y" and “X black Y -&gt; X Y".
      </p>
      <p>We identified the patterns responsible for the noisy
center and outliers by computing weighted inertia for each
pattern (the sum of squared distances of samples to their
cluster center divided by the size of sample). The
clusters with highest inertia consists of patterns representing a
change of word order and/or adding or removing
punctuation. These patterns are:</p>
      <p>X is Y . ! Y is X
X , Y . ! Y X .</p>
      <p>X Y . ! Y , X .</p>
      <p>X Y . ! Y X .</p>
      <p>X , Y . ! Y , X .</p>
      <p>X . ! X</p>
      <p>X ! X .</p>
      <p>
        To see if the space of operations can be interpreted also
automatically, i.e. if the sentence relations are
generalizable, we remove the noisy patterns as above and apply
fully unsupervised clustering: we do not even disclose the
expected number of patterns, i.e. clusters. We try two
metrics for finding the optimal number of clusters:
DaviesBouldin’s index [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ] and Silhouette Coefficient [
        <xref ref-type="bibr" rid="ref20">20</xref>
        ]. They
are both designed to measure compactness and separation
of the clusters, i.e. they award dense clusters that are far
from each other. Both Davies-Bouldin index and
Silhouette Coefficient agree that the best separation is achieved
3:[X Y -&gt; X sad Y (680/703),
      </p>
      <p>X young Y -&gt; X sad Y (68/68),
X -&gt; sad X (50/51),
X little Y -&gt; X sad Y (19/21),
X -&gt; there is X (9/25),
X Y -&gt; X big Y (1/122)]
4: [two X -&gt; X (52/56),
a group of X -&gt; X (36/38),
three X -&gt; X (24/24)]
5: [X man Y -&gt; X person Y (218/227),</p>
      <p>X Y -&gt; X not Y (1/56)]
6: [X man Y -&gt; X woman Y (414/414),</p>
      <p>X men Y -&gt; X women Y (109/111),
X boy Y -&gt; X girl Y (107/109),
man X -&gt; woman X (31/31),
X boys Y -&gt; X girls Y (21/27),
X boy Y -&gt; X person Y (1/65),
X man Y -&gt; X person Y (1/227)]
at 9 clusters. Running k-Means with 9 clusters, we get the
result as plotted in Figure 5.</p>
      <p>Manually inspecting the contents of the automatically
identified clusters, we see that many clusters are
meaningful in some way. For instance, Cluster 1 captures 90%
(altogether 264 out of 292) sentence pairs exerting the pattern
of generalizing women, boys or girls to people. The
counterpart for men belonging to people is spread into Cluster 5
(218 out of 227 pairs) for the singular case and not so clean
Cluster 7 containing 57/57 of the plural pairs “X men Y !
X people Y” together with various oppositions. Cluster 2
covers all sentence pairs where a person is replaced with a
dog. Cluster 3 is primarily connected with sentence pairs
introducing bad mood. Cluster 4 unites patterns that
represent omitting a numeral/group. Cluster 6 covers gender
oppositions in one direction and Cluster 9 adds the other
direction (with some noise for child/man and person/man
and similar), etc.
7</p>
    </sec>
    <sec id="sec-7">
      <title>Conclusion and Future Work</title>
      <p>We examined vector spaces of sentence representations
as inferred automatically by sentence embedding
methods such as InferSent or ELMo. Our goal was to find out
if some simple arithmetic operations in the vector space
correspond to meaningful edit operations on the sentence
strings.</p>
      <p>Our first explorations of 60 sentence edit patterns
document that this is indeed the case. Automatically identified
frequent patterns with 20 or more occurrences in the SNLI
and MultiNLI datasets correspond to simple vector
differences. The ELMo space (and others such as Infersent_1,
LASER and USE-T, which are omitted due to paper length
requirements) exerts this property very well.</p>
      <p>Unfortunately, choosing ELMo as example might not
have been the best option – we compute ELMo
embeddings by averaging contextualized word embeddings and
majority of the patterns are just
removing/adding/changing a single word. Difference between two such sentence
embeddings may be a simple difference between the
embeddings of the words substituted, depending on the effect
of the contextualization. Thus, the differences in vector
space would show rather the word embeddings than the
sentence embeddings.</p>
      <p>It should be noted that our search made use of only
about 0.5% of the sentence pairs available in SNLI and
MultiNLI. The remaining sentence pairs differ beyond
what was extractable automatically using our simple
pattern method. A different approach for a fine-grained
description of the semantic relation between two sentences
would have to be taken for a better exploitation of the
available data.</p>
      <p>Our plans for the long term are to further verify these
observations using a more diverse set of vector operations
and a larger set of sentence alternations, primarily by
extending the set of alternation types. We also plan to
examine the possibilities of generating sentence strings back
from the sentence embedding space. If successful, our
method could lead to controlled paraphrasing via the
continuous space: take an input sentence, embed it, modify
the embedding using a vector operation and generate the
target sentence in the standard textual from.</p>
    </sec>
    <sec id="sec-8">
      <title>Acknowledgment</title>
      <p>This work has been supported by the grant No.
1824210S of the Czech Science Foundation. It has been
using language resources and tools stored and distributed
by the LINDAT/CLARIN project of the Ministry of
Education, Youth and Sports of the Czech Republic (project
LM2015071).</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>M.</given-names>
            <surname>Artetxe</surname>
          </string-name>
          and
          <string-name>
            <given-names>H.</given-names>
            <surname>Schwenk</surname>
          </string-name>
          .
          <article-title>Massively multilingual sentence embeddings for zero-shot cross-lingual transfer and beyond</article-title>
          . CoRR, abs/
          <year>1812</year>
          .10464,
          <year>2018</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>P.</given-names>
            <surname>Bojanowski</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Grave</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Joulin</surname>
          </string-name>
          , and
          <string-name>
            <given-names>T.</given-names>
            <surname>Mikolov</surname>
          </string-name>
          .
          <article-title>Enriching word vectors with subword information</article-title>
          .
          <source>CoRR</source>
          , abs/1607.04606,
          <year>2016</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>S. R.</given-names>
            <surname>Bowman</surname>
          </string-name>
          , G. Angeli,
          <string-name>
            <given-names>C.</given-names>
            <surname>Potts</surname>
          </string-name>
          , and
          <string-name>
            <given-names>C. D.</given-names>
            <surname>Manning</surname>
          </string-name>
          .
          <article-title>A large annotated corpus for learning natural language inference</article-title>
          .
          <source>In Proceedings of the 2015 Conference on Empirical Methods in Natural Language Processing (EMNLP)</source>
          .
          <source>Association for Computational Linguistics</source>
          ,
          <year>2015</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>D.</given-names>
            <surname>Cer</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Yang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Kong</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Hua</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Limtiaco</surname>
          </string-name>
          ,
          <string-name>
            R. S. John,
            <given-names>N.</given-names>
            <surname>Constant</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Guajardo-Cespedes</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Yuan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Tar</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Sung</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Strope</surname>
          </string-name>
          , and
          <string-name>
            <given-names>R.</given-names>
            <surname>Kurzweil</surname>
          </string-name>
          .
          <article-title>Universal sentence encoder</article-title>
          .
          <source>CoRR</source>
          , abs/
          <year>1803</year>
          .11175,
          <year>2018</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>W.</given-names>
            <surname>Che</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Liu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Zheng</surname>
          </string-name>
          , and
          <string-name>
            <given-names>T.</given-names>
            <surname>Liu</surname>
          </string-name>
          .
          <article-title>Towards better UD parsing: Deep contextualized word embeddings, ensemble, and treebank concatenation</article-title>
          .
          <source>CoRR</source>
          , abs/
          <year>1807</year>
          .03121,
          <year>2018</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>A.</given-names>
            <surname>Conneau</surname>
          </string-name>
          and
          <string-name>
            <given-names>D.</given-names>
            <surname>Kiela</surname>
          </string-name>
          . Senteval:
          <article-title>An evaluation toolkit for universal sentence representations</article-title>
          .
          <source>arXiv preprint arXiv:1803.05449</source>
          ,
          <year>2018</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>A.</given-names>
            <surname>Conneau</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Kiela</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Schwenk</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Barrault</surname>
          </string-name>
          ,
          <article-title>and</article-title>
          <string-name>
            <given-names>A.</given-names>
            <surname>Bordes</surname>
          </string-name>
          .
          <article-title>Supervised learning of universal sentence representations from natural language inference data</article-title>
          .
          <source>CoRR, abs/1705.02364</source>
          ,
          <year>2017</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>A.</given-names>
            <surname>Conneau</surname>
          </string-name>
          , G. Kruszewski,
          <string-name>
            <given-names>G.</given-names>
            <surname>Lample</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Barrault</surname>
          </string-name>
          , and
          <string-name>
            <given-names>M.</given-names>
            <surname>Baroni</surname>
          </string-name>
          .
          <article-title>What you can cram into a single vector: Probing sentence embeddings for linguistic properties</article-title>
          . CoRR, abs/
          <year>1805</year>
          .01070,
          <year>2018</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>D. L.</given-names>
            <surname>Davies</surname>
          </string-name>
          and
          <string-name>
            <given-names>D. W.</given-names>
            <surname>Bouldin</surname>
          </string-name>
          .
          <article-title>A cluster separation measure</article-title>
          .
          <source>IEEE Trans. Pattern Anal. Mach</source>
          . Intell.,
          <volume>1</volume>
          (
          <issue>2</issue>
          ):
          <fpage>224</fpage>
          -
          <lpage>227</lpage>
          , Feb.
          <year>1979</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>J.</given-names>
            <surname>Devlin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Chang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Lee</surname>
          </string-name>
          , and
          <string-name>
            <given-names>K.</given-names>
            <surname>Toutanova</surname>
          </string-name>
          . BERT:
          <article-title>pre-training of deep bidirectional transformers for language understanding</article-title>
          .
          <source>CoRR</source>
          , abs/
          <year>1810</year>
          .04805,
          <year>2018</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>L.</given-names>
            <surname>Hubert</surname>
          </string-name>
          and
          <string-name>
            <given-names>P.</given-names>
            <surname>Arabie</surname>
          </string-name>
          .
          <article-title>Comparing partitions</article-title>
          .
          <source>Journal of Classification</source>
          ,
          <volume>2</volume>
          (
          <issue>1</issue>
          ):
          <fpage>193</fpage>
          -
          <lpage>218</lpage>
          ,
          <year>Dec 1985</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>M.</given-names>
            <surname>Iyyer</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Manjunatha</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Boyd-Graber</surname>
          </string-name>
          , and H. Daumé III.
          <article-title>Deep unordered composition rivals syntactic methods for text classification</article-title>
          .
          <source>In Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing (Volume 1: Long Papers)</source>
          , pages
          <fpage>1681</fpage>
          -
          <lpage>1691</lpage>
          , Beijing, China,
          <year>July 2015</year>
          .
          <article-title>Association for Computational Linguistics</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <given-names>T.</given-names>
            <surname>Kocmi</surname>
          </string-name>
          and
          <string-name>
            <given-names>O.</given-names>
            <surname>Bojar</surname>
          </string-name>
          . Subgram:
          <article-title>Extending skipgram word representation with substrings</article-title>
          .
          <source>CoRR</source>
          , abs/
          <year>1806</year>
          .06571,
          <year>2018</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <given-names>E.</given-names>
            <surname>Loper</surname>
          </string-name>
          and
          <string-name>
            <given-names>S.</given-names>
            <surname>Bird</surname>
          </string-name>
          .
          <article-title>Nltk: The natural language toolkit</article-title>
          .
          <source>In In Proceedings of the ACL Workshop on Effective Tools and Methodologies for Teaching Natural Language Processing and Computational Linguistics</source>
          . Philadelphia: Association for Computational Linguistics,
          <year>2002</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <given-names>T.</given-names>
            <surname>Mikolov</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Chen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G. S.</given-names>
            <surname>Corrado</surname>
          </string-name>
          , and
          <string-name>
            <given-names>J.</given-names>
            <surname>Dean</surname>
          </string-name>
          .
          <article-title>Efficient estimation of word representations in vector space</article-title>
          ,
          <year>2013</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [16]
          <string-name>
            <given-names>F.</given-names>
            <surname>Pedregosa</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Varoquaux</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Gramfort</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Michel</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Thirion</surname>
          </string-name>
          ,
          <string-name>
            <given-names>O.</given-names>
            <surname>Grisel</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Blondel</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Prettenhofer</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Weiss</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Dubourg</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Vanderplas</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Passos</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Cournapeau</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Brucher</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Perrot</surname>
          </string-name>
          , and
          <string-name>
            <given-names>E.</given-names>
            <surname>Duchesnay</surname>
          </string-name>
          .
          <article-title>Scikit-learn: Machine learning in Python</article-title>
          .
          <source>Journal of Machine Learning Research</source>
          ,
          <volume>12</volume>
          :
          <fpage>2825</fpage>
          -
          <lpage>2830</lpage>
          ,
          <year>2011</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [17]
          <string-name>
            <given-names>J.</given-names>
            <surname>Pennington</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Socher</surname>
          </string-name>
          , and
          <string-name>
            <given-names>C.</given-names>
            <surname>Manning</surname>
          </string-name>
          . Glove:
          <article-title>Global vectors for word representation</article-title>
          .
          <source>In Proceedings of the 2014 Conference on Empirical Methods in Natural Language Processing (EMNLP)</source>
          , pages
          <fpage>1532</fpage>
          -
          <lpage>1543</lpage>
          , Doha, Qatar, Oct.
          <year>2014</year>
          .
          <article-title>Association for Computational Linguistics</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          [18]
          <string-name>
            <given-names>C. S.</given-names>
            <surname>Perone</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Silveira</surname>
          </string-name>
          , and
          <string-name>
            <given-names>T. S.</given-names>
            <surname>Paula</surname>
          </string-name>
          .
          <article-title>Evaluation of sentence embeddings in downstream and linguistic probing tasks</article-title>
          .
          <source>CoRR</source>
          , abs/
          <year>1806</year>
          .06259,
          <year>2018</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          [19]
          <string-name>
            <given-names>A.</given-names>
            <surname>Rosenberg</surname>
          </string-name>
          and
          <string-name>
            <given-names>J.</given-names>
            <surname>Hirschberg</surname>
          </string-name>
          .
          <article-title>V-measure: A conditional entropy-based external cluster evaluation measure</article-title>
          .
          <source>In Proceedings of the 2007 Joint Conference on Empirical Methods in Natural Language Processing and Computational Natural Language Learning (EMNLP-CoNLL)</source>
          , pages
          <fpage>410</fpage>
          -
          <lpage>420</lpage>
          , Prague, Czech Republic,
          <year>June 2007</year>
          .
          <article-title>Association for Computational Linguistics</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          [20]
          <string-name>
            <given-names>P.</given-names>
            <surname>Rousseeuw. Silhouettes</surname>
          </string-name>
          :
          <article-title>A graphical aid to the interpretation and validation of cluster analysis</article-title>
          .
          <source>J. Comput. Appl</source>
          . Math.,
          <volume>20</volume>
          (
          <issue>1</issue>
          ):
          <fpage>53</fpage>
          -
          <lpage>65</lpage>
          , Nov.
          <year>1987</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          [21]
          <string-name>
            <given-names>A.</given-names>
            <surname>Strehl</surname>
          </string-name>
          and
          <string-name>
            <given-names>J.</given-names>
            <surname>Ghosh</surname>
          </string-name>
          .
          <article-title>Cluster ensembles: A knowledge reuse framework for combining partitionings</article-title>
          .
          <source>In Eighteenth National Conference on Artificial Intelligence</source>
          , pages
          <fpage>93</fpage>
          -
          <lpage>98</lpage>
          , Menlo Park, CA, USA,
          <year>2002</year>
          . American Association for Artificial Intelligence.
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          [22]
          <string-name>
            <given-names>L. van der Maaten and G.</given-names>
            <surname>Hinton</surname>
          </string-name>
          .
          <article-title>Visualizing data using t-SNE</article-title>
          .
          <source>Journal of Machine Learning Research</source>
          ,
          <volume>9</volume>
          :
          <fpage>2579</fpage>
          -
          <lpage>2605</lpage>
          ,
          <year>2008</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          [23]
          <string-name>
            <given-names>A.</given-names>
            <surname>Vaswani</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Shazeer</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Parmar</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Uszkoreit</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Jones</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A. N.</given-names>
            <surname>Gomez</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Kaiser</surname>
          </string-name>
          ,
          <string-name>
            <surname>and I. Polosukhin.</surname>
          </string-name>
          <article-title>Attention is all you need</article-title>
          .
          <source>In NIPS</source>
          ,
          <year>2017</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          [24]
          <string-name>
            <given-names>N. X.</given-names>
            <surname>Vinh</surname>
          </string-name>
          and
          <string-name>
            <given-names>J.</given-names>
            <surname>Epps</surname>
          </string-name>
          .
          <article-title>A novel approach for automatic number of clusters detection in microarray data based on consensus clustering</article-title>
          .
          <source>In 2009 Ninth IEEE International Conference on Bioinformatics and BioEngineering</source>
          , pages
          <fpage>84</fpage>
          -
          <lpage>91</lpage>
          ,
          <year>June 2009</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref25">
        <mixed-citation>
          [25]
          <string-name>
            <given-names>R.</given-names>
            <surname>Voleti</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. M.</given-names>
            <surname>Liss</surname>
          </string-name>
          , and
          <string-name>
            <given-names>V.</given-names>
            <surname>Berisha</surname>
          </string-name>
          .
          <article-title>Investigating the effects of word substitution errors on sentence embeddings</article-title>
          .
          <source>CoRR</source>
          , abs/
          <year>1811</year>
          .07021,
          <year>2018</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref26">
        <mixed-citation>
          [26]
          <string-name>
            <given-names>E.</given-names>
            <surname>Vylomova</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Rimell</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Cohn</surname>
          </string-name>
          , and
          <string-name>
            <given-names>T.</given-names>
            <surname>Baldwin</surname>
          </string-name>
          .
          <article-title>Take and took, gaggle and goose, book and read: Evaluating the utility of vector differences for lexical relation learning</article-title>
          .
          <source>CoRR, abs/1509.01692</source>
          ,
          <year>2015</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref27">
        <mixed-citation>
          [27]
          <string-name>
            <given-names>A.</given-names>
            <surname>Williams</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Nangia</surname>
          </string-name>
          , and
          <string-name>
            <given-names>S. R.</given-names>
            <surname>Bowman</surname>
          </string-name>
          .
          <article-title>A broad-coverage challenge corpus for sentence understanding through inference</article-title>
          .
          <source>In NAACL-HLT</source>
          ,
          <year>2018</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>