<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>NILC at ASSIN 2: Exploring Multilingual Approaches</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Marco A. Sobrevilla Cabezudo</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Marcio In´acio</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Ana Carolina Rodrigues</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Edresson Casanova</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Rog´erio Figueredo de Sousa</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>NILC - Interinstitutional Center for Computational Linguistics Instituto de Ciˆencias Matem ́aticas e de Computac ̧a ̃o, Universidade de Sa ̃o Paulo</institution>
          ,
          <addr-line>Sa ̃o Carlos SP 13566-590</addr-line>
          ,
          <country country="BR">Brazil</country>
        </aff>
      </contrib-group>
      <fpage>48</fpage>
      <lpage>57</lpage>
      <abstract>
        <p>Recognizing Textual Entailment, also known as Natural Language Inference recognition, aims to identify if it is possible to infer the meaning of a text from another fragment of text. In this work, we investigate the use of multilingual models, through BERT, for recognizing inference and similarity in the ASSIN 2 dataset, an entailment recognition and sentence similarity corpus for Portuguese. We also investigate possible features that could enhance the results, such as similarity scores or WordNet relations. Our results show that a multilingual pre-trained BERT model may be su cient to outperform the current state-of-theart in this task for the Portuguese Language. We also show that using other features did not necessarily improve the performance of the model, however deeper studies are needed to investigate the causes for this.</p>
      </abstract>
      <kwd-group>
        <kwd>Natural Language Inference</kwd>
        <kwd>BERT</kwd>
        <kwd>Multilingual Training</kwd>
        <kwd>Cross-lingual Training</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>Recognizing meaning connections such as entailment relations and content
similarity among di↵ erent statements is part of daily communication and usually
done e↵ ortless by humans. However, automatizing such communication
component has been a challenge. Overcome it can help many Natural Language
Processing (NLP) applications such as Machine Translation, Question Answering,
Semantic Search and Information Extraction.</p>
      <p>
        Particularly, the task of recognizing textual entailment (RTE) has been widely
explored in natural language processing field. Inference recognition in NLP, also
called text entailment recognition, consist recognizing a directional relationship
between pairs of text expressions, in which a human reading the first text would
infer the second one is likely true.[
        <xref ref-type="bibr" rid="ref7">7</xref>
        ].
      </p>
      <p>
        Initially spread by the Pascal Challenge [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ], several inference-annotated
corpus for English have been released in the last decade, such as MultiNLI [
        <xref ref-type="bibr" rid="ref19">19</xref>
        ],
XNLI [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ], and SICK [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ]. Specifically for Portuguese, multiple e↵ orts have been
made to develop an inference-annotated corpus[
        <xref ref-type="bibr" rid="ref11">11</xref>
        ][
        <xref ref-type="bibr" rid="ref16">16</xref>
        ]. In 2016 the first shared
task for inference recognition for Portuguese, ASSIN [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ], took place, followed
by the second edition in 2019 (ASSIN 2) [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ].
      </p>
      <p>
        In addition to the growth of available corpora, several techniques have been
tested to improve inference recognition in NLP, including probabilistic models
and rule-based approaches [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ]. Recently, with the expansion of machine learning
applications (and neural networks in particular), it has been tested on the lights
of pre-trained language representations [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ][
        <xref ref-type="bibr" rid="ref18">18</xref>
        ].
      </p>
      <p>
        A broadly known one is the Bidirectional Encoder Representations from
Transformers (BERT). BERT has been used e↵ ectively in multiple tasks like
Semantic Text Similarity, Paraphrase detection, among others[
        <xref ref-type="bibr" rid="ref8">8</xref>
        ]1. BERT has
improved fine-tuning approaches using a masked language model (MLM), in
which parts of the input are randomly hidden to be predicted based only on
their context.
      </p>
      <p>In addition to MLM, the authors also use next sentence prediction task that
jointly pre-trains text-pair representations. BERT was the first fine-tuned
representation model to achieve state-of-the-art performance for a large number of
token and phrase-level tasks, outperforming models developed specifically for
these tasks. As far as we know, there are two monolingual pre-trained models
publicly available: one trained for English and one for Chinese, along with one
multilingual model. The multilingual model was trained for the 100 languages
with most articles on Wikipedia2.</p>
      <p>This work presents the results achieved by the NILC group for the ASSIN
2 shared task. We firstly analyse the corpus in order to find correlation
between features and classification labels. Then we fine-tune multilingual BERT
on sentence-pairs from ASSIN 2 corpus for the RTE task and finally, use the
generated embeddings and incorporate some linguistic features for the Semantic
Textual similarity (STS) evaluation. In general, we rank 3rd place for the RTE
task and 5th place for STS task.</p>
      <p>This paper is organized as follows. Firstly, we discuss previous related work
in section 2. Afterwards, we describe the ASSIN dataset in section 3, followed by
our experiments and results in section 5. Finally, some conclusions are presented
in section 6.
2</p>
    </sec>
    <sec id="sec-2">
      <title>Related Work</title>
      <p>There are several works on textual inference for multiple languages. However,
due to di↵ erences in corpora for other languages, we report mainly works done in
Brazilian Portuguese. This way, the works reported here use the corpus ASSIN,
thus making a closer comparison to our work.</p>
      <p>
        Rocha and Lopes Cardoso [
        <xref ref-type="bibr" rid="ref17">17</xref>
        ] reported only the result for European
Portuguese (PT-PT) and obtained F1 of 0.73. they explored the use of named
en1 Available at https://github.com/google-research/bert.
2 Available at https://github.com/google-research/bert/blob/master/multilingual.
md.
tities as a feature along with word similarity, number of semantically related
tokens, and whether both sentences have the same verb tense and voice.
      </p>
      <p>
        Fialho et al. [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ] (from the INESC-ID group) obtained an F1 score of 0.71 in
Brazilian Portuguese (PT-BR) and 0.66 in PT-PT for textual inference in their
best experiment. The authors trained a Support Vector Machine (SVM) model
using 96 lexical features, including editing distance, BLEU score, word overlap,
ROUGE score, among others.
      </p>
      <p>
        Barbosa et al. [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ] (from the Blue Man Group) obtained, in their best
experiment, an F1 score of 0.52 for PT-PB and 0.61 for PT-PT exploring the use
of word embeddings similarity. For classification, the authors used SVM and
Siamese Networks [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ].
      </p>
      <p>
        Reciclagem and ASAPP were proposed by Alves et al. [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ] (from the ASAPP
group). Reciclagem is based only on heuristics in semantic networks. While
ASAPP explores the use of lexical, syntactic, and semantic features extracted from
texts. Their best results were an F1 score of approximately 0.5 for PT-PB and
0.59 for PT-PT.
      </p>
      <p>
        Finally, Fonseca and Alu´ısio proposed the Infernal system [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ]. The authors
explored some features such as syntactic knowledge, embedding-based
similarity, and well-established features that deal with word alignments, totalizing 28
features. Their best experiment for PT-BR achieved F1 score of 0.71, similarly
to the previously reported INESC-ID system. On the other hand, for PT-PT,
F1 score of 0.72 has been reported, lower than the one obtained by Rocha and
Lopes Cardoso [
        <xref ref-type="bibr" rid="ref17">17</xref>
        ]. When considering the entire dataset, i.e. both PT-BR and
PT-PT, the Infernal system reached F1 score of 0.72, currently the best result
reported for this dataset.
3
      </p>
    </sec>
    <sec id="sec-3">
      <title>ASSIN 2 Dataset</title>
      <p>The ASSIN 2 corpus consists of 10,000 pairs of sentences tagged with similarity
grading and entailment classification. Tags for the Recognizing Textual
Entailment (RTE) task are None, when both sentences are not related in any way and
Entailment when the second sentence is a direct inference of the first one.</p>
      <p>ASSIN 2 was also manually annotated for semantic textual similarity in a
range from 0 to 5, for which the pair was considered more similar by the
annotators as higher is the number.</p>
      <p>The corpus was split in two parts for the shared task, 7000 pairs were
provided in advance as a training set and the other 3000 pairs later. The dataset
provided for training was balanced, with 3500 pairs labeled as None and 3500
as Entailment.
4</p>
    </sec>
    <sec id="sec-4">
      <title>Corpus Analysis</title>
      <p>As the dataset provided for training was balanced for the two entailment classes,
we verified if it was equally balanced for other features. We use
OpenWordnetPT to extract the number of synonyms and hypernyms in each pair per class.
Hypernyms were counted when the second phrase contained any hypernym of
any word in the first one. Synonym values were calculated in a similar way.
Results are shown in figures 1 and 2.</p>
      <p>As can be seen, sentence pairs with entailment tend to have more hypernyms
(as can be seen by the bars representing counts of 0 and 2). The same can be
observed for synonyms: there are more sentence pairs without entailment with
no synonyms than those with an entailment relation.</p>
      <p>Since the corpus was also annotated for semantic similarity in a continuous
range from 0 to 5, we investigate the relation between similarity index and
entailment classes. As a result, we find that entailment pairs have lower dispersion
than none ones, and, di↵ erently from none, its range is mostly concentrated in
higher similarity values, as can be seen in Figure 3.</p>
      <p>
        Although similarity values shows notable correlation to entailment classes,
it was not possible to considered them for the entailment recognition task, once
it was part of the expected predicted result. In other words, textual similarity
annotation were hidden in the test set. Therefore, in order to incorporate this
kind of knowledge into the model, some metrics have been explored, namely
BLEU [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ] and Levenshtein’s Edit Distance [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ].
      </p>
      <p>Both metrics were calculated for each pair of sentences resulting in the
distributions presented in figures 4 and 5. The figures show the distributions of the
metrics according to the classes in the corpus. The results for the BLEU metric,
as it is based on the calculation of string overlaps between texts, show greater
values for entailment (higher similarity). Accordingly, Edit Distance values are
lower for this class, as it computes how di↵ erent a pair of sentences is, i.e. their
dissimilarity.</p>
      <p>From these analyses, these four features (synonyms and hypernyms counts,
BLEU and Edit Distance scores) seem applicable to text inference prediction.
Thus, we try to combine these features with the model selected for this task: the
BERT language model, as will be discussed later.
5</p>
    </sec>
    <sec id="sec-5">
      <title>Experiments and Results</title>
      <p>We analyze how BERT performs on Recognizing Textual Entailment (RTE) and
Semantic Textual Similarity (STS) for Brazilian Portuguese in ASSIN 2 corpus.</p>
      <p>It is worth noting that our methods were tested on three variations of the
original dataset for RTE, which are the runs submitted to the shared task. On
the other hand, our submission to the Semantic Textual Similarity task only was
tested on the original dataset.
5.1</p>
      <sec id="sec-5-1">
        <title>Recognizing Textual Entailment (RTE)</title>
        <p>We use BERT to train the RTE classifier. Specifically, we use a pre-trained BERT
model, add an untrained layer of neurons at the end, and train the new model
for the RTE task. For this purpose, we use the pre-trained BERT multilingual
model that includes Brazilian Portuguese3 along with other 103 languages. The
model was trained for 7 epochs with a learning rate of 0.00002, a batch size of
22 and a maximum sequence length of 128 tokens4.</p>
        <p>As mentioned previously, linguistic features like BLEU, Edit-distance,
number of hypernymns and synonymns between the sentence in a sentence-pair were
also used in the model, however, the introduction of these features did not
contribute positively to the final result.
5.2</p>
      </sec>
      <sec id="sec-5-2">
        <title>Semantic Textual Similarity</title>
        <p>Due to the high correlation between the Entailment class and the similarity
values, we used the previous trained model for RTE to obtain embeddings for each
sentence-pair. Initially, we experiment using separate embeddings (an embedding
for each sentence, obtained through BERT), but the results were poor. Thus, we
use joint embeddings (size of 768) as input, obtained by providing both sentences
in the pair as an input to the BERT model, which creates a single representation
of the whole pair. Additionally, we incorporate the features BLEU, number of
synonyms and hypernyms in common for each sentence-pair.</p>
        <p>For experiments, we use a multilayer perceptron with a hidden layer (64
neurons), the logistic function as the activation function, the adam optimizer, a
learning rate of 0.001 and a maximum number of iterations of 1,000.
5.3</p>
      </sec>
      <sec id="sec-5-3">
        <title>Results</title>
        <p>As mentioned, our methods were trained and fine-tuned on three variations of
the original dataset. The first variation (”own” in Table 1), consists in splitting
the original training dataset (7,000 sentence-pairs) into 6,300 for training and
700 for development. This split has been done by a stratified sampling according
to both entailment and similarity values, to guarantee that their distribution is
also represented in the development set. The second one was provided by the
organization and it contains 6,500 sentence-pairs for training and 500 for
development (”assin-2” in Table 1). The last variation comprised all 7000 sentences
(”all” in Table 1).</p>
        <p>Table 1 shows the results of the best three teams (excluding our team), our
results and the baseline results for RTE. In general, our proposal obtained the
third place in the RTE task (being only surpassed by the Deep Learning Brasil
and IPR teams) and the di↵ erence between our proposal and their proposals is
small.</p>
        <p>Concerning our proposal, it is important to highlight three regards. Firstly,
the split performed by us shows the best results even containing fewer instances
3 The model is available at https://storage.googleapis.com/bert models/2018 11 23/
multi cased L-12 H-768 A-12.zip
4 It is worth noting that other hyperparameters were tested. However, the results were
not shown improvements and they are no reported in this paper.
Team</p>
        <p>Run
Deep Learning Brasil
IPR
Stilingue
NILC
NILC
NILC
Baseline
Baseline
Baseline</p>
        <p>Ensemble
1
2
own
assin-2
all
BoW sentence 2
Word Overlap</p>
        <p>Infernal
in the training set than the original split, which could note the relevance of the
splitting strategy. However, we cannot a rm this due to the small improvements.</p>
        <p>Secondly, the introduction of linguistic features (BLEU, edit-distance, among
others) did not contribute positively to the final result. Thus, we only fine-tuned
the multilingual BERT on our RTE task. In principle, it could be thought that
multilingual BERT learned all these features, this way by including them, the
results did not improve. Another explanation could be that it is necessary to
explore other ways to integrate this kind of information. However, a deeper
study must be performed.</p>
        <p>Finally, it is worth noting the potential of multilingual BERT. This model
made our proposal easier in comparison to other approaches as it only needs
the pre-trained model and adding an untrained layer to perform the fine-tuning
process on the RTE task.</p>
        <p>Concerning the Semantic Textual Similarity (STS) task, our proposal
obtained smaller results than the other proposals. However, our results
outperformed all baselines. It is interesting to note that fine-tuning multilingual BERT
on the RTE task contributes positively to the STS task. However, some
linguistic features had to be incorporated to obtain better results, showing that BERT
could not learn this kind of information. In that sense, an interesting direction
could be experimenting fine-tuning on STS task and then using this
information to apply to RTE task or trying to learn both tasks in a multi-task learning
approach.</p>
        <p>Results</p>
        <p>Acc.
88.32%
87.58%
86.64%
87.17%
86.85%
86.56%
56.74%
66.71%
74.18%
This work presents the results obtained by the NILC group for the shared task
ASSIN 2 on entailment and textual similarity recognition. We analyzed
characteristics of the ASSIN 2 dataset according to its entailment classes and tested
two approaches of classification using BERT.</p>
        <p>Fine-tuning BERT, on the ASSIN 2 corpus without any extra feature
presented the best results, largely outperforming the baselines. Therefore, we show
that using a simple BERT model can provide satisfactory results in these tasks.</p>
        <p>As shown in the corpus analyses, similarity, BLEU and Edit Distance metrics
seem to be suitable for discriminating the entailment classes of the ASSIN 2
corpus. In particular, entailment pairs have higher similarity values and higher
BLEU scores, as well as lower Edit Distance values, than none class.</p>
        <p>The number of hypernyms and synonyms calculated consulting
OpenWordNet-PT for each pair also indicates some level of distinction between the two
classes, there were considerable more entailment pairs containing these
relations. However, incorporating these features as values concatenated to BERT
embedding vectors achieved poorer results.</p>
        <p>We considered some possibilities for these negative results, such as the way
features were incorporated, as a concatenation to BERT embeddings. Another
reason may be that BERT embeddings already incorporate such knowledge
(similarity, synonym and hypernym relations) within their representation. A future
deeper analysis about the incorporation of these features may lead to further
conclusions about these hypothesis.</p>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Alves</surname>
            ,
            <given-names>A.O.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rodrigues</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Oliveira</surname>
            ,
            <given-names>H.G.</given-names>
          </string-name>
          :
          <article-title>ASAPP: alinhamento semaˆntico automa´tico de palavras aplicado ao portuguˆes</article-title>
          .
          <source>Linguama´tica 8(2)</source>
          ,
          <fpage>43</fpage>
          -
          <lpage>58</lpage>
          (
          <year>2016</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Barbosa</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Cavalin</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Guimaraes</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kormaksson</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          : Blue Man Group no ASSIN:
          <article-title>Usando representac¸o˜es distribu´ıdas para similaridade semaˆntica e inferˆencia textual</article-title>
          .
          <source>Linguama´tica 8(2)</source>
          ,
          <fpage>15</fpage>
          -
          <lpage>22</lpage>
          (
          <year>2016</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Chopra</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hadsell</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          , LeCun, Y., et al.:
          <article-title>Learning a similarity metric discriminatively, with application to face verification</article-title>
          .
          <source>In: CVPR (1)</source>
          . pp.
          <fpage>539</fpage>
          -
          <lpage>546</lpage>
          (
          <year>2005</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Conneau</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lample</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rinott</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Williams</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bowman</surname>
            ,
            <given-names>S.R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Schwenk</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Stoyanov</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          :
          <article-title>Xnli: Evaluating cross-lingual sentence representations</article-title>
          .
          <source>arXiv preprint arXiv:1809</source>
          .
          <volume>05053</volume>
          (
          <year>2018</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Dagan</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Glickman</surname>
            ,
            <given-names>O.</given-names>
          </string-name>
          :
          <article-title>Probabilistic textual entailment: Generic applied modeling of language variability</article-title>
          .
          <source>Learning Methods for Text Understanding and Mining</source>
          <year>2004</year>
          ,
          <fpage>26</fpage>
          -
          <lpage>29</lpage>
          (
          <year>2004</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Dagan</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Glickman</surname>
            ,
            <given-names>O.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Magnini</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          :
          <article-title>The pascal recognising textual entailment challenge</article-title>
          .
          <source>In: Machine Learning Challenges Workshop</source>
          . pp.
          <fpage>177</fpage>
          -
          <lpage>190</lpage>
          . Springer (
          <year>2005</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>Dagan</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Roth</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sammons</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zanzotto</surname>
            ,
            <given-names>F.M.:</given-names>
          </string-name>
          <article-title>Recognizing textual entailment: Models and applications</article-title>
          .
          <source>Synthesis Lectures on Human Language Technologies</source>
          <volume>6</volume>
          (
          <issue>4</issue>
          ),
          <fpage>1</fpage>
          -
          <lpage>220</lpage>
          (
          <year>2013</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>Devlin</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chang</surname>
            ,
            <given-names>M.W.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lee</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Toutanova</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          :
          <article-title>Bert: Pre-training of deep bidirectional transformers for language understanding</article-title>
          . arXiv preprint arXiv:
          <year>1810</year>
          .
          <volume>04805</volume>
          (
          <year>2018</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <surname>Fialho</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Marques</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Martins</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Coheur</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Quaresma</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          : INESCID at ASSIN:
          <article-title>Measuring semantic similarity and recognizing textual entailment</article-title>
          .
          <source>Linguama´tica 8(2)</source>
          ,
          <fpage>33</fpage>
          -
          <lpage>42</lpage>
          (
          <year>2016</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <surname>Fonseca</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          , Alu´ısio, S.M.:
          <article-title>Syntactic knowledge for natural language inference in portuguese</article-title>
          .
          <source>In: International Conference on Computational Processing of the Portuguese Language</source>
          . pp.
          <fpage>242</fpage>
          -
          <lpage>252</lpage>
          . Springer (
          <year>2018</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <string-name>
            <surname>Fonseca</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Santos</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Criscuolo</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          , Alu´ısio, S.:
          <article-title>Vis˜ao geral da avaliac¸a˜o de similaridade semaˆntica e inferˆencia textual</article-title>
          .
          <source>Linguama´tica 8(2)</source>
          ,
          <fpage>3</fpage>
          -
          <lpage>13</lpage>
          (
          <year>2016</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          12.
          <string-name>
            <surname>Levenshtein</surname>
            ,
            <given-names>V.I.</given-names>
          </string-name>
          :
          <article-title>Binary codes capable of correcting deletions, insertions, and reversals</article-title>
          .
          <source>In: Soviet physics doklady</source>
          . vol.
          <volume>10</volume>
          , pp.
          <fpage>707</fpage>
          -
          <lpage>710</lpage>
          (
          <year>1966</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          13.
          <string-name>
            <surname>Marelli</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bentivogli</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Baroni</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bernardi</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Menini</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zamparelli</surname>
          </string-name>
          , R.:
          <article-title>Semeval-2014 task 1: Evaluation of compositional distributional semantic models on full sentences through semantic relatedness and textual entailment</article-title>
          .
          <source>In: Proceedings of the 8th international workshop on semantic evaluation (SemEval</source>
          <year>2014</year>
          ). pp.
          <fpage>1</fpage>
          -
          <lpage>8</lpage>
          (
          <year>2014</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          14.
          <string-name>
            <surname>Papineni</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Roukos</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ward</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zhu</surname>
          </string-name>
          , W.J.:
          <article-title>Bleu: a method for automatic evaluation of machine translation</article-title>
          .
          <source>In: Proceedings of the 40th annual meeting on association for computational linguistics</source>
          . pp.
          <fpage>311</fpage>
          -
          <lpage>318</lpage>
          . Association for Computational Linguistics (
          <year>2002</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          15.
          <string-name>
            <surname>Real</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Fonseca</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          , Gonc¸alo Oliveira, H.:
          <article-title>The ASSIN 2 shared task: Evaluating Semantic Textual Similarity and Textual Entailment in Portuguese</article-title>
          .
          <source>In: Proceedings of the ASSIN 2 Shared Task: Evaluating Semantic Textual Similarity and Textual Entailment in Portuguese</source>
          . p. [In this volume].
          <source>CEUR Workshop Proceedings</source>
          , CEUR-WS.org (
          <year>2020</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          16.
          <string-name>
            <surname>Real</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rodrigues</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          , e
          <string-name>
            <surname>Silva</surname>
            ,
            <given-names>A.V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Albiero</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Thalenberg</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Guide</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Silva</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>de Oliveira Lima</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          , Caˆmara,
          <string-name>
            <surname>I.C.</surname>
          </string-name>
          , Stanojevi´c, M., et al.:
          <article-title>Sick-br: a portuguese corpus for inference</article-title>
          .
          <source>In: International Conference on Computational Processing of the Portuguese Language</source>
          . pp.
          <fpage>303</fpage>
          -
          <lpage>312</lpage>
          . Springer (
          <year>2018</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          17.
          <string-name>
            <surname>Rocha</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lopes</surname>
            <given-names>Cardoso</given-names>
          </string-name>
          , H.:
          <article-title>Recognizing textual entailment: challenges in the portuguese language</article-title>
          .
          <source>Information</source>
          <volume>9</volume>
          (
          <issue>4</issue>
          ),
          <volume>76</volume>
          (
          <year>2018</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          18.
          <string-name>
            <surname>Singh</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>McCann</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Keskar</surname>
            ,
            <given-names>N.S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Xiong</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Socher</surname>
          </string-name>
          , R.: Xlda:
          <article-title>Cross-lingual data augmentation for natural language inference and question answering</article-title>
          . arXiv preprint arXiv:
          <year>1905</year>
          .
          <volume>11471</volume>
          (
          <year>2019</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          19.
          <string-name>
            <surname>Williams</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Nangia</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bowman</surname>
            ,
            <given-names>S.:</given-names>
          </string-name>
          <article-title>A broad-coverage challenge corpus for sentence understanding through inference</article-title>
          .
          <source>In: Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies</source>
          , Volume
          <volume>1</volume>
          (Long Papers). pp.
          <fpage>1112</fpage>
          -
          <lpage>1122</lpage>
          . Association for Computational Linguistics (
          <year>2018</year>
          )
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>