<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <journal-title-group>
        <journal-title>CLiC-it</journal-title>
      </journal-title-group>
      <issn pub-type="ppub">1613-0073</issn>
    </journal-meta>
    <article-meta>
      <title-group>
        <article-title>Textual Entailment with Natural Language Explanations: The Italian e-RTE-3 Dataset</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Andrea Zaninello</string-name>
          <email>azaninello@fbk.eu</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Sofia</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Brenna</string-name>
          <email>sbrenna@fbk.eu</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Bernardo Magnini</string-name>
          <email>magnini@fbk.eu</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Commons License Attribution 4.0 International</institution>
          ,
          <addr-line>CC BY 4.0</addr-line>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Fondazione Bruno Kessler</institution>
          ,
          <addr-line>Trento</addr-line>
          ,
          <country country="IT">Italy</country>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>Free University of Bozen-Bolzano</institution>
          ,
          <country country="IT">Italy</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2023</year>
      </pub-date>
      <volume>9</volume>
      <fpage>0000</fpage>
      <lpage>0001</lpage>
      <abstract>
        <p>understanding. We introduce the 'e-RTE-3-it' dataset, an enriched version of the Italian RTE-3 dataset, where each text-hypothesis pair, in addition to the 'entailment', 'contradiction', or 'neutrality' label, has been enriched with an explanation for the label itself. Moreover, the dataset includes the level of confidence with which the annotators could write the explanation, and in cases where the annotators did not agree with the original label, an alternative label, along with an explanation for the new label. This ofers the opportunity to analyse cases of uncertainty in annotation and delve into diferent perspectives on language Explanations, recognizing textual entailment, lexical resources</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>CEUR
ceur-ws.org</p>
    </sec>
    <sec id="sec-2">
      <title>1. Introduction</title>
    </sec>
    <sec id="sec-3">
      <title>2. Background and Related</title>
    </sec>
    <sec id="sec-4">
      <title>Work</title>
      <p>
        Recently, Large Language Models (LLMs) like T5 [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ],
Recognising Textual Entailment (RTE) emerged as a task
GPT-3.5/4 [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ], LLama-2 [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ], It5 [4], and Camoscio [5]
in 2005 [10], aiming to determine if two sentences have an
have demonstrated impressive performance across
varentailment, contradiction, or neutrality relationship. An
ious natural language processing tasks. Despite their
Italian version of the RTE-3 dataset was later developed to
success, these LLMs also face limitations and risks, such
explore language comprehension and textual entailment
as lack of factuality [6], hallucinations [7], and poor trans- [11].
parency [8]. As a result, there is a growing demand for
The significance of free-form explanations in
enhanc”inherent explainability,” which refers to the ability of ing understanding and interpretability has led to the
models to provide human-like, natural language
explacreation of various datasets. For example, the CODAH
nations for their predictions. Many studies have thus
dataset presents commonsense reasoning problems with
focused on natural language explanations, and numerous
adversarially constructed explanations [12]. Similarly,
datasets have been created for this purpose, primarily
the COPA-SSE dataset ofers crowd-sourced
explanain English [9]. However, there is a notable gap for non- tions for commonsense reasoning tasks [13]. The COS-E
      </p>
      <sec id="sec-4-1">
        <title>English languages, including Italian. dataset couples commonsense reasoning problems with</title>
      </sec>
      <sec id="sec-4-2">
        <title>To fill this void, this paper introduces the ’e-RTE-3explanations [14], providing valuable insights into huit’ dataset, the first Italian dataset for natural language man approaches to these tasks.</title>
        <p>inference enriched with free-form, human-written
ex</p>
      </sec>
      <sec id="sec-4-3">
        <title>The e-SNLI dataset is a relevant resource, as an en</title>
        <p>planations for the relationship between two sentences. riched version of the Stanford Natural Language
InferAdditionally, the dataset includes alternative labels and
ence (SNLI) corpus, containing human-written
explanaconfidence scores from annotators to account for the
tions for entailment decisions [15]. However, this dataset,
variability in human judgments. This aspect of the anno- while valuable for tasks requiring extensive training data,
tation scheme enhances the ’e-RTE-3-it’ dataset, making
is not manually curated and focuses exclusively on the
ability in language understanding1.
it a valuable resource for exploring subjectivity and vari- English language.
nEvelop-O
CEUR
Workshop
Proce dings
htp:/ceur-ws.org
ISN1613-073
1We
© 2023 Copyright for this paper by its authors. Use permitted under Creative</p>
        <p>CEUR</p>
        <p>Workshop Proceedings (CEUR-WS.org)
make the e-RTE-3-it dataset available at the
following
link:</p>
        <p>https://nlplab.fbk.eu/tools-and-resources/
lexical-resources-and-corpora/e-rte-3-ita
3.</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>Methodology</title>
      <p>3.1. Annotation layers
For each text-hypothesis pair in the original Italian RTE
dataset, annotators were asked to provide an explanation
(&lt;e&gt;) for the given label and rate their confidence in
providing that explanation on a 5-point Likert scale. We
also encouraged diversity in perspectives by allowing
annotators to disagree with the original label. 2 &lt;t&gt;Basandosi su uno studio mondiale [...] gli</p>
      <p>In such instances, they provided an explanation for the epidemiologi [...] dimostrano che il fumo e' la
original label as well as an alternative label (&lt;a&gt;) and a causa principale degli incendi e delle morti
corresponding explanation for the new label, along with per incendi nel mondo.&lt;/t&gt;
the level of confidence for the second explanation. 3 &lt;h&gt;Gli incendi domestici sono una causa importante
delle morti da incendio.&lt;/h&gt;
4 &lt;e confidence="4"&gt;Il fumo e' la causa principale
3.2. Data collection and guidelines degli incendi domestici.&lt;/e&gt;&lt;/pair&gt;
We recruited 40 annotators among students (from un- In this case, the explanation was stating something
dergraduate to PhD level) at the University of Bologna. that could not be inferred from the input sentences, and
Annotators were native Italian speakers fluent in at least was re-written by another annotator in the following
one other language, and each took at least one linguistics way.
course, ensuring meta-linguistic proficiency as well as
broader cultural understanding. Annotators were pro1- &lt;e confidence="3"&gt;Il fumo e' la causa principale
vided with 50 text-hypothesis pairs each labeled with an degli incendi e delle morti per incendio, ma
entailment relationship. non e' specificato se un'altra causa importante</p>
      <p>They were asked to write one free-form, natural lan- di morti da incendio siano proprio gli incendi
guage explanation in Italian explaining why the two sen- domestici.&lt;/e&gt;
tences stood in that particular entailment, contradiction,
or neutrality relationship. To ensure language variety as
well as uniformity across labelers, the following guide- 3.4. Original dataset correction
lines were given:
While editing the explanations, the experts also detected
and corrected some errors in the original dataset. In
• please write an explanation in the form of one few cases, these included missing information that
or two self-contained sentences for each &lt;pair, made it impossible to infer the right label for t and h.
label&gt;; For example, consider the following text-hypothesis
• you can refer back to, quote, or paraphrase pair from the test set (id: 52). Text: Oscar Chisini
chunks of both the text and hypothesis; (nato il 4 marzo 1889 a Bergamo, morto il 10
• use case marking and punctuation consistently aprile **1967** a Milano) fu un matematico
with the original sentences; italiano. Lui introdusse la media Chisini
• you can use metalanguage to refer back to the nel 1929. ; hypothesis: Oscar Chisini morì nel
original sentences with phrases such as “in the 1967. ; entailment: "YES". The information within
text, it is stated that...”, “the hypothesis does not stars ** (the year of death) was missing from the Italian
mention...”, etc.; dataset but was present in the original English RTE-3
• please provide your level of confidence (i.e. how dataset. This information was essential to infer the
sure you are about the reasons provided in your entailment relationship, and was re-introduced by
explanation) on a scale from 1 to 5; checking the original English version.
• if you (even partially) disagree with the given Moreover, the Italian RTE-3 dataset, as reported in the
label, provide a new label for the pair, an expla- description2 changed the original label (from ”YES”:
ennation for the new label, and your level of confi- tailment, to ”NO”: contradiction) in 15 pairs, creating a
dence in the new explanation. mismatch with the English dataset. To ensure
comparability, we decided to restore the original label provided
3.3. Post-editing of the explanations by the English dataset, as our annotators were still able
to express an alternative label in case they did not agree
Finally, two diferent linguistics experts post edited the with it. They did so only in the dev set, where they
proexplanations to proofread them, validate them and ad- vided an alternative label ”NO” in pairs 51, 490, 549, and
dress any discrepancies, as well as ensure uniformity and a label ”UNKNOWN” (neutrality) in pair 604. In all other
coherence in spelling. cases, they agreed with the original label.</p>
      <p>In some cases, the experts discarded some explanations For these reasons, the e-RTE-3-it dataset can also be
because of logical errors, in cases when explanations only regarded as an emended, manually curated version of the
paraphrased the input texts or included information not original RTE-3-it dataset.</p>
      <p>originally conveyed by the input texts. For example:
1 &lt;pair id="224" entailment="UNKNOWN" task="IR" length=
"short"&gt;</p>
      <sec id="sec-5-1">
        <title>2The original RTE 3 Italian dataset description can</title>
        <p>be found at https://nlplab.fbk.eu/tools-and-resources/
lexical-resources-and-corpora/rte-3-ita
Original label
New labels in &lt;a&gt;
New labels from entailment
New labels from contradiction
New labels from neutrality
Confidence (mean) in &lt;e&gt;
Confidence (mean) in &lt;e&gt; w/o &lt;a&gt;
Confidence (mean) in &lt;e&gt; with &lt;a&gt;
Confidence (mean) in &lt;a&gt;
4. Dataset Description
1 &lt;?xml version="1.0" encoding="UTF-8"?&gt;
2 &lt;pair id="201" entailment="YES" task="IR" length="</p>
        <p>short"&gt;
3 &lt;t&gt;Berlino ha un nuovo punto di riferimento. Sopra le
gru che ancora dominano l'orizzonte della
nuova capitale dell'Europa adesso c'e' una
cancelleria, dove vivra' il capo del governo
Gerhard Schroeder e il governo tedesco terra' i
suoi incontri regolari.&lt;/t&gt;
4 &lt;h&gt;Nuovi edifici sono stati eretti a Berlino.&lt;/h&gt;
5 &lt;e confidence="4"&gt;La frase "sopra le gru... adesso c'
e' una cancelleria" e' da intendersi in modo
figurato, e indica che e' stato costruito un
nuovo edificio dove ha sede la cancelleria.&lt;/e&gt;
6 &lt;a confidence="5" new_label="UNKNOWN"&gt;Il fatto che
ora sopra le gru c'e' una cancelleria, non
implica che nuovi edifici sono stati eretti a</p>
        <p>Berlino.&lt;/a&gt;
7 &lt;/pair&gt;
an alternative one. If we consider disagreements from
the original label (147 pairs), we observe an increase of
The final dataset comprises 1600 text-hypothesis pairs, the contradiction relationship to 13% and a decrease to
divided into two dev/test splits of 800 pairs each. Each 37% of the neutrality label. As an example, consider the
pair inherits the original dataset’s attributes indicating following ’t-h’ pair:
the pair’s ID, the entailment relation (yes, no, unknown), Text: ”Finora non ci sono segnalazioni di qualche parente
the original task for which the pair was collected, and che abbia reclamato i corpi dei quattro uomini delle forze
whether the text is long or short. Each pair is comple- armate che sono presumibilmente morti quando l’aereo
mented with one explanation and a confidence score. It si schiantò.”
also provides 147 alternative labels with their respective Hypothesis: ”Quattro uomini delle forze armate morirono
explanation and confidence score. In the following, we in uno schianto aereo.”
provide a snippet of the test set. The original label in the Italian RTE-3 dataset is ’YES’.
However, an annotator disagrees and assigns the
alternative label ’NO’, explaining: ”Afermando che quattro
uomini delle forze armate sono presumibilmente morti
quando l’aereo si schiantò, si manifesta una mancata
certezza totale dell’episodio.” The annotator rated their
confidence in this explanation as 4.</p>
        <p>We observed that among the cases where annotators
disagreed with the original label, the ”neutrality” label
’UNKNOWN’ was most frequently revised to
”contradiction” (’NO’, 52) and to ”entailment” (’YES’, 48). Upon
examining the explanations provided for these revised
labels, a common theme emerged: they often stated that
the interpretation of ’h’ needed to assign a neutrality
label was too narrow and did not match with commonsense
reasoning and inferences often made in discourse.</p>
        <p>For example, in a case when an annotator changed the
label from neutrality to entailment, ’t’ and ’h’ stated that
Text: [...] Michael Howard non riuscì a scalzare il Governo
Laburista, sebbene i Conservatori avessero guadagnato 33
seggi”
5. Data analysis Hypothesis: i Conservatori ottennero 33 seggi. Here, the
usual interpretation would be that they obtained at least,
Annotators’ Agreement with Original Labels. Ta- and not exactly 33 seats, explaining that “Guadagnare in
ble 1 reports a detailed description of the dataset. The questo caso è sinonimo di ottenere”.
original labels in the dataset exhibit a distribution of 50% In a case when the annotator changed the label from
’entailment’, 10% ’contradiction’, and 40% ’neutrality’. A ”UNKNOWN” to ”NO” with confidence 4, ’t’ and ’h’ stated
characteristic of our dataset is the allowance for anno- Text: I proprietari di Phinda, l’Ente per la Conservazione
tators to disagree with the original labels and propose con base in Sud Africa, non avrebbero potuto pagare per
una pubblicità migliore per la loro filosofia di tutela della
natura: un approccio alla tutela basato sulle persone, che
sta lentamente guadagnando terreno in Africa poiché le
riserve di caccia sono sempre piü minacciate dalle popolazioni
locali afamate, povere e arrabbiate
Hypothesis: L’Ente per la Conservazione con base in Sud
Africa minaccia la popolazione locale. The explanation
given was that “L’Ente per la Conservazione con base in
Sud Africa basa il suo rapporto di tutuela sulle persone,
sulle popolazioni povere e afamate, quindi aiutandole
non minacciandole”.</p>
        <p>Cases like these underline the subtleties involved in
the inference process, and how tightly it connects to the
interpretation of words in context, which may also be
influenced by some level of subjectivity, an observation
that paves the way for further investigation.</p>
        <p>Lexical variety. We were also interested in the lexical
variety of both the original sentences and the collected
explanations. We noted that while the type/token
ratio for each sentence is very high, indicating that few
words are repeated in the same sentence, if we look at the
lexical overlap between the sentences, we noted a high
overlap between the alternative label explanation and
the hypothesis, even compared to the text. This seems
to indicate that the alternative explanations may rely on
the information in the hypothesis more than the
explanations for the original label, or and that they may be more
’metalinguistic’ in nature, with a tendency to repeat the
whole hypothesis literally.</p>
        <p>mean length
length (tokens)
length (types)
types/tokens ratio
lexical overlapping
t
h
e
a</p>
        <p>Confidence in Explanations. As can be seen in Table
2, the confidence scores assigned by annotators to
explanations were generally high, with a mean score of ’e’ of
4 on a 5-point Likert scale and the highest score being
given to the entailment label. However, when annotators
disagreed with the original label (and no alternative label
was given) the mean confidence score for &lt;e&gt; decreased
to 3.14 and the entailment label became the label with
the lowest score (2.78). The fact that overall confidence
in ’a’ is lower than that of ’e’ seems to indicate that while
annotators felt confident in their judgments when they
agreed with the label, cases involving label revision posed
more challenges and perhaps involved a higher degree
of uncertainty.</p>
      </sec>
    </sec>
    <sec id="sec-6">
      <title>6. Conclusion and future work</title>
      <p>The insights derived from the ’e-RTE-3-it’ dataset pave
the way for multifaceted research directions. The
provided explanations can serve as a gold standard for
training models to generate human-like explanations. Further,
the alternative labels and explanations open avenues for
investigating the subjectivity in language understanding.
The rich layers of the dataset also allow for the study of
correlation between the original and alternative labels,
the confidence score, and the degree of disagreement
among annotators. Future work includes utilizing the
data to develop models capable of providing explanations
for their entailment decisions and conducting a deeper
analysis into the dynamics of subjectivity in the
entailment task.</p>
    </sec>
    <sec id="sec-7">
      <title>Acknowledgments</title>
      <p>This work has been partially supported by the PNRR
project FAIR - Future AI Research (PE00000013), under
the NRRP MUR program funded by the
NextGenerationEU and by the ANTIDOTE project (CHIST-ERA grant
of the Call XAI 2019 of the ANR with the grant number
Project-ANR-21-CHR4- 0002).
P. Mishra, I. Molybog, Y. Nie, A. Poulton, J. Reizen- arXiv:1906.02361 (2019).
stein, R. Rungta, K. Saladi, A. Schelten, R. Silva, E. M. [15] O.-M. Camburu, T. Rocktäschel, T. Lukasiewicz,
Smith, R. Subramanian, X. E. Tan, B. Tang, R. Tay- P. Blunsom, e-snli: Natural language inference
lor, A. Williams, J. X. Kuan, P. Xu, Z. Yan, I. Zarov, with natural language explanations, Advances in
Y. Zhang, A. Fan, M. Kambadur, S. Narang, A. Ro- Neural Information Processing Systems 31 (2018).
driguez, R. Stojnic, S. Edunov, T. Scialom, Llama 2:
Open foundation and fine-tuned chat models, 2023.</p>
      <p>arXiv:2307.09288.
[4] G. Sarti, M. Nissim, It5: Large-scale text-to-text
pretraining for italian language understanding and
generation, ArXiv preprint 2203.03759 (2022). URL:
https://arxiv.org/abs/2203.03759.
[5] A. Santilli, Camoscio: An italian instruction-tuned</p>
      <p>llama, https://github.com/teelinsan/camoscio, 2023.
[6] O. Honovich, R. Aharoni, J. Herzig, H. Taitelbaum,</p>
      <p>D. Kukliansy, V. Cohen, T. Scialom, I. Szpektor,
A. Hassidim, Y. Matias, True: Re-evaluating
factual consistency evaluation, arXiv preprint
arXiv:2204.04991 (2022).
[7] Z. Ji, N. Lee, R. Frieske, T. Yu, D. Su, Y. Xu, E. Ishii,</p>
      <p>Y. Bang, A. Madotto, P. Fung, Survey of
hallucination in natural language generation, ACM
Computing Surveys (2022).
[8] R. Guidotti, A. Monreale, S. Ruggieri, F. Turini, F.
Giannotti, D. Pedreschi, A survey of methods for
explaining black box models, ACM Comput. Surv.
51 (2019) 93:1–93:42. URL: https://doi.org/10.1145/
3236009. doi:10.1145/3236009.
[9] S. Wiegrefe, A. Marasović, Teach me to explain: A
review of datasets for explainable nlp, in:
Proceedings of NeurIPS, 2021. URL: https://arxiv.org/abs/
2102.12060.
[10] I. Dagan, O. Glickman, B. Magnini, The pascal
recognising textual entailment challenge, in:
Machine learning challenges workshop, Springer, 2005,
pp. 177–190.
[11] B. Magnini, A. Lavelli, S. Magnolini, Comparing
machine learning and deep learning approaches on
NLP tasks for the Italian language, in: Proceedings
of the Twelfth Language Resources and Evaluation
Conference, European Language Resources
Association, Marseille, France, 2020, pp. 2110–2119. URL:
https://aclanthology.org/2020.lrec-1.259.
[12] M. Chen, M. D’Arcy, A. Liu, J. Fernandez,</p>
      <p>D. Downey, Codah: An adversarially authored
question-answer dataset for common sense, arXiv
preprint arXiv:1904.04365 (2019).
[13] A. Brassard, B. Heinzerling, P. Kavumba, K. Inui,</p>
      <p>COPA-SSE: semi-structured explanations for
commonsense reasoning, CoRR abs/2201.06777
(2022). URL: https://arxiv.org/abs/2201.06777.</p>
      <p>arXiv:2201.06777.
[14] N. F. Rajani, B. McCann, C. Xiong, R. Socher,</p>
      <p>Explain yourself! leveraging language models
for commonsense reasoning, arXiv preprint</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>C.</given-names>
            <surname>Rafel</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Shazeer</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Roberts</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Lee</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Narang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Matena</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Zhou</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P. J.</given-names>
            <surname>Liu</surname>
          </string-name>
          ,
          <article-title>Exploring the limits of transfer learning with a unified text-totext transformer</article-title>
          ,
          <source>Journal of Machine Learning Research</source>
          <volume>21</volume>
          (
          <year>2020</year>
          )
          <fpage>1</fpage>
          -
          <lpage>67</lpage>
          . URL: http://jmlr.org/papers/ v21/
          <fpage>20</fpage>
          -
          <lpage>074</lpage>
          .html.
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2] OpenAI, Gpt-4
          <source>technical report</source>
          ,
          <year>2023</year>
          . arXiv:
          <volume>2303</volume>
          .
          <fpage>08774</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>H.</given-names>
            <surname>Touvron</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Martin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Stone</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Albert</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Almahairi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Babaei</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Bashlykov</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Batra</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Bhargava</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Bhosale</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Bikel</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Blecher</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C. C.</given-names>
            <surname>Ferrer</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Chen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Cucurull</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Esiobu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Fernandes</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Fu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W.</given-names>
            <surname>Fu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Fuller</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Gao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Goswami</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Goyal</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Hartshorn</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Hosseini</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Hou</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Inan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Kardas</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Kerkez</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Khabsa</surname>
          </string-name>
          ,
          <string-name>
            <surname>I. Kloumann</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Korenev</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P. S.</given-names>
            <surname>Koura</surname>
          </string-name>
          , M.
          <article-title>-</article-title>
          <string-name>
            <surname>A. Lachaux</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          <string-name>
            <surname>Lavril</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          <string-name>
            <surname>Lee</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          <string-name>
            <surname>Liskovich</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          <string-name>
            <surname>Lu</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          <string-name>
            <surname>Mao</surname>
            ,
            <given-names>X.</given-names>
          </string-name>
          <string-name>
            <surname>Martinet</surname>
          </string-name>
          , T. Mihaylov,
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>