<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Deep Bidirectional Transformers for Italian Question Answering</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Danilo Croce</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Giorgio Brandi</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Roberto Basili</string-name>
          <email>basilig@info.uniroma2.it</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Department Of Enterprise Engineering University of Roma</institution>
          ,
          <addr-line>Tor Vergata Via del Politecnico 1, 00133 Roma</addr-line>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2017</year>
      </pub-date>
      <fpage>24</fpage>
      <lpage>26</lpage>
      <abstract>
        <p>English. Deep learning continues to achieve state-of-the-art results in several NLP tasks, such as Question Answering (QA). Unfortunately, the requirements of neural QA systems are very strict in the size of the involved training datasets. Recent works show that the application of Automatic Machine Translation is an enabling factor for the acquisition of large scale QA training sets in resource poor languages such as Italian. In this work, we show how these resources can be used to train a state-of-the-art deep architecture, based on effective techniques recently proposed within the Bidirectional Encoder Representations from Transformers (BERT) paradigm.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>Italiano. I recenti studi sull’applicazione
di metodi di Deep Learning hanno
portato a risultati importanti rispetto a
diversi problemi di Natural Language
Processing, come il Question Answering (QA)
task. Sfortunatamente, i requisiti di tali
sistemi di QA neurali sono molto
stringenti per quanto riguarda le dimensioni
dei dataset necessari per addestrare i
modelli piu´ complessi. Tuttavia, recenti
lavori hanno dimostrato che e´ possibile
applicare tecniche di traduzione
automatica al fine di acquisire collezioni di
esempi di larga scala e addestrare
architetture neurali per il Question Answering
nelle lingue in cui i dati di training sono
scarsi, come l’italiano. In questo
lavoro, mostriamo come queste risorse
permettono l’addestramento di una
architettura neurale molto efficace, basata sul
“Copyright c 2019 for this paper by its authors. Use
permitted under Creative Commons License Attribution 4.0
International (CC BY 4.0).”
paradigma noto come Bidirectional
Encoder Representations from Transformers
(BERT), con risultati che costituiscono lo
stato dell’arte.</p>
    </sec>
    <sec id="sec-2">
      <title>1 Introduction</title>
      <p>
        Question Answering (QA) (
        <xref ref-type="bibr" rid="ref12">(Hirschman and
Gaizauskas, 2001)</xref>
        ) tackles the problem of
returning one or more answers to a question posed by a
user in natural language, using as source a large
knowledge base or, even more often, a large scale
text collection: in this setting, the answers
correspond to sentences (or their fragments) stored in
the text collection. A typical QA process consists
of three main steps: the question processing that
aims at extracting requirements and objectives of
the user’s query, the retrieval phase where
documents and sentences that include the answers are
retrieved from the text collection and the answer
extraction phase that locates the answer within the
candidate sentences
        <xref ref-type="bibr" rid="ref11 ref13">(Harabagiu et al., 2000; Kwok
et al., 2001)</xref>
        .
      </p>
      <p>
        Various QA architectures have been proposed
so far. Some of these rely on structured resources,
such as Freebase, while others use unstructured
information from sources such as Wikipedia (an
example of such a system is the Microsoft’s
AskMSR
        <xref ref-type="bibr" rid="ref3">(Brill et al., 2002)</xref>
        ), or generic Web
pages, e.g. the QuASE system (Sun et al., 2015).
Hybrid models exist as well, that make use of
both the structured and the unstructured
information. These include IBM’s DeepQA
        <xref ref-type="bibr" rid="ref9">(Ferrucci et
al., 2010)</xref>
        and YodaQA
        <xref ref-type="bibr" rid="ref1">(Baudisˇ and Sˇ edivy´, 2015)</xref>
        .
      </p>
      <p>
        In order to initialize such systems, a
manually constructed and annotated dataset is crucial,
from which the mapping between questions and
answers can be learned. Datasets designed for
structured-knowledge based systems, such as
WebQuestions
        <xref ref-type="bibr" rid="ref2">(Berant et al., 2013)</xref>
        , usually contain
the questions, their logical forms and the answers.
On the other side, datasets over unstructured
information are usually composed of question-answer
pairs: WikiMovies
        <xref ref-type="bibr" rid="ref14">(Miller et al., 2016)</xref>
        is an
example of this class of systems and it is made of a
collection of texts from the movie domain. Finally,
some datasets contain the entire triplets made of
the questions, the paragraphs and the answers, that
are expressed as specific spans of the paragraph
and thus located in the paragraph. This is the
case of the recently proposed SQuAD dataset
        <xref ref-type="bibr" rid="ref16">(Rajpurkar et al., 2016)</xref>
        .
      </p>
      <p>
        State-of-the-art approaches proposed in
literature
        <xref ref-type="bibr" rid="ref15 ref18 ref4 ref5 ref6">(Chen et al., 2017; Seo et al., 2017; Clark
and Gardner, 2018; Peters et al., 2018)</xref>
        are based
on neural paradigms and are often portable across
different languages. Among them, the neural
approach presented in
        <xref ref-type="bibr" rid="ref8">(Devlin et al., 2019)</xref>
        , beside
achieving state-of-the-art results in several NLP
tasks, is shown competitive in QA even with
respect to human annotators.
      </p>
      <p>Unfortunately, the limited availability of
training data for languages different from English still
remains an important problem. Even though
multilingual data collections, such as Wikipedia, do
exist for many languages, the portability of the
corresponding annotated resources for supervised
learning algorithms remains limited: large-scale
annotated data mostly exist only for the English
language.</p>
      <p>
        Recent works show that the application of
Automatic Machine Translation enables the acquisition
of large corpora for QA in resource poor languages
such as Italian
        <xref ref-type="bibr" rid="ref6 ref7">(Croce et al., 2018; Croce et al.,
2019)</xref>
        . As a result, SQuAD-IT, i.e., a large scale
dataset made of about 50,000 questions/answer
pairs has been made available. It was not fully
manually validated but still represents a valuable
resource for training neural approaches.
      </p>
      <p>
        In this work, we show how these resources
enable the training of a recent and promising
deep neural architecture, based on the effective
techniques recently justified within the
Bidirectional Encoder Representations from
Transformers (BERT) paradigm
        <xref ref-type="bibr" rid="ref19 ref8">(Vaswani et al., 2017;
Devlin et al., 2019)</xref>
        . The experimental evaluation
carried out with respect to SQuAD-IT confirm the
impressive results of BERT even in Italian QA,
providing state-of-the-art results which are far higher
with respect to previous methods.
      </p>
      <p>In the rest of the paper, section 2 introduces the
BERT architecture for QA. Section 3 report the
experimental evaluation, while Section 4 draws some
conclusions.</p>
    </sec>
    <sec id="sec-3">
      <title>Bidirectional Encoder Representations for QA</title>
      <p>
        In the field of computer vision, researchers have
repeatedly shown the beneficial contribution of
transfer learning, i.e., the pre-training a neural
network model on a known task, for instance image
classification with respect to the ImageNet dataset,
and then performing fine-tuning using the trained
neural network as the basis of a new
purposespecific model, e.g.,
        <xref ref-type="bibr" rid="ref10">(Girshick et al., 2013)</xref>
        .
      </p>
      <p>
        The approach proposed in
        <xref ref-type="bibr" rid="ref8">(Devlin et al., 2019)</xref>
        ,
namely Bidirectional Encoder Representations
from Transformers (BERT) provides a very
effective model to pre-train a deep and complex neural
network over very large scale of unannotated texts
and to apply it to a large variety of NLP task by
simply extending it to each new problem by
finetuning the entire architecture.
      </p>
      <p>
        The building block of BERT is the Transformer
element, an attention-based mechanism that learns
contextual relations between words (or sub-words,
i.e. word pieces,
        <xref ref-type="bibr" rid="ref17">(Schuster and Nakajima, 2012)</xref>
        )
in a text. In its original form, proposed in
        <xref ref-type="bibr" rid="ref19">(Vaswani et al., 2017)</xref>
        , Transformer includes two
separate mechanisms, an encoder that reads the
text input and a decoder that produces a prediction
for the targeted Machine Translation tasks.
      </p>
      <p>
        In line with
        <xref ref-type="bibr" rid="ref15">(Peters et al., 2018)</xref>
        , BERT aims
at providing a sentence embedding (as well as
the contextualized embeddings of each word
composing the sentence) where the pre-training stage
aims at acquiring an expressive and robust
language model, where only the encoder is used. As
shown in Figure 1 (on the left) the Transformer
encoder reads the entire sequence of words at once
and acquire a language model by reconstructing
the original sentence applying a MLM (masked
language model) pre-training objective: the MLM
randomly masks some of the tokens from the
input, and the objective is to predict the original
masked word based only on its context. In addition
to the masked language model, BERT also uses a
next sentence prediction task that jointly pre-trains
text-pair representations. This last objective is
crucial to improve the network capability of modeling
relational information between text pairs, which is
particularly important in tasks such as QA in order
to relate an answer to a question.
      </p>
      <p>After the language model is trained over a
generic document collection, the BERT
architecture allows encoding (i) specific words
belonging to a sentence, (ii) the entire sentence and (iii)
sentence pairs with dedicated embeddings. These
can be used in input to further deep architectures
to solve sentence classification, sequence
labeling or relational learning tasks by simply adding
simple layers and fine-tuning the entire
architecture. On top of such embeddings, fine-tuning is
applied by adding task specific and simple layers
on top of the architecture acquiring the language
model. In a nutshell, this layer introduces
minimal task-specific parameters, and is trained on
the targeted tasks by simply fine-tuning all
pretrained parameters, optimizing the performance on
the specific problem. The straightforward
application of BERT has shown better results than
previous state-of-the-art models on a wide spectrum of
natural language processing tasks.</p>
      <p>
        One of the most impressive results was achieved
with respect to the Question Answering task
proposed by
        <xref ref-type="bibr" rid="ref16">(Rajpurkar et al., 2016)</xref>
        : given a question
and a passage from Wikipedia containing the
answer, the task is to predict the answer text span
in the passage. An example of paragraph,
showing the Wikipedia answer to the question “What
was Marie Curie the first female recipient of?”
is reported in Figure 2. This specific task
originated the Stanford Question Answering Dataset
(SQuAD), a collection of 100k crowd-sourced
question/answer pairs.
      </p>
      <p>The fine-tuning process of BERT in the QA task
(shown on the right side of Figure 1) requires to
encode the input question and passage as a generic
text pair, such as the ones used for the next
sentence prediction task used in the initial training
stages.</p>
      <p>
        In order to determine the correct span for the
answer,
        <xref ref-type="bibr" rid="ref8">(Devlin et al., 2019)</xref>
        introduces on top
of embeddings encoding the words of the
question/answer pairs a so-called start vector S 2 RH
(with H the dimensionality of the embedding
produced for each wordpiece Ti) and an end vector
S 2 RH . Then, the probability of word i being
the start of the answer span is computed as a dot
product between the associated embedding Ti and
S followed by a softmax layer over all of the words
in the paragraph: Pi = PeSeTSiTj . The analogous
j
formula is used for the end of the answer span.
The score of a candidate span from position i to
position j is defined as S Ti +E Tj , and the
maximum scoring span where j i is used as a
prediction. The training objective is the sum of the
loglikelihoods of the correct start and end positions.
The above fine-tuning of BERT achieved
state-ofthe-art results over the official benchmarking
campaign related to SQuAD and, most noticeably, its
accuracy is comparable to the one observed in
human annotators1.
      </p>
      <p>It is worth noting that no bias over the input
lan1https://rajpurkar.github.io/SQuAD-explorer/
guage exists, so that the language model
underlying BERT can be acquired over any text collection
independently from the input language. As a
consequence a pre-trained model acquired over
documents written in more than one hundred languages
exists. It will be applied in the next section to train
and evaluate such a QA model over a dataset of
examples in Italian.
3</p>
    </sec>
    <sec id="sec-4">
      <title>Experimental Evaluation</title>
      <p>
        In order to assess the applicability of the BERT
architecture against the targeted QA task, a
multilingual pre-trained model has been downloaded2:
in particular, this model has been acquired over
documents written in one hundred languages, it is
composed of 12 layers of Transformers and
associates each token in input to a word embedding
made of 768 dimensions. For consistency with
        <xref ref-type="bibr" rid="ref8">(Devlin et al., 2019)</xref>
        , 5 epochs have been
considered to fine-tune the model.
      </p>
      <p>
        We trained the architecture over SQuAD-IT3,
2https://storage.googleapis.com/bert models/
2018 11 23/multi cased L-12 H-768 A-12.zip
3https://github.com/crux82/squad-it
a dataset made available by
        <xref ref-type="bibr" rid="ref7">(Croce et al., 2019)</xref>
        .
This dataset includes more than 50,000
question/paragraph pairs obtained by automatic
translating the original SQuAD dataset. The details
about the number of sentences is reported in Table
1 where a comparison with the original SQuAD in
English is reported.
      </p>
      <p>The parameters of the neural network were set
equal to those of the original work, including the
word embeddings resource. Two evaluation
metrics are used: exact string match (EM) and the
F1 score, which measures the weighted average of
precision and recall at the token level. EM is a
stricter measure evaluated as the percentage of
answers perfectly retrieved by the systems, i.e. the
text extracted by the span produced by the
system is exactly the same as the gold-standard. The
adopted token-based F1 score smooths this
constraint by measuring the overlap (the number of
shared tokens) between the provided answers and
the gold standard.</p>
      <p>
        Performances are reported in Table 2 together
with the results achieved by a variant of the DrQA
system
        <xref ref-type="bibr" rid="ref4">(Chen et al., 2017)</xref>
        , evaluated against the
same SQuAD-IT dataset, as from
        <xref ref-type="bibr" rid="ref7">(Croce et al.,
2019)</xref>
        . Improvements are impressive, as both EM
and F1 are improved of more than 10%. Anyway,
these results are in line with the impact of BERT
over the original English dataset. In the final
version of this paper we will provide an in depth
comparison between DrQA and BERT.
      </p>
    </sec>
    <sec id="sec-5">
      <title>Conclusions</title>
      <p>This paper explores the application of
Bidirectional Encoder Representations within the QA task
in Italian, enabled by the recent availability of a
large-scale annotated corpus, SQuAD-IT. The
experimental results confirm the robustness of the
adopted Transformer-based architecture, with a
significant improvement with respect to earlier
neural architectures. This result paves the way to
the development of portable, robust and accurate
neural models for QA in Italian, and future work
will certainly consider other possible extensions of
the adopted model.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          <article-title>Petr Baudisˇ and Jan Sˇ edivy´</article-title>
          .
          <year>2015</year>
          .
          <article-title>Modeling of the Question Answering Task in the YodaQA System</article-title>
          . In Josanne Mothe, Jacques Savoy, Jaap Kamps, Karen Pinel-Sauvagnat, Gareth Jones, Eric San Juan, Linda Capellato, and Nicola Ferro, editors,
          <source>Experimental IR Meets Multilinguality, Multimodality, and Interaction</source>
          , pages
          <fpage>222</fpage>
          -
          <lpage>228</lpage>
          , Cham. Springer International Publishing.
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          <string-name>
            <given-names>Jonathan</given-names>
            <surname>Berant</surname>
          </string-name>
          , Andrew Chou, Roy Frostig, and
          <string-name>
            <given-names>Percy</given-names>
            <surname>Liang</surname>
          </string-name>
          .
          <year>2013</year>
          .
          <article-title>Semantic parsing on freebase from question-answer pairs</article-title>
          .
          <source>In EMNLP</source>
          , pages
          <fpage>1533</fpage>
          -
          <lpage>1544</lpage>
          . ACL.
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          <string-name>
            <given-names>E.</given-names>
            <surname>Brill</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Dumais</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Banko</surname>
          </string-name>
          , Eric Brill, Michele Banko, and
          <string-name>
            <given-names>Susan</given-names>
            <surname>Dumais</surname>
          </string-name>
          .
          <year>2002</year>
          .
          <article-title>An Analysis of the AskMSR Question-Answering System</article-title>
          .
          <source>In Proceedings of EMNLP</source>
          <year>2002</year>
          ,
          <article-title>January</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          <string-name>
            <given-names>Danqi</given-names>
            <surname>Chen</surname>
          </string-name>
          , Adam Fisch, Jason Weston, and
          <string-name>
            <given-names>Antoine</given-names>
            <surname>Bordes</surname>
          </string-name>
          .
          <year>2017</year>
          .
          <article-title>Reading wikipedia to answer opendomain questions</article-title>
          .
          <source>In Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)</source>
          , pages
          <fpage>1870</fpage>
          -
          <lpage>1879</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          <string-name>
            <given-names>Christopher</given-names>
            <surname>Clark</surname>
          </string-name>
          and
          <string-name>
            <given-names>Matt</given-names>
            <surname>Gardner</surname>
          </string-name>
          .
          <year>2018</year>
          .
          <article-title>Simple and effective multi-paragraph reading comprehension</article-title>
          .
          <source>In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)</source>
          , pages
          <fpage>845</fpage>
          -
          <lpage>855</lpage>
          , Melbourne, Australia, July. Association for Computational Linguistics.
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          <string-name>
            <given-names>Danilo</given-names>
            <surname>Croce</surname>
          </string-name>
          , Alexandra Zelenanska, and
          <string-name>
            <given-names>Roberto</given-names>
            <surname>Basili</surname>
          </string-name>
          .
          <year>2018</year>
          .
          <article-title>Neural learning for question answering in italian</article-title>
          . In Chiara Ghidini, Bernardo Magnini, Andrea Passerini, and Paolo Traverso, editors,
          <source>AI*IA 2018 - Advances in Artificial Intelligence</source>
          , pages
          <fpage>389</fpage>
          -
          <lpage>402</lpage>
          , Cham. Springer International Publishing.
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          <string-name>
            <given-names>Danilo</given-names>
            <surname>Croce</surname>
          </string-name>
          , Alexandra Zelenanska, and
          <string-name>
            <given-names>Roberto</given-names>
            <surname>Basili</surname>
          </string-name>
          .
          <year>2019</year>
          .
          <article-title>Enabling deep learning for large scale question answering in italian</article-title>
          .
          <source>Intelligenza Artificiale</source>
          ,
          <volume>13</volume>
          (
          <issue>1</issue>
          ):
          <fpage>49</fpage>
          -
          <lpage>61</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          <string-name>
            <given-names>Jacob</given-names>
            <surname>Devlin</surname>
          </string-name>
          ,
          <string-name>
            <surname>Ming-Wei</surname>
            <given-names>Chang</given-names>
          </string-name>
          ,
          <string-name>
            <given-names>Kenton</given-names>
            <surname>Lee</surname>
          </string-name>
          ,
          <string-name>
            <given-names>and Kristina</given-names>
            <surname>Toutanova</surname>
          </string-name>
          .
          <year>2019</year>
          .
          <article-title>BERT: Pre-training of deep bidirectional transformers for language understanding</article-title>
          .
          <source>In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies</source>
          , Volume
          <volume>1</volume>
          (Long and Short Papers), pages
          <fpage>4171</fpage>
          -
          <lpage>4186</lpage>
          , Minneapolis, Minnesota, June. Association for Computational Linguistics.
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          <string-name>
            <given-names>David A.</given-names>
            <surname>Ferrucci</surname>
          </string-name>
          , Eric W. Brown, Jennifer ChuCarroll, James Fan, David Gondek,
          <string-name>
            <given-names>Aditya</given-names>
            <surname>Kalyanpur</surname>
          </string-name>
          , Adam Lally,
          <string-name>
            <given-names>J. William</given-names>
            <surname>Murdock</surname>
          </string-name>
          , Eric Nyberg, John M. Prager, Nico Schlaefer, and
          <string-name>
            <surname>Christopher</surname>
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Welty</surname>
          </string-name>
          .
          <year>2010</year>
          .
          <article-title>Building Watson: An Overview of the DeepQA Project</article-title>
          .
          <source>AI Magazine</source>
          ,
          <volume>31</volume>
          (
          <issue>3</issue>
          ):
          <fpage>59</fpage>
          -
          <lpage>79</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          <string-name>
            <surname>Ross B. Girshick</surname>
            , Jeff Donahue, Trevor Darrell, and
            <given-names>Jitendra</given-names>
          </string-name>
          <string-name>
            <surname>Malik</surname>
          </string-name>
          .
          <year>2013</year>
          .
          <article-title>Rich feature hierarchies for accurate object detection and semantic segmentation</article-title>
          .
          <source>CoRR, abs/1311</source>
          .2524.
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          <string-name>
            <surname>Sanda M. Harabagiu</surname>
            ,
            <given-names>Dan I. Moldovan</given-names>
          </string-name>
          , Marius Pasca, Rada Mihalcea, Mihai Surdeanu, Razvan C. Bunescu, Roxana Girju, Vasile Rus, and
          <string-name>
            <given-names>Paul</given-names>
            <surname>Morarescu</surname>
          </string-name>
          .
          <year>2000</year>
          .
          <article-title>FALCON: boosting knowledge for answer engines</article-title>
          .
          <source>In Proceedings of The Ninth Text REtrieval Conference</source>
          , TREC 2000, Gaithersburg, Maryland, USA, November
          <volume>13</volume>
          -
          <issue>16</issue>
          ,
          <year>2000</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          <string-name>
            <given-names>L.</given-names>
            <surname>Hirschman</surname>
          </string-name>
          and
          <string-name>
            <given-names>R.</given-names>
            <surname>Gaizauskas</surname>
          </string-name>
          .
          <year>2001</year>
          .
          <article-title>Natural language question answering: the view from here</article-title>
          .
          <source>Natural Language Engineering</source>
          ,
          <volume>7</volume>
          (
          <issue>4</issue>
          ):
          <fpage>275</fpage>
          -
          <lpage>300</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          <string-name>
            <surname>Cody C. T. Kwok</surname>
            ,
            <given-names>Oren</given-names>
          </string-name>
          <string-name>
            <surname>Etzioni</surname>
          </string-name>
          , and
          <string-name>
            <surname>Daniel</surname>
            <given-names>S.</given-names>
          </string-name>
          <string-name>
            <surname>Weld</surname>
          </string-name>
          .
          <year>2001</year>
          .
          <article-title>Scaling question answering to the web</article-title>
          .
          <source>In WWW</source>
          , pages
          <fpage>150</fpage>
          -
          <lpage>161</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          <string-name>
            <surname>Alexander H. Miller</surname>
            , Adam Fisch, Jesse Dodge, AmirHossein Karimi, Antoine Bordes, and
            <given-names>Jason</given-names>
          </string-name>
          <string-name>
            <surname>Weston</surname>
          </string-name>
          .
          <year>2016</year>
          .
          <article-title>Key-value memory networks for directly reading documents</article-title>
          .
          <source>In EMNLP.</source>
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          <string-name>
            <given-names>Matthew</given-names>
            <surname>Peters</surname>
          </string-name>
          , Mark Neumann, Mohit Iyyer, Matt Gardner, Christopher Clark,
          <string-name>
            <given-names>Kenton</given-names>
            <surname>Lee</surname>
          </string-name>
          ,
          <string-name>
            <given-names>and Luke</given-names>
            <surname>Zettlemoyer</surname>
          </string-name>
          .
          <year>2018</year>
          .
          <article-title>Deep contextualized word representations</article-title>
          .
          <source>In Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies</source>
          , Volume
          <volume>1</volume>
          (
          <issue>Long Papers)</issue>
          , pages
          <fpage>2227</fpage>
          -
          <lpage>2237</lpage>
          , New Orleans, Louisiana, June. Association for Computational Linguistics.
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          <string-name>
            <given-names>Pranav</given-names>
            <surname>Rajpurkar</surname>
          </string-name>
          , Jian Zhang, Konstantin Lopyrev, and
          <string-name>
            <given-names>Percy</given-names>
            <surname>Liang</surname>
          </string-name>
          .
          <year>2016</year>
          . SQuAD:
          <volume>100</volume>
          .000+
          <article-title>Questions for Machine Comprehension of Text</article-title>
          . CoRR, abs/1606.05250.
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          <string-name>
            <given-names>Mike</given-names>
            <surname>Schuster</surname>
          </string-name>
          and
          <string-name>
            <given-names>Kaisuke</given-names>
            <surname>Nakajima</surname>
          </string-name>
          .
          <year>2012</year>
          .
          <article-title>Japanese and korean voice search</article-title>
          .
          <source>In International Conference on Acoustics, Speech and Signal Processing</source>
          , pages
          <fpage>5149</fpage>
          -
          <lpage>5152</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          <string-name>
            <given-names>Min</given-names>
            <surname>Joon</surname>
          </string-name>
          <string-name>
            <surname>Seo</surname>
          </string-name>
          , Aniruddha Kembhavi, Ali Farhadi, and
          <string-name>
            <given-names>Hannaneh</given-names>
            <surname>Hajishirzi</surname>
          </string-name>
          .
          <year>2017</year>
          .
          <article-title>Bidirectional attention flow for machine comprehension</article-title>
          .
          <source>In 5th International Conference on Learning Representations, Huan Sun</source>
          , Hao Ma, Wen tau Yih,
          <string-name>
            <surname>Chen-Tse</surname>
            <given-names>Tsai</given-names>
          </string-name>
          , Jingjing Liu, and
          <string-name>
            <surname>Ming-Wei Chang</surname>
          </string-name>
          .
          <year>2015</year>
          .
          <article-title>Open domain question answering via semantic enrichment</article-title>
          .
          <source>In WWW.</source>
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          <string-name>
            <given-names>Ashish</given-names>
            <surname>Vaswani</surname>
          </string-name>
          , Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez,
          <article-title>Ł ukasz Kaiser, and</article-title>
          <string-name>
            <given-names>Illia</given-names>
            <surname>Polosukhin</surname>
          </string-name>
          .
          <year>2017</year>
          .
          <article-title>Attention is all you need</article-title>
          . In I. Guyon,
          <string-name>
            <given-names>U. V.</given-names>
            <surname>Luxburg</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Bengio</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Wallach</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Fergus</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Vishwanathan</surname>
          </string-name>
          , and R. Garnett, editors,
          <source>Advances in Neural Information Processing Systems</source>
          <volume>30</volume>
          , pages
          <fpage>5998</fpage>
          -
          <lpage>6008</lpage>
          . Curran Associates, Inc.
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>