<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Enhancing a Text Summarization System with ELMo</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Claudio Mastronardo</string-name>
          <email>claudio.mastronardo@studio.unibo.it</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Fabio Tamburini</string-name>
          <email>fabio.tamburini@unibo.it</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>DISI - University of Bologna</institution>
          ,
          <country country="IT">Italy</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>FICLIT - University of Bologna</institution>
          ,
          <country country="IT">Italy</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>Text summarization has gained a considerable amount of research interest due to deep learning based techniques. We leverage recent results in transfer learning for Natural Language Processing (NLP) using pre-trained deep contextualized word embeddings in a sequence-to-sequence architecture based on pointer-generator networks. We evaluate our approach on the two largest summarization datasets: CNN/Daily Mail and the recent Newsroom dataset. We show how using pre-trained contextualized embeddings on Newsroom improves significantly the state-of-the-art ROUGE-1 measure and obtains comparable scores on the other ROUGE values.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1 Introduction</title>
      <p>The amount of human generated data is
outstanding: every day we generate about 2 quintillion
bytes of unstructured data and this number is
expected to grow. With such a huge amount of
information, swiftly accessing and comprehending
large piece of textual data is becoming more and
more difficult. Automatic text summarization
constitutes a powerful tool which can provide a useful
solution to this problem.</p>
      <p>
        In recent years, automatic text summarization
systems have gained a considerable amount of
research interest due to deep learning powered NLP
impressive results
        <xref ref-type="bibr" rid="ref1 ref16 ref25 ref34 ref38 ref41 ref43 ref7">(Mikolov et al., 2013;
Bahdanau et al., 2015; Yang et al., 2017; Vaswani
et al., 2017; Jo´zefowicz et al., 2016; Devlin et al.,
2019)</xref>
        . Neural network (NN) based approaches
have always been considered data hungry
techniques because they often require a large amount
      </p>
      <p>
        Copyright c 2019 for this paper by its authors. Use
permitted under Creative Commons License Attribution 4.0
International (CC BY 4.0).
of training data, but, in the latest years, several
works have made a huge contribution in this
direction
        <xref ref-type="bibr" rid="ref13 ref26 ref27 ref28 ref29 ref34">(Grusky et al., 2018; Nallapati et al., 2016a;
Napoles et al., 2012)</xref>
        .
      </p>
      <p>
        Text summarization systems can be divided into
two main categories: Extractive and Abstractive
        <xref ref-type="bibr" rid="ref36">(Shi et al., 2018)</xref>
        . The first generate summaries
by purely copying the most representative chunks
from the source text
        <xref ref-type="bibr" rid="ref26 ref27 ref28 ref34 ref8">(Dorr et al., 2003; Nallapati
et al., 2016b)</xref>
        , while in the second summarization
algorithms make up summaries by using novel
phrases and words in order to rephrase and
compress the information in the source text
        <xref ref-type="bibr" rid="ref34 ref6">(Chopra
et al., 2016)</xref>
        . Some works shed light on using both
approaches through hybrid neural architectures
attempting to gather the best characteristics of each
world
        <xref ref-type="bibr" rid="ref18 ref35">(See et al., 2017; Khatri et al., 2017)</xref>
        .
      </p>
      <p>
        NLP has seen a tremendous amount of attention
after several deep learning based important results
        <xref ref-type="bibr" rid="ref14 ref16 ref20 ref34 ref34">(Lample et al., 2016; Jo´zefowicz et al., 2016;
Hermann et al., 2015)</xref>
        . Most of them relied on the
concept of distributed representation of words,
defining them as real-valued vectors learned from data
        <xref ref-type="bibr" rid="ref15 ref15 ref2 ref2 ref25 ref31">(Mikolov et al., 2013; Pennington et al., 2014;
Bojanowski et al., 2017; Joulin et al., 2017)</xref>
        . Recent
results were able to generate richer word
embeddings by exploiting their linguistic context in order
to model word polysemy
        <xref ref-type="bibr" rid="ref22 ref32 ref33">(Peters et al., 2018;
McCann et al., 2017; Peters et al., 2017)</xref>
        .
      </p>
      <p>
        In this paper, we build upon the work
        <xref ref-type="bibr" rid="ref23">of See
et al. (2017</xref>
        ) on the Pointer-Generator Network
for text summarization by integrating it with
recent advances in transfer learning for NLP with
deep contextualized word embeddings, namely an
ELMo model
        <xref ref-type="bibr" rid="ref33">(Peters et al., 2018)</xref>
        . We show that,
using pre-trained deep contextualized word
embeddings, integrating them with pointer-generator
networks and learning the ELMo parameters for
combining the various model layers together with
the text summarization model, we can improve
substantially some of the ROUGE evaluation
metrics. Our experiments were based on two datasets
commonly used to evaluate this task: CNN/Daily
Mail (
        <xref ref-type="bibr" rid="ref17">Nallapati et al., 2016</xref>
        a) and Newsroom
        <xref ref-type="bibr" rid="ref13">(Grusky et al., 2018)</xref>
        .
2
      </p>
    </sec>
    <sec id="sec-2">
      <title>Related work</title>
      <p>
        One of the first neural encoder-decoder
approaches to text summarization has been presented
by
        <xref ref-type="bibr" rid="ref17">Nallapati et al. (2016</xref>
        a) where they show that
an off-the-shelf encoder-decoder framework, used
for machine translation, already outperforms the
previous systems for text summarization. They
also augment input data by concatenating to
classical word embeddings part-of-speech tags,
namedentity tags and tf-idf statistics. They leverage the
hierarchical attention mechanism where less
important chunks of text are less attended with a
chunk-level mechanism attention.
      </p>
      <p>Zhou et al. (2017) propose selective encoding
for text summarization by introducing a selective
gate network into the encoder with the purpose of
distilling salient information from source articles.
Then a second layer called “distilled
representation” is constructed by multiplying the selective
gate to the hidden state of the first layer. Such
gate network can control information flow from
encoder to the decoder and select salient
information, boosting the performances of the sentence
summarization task.</p>
      <p>
        Read-Again Encoding
        <xref ref-type="bibr" rid="ref34 ref42">(Zeng et al., 2016)</xref>
        follows the human approach of reading several times
before writing a summary by using two LSTM
encoders reading the source article and a
transformed version of the first LSTM output
respectively. Another original approach is presented
by Xia et al. (2017) where they follow another
human-driven approach by first writing a draft and
then polishing it looking at the global context. In
an encoder-decoder framework there are two
decoders, the first attends to encoder states and
generates a draft while the second attends to both
the encoder and first decoder outputs generating a
summary by exploiting information from two
context vectors. This approach, called deliberation
network, boosted the performances for both text
summarization and machine translation.
      </p>
      <p>Another set of approaches uses reinforcement
learning as in Chen and Bansal (2018), where they
use two sequence-to-sequence models. The first
is defined as an extractive model with the goal of
extracting salient sentences from the input source.
The second is an abstractive model which
paraphrases and compresses the extracted sentences
into a short summary. They make use of
convolutional neural networks (ConvNet) to encode
tokens and train the two models by using
standard policy gradient methods treating them as
reinforcement learning agents.</p>
      <p>Paulus et al. (2018) presented a new
abstractive summarization model achieving
state-of-theart on the New York Times dataset by
introducing intra-temporal attention in both encoder
and decoder. They use a new objective function
by combining maximum-likelihood cross-entropy
loss and rewards from policy gradient
reinforcement learning in order to reduce the exposure bias
and train their architectures by directly optimizing
the ROUGE score.</p>
      <p>
        Another research direction goes beyond RNNs
to avoid their computational and memory costs
by using ConvNet-based encoder-decoder models.
Kalchbrenner et al. (2016) adopt one-dimensional
convolutions stacking on top of the hidden
representation on the encoder/decoder ConvNet.
QuasiRecurrent Neural Networks
        <xref ref-type="bibr" rid="ref22 ref3">(Bradbury et al.,
2017)</xref>
        use encoders and decoders made of
convolutional layers and dynamic average pooling
layers, requiring less amount of computational time
when compared with LSTMs. Several other
approaches attempted to use ConvNets for NLP.
      </p>
      <p>
        It is also relevant the transformer model
proposed in
        <xref ref-type="bibr" rid="ref38">(Vaswani et al., 2017)</xref>
        which uses only
feed-forward NN and multi-head attention.
3
      </p>
    </sec>
    <sec id="sec-3">
      <title>Datasets</title>
      <p>
        All the experiments in this work have been
conducted on two datasets. The first, the CNN/Daily
Mail dataset (
        <xref ref-type="bibr" rid="ref17">Nallapati et al., 2016</xref>
        a), has been
created by scraping news articles from the cnn.com
website and concatenating news highlights in
order to form a multi-sentence summary. It is
composed of about 300,000 examples. The second, the
recently released Newsroom dataset
        <xref ref-type="bibr" rid="ref13">(Grusky et al.,
2018)</xref>
        consists of 1.3 million article-summaries
pairs. It is the largest and most diverse dataset
known in literature. Compared to CNN/Daily
Mail dataset, Newsroom has been created with
the explicit goal of summarizing articles over two
decades by using 38 major publishers as sources.
Authors in
        <xref ref-type="bibr" rid="ref13">(Grusky et al., 2018)</xref>
        also
demonstrate that CNN/Daily Mail dataset is skewed
towards extractive summaries, while the
Newsroom dataset covers a wider range of
summarization styles, highly abstractive/extractive
summaries and several article-summary compression
ratios. For these reasons, even if we will provide
the results for both datasets, we will mainly
comment them only for the Newsroom dataset.
4
      </p>
    </sec>
    <sec id="sec-4">
      <title>The Proposed Model</title>
      <p>
        Our approach builds upon the work made by See
et al. (2017) on pointer-generator networks
applied to text summarization. The pointer-generator
network is based on the architecture presented in
(
        <xref ref-type="bibr" rid="ref17">Nallapati et al., 2016</xref>
        c).
4.1
      </p>
      <sec id="sec-4-1">
        <title>Pointer-Generator Network</title>
        <p>
          It is an encoder-decoder architecture where tokens
of a source text are fed one-by-one to an encoder
network (a single layer LSTM) which also
generates a sequence of hidden states. The decoder
network (a single layer LSTM), at each step t receives
the embedding of the emitted word at time t 1
and the current decoder’s hidden state. This
architecture makes use of Bahdanau attention
          <xref ref-type="bibr" rid="ref1">(Bahdanau et al., 2015)</xref>
          using:
eit = vT tanh (Whhi + Wsst + battn)
at = softmax et
where st represents the decoder’s hidden state at
step t, hi represents the encoder’s hidden state at
timestep i and eit represents the weight given to hi
at decoder’s timestep t not yet normalized.
Capital letters mark trainable parameters. The
tensor a represents a probability distribution over
encoder’s hidden states and encodes how much to
attend each state in order to alleviate the encoder
from the responsibility of encoding all the
information into a fixed vector. The tensor a is used
to produce a weighted sum of the encoder hidden
states called h which is concatenated to the
decoder’s current hidden state making up the input
tensor for the LSTM cell that produces a
distribution of probability over the vocabulary.
        </p>
        <p>
          Pointer-generator networks extend this
architecture by leveraging ideas from pointer networks
          <xref ref-type="bibr" rid="ref39">(Vinyals et al., 2015)</xref>
          : it is a special kind of
architecture being able to point to a specific input token
and copy it from the source text to the output
sequence. At each time-step t the network produces
a generation probability value pgen 2 [0; 1]
calculated from the context vector h , the decoder’s
state st and the decoder’s input xt:
pgen =
        </p>
        <p>whT ht + wsT st + wxT xt + bptr
again capital letters represent learnable
parameters and indicates the sigmoid function. pgen is
used as a soft switch to choose whether to
generate a word from the network’s vocabulary or copy
a word from the source text. So, given pgen, the
probability of outputting a word w is:
P (w) = pgenPvocab(w) + (1
pgen) Pi:wi=w ait
where Pvocab represents the probability value for
the word w at the output layer of the LSTM
decoder, Pi:wi=w ait is the sum of the attention
values given to the hidden states at time t whose input
word was the specific word w. In the case of an
extremely low pgen, the decoder gives a higher
probability value to the input words which produced
hidden states who had been attended the most.</p>
        <p>At a given time-step t the loss value is computed
as the negative log-likelihood of the ground truth
word wt for that time-step</p>
        <p>losst = log P (wt )
and for a given sequence the loss value is
computed by averaging the losses for each word.</p>
        <p>
          In order to cope with the common repetition
problem
          <xref ref-type="bibr" rid="ref24 ref24 ref34 ref34 ref34 ref37">(Mi et al., 2016; Tu et al., 2016; Sankaran
et al., 2016)</xref>
          , the coverage loss
          <xref ref-type="bibr" rid="ref34 ref37">(Tu et al., 2016)</xref>
          is
used to penalize source-document words attended
too much. It is implemented by maintaining a
coverage vector ct: ct = Ptt0=10 at0 which tracks the
degree of coverage that words have received from
the attention mechanism so far. This leads to the
augmented version of the attention mechanism
including the coverage loss
        </p>
        <p>eit = vT tanh Whhi + Wsst + Wccit + battn
with Wc as learnable parameter. Hence, coverage
loss is computed by:</p>
        <p>covloss t = Pi min ait; cit
in order to prevent repeated attention.
4.2</p>
      </sec>
      <sec id="sec-4-2">
        <title>Deep Contextualized Word Embeddings</title>
        <p>The original pointer-generator network does not
use pre-trained word embeddings, but it learns
128-dimensional word embeddings from scratch
during training. Even though learning
specialized word embeddings for the summarization task
might seem a reasonable approach, we think that
using pre-trained word embeddings could improve
the overall network performance.</p>
        <p>Following Peters et al. (2018) we adopt a
transfer learning approach by leveraging the power of
a
zoo</p>
        <p>V
o
c
a
b
u
l
a
r
y
D
i
s
t
r
i
b
u
t
i
o
n
D
e
c
o
d
e
r
H
i
d
d
e
n
S
t
a
t
e
s
itnno itrunbo
e i
tt t</p>
        <p>A isD
redo end tse</p>
        <p>d ta
cnE iH S
"Argentina"</p>
        <sec id="sec-4-2-1">
          <title>Context Vector</title>
          <p>a
"2-0"
zoo
...</p>
          <p>Germany emerge victorious in
2-0
win against Argentina on Saturday ...
&lt;START&gt; Germany beat</p>
        </sec>
        <sec id="sec-4-2-2">
          <title>Source Text</title>
        </sec>
        <sec id="sec-4-2-3">
          <title>Partial Summary</title>
          <p>pre-trained deep contextualized word embeddings.
Embedding from Language Model (ELMo) is a
particular type of embedding where word
representation is a function of the entire input sequence.
ELMo trains a bidirectional language modeling
architecture inspired by Jo´zefowicz et al. (2016) and
Kim et al. (2016), on a large corpus. In order
to compute the probability for the token tk, the
language model architecture computes a
contextindependent token representation via a ConvNet
over characters and passes the output to a
Llayer bidirectional LSTM. An ELMo
representation is the result of a weighted combination
of the hidden states of the language modeling
architecture. For each token tk, this
architecture computes a set of 2L + 1 representations:
nhkL;Mj jj = 0; : : : ; Lo where hkL;M0 is the
Rk =
output of the ConvNet token layer and hkL;Mj =
h!h kL;Mj ; h kL;Mj i j &gt; 0, for each bi-LSTM layer.</p>
          <p>More generally, in order to use ELMo for a
specific downstream task, word representations are
computed by a weighted sum of each
intermediate network representation:</p>
          <p>ELMo ktask = task PjL=0 sjtask hkL;Mj
where stask are softmax-normalized learnable
weights and task allows to scale the entire
produced vector with respect to the downstream task.</p>
          <p>Our method feeds ELMo embeddings into a
pointer-generator model: as the encoder reads
the source text, a pre-trained ELMo model
generates contextualized word embeddings.
Pointergenerator encoder has two main sources to keep
track of what has been read: its own memory
and the inner information about past and
following words injected into the current word
embedding. We learn the stask and task weights during
training.</p>
          <p>We used the “Original (5.5B)” ELMo
embeddings1. The encoder gets 1024 dimensional
embeddings which are fed into an LSTM cell of 512
neurons followed by a linear layer. Between the
encoder and the decoder there is a neural network
called reduce state with the aim of reducing the
dimensionality of the passed tensors. The decoder
is a bidirectional LSTM with size 256 followed by
two linear layers of 256 neurons. We use an
attention network with Bahdanau’s formula and the
coverage mechanism. Decoder’s vocabulary size
is set to the first most common 50,000 tokens in
the training set. Freezing the model from
learning embeddings from scratch reduces the number
1https://allennlp.org/elmo</p>
        </sec>
      </sec>
      <sec id="sec-4-3">
        <title>Paper R-1</title>
        <p>
          <xref ref-type="bibr" rid="ref35">(See et al., 2017)</xref>
          39.53
          <xref ref-type="bibr" rid="ref30">(Paulus et al., 2018)</xref>
          41.16
          <xref ref-type="bibr" rid="ref12">(Gehrmann et al., 2018)</xref>
          41.22
          <xref ref-type="bibr" rid="ref21">(Liu, 2019)</xref>
          43.25
This work 38.96
of parameters of 2,150,011. We trained our
architecture on both CNN/Daily Mail and Newsroom
datasets using Adagrad as the optimization
algorithm
          <xref ref-type="bibr" rid="ref9">(Duchi et al., 2011)</xref>
          with an initial learning
rate of 0.15 and the initial accumulator set to 0.1.
During training the batch size has been fixed to 8
and we run the decoder for at least 35 steps. As
pre-processing step we just lowercased and
tokenized texts using the nltk python package. The
loss function remained unchanged since we used
the negative log-likelihood for the ground truth
word with coverage loss.
5
        </p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>Experimental Results</title>
      <p>We trained our model for 455,000 iterations on
CNN/Daily Mail and for 520,000 iterations on
Newsroom. The best performing models have
been tested on both CNN/Daily Mail and
Newsroom test sets and the ROUGE metrics are
reported in Table 1 and 2 respectively.</p>
      <p>The proposed approach achieves state-of-the-art
ROUGE-1 value for the Newsroom dataset and
competitive values for ROUGE-2 and
ROUGEL. ELMo addition causes an increase of +14:45,
+13:91 and +11:66 for the three metrics with
respect to basic pointer-generator from Grusky et al.
(2018). ELMo stask learned weights are,
respectively, 0:4140, 0:4690, 0:1169 and task = 0:35.
This shows that the model favours syntactic
information (captured at lower LSTM layers) instead
of semantic information when generating text
embeddings. From a qualitatively point of view we
report some network generated summaries as
supplementary material2. As we can see the model
can generate fairly reasonable summaries, which
can differ from the ground truth but still represent
valid alternatives. This can explain the high value
for ROUGE-1, meaning that summaries’ words
have been covered anyway but in a different order
(causing a lower ROUGE-L).
6</p>
    </sec>
    <sec id="sec-6">
      <title>Discussion and Conclusions</title>
      <p>In this work we leveraged recent results in transfer
learning for NLP with deep contextualized word
embeddings in conjunction with pointer-generator
NN for automatic abstractive text summarization.
We noticed a considerable increase of model’s
performance in terms of the ROUGE score, achieving
state-of-the-art on the Newsroom dataset for the
ROUGE-1 metric. This is a dataset designed for
testing abstractive systems while the other dataset
(CNN/Daily Mail) contains summaries formed by
sentences extracted from the original texts and it is
more suitable for testing extractive systems. Then,
it is reasonable that we got improvements only
when using the Newsroom dataset.</p>
      <p>
        Intrinsic, corpus-based metrics based on string
overlap, string distance, or content overlap, such
as BLEU and ROUGE, suffer from the need to
have a reference output provided by the gold
standard corpus in order to evaluate the system
outputs. That seems very problematic
        <xref ref-type="bibr" rid="ref10 ref12 ref13 ref30 ref33 ref36 ref5">(e.g. see Gatt
and Krahmer 2018)</xref>
        because the reference
summary is only one of the possible summaries that
humans can produce. By looking at the
supplementary material regarding some examples of
our system output, one can immediately recognize
that, even if very different from the reference one,
the summaries produced by the proposed system
are in most cases acceptable.
      </p>
      <p>The definition of proper metrics capturing in the
right way the correctness of system outputs
remains, in our opinion, a critical open issue. As
discussed also in the recent review by Chatzikoumi
(2019) about Machine Translation (MT) metrics,
“When reference translations are used [...] MT
outputs that are very similar to the reference
translation are boosted and not similar MT outputs are
penalised even if they are good”, the so-called
“reference bias”. The same metrics are currently
used also in text summarization leading to similar
problems.</p>
      <p>2https://bit.ly/2XUJvbd</p>
    </sec>
    <sec id="sec-7">
      <title>Acknowledgments</title>
      <p>We gratefully acknowledge the support of
NVIDIA Corporation with the donation of the
Titan Xp GPU used for this research.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          <string-name>
            <given-names>Dzmitry</given-names>
            <surname>Bahdanau</surname>
          </string-name>
          , Kyunghyun Cho, and
          <string-name>
            <given-names>Yoshua</given-names>
            <surname>Bengio</surname>
          </string-name>
          .
          <year>2015</year>
          .
          <article-title>Neural machine translation by jointly learning to align and translate</article-title>
          .
          <source>In Proc. of ICLR</source>
          <year>2015</year>
          , San Diego, CA.
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          <string-name>
            <given-names>Piotr</given-names>
            <surname>Bojanowski</surname>
          </string-name>
          , Edouard Grave, Armand Joulin, and
          <string-name>
            <given-names>Tomas</given-names>
            <surname>Mikolov</surname>
          </string-name>
          .
          <year>2017</year>
          .
          <article-title>Enriching word vectors with subword information</article-title>
          .
          <source>Transactions of the Association for Computational Linguistics</source>
          ,
          <volume>5</volume>
          :
          <fpage>135</fpage>
          -
          <lpage>146</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          <string-name>
            <given-names>James</given-names>
            <surname>Bradbury</surname>
          </string-name>
          , Stephen Merity, Caiming Xiong, and Richard Socher.
          <year>2017</year>
          .
          <article-title>Quasi-recurrent neural networks</article-title>
          .
          <source>In Proc. of ICLR</source>
          <year>2017</year>
          , Toulon, France.
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          <string-name>
            <given-names>Eirini</given-names>
            <surname>Chatzikoumi</surname>
          </string-name>
          .
          <year>2019</year>
          .
          <article-title>How to evaluate machine translation: A review of automated and human metrics</article-title>
          .
          <source>Natural Language Engineering</source>
          , pages
          <fpage>1</fpage>
          -
          <lpage>25</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          <string-name>
            <surname>Yen-Chun Chen</surname>
            and
            <given-names>Mohit</given-names>
          </string-name>
          <string-name>
            <surname>Bansal</surname>
          </string-name>
          .
          <year>2018</year>
          .
          <article-title>Fast abstractive summarization with reinforce-selected sentence rewriting</article-title>
          .
          <source>In Proc. of ACL</source>
          <year>2018</year>
          , pages
          <fpage>675</fpage>
          -
          <lpage>686</lpage>
          , Melbourne, Australia.
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          <string-name>
            <given-names>Sumit</given-names>
            <surname>Chopra</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Michael</given-names>
            <surname>Auli</surname>
          </string-name>
          , and
          <string-name>
            <surname>Alexander</surname>
            <given-names>M.</given-names>
          </string-name>
          <string-name>
            <surname>Rush</surname>
          </string-name>
          .
          <year>2016</year>
          .
          <article-title>Abstractive sentence summarization with attentive recurrent neural networks</article-title>
          .
          <source>In Proc. of NAACL-HLT</source>
          <year>2016</year>
          , pages
          <fpage>93</fpage>
          -
          <lpage>98</lpage>
          , San Diego, California.
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          <string-name>
            <given-names>Jacob</given-names>
            <surname>Devlin</surname>
          </string-name>
          ,
          <string-name>
            <surname>Ming-Wei</surname>
            <given-names>Chang</given-names>
          </string-name>
          ,
          <string-name>
            <given-names>Kenton</given-names>
            <surname>Lee</surname>
          </string-name>
          ,
          <string-name>
            <given-names>and Kristina</given-names>
            <surname>Toutanova</surname>
          </string-name>
          .
          <year>2019</year>
          .
          <article-title>BERT: pre-training of deep bidirectional transformers for language understanding</article-title>
          .
          <source>In Proc. of NAACL</source>
          <year>2019</year>
          , pages
          <fpage>4171</fpage>
          -
          <lpage>4186</lpage>
          , Minneapolis, Minnesota.
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          <string-name>
            <given-names>Bonnie</given-names>
            <surname>Dorr</surname>
          </string-name>
          , David Zajic, and Richard Schwartz.
          <year>2003</year>
          .
          <article-title>Hedge trimmer: A parse-and-trim approach to headline generation</article-title>
          .
          <source>In Proceedings of the HLT-NAACL 03 Text Summarization Workshop</source>
          , pages
          <fpage>1</fpage>
          -
          <lpage>8</lpage>
          , Edmonton, Canada.
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          <string-name>
            <surname>John Duchi</surname>
            , Elad Hazan, and
            <given-names>Yoram</given-names>
          </string-name>
          <string-name>
            <surname>Singer</surname>
          </string-name>
          .
          <year>2011</year>
          .
          <article-title>Adaptive subgradient methods for online learning and stochastic optimization</article-title>
          .
          <source>J. Mach. Learn. Res.</source>
          ,
          <volume>12</volume>
          :
          <fpage>2121</fpage>
          -
          <lpage>2159</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          <string-name>
            <given-names>Albert</given-names>
            <surname>Gatt</surname>
          </string-name>
          and
          <string-name>
            <given-names>Emiel</given-names>
            <surname>Krahmer</surname>
          </string-name>
          .
          <year>2018</year>
          .
          <article-title>Survey of the state of the art in natural language generation: Core tasks, applications and evaluation</article-title>
          . J.
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          <string-name>
            <surname>Artif. Int. Res.</surname>
          </string-name>
          ,
          <volume>61</volume>
          (
          <issue>1</issue>
          ):
          <fpage>65</fpage>
          -
          <lpage>170</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          <string-name>
            <given-names>Sebastian</given-names>
            <surname>Gehrmann</surname>
          </string-name>
          , Yuntian Deng, and
          <string-name>
            <surname>Alexander</surname>
            <given-names>M.</given-names>
          </string-name>
          <string-name>
            <surname>Rush</surname>
          </string-name>
          .
          <year>2018</year>
          .
          <article-title>Bottom-up abstractive summarization</article-title>
          .
          <source>In Proc. of EMNLP</source>
          <year>2018</year>
          , pages
          <fpage>4098</fpage>
          -
          <lpage>4109</lpage>
          , Brussels, Belgium.
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          <string-name>
            <given-names>Max</given-names>
            <surname>Grusky</surname>
          </string-name>
          , Mor Naaman, and
          <string-name>
            <given-names>Yoav</given-names>
            <surname>Artzi</surname>
          </string-name>
          .
          <year>2018</year>
          .
          <article-title>Newsroom: A dataset of 1.3 million summaries with diverse extractive strategies</article-title>
          .
          <source>In Proc. of NAACL2018</source>
          , pages
          <fpage>708</fpage>
          -
          <lpage>719</lpage>
          , New Orleans, Louisiana.
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          <string-name>
            <surname>Karl Moritz</surname>
            <given-names>Hermann</given-names>
          </string-name>
          , Toma´s Kocisky´,
          <string-name>
            <surname>Edward</surname>
            <given-names>Grefenstette</given-names>
          </string-name>
          , Lasse Espeholt, Will Kay, Mustafa Suleyman, and
          <string-name>
            <given-names>Phil</given-names>
            <surname>Blunsom</surname>
          </string-name>
          .
          <year>2015</year>
          .
          <article-title>Teaching machines to read and comprehend</article-title>
          .
          <source>In Proc. of NIPS</source>
          <year>2015</year>
          , pages
          <fpage>1693</fpage>
          -
          <lpage>1701</lpage>
          , Montreal, Canada.
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          <string-name>
            <given-names>Armand</given-names>
            <surname>Joulin</surname>
          </string-name>
          , Edouard Grave, Piotr Bojanowski, and
          <string-name>
            <given-names>Tomas</given-names>
            <surname>Mikolov</surname>
          </string-name>
          .
          <year>2017</year>
          .
          <article-title>Bag of tricks for efficient text classification</article-title>
          .
          <source>In Proc. of EACL</source>
          <year>2017</year>
          , pages
          <fpage>427</fpage>
          -
          <lpage>431</lpage>
          , Valencia, Spain.
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          <string-name>
            <given-names>Rafal</given-names>
            <surname>Jo</surname>
          </string-name>
          ´zefowicz, Oriol Vinyals, Mike Schuster, Noam Shazeer, and
          <string-name>
            <given-names>Yonghui</given-names>
            <surname>Wu</surname>
          </string-name>
          .
          <year>2016</year>
          .
          <article-title>Exploring the limits of language modeling</article-title>
          .
          <source>CoRR, abs/1602</source>
          .02410.
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          <string-name>
            <given-names>Nal</given-names>
            <surname>Kalchbrenner</surname>
          </string-name>
          , Lasse Espeholt, Karen Simonyan, Aa¨ron van den Oord, Alex Graves, and
          <string-name>
            <given-names>Koray</given-names>
            <surname>Kavukcuoglu</surname>
          </string-name>
          .
          <year>2016</year>
          .
          <article-title>Neural machine translation in linear time</article-title>
          .
          <source>CoRR, abs/1610</source>
          .10099.
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          <string-name>
            <given-names>Chandra</given-names>
            <surname>Khatri</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Gyanit</given-names>
            <surname>Singh</surname>
          </string-name>
          ,
          <string-name>
            <given-names>and Nish</given-names>
            <surname>Parikh</surname>
          </string-name>
          .
          <year>2017</year>
          .
          <article-title>Abstractive and extractive text summarization using document context vector and recurrent neural networks</article-title>
          .
          <source>In Proc. of CoNLL</source>
          <year>2016</year>
          , pages
          <fpage>280</fpage>
          -
          <lpage>290</lpage>
          , Berlin, Germany.
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          <string-name>
            <given-names>Yoon</given-names>
            <surname>Kim</surname>
          </string-name>
          , Yacine Jernite, David Sontag, and
          <string-name>
            <surname>Alexander</surname>
            <given-names>M.</given-names>
          </string-name>
          <string-name>
            <surname>Rush</surname>
          </string-name>
          .
          <year>2016</year>
          .
          <article-title>Character-aware neural language models</article-title>
          .
          <source>In Proc. of AAAI</source>
          <year>2016</year>
          , pages
          <fpage>2741</fpage>
          -
          <lpage>2749</lpage>
          , Phoenix, Arizona.
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          <string-name>
            <given-names>Guillaume</given-names>
            <surname>Lample</surname>
          </string-name>
          , Miguel Ballesteros, Sandeep Subramanian, Kazuya Kawakami, and
          <string-name>
            <given-names>Chris</given-names>
            <surname>Dyer</surname>
          </string-name>
          .
          <year>2016</year>
          .
          <article-title>Neural architectures for named entity recognition</article-title>
          .
          <source>In Proc. of NAACL-HLT</source>
          <year>2016</year>
          , pages
          <fpage>260</fpage>
          -
          <lpage>270</lpage>
          , San Diego, California.
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          <string-name>
            <given-names>Yang</given-names>
            <surname>Liu</surname>
          </string-name>
          .
          <year>2019</year>
          .
          <article-title>Fine-tune BERT for extractive summarization</article-title>
          .
          <source>CoRR</source>
          , abs/
          <year>1903</year>
          .10318.
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          <string-name>
            <surname>Bryan</surname>
            <given-names>McCann</given-names>
          </string-name>
          ,
          <string-name>
            <given-names>James</given-names>
            <surname>Bradbury</surname>
          </string-name>
          , Caiming Xiong, and Richard Socher.
          <year>2017</year>
          .
          <article-title>Learned in translation: Contextualized word vectors</article-title>
          .
          <source>In Proc.</source>
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          <source>of NIPS</source>
          <year>2017</year>
          , pages
          <fpage>6297</fpage>
          -
          <lpage>6308</lpage>
          , Long Beach, CA.
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          <string-name>
            <given-names>Haitao</given-names>
            <surname>Mi</surname>
          </string-name>
          , Baskaran Sankaran,
          <string-name>
            <given-names>Zhiguo</given-names>
            <surname>Wang</surname>
          </string-name>
          , and
          <string-name>
            <given-names>Abe</given-names>
            <surname>Ittycheriah</surname>
          </string-name>
          .
          <year>2016</year>
          .
          <article-title>Coverage embedding models for neural machine translation</article-title>
          .
          <source>In Proc. of EMNLP</source>
          <year>2016</year>
          , pages
          <fpage>955</fpage>
          -
          <lpage>960</lpage>
          , Austin, Texas.
        </mixed-citation>
      </ref>
      <ref id="ref25">
        <mixed-citation>
          <string-name>
            <given-names>Tomas</given-names>
            <surname>Mikolov</surname>
          </string-name>
          , G.s Corrado, Kai Chen, and
          <string-name>
            <given-names>Jeffrey</given-names>
            <surname>Dean</surname>
          </string-name>
          .
          <year>2013</year>
          .
          <article-title>Efficient estimation of word representations in vector space</article-title>
          .
          <source>In Proc. of ICLR</source>
          <year>2013</year>
          , Scottsdale, Arizona.
        </mixed-citation>
      </ref>
      <ref id="ref26">
        <mixed-citation>
          <string-name>
            <given-names>Ramesh</given-names>
            <surname>Nallapati</surname>
          </string-name>
          , Bing Xiang, and
          <string-name>
            <given-names>Bowen</given-names>
            <surname>Zhou</surname>
          </string-name>
          . 2016a.
          <article-title>Sequence-to-sequence rnns for text summarization</article-title>
          .
          <source>In Proc. of Workshop track - ICLR</source>
          <year>2016</year>
          , San Juan, Puerto Rico.
        </mixed-citation>
      </ref>
      <ref id="ref27">
        <mixed-citation>
          <string-name>
            <given-names>Ramesh</given-names>
            <surname>Nallapati</surname>
          </string-name>
          , Feifei Zhai, and
          <string-name>
            <given-names>Bowen</given-names>
            <surname>Zhou</surname>
          </string-name>
          . 2016b.
          <article-title>Summarunner: A recurrent neural network based sequence model for extractive summarization of documents</article-title>
          .
          <source>In Proc. of AAAI</source>
          <year>2016</year>
          , Phoenix, Arizona.
        </mixed-citation>
      </ref>
      <ref id="ref28">
        <mixed-citation>
          <string-name>
            <given-names>Ramesh</given-names>
            <surname>Nallapati</surname>
          </string-name>
          , Bowen Zhou,
          <article-title>Cicero dos Santos, C¸ag˘lar Gulc¸ehre, and Bing Xiang</article-title>
          . 2016c.
          <article-title>Abstractive text summarization using sequenceto-sequence RNNs and beyond</article-title>
          .
          <source>In Proc. of The 20th SIGNLL Conference on Computational Natural Language Learning</source>
          , pages
          <fpage>280</fpage>
          -
          <lpage>290</lpage>
          , Berlin, Germany.
        </mixed-citation>
      </ref>
      <ref id="ref29">
        <mixed-citation>
          <string-name>
            <given-names>Courtney</given-names>
            <surname>Napoles</surname>
          </string-name>
          , Matthew Gormley, and Benjamin Van Durme.
          <year>2012</year>
          .
          <article-title>Annotated gigaword</article-title>
          .
          <source>In Proceedings of the Joint Workshop on Automatic Knowledge Base Construction and Web-scale Knowledge Extraction</source>
          , pages
          <fpage>95</fpage>
          -
          <lpage>100</lpage>
          , Montreal, Canada.
        </mixed-citation>
      </ref>
      <ref id="ref30">
        <mixed-citation>
          <string-name>
            <given-names>Romain</given-names>
            <surname>Paulus</surname>
          </string-name>
          , Caiming Xiong, and Richard Socher.
          <year>2018</year>
          .
          <article-title>A deep reinforced model for abstractive summarization</article-title>
          .
          <source>In Proc. of ICLR</source>
          <year>2018</year>
          , Vancouver, Canada.
        </mixed-citation>
      </ref>
      <ref id="ref31">
        <mixed-citation>
          <string-name>
            <surname>Jeffrey</surname>
            <given-names>Pennington</given-names>
          </string-name>
          , Richard Socher, and
          <string-name>
            <given-names>Christoper</given-names>
            <surname>Manning</surname>
          </string-name>
          .
          <year>2014</year>
          .
          <article-title>Glove: Global vectors for word representation</article-title>
          .
          <source>In Proc. of EMNLP</source>
          <year>2014</year>
          , pages
          <fpage>1532</fpage>
          -
          <lpage>1543</lpage>
          , Doha, Qatar.
        </mixed-citation>
      </ref>
      <ref id="ref32">
        <mixed-citation>
          <string-name>
            <given-names>Matthew E.</given-names>
            <surname>Peters</surname>
          </string-name>
          , Waleed Ammar, Chandra Bhagavatula, and
          <string-name>
            <given-names>Russell</given-names>
            <surname>Power</surname>
          </string-name>
          .
          <year>2017</year>
          .
          <article-title>Semisupervised sequence tagging with bidirectional language models</article-title>
          .
          <source>In Proc. of ACL</source>
          <year>2017</year>
          , pages
          <fpage>1756</fpage>
          -
          <lpage>1765</lpage>
          , Vancouver, Canada.
        </mixed-citation>
      </ref>
      <ref id="ref33">
        <mixed-citation>
          <string-name>
            <given-names>Matthew E.</given-names>
            <surname>Peters</surname>
          </string-name>
          , Mark Neumann, Mohit Iyyer, Matt Gardner, Christopher Clark,
          <string-name>
            <given-names>Kenton</given-names>
            <surname>Lee</surname>
          </string-name>
          ,
          <string-name>
            <given-names>and Luke</given-names>
            <surname>Zettlemoyer</surname>
          </string-name>
          .
          <year>2018</year>
          .
          <article-title>Deep contextualized word representations</article-title>
          .
          <source>In Proc. of NAACLHLT</source>
          <year>2018</year>
          , pages
          <fpage>2227</fpage>
          -
          <lpage>2237</lpage>
          , New Orleans, Louisiana.
        </mixed-citation>
      </ref>
      <ref id="ref34">
        <mixed-citation>
          <string-name>
            <given-names>Baskaran</given-names>
            <surname>Sankaran</surname>
          </string-name>
          , Haitao Mi, Yaser Al-Onaizan, and
          <string-name>
            <given-names>Abe</given-names>
            <surname>Ittycheriah</surname>
          </string-name>
          .
          <year>2016</year>
          .
          <article-title>Temporal attention model for neural machine translation</article-title>
          .
          <source>CoRR, abs/1608</source>
          .02927.
        </mixed-citation>
      </ref>
      <ref id="ref35">
        <mixed-citation>
          <string-name>
            <given-names>Abigail</given-names>
            <surname>See</surname>
          </string-name>
          ,
          <string-name>
            <surname>Peter J. Liu</surname>
            , and
            <given-names>Christopher D.</given-names>
          </string-name>
          <string-name>
            <surname>Manning</surname>
          </string-name>
          .
          <year>2017</year>
          .
          <article-title>Get to the point: Summarization with pointer-generator networks</article-title>
          .
          <source>In Proc. of ACL</source>
          <year>2017</year>
          , pages
          <fpage>1073</fpage>
          -
          <lpage>1083</lpage>
          , Vancouver, Canada.
        </mixed-citation>
      </ref>
      <ref id="ref36">
        <mixed-citation>
          <string-name>
            <given-names>Tian</given-names>
            <surname>Shi</surname>
          </string-name>
          , Yaser Keneshloo, Naren Ramakrishnan, and
          <string-name>
            <surname>Chandan</surname>
            <given-names>K.</given-names>
          </string-name>
          <string-name>
            <surname>Reddy</surname>
          </string-name>
          .
          <year>2018</year>
          .
          <article-title>Neural abstractive text summarization with sequence-tosequence models</article-title>
          .
          <source>CoRR</source>
          , abs/
          <year>1812</year>
          .02303.
        </mixed-citation>
      </ref>
      <ref id="ref37">
        <mixed-citation>
          <string-name>
            <given-names>Zhaopeng</given-names>
            <surname>Tu</surname>
          </string-name>
          , Zhengdong Lu, Yang Liu, Xiaohua Liu, and
          <string-name>
            <given-names>Hang</given-names>
            <surname>Li</surname>
          </string-name>
          .
          <year>2016</year>
          .
          <article-title>Modeling coverage for neural machine translation</article-title>
          .
          <source>In Proc. ACL</source>
          <year>2016</year>
          , pages
          <fpage>76</fpage>
          -
          <lpage>85</lpage>
          , Berlin, Germany.
        </mixed-citation>
      </ref>
      <ref id="ref38">
        <mixed-citation>
          <string-name>
            <given-names>Ashish</given-names>
            <surname>Vaswani</surname>
          </string-name>
          , Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones,
          <string-name>
            <given-names>Aidan N.</given-names>
            <surname>Gomez</surname>
          </string-name>
          , Lukasz Kaiser, and
          <string-name>
            <given-names>Illia</given-names>
            <surname>Polosukhin</surname>
          </string-name>
          .
          <year>2017</year>
          .
          <article-title>Attention is all you need</article-title>
          .
          <source>In Proc. of NIPS</source>
          <year>2017</year>
          , Long Beach, CA.
        </mixed-citation>
      </ref>
      <ref id="ref39">
        <mixed-citation>
          <string-name>
            <given-names>Oriol</given-names>
            <surname>Vinyals</surname>
          </string-name>
          , Meire Fortunato, and
          <string-name>
            <given-names>Navdeep</given-names>
            <surname>Jaitly</surname>
          </string-name>
          .
          <year>2015</year>
          .
          <article-title>Pointer networks</article-title>
          .
          <source>In Proc. of NIPS</source>
          <year>2015</year>
          , pages
          <fpage>2692</fpage>
          -
          <lpage>2700</lpage>
          , Montreal, Canada.
        </mixed-citation>
      </ref>
      <ref id="ref40">
        <mixed-citation>
          <string-name>
            <given-names>Yingce</given-names>
            <surname>Xia</surname>
          </string-name>
          , Fei Tian, Lijun Wu,
          <string-name>
            <surname>Jianxin Lin</surname>
            , Tao Qin,
            <given-names>Nenghai</given-names>
          </string-name>
          <string-name>
            <surname>Yu</surname>
          </string-name>
          , and
          <string-name>
            <surname>Tie-Yan Liu</surname>
          </string-name>
          .
          <year>2017</year>
          .
          <article-title>Deliberation networks: Sequence generation beyond one-pass decoding</article-title>
          .
          <source>In Proc. NIPS</source>
          <year>2017</year>
          , pages
          <fpage>1784</fpage>
          -
          <lpage>1794</lpage>
          , Long Beach, CA.
        </mixed-citation>
      </ref>
      <ref id="ref41">
        <mixed-citation>
          <string-name>
            <given-names>Zichao</given-names>
            <surname>Yang</surname>
          </string-name>
          , Zhiting Hu, Yuntian Deng, Chris Dyer, and
          <string-name>
            <given-names>Alex</given-names>
            <surname>Smola</surname>
          </string-name>
          .
          <year>2017</year>
          .
          <article-title>Neural machine translation with recurrent attention modeling</article-title>
          .
          <source>In Proc. of EACL</source>
          <year>2017</year>
          , pages
          <fpage>383</fpage>
          -
          <lpage>387</lpage>
          , Valencia, Spain.
        </mixed-citation>
      </ref>
      <ref id="ref42">
        <mixed-citation>
          <string-name>
            <given-names>Wenyuan</given-names>
            <surname>Zeng</surname>
          </string-name>
          , Wenjie Luo, Sanja Fidler, and
          <string-name>
            <given-names>Raquel</given-names>
            <surname>Urtasun</surname>
          </string-name>
          .
          <year>2016</year>
          .
          <article-title>Efficient summarization with read-again and copy mechanism</article-title>
          . CoRR, abs/1611.03382.
        </mixed-citation>
      </ref>
      <ref id="ref43">
        <mixed-citation>
          <string-name>
            <given-names>Qingyu</given-names>
            <surname>Zhou</surname>
          </string-name>
          , Nan Yang,
          <string-name>
            <given-names>Furu</given-names>
            <surname>Wei</surname>
          </string-name>
          , and
          <string-name>
            <given-names>Ming</given-names>
            <surname>Zhou</surname>
          </string-name>
          .
          <year>2017</year>
          .
          <article-title>Selective encoding for abstractive sentence summarization</article-title>
          .
          <source>In Proc. of ACL</source>
          <year>2017</year>
          , pages
          <fpage>1095</fpage>
          -
          <lpage>1104</lpage>
          , Vancouver, Canada.
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>