<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Neural Abstractive Text Summarization</article-title>
      </title-group>
      <contrib-group>
        <aff id="aff0">
          <label>0</label>
          <institution>Department of Computer Science University of Bari “Aldo Moro”</institution>
        </aff>
      </contrib-group>
      <abstract>
        <p>ive text summarization, where our aim is to investigate methods to infuse prior knowledge into deep neural networks. We believe that these approaches can obtain better performance than the state-of-the-art models for generating well-formed and meaningful summaries.</p>
      </abstract>
      <kwd-group>
        <kwd>Natural Language Processing</kwd>
        <kwd>Abstractive Text Summarization</kwd>
        <kwd>Recurrent Neural Networks</kwd>
        <kwd>Deep Learning</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Introduction</title>
      <p>Information overload is a problem in modern digital society caused by the
explosion of the amount of information produced on both the World Wide
Web and the enterprise environments. For textual information, this problem is
even more significant due to the high cognitive load required for reading and
understanding a text. Automatic text summarization tools are thus useful to
quickly understand a large amount of information.</p>
      <p>The goal of summarization is to produce a shorter version of a source text
by preserving the meaning and the key contents of the original. This is a very
complex problem since it requires to emulate the cognitive capacity of human
beings to generate summaries. For this reason, text summarization poses open
challenges in both natural language understanding and generation. Due to the
difficulty of this task, research works in literature focused on the extractive
aspect of summarization, where the generated summary is a selection of relevant
sentences from the source text in a copy-paste fashion [ ] [ ]. Over the past years,
few works have been proposed to solve the abstractive problem of summarization,
which aims to produce from scratch a new cohesive text not necessarily present
in the original source [ ] [ ].</p>
      <p>Abstractive summarization requires deep understanding and reasoning over
the text, determining the explicit or implicit meaning of each element, such as
words, phrases, sentences and paragraphs, and making inferences about their
properties [ ] in order to generate new sentences which compose the summary.</p>
      <p>Recently, riding the wave of prominent results of modern deep learning models
in many natural language processing tasks [ ], several groups have started to
exploit deep neural networks for abstractive text summarization [ ] [ ]. These
deep architectures share the idea of casting the summarization task as a neural
machine translation problem [ ], where the models, trained on a large amount
of data, learn the alignments between the input text and the target summary
through an attention encoder-decoder paradigm. In detail, in [ ] the authors
propose a feed-forward neural network based on neural language model [ ] with
an attention-based encoder, while the models proposed in [ ] and [ ] use the
attention encoder into a sequence-to-sequence framework modeled by RNNs [ ].
Once parametric models are trained, a decoder module greedily generates a
summary, word by word, through a beam search algorithm.</p>
      <p>The aim of these works based on neural networks is to provide a fully
datadriven approach to solve the abstractive summarization task, where the models
learn automatically the representation of relationships between the words in
the input document and those in the output summary without using complex
handcrafted linguistic features. Indeed, the experiments highlight significant
improvements of these deep architectures compared to extractive and
abstractive state-of-the-art methods evaluated on various datasets, including the
goldstandard DUC- [ ] using several variants of ROUGE metric [ ]. These
results prove the effectiveness of the approaches based on modern deep
learning architectures to solve the abstractive summarization task and this lays the
foundation for a new promising area of research that we want to explore.</p>
    </sec>
    <sec id="sec-2">
      <title>Motivation and Research Questions</title>
      <p>The proposed neural attention-based models for abstractive summarization are
still in an early stage, thus they show some limitations. Firstly, they require a large
amount of training data in order to capture a good representation that properly
maps good (soft) alignments between original text and the related summary.
Moreover, since these deep models learn the linguistic regularities relying only
on statistical co-occurrences of words over the training set, some grammar and
semantic errors can occur in the generated summaries. Finally, these models
work only at sentence level and are effective for sentence compression rather than
document summarization, where both input text and target summary consist of
several sentences.</p>
      <p>In this work we argue about our ongoing research on abstractive text
summarization. Taking up the idea of setting the summarization task as a
sequence-tosequence learning problem [ ], we study approaches to infuse prior knowledge
into a RNN in a unified manner in order to overtake the aforementioned limits.</p>
      <p>In the first stage of our research we focus on methodologies to introduce
linguistic features, such as part-of-speech and named entities tags, and relational
semantic information coming from knowledge bases and thesaurus, such as
DBpedia and WordNet. We believe that informing the neural network about the
specific syntactic and semantic role of each word (or concept) during the training
phase may led several advantages as described below.</p>
      <p>Introducing information about the syntactical role of each word, the neural
network can tend to learn the right collocation of words by belonging to a certain
part-of-speech class. Besides, for standard neural language models, the named
entities are usually considered as rare words. This brings some shortcomings
during new text generation, where the entities are treated as unknown words [ ].
Thus, the linguistic features can improve the model avoiding grammar errors and
producing well-formed summaries. Also, a jointly learning of word and knowledge
embedding can induce the model to produce more meaningful summaries. Finally,
the summarization task lacks of availability of data required to train the models,
especially in specific domains. The introduction of prior knowledge can help to
reduce the amount of data needed in the training phase.</p>
      <p>Concretely, our research wants to answer the following questions:
RQ
RQ
RQ
RQ
RQ
RQ</p>
      <p>How to formalize the abstractive text summarization as a learning problem
using deep neural networks?</p>
      <p>What is an effective way to introduce prior knowledge into deep neural
models?</p>
      <p>How to combine the distributional and relational semantic in a unified
model?</p>
      <p>Can the proposed models reduce the syntactic and semantic errors in the
generated summaries?</p>
      <p>How to extend the models so that they can be applied to summarization
of paragraphs, documents and multi-documents?</p>
      <p>How to evaluate the proposed models?</p>
    </sec>
    <sec id="sec-3">
      <title>Methodology</title>
      <p>At current state of our research, we are designing a novel approach to incorporate
prior knowledge into deep neural networks. Our general deep architecture is
inspired by [ ], where we formalize the abstractive summarization as a
sequenceto-sequence problem by adopting an encoder-decoder paradigm using RNNs.
Figure shows a graphical example.</p>
      <p>These models learn soft alignments between source and target sentences from
a training set composed by original text and target summary pairs. In detail,
the encoder is a RNN that reads one token at time from the input source and
returns a fixed-size vector representing the input text. The decoder is another
RNN that generates words for the summary and it is conditioned by the vector
representation returned by the first network. In order to find the best sequence
of words that represent a summary, a beam search algorithm is commonly used.
The RNNs of the encoder and decoder can be implemented with Elman RNN [ ]
or with a more sophisticated variants as Long-Short Term Memory (LSTM) [ ]
and Gated Recurrent Unit (GRU) [ ] networks. In tasks involving language
modeling, these variants have shown impressive performance and they solve
the vanishing gradient problem that typically involves RNNs. Several tricks,
such as to read the input sequence in reverse or bidirectional mode, are applied
to the encoder to improve the performance. Moreover, some attention-based
mechanisms [ ] [ ] [ ] [ ] are integrated into the encoder to help the network
to remember certain aspects about the input. The good performance of the
whole architecture often depends on how these attention-based components are
modeled.</p>
      <p>These models take in account only the distribution of the words in the training
corpus. At first stage of our research, we want to incorporate lexical and syntactic
features, such as part-of-speech and named entities tags, into RNNs. The core
idea is to replace the softmax of each RNN layer with a log-linear model or a
probabilistic graphical model, like factor graphs. This replacement does not arise
any problem because the softmax function converts the output of the network
into probability values, where the softmax can be seen as a special case of the
extended version of RNN [ ]. Thus, the use of probabilistic models allows to
condition the probability value, given an extra feature vector that represents the
lexical and syntactic information of each word. We believe that this approach
can learn a better representation of the input vector during the training and it
can help the decoder in the generation phase. In this way, the decoder can assign
to the next word a probability value which is related to the specific lexical role
of that word in the generated summary. This can allow the model to decrease
the number of grammar errors in the summary, even using a smaller training set
since the linguistic regularities are supported by the extra vector of syntactic
features.</p>
      <p>In the next step, we want to explore approaches that combine the distributional
and relational semantic in a unified model. For this purpose, in the recent
literature, several works [ ] [ ] [ ] have been proposed principally to solve the
automatic knowledge base construction task. Our idea is to adopt these methods
into a sequence-to-sequence model to solve the abstractive text summarization
problem. We believe that a jointly learning of the text and knowledge can improve
the model by generating more abstractive and meaningful summaries.</p>
      <p>Another promising direction that we want to investigate is the generation
of abstractive summaries from documents or multiple documents using deep
learning models. Broadly, the idea can involve a first extractive phase, where the
relevant sentences are extracted from source text, followed by the abstractive
phase, where a summary is generated only from these relevant sentences.</p>
    </sec>
    <sec id="sec-4">
      <title>Evaluation Plans</title>
      <p>We plan to evaluate our models on gold-standard datasets for the summarization
task, such as DUC- [ ], Gigaword [ ] and CNN/DailyMail [ ] corpus,
as well as on a local government dataset of documents made available by
InnovaPuglia S.p.A. (consisting of projects and funding proposals) using several
variants of ROUGE [ ] metric.</p>
      <p>ROUGE is a recall-based metric which assesses how many n-grams in generated
summaries appear in the human reference summaries. This metric is designed to
evaluate extractive methods rather than abstractive ones, thus the former would
be advantaged. The evaluation in summarization is a complex problem and it is
still an open challenge for three main reasons. First, given an input text, there
are different summaries that preserve the original meaning. Furthermore, the
words that compose the summary could not appear at all in the original source.
Finally, ROUGE metric cannot measure the quality of grammar structure of the
generated summary. To overcome these issues we plan an in-vivo experiment
with a user study.
. S. Ahn, H. Choi, T. Pärnamaa, and Y. Bengio. A Neural Knowledge Language</p>
      <p>Model. CoRR, abs/ . , .
. D. Bahdanau, K. Cho, and Y. Bengio. Neural machine translation by jointly
learning to align and translate. CoRR, abs/ . , .
. Y. Bengio, R. Ducharme, P. Vincent, and C. Janvin. A neural probabilistic language
model. J. Mach. Learn. Res., : – , .
. S. Chopra, M. Auli, A. M. Rush, and S. Harvard. Abstractive sentence
summarization with attentive recurrent neural networks. .
. J. Chung, Ç. Gülçehre, K. Cho, and Y. Bengio. Empirical evaluation of gated
recurrent neural networks on sequence modeling. CoRR, abs/ . , .
. M. Dymetman and C. Xiao. Log-linear rnns: Towards recurrent neural networks
with flexible prior knowledge. CoRR, abs/ . , .
. J. L. Elman. Finding structure in time. Cognitive Science, ( ): – , .
. S. Hochreiter and J. Schmidhuber. Long short-term memory. Neural Comput.,
( ): – , .
. K. S. Jones. Automatic summarising: The state of the art. Information Processing
&amp; Management, ( ): – , .
. Y. LeCun, Y. Bengio, and G. Hinton. Deep learning. Nature, : – , .
. C.-Y. Lin. ROUGE: A Package for Automatic Evaluation of Summaries. In Proc.</p>
      <p>of the ACL- Workshop, page . Association for Computational Linguistics, .
. K. C. Litkowski. Summarization experiments in duc. In Proc. of DUC , .
. R. Nallapati, B. Xiang, and B. Zhou. Sequence-to-sequence RNNs for text
summarization. CoRR, abs/ . , .
. P. Norvig. Inference in text understanding. In AAAI, pages – , .
. A. M. Rush, S. Chopra, and J. Weston. A neural attention model for abstractive
sentence summarization. In Proc. of EMNLP , Lisbon, Portugal, pages – ,
.
. H. Saggion and T. Poibeau. Automatic text summarization: Past, present and
future. In Multi-source, Multilingual Information Extraction and Summarization,
pages – . Springer, .
. N. Salim. A review on abstractive summarization methods. Journal of Theoretical
and Applied Information Technology, ( ): – , .
. I. Sutskever, O. Vinyals, and Q. V. Le. Sequence to sequence learning with neural
networks. In Proc of NIPS, pages – , .
. P. Verga and A. McCallum. Row-less universal schema. CoRR, abs/ . ,
.
. Z. Wang, J. Zhang, J. Feng, and Z. Chen. Knowledge graph and text jointly
embedding. In In Proc. of EMNLP. ACL, .</p>
    </sec>
  </body>
  <back>
    <ref-list />
  </back>
</article>