<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Multilingual Natural Language Generation within Abstractive Summarization</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Simon Mille</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Miguel Ballesteros</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Alicia Burga</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Gerard Casamayor</institution>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>ICREA</institution>
          ,
          <addr-line>Passeig Llu ́ıs Companys, 23, 08010 Barcelona</addr-line>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>TALN, Pompeu Fabra University</institution>
          ,
          <addr-line>C/ Roc Boronat, 138, 08018 Barcelona</addr-line>
        </aff>
      </contrib-group>
      <abstract>
        <p>With the tremendous amount of textual data available in the Internet, techniques for abstractive text summarization become increasingly appreciated. In this paper, we present work in progress that tackles the problem of multilingual text summarization using semantic representations. Our system is based on abstract linguistic structures obtained from an analysis pipeline of disambiguation, syntactic and semantic parsing tools. The resulting structures are stored in a semantic repository, from which a text planning component produces content plans that go through a multilingual generation pipeline that produces texts in English, Spanish, French, or German. In this paper we focus on the lingusitic components of the summarizer, both analysis and generation.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>With the tremendous amount of multilingual textual data available in
the Internet, techniques for intelligent abstractive text summarization
in the language of the preference of the user enjoy a steadily
increasing demand for different applications, among them journalism and
media monitoring. Thus, journalists and media monitors have to
review a large number of press articles on a daily basis, a considerable
number of which may not be available in their native language. We
present work in progress that tackles the problem of multilingual text
summarization using semantic representations.</p>
      <p>
        The most popular summarization strategy is still
“extraction”oriented. Text fragments (in general, entire sentences, but in some
cases also phrases), are selected from one or more source documents,
based on some relevance metric, and the most relevant fragments are
put together in a summary (see, e.g., [
        <xref ref-type="bibr" rid="ref19">19</xref>
        ] for an overview). Although
extractive summarization can be addressed with little linguistic
analysis and, in the case of sentence-based selection metrics, the
resulting summaries are always grammatically correct, it is known to have
some significant shortcomings. For instance, the selection of the
content to be included into the summary is rather coarse-grained and
surface-(instead of knowledge-)oriented, and the summaries tend to
lack internal coherence between the selected text fragments.
Furthermore, in general, the summaries are monolingual, i.e., the original
text and the summary are in the same language.
      </p>
      <p>
        Opposed to extractive summarization is “abstractive
summarization”. Abstractive (i.e., concept-based) summarization analyzes the
original textual material using language parsing and/or Information
Extraction into intermediate linguistic or conceptual representations.
Content selection relevance-driven techniques are then applied to
these representations to choose the content elements that are to be
communicated in the summary. From the chosen content elements,
a summary is generated using deep Natural Language Generation
(NLG) techniques. A number of approaches to abstractive
summarization have been proposed. Some attempt to adapt extractive
techniques to abstractive summarization [
        <xref ref-type="bibr" rid="ref23">23</xref>
        ]. Others do not use abstract
representations and remain at a superficial level [
        <xref ref-type="bibr" rid="ref15 ref26">15, 26</xref>
        ], or use
partially abstract structures, be it because not all the content of the input
text is represented [
        <xref ref-type="bibr" rid="ref17 ref20">20, 17</xref>
        ], or because some idiosyncratic features
are maintained [
        <xref ref-type="bibr" rid="ref37">37</xref>
        ]. It is also not always the case that deep
generation is used. For instance, Genest and Lapalme [
        <xref ref-type="bibr" rid="ref17">17</xref>
        ] and Saggion and
Lapalme [
        <xref ref-type="bibr" rid="ref35">35</xref>
        ] start from templates, Ganesan et al. [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ] from word
lattices, and Cheung and Penn [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ] and Genest and Lapalme [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ] from
syntactic structures. Liu et al. [
        <xref ref-type="bibr" rid="ref22">22</xref>
        ] do not have a proper generation
component at all. Liu et al. [
        <xref ref-type="bibr" rid="ref22">22</xref>
        ] and Cheung and Penn [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ] apply
sentence fusion, rather than content selection.
      </p>
      <p>We developed techniques for abstractive summarization that are
capable of generating multilingual summaries in response to a user
query on a specific content element using full state-of-the-art (deep)
language analysis and language generation mechanisms, combining
statistic and rule-based techniques. In this paper, we focus on the
general architecture of the summarizer and its generation module.</p>
    </sec>
    <sec id="sec-2">
      <title>An architecture for abstractive summarization 2 2.1</title>
    </sec>
    <sec id="sec-3">
      <title>Theoretical framework</title>
      <p>
        The theoretical framework that underlies our system is the
MeaningText Theory [
        <xref ref-type="bibr" rid="ref27">27</xref>
        ]. MTT is based on the notion of dependency, which
establishes a relation of “governance” between two elements.
      </p>
      <p>The MTT model supports high expressiveness at the three main
levels of the linguistic description of written language: semantics,
syntax and morphology, while facilitating a coherent transition
between them via intermediate levels of deep syntax and deep
morphology. In total, the model foresees five strata; at each stratum, a
clearly defined type of linguistic phenomena is described in terms of
distinct dependency structures.</p>
      <p>Semantic Structures (SemSs) are predicate-argument structures
in which the relations between predicates and their arguments are
numbered in accordance with the order of the arguments.
Deep-syntactic structures (DSyntSs) are dependency trees, with
the nodes labeled by meaningful (“deep”) lexical units (LUs) and
the edges by actant relations I, II, III, ..., VI (in accordance with
the syntactic valency pattern of the governing LU) or one of
the following three non-argumental relations: ATTR(ibute),
COORD(ination), APPEND(itive).</p>
      <p>Surface-Syntactic Structures (SSyntSs) are dependency trees in
which the nodes are labeled by open or closed class lexemes and
the edges by grammatical function relations of the type subject,
oblique object, adverbial, modifier, etc.</p>
      <sec id="sec-3-1">
        <title>Deep-Morphological Structures (DMorphSs) are chains of lex</title>
        <p>emes in their base form (with inflectional and PoS features
being associated to them in terms of attribute-feature pairs) between
which a precedence relation is defined and which are grouped in
terms of constituents.</p>
      </sec>
      <sec id="sec-3-2">
        <title>Surface-Morphological Structures (SMorphSs) are chains of in</title>
        <p>flected word forms, i.e., sentences as they appear in the corpus,
except that orthographic contractions still did not take place.</p>
        <p>The analysis and generation modules in our abstractive pipeline
draw upon these strata. In particular, the tasks of language analysis
and language generation can be seen as a sequence of mappings
between adjacent strata; for analysis, starting from text and arriving at a
semantic (or conceptual) representation , and for generation, starting
from a semantic representation up to the text surface.
2.2</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>A pipeline for abstractive summarization</title>
      <p>
        The implementation of our abstractive summarizer is based on a
sequence of modules that realize the sequence of transitions between
the different strata of the MTT model. The pipeline shown in Figure
1 can be divided into three main parts:
1. Language analysis: Language analysis is carried out by a text
analysis pipeline that takes as input the textual content of a
document in a given language. This content is first analyzed and
represented as a forest of DSyntSs. In the case that the input language
is different from English, every lexeme in the DSyntSs is mapped
onto an English lexeme using bilingual dictionaries in order to
arrive at a kind of interlingua structure that facilitates
languageneutral representations (see Subsection 3.3 for a justification).
These English “interlingua” structures are then mapped onto
semantic structures, enriched with Frames from the FrameNet
lexicon [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ], modeled as RDF triples, and stored in a semantic
repository.
2. Text planning: Conceptual summarization is approached by
assessing the relevance of the semantic structures produced by the
language analysis step in relation to a specific entity which
constitutes a topic of interest for the end user and to which the generated
summary is tailored. In addition to determining the relevance of
contents, our text planning component also attempts to guarantee
a degree of coherence in the summary to be generated by
sorting relevant contents in a sequence that satisfies certain choerence
constraints, e.g. grouping together in the text contents making
reference to the same entities. Relevance calculations are based on
relative cooccurrence metrics of word senses and references to
entities detected in the original documents during language
analysis, the cooccurrence metrics being obtained from pre-existing
corpora of annotated documents.
3. Natural language generation: Following this planning step,
linguistic generation starts by transferring the lexemes associated to
the semantic structures to the desired target language, using
available multilingual lexical resources. Then, the structure of the
sentence is determined and all grammatical words are introduced and
linked with syntactic relations. Finally, all morphological
agreements between the words are resolved, the words are ordered and
punctuation signs are introduced.
3
3.1
      </p>
    </sec>
    <sec id="sec-5">
      <title>Language analysis</title>
    </sec>
    <sec id="sec-6">
      <title>Tokenization and disambiguation</title>
      <p>
        Language analysis starts by determining sentence and token
boundaries using Bohnet et al.’s [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ] tools. Rather than addressing
tokenization at word level, however, our analysis pipeline treats each
sequence of words referring to a specific entity as an atomic unit of
meaning. In doing so, we seek to avoid unnecessary internal analysis
of multiword expressions which may not even have a strictly
compositional meaning (as, e.g., United States of America), and also to
eventually obtain predicate-argument structures in which the
arguments are not just words, but expressions with an atomic meaning.
      </p>
      <p>
        To determine the disambiguated senses of individual words and
the entities referred to by single words or phrases, we use Babelfy3
[
        <xref ref-type="bibr" rid="ref29">29</xref>
        ]. Babelfy addresses both Word Sense Disambiguation and
Entity Linking against BabelNet [
        <xref ref-type="bibr" rid="ref30">30</xref>
        ], a large multilingual semantic
network organized around Babel synsets resulting from mapping
WordNet synsets and Wikipedia pages. The large coverage of BabelNet
allows Babelfy to annotate both Named Entities and conceptual
meanings. All multiword expressions annotated by Babelfy are considered
by the following modules as a single token.
3.2
      </p>
    </sec>
    <sec id="sec-7">
      <title>Deep-syntactic parsing</title>
      <p>Once the texts are clean, tokenized and the words are disambiguated
against BabelNet, they are sent to a parsing module that carries out
in sequence Surface-Syntactic and Deep-Syntactic Parsing.</p>
      <p>
        For Surface-Syntactic Parsing, we use Bohnet et al.’s [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ] joint
lemmatizer, part of speech tagger, morphology tagger and
dependency parser, which follows a transition-based approach with
beamsearch. Trained on a surface syntactic treebank, the joint parser
produces the surface syntactic tree for an unseen sentence.4
      </p>
      <p>For Deep-Syntactic Parsing, we use a SSynt-DSynt transducer.
The objective of the transducer is to identify and remove all
functional words (auxiliaries, determiners, void prepositions and
conjunctions) in the surface-syntactic tree and to generalize the
syntactic dependencies obtained during the previous stage, while adding
subcategorization information for lexical predicates. Two different
transducers have been developed. One is based on a statistical model
3 www.babelfy.org
4 The details on Bohnet et al.’s system can be obtained from the original work.</p>
      <p>
        It suffices to note here that it produces very competitive scores for all the
tasks it performs, for a wide range of languages.
and the other is rule-based. The statistical transducer (see [
        <xref ref-type="bibr" rid="ref1 ref2">1, 2</xref>
        ] for
details) is trained on parallel SSynt and DSynt corpora (see for
instance [
        <xref ref-type="bibr" rid="ref24">24</xref>
        ] for an example in Spanish). The DSyntS-SSyntS
transducer has the potential to be trained for any language in which there
are parallel DSynt and SSynt available treebanks. Currently, it is the
case for English and Spanish. The rule-based transducer is
implemented a graph-transduction grammars that have access to
languagespecific lexicons to remove the void prepositions and conjunctions
[
        <xref ref-type="bibr" rid="ref28">28</xref>
        ], when any is available. The rule-based version is available for
English, Spanish, German, and French.
3.3
      </p>
    </sec>
    <sec id="sec-8">
      <title>Mapping to abstract representations and frame assignment</title>
      <p>
        For mapping deep-syntactic structures to more abstract linguistic
representations, large-scale lexical resources are needed.
Unfortunately, such resources are available, at this point, only for English;
see, e.g., PropBank [
        <xref ref-type="bibr" rid="ref21">21</xref>
        ], FrameNet [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ], VerbNet [
        <xref ref-type="bibr" rid="ref36">36</xref>
        ], and the
mappings between them (SemLink [
        <xref ref-type="bibr" rid="ref32">32</xref>
        ]). For this reason, we chose to
map all input languages to English.
      </p>
      <p>After the SSynt-DSynt transduction, the obtained structure does
not contain any functional words, which tend to be idiosyncratic. The
nodes are labeled with meaningful lexemes.5 Using multilingual
resources such as BabelNet (see Sections 3.1 and 5.1), it is possible
to obtain the translations of these words into English. Once this is
done, the combination of the subcategorization information in the
deep-syntactic structure and SemLink allows us to obtain Frame
annotations on top of connected predicate-argument structures. The
latter follow the principles of the Meaning-Text Theory model, with the
addition of a subset of relations such as Location, Time, etc., which
facilitate the further processing. During this step, shared argumental
positions are made explicit and idiosyncratic structuring such as the
representation of raising and control verbs is generalized.
4</p>
    </sec>
    <sec id="sec-9">
      <title>Text Planning: Planning the Summary</title>
      <p>
        Before a summary can be generated, it is necessary to determine, on
the one hand, the content that is to be communicated to the user and,
on the other hand, the discourse structure of the determined content.
These two tasks are commonly referred to in NLG literature as text
or document planning [
        <xref ref-type="bibr" rid="ref34">34</xref>
        ].
      </p>
      <p>The production of summaries in our system assumes that
summaries are generated in response to a user query in which an entity in
the semantic repository is specified. The entity must correspond to a
BabelNet synset, identified by any of the multilingual lexicalizations
associated to it in BabelNet. If multiple synsets match the string
introduced by the user, the user is asked to choose from the available
meanings. The summary to be produced should contain content in
the semantic repository that is relevant to the queried entity. The task
of the text planning module is to determine the set of most relevant
content elements and generate an ordered list out of them.
4.1</p>
    </sec>
    <sec id="sec-10">
      <title>Ranking semantic structures</title>
      <p>
        Following previous graph-based approaches to text planning [
        <xref ref-type="bibr" rid="ref10 ref31">31, 10</xref>
        ],
our approach adopts a graph view on the content of the semantic
repository where nodes correspond to predicate-argument structures
produced by the analysis pipeline, and edges indicate entity-sharing
5 This assumption is not entirely true since our DSyntSs still contain, e.g.,
support verbs (as deliver in John delivered his first speech in the Congress),
which are generally assumed to be void of meaning as well. In genuine
MTT DSyntSs, support verbs do not appear as such either. However, we
think that this simplification can be tolerated without a too significant loss
of quality.
relations between nodes. For each query, a query graph is created that
contains as nodes all predicates that have the user-specified entity
as one of its arguments. This initial set is extended recursively with
other nodes that share at least one argument with relations already in
the graph, up to a fixed depth. The resulting graph serves to constrain
the planning task to a set of related contents, in a similar fashion to
past works such as [
        <xref ref-type="bibr" rid="ref25 ref7 ref9">25, 9, 7</xref>
        ].
      </p>
      <p>
        Given a query graph, we formulate the planning of summaries as
a ranking problem, similar in spirit to other text planning
implementations based on ranking contents [
        <xref ref-type="bibr" rid="ref14 ref7 ref9">9, 7, 14</xref>
        ]. In our ranking
formulation all nodes in a query graph must be ranked according to some
function that indicates their relevance. Consequently, the text plans
produced by our method are sorted lists of nodes in a query graph.
The ranking starts from an initial distribution of relevance obtained
from co-occurrence counts in a corpus of texts analyzed with
Babelfy. For each entity annotated in the corpus, we thus estimate its
probability of being annotated in the same document as any other
entity. Predicate-argument structures are then assigned an initial rank
according to the probabilities of their arguments co-occurring with
the queried entity.
      </p>
      <p>
        Some predicates may have none of their arguments annotated in
the Babelfy corpus. In this case, we have no empirical basis to assess
their relevance. In order to ameliorate this situation, we distribute the
relevance from those nodes that do have some probability assigned to
them to their neighboring nodes in the query graph. This is achieved
by iteratively multiplying the initial distribution of relevance to nodes
of the query graph with an adjacency matrix of the query graph which
has been modified so that it can be interpreted as a Markov Chain.
That is, given a node, its transition probabilities are calculated from
the relevance scores of its target nodes and normalized according to
the sum of relevance of all nodes reachable from the initial node.
This procedure, which is similar to web ranking [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ], produces a
near-stationary distribution of relevance in which the initial relevance
scores have been adjusted according to the graph topology.
4.2
      </p>
    </sec>
    <sec id="sec-11">
      <title>Producing a coherent text plan</title>
      <p>
        As pointed out above, the goal of the text planning module is not
only to determine the relevance of content elements with respect to
the user query, but also to define a discourse structure for these
elements, i.e., an ordering of the elements that enforces a certain
degree of coherence in the resulting text. We do so by ensuring that
new predicate-argument structures are added to a text plan only if
they are semantically related to the content elements already in the
plan. More precisely, we guarantee entity-coherence by performing a
graph exploration of the query graph, which consists in visiting only
those nodes that are connected to nodes that have been already
visited. This notion of entity-based coherence is inspired by theories of
local coherence such as the Centering Theory [
        <xref ref-type="bibr" rid="ref18 ref33">18, 33</xref>
        ].
      </p>
      <p>Since the edges in the content graph capture argument-sharing
relations between predicates, the sequence of visited nodes is such that,
for every node, at least one of its arguments is either the requested
entity, or an argument that has already been introduced into the plan.
The traversal of the query graph is done in a greedy way. Starting
from the set of predicates that have the queried entity as their
argument, the most relevant node from all those that available is always
selected. Once a node is selected, all the nodes in the query graph
connected to it become available for selection. The traversal of the
graph produces the ordered sequence that constitutes the input of the
surface generation pipeline.</p>
    </sec>
    <sec id="sec-12">
      <title>Multilingual text generation</title>
      <p>The predicate-argument structures produced by the text planning
module are obtained from translating words from the source
documents into English (see Section 3.3). In order to generate
multilingual text, however, it is necessary to map them to linguistic structures
that serve as a starting point for multilingual linguistic generation,
which, in turn, requires language-specific lexical resources that
capture the lexical and syntactic characteristics of each language. In the
next subsection, the creation of such multilingual lexical resources is
explained in detail.
5.1</p>
    </sec>
    <sec id="sec-13">
      <title>Multilingual lexical resources</title>
      <p>For multilingual generation, we need to create lexicons for each
language we cover. These lexicons must not only contain
languagespecific vocabulary, but also be linked to our pivot language, namely
English. Given that BabelNet senses annotated during the analysis
stage are language-independent, we use them as the cross-linguistic
link. Below, we detail the creation procedure and structure of the
language-specific lexicons used to go from predicate-argument
structures and BabelNet synsets to each language.</p>
      <p>The languages supported by our multilingual generation pipeline
(English, French, Spanish and German) have a satisfactory amount
of NLP resources. The experimental compilation of the
corresponding language-specific lexicons was done in different stages. First of
all, three texts in each language were randomly selected. Thus, a
set of eight texts (around 2,400 tokens) was used as base for the
language-specific lexicons.6 Given that word sense ambiguity is a
problem inherent to any language, it was necessary to disambiguate
and recognize the right sense of a lexical unit before assigning any
specific BabelNet id to it. Babelfy, which as explained in Section
3.1, is connected to BabelNet, was used for disambiguation, using
the API offered to remotely access the service. As output of this step,
a list of unique BabelNet ids (1,013 items in total) was obtained,
which served as the basis for creating the lexicons. This list has then
been locally enriched with the word form linked to each id in each
language. Using this list as base, for each LU, its part of speech,
its lemma, its BabelNet id and its government pattern, i.e., its
subcategorization frame, are stored. Within the government pattern, the
information collected for each argument includes its part of speech,
the preposition introducing it (if it is required by the described LU)
and the corresponding case. Below, the entries for the same specific
BabelNet id in German (a language with case) and in Spanish are
shown.</p>
      <p>SPANISH
“contar VV 01”: verb f
lemma = “contar”
bn = bn:00091011v</p>
      <p>GERMAN
“sagen VV 01”: verb f
lemma = “sagen”
bn = bn:00091011v
gp = f</p>
      <p>I = fdpos = “N”g
II = fdpos = “N”g
III = fdpos = “N” prep = “a”ggg
gp = f
I = fdpos = “N” case = “nom”g
II = fdpos = “N” case = “acc”g</p>
      <p>III = fdpos = “N” case = “dat”ggg</p>
      <p>From the English structure, the system thus turns to the lexicons
to obtain information about the specific characteristics of the
sentences to be generated in each language. If no specific information
is added, the system interprets that there are no restrictions with
respect to the argument in question. Thus, the four compiled parallel
6 Although it can be argued that the work is based on a small sample of
vocabulary, the sample is big enough to test the adopted methodology.
language-specific lexicons serve in a direct way for the multilingual
generation pipeline, allowing the mapping from English to any of
the other languages involved. Potentially, the mapping could be even
done not only from English to other language, but from any other
language included in the system to each other.
5.2</p>
    </sec>
    <sec id="sec-14">
      <title>Hybrid NLG system</title>
      <p>The lexical resources described in Section 5.1 are meant to be used
together with generation grammars, which are rule systems that
produce successively the different layers of representation mentioned in
Section 2. In this section, we describe the different submodules of
the NLG pipeline, together with their alternative Machine Learning
implementations. In order to understand better the process, Figure 2
includes some intermediate structures of this pipeline.</p>
      <sec id="sec-14-1">
        <title>1. Mapping to output language predicate-argument structures</title>
        <p>Starting from the structures provided by the text planning module
(see Section 4), first, some idiosyncratic transformations are made to
adjust the structures to the predicate-argument format understood by
our generation pipeline, and then, the English labels of the nodes are
translated into the desired target language using the lexicons detailed
in Section 5.1.</p>
      </sec>
      <sec id="sec-14-2">
        <title>2. Mapping to syntactic structures</title>
        <p>Once genuine predicate-argument structures in the target language
are available, the first task is to find which node in each structure is
most likely to be the root of the dependency tree. That is, we want to
identify what will be the main verb of the sentence, or the word that
triggers its appearance. The main node is typically a word (i) that is
predicate, (ii) that has more participants than any other predicate of
the structure, and (iii) that is not involved in a semantic relation of
secondary relevance. Adjectives, adverbs, prepositions and nouns are
possible alternatives to verbs when no verb is available. Around the
main node, the deep-syntacticization module builds the rest of the
syntactic structure of the sentence. In particular, it is able to decide if
a main predicate has to be introduced, or what will be realized as an
argument, an attribute, or a coordination.</p>
        <p>The procedure of the retrieval of deep-syntactic target structures
has been successfully tested on around 39,000 sentences: more than
99% of the semantic structures are mapped to well-formed
deepsyntactic structures. In the rest of the cases, the generator is unable
to produce any syntactic tree and a fallback message is returned.</p>
        <p>The next step in the procedure is to obtain surface-syntactic
structures, i.e., to generate all functional words and labeling the
dependencies with SSynt relations. In the same fashion as for SSynt–
DSynt transduction in the case of analysis, we use two alternative
approaches for DSynt–SSynt transduction in the case of generation.
For languages with limited amount of annotated data (as, e.g., French
or German), a rule-based system is preferred, but if multilayered
corpora of reasonable size are available (as Spanish and English),
training statistical tools is also possible.</p>
        <p>
          For rule-based transduction, we use an adapted version of the
MARQUIS generator [
          <xref ref-type="bibr" rid="ref38">38</xref>
          ]. MARQUIS had been designed for
datato-text generation. It starts from air quality and meteorology time
series, and uses language-specific resources that contain a fine-grained
description of all the concepts and words in the air quality domain.
Generation in the context of abstractive summarization is a case of
text-to-text generation. That is, we cannot focus on the concepts
of a specific domain. Rather, any concept can be present in a
semantic structure, and there are no lexical resources that are
complete enough to contain all of them. As a consequence, MARQUIS’s
graph-transduction grammars had to be adapted.
For machine learning-based transduction, we developed a series of
Support Vector Machine-based transducers; cf., [
          <xref ref-type="bibr" rid="ref3">3</xref>
          ] for details.
        </p>
      </sec>
      <sec id="sec-14-3">
        <title>3. Morphological agreement resolution and surface form retrieval</title>
        <p>During the generation of syntactic structures, morphological features
of individual words are already inserted (e.g., nominative case for a
German subject). During the transition to the morphological
structure, agreement is established (using the introduced morphological
features and the fine-grained syntactic relations in the SSyntSs) and
surface forms of the words are retrieved using a full-form dictionary.</p>
        <p>In order to obtain the full-form dictionary, we run the
morphological tagger of our surface syntactic parser on a large collection of
texts and store each possible combination of surface form, lemma
and morphological features. We can therefore retrieve a surface form
given a lemma and a set of morphological features. The size of the
text collection is crucial in order to ensure a large coverage. For
instance, for English, we use the entire Gigaword corpus.7</p>
      </sec>
      <sec id="sec-14-4">
        <title>4. Linearization of Unordered Syntactic Dependency Trees</title>
        <p>
          The linearization of unordered syntactic dependency trees, i.e., word
order determination, is performed with the state-of-the-art Bohnet
et al. [
          <xref ref-type="bibr" rid="ref5">5</xref>
          ] linearizer. This linearizer is trained on a surface-syntactic
treebank. It produces a statistical model that is capable of
determining the word order in a sentence by using mainly surface-syntactic
relations and part-of-speech tags.
6
        </p>
      </sec>
    </sec>
    <sec id="sec-15">
      <title>Conclusions and Future Work</title>
      <p>In this paper, we presented work in progress for the production of
abstractive summaries about specific entities from contents obtained
from the analysis of multiple texts. We have covered the resources,
tools and techniques applied to obtain the summaries, placing special
emphasis on text planning and the multilingual generation
component.</p>
      <p>In the future, we plan to evaluate the pipeline in one or more
domains and use the results to determine what components require
improvement. Our first application domain will be the production of
multilingual summaries from news articles in the scope of the
MULTISENSOR project 8. We expect our natural language processing
tools to perform better in journalistic texts than in more specialized
domains that may require models obtained from domain-specific
corpora. Crucially, the entities and concepts found in press articles are
7 https://catalog.ldc.upenn.edu/LDC2003T05
8 http://www.multisensorproject.eu/
also more likely to be covered by BabelNet, which plays a crucial
role in deciding what contents go into the summary. Specialized
domains, e.g. medical or legal texts, may use terminology and make
reference to entities only found in specialized kwnowledge and
lexical resources. Creating multilingual resources and tools for specific
domains is one of the major limitations of applying an abstractive
approach to summarization.</p>
      <p>The evaluation of our approach will involve both a quantitative
evaluation where system-produced summaries are compared to a
gold standard of manually written (abstractive) summaries, and a
qualitative evaluation in which users will be handed a
questionnaire designed at reviewing various facets of the texts: relevance,
coherence, grammaticality, readability, etc. Considering the pipeline
architecture of our system and the problems introduced by
errorpropagation, an individual evaluation of each module will also be
conducted to identify the most problematic areas. We are
particularly interested in finding ways to cope with noisy output from the
text analysis component during text planning and linguistic
generation, in order to avoid generating ungrammatical or meaningless
sentences.</p>
      <p>With respect to the lexicons used in the surface generation
module, although BabelNet seems very useful in order to obtain
interconnected language-specific resources, some issues have been identified
which will have to be dealt with in the future. First of all, languages
with a very productive compositional process (e.g., German) have
BabelNet synsets for which there is no direct correspondence in other
languages (in other words, they correspond to more than one synset).
Second, and partly as a consequence of the first issue, not all
BabelNet synsets correspond to a term in a specific language. Third and
last, the procedure for compiling BabelNet synsets can be optimized:
if a sequence of lexical units is considered as a multiword unit, then
synsets are duplicated (one synset is assigned for each single unit and
another one for the multiword unit).</p>
      <p>
        As far as analysis is concerned, we plan to incorporate alternative
surface-syntactic parsers based on recurrent neural networks [
        <xref ref-type="bibr" rid="ref12 ref4">12, 4</xref>
        ],
which have been found to be particularly beneficiary for
out-ofvocabulary words.
      </p>
    </sec>
    <sec id="sec-16">
      <title>Acknowledgements</title>
      <p>This work has been supported by the European Commission under
the contract number FP7-ICT-610411.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>M.</given-names>
            <surname>Ballesteros</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Bohnet</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Mille</surname>
          </string-name>
          , and L. Wanner, '
          <article-title>Deep-syntactic parsing'</article-title>
          ,
          <source>in Proceedings of the 25th International Conference on Computational Linguistics (COLING)</source>
          , Dublin, Ireland, (
          <year>2014</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>M.</given-names>
            <surname>Ballesteros</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Bohnet</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Mille</surname>
          </string-name>
          , and L. Wanner, '
          <article-title>Data-driven deepsyntactic dependency parsing'</article-title>
          ,
          <source>Natural Language Engineering</source>
          ,
          <fpage>1</fpage>
          -
          <lpage>36</lpage>
          , (
          <year>2015</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>M.</given-names>
            <surname>Ballesteros</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Bohnet</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Mille</surname>
          </string-name>
          , and L. Wanner, '
          <article-title>Data-driven sentence generation with non-isomorphic trees'</article-title>
          ,
          <source>in Proceedings of the 2015 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies</source>
          , pp.
          <fpage>387</fpage>
          -
          <lpage>397</lpage>
          , Denver, Colorado, (May-June
          <year>2015</year>
          ).
          <article-title>Association for Computational Linguistics</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>M.</given-names>
            <surname>Ballesteros</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Dyer</surname>
          </string-name>
          , and
          <string-name>
            <given-names>N.A.</given-names>
            <surname>Smith</surname>
          </string-name>
          , '
          <article-title>Improved transition-based parsing by modeling characters instead of words with lstms'</article-title>
          ,
          <source>in Proceedings of the 2015 Conference on Empirical Methods in Natural Language Processing</source>
          , pp.
          <fpage>349</fpage>
          -
          <lpage>359</lpage>
          , Lisbon, Portugal, (
          <year>September 2015</year>
          ).
          <article-title>Association for Computational Linguistics</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>B.</given-names>
            <surname>Bohnet</surname>
          </string-name>
          ,
          <string-name>
            <surname>A</surname>
          </string-name>
          . Bjo¨rkelund, J. Kuhn,
          <string-name>
            <given-names>W.</given-names>
            <surname>Seeker</surname>
          </string-name>
          , and
          <string-name>
            <given-names>S.</given-names>
            <surname>Zarriess</surname>
          </string-name>
          , '
          <article-title>Generating non-projective word order in statistical linearization'</article-title>
          ,
          <source>in Proceedings of the 2012 Joint Conference on EMNLP and CoNLL</source>
          , pp.
          <fpage>928</fpage>
          -
          <lpage>939</lpage>
          ,
          <string-name>
            <surname>Jeju</surname>
            <given-names>Island</given-names>
          </string-name>
          , Korea, (
          <year>July 2012</year>
          ).
          <article-title>Association for Computational Linguistics</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>B.</given-names>
            <surname>Bohnet</surname>
          </string-name>
          and
          <string-name>
            <given-names>J.</given-names>
            <surname>Nivre</surname>
          </string-name>
          , '
          <article-title>A transition-based system for joint Partof-Speech tagging and labeled non-projective dependency parsing'</article-title>
          ,
          <source>in Proceedings of the 2012 Joint Conference on Empirical Methods in Natural Language Processing and Computational Natural Language Learning (EMNLP-CoNLL)</source>
          , pp.
          <fpage>1455</fpage>
          -
          <lpage>1465</lpage>
          ,
          <string-name>
            <surname>Jeju</surname>
            <given-names>Island</given-names>
          </string-name>
          , Korea, (
          <year>2012</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>N.</given-names>
            <surname>Bouayad-Agha</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Casamayor</surname>
          </string-name>
          , and L. Wanner, '
          <article-title>Content Determination from an Ontology-based Knowledge Base for the Generation of Football Summaries'</article-title>
          ,
          <source>in Proceedings of the 13th European Natural Language Generation Workshop (ENLG)</source>
          , pp.
          <fpage>27</fpage>
          -
          <lpage>81</lpage>
          , Stroudsburg, PA, USA, (
          <year>2011</year>
          ).
          <article-title>Association for Computational LInguistics</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>J.C.K.</given-names>
            <surname>Cheung</surname>
          </string-name>
          and G. Penn, '
          <article-title>Unsupervised sentence enhancement for automatic summarization'</article-title>
          ,
          <source>in Proceedings of the 2014 Conference on EMNLP</source>
          , pp.
          <fpage>775</fpage>
          -
          <lpage>786</lpage>
          , Doha, Qatar, (
          <year>October 2014</year>
          ).
          <article-title>Association for Computational Linguistics</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>D.</given-names>
            <surname>Danne</surname>
          </string-name>
          <article-title>´lls, 'The value of weights in automatically generated text structures'</article-title>
          ,
          <source>Lecture Notes in Computer Science</source>
          ,
          <volume>5449</volume>
          LNCS,
          <fpage>233</fpage>
          -
          <lpage>244</lpage>
          , (
          <year>2009</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>S.</given-names>
            <surname>Demir</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Carberry</surname>
          </string-name>
          , and
          <string-name>
            <surname>K.F. McCoy</surname>
          </string-name>
          , '
          <article-title>A Discourse-aware Graphbased Content Selection Framework'</article-title>
          ,
          <source>in Proceedings of the 6th International Natural Language Generation Conference</source>
          , pp.
          <fpage>17</fpage>
          -
          <lpage>27</lpage>
          , Stroudsburg, PA, USA, (
          <year>2010</year>
          ).
          <article-title>Association for Computational LInguistics</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <surname>Michelangelo</surname>
            <given-names>Diligenti</given-names>
          </string-name>
          , Marco Gori, and Marco Maggini, '
          <article-title>A unified probabilistic framework for web page scoring systems'</article-title>
          ,
          <source>IEEE Transactions on knowledge and data engineering</source>
          ,
          <volume>16</volume>
          (
          <issue>1</issue>
          ),
          <fpage>4</fpage>
          -
          <lpage>16</lpage>
          , (
          <year>2004</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>C.</given-names>
            <surname>Dyer</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Ballesteros</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W.</given-names>
            <surname>Ling</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Matthews</surname>
          </string-name>
          , and
          <string-name>
            <given-names>N.A.</given-names>
            <surname>Smith</surname>
          </string-name>
          , '
          <article-title>Transition-based dependency parsing with stack long short-term memory'</article-title>
          ,
          <source>in Proceedings of the 53rd Annual ACL Meeting and the 7th International Joint Conference on Natural Language Processing</source>
          , pp.
          <fpage>334</fpage>
          -
          <lpage>343</lpage>
          , Beijing, China, (
          <year>July 2015</year>
          ).
          <article-title>Association for Computational Linguistics</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <surname>Charles</surname>
            <given-names>J.</given-names>
          </string-name>
          <string-name>
            <surname>Fillmore</surname>
            ,
            <given-names>Collin F.</given-names>
          </string-name>
          <string-name>
            <surname>Baker</surname>
          </string-name>
          , and Hiroaki Sato, '
          <article-title>The FrameNet database and software tools'</article-title>
          ,
          <source>in Proceedings of the 3rd International Conference on Language Resources and Evaluation (LREC)</source>
          , pp.
          <fpage>1157</fpage>
          -
          <lpage>1160</lpage>
          ,
          <string-name>
            <surname>Las</surname>
            <given-names>Palmas</given-names>
          </string-name>
          , Canary Islands, Spain, (
          <year>2002</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <given-names>A.</given-names>
            <surname>Freitas</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.G.</given-names>
            <surname>Oliveira</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Curry</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S. O</given-names>
            <surname>'Riain</surname>
          </string-name>
          ,
          <string-name>
            <surname>and J.C.</surname>
          </string-name>
          <article-title>Pereira da Silva, 'Treo: Combining entity-search, spreading activation and semantic relatedness for querying linked data'</article-title>
          ,
          <source>in In: 1st Workshop on Question Answering over Linked Data (QALD-1) Workshop at 8th Extended Semantic Web Conference (ESWC</source>
          , (
          <year>2011</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <given-names>K.</given-names>
            <surname>Ganesan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Zhai</surname>
          </string-name>
          , and J. Han, '
          <article-title>Opinosis: A graph based approach to abstractive summarization of highly redundant opinions'</article-title>
          ,
          <source>in Proceedings of the 23rd International Conference on Computational Linguistics (Coling</source>
          <year>2010</year>
          ), pp.
          <fpage>340</fpage>
          -
          <lpage>348</lpage>
          , Beijing, China, (
          <year>August 2010</year>
          ).
          <article-title>Coling 2010 Organizing Committee</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [16]
          <string-name>
            <given-names>P.-E.</given-names>
            <surname>Genest</surname>
          </string-name>
          and G. Lapalme, '
          <article-title>Framework for abstractive summarization using text-to-text generation'</article-title>
          ,
          <source>in Proceedings of the Workshop on Monolingual Text-To-Text Generation</source>
          , pp.
          <fpage>64</fpage>
          -
          <lpage>73</lpage>
          . Association for Computational Linguistics, (
          <year>2011</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [17]
          <string-name>
            <given-names>P.-E.</given-names>
            <surname>Genest</surname>
          </string-name>
          and G. Lapalme, '
          <article-title>Fully abstractive approach to guided summarization'</article-title>
          ,
          <source>in Proceedings of the 50th Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers)</source>
          , pp.
          <fpage>354</fpage>
          -
          <lpage>358</lpage>
          ,
          <string-name>
            <surname>Jeju</surname>
            <given-names>Island</given-names>
          </string-name>
          , Korea, (
          <year>July 2012</year>
          ).
          <article-title>Association for Computational Linguistics</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          [18]
          <string-name>
            <given-names>B. J</given-names>
            <surname>Grosz</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Weinstein</surname>
          </string-name>
          ,
          <article-title>and</article-title>
          <string-name>
            <surname>A. K Joshi</surname>
          </string-name>
          , '
          <article-title>Centering: A framework for modeling the local coherence of discourse'</article-title>
          ,
          <source>Computational linguistics</source>
          ,
          <volume>21</volume>
          (
          <issue>2</issue>
          ),
          <fpage>203</fpage>
          -
          <lpage>225</lpage>
          , (
          <year>1995</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          [19]
          <string-name>
            <given-names>V.</given-names>
            <surname>Gupta</surname>
          </string-name>
          and
          <string-name>
            <given-names>G.S.</given-names>
            <surname>Lehal</surname>
          </string-name>
          , '
          <article-title>A survey of text summarization extractive techniques'</article-title>
          ,
          <source>Journal of Emerging Technologies in Web Intelligence</source>
          ,
          <volume>2</volume>
          (
          <issue>3</issue>
          ),
          <fpage>258</fpage>
          -
          <lpage>268</lpage>
          , (
          <year>2010</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          [20]
          <string-name>
            <surname>Atif</surname>
            <given-names>Khan</given-names>
          </string-name>
          , Naomie Salim, and Yogan Jaya Kumar, '
          <article-title>A framework for multi-document abstractive summarization based on semantic role labelling'</article-title>
          , Applied Soft Computing,
          <volume>30</volume>
          ,
          <fpage>737</fpage>
          -
          <lpage>747</lpage>
          , (
          <year>2015</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          [21]
          <string-name>
            <given-names>P.</given-names>
            <surname>Kingsbury</surname>
          </string-name>
          and
          <string-name>
            <given-names>M.</given-names>
            <surname>Palmer</surname>
          </string-name>
          , 'From TreeBank to PropBank',
          <source>in Proceedings of the 3rd International Conference on Language Resources and Evaluation (LREC)</source>
          , pp.
          <fpage>1989</fpage>
          -
          <lpage>1993</lpage>
          ,
          <string-name>
            <given-names>Las</given-names>
            <surname>Palmas</surname>
          </string-name>
          , Canary Islands, Spain, (
          <year>2002</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          [22]
          <string-name>
            <given-names>F.</given-names>
            <surname>Liu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Flanigan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Thomson</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Sadeh</surname>
          </string-name>
          , and
          <string-name>
            <given-names>N.A.</given-names>
            <surname>Smith</surname>
          </string-name>
          , '
          <article-title>Toward abstractive summarization using semantic representations'</article-title>
          ,
          <source>in Proceedings of the 2015 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies</source>
          , pp.
          <fpage>1077</fpage>
          -
          <lpage>1086</lpage>
          , Denver, Colorado, (May-June
          <year>2015</year>
          ).
          <article-title>Association for Computational Linguistics</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          [23]
          <string-name>
            <given-names>Fei</given-names>
            <surname>Liu</surname>
          </string-name>
          and Yang Liu, '
          <article-title>From extractive to abstractive meeting summaries: Can it be done by sentence compression?'</article-title>
          ,
          <source>in Proceedings of the ACL-IJCNLP 2009 Conference Short Papers</source>
          , pp.
          <fpage>261</fpage>
          -
          <lpage>264</lpage>
          , Suntec, Singapore, (
          <year>August 2009</year>
          ).
          <article-title>Association for Computational Linguistics</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          [24]
          <string-name>
            <surname>Simon</surname>
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Burga</surname>
          </string-name>
          , and L. Wanner, '
          <article-title>AnCora-UPF: A multi-level annotation of Spanish'</article-title>
          ,
          <source>in Proceedings of the 2nd International Conference on Dependency Linguistics (DepLing)</source>
          , pp.
          <fpage>217</fpage>
          -
          <lpage>226</lpage>
          , Prague, Czech Republic, (
          <year>2013</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref25">
        <mixed-citation>
          [25]
          <string-name>
            <surname>Kathleen R. McKeown</surname>
          </string-name>
          , '
          <article-title>The text system for natural language generation: An overview'</article-title>
          ,
          <source>in Proceedings of the 20th Annual Meeting on Association for Computational Linguistics, ACL '82</source>
          , pp.
          <fpage>113</fpage>
          -
          <lpage>120</lpage>
          , Stroudsburg, PA, USA, (
          <year>1982</year>
          ).
          <article-title>Association for Computational Linguistics</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref26">
        <mixed-citation>
          [26]
          <string-name>
            <given-names>Y.</given-names>
            <surname>Mehdad</surname>
          </string-name>
          , G. Carenini, and R.T. Ng, '
          <article-title>Abstractive summarization of spoken and written conversations based on phrasal queries', in Proceedings of the 52nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)</article-title>
          , pp.
          <fpage>1220</fpage>
          -
          <lpage>1230</lpage>
          , Baltimore, Maryland, (
          <year>June 2014</year>
          ).
          <article-title>Association for Computational Linguistics</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref27">
        <mixed-citation>
          [27]
          <string-name>
            <given-names>Igor</given-names>
            <surname>Mel'cˇuk</surname>
          </string-name>
          ,
          <source>Dependency Syntax: Theory and Practice</source>
          , State University of New York Press, Albany,
          <year>1988</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref28">
        <mixed-citation>
          [28]
          <string-name>
            <given-names>Simon</given-names>
            <surname>Mille</surname>
          </string-name>
          and Leo Wanner, '
          <article-title>Towards large-coverage detailed lexical resources for data-to-text generation'</article-title>
          ,
          <source>in Proceedings of the First International Workshop on Data-to-text Generation</source>
          , Edinburgh, Scotland, (
          <year>2015</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref29">
        <mixed-citation>
          [29]
          <string-name>
            <given-names>A.</given-names>
            <surname>Moro</surname>
          </string-name>
          ,
          <string-name>
            <surname>A.</surname>
          </string-name>
          l Raganato,
          <article-title>and</article-title>
          <string-name>
            <given-names>R.</given-names>
            <surname>Navigli</surname>
          </string-name>
          , '
          <article-title>Entity linking meets word sense disambiguation: a unified approach', Transactions of the Association for Computational Linguistics</article-title>
          ,
          <volume>2</volume>
          ,
          <fpage>231</fpage>
          -
          <lpage>244</lpage>
          , (
          <year>2014</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref30">
        <mixed-citation>
          [30]
          <string-name>
            <given-names>Roberto</given-names>
            <surname>Navigli</surname>
          </string-name>
          and Simone Paolo Ponzetto, 'Babelnet:
          <article-title>The automatic construction, evaluation and application of a wide-coverage multilingual semantic network'</article-title>
          ,
          <source>Artificial Intelligence</source>
          ,
          <volume>193</volume>
          ,
          <fpage>217</fpage>
          -
          <lpage>250</lpage>
          , (
          <year>2012</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref31">
        <mixed-citation>
          [31]
          <string-name>
            <surname>M. O'Donnell</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          <string-name>
            <surname>Mellish</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          <string-name>
            <surname>Oberlander</surname>
          </string-name>
          ,
          <article-title>and</article-title>
          <string-name>
            <given-names>A.</given-names>
            <surname>Knott</surname>
          </string-name>
          , '
          <article-title>ILEX: an Architecture for a Dynamic Hypertext Generation System'</article-title>
          ,
          <source>Natural Language Engineering</source>
          ,
          <volume>7</volume>
          ,
          <fpage>225</fpage>
          -
          <lpage>250</lpage>
          , (
          <year>2001</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref32">
        <mixed-citation>
          [32]
          <string-name>
            <surname>Martha</surname>
            <given-names>Palmer</given-names>
          </string-name>
          , 'Semlink: Linking Propbank, VerbNet and FrameNet',
          <source>in Proceedings of the Generative Lexicon Conference (GenLex-09)</source>
          , Pisa, Italy, (
          <year>2009</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref33">
        <mixed-citation>
          [33]
          <string-name>
            <surname>Massimo</surname>
            <given-names>Poesio</given-names>
          </string-name>
          , Rosemary Stevenson, Barbara Di Eugenio, and Janet Hitzeman, '
          <article-title>Centering: A Parametric Theory</article-title>
          and Its Instantiations',
          <source>Computational Linguistics</source>
          ,
          <volume>30</volume>
          (
          <issue>3</issue>
          ),
          <fpage>309</fpage>
          -
          <lpage>363</lpage>
          , (
          <year>2004</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref34">
        <mixed-citation>
          [34]
          <string-name>
            <given-names>Ehud</given-names>
            <surname>Reiter</surname>
          </string-name>
          and
          <string-name>
            <given-names>Robert</given-names>
            <surname>Dale</surname>
          </string-name>
          ,
          <source>Building Natural Language Generation Systems</source>
          , Cambridge University Press, New York, NY, USA,
          <year>2000</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref35">
        <mixed-citation>
          [35]
          <string-name>
            <given-names>H.</given-names>
            <surname>Saggion</surname>
          </string-name>
          and G. Lapalme, '
          <article-title>Generating indicative-informative summaries with sumum', Computational linguistics</article-title>
          ,
          <volume>28</volume>
          (
          <issue>4</issue>
          ),
          <fpage>497</fpage>
          -
          <lpage>526</lpage>
          , (
          <year>2002</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref36">
        <mixed-citation>
          [36]
          <string-name>
            <given-names>Karin</given-names>
            <surname>Kipper</surname>
          </string-name>
          <string-name>
            <surname>Schuler</surname>
          </string-name>
          ,
          <article-title>VerbNet: A broad-coverage, comprehensive verb lexicon</article-title>
          ,
          <source>Ph.D. dissertation</source>
          , University of Pennsylvania,
          <year>2005</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref37">
        <mixed-citation>
          [37]
          <string-name>
            <surname>Lucy</surname>
            <given-names>Vanderwende</given-names>
          </string-name>
          , Michele Banko, and Arul Menezes, '
          <article-title>Event-centric summary generation'</article-title>
          ,
          <source>Working notes of DUC</source>
          ,
          <fpage>127</fpage>
          -
          <lpage>132</lpage>
          , (
          <year>2004</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref38">
        <mixed-citation>
          [38]
          <string-name>
            <given-names>L.</given-names>
            <surname>Wanner</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Bohnet</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Bouayad-Agha</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Lareau</surname>
          </string-name>
          , and
          <string-name>
            <given-names>D.</given-names>
            <surname>Nicklaß</surname>
          </string-name>
          , '
          <article-title>MARQUIS: Generation of user-tailored multilingual air quality bulletins'</article-title>
          ,
          <source>Applied Artificial Intelligence</source>
          ,
          <volume>24</volume>
          (
          <issue>10</issue>
          ),
          <fpage>914</fpage>
          -
          <lpage>952</lpage>
          , (
          <year>2010</year>
          ).
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>