<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Trainable Citation-enhanced Summarization of Scienti c Articles</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Horacio Saggion</string-name>
          <email>horacio.saggion@upf.edu</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Ahmed AbuRa'ed</string-name>
          <email>ahmed.aburaed@upf.edu</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Francesco Ronzano</string-name>
          <email>francesco.ronzano@upf.edu</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>TALN - DTIC Universitat Pompeu Fabra Barcelona</institution>
          ,
          <country country="ES">Spain</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2016</year>
      </pub-date>
      <fpage>175</fpage>
      <lpage>186</lpage>
      <abstract>
        <p>In order to cope with the growing number of relevant scienti c publications to consider at a given time, automatic text summarization is a useful technique. However, summarizing scienti c papers poses important challenges for the natural language processing community. In recent years a number of evaluation challenges have been proposed to address the problem of summarizing a scienti c paper taking advantage of its citation network (i.e., the papers that cite the given paper). Here, we present our trainable technology to address a number of challenges in the context of the 2nd Computational Linguistics Scienti c Document Summarization Shared Task.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Introduction</title>
      <p>
        During the last decade the amount of scienti c information available on-line
increased at an unprecedented rate with recent estimates reporting a new paper
published every 20 seconds [
        <xref ref-type="bibr" rid="ref17">17</xref>
        ]. In this scenario of scienti c information overload,
researchers are overwhelmed by an enormous and continuously growing number
of articles to consider in their research work: from the exploration of advances
in speci c topics, to peer reviewing, writing and evaluation. In order to cope
with the growing number of relevant publications to consider at a given time,
automatic text summarization is a useful technique [
        <xref ref-type="bibr" rid="ref23">23</xref>
        ]. However, generic text
summarization techniques may not work well in specialized genres such as the
scienti c genre and domain speci c techniques may be needed [
        <xref ref-type="bibr" rid="ref25 ref26">25, 26</xref>
        ]. Scienti c
publications are characterized by several structural, linguistic and semantic
peculiarities. Articles include common structural elements (title, authors, abstract,
sections, gures, tables, citations, bibliography) that often require speci c text
processing tools. Additionally, scienti c documents have speci c discourse
structure [
        <xref ref-type="bibr" rid="ref12 ref27">27, 12</xref>
        ]. Another important aspect of scienti c papers is their network of
citations that identi es links among research works, making them also
particularly interesting from the social viewpoint. Although citation counts had been
used to assess some aspects of research output for a long time, citation semantics
has started to be exploited in several context including opinion mining [
        <xref ref-type="bibr" rid="ref1 ref28">28, 1</xref>
        ]
and scienti c text summarization [
        <xref ref-type="bibr" rid="ref18 ref2">18, 2</xref>
        ].
Considering the urgent need for new, automated approaches to browse and
aggregate scienti c information, in recent years a number of natural language
processing challenges have been proposed: the Biomedical Summarization Task
(BioSumm2014) carried out in the context of the Text Analysis Conferences1
provided a forum for researchers interested in exploring the summarization of
clusters of documents where one of the documents is a reference paper and the
rest of the documents in the cluster are citing papers which cite the reference
paper. In particular, the BioSumm2014 evaluation released a dataset consisting
of 20 collections of annotated papers (i.e., clusters), each one including a
reference article and 10 citing articles. Similarly, a pilot task on summarization of
Computational Linguistic papers was proposed in 2014 [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ]. Unfortunately, none
of the evaluation contests provided with o cial evaluation results.
In this paper, we report our e orts to develop a system to participate in the
CLSciSumm 2016 evaluation [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ] which is a renewed e ort to address the challenges
proposed in 2014. In a nutshell, participants were given a set of clusters, each
one composed of n documents where one is a reference paper (RP) and the n-1
remaining documents are referred to as citing papers (CPs) since they cite the
reference paper. Participants have to develop automatic procedures to perform
the following tasks:
{ Task 1A: For each citance (i.e., a reference to the RP), identify the spans
of text (cited text spans) in the RP that most accurately re ect the citance.
{ Task 1B: For each cited text span, identify what facet of the paper it
belongs to, from a prede ned set of facets, namely: Aim, Hypothesis,
Implication, Results or Method.
{ Task 2: Finally, an optional task consists on generating a structured (of up
to 250 words) summary of the RP from the cited text spans of the RP.
      </p>
      <p>In the rest of this paper we rst present related work on summarization of
research articles to then explain how we have addressed the di erent
summarization tasks.
2</p>
    </sec>
    <sec id="sec-2">
      <title>Related work</title>
      <p>
        Although research in summarization can be traced back to the 50s [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ] and
although a number of important discoveries have been produced in this area,
automatic text summarization still faces many challenges given its inherent
complexity. Scienti c text summarization is of paramount importance and scienti c
texts were automatic summarization's rst application domain [
        <xref ref-type="bibr" rid="ref14 ref5">14, 5</xref>
        ]. Several
methods and techniques have been already reported in the literature to produce
text summaries by automatic means [
        <xref ref-type="bibr" rid="ref23">23</xref>
        ]. Summarization of scienti c documents
has been addressed from di erent angles: in [
        <xref ref-type="bibr" rid="ref26">26</xref>
        ] summarization is treated as
      </p>
      <sec id="sec-2-1">
        <title>1 http://www.nist.gov/tac/2014/</title>
        <p>
          BiomedSumm/
a rhetorical classi cation task where each sentence in an input text is
classied as belonging to speci c rhetorical categories (background, objective, etc.).
Although the approach is interesting from the point of view of document
interpretation, it is not a proper summarization task since no summary of the
input is produced. [
          <xref ref-type="bibr" rid="ref25">25</xref>
          ] addressed the summarization problem as one of
information extraction and text generation: the idea behind the approach is that a
number of important concepts and relations should be extracted from text in
order to create a summary independently of the particular scienti c domain of
the text. In recent years new generations of scienti c summarization approaches
have emerged which take advantage of the citations that a research paper has in
order to extract and summarize its main contributions [
          <xref ref-type="bibr" rid="ref19">19</xref>
          ]. Methods to improve
the coherence of the generated citation summaries use sentence classi cation to
decide what type of information a sentence is conveying [
          <xref ref-type="bibr" rid="ref1">1</xref>
          ].
3
        </p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>Transforming the Source Documents into GATE</title>
    </sec>
    <sec id="sec-4">
      <title>Language Resources</title>
      <p>CL-SciSumm 2016 Challenge organizers have provided training data structured
in clusters of reference and citing papers together with manual annotations
indicating for each citance to the reference paper, the facet of this citance and the
text span(s) in the reference paper that best represent the citance. In order to
properly analize the provided training and testing documents, we transformed
the provided clusters into GATE documents. Given the manual annotations
provided in text les, we automatically annotated the training set by creating in
the reference paper a References Annotation set which contains the annotations
corresponding to the text spans being cited by each citing paper. On the other
hand, an Annotation set for the citances in each of the citing papers was also
created. The link between citing paper annotations and reference paper
annotations is implemented through a unique identi er (a concatenation of citance
number, reference paper, citing paper, and annotator).</p>
      <p>Such annotations are helpful in order to retrieve the necessary information
from the documents. In this way, in each citing paper we are able to identify
for each sentence that belongs to a citance, which are the sentences of the
corresponding reference paper that most accurately re ect the citance. Thanks to
this information, we can build pairs of matching sentences (Citing Paper
Sentence, Reference Paper Sentence) and associate to each pair the facet that each
annotator considers the citation is referring to (see Task 1B).</p>
      <p>An example of the representation can be seen in Figure 1 where it is shown a
reference paper (on the left side of the Figure) annotated with information from
the citing papers (on the right side of the Figure).
3.1</p>
      <sec id="sec-4-1">
        <title>Text Processing</title>
        <p>
          Each document was annotated using processing resources from the GATE
system [
          <xref ref-type="bibr" rid="ref15">15</xref>
          ] and the SUMMA library [
          <xref ref-type="bibr" rid="ref22">22</xref>
          ]. Additionally, and in order to further
enrich the documents, some components from the freely-available Dr Inventor
library2 (DRI Framework) were used [
          <xref ref-type="bibr" rid="ref21">21</xref>
          ]. The GATE system was used to
tokenize, sentence split, part of speech tag, and lemmatize each document. The
SUMMA library was used to produce normalized term vectors for each
document (see Section 5). Although the Dr Inventor's library produces very rich
information, for the experiments we present here we rely only on its rhetorical
sentence classi cation capabilities. The DRI Framework classify each sentence
of a paper as belonging to a rethorical category of scienti c discourse among:
Approach, Background, Challenge, Outcome and FutureWork. In particular, the
framework computes for each sentence the probability the sentece has to belong
to each rhetorical category. See [
          <xref ref-type="bibr" rid="ref6">6</xref>
          ] for details about the corpus used for training
the classi er. For each sentence in the reference paper, we computed the cosine
similarity between its sentence vector (see Section 5) and the vectors
corresponding to the sentences citing the reference paper in the citing articles. These values
were stored in the reference paper for further processing.
4
        </p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>Method</title>
      <p>In order to identify reference paper text spans for each citance (Task 1A), we
modeled pairs of reference and citance sentences as a feature vector. Then, we
used such pair representation to enable the training of distinct binary classi
cation algorithms tailored to determine whether they are a match.</p>
      <p>On the other hand we used the same representation of pairs of sentences
for identifying to what facet of the reference paper a cited text span belongs to</p>
      <sec id="sec-5-1">
        <title>2 http://backingdata.org/dri/</title>
        <p>library/
(Task 1B): we classi ed each pair of sentences in one out of 5 prede ned facets
Aim, Hypothesis, Implication, Results or Method.</p>
        <p>
          To this end, we relied on the Weka machine learning framework [
          <xref ref-type="bibr" rid="ref29">29</xref>
          ]. We
evaluated the performance of six classi cation algorithms: SMO, Naive Bayes,
J48, Lazy IBK, Decision table and Random Forest for both tasks. We performed
10-fold cross validation experiments with the training data in order to decide
which algorithm to use during testing.
        </p>
        <p>In the remainder of this Section, we describe the set of sentence pair features
we used, and motivate their relevance with respect to the characterization of
sentences similarity. When presenting the features, we group subsets of related
features in the same subsection (Position features, Similarity features, etc.).
4.1</p>
        <sec id="sec-5-1-1">
          <title>Position Features</title>
          <p>We exploited the following set of position related features for both Task 1A (text
spans for citance sentence) and 1B (the facet such text span belongs to):
{ Sentence position (sentence position): the normalized position of the
sentence in the reference paper.
{ Sentence section position (sentence section position): the normalized
position of the sentence in the section of the reference paper.
{ Facet position (facet aim, facet hypothesis, facet implication, facet method
and facet result): ve features were generated to indicate which facet a
cited text span belongs to. Binary values were calculated by analyzing the
reference paper sentence's section title and looking for any words which
could indicate the feature facet: aim, hypothesis, implication, method or
result. The value of the feature is 1 for section titles containing a word that
indicate such facet and 0 otherwise.
4.2</p>
        </sec>
        <sec id="sec-5-1-2">
          <title>WordNet Semantic Similarity Measures features</title>
          <p>
            The following set of Semantic Similarity features were exploited for task 1A (text
spans for citance sentence) only with the exception of the cosine similarity which
was used for both task 1A and task 1B. We used WS4J (WordNet Similarity for
Java) library which includes several semantic relatedness algorithms that rely
on WordNet 3.0. Given a pair of sentences (reference and citance), we retrieve
all the synsets associated to nouns and verbs in each one of them. Then, by
considering all the pairs of synsets belonging to di erent sentences, we compute
similarity values between citance sentence and reference sentence as follows3:
{ Path similarity [
            <xref ref-type="bibr" rid="ref7">7</xref>
            ] (path similarity): The shorter the path between two
words/senses in WordNet, the more similar they are.
3 We calculated similarity values between
each token in the citance sentence and
each and every token in the reference
sentence. Finally averaging all the
similaries for the given sentence pair.
{ JCN similarity [
            <xref ref-type="bibr" rid="ref10">10</xref>
            ] (jiangconrath similarity): the conditional probability
of encountering an instance of a child-synset given an instance of a parent
synset.
{ LCH similarity [
            <xref ref-type="bibr" rid="ref11">11</xref>
            ] (lch similarity): the length of the shortest path between
two synsets for their measure of similarity.
{ LESK similarity [
            <xref ref-type="bibr" rid="ref3">3</xref>
            ] (lesk similarity): Similarity of two concepts is de ned
as a function of the overlap between the corresponding de nitions (i.e., their
WordNet glosses).
{ LIN similarity [
            <xref ref-type="bibr" rid="ref13">13</xref>
            ] (lin similarity): The Similarity between A and B is
measured by the ratio between the amount of information needed to state
the commonality of A and B and the information needed to fully describe
what A and B are.
{ RESNIK similarity [
            <xref ref-type="bibr" rid="ref20">20</xref>
            ] (resnik similarity): The probability of
encountering an instance of concept c in a large corpus.
{ WUP similarity [
            <xref ref-type="bibr" rid="ref30">30</xref>
            ] (wup similarity): The depths of the two synsets in the
          </p>
          <p>WordNet taxonomies, along with the depth of the lowest common subsumer.
{ Cosine similarity (cosine similarity): The cosine similarity between the
normalized vectors of the two sentences in the instance pair (this
computation is di erent from the other similarity features).
4.3</p>
        </sec>
        <sec id="sec-5-1-3">
          <title>Rhetorical Category Probability Features</title>
          <p>We exploited the following set of rhetorical features for both task 1A (text spans
for citance sentence) and 1B (the facet which the text span belongs to):
{ Rhetorical Category Probability (probability approach, probability background,
probability challenge, probability future work and probability outcome):
ve features were exploited to represent the probability of the reference text
span to belong to such facet (from the Dr Inventor corpus and computed
from the Dr Inventor library).</p>
          <p>We also added both the reference sentence string and the citance sentence
string to the set of features and then converted them to word vectors by using
WEKA (i.e., bag-of-words).
4.4</p>
        </sec>
        <sec id="sec-5-1-4">
          <title>Matching Citations to Reference Papers</title>
          <p>The training data was prepared as follows: positive instances of the problem
were the pairs of sentences from the citance which where matched with cited
text spans from the references (according to information given in the gold
annotations). Negative instances, instead, were pairs of sentences from citances
to identi ed cited text spans which were not annotated as matches by the
annotators (complementary information). As a consequence we casted the Task
1A as a binary classi cation problem where we decide for each pairs of citance
sentence and reference paper sentence whether they match or not, or in other
words whether the reference paper sentence re ects the reason of that speci c
citations. These procedure, which was decided upon to reduce the number of
negative cases, produced 3,786 instances unevenly distributed (3,356 no matches vs
430 matches). After testing several algorithms from WEKA, we opted for the J48
implementation of decision threes. Ten fold cross-validation results are presented
in Table 1.
The training data was prepared as follows: similarly to the previous task, pairs of
citing sentences and matched cited sentences (according to the gold annotations)
were used to create instances. The facet of each instance was also given by the
gold standard. This procedure produced just 432 instances with the following
distribution: Aim (72), Implication (26), Result (76), Hypothesis (1), Method
(257). After testing several algorithms from WEKA, we opted for the Support
Vector Machines (SMO) implementation provided by the tool. We used
polynomial Kernels and performed no parameter optimization due to time constraints.
Ten fold cross-validation results are presented in Table 2.</p>
        </sec>
      </sec>
    </sec>
    <sec id="sec-6">
      <title>Summarizing Scienti c Articles</title>
      <p>
        In order to summarize the reference paper by taking into account how it is
mentioned in the citing papers, we combined information from the reference and
citing papers. We have implemented, using the resources of the freely available
text summarization library SUMMA [
        <xref ref-type="bibr" rid="ref22 ref24">22, 24</xref>
        ], a series of sentence relevance
features, all numeric, which are used to train a linear regression model following
the methodology that was already used in [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ].
      </p>
      <p>In addition to rich set of features provided by the DRI Framework, document
processing for summarization is carried out with SUMMA on reference and citing
papers. More speci cally, the following computations with the library are carried
out to enable the summarization of scienti c documents:
{ Each token (i.e., lemma) is weighted by its term frequency* inverted
document frequency, where inverted document values are computed from training
data previously analysed (test documents in the CL-SciSumm 2016 dataset);
{ For each sentence a vector of terms and normalized weights is created using
the previously computed weights;
{ For the title, a single vector of terms and normalized weights is also created
(title vector);
{ Using the normalized sentence term vectors in the whole document a centroid
vector of terms is computed (document centroid);
{ Using the normalized sentence term vectors of the abstracts a centroid vector
of terms is computed (abstract centroid);
{ All vectors corresponding to sentences citing the reference paper (from all
citing papers) are used to create a centroid (citances vector).</p>
      <p>
        The following is the set of sentence relevance features we have used for
training a linear regression summarization system. Note that all text-based
similarities we mention are the result of comparing two vectors using the cosine similarity
function implemented in SUMMA. The reference paper features are as follows:
{ Sentence Abstract Similarity (abs sim): the similarity of a sentence to the
author abstract;
{ Sentence Centroid Similarity (centroid sim): the similarity of a sentence
to the document centroid (e.g., the average of all sentence vectors in the
document);
{ First Sentence Similarity ( rt sim): the similarity of a sentence to the title
vector;
{ Position Score (position score): the SUMMA implementation of the
position method where sentences at the beginning of the document have high
scores and sentence at the end of the document have low scores;
{ Position in Section Score (in sec): an score representing the position of the
sentence in the section of the document. Sentences in rst section get higher
scores, sentences in last section get low scores;
{ Sentence Position in Section Score (in sec sent): a position method applied
to sentences in each section of the document (sentence at the beginning of
the section get higher scores and sentences at the end of the section get lower
scores);
{ Normalised Cue-phrase Score(norm cue): we produce a normalized score
for each sentence which is the total number of cue-words in the sentence
divided by the total number of cue-words in the document. We have relied
on [
        <xref ref-type="bibr" rid="ref26">26</xref>
        ] formulaic expressions to implement our cue-phrase gazetteer lookup
procedure;
{ TextRank Normalized Score (textrank score): the SUMMA
implementation of the TextRank algorithm [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ] but with a normalization procedure
which yields values for sentences between 0 and 1.
      </p>
      <p>The cluster-based features are as follows:
{ Citing Paper Maximum Similarity (cps max): each reference paper sentence
vector is compared (using cosine) to each citance vector in each citing paper
to obtain the maximum possible cosine similarity;
{ Citing Paper Average Similarity (cps avg): the average cosine similarity
between a reference paper vector and all citance vectors in the cluster is
produced;
{ Citing Paper Citances Similarity (cps sim): the similary of the sentence
vector to the centroid of the citance vectors.</p>
      <p>
        The approach taken to score sentence is to produce a cumulative score of the
weighted values of summarization features f1; :::fn using the following formula:
score(S) =
n
X wi fi
i=0
(1)
with S as the sentence to score, fi as the value of feature i and wi as the
weight assigned to feature i. As we stated before, the weights of each features
in the formula are learned from training data. We t a linear regression model
using the 10 testing documents from the provided annotated document for a
total of 2,585 instances. The target numerical value to learn is computed from
two sources (giving rise to two di erent systems): On the one hand, we compute
the similarity of each reference paper sentence (i.e. vector) to the combined
vectors of texts fragments identi ed as the annotators as cited text spans; on
the other hand, we compute the similarity of each reference paper sentence (i.e.
vector) to a vector of the community-based summary provided for training by
the organizers. Table 3 shows the weights of the features learnt by the linear
regression implementation from WEKA [
        <xref ref-type="bibr" rid="ref29">29</xref>
        ].
6
      </p>
    </sec>
    <sec id="sec-7">
      <title>The Final System</title>
      <p>The nal system was assembled as follows. Given a cluster of documents with
reference and citing papers, the following pipeline was applied for tasks 1A and
1B.
1. The documents were annotated with the citance information (no matched
reference sentences were annotated);
2. All the document processing algorithms were applied to reference and citing
papers as described in Section 3.1 and the features computed;
3. Instances were created using a citance sentence from each citing paper and
each sentence from the reference paper;
4. The instances were sent to the matching classi er which returned a match/no
match class and a con dence value;
5. The matched instances according to the previous steps were sent to the facet
classi er to obtain the predicted citation facet.</p>
      <p>Two runs were produced for tasks 1A and 1B. In one run, all matched
sentences for a given citance were returned. In a second run, only top matches (with
higher con dence) were returned. In order to produce the summaries for each
cluster, summarization features were computed using the procedure described in
Section 5, and SUMMA was exploited to score and extract top scored sentences
based on formula (1). Two 250-word text summaries were produced per cluster
using the models described in Section 5.
7</p>
    </sec>
    <sec id="sec-8">
      <title>Outlook</title>
      <p>In this paper, we have presented the techniques used to participate in the
Computational Linguistics Summarization challenge. We have relied on competitive text
processing and summarization tools to compute features to create rich document
representations for dealing with the proposed tasks. Our approach is supervised
combining evidence from several sources. Due to time limitations we could not
carry out an exhaustive performance and feature analysis on test/development
data, which we intend to carry out as future work. We look forward to know the
o cial results of the evaluation so as to better understand the pros and cons of
our approach.</p>
    </sec>
    <sec id="sec-9">
      <title>Acknowledgements</title>
      <p>This work is (partly) supported by the Spanish Ministry of Economy and
Competitiveness under the Maria de Maeztu Units of Excellence Programme
(MDM2015-0502), the TUNER project (TIN2015-65308-C5-5-R, MINECO/FEDER,
UE) and the European Project Dr. Inventor (FP7-ICT-2013.8.1 - Grant no:
611383).</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Abu-Jbara</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ezra</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Radev</surname>
            ,
            <given-names>D.R.</given-names>
          </string-name>
          :
          <article-title>Purpose and polarity of citation: Towards nlp-based bibliometrics</article-title>
          .
          <source>In: HLT-NAACL</source>
          . pp.
          <volume>596</volume>
          {
          <issue>606</issue>
          (
          <year>2013</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Abu-Jbara</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Radev</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          :
          <article-title>Coherent citation-based summarization of scienti c papers</article-title>
          .
          <source>In: Proceedings of the 49th Annual Meeting of the Association for Computational Linguistics: Human Language Technologies - Volume 1</source>
          . pp.
          <volume>500</volume>
          {
          <fpage>509</fpage>
          . HLT '
          <volume>11</volume>
          ,
          <string-name>
            <surname>Association</surname>
          </string-name>
          for Computational Linguistics, Stroudsburg, PA, USA (
          <year>2011</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Banerjee</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Pedersen</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          :
          <article-title>An adapted lesk algorithm for word sense disambiguation using wordnet</article-title>
          .
          <source>In: Computational linguistics and intelligent text processing</source>
          , pp.
          <volume>136</volume>
          {
          <fpage>145</fpage>
          . Springer (
          <year>2002</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4. Brugmann,
          <string-name>
            <given-names>S.</given-names>
            ,
            <surname>Bouayad-Aghab</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            ,
            <surname>Burga</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            ,
            <surname>Carrascosa</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            ,
            <surname>Ciaramella</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            ,
            <surname>Ciaramella</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            ,
            <surname>Codina-Filba</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            ,
            <surname>Escorsa</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            ,
            <surname>Judea</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            ,
            <surname>Mille</surname>
          </string-name>
          ,
          <string-name>
            <surname>S.</surname>
          </string-name>
          , Muller,
          <string-name>
            <given-names>A.</given-names>
            ,
            <surname>Saggion</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            ,
            <surname>Ziering</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            , Schutze, H.,
            <surname>Wanner</surname>
          </string-name>
          ,
          <string-name>
            <surname>L.</surname>
          </string-name>
          :
          <article-title>Towards content-oriented patent document processing: Intelligent patent analysis and summarization</article-title>
          .
          <source>World Patent Information</source>
          <volume>40</volume>
          ,
          <issue>30</issue>
          {
          <fpage>42</fpage>
          (
          <year>2015</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Edmundson</surname>
            ,
            <given-names>H.P.</given-names>
          </string-name>
          :
          <article-title>New methods in automatic extracting</article-title>
          .
          <source>J. ACM</source>
          <volume>16</volume>
          (
          <issue>2</issue>
          ),
          <volume>264</volume>
          {285 (Apr
          <year>1969</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Fisas</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ronzano</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Saggion</surname>
          </string-name>
          , H.:
          <article-title>A Multi-Layered Annotated Corpus of Scienti c Papers</article-title>
          . In: Calzolari,
          <string-name>
            <given-names>N.</given-names>
            ,
            <surname>Choukri</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            ,
            <surname>Declerck</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            ,
            <surname>Grobelnik</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            ,
            <surname>Maegaard</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            ,
            <surname>Mariani</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            ,
            <surname>Moreno</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            ,
            <surname>Odijk</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            ,
            <surname>Piperidis</surname>
          </string-name>
          , S. (eds.)
          <source>Proceedings of the Tenth International Conference on Language Resources and Evaluation (LREC</source>
          <year>2016</year>
          ).
          <source>European Language Resources Association (ELRA)</source>
          , Paris, France (may
          <year>2016</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>Hirst</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>St-Onge</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          :
          <article-title>Lexical chains as representations of context for the detection and correction of malapropisms</article-title>
          .
          <source>WordNet: An electronic lexical database 305</source>
          , 305{
          <fpage>332</fpage>
          (
          <year>1998</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>Jaidka</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chandrasekaran</surname>
            ,
            <given-names>M.K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Elizalde</surname>
            ,
            <given-names>B.F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Jha</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Jones</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kan</surname>
            ,
            <given-names>M.Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Khanna</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Molla-Aliod</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Radev</surname>
            ,
            <given-names>D.R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ronzano</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Saggion</surname>
          </string-name>
          , H.:
          <article-title>The computational linguistics summarization pilot task</article-title>
          .
          <source>In: Proceedings of TAC</source>
          <year>2014</year>
          (
          <year>2014</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <surname>Jaidka</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chandrasekaran</surname>
            ,
            <given-names>M.K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rustagi</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kan</surname>
          </string-name>
          , M.Y.:
          <article-title>Overview of the 2nd Computational Linguistics Scienti c Document Summarization Shared Task (CLSciSumm 2016)</article-title>
          . In: To appear
          <source>in the Proceedings of the Joint Workshop on Bibliometric-enhanced Information Retrieval and Natural Language Processing for Digital Libraries (BIRNDL</source>
          <year>2016</year>
          )
          <article-title>(</article-title>
          <year>2016</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <surname>Jiang</surname>
            ,
            <given-names>J.J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Conrath</surname>
            ,
            <given-names>D.W.</given-names>
          </string-name>
          :
          <article-title>Semantic similarity based on corpus statistics and lexical taxonomy</article-title>
          .
          <source>arXiv preprint cmp-lg/9709008</source>
          (
          <year>1997</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <string-name>
            <surname>Leacock</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chodorow</surname>
            ,
            <given-names>M.:</given-names>
          </string-name>
          <article-title>Combining local context and wordnet similarity for word sense identi cation</article-title>
          .
          <source>WordNet: An electronic lexical database 49(2)</source>
          ,
          <volume>265</volume>
          {
          <fpage>283</fpage>
          (
          <year>1998</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          12.
          <string-name>
            <surname>Liakata</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Teufel</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Siddharthan</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Batchelor</surname>
            ,
            <given-names>C.R.</given-names>
          </string-name>
          , et al.:
          <article-title>Corpora for the conceptualisation and zoning of scienti c papers</article-title>
          .
          <source>In: LREC</source>
          (
          <year>2010</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          13.
          <string-name>
            <surname>Lin</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          :
          <article-title>An information-theoretic de nition of similarity</article-title>
          .
          <source>In: ICML</source>
          . vol.
          <volume>98</volume>
          , pp.
          <volume>296</volume>
          {
          <issue>304</issue>
          (
          <year>1998</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          14.
          <string-name>
            <surname>Luhn</surname>
            ,
            <given-names>H.P.:</given-names>
          </string-name>
          <article-title>The automatic creation of literature abstracts</article-title>
          .
          <source>IBM J. Res. Dev</source>
          .
          <volume>2</volume>
          (
          <issue>2</issue>
          ),
          <volume>159</volume>
          {165 (Apr
          <year>1958</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          15.
          <string-name>
            <surname>Maynard</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Tablan</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Cunningham</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ursu</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Saggion</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bontcheva</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Wilks</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          :
          <article-title>Architectural Elements of Language Engineering Robustness</article-title>
          .
          <source>Journal of Natural Language Engineering { Special Issue on Robust Methods in Analysis of Natural Language Data</source>
          <volume>8</volume>
          (
          <issue>2</issue>
          /3),
          <volume>257</volume>
          {
          <fpage>274</fpage>
          (
          <year>2002</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          16.
          <string-name>
            <surname>Mihalcea</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Tarau</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          : TextRank:
          <article-title>Bringing order into texts</article-title>
          .
          <source>In: Proceedings of EMNLP-04and the 2004 Conference on Empirical Methods in Natural Language Processing (July</source>
          <year>2004</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          17.
          <string-name>
            <surname>Munroe</surname>
            ,
            <given-names>R.:</given-names>
          </string-name>
          <article-title>The rise of open access</article-title>
          .
          <source>Science</source>
          <volume>342</volume>
          (
          <issue>6154</issue>
          ),
          <volume>58</volume>
          {
          <fpage>59</fpage>
          (
          <year>2013</year>
          ), https: //www.sciencemag.org/content/342/6154/58.full
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          18.
          <string-name>
            <surname>Qazvinian</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Radev</surname>
            ,
            <given-names>D.R.</given-names>
          </string-name>
          :
          <article-title>Identifying non-explicit citing sentences for citationbased summarization</article-title>
          .
          <source>In: ACL</source>
          <year>2010</year>
          ,
          <article-title>Proceedings of the 48th Annual Meeting of the Association for Computational Linguistics</article-title>
          ,
          <source>July 11-16</source>
          ,
          <year>2010</year>
          , Uppsala, Sweden. pp.
          <volume>555</volume>
          {
          <issue>564</issue>
          (
          <year>2010</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          19.
          <string-name>
            <surname>Qazvinian</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Radev</surname>
            ,
            <given-names>D.R.</given-names>
          </string-name>
          :
          <article-title>Identifying non-explicit citing sentences for citationbased summarization</article-title>
          . In:
          <article-title>Proceedings of the 48th annual meeting of the association for computational linguistics</article-title>
          . pp.
          <volume>555</volume>
          {
          <fpage>564</fpage>
          .
          <article-title>Association for Computational Linguistics (</article-title>
          <year>2010</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          20.
          <string-name>
            <surname>Resnik</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          :
          <article-title>Using information content to evaluate semantic similarity in a taxonomy</article-title>
          .
          <source>arxiv preprint cmplg/9511007</source>
          (
          <year>1995</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          21.
          <string-name>
            <surname>Ronzano</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Saggion</surname>
          </string-name>
          , H.:
          <article-title>Dr. inventor framework: Extracting structured information from scienti c publications</article-title>
          .
          <source>In: Discovery Science - 18th International Conference, DS</source>
          <year>2015</year>
          ,
          <article-title>Ban</article-title>
          ,
          <string-name>
            <surname>AB</surname>
          </string-name>
          , Canada, October 4-
          <issue>6</issue>
          ,
          <year>2015</year>
          , Proceedings. pp.
          <volume>209</volume>
          {
          <issue>220</issue>
          (
          <year>2015</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          22.
          <string-name>
            <surname>Saggion</surname>
          </string-name>
          , H.:
          <article-title>SUMMA: A Robust and Adaptable Summarization Tool</article-title>
          .
          <source>Traitement Automatique des Langues</source>
          <volume>49</volume>
          (
          <issue>2</issue>
          ),
          <volume>103</volume>
          {
          <fpage>125</fpage>
          (
          <year>2008</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          23.
          <string-name>
            <surname>Saggion</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Poibeau</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          :
          <article-title>Automatic text summarization: Past, present andfuture</article-title>
          . In: Poibeau,
          <string-name>
            <given-names>T.</given-names>
            ,
            <surname>Saggion</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            ,
            <surname>Piskorski</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            ,
            <surname>Yangarber</surname>
          </string-name>
          ,
          <string-name>
            <surname>R</surname>
          </string-name>
          . (eds.) Multi-source,
          <source>Multilingual Information Extraction and Summarization</source>
          . Springer Verlag, Berlin (
          <year>2013</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          24.
          <string-name>
            <surname>Saggion</surname>
          </string-name>
          , H.:
          <article-title>Creating summarization systems with SUMMA</article-title>
          .
          <source>In: Proceedings of the Ninth International Conference on Language Resources and Evaluation (LREC2014)</source>
          , Reykjavik, Iceland, May
          <volume>26</volume>
          -31,
          <year>2014</year>
          . pp.
          <volume>4157</volume>
          {
          <issue>4163</issue>
          (
          <year>2014</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref25">
        <mixed-citation>
          25.
          <string-name>
            <surname>Saggion</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lapalme</surname>
          </string-name>
          , G.:
          <article-title>Generating indicative-informative summaries with sumum</article-title>
          .
          <source>Comput. Linguist</source>
          .
          <volume>28</volume>
          (
          <issue>4</issue>
          ),
          <volume>497</volume>
          {526 (Dec
          <year>2002</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref26">
        <mixed-citation>
          26.
          <string-name>
            <surname>Teufel</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Moens</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          :
          <article-title>Summarizing scienti c articles: Experiments with relevance and rhetorical status</article-title>
          .
          <source>Comput. Linguist</source>
          .
          <volume>28</volume>
          (
          <issue>4</issue>
          ),
          <volume>409</volume>
          {445 (Dec
          <year>2002</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref27">
        <mixed-citation>
          27.
          <string-name>
            <surname>Teufel</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Siddharthan</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Batchelor</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          :
          <article-title>Towards discipline-independent argumentative zoning: evidence from chemistry and computational linguistics</article-title>
          .
          <source>In: Proceedings of the 2009 Conference on Empirical Methods in Natural Language Processing: Volume 3-</source>
          Volume 3. pp.
          <volume>1493</volume>
          {
          <fpage>1502</fpage>
          .
          <article-title>Association for Computational Linguistics (</article-title>
          <year>2009</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref28">
        <mixed-citation>
          28.
          <string-name>
            <surname>Teufel</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Siddharthan</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Tidhar</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          :
          <article-title>Automatic classi cation of citation function</article-title>
          .
          <source>In: Proceedings of the 2006 conference on empirical methods in natural language processing</source>
          . pp.
          <volume>103</volume>
          {
          <fpage>110</fpage>
          .
          <article-title>Association for Computational Linguistics (</article-title>
          <year>2006</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref29">
        <mixed-citation>
          29.
          <string-name>
            <surname>Witten</surname>
            ,
            <given-names>I.H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Frank</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hall</surname>
            ,
            <given-names>M.A.</given-names>
          </string-name>
          :
          <article-title>Data Mining: Practical Machine Learning Tools and Techniques</article-title>
          . Morgan Kaufmann Publishers Inc., San Francisco, CA, USA, 3rd edn. (
          <year>2011</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref30">
        <mixed-citation>
          30.
          <string-name>
            <surname>Wu</surname>
            ,
            <given-names>Z.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Palmer</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          :
          <article-title>Verbs semantics and lexical selection</article-title>
          .
          <source>In: Proceedings of the 32nd annual meeting on Association for Computational Linguistics</source>
          . pp.
          <volume>133</volume>
          {
          <fpage>138</fpage>
          .
          <article-title>Association for Computational Linguistics (</article-title>
          <year>1994</year>
          )
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>