<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>University of Mannheim @ CLSciSumm-17: Citation-Based Summarization of Scienti c Articles Using Semantic Textual Similarity</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Anne Lauscher</string-name>
          <email>anne@informatik.uni-mannheim.de</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Goran Glavas</string-name>
          <email>goran@informatik.uni-mannheim.de</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Kai Eckert</string-name>
          <email>eckert@hdm-stuttgart.de</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Stuttgart Media University, Web-based Information Systems and Services</institution>
          ,
          <addr-line>Nobelstra e 10, 70569 Stuttgart</addr-line>
          ,
          <country country="DE">Germany</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>University of Mannheim, Data and Web Science Research Group</institution>
          ,
          <addr-line>B6 26, 68159 Mannheim</addr-line>
          ,
          <country country="DE">Germany</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>The number of publications is rapidly growing and it is essential to enable fast access and analysis of relevant articles. In this paper, we describe a set of methods based on measuring semantic textual similarity, which we use to semantically analyze and summarize publications through other publications that cite them. We report the performance of our approach in the context of the third CL-SciSumm shared task and show that our system performs favorably to competing systems in terms of produced summaries.</p>
      </abstract>
      <kwd-group>
        <kwd>Scienti c Publication Mining</kwd>
        <kwd>Scienti c Summarization</kwd>
        <kwd>Information Retrieval</kwd>
        <kwd>Text Classi cation</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>
        Citations play an important role in the interpretation of scienti c literature and
help retrace the evolution of scienti c ideas. The text surrounding the citation
often reveals some important aspect of the cited publication, e.g., the purpose,
polarity, or function (e.g., method or hypothesis) of the citation [
        <xref ref-type="bibr" rid="ref1 ref10 ref8">1, 10, 8</xref>
        ].
      </p>
      <p>
        The citances of a publication, i.e., the sentences of the citing articles
containing the citation to the publication in focus [
        <xref ref-type="bibr" rid="ref17">17</xref>
        ], are often very useful for
higher level analyses of referenced publications, and may contribute to better
summarization of scienti c articles. The Computational Linguistics Scienti c
Document Summarization Shared Task (CL-SciSumm) has been designed precisely
to encourage the exploitation of citing contexts in automatic summarization of
scienti c publications [
        <xref ref-type="bibr" rid="ref6 ref8">8, 6</xref>
        ]. The overall aim of the shared task can be summarized
as follows: Given a referenced paper (RP) and a set of its citing papers (CPs),
create a (community) summary of the RP. The overall task is divided into the
following subtasks:
1a) For each citance, retrieve the RP text span to which the citation refers;
1b) Assign to every citation one or more discourse facets (Method, Aim, Result,
      </p>
      <p>
        Implication, and Hypothesis ), based on the retrieved RP text (result of 1a);
2) Summarize (max. 250 words) the RP, using RP spans retrieved for all citances
(with assigned discourse facets), i.e., using the results from 1a) and 1b).
Similar to most systems from previous task editions, we frame (1a) as an
information retrieval (IR) task. Given the citance, we rank all RP sentences according
to their relevance for the citance. We train a learning to rank (L2R) model with
features indicating lexical overlap and semantic similarity between sentences. We
then augment the top-ranked RP sentence with its adjacent RP sentences, if they
also appear high in the L2R model's ranking. For the discourse facet classi cation
task (subtask 1b), we train one binary classi er for each label. We experimented
with Support Vector Machines (SVM) [
        <xref ref-type="bibr" rid="ref22">22</xref>
        ] and Convolutional Neural Network
(CNN) [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ] as learning models. Finally, we generate the summary of the RP
by (1) clustering the RP segments retrieved for individual citances according to
their semantic textual similarity, and (2) selecting the most informative sentence
from each cluster, according to the TextRank score [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ]. The o cial shared task
evaluation results show that our system performs favorably to competing systems
in terms of quality of the produced summaries.
2
      </p>
    </sec>
    <sec id="sec-2">
      <title>Related Work</title>
      <p>
        Here, we brie y discuss the best performing systems from the previous editions
of the CL-SciSumm shared task [
        <xref ref-type="bibr" rid="ref7 ref8">7, 8</xref>
        ].
      </p>
      <p>
        Moraes et al. [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ] propose two methods for detecting RP spans corresponding
to citances: (1) the cosine similarity between the citance and RP candidate text's
sparse TF-IDF vectors (2) SVM with tree kernels. Surprisingly, the simple cosine
similarity between bag-of-words (BoW) vectors performed better, but the authors
still summarized based on the SVM tree kernel ranking.
      </p>
      <p>
        Li et al. [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ] combine lexical overlap and semantic similarity scores (e.g.,
bagof-words similarity, unigram overlap, and cosine between word2vec vectors) in a
rule-based fashion to select the RP text spans for citances. For summarization,
the authors cluster the candidate sentences using hierarchical LDA and compute
many features to select cluster representatives for the summary.
      </p>
      <p>
        On the other hand, Conroy et al.'s [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ] summarization method is based on a
vector space model, in which they use term frequency and nonnegative matrix
factorization to obtain term weights which they then use to create a summary.
3
      </p>
    </sec>
    <sec id="sec-3">
      <title>Methodology</title>
      <p>In this Section, we provide methodological details of the approaches we used for
solving di erent subtasks of the shared task.</p>
      <sec id="sec-3-1">
        <title>Task 1a: Retrieval of Referenced Text Spans</title>
        <p>We cast the identi cation of the referenced RP text span for the CP citance as
an IR task, divided into two steps:
1. Ranking RP sentences according to their relevance for the citance;
2. Selecting the sentences for the RP span, based on the above ranking.
Ranking of Candidate Sentences. We resort to the supervised L2R paradigm.
Concretely, we train the Coordinate Ascent model optimizing the mean average
precision (MAP) from the RankLib library3, with the following features:
Lexical similarity features. Two features capture the lexical overlap between
an RP sentence and a CP citance: (1) vector space similarity (VSS) is the cosine
between TF-IDF-weighted BoW vectors; (2) unigram overlap (UO) is the Jaccard
coe cient computed over term sets of the RP sentence and the citance.
Semantic similarity features. An RP sentence can be semantically similar to
the citance, but with little or no lexical overlap. We exploit word embeddings (i.e.,
semantic word vectors) to compute two measures of semantic textual similarity:
Aggregate sentence embedding similarity (AGG) is the cosine between the
aggregate sentence embeddings. The aggregate embedding vector of a sentence is
obtained simply as the weighted average (with TF-IDF scores of terms as weights)
of embeddings of the terms that the sentence contains.</p>
        <p>
          Word mover's similarity (or distance, WMS) [
          <xref ref-type="bibr" rid="ref9">9</xref>
          ] is the measure of semantic
similarity that aims to compute the maximal similarity (i.e., minimal distance)
in meaning between two texts. Let A 2 R2 be the matrix in which rows denote
set of unique tokens S of some RP sentence s, and columns represent set of
distinct tokens C of the citance c. The WMD score is then the solution to the
optimization problem
subject to the constraints
        </p>
        <p>jSj jCj
WMD (s; c) = max X X Ai;j sim(wi; wj0);</p>
        <p>
          A i21 j21
jCj
X Ai;j = freq(wi ; s); 8i 2 f1; : : : jSjg; and
j=1
jSj
X Ai;j = freq(wj0; c); 8j 2 f1; : : : jCjg;
i=1
3 Online available at: https://sourceforge.net/p/lemur/wiki/RankLib/.
where sim(wi; wj0) is the cosine similarity of embedding vectors of words wi and
wj0 and freq(w; s) is the frequency with which the word w appears in sentence s.
We compute both of the above features (AGG and WMS) using two di erent sets
of word embedding vectors. We experiment with (1) 300-dimensional Skip-Gram
embeddings [
          <xref ref-type="bibr" rid="ref15">15</xref>
          ], pre-trained on the Google News dataset4 and (2) 300-dimensional
domain-speci c embeddings obtained by running the CBOW model [
          <xref ref-type="bibr" rid="ref15">15</xref>
          ] on the
ACL Reference Corpus [
          <xref ref-type="bibr" rid="ref2">2</xref>
          ].
        </p>
        <p>
          Entity-based features. We run the TagMe entity linker [
          <xref ref-type="bibr" rid="ref5">5</xref>
          ] over the citances
and the RP candidate sentences to link mentions to Wikipedia concepts. We
compute the entity overlap (EO) feature as the Jaccard coe cient over sets of
linked entities from the citance and the RP sentence. We add a binary feature
indicating whether the RP sentence contains any linked entities.
Positional features. We compute the relative sentence position (absolute
position normalized by the document length) for the candidate RP sentence
in the RP and for the citance in the CP. Similarly, for both the RP candidate
within the RP and citance within the CP we extract the relative section positions
(section number of the sentence divided by the total number of sections). Finally,
we compute the ratio between relative positions of RP sentence and CP citance.
Adjacency-Based Postprocessing. Our L2R model ranks individual RP
sentences. Although most often the relevant RP texts of citances have one sentence,
reasonably often they also contain two or more sentences. To account for such
cases, we perform a postprocessing step where we decide whether to add additional
sentences to the output. We evaluated three postprocessing strategies:
1. Top-rank returns the top-ranked sentence from the L2R model's ranking;
2. Top-K neighbours extends the output with RP sentences adjacent to the
top-ranked sentence if these are found within the K top-ranked sentences in
the L2R model's ranking;
3. Iterative Top-K neighbours extends the Top-K neighbours by repeatedly
adding adjacent sentences of those already in the output if adjacent sentences
are among the top K in the ranking.
3.2
        </p>
      </sec>
      <sec id="sec-3-2">
        <title>Task 1b: Discourse Facet Classi cation</title>
        <p>The second subtask (subtask 1b), discourse facet classi cation, is a multi-label
classi cation task. Each RP text span retrieved as relevant for a citance needs
to be annotated with appropriate discourse facet labels. Since the retrieved
RP text snippet may be labeled with more than one discourse facet, we train
one binary classi er for each discourse facet label. For each of the ve binary
classi cation tasks, we experimented with two supervised machine learning models:
Convolutional Neural Networks (CNN) and Support Vector Machines (SVM).
4 Available at https://drive.google.com/file/d/0B7XkCwpI5KDYNlNUTTlSS21pQmM/
edit?usp=sharing.</p>
        <p>
          CNNs [
          <xref ref-type="bibr" rid="ref11">11</xref>
          ], rst applied on NLP tasks by Collobert and Weston [
          <xref ref-type="bibr" rid="ref3">3</xref>
          ], have
been shown to be successful on a range of short text classi cation task (see
[
          <xref ref-type="bibr" rid="ref20 ref21">21, 20</xref>
          ], inter alia). The architecture of the CNN we apply consists of a single
convolutional layer, followed by a single max-pooling layer. We use recti ed linear
unit (ReLU) as the non-linear activation function. In addition to the vanilla CNN
model that makes predictions based purely on the retrieved RP text span, we
also evaluate a CNN variant in which we introduce hand-crafted features that
are, for each instance, concatenated to the latent CNN features (i.e., the output
of the max-pooling layer) and fed to the last (feed-forward) layer of the network,
which makes the nal label prediction. Let xCNN be the latent CNN vector for
some input example, and xHF be the vector of hand-crafted features for the same
RP text span instance. The output vector y (a probability distribution over the
two labels in binary classi cation tasks) is then computed as follows:
y = softmax (W
        </p>
        <p>
          (xCNN kxHF ) + b) ;
where W and b are the weights matrix and biases vector of a feed-forward network
with a single hidden layer and linear activation. For the CNN classi cation
tasks we represent the input tokens with 300-dimensional domain-speci c word
embeddings, trained on the ACL Reference Corpus [
          <xref ref-type="bibr" rid="ref2">2</xref>
          ] (see Section 3.1).
        </p>
        <p>
          We also experimented with binary SVM [
          <xref ref-type="bibr" rid="ref22">22</xref>
          ] classi ers, employing the following
set of hand-crafted features (all features except lexical are also used as additional
hand-crafted features for the hybrid CNN model):
Lexical features. The sparse TF-IDF weighted BoW vector of the RP span;
Positional features. The relative sentence position and the relative section position
of the retrieved RP text span;
Other features. Two binary features indicating whether the retrieved RP span (1)
contains numbers and (2) consists of multiple sentences.
3.3
        </p>
      </sec>
      <sec id="sec-3-3">
        <title>Task 2: Citation-Based Summarization</title>
        <p>
          Finally, we create the RP summary by exploiting the output of the rst task {
the retrieved RP text spans for all citances. To build a non-redundant summary
re ecting the most important aspects of the RP, we propose the following
rulebased approach:
1. We cluster the RP text spans (retrieved for all of the citances) using the
simple single-pass clustering algorithm employing word mover's similarity
(cf. Section 3.1) as the similarity score between di erent RP text spans;
2. In order to select the most informative sentences for the summary, we compute
the TextRank score [
          <xref ref-type="bibr" rid="ref5">5</xref>
          ] for each retrieved RP span and order the RP spans
within clusters according to their TextRank scores;
3. We rank the clusters according to the average TextRank scores of the RP
text spans they contain. We then rst select for the summary the most
informative text span from the most informative cluster, then the most
informative text span from the second most informative cluster, etc., until
we reach the summary limit of 250 words.
        </p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>Evaluation</title>
      <p>We rst describe the datasets we used to train and optimize our models. Next, we
explain the evaluation setting and describe di erent con gurations we submitted
for the nal evaluation.
4.1</p>
      <sec id="sec-4-1">
        <title>Dataset</title>
        <p>
          The training set provided by the shared task organizers5 is an annotated subset
of the ACL Anthology [
          <xref ref-type="bibr" rid="ref18 ref19">19, 18</xref>
          ] consisting of 30 topics, each of which consists of
one referenced paper (RP) and its corresponding citing papers (CPs). In total,
the training set consists of 594 instances, i.e., citances paired (i.e., annotated)
with relevant references text spans (indicated as sentence o sets in the RP) and
the corresponding discourse facet labels.
4.2
        </p>
      </sec>
      <sec id="sec-4-2">
        <title>Evaluation Setting</title>
        <p>For subtask (1a), i.e., retrieval of relevant RP spans for given citances, we
evaluated di erent model variants via the 10-folded cross validation (CV) on the
training set. Positive instances for the L2R model are given directly as citance{RP
text span pairs. For the negative instances, one could couple the citance with any
other portion of text from the RP. In order to prevent excessive skewness of the
training set in favor of the negative examples, we sampled 10 negative instances
for each positive instance, picking both sentences adjacent to the relevant RP
text span and RP sentences from other article sections. We optimized the L2R
model for mean average precision (MAP), i.e., we searched (via greedy feature
selection) for the combination of features with the largest MAP performance.</p>
        <p>
          To evaluate the e ects of di erent postprocessing strategies (see Section 3.1),
we ran the evaluation script provided by the task organizers on outputs produced
by di erent RP span selection models (we used the gold discourse facet labels in
this case). We estimated our performance on the discourse facet classi cation task
in a 5-fold CV setting in terms of precision, recall and F1-score, micro-averaged
over the folds. Finally, our system's output for the nal summarization task was
evaluated, also in CV setting on the train set, in terms of the ROUGE-2 score
against three types of gold summaries { RP abstract, expert human summary,
and community summary [
          <xref ref-type="bibr" rid="ref13">13</xref>
          ]. We experimented with di erent WMS thresholds
for the single-pass clustering. We omit the CV performance on the train set due
to space constraints.
4.3
        </p>
      </sec>
      <sec id="sec-4-3">
        <title>Submitted Runs</title>
        <p>In total, we submitted 9 di erent runs, composed as follows:
5 Online available at https://github.com/WING-NUS/scisumm-corpus/tree/master/
data/Training-Set-2017.
Task (1a). According to CV evaluation on the training set, the best feature
combination consisted of: unigram overlap (UO), vector space similarity (VSS),
aggregated embedding similarity using domain-speci c word embeddings
(AGGACL), and WMS based on the domain-speci c word embeddings (WMS-ACL). We
coupled the predictions of the L2R model trained with this feature combination
with three post-processing strategies: top-rank, top-5 neighbours, and top-10
neighbours. This gave us three runs for retrieving relevant RP spans.
Task (1b). For the discourse facet classi cation we considered three variants:
(1) predict with SVM classi ers for all ve facet labels, (2) predict with CNN
classi ers for all ve labels, and (3) for each label, predict with the classi er that
yielded best results in the CV setting on the training set. These three variants,
combined with three retrieval variants for (1a) resulted in total of nine runs for
tasks 1a and 1b together.</p>
        <p>Task (2). CV experiments on the training set suggested the value of 0:85 to be
the optimal WMS threshold for the single-pass clustering. Using only the results
of subtask (1a) for summarization (i.e., our summarization algorithm does not
use discourse facets), we submit three summaries for each topic.
4.4</p>
      </sec>
      <sec id="sec-4-4">
        <title>Final Results</title>
        <p>
          The nal evaluation of the submissions was performed by the organizers of the
shared task. Here, we report the results of our submitted runs as well as the
average and winning scores across all participants. For more information please
refer to the overview paper of the shared task [
          <xref ref-type="bibr" rid="ref6">6</xref>
          ].
        </p>
        <p>The results of the referenced text span identi cation task (task 1a) are listed
in table 1. Just picking the top-ranked sentence outputted by our L2R model is
numerically above the average performance across all submissions. Our best result
is reached using the top-5 neighbors postprocessing, i.e., by searching among the
top-5 ranked candidate sentences for neighbors of the top-ranked sentence. Due
to the very limited size of the test set, the performance di erences are most likely
not statistically signi cant.6</p>
        <p>Table 2 shows the results of task (1b), the discourse facet classi cation. The
results heavily depend on the output of task 1a, thus it is not surprising that
the classi cations produced on top of the sentences retrieved using our L2R
model with the top-5 neighbors postprocessing strategy exhibit best for all three
classi cation strategies applied. The best score was achieved by applying only
SVM classi ers. The limited size of the provided training data is the most likely
explanation for the CNN classi ers performing worse than the SVM classi ers.</p>
        <p>The evaluation of the nal output summaries performed by the organizers
(see table 3) shows that our approach performs best compared to those of the
other participants in terms of ROUGE-SU4 F1 score when compared against the
community and the abstract gold summaries. Moreover, for all three variants of
6 The shared task organizers provided no information on statistical signi cance of the
performance di erences between submissions.
the gold summaries { human expert summary, author abstract, and community
summary { our approaches reach the highest precision. These results suggest that
the errors in identifying the referenced text spans (task 1a) do not necessarily
propagate to summary composition, which, in turn, suggests that referenced RP
texts might not be the best source of text for constructing the summaries.
5</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>Conclusion</title>
      <p>In this work, we presented a combination of methods for citation-based
semantic analysis and summarization of scienti c publications. For the retrieval of
referenced text spans we employed a supervised learning to rank model with a
number of features capturing semantic textual similarity between the citation
context and reference paper sentences. Next, we experimented with SVM and
CNN classi ers for discourse facet classi cation. Finally, we proposed a simple
summarization approach based on clustering of the referenced sentences, again
by exploiting measures of semantic textual similarity. The o cial evaluation of
the automatically created publication summaries shows that our system produces
higher quality than competing systems in several evaluation settings.</p>
    </sec>
    <sec id="sec-6">
      <title>Acknowledgements</title>
      <p>This research was partly funded by the German Research Foundation (DFG),
grant number EC 477/5-1 (LOC-DB). We thank the NVIDIA Corporation for
donating the GeForce Titan X GPU used to carry out some of our experiments.</p>
      <p>Lauscher et al.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Abu-Jbara</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ezra</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Radev</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          :
          <article-title>Purpose and polarity of citation: Towards nlp-based bibliometrics</article-title>
          . In:
          <article-title>Proceedings of the Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (NAACL-HTL)</article-title>
          . pp.
          <volume>596</volume>
          {
          <fpage>606</fpage>
          .
          <string-name>
            <surname>ACL</surname>
          </string-name>
          (
          <year>2013</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Bird</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Dale</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Dorr</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gibson</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Joseph</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kan</surname>
            ,
            <given-names>M.Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lee</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Powley</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Radev</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Tan</surname>
            ,
            <given-names>Y.F.</given-names>
          </string-name>
          :
          <article-title>The acl anthology reference corpus: A reference dataset for bibliographic research in computational linguistics</article-title>
          .
          <source>In: Proceedings of the Sixth International Conference on Language Resources and Evaluation (LREC)</source>
          . pp.
          <volume>1755</volume>
          {
          <issue>1759</issue>
          (
          <year>2008</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Collobert</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Weston</surname>
            ,
            <given-names>J.:</given-names>
          </string-name>
          <article-title>A uni ed architecture for natural language processing: deep neural networks with multitask learning</article-title>
          . In: Cohen,
          <string-name>
            <given-names>W.W.</given-names>
            ,
            <surname>McCallum</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            ,
            <surname>Roweis</surname>
          </string-name>
          , S.T. (eds.)
          <source>Proceedings of the 25th International Conference on Machine Learning (ICML)</source>
          . pp.
          <volume>160</volume>
          {
          <issue>167</issue>
          (
          <year>2008</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Conroy</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Davis</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          :
          <article-title>Vector space and language models for scienti c document summarization</article-title>
          .
          <source>In: Proceedings of NAACL-HLT</source>
          . pp.
          <volume>186</volume>
          {
          <issue>191</issue>
          (
          <year>2015</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Ferragina</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Scaiella</surname>
            ,
            <given-names>U.</given-names>
          </string-name>
          : Tagme:
          <article-title>On-the- y annotation of short text fragments (by wikipedia entities)</article-title>
          .
          <source>In: Proceedings of the 19th ACM International Conference on Information and Knowledge Management (CIKM)</source>
          . pp.
          <volume>1625</volume>
          {
          <fpage>1628</fpage>
          .
          <string-name>
            <surname>ACM</surname>
          </string-name>
          , New York, NY, USA (
          <year>2010</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Jaidka</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chandrasekaran</surname>
            ,
            <given-names>M.K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Jain</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kan</surname>
          </string-name>
          , M.Y.:
          <article-title>Overview of the cl-scisumm 2017 shared task</article-title>
          .
          <source>In: Proceedings of the Joint Workshop on Bibliometric-enhanced Information Retrieval and Natural Language Processing for Digital Libraries</source>
          <year>2017</year>
          (
          <article-title>BIRNDL)</article-title>
          .
          <source>CEUR</source>
          , Tokyo, Japan (
          <year>2017</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>Jaidka</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chandrasekaran</surname>
            ,
            <given-names>M.K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Elizalde</surname>
            ,
            <given-names>B.F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Jha</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Jones</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kan</surname>
            ,
            <given-names>M.Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Khanna</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Radev</surname>
            ,
            <given-names>D.R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ronzano</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Saggion</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kim</surname>
            ,
            <given-names>W.:</given-names>
          </string-name>
          <article-title>The computational linguistics summarization pilot task (</article-title>
          <year>2014</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>Jaidka</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chandrasekaran</surname>
            ,
            <given-names>M.K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rustagi</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kan</surname>
          </string-name>
          , M.Y.:
          <article-title>Insights from cl-scisumm 2016: the faceted scienti c document summarization shared task</article-title>
          .
          <source>International Journal on Digital Libraries (Jun</source>
          <year>2017</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <surname>Kusner</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sun</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kolkin</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Weinberger</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          :
          <article-title>From word embeddings to document distances</article-title>
          . In: Bach,
          <string-name>
            <given-names>F.</given-names>
            ,
            <surname>Blei</surname>
          </string-name>
          ,
          <string-name>
            <surname>D</surname>
          </string-name>
          . (eds.)
          <source>Proceedings of ICML 2015</source>
          . vol.
          <volume>37</volume>
          , pp.
          <volume>957</volume>
          {
          <fpage>966</fpage>
          . PMLR, Lille,
          <source>France (07{09 Jul</source>
          <year>2015</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <surname>Lauscher</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Glavas</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ponzetto</surname>
            ,
            <given-names>S.P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Eckert</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          :
          <article-title>Investigating convolutional networks and domain-speci c embeddings for semantic classi cation of citations</article-title>
          .
          <source>In: Proceedings of the Workshop on Mining Scienti c Publications (WOSP '17)</source>
          . p. in press.
          <source>ACM</source>
          (
          <year>2017</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <string-name>
            <surname>LeCun</surname>
          </string-name>
          , Y.,
          <string-name>
            <surname>Bengio</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          :
          <article-title>The handbook of brain theory and neural networks</article-title>
          .
          <source>chap. Convolutional Networks for Images, Speech, and Time Series</source>
          , pp.
          <volume>255</volume>
          {
          <fpage>258</fpage>
          . MIT Press, Cambridge, MA, USA (
          <year>1998</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          12.
          <string-name>
            <surname>Li</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Mao</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zhang</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chi</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Huang</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Cong</surname>
            ,
            <given-names>X.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Peng</surname>
          </string-name>
          , H.:
          <article-title>Cist system for cl-scisumm 2016 shared task</article-title>
          . In: Cabanac,
          <string-name>
            <given-names>G.</given-names>
            ,
            <surname>Chandrasekaran</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.K.</given-names>
            ,
            <surname>Frommholz</surname>
          </string-name>
          ,
          <string-name>
            <given-names>I.</given-names>
            ,
            <surname>Jaidka</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            ,
            <surname>Kan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.Y.</given-names>
            ,
            <surname>Mayr</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            ,
            <surname>Wolfram</surname>
          </string-name>
          ,
          <string-name>
            <surname>D</surname>
          </string-name>
          . (eds.)
          <source>Proceedings of BIRNDL 2016. CEUR Workshop Proceedings</source>
          , vol.
          <volume>1610</volume>
          , pp.
          <volume>156</volume>
          {
          <fpage>167</fpage>
          .
          <string-name>
            <surname>CEUR-WS.org</surname>
          </string-name>
          (
          <year>2016</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          13.
          <string-name>
            <surname>Lin</surname>
            ,
            <given-names>C.Y.</given-names>
          </string-name>
          :
          <article-title>Rouge: A package for automatic evaluation of summaries</article-title>
          . In: MarieFrancine Moens, S.S. (ed.)
          <source>Text Summarization Branches Out: Proceedings of the ACL-04 Workshop</source>
          . pp.
          <volume>74</volume>
          {
          <fpage>81</fpage>
          . ACL, Barcelona,
          <source>Spain (July</source>
          <year>2004</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          14.
          <string-name>
            <surname>Mihalcea</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Tarau</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          : Textrank:
          <article-title>Bringing order into texts</article-title>
          . In: Lin,
          <string-name>
            <given-names>D.</given-names>
            ,
            <surname>Wu</surname>
          </string-name>
          ,
          <string-name>
            <surname>D</surname>
          </string-name>
          . (eds.)
          <source>Proceedings of the Conference on Empirical Methods in Natural Language Processing</source>
          <year>2004</year>
          (EMNLP). pp.
          <volume>404</volume>
          {
          <fpage>411</fpage>
          . ACL, Barcelona,
          <source>Spain (July</source>
          <year>2004</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          15.
          <string-name>
            <surname>Mikolov</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sutskever</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chen</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Corrado</surname>
            ,
            <given-names>G.S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Dean</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          :
          <article-title>Distributed representations of words and phrases and their compositionality</article-title>
          .
          <source>In: NIPS</source>
          , pp.
          <volume>3111</volume>
          {
          <fpage>3119</fpage>
          . Curran Associates, Inc. (
          <year>2013</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          16.
          <string-name>
            <surname>Moraes</surname>
            ,
            <given-names>L.F.T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Baki</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Verma</surname>
            ,
            <given-names>R.M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lee</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          : University of houston at clscisumm 2016:
          <article-title>Svms with tree kernels and sentence similarity</article-title>
          . In: Cabanac,
          <string-name>
            <given-names>G.</given-names>
            ,
            <surname>Chandrasekaran</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.K.</given-names>
            ,
            <surname>Frommholz</surname>
          </string-name>
          ,
          <string-name>
            <given-names>I.</given-names>
            ,
            <surname>Jaidka</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            ,
            <surname>Kan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.Y.</given-names>
            ,
            <surname>Mayr</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            ,
            <surname>Wolfram</surname>
          </string-name>
          ,
          <string-name>
            <surname>D</surname>
          </string-name>
          . (eds.)
          <source>Proceedings of BIRNDL 2016. CEUR Workshop Proceedings</source>
          , vol.
          <volume>1610</volume>
          , pp.
          <volume>113</volume>
          {
          <fpage>121</fpage>
          .
          <string-name>
            <surname>CEUR-WS.org</surname>
          </string-name>
          (
          <year>2016</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          17.
          <string-name>
            <surname>Nakov</surname>
            ,
            <given-names>P.I.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Schwartz</surname>
            ,
            <given-names>A.S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hearst</surname>
            ,
            <given-names>M.A.</given-names>
          </string-name>
          :
          <article-title>Citances: Citation sentences for semantic analysis of bioscience text</article-title>
          .
          <source>In: In Proceedings of the SIGIR04 workshop on Search and Discovery in Bioinformatics</source>
          (
          <year>2004</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          18.
          <string-name>
            <surname>Radev</surname>
            ,
            <given-names>D.R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Muthukrishnan</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Qazvinian</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          :
          <article-title>The ACL anthology network corpus</article-title>
          .
          <source>In: Proceedings, ACL Workshop on Natural Language Processing and Information Retrieval for Digital Libraries</source>
          . pp.
          <volume>54</volume>
          {
          <fpage>61</fpage>
          .
          <string-name>
            <surname>Singapore</surname>
          </string-name>
          (
          <year>2009</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          19.
          <string-name>
            <surname>Radev</surname>
            ,
            <given-names>D.R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Muthukrishnan</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Qazvinian</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Abu-Jbara</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          :
          <article-title>The acl anthology network corpus</article-title>
          .
          <source>Language, Resources and Evaluation</source>
          <volume>47</volume>
          (
          <issue>4</issue>
          ),
          <volume>919</volume>
          {944 (Dec
          <year>2013</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          20.
          <string-name>
            <surname>Severyn</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Moschitti</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          :
          <article-title>Twitter sentiment analysis with deep convolutional neural networks</article-title>
          .
          <source>In: Proceedings of the 38th International ACM SIGIR Conference on Research and Development in Information Retrieval</source>
          . pp.
          <volume>959</volume>
          {
          <fpage>962</fpage>
          . SIGIR '15,
          <string-name>
            <surname>ACM</surname>
          </string-name>
          , New York, NY, USA (
          <year>2015</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          21.
          <string-name>
            <surname>Shrestha</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sierra</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gonzalez</surname>
            ,
            <given-names>F.A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rosso</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Montes-y Gomez</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Solorio</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          :
          <article-title>Convolutional neural networks for authorship attribution of short texts</article-title>
          .
          <source>In: Proceedings of the 2017 Conference of the European Chapter of the Association of Computational Linguistics (EACL)</source>
          (
          <year>2017</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          22.
          <string-name>
            <surname>Vapnik</surname>
          </string-name>
          , V.:
          <source>Estimation of Dependences Based on Empirical Data</source>
          : Springer Series in Statistics (Springer Series in Statistics). Springer-Verlag New York, Inc., Secaucus, NJ, USA (
          <year>1982</year>
          )
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>