<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Identifying Referenced Text in Scienti c Publications by Summarisation and Classi cation Techniques</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Know-Center GmbH In eldgasse</string-name>
          <email>arexha@know-center.at</email>
          <email>rkern@know-center.at</email>
          <email>sklampfl@know-center.at</email>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Austria</string-name>
        </contrib>
      </contrib-group>
      <pub-date>
        <year>2016</year>
      </pub-date>
      <fpage>122</fpage>
      <lpage>131</lpage>
      <abstract>
        <p>This report describes our contribution to the 2nd Computational Linguistics Scienti c Document Summarization Shared Task (CLSciSumm 2016), which asked to identify the relevant text span in a reference paper that corresponds to a citation in another document that cites this paper. We developed three di erent approaches based on summarisation and classi cation techniques. First, we applied a modi ed version of an unsupervised summarisation technique, TextSentenceRank, to the reference document, which incorporates the similarity of sentences to the citation on a textual level. Second, we employed classi cation to select from candidates previously extracted through the original TextSentenceRank algorithm. Third, we used unsupervised summarisation of the relevant sub-part of the document that was previously selected in a supervised manner.</p>
      </abstract>
      <kwd-group>
        <kwd>text summarisation</kwd>
        <kwd>key sentence extraction</kwd>
        <kwd>citation analysis</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>
        Extractive summarisation of a textual document is the process of nding a
representative subset of the document text that captures as much information about
the original document as possible. A promising idea in the realm of scienti c
publications is to consider the set of sentences that cite a paper as a summary
created by the research community. Here we describe our contribution to the
2nd Computational Linguistics Scienti c Document Summarization Shared Task
(CL-SciSumm 2016)1 [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ], which aims at exploring and encouraging novel
techniques for scienti c paper summarisation along this direction. This task takes
place at the Joint Workshop on Bibliometric-enhanced Information Retrieval
and Natural Language Processing for Digital Libraries (BIRNDL 2016)2 [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ] at
the Joint Conference on Digital Libraries (JCDL '16) and is a follow-up on the
2
      </p>
      <p>
        Stefan Klamp , Andi Rexha, and Roman Kern
CL Pilot Task that has been conducted as a part of the BiomedSumm Track at
the Text Analysis Conference 2014 (TAC 2014)3 [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ].
      </p>
      <p>The dataset provided for this year's task consists of a set of reference papers
(RP), each of which is accompanied by a set of citing papers (CP). The goal is to
identify the text span in the RP which corresponds to the citations in CP (Task
1A) as well as the discourse facet of the RP this text span belongs to (Task 1B).
Human annotators have created the ground truth in the form of pairs of citation
text and cited text on the granularity level of sentences.</p>
      <p>For Task 1A, we implemented a number of approaches that employ both
unsupervised and supervised techniques that di er in the way how information
from the citing sentence in the CP is incorporated into the process. In total, we
submitted three runs, corresponding to our three approaches for Task 1A: (i)
modi ed-tsr, (ii) tsr-sent-class, and (iii) sect-class-tsr.</p>
      <p>
        First, in a completely unsupervised setting, we applied a modi ed variant of
our TextSentenceRank algorithm to the RP (modi ed-tsr ). TextSentenceRank
[
        <xref ref-type="bibr" rid="ref9">9</xref>
        ] is a graph based ranking algorithm, a re nement of the well-known TextRank
algorithm, which is applied to text in order to extract key words and/or key
sentences. For the task at hand, we investigated a speci c weighting that takes
into account the information provided by the CP.
      </p>
      <p>As a second approach for Task 1A, we employed a supervised classi cation
setting following an unsupervised preprocessing (tsr-sent-class ). Through the
original version of TextSentenceRank we pre-selected candidate sentences from
the RP, independently from the CP, that are potential text spans for being cited.
For a given citation in the CP, we then selected the corresponding candidate
through supervised classi cation.</p>
      <p>In our third option, we took a dual approach (sect-class-tsr ). We rst used
supervised learning to identify the relevant sub-part (section) of the RP that
corresponds to the citing sentence in the CP. Once the relevant section has
been found, we used the original version of TextSentenceRank on this sub-part,
independent from the CP, to identify referenced text spans.</p>
      <p>For Task 1B, we used a similar classi er as for identi cation of the section.
We used features derived from the citing sentence as well es from the extracted
text span in the reference document to determine the discourse facet. This is
applied in all three approaches to Task 1A.</p>
      <p>In principle, TextSentenceRank is able to extract multiple candidate text
spans scattered across the document, but since the task description required the
extraction of consecutive referenced text, we decided to output a single sentence
as the extracted reference span in all of our approaches.</p>
      <p>This report is structured as follows. In sections 2 and 3 we explain our
approaches for both Task 1A and Task 1B in detail. In section 4 we individually
evaluate the classi ers that we used in our approaches and present our results
for the overall task. In the end, in section 5 we conclude, discuss our ndings,
and give an outlook to potential future work.
3 http://www.nist.gov/tac/2014/BiomedSumm/</p>
    </sec>
    <sec id="sec-2">
      <title>Task 1A: Identi cation of the Referenced Text Spans</title>
      <p>In this subsection we detail our three approaches for Task 1A, the identi cation
of referenced text spans. We also introduce TextSentenceRank as a base method
which is used in all three runs.
2.1</p>
      <sec id="sec-2-1">
        <title>TextSentenceRank as a Method for Extracting Candidate Text</title>
      </sec>
      <sec id="sec-2-2">
        <title>Spans</title>
        <p>
          This approach is inspired by graph based ranking algorithms, such as Google's
PageRank [
          <xref ref-type="bibr" rid="ref1">1</xref>
          ], where vertices in a graph are ranked based on their importance
given by the connectedness within the graph: the more likely a vertex is visited
by random walk, the higher is its score in the ranking. The TextSentenceRank
algorithm [
          <xref ref-type="bibr" rid="ref9">9</xref>
          ] is an application of such a graph based ranking method to natural
language text, returning a list of relevant key terms and/or key sentences ordered
by descending scores. It builds a graph where vertices correspond to sentences
or tokens and connects them with edges weighted according to their similarity
(textual similarity between sentences or textual distance between words).
Furthermore, each sentence vertex is connected to the vertices corresponding to the
tokens it contains4.
        </p>
        <p>The score s(vi) of a vertex vi is given by
s(vi) = (1
d) + d X
j2I(vi)</p>
        <p>
          wij
P
vk2O(vj) wjk
s(vj );
(1)
where d is a parameter accounting for latent transitions between non-adjacent
vertices (set to 0.85 as in [
          <xref ref-type="bibr" rid="ref1 ref7">1, 7</xref>
          ]), wij is the edge weight between vi and vj , and
I(v) and O(v) are the predecessors and successors of v, respectively. The scores
can be obtained algorithmically by an iterative procedure, or alternatively by
solving an eigenvalue problem on the weighted adjacency matrix.
        </p>
        <p>
          The TextSentenceRank algorithm is an extension to the original TextRank
algorithm [
          <xref ref-type="bibr" rid="ref7">7</xref>
          ], which could either compute key terms or key sentences. It has
been shown in [
          <xref ref-type="bibr" rid="ref9">9</xref>
          ] that by computing both key terms and key sentences at the
same time, the performance of key term extraction can be improved.
        </p>
        <p>When used for key sentence extraction, TextSentenceRank extracts the \most
relevant" sentences in terms of how they are connected to other sentences in the
document via co-occurring words. Here we pursue our intuition that such relevant
sentences are also more likely to be cited and use TextSentenceRank as a base
algorithm for extracting candidates for referenced text.
2.2</p>
      </sec>
      <sec id="sec-2-3">
        <title>Run 1: A Modi ed Version of TextSentenceRank</title>
        <p>Run 1 (modi ed-tsr ) follows a completely unsupervised setting, where we applied
a modi ed variant of our TextSentenceRank algorithm to the RP. The idea here
4 Here, we restrict the set of relevant tokens to adjectives, nouns, and proper nouns.</p>
        <p>This information is obtained through part-of-speech tagging.
4</p>
        <p>Stefan Klamp , Andi Rexha, and Roman Kern
was to use a speci c weighting of the underlying graph that takes into account
the information provided by the CP, in particular, the citing sentence. More
precisely, the weight of an edge adjacent to a node corresponding to sentence S
is modi ed by
wnew = wold
[1 + sim(S; C)] ;
(2)
where sim(S; C) is a similarity measure between sentence S of the RP and citing
sentence C of the CP.</p>
        <p>
          In our approach we used the Jaccard similarity [
          <xref ref-type="bibr" rid="ref4">4</xref>
          ] on the sets of tokens
contained in the respective sentences. Di erent similarity measures are possible,
e.g., a measure capturing the semantic similarity of words in the sentences, but
were not applied in the scope of this task. From the resulting list of the most
relevant sentences in the RP, we selected the one with the largest similarity to
the citing sentence in the CP.
2.3
        </p>
      </sec>
      <sec id="sec-2-4">
        <title>Run 2: TextSentenceRank and Sentence Classi cation</title>
        <p>In Run 2 (tsr-sent-class ) we employed a supervised classi cation setting following
an unsupervised preprocessing. First, we pre-selected candidate sentences from
the RP through the original version of TextSentenceRank. This selection is thus
independent from the CP and consists of potential text spans for being cited.
We then selected the corresponding candidate through supervised classi cation
that takes into account the information from the citing sentence in the CP.</p>
        <p>
          We used a Random Forest classi er [
          <xref ref-type="bibr" rid="ref3">3</xref>
          ], an ensemble method based on decision
trees, with the following features:
{ Section features: title and number of the section enclosing the candidate
sentence,
{ Sentence position features: relative positions of the candidate sentence
within the RP and within the enclosing section,
{ Discriminative term features: information about tokens shared between
the candidate sentence in the RP and the citing sentence in the CP.
This classi er is a binary classi er that decides for each candidate sentence in
the RP whether it is an actual referenced text span, based on information from
the citing sentence. For training the classi er we used the information provided
by the training set. Because of the unbalanced nature of positive and negative
training examples we used TextSentenceRank also to pre-select the sentences for
training.
2.4
        </p>
      </sec>
      <sec id="sec-2-5">
        <title>Run 3: Section classi cation and TextSentenceRank</title>
        <p>In Run 3 (sect-class-tsr ), we took a dual approach to Run 2. Instead of applying
an unsupervised preprocessing followed by a supervised classi cation, we rst
used supervised learning to identify the relevant sub-part of the RP that
corresponds to the citing sentence in the CP. This sub-part can in principle be of
any granularity, but as sections are annotated in the provided dataset, we chose
Identifying Referenced Text by Summarisation and Classi cation
5
the granularity of sections. Once the relevant section has been found, we used
the original version of TextSentenceRank on this sub-part, independent from the
CP, to identify referenced text spans.</p>
        <p>
          Again, we used a Random Forest classi er, now with these features:
{ Section features: title and number of the section enclosing the candidate
sentence,
{ Tf-Idf features: information about the frequency of tokens from the citing
sentence within the section, normalized by the inverse frequency across all
the sections of the RP. This feature is motivated by the standard Tf-Idf
measure [
          <xref ref-type="bibr" rid="ref8">8</xref>
          ], applied to the sections of the RP. It emphasizes sections that
exclusively share tokens with the citing sentence.
        </p>
        <p>This classi er is a binary classi er that decides for each candidate section in the
RP whether it is a section containing a referenced text span, based on
information from the citing sentence. The cited text span is then selected through the
original version of TextSentenceRank, applied to the sub-document spanning the
selected section, independently from the CP.
3</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>Task 1B: Identi cation of the Discourse Facet</title>
      <p>The discourse facet takes the following values in the training set: Implication,
Method, Aim, Results, and Hypothesis. We used a Random Forest classi er with
section features, sentence position features, and discriminative term features to
distinguish between these classes. That is, we took into account information
from both the citing sentence as well es from the extracted text span in the
reference document to determine the discourse facet. The same model, which
was previously trained on the training set, was applied in all three runs.
4</p>
    </sec>
    <sec id="sec-4">
      <title>Evaluation</title>
      <p>In our evaluation, we rst determine the isolated performance of individual
components that we used in our approaches. Then we present our results for the
overall task. We performed these evaluations on both the provided development
set and the training set, as the results on the test corpus have not yet been made
available.
4.1</p>
      <sec id="sec-4-1">
        <title>Performance of individual classi ers</title>
        <p>In our contribution we used three di erent classi ers, which we evaluate
separately in this section. For the full system runs we trained the classi er on the
training set, and submitted the results produced by applying the trained
classiers on the test set. Here, we evaluate the classi ers on both the development
set and on the training set using 10-fold cross validation.</p>
        <p>In Run 2 of Task 1A (tsr-sent-class ) we used a binary classi er that decides
for each candidate sentence whether it serves as a referenced text span for a given
6</p>
        <p>Stefan Klamp , Andi Rexha, and Roman Kern
citing sentence. Table 1 shows the cross-validation performance of this sentence
classi er on both the development set and on the training set. Since it is a
binary classi er, we report here the precision, recall, and F1 values with respect
to the positive class, i.e., whether the sentence is classi ed as a referenced text
span. A reasonable precision is achieved, which means that there are relatively
few false positives, however, at the expense of low recall, indicating that many
true positives are missed. Even though we pre- ltered the negative instances with
TextSentenceRank, the classi cation problem is still quite unbalanced: There are
about 10 times as many negative as positive examples. This unbalance might
induce a certain bias in the classi er; still, the accuracy, i.e., the fraction of
correctly classi ed instances, across both classes is above 90% for both datasets.</p>
        <p>In Run 3 of Task 1A (sect-class-tsr ) we employed a binary classi er that
decides for each section whether it contains a referenced text span corresponding
to a given citing sentence. Table 2 shows the cross-validation performance of this
section classi er on both the development set and on the training set. Again,
we report here the precision, recall, and F1 values with respect to the positive
class, i.e., whether the section is classi ed as containing a referenced text span.
It can be seen that the performance is quite lower than for the sentence classi er,
but the accuracy is still above 80%. Here the unbalance between positive and
negative classes is given by the number of sections in the reference paper (on
average about 5 to 6 in the training set).</p>
        <p>In Task 1B, for the identi cation of the discourse facet, we used a multi-label
classi er to categorise a referenced text span into one of the following classes:
Identifying Referenced Text by Summarisation and Classi cation
7
Implication, Method, Aim, Results, and Hypothesis. Table 3 shows the
crossvalidation performance of this discourse facet classi er on both the development
set and on the training set. Precision and recall are given as micro-averages over
all labels and lie between 60% and 70%. Classi cation accuracy is around 70%
for both datasets. Table 4 shows the confusion matrix obtained by the classi er
on the training set. Method is by far the most occurring label in the datasets,
and the classi er might have a certain bias of generating this label, but also the
quality of retrieving the label Results is reasonable. Hypothesis is the rarest label
and the one with the lowest accuracy.
4.2</p>
      </sec>
      <sec id="sec-4-2">
        <title>Overall task performance</title>
        <p>We evaluated all of our three approaches for Task 1A on both the development
set and the training set. We compared the extracted reference spans in our
system output with the corresponding reference spans provided by the human
annotators in terms of overlap and distance. In Table 5 we show for each of the
three runs and for each topic in the development set the number of citances for
which the extracted reference span lies within 10 sentences of the true reference
span and for which both reference spans actually overlap. Table 6 shows the
same information for the training set.
8</p>
        <p>Stefan Klamp , Andi Rexha, and Roman Kern</p>
        <p>It can be seen that only in few cases a referenced text span is extracted
that overlaps the text segment identi ed by the human annotator. If we allow
a certain neighborhood around the true spans, considerably more matches are
found. Interestingly, Run 1, the modi ed TextSentenceRank, achieves the best
results, followed by Run 3, the variant with section classi cation, which slightly
outperforms Run 2, the version with sentence classi cation.</p>
        <p>The nding that the modi ed TextSentenceRank works best suggests that
considering the document as a whole might be bene cial for extracting relevant
key sentences. The low performance of the classi cation approaches might be due
to a lack of representative features that are relevant for the task at hand. In Run 2
it is possible that the set of sentences provided by the original TextSentenceRank
algorithm is already too limited before the sentence classi er can select suitable
text spans. The dual approach of Run 3, rst selecting a sub-part and then
applying TextSentenceRank, works slightly better.</p>
        <p>These results also demonstrate the di culty of this task. It is worth
mentioning that all our approaches are solely based on statistics of words and
sentences in both the reference document and the citing sentence as well as in their
comparison. Currently, we do not incorporate any semantic information, but in
principle our approaches can be easily adapted through additional features for
classi cation and a di erent weighting strategy for TextSentenceRank.
5</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>Discussion</title>
      <p>In this report we have described our contribution to the 2nd Computational
Linguistics Scienti c Document Summarization Shared Task (CL-SciSumm 2016),
Identifying Referenced Text by Summarisation and Classi cation
9
which asked participants to identify the relevant text span in a reference paper
that corresponds to a citation in another document that cites this paper. We
developed three di erent approaches based on summarisation and classi cation
techniques. They employ both unsupervised and supervised techniques that
differ in the way how information from the citing sentence is incorporated into the
process. First, we applied a modi ed version of an unsupervised summarisation
technique, TextSentenceRank, to the reference document, which incorporates
the similarity of sentences to the citation on a textual level. Second, we
employed classi cation to select from candidates previously extracted through the
original TextSentenceRank algorithm. Third, we used unsupervised
summarisation of the relevant sub-part of the document that was previously selected in a
supervised manner.</p>
      <p>We evaluated both the individual classi ers used in our approaches as well
as the performance in the overall task. We believe that the performance of our
systems could be improved by incorporating di erent similarity measures, e.g.,
measures capturing the semantic similarity of citing and cited sentences, not
only in the modi ed weighting of TextSentenceRank, but also into the set of
features used for classi cation. Furthermore, the relative strength of the in uence
of the citing sentence can be optimised. For example, in the modi ed
TextSentenceRank algorithm there could be a trade-o parameter in equation 2 that
weights the relative in uences of wold and sim(S; C). Finally, the inclusion of
multiple, non-consecutive sentences in the output would likely include candidate
text spans of better quality.</p>
      <p>Another aspect that likely in uences our system is the fact that the text of
the given documents was extracted with OCR methods. These methods
some</p>
      <p>Stefan Klamp , Andi Rexha, and Roman Kern
times yield noisy and erroneous words, and many of our methods rely on the
statistics of terms within and across documents. In our experience, in some cases
TextSentenceRank seems to prefer sentences containing such noisy or invalid
tokens.</p>
      <p>As a contribution to the CL-SciSumm 2016 task our work aimed at
facilitating the summarisation of scienti c publications. In addition, we hope that our
contribution will also provide further insight into the scienti c writing habits
of researchers, both in terms of how they structure their papers and how they
reference the work of others.</p>
    </sec>
    <sec id="sec-6">
      <title>Acknowledgements</title>
      <p>The Know-Center is funded within the Austrian COMET Program {
Competence Centers for Excellent Technologies { under the auspices of the Austrian
Federal Ministry of Transport, Innovation and Technology, the Austrian Federal
Ministry of Economy, Family and Youth and by the State of Styria. COMET is
managed by the Austrian Research Promotion Agency FFG.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Brin</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Page</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          :
          <article-title>The anatomy of a large-scale hypertextual web search engine</article-title>
          .
          <source>Computer Networks</source>
          <volume>56</volume>
          (
          <issue>18</issue>
          ),
          <volume>3825</volume>
          {
          <fpage>3833</fpage>
          (
          <year>2012</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Cabanac</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chandrasekaran</surname>
            ,
            <given-names>M.K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Frommholz</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Jaidka</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kan</surname>
            ,
            <given-names>M.Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Mayr</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Wolfram</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          : Joint Workshop on Bibliometric-enhanced
          <source>Information Retrieval and Natural Language Processing for Digital Libraries (BIRNDL</source>
          <year>2016</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Ho</surname>
          </string-name>
          , T.K.:
          <article-title>Random decision forests</article-title>
          .
          <source>In: Proceedings of the 3rd International Conference on Document Analysis and Recognition</source>
          . pp.
          <volume>278</volume>
          {
          <fpage>282</fpage>
          . Montreal, QC (
          <year>1995</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Jaccard</surname>
            ,
            <given-names>P.:</given-names>
          </string-name>
          <article-title>The distribution of the ora in the alpine zone</article-title>
          .
          <source>New Phytologist</source>
          <volume>11</volume>
          ,
          <issue>37</issue>
          {
          <fpage>50</fpage>
          (
          <year>1912</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Jaidka</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chandrasekaran</surname>
            ,
            <given-names>M.K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Elizalde</surname>
            ,
            <given-names>B.F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Jha</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Jones</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kan</surname>
            ,
            <given-names>M.Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Khanna</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Molla-Aliod</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Radev</surname>
            ,
            <given-names>D.R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ronzano</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          , et al.:
          <article-title>The computational linguistics summarization pilot task</article-title>
          .
          <source>In: Proceedings of TAC. Gaithersburg</source>
          , USA (
          <year>2014</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Jaidka</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chandrasekaran</surname>
            ,
            <given-names>M.K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rustagi</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kan</surname>
          </string-name>
          , M.Y.:
          <article-title>Overview of the 2nd Computational Linguistics Scienti c Document Summarization Shared Task (CL-SciSumm 2016)</article-title>
          .
          <source>In: Proceedings of the Joint Workshop on Bibliometricenhanced Information Retrieval and Natural Language Processing for Digital Libraries (BIRNDL</source>
          <year>2016</year>
          ). Newark, New Jersey, USA (
          <year>2016</year>
          ), to appear
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>Mihalcea</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Tarau</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          : Textrank:
          <article-title>Bringing order into texts</article-title>
          .
          <source>In: Conference on Empirical Methods in Natural Language Processing</source>
          . Barcelona,
          <string-name>
            <surname>Spain</surname>
          </string-name>
          (
          <year>2004</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>Salton</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>McGill</surname>
            ,
            <given-names>M.J.</given-names>
          </string-name>
          :
          <article-title>Introduction to Modern Information Retrieval. McGrawHill, Inc</article-title>
          ., New York, NY, USA (
          <year>1986</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <surname>Seifert</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ulbrich</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kern</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Granitzer</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          :
          <article-title>Text representation for e cient document annotation</article-title>
          .
          <source>Journal of Universal Computer Science</source>
          <volume>19</volume>
          (
          <issue>3</issue>
          ),
          <volume>383</volume>
          {
          <fpage>405</fpage>
          (
          <year>2013</year>
          )
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>