<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Recognizing reference spans and classifying their discourse facets</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Kun Lu</string-name>
          <email>kunlu@ou.edu</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Gang Li</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Jin Mao</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Jian Xu</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>School of Information Management Wuhan University</institution>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>School of Information University of Arizona</institution>
          ,
          <country country="US">USA</country>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>School of Library and Information Studies University of Oklahoma</institution>
          ,
          <country country="US">USA</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2016</year>
      </pub-date>
      <fpage>139</fpage>
      <lpage>145</lpage>
      <abstract>
        <p>In this shared task, we applied “Learning to Rank” algorithm with multiple features, including lexical features, topic features, knowledge-based features and sentence importance, to Task 1A by regarding reference span finding as an information retrieval problem. Task 1B, discourse facet identifying, is treated as a text classification problem by considering features of both citation contexts and cited spans.</p>
      </abstract>
      <kwd-group>
        <kwd>Learning to rank</kwd>
        <kwd>Topic model</kwd>
        <kwd>Facet classification</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Introduction</title>
      <p>
        The 2nd CL-SciSumm Shared task follows the TAC 2014 Biomedical Summarization
Track on scientific paper summarization. An overview of the shared task, including
specific details on the dataset, the competitive results and subsequent analyses for each
task can be found in the shared task overview paper[
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. In this report, we provide a
detailed description of the methods we used for the Task 1A and Task 1B. Our methods
are introduced in Section 2 followed by results in Section 3. Some conclusions are
presented in the last section.
2.1
      </p>
    </sec>
    <sec id="sec-2">
      <title>Methods</title>
      <p>Task 1A
We considered Task 1A as an information retrieval problem. A citance (citation
context) is regarded as a query and sentences from reference paper (cited spans) are treated
as candidate documents. Then, the problem becomes how to rank the sentences of the
reference paper (i.e., candidate cited spans in this report) for a given citance. The most
relevant sentences of the reference paper to a citation context are selected as the golden
sentences of the citation context. We apply “learning to rank” algorithms to address this
problem and exploit multiple features. The explored features are as follows:
• Lexical Features. Bag of words is a widely used text representation method. By
representing citation contexts and candidate cited spans with the bag of words model,
the lexical similarities between citation contexts and candidate cited spans can be
obtained. We chose four candidate lexical similarity features including Cosine
similarity, Jaccard similarity, Dice similarity and LCS (longest common subsequence).
When computing cosine similarity, TFIDF term weighing was applied. In the
formula below, TF is the number of times a term occurring in a given sentence, while
the IDF value of a term is computed from all the 437 papers in the training set (93
papers), development set (155
papers), and test set(229
papers). Thus,


curs.</p>
      <p>__
_
      is 437,</p>
      <p>
        which is the same for all terms. And
) is the number of the different papers that a term
ocTFIDF(Term) = TF(Term) ∗ ln


 
   
(   )
(1)
• Topic Features. Bag of words model is insufficient in handling polysemy and
synonym problems. Taking topics into account can relieve this problem. Topic
modeling method[
        <xref ref-type="bibr" rid="ref2">2</xref>
        ] was used to identify latent topics from 10,921 articles of the ACL
Anthology Reference Corpus. The topic distributions of both citation contexts and
candidate cited spans were predicted through the LDA models. Cosine similarity was
then used to measure their topic similarities.
• Knowledge Based Features. WordNet[
        <xref ref-type="bibr" rid="ref3">3</xref>
        ] was used to compute the concept
similarity between citation contexts and candidate cited spans. Lin similarity[
        <xref ref-type="bibr" rid="ref4">4</xref>
        ] that
measures the similarity of two words was applied. The similarity of two sentences
can be compute by cumulating the similarities between their words. We used
different combinations of nouns and verbs to calculate similarities: WordNet (N)-only
using nouns, WordNet (V)-only using verbs, WordNet (N,V)-using both nouns and
verbs, and WordNet (N,Vsep) which is obtained by (WordNet(N)+ WordNet(V))/2.
• Sentence Importance. The importance of candidate cited spans in the reference
paper is considered as a factor influencing whether they are being cited or not.
TextRank[
        <xref ref-type="bibr" rid="ref5">5</xref>
        ], which is a widely used unsupervised method to extract the keyword or rank
the sentences of a given document, was applied to measure the importance of a
sentence in the reference paper. The assumption is that the more important a sentence
is in the reference paper, the more likely the sentence belongs to the cited span.
Each pair of a citation context and a candidate sentence from the cited article is an
instance. Positive instances (i.e., the candidate cited sentence is one of the gold standard
cited sentences) were assigned higher scores than negative instances. The above
features were calculated and fed into “learning to rank” methods. For topic similarity, the
number of topics varied from 20 to 200 with a step of 20. The features showing high
performance were selected as the final set of features. Then, five learning to rank
algorithms from RankLib[
        <xref ref-type="bibr" rid="ref6">6</xref>
        ], including RankBoost7, RankNet[
        <xref ref-type="bibr" rid="ref8">8</xref>
        ], AdaRank[
        <xref ref-type="bibr" rid="ref10">9</xref>
        ], and
Coordinate Ascent[
        <xref ref-type="bibr" rid="ref11">10</xref>
        ], were compared.
2.2
      </p>
      <p>
        Task 1B
Task 1B is considered as a text classification problem. The five discourse facets of a
sentence in a reference paper are Aim, Method, Result, Implication, and Hypothesis.
The features adopted by the classifiers are as follows.
• The text of the reference sentence
• The title of the section that the reference sentence belongs to
• The section type of the section that the reference sentence belongs to
• The text of the citation context
• The title of the section that the citation context belongs to
• The section type of the section that the citation context belongs to
Three classifiers from Weka[
        <xref ref-type="bibr" rid="ref13">11</xref>
        ] including Naïve Bayes, Decision Tree and Supporting
Vector Machine were applied and compared for Task 1B. The algorithm behind these
classifiers are Naive Bayesian classification, C4.5, and Sequential Minimal
Optimization. Default parameter settings were used to train the classifiers implemented in Weka.
2.3
      </p>
      <p>Evaluation</p>
      <sec id="sec-2-1">
        <title>2.3.1 Task 1A metrics.</title>
        <p>Two groups of precision (P), recall (R) and F_1 measures were used to evaluate the
performance in Task 1A. The first group counted the number of sentences returned by
our methods that match the gold standard sentences annotated by the task organizers.
 =
| ⋂  |
| |
 =
| ⋂  |
| |</p>
        <p>2 ∗  ∗ 
F1 = ( +  )
(2)
where G indicates the gold standard sentences, S denotes the sentences returned by our
methods.</p>
        <p>
          The second group of measures are ROUGE_1[
          <xref ref-type="bibr" rid="ref14">12</xref>
          ] (Recall-Oriented Understudy for
Gisting Evaluation) measures used as the official comparison measures.
        </p>
      </sec>
      <sec id="sec-2-2">
        <title>2.3.2 Task 1B metrics.</title>
        <p>Precision (P), recall (R) and F_1 measures were calculated to evaluate the performance
for each facet. For one facet, a is denoted as the number of correct predictions for this
facet, b is the number of wrong predictions for this facet, and c is the number of cited
sentences from gold standard sentences for this facet that are not predicted as belonging
to this facet.</p>
        <p>R = a+ac P = a+ab  1 = (2R∗R+∗PP)
(3)</p>
        <p>Then, Macro average and Micro average were used to evaluate the system on all
facets. Macro average measures are computed as:
Micro average measures are computed as:</p>
        <p>Rmacro = ∑ N        = ∑ N   1     = 2∗R      +∗</p>
        <p>∑   ∑   2∗R    ∗    
     = ∑  f+∑  f       = ∑   +∑     1    =      +    
(4)
(5)
where f is the facet index, N is the number of facets (i.e., 5). A good classifier should
have both high Macro and Micro average measures.
3</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>Results</title>
      <p>In this section, the results where the training set is used for training and the development
set is used for testing.
3.1 Task 1A results
According to Fig.1, Jaccard similarity and Dice similarity achieved better F_1 measures
than Cosine similarity and LCS. Topic similarity feature (indicated as Topic_number
in Fig.1) showed slight different performance among different numbers of topics.
WordNet based similarity achieved best results when combining the results of
standalone nouns similarities and verb similarities. TextRank showed similar results to
topic similarity and WordNet based similarity, however, this feature did not improve
the performance of learning to rank. Finally, Jaccard similarity, Topic similarity with
200 topics, WordNet (N,Vsep), and TextRank were kept. In the development set,
RankNet and AdaRank achieved best performance with Overlap/ ROUGE_1 F_1 of
0.057/0.211 (Table 1).
0.3
0.25
0.2
0.15
0.1
0.05
0</p>
      <sec id="sec-3-1">
        <title>Ovelap_F1</title>
      </sec>
      <sec id="sec-3-2">
        <title>Rouge_F1</title>
        <p>Micro Avg
Macro Avg</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>Final run methods and conclusions</title>
      <p>For the last run, we used both training set and development set to train instances. We
chose Jaccard, Topic_200, WordNet (N+V) and TextRank as the final ranking features
and applied AdaRank learning to rank algorithm on Task 1A. For Task 1B, Naïve Bayes
was used as the classifier.</p>
      <p>Recognizing the cited spans and determining their discourse facets are very challenging
for the summarization of scientific papers. There are some issues to be addressed during
our study on the tasks. First, either cited span discovery or discourse facet classification
is an unbalanced problem—negative examples are more prevalent than positive
examples. That would be our future work. Second, Task 1A involves much more than a
similarity problem. This is because the underlying citation intentions are complex. More
features that reflect the citation intentions should be explored.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>Kokil</given-names>
            <surname>Jaidka</surname>
          </string-name>
          , Muthu Kumar Chandrasekaran, Sajal Rustagi, and
          <string-name>
            <surname>Min-Yen Kan</surname>
          </string-name>
          (
          <year>2016</year>
          ).
          <article-title>Overview of the 2nd Computational Linguistics Scientific Document Summarization Shared Task (CL-SciSumm 2016)</article-title>
          , To appear
          <source>in the Proceedings of the Joint Workshop on Bibliometric-enhanced Information Retrieval and Natural Language Processing for Digital Libraries (BIRNDL</source>
          <year>2016</year>
          ), Newark, New Jersey, USA.
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <surname>Blei</surname>
            ,
            <given-names>D. M.</given-names>
          </string-name>
          (
          <year>2012</year>
          ).
          <article-title>Probabilistic topic models</article-title>
          .
          <source>Communications of the ACM</source>
          ,
          <volume>55</volume>
          (
          <issue>4</issue>
          ),
          <fpage>77</fpage>
          -
          <lpage>84</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <surname>George</surname>
            <given-names>A</given-names>
          </string-name>
          .
          <string-name>
            <surname>Miller</surname>
          </string-name>
          (
          <year>1995</year>
          ).
          <article-title>WordNet: A Lexical Database for English</article-title>
          .
          <source>Communications of the ACM</source>
          Vol.
          <volume>38</volume>
          , No.
          <volume>11</volume>
          :
          <fpage>39</fpage>
          -
          <lpage>41</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <surname>Lin</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          (
          <year>1998</year>
          ,
          <article-title>July)</article-title>
          .
          <article-title>An information-theoretic definition of similarity</article-title>
          . In ICML (Vol.
          <volume>98</volume>
          , pp.
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <surname>Mihalcea</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>Tarau</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          (
          <year>2004</year>
          ,
          <article-title>July)</article-title>
          .
          <source>TextRank: Bringing order into texts. Association for Computational Linguistics.</source>
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <surname>Dang</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          <string-name>
            <surname>The Lemur Project-Wiki-RankLib. Lemur Project</surname>
          </string-name>
          ,[Online]. Available: https://sourceforge.net/p/lemur/wiki/RankLib.
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <surname>Freund</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Iyer</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Schapire</surname>
            ,
            <given-names>R. E.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>Singer</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          (
          <year>2003</year>
          ).
          <article-title>An efficient boosting algorithm for combining preferences</article-title>
          .
          <source>The Journal of machine learning research</source>
          ,
          <volume>4</volume>
          ,
          <fpage>933</fpage>
          -
          <lpage>969</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <surname>Burges</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Shaked</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Renshaw</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lazier</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Deeds</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hamilton</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>Hullender</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          (
          <year>2005</year>
          ,
          <article-title>August)</article-title>
          .
          <article-title>Learning to rank using gradient descent</article-title>
          .
          <source>In Proceedings of the 22nd international conference on Machine learning</source>
          (pp.
          <fpage>89</fpage>
          -
          <lpage>96</lpage>
          ). ACM.
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [9]
          <string-name>
            <surname>Xu</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>Li</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          (
          <year>2007</year>
          ,
          <article-title>July)</article-title>
          .
          <article-title>Adarank: a boosting algorithm for information retrieval</article-title>
          .
          <source>In Proceedings of the 30th annual international ACM SIGIR conference on Research and development in information retrieval</source>
          (pp.
          <fpage>391</fpage>
          -
          <lpage>398</lpage>
          ). ACM.
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [10]
          <string-name>
            <surname>Metzler</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>Croft</surname>
            ,
            <given-names>W. B.</given-names>
          </string-name>
          (
          <year>2007</year>
          ).
          <article-title>Linear feature-based models for information retrieval</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          <string-name>
            <given-names>Information</given-names>
            <surname>Retrieval</surname>
          </string-name>
          ,
          <volume>10</volume>
          (
          <issue>3</issue>
          ),
          <fpage>257</fpage>
          -
          <lpage>274</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [11]
          <string-name>
            <surname>Hall</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Frank</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Holmes</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Pfahringer</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Reutemann</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>Witten</surname>
            ,
            <given-names>I. H.</given-names>
          </string-name>
          (
          <year>2009</year>
          ).
          <article-title>The WEKA data mining software: an update</article-title>
          .
          <source>ACM SIGKDD explorations newsletter</source>
          ,
          <volume>11</volume>
          (
          <issue>1</issue>
          ),
          <fpage>10</fpage>
          -
          <lpage>18</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [12]
          <string-name>
            <surname>Lin</surname>
            ,
            <given-names>C. Y.</given-names>
          </string-name>
          (
          <year>2004</year>
          ,
          <article-title>July)</article-title>
          .
          <article-title>Rouge: A package for automatic evaluation of summaries</article-title>
          .
          <source>In Text summarization branches out: Proceedings of the ACL-04 workshop</source>
          (Vol.
          <volume>8</volume>
          ).
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>