<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>ATHENA@CL-SciSumm 2019: Siamese recurrent bi-directional neural network for identifying cited text spans</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Aris Fergadis</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Dimitris Pappas</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Haris Papageorgiou</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Athena Research and Innovation Center</institution>
          ,
          <country country="GR">Greece</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Department of Informatics, Athens University of Economics and Business</institution>
          ,
          <country country="GR">Greece</country>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>School of Electrical and Computer Engineering, National Technical University of Athens</institution>
          ,
          <country country="GR">Greece</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>In this paper we describe our participation to the Task1 of the CL-SciSumm 2019. The task is on automatic paper summarization in the research area of Computational Linguistics. Our approach is a two step binary sentence pair classi cation between the so-called citances and candidate sentences. Firstly, we classify sentences in the abstracts to prede ned classes we call \zones". These zones capture the discourse structure of a scienti c publication. We then expand these zones with additional, similar sentences which are found in the main sections of the publication body. We train a Siamese bi-directional GRU neural network with a logistic regression layer to decide if a citance alludes to a candidate sentence. The cited sentences are also assigned one or more discourse facets (i.e., categories de ned in the Task) using a multi-class SVM. We ran extensive experiments in three different datasets achieving promising results.</p>
      </abstract>
      <kwd-group>
        <kwd>Candidate Cited Sentence Selection</kwd>
        <kwd>Siamese Neural Net- work</kwd>
        <kwd>Discourse Facet</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>
        Researchers are confronted with a continuously increasing volume of scienti c
publications, facing difficulties to monitor and track [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ]. The ability to create
synopsis of the key-points, contribution and importance of a paper within an
academic community is an important step [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ]. This synopsis, can be created
by using citation sentences (i.e., the citances) that reference a speci c paper
and can be considered as a community-created summary of a topic or a paper.
Scienti c summaries offer an overview of the cited paper useful to scholars,
writers or literature reviewers [
        <xref ref-type="bibr" rid="ref10 ref7">7, 10</xref>
        ]. The CL-SciSumm Shared Task focuses
on the scienti c summarization of papers [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ], organized into two tasks. For both
tasks the organizers provide several Reference Papers (RPs) called \topics".
Task1A: For each citance, identify the spans of text (cited text spans) in the RP
that most accurately re ect the citance. These spans are of the granularity of
a sentence fragment, a full sentence, or several consecutive sentences (no more
than 5).
      </p>
      <p>Task1B: For each cited text span, identify what facet of the paper it belongs to,
from a prede ned set of facets.</p>
      <p>Task2 (optional): Generate a structured summary of the RP from the cited texts
pans of the RP. The length of the summary should not exceed 250 words.</p>
      <p>
        We participated on the tasks 1A and 1B of 2019 shared task [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ] and present
here our methodology.
2
      </p>
    </sec>
    <sec id="sec-2">
      <title>Methodology</title>
      <p>
        We approach Task 1A as a binary sentence pair classi cation problem. We
create pairs of citance and a candidate sentence extracted from the Citing Papers
(CP) and the RPs respectively. Word embeddings are used to select candidate
sentences. A Siamese neural network process these pairs to decide whether or
not the candidate sentence is a cited text span of the citance. For the Task 1B
a multi-class SVM [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ] model assigns a discourse facet to the cited text spans.
2.1
      </p>
      <p>
        System Components
Word Embeddings We use word embeddings for both the candidate
sentences selection and for the embedding layer of our network. Embedding vectors
are trained on the ACL Corpus dump4 using the CBOW implementation of
word2vec [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ] of the gensim5 tool, with negative sampling set to 5 and 100 for
dimensionality of the word vectors. All words are converted to lowercase.
Candidate Sentence Selection We select sentences from the RP as candidate
sentences. The intuition is that not all sentences are equally important as cited
text spans. Thus, we try to select sentences that are about methodology, results
and conclusions and discard sentences about background and related work. This
is also supported by the fact that the cited text is assigned a discourse facet. Our
approach tries to eliminate sentences that potentially would be false positives.
      </p>
      <p>
        To select the candidate sentences of the RP we split the abstract into zones [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ].
Each sentence is classi ed to one of the following zones: Background, Method,
Result and Conclusion. We keep only the sentences that belong to Method,
Result and Conclusion zones (referred to as zone sentences ). Sentences are split
into words6, punctuation and numbers are removed and each word is assigned
its embedding vector. For each zone sentence and the rest of the RP sentences
we calculate sentence embeddings by averaging the word embedding vectors.
4 http://acl-arc.comp.nus.edu.sg/archives/acl-arc-160301-parscit/
5 https://radimrehurek.com/gensim/, version 3.7.3
6 Using the tokenization tools of the gensim module
      </p>
      <p>Using the embedding vectors of the N zone sentences and all the other
embedding vectors of the M RP sentences, we calculate a similarity matrix S 2 RN M
using cosine similarity measure. To get the most similar sentences to the zone
sentences we de ne a threshold ts. The RP sentences Si;j that pass the
similarity threshold ts and the zone sentences are kept as candidate sentences. The
decision of the ts value is discussed into section 3.</p>
      <p>
        Siamese Neural Network The Siamese neural network is composed of two
bi-directional GRUs (biGRU) [
        <xref ref-type="bibr" rid="ref14 ref2">14, 2</xref>
        ] and a logistic regression layer, as depicted in
Figure 1. Each biGRU processes one sentence at a time. For each citance and a set
of candidate sentences of the RP, the left biGRU takes as input the citance and
the right biGRU takes as input one of the candidate sentences. We use w1:n to
denote a sequence of words w1:n = w1; : : : wn, each with their corresponding demb
dimensional word embedding ei = E[wi]. The embedding matrix E 2 RjV j demb
associates words from the vocabulary V with demb dimensional dense vectors.
      </p>
      <p>
        The left biGRU applies additive zero-centered Gaussian noise [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ] to word
embeddings with = 0:05 as a regularization layer at the training phase. The
outputs y1b and ynf of the backward GRUb and the forward GRUf respectively
are concatenated in one vector
y1b = GRUb(en:1)
ynf = GRUf (e1:n)
xl = [y1b; ynf]
y1′b = GRUb(en:1)
yn′f = GRUf (e1:n)
xr = [y1′b; yn′f ]
We use xl to denote the output of the left input and xr of the right input and [ ; ]
to denote concatenation. The two output vectors are element wise multiplied to
give a vector x. A logistic regression layer (LR) with a sigmoid activation
function ( ) is used to make the nal prediction y^. To summarize the architecture
p(y = kjw1:n) = y^; k 2 f0; 1g
y^ = LR(x); with ( ) activation
x = [xl
      </p>
      <p>xr]</p>
      <p>The described model considers one sentence at a time. In order to nd if a
citance references more than one sentences in the RP, we take the predictions
of all the candidate sentences and keep the maximum score smax. We de ne a
threshold as st = 0:98 smax. Any candidate sentence that has score s such as
st s smax is selected as a cited sentence.</p>
      <p>b
y1</p>
      <sec id="sec-2-1">
        <title>GRUb GRUb</title>
      </sec>
      <sec id="sec-2-2">
        <title>GRUf GRUf</title>
        <p>GN
e1
GN
e2</p>
        <p>Citance
y1b; ynf</p>
      </sec>
      <sec id="sec-2-3">
        <title>GRUb</title>
      </sec>
      <sec id="sec-2-4">
        <title>GRUf</title>
        <p>GN
en
y^
LR
y1b; ynf
y1′b; yn′f
f
yn
y1′b; yn′f
y1′b</p>
      </sec>
      <sec id="sec-2-5">
        <title>GRUb</title>
      </sec>
      <sec id="sec-2-6">
        <title>GRUb</title>
      </sec>
      <sec id="sec-2-7">
        <title>GRUf</title>
      </sec>
      <sec id="sec-2-8">
        <title>GRUf</title>
        <p>yn′f</p>
      </sec>
      <sec id="sec-2-9">
        <title>GRUb</title>
      </sec>
      <sec id="sec-2-10">
        <title>GRUf</title>
        <p>e′
1
e′
2</p>
        <p>e′n
Candidate Sentence</p>
        <p>Discourse Facet Task 1B asks \for each cited text span, identify what facet of
the paper it belongs to, from a prede ned set of facets". The prede ned facets
are Aim, Hypothesis, Method, Result and Implication. We approach this task as
a multi-class classi cation problem due to the fact that some cited text spans
may have up to two facets. We build a bag-of-terms representation of all n-grams
with n = 1; 2; 3 and calculate their tf-idf values using L1-norm. Five one-vs-rest
SVM classi ers were trained assigning a cited text span to each of the ve facets.</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>Experiments and Results</title>
      <p>The dataset provided was split into a training and a development set. The
training set consists of a set of 40 RPs with their CPs annotated by humans and 1000
RPs and their corresponding CPs that were automatically annotated. For our
experiments we only used the human part of the training set (the TR-H set ).</p>
      <p>As a rst step for our experiments we selected candidate sentences from the
RPs. By keeping only the candidate sentences we might miss cited sentences
in the RP which were not selected by our method. In Table 1, coverage is the
number of RP sentences we kept and the hits metric is the number of the cited
sentences in our candidate list (in percentage). Our target is to get minimum
coverage with maximum hits. Minimum coverage means that we have kept all
the good candidates while maximum hits denotes that the cited sentences are
within our candidate list.</p>
      <p>Table 1 displays the average coverage and hits for the 40 RPs of the training
set and the 10 RPs of the development set for different thresholds. Based on
the results, ts was set to 0.5. Using this threshold, we keep about 70% of the
RP sentences on the training set and 60% on the development set, on average.
Despite the fact that we discarded about 30% and 40% of the candidate sentences
we only lose 15% and 20% of the cited text spans, respectively.</p>
      <p>
        We evaluated our system in three different versions of the dataset; for each
version, we used for testing the development set (Dev), the 2016 test set (2016)
and the 2017 test set (2017) respectively. For training, we used the TR-H set
provided that we have excluded all papers in the relevant testing set for obvious
reasons. The results shown in Table 2 are comparable to those of the previous
shared tasks [
        <xref ref-type="bibr" rid="ref5 ref6 ref8">6, 5, 8</xref>
        ].
4
      </p>
    </sec>
    <sec id="sec-4">
      <title>Conclusions and Future Work</title>
      <p>
        Scienti c summarization is a challenging task as it is evident from the results of
the previous shared tasks [
        <xref ref-type="bibr" rid="ref5 ref6 ref8">6, 5, 8</xref>
        ]. In our methodology we create pairs of citance
and a candidate sentence extracted from the CP and the RP respectively. These
pairs are classi ed from a Siamese neural network as positive if a citance indeed
cites a sentences, otherwise as negative. The cited sentences are also assigned one
or more discourse facets. We applied our methods on the dataset of the 2019
CLSciSumm shared task. The evaluation of our system indicates that the Siamese
neural network performs comparable to other machine learning methods.
      </p>
      <p>In future work we will investigate the impact of replacing the logistic
regression layer with other similarity functions, such as cosine similarity. We also plan
to select the best value for the st threshold via hyper-parameter tuning. Finally,
we will experiment with different methods for cited sentences selection which
take into account the scores of neighboring sentences.
5</p>
    </sec>
    <sec id="sec-5">
      <title>Acknowledgement</title>
      <p>We acknowledge support of this work by the Data4Impact Project which received
funding from the European Union's Horizon 2020 research and innovation
programme under grant agreement No 770531.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Chandrasekaran</surname>
            ,
            <given-names>M.K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Yasunaga</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Radev</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kan</surname>
          </string-name>
          , M.Y.:
          <article-title>Overview and results: Cl-scisumm shared task 2019</article-title>
          .
          <source>In: Proceedings of the 4th Joint Workshop on Bibliometric-enhanced Information Retrieval and Natural Language Processing for Digital Libraries (BIRNDL</source>
          <year>2019</year>
          ) @
          <source>SIGIR</source>
          <year>2019</year>
          , Paris, France. (
          <year>2019</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Cho</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          , Van Merrienboer,
          <string-name>
            <given-names>B.</given-names>
            ,
            <surname>Bahdanau</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            ,
            <surname>Bengio</surname>
          </string-name>
          ,
          <string-name>
            <surname>Y.</surname>
          </string-name>
          :
          <article-title>On the properties of neural machine translation: Encoder-decoder approaches</article-title>
          .
          <source>arXiv preprint arXiv:1409.1259</source>
          (
          <year>2014</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Cortes</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Vapnik</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          :
          <article-title>Support-vector networks</article-title>
          .
          <source>Machine learning 20(3)</source>
          ,
          <volume>273</volume>
          {
          <fpage>297</fpage>
          (
          <year>1995</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Hinton</surname>
          </string-name>
          , G.,
          <string-name>
            <surname>Van Camp</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          :
          <article-title>Keeping neural networks simple by minimizing the description length of the weights</article-title>
          .
          <source>In: in Proc. of the 6th Ann. ACM Conf. on Computational Learning Theory. Citeseer</source>
          (
          <year>1993</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Jaidka</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chandrasekaran</surname>
            ,
            <given-names>M.K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Jain</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kan</surname>
          </string-name>
          , M.Y.:
          <article-title>The cl-scisumm shared task 2017: Results and key insights</article-title>
          .
          <source>In: BIRNDL@ SIGIR (2)</source>
          . pp.
          <volume>1</volume>
          {
          <issue>15</issue>
          (
          <year>2017</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Jaidka</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chandrasekaran</surname>
            ,
            <given-names>M.K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rustagi</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kan</surname>
          </string-name>
          , M.Y.:
          <article-title>Overview of the clscisumm 2016 shared task</article-title>
          .
          <source>In: Proceedings of the Joint Workshop on Bibliometricenhanced Information Retrieval and Natural Language Processing for Digital Libraries (BIRNDL)</source>
          . pp.
          <volume>93</volume>
          {
          <issue>102</issue>
          (
          <year>2016</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>Jaidka</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Khoo</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Na</surname>
            ,
            <given-names>J.C.</given-names>
          </string-name>
          :
          <article-title>Deconstructing human literature reviews{a framework for multi-document summarization</article-title>
          .
          <source>In: proceedings of the 14th European workshop on natural language generation</source>
          . pp.
          <volume>125</volume>
          {
          <issue>135</issue>
          (
          <year>2013</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>Jaidka</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Yasunaga</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chandrasekaran</surname>
            ,
            <given-names>M.K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Radev</surname>
            ,
            <given-names>D.R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kan</surname>
          </string-name>
          , M.Y.:
          <article-title>The cl-scisumm shared task 2018: Results and key insights</article-title>
          .
          <source>In: BIRNDL@SIGIR</source>
          (
          <year>2018</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <surname>Jin</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Szolovits</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          :
          <article-title>Hierarchical neural networks for sequential sentence classi - cation in medical scienti c abstracts</article-title>
          . ArXiv abs/
          <year>1808</year>
          .06161 (
          <year>2018</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <surname>Jones</surname>
            ,
            <given-names>K.S.:</given-names>
          </string-name>
          <article-title>Automatic summarising: The state of the art</article-title>
          .
          <source>Inf. Process. Manage</source>
          .
          <volume>43</volume>
          ,
          <issue>1449</issue>
          {
          <fpage>1481</fpage>
          (
          <year>2007</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <string-name>
            <surname>Mikolov</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sutskever</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chen</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Corrado</surname>
            ,
            <given-names>G.S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Dean</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          :
          <article-title>Distributed representations of words and phrases and their compositionality</article-title>
          .
          <source>In: Advances in neural information processing systems</source>
          . pp.
          <volume>3111</volume>
          {
          <issue>3119</issue>
          (
          <year>2013</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          12.
          <string-name>
            <surname>Nakov</surname>
            ,
            <given-names>P.I.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Schwartz</surname>
            ,
            <given-names>A.S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hearst</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          : Citances:
          <article-title>Citation sentences for semantic analysis of bioscience text</article-title>
          .
          <source>In: Proceedings of the SIGIR</source>
          . vol.
          <volume>4</volume>
          , pp.
          <volume>81</volume>
          {
          <issue>88</issue>
          (
          <year>2004</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          13.
          <string-name>
            <surname>Qazvinian</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Radev</surname>
            ,
            <given-names>D.R.</given-names>
          </string-name>
          :
          <article-title>Scienti c paper summarization using citation summary networks</article-title>
          .
          <source>In: Proceedings of the 22nd International Conference on Computational Linguistics-Volume</source>
          <volume>1</volume>
          . pp.
          <volume>689</volume>
          {
          <fpage>696</fpage>
          .
          <article-title>Association for Computational Linguistics (</article-title>
          <year>2008</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          14.
          <string-name>
            <surname>Schuster</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Paliwal</surname>
            ,
            <given-names>K.K.:</given-names>
          </string-name>
          <article-title>Bidirectional recurrent neural networks</article-title>
          .
          <source>IEEE Transactions on Signal Processing</source>
          <volume>45</volume>
          (
          <issue>11</issue>
          ),
          <volume>2673</volume>
          {
          <fpage>2681</fpage>
          (
          <year>1997</year>
          )
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>