<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Attention: to Better Stand on the Shoulders of Giants</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Sha Yuan</string-name>
          <email>yuansha@baai.ac.cn</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Zhou Shao</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Yu Zhang</string-name>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Tong Xiao</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Yifan Wang</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Beijing Academy of Arti cial Intelligence</institution>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Department of Computer Science and Technology, Tsinghua University</institution>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>Institute of Medical Information, Peking Union Medical College, Chinese Academy of Medical Sciences</institution>
        </aff>
        <aff id="aff3">
          <label>3</label>
          <institution>Nanjing University of Science and Technology</institution>
        </aff>
      </contrib-group>
      <abstract>
        <p>Science of Science (SciSci) is an emerging discipline wherein science is used to study the structure and evolution of science itself using large data sets. The increasing availability of digital data on scholarly outcomes o ers unprecedented opportunities to explore SciSci. In the progress of science, the previously discovered knowledge principally inspires new scienti c ideas, and citation is a reasonably good re ection of this cumulative nature of scienti c research. The researchers that choose potentially in uential references will have a lead over the emerging publications. Although the peer-review process is the mainly reliable way of predicting a paper's future impact, the ability to foresee the lasting impact based on citation records is increasingly essential in the scienti c impact analysis in the era of big data. This paper develops an attention mechanism for the long-term scienti c impact prediction and validates the method based on a real large-scale citation data set. The results break conventional thinking. Instead of accurately simulating the original power-law distribution, emphasizing the limited attention can better stand on the shoulders of giants.</p>
      </abstract>
      <kwd-group>
        <kwd>Science of Science</kwd>
        <kwd>Scienti c Impact</kwd>
        <kwd>Attention</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>
        With the advent of the era of big data, people pay more and more attention
to the value of data. The massive volume of publications created every year
has grown into big data that we can't ignore. With the development of SciSci, it
provides a quantitative understanding of scienti c discovery, creativity, and
practice [
        <xref ref-type="bibr" rid="ref11 ref12 ref23 ref29 ref34 ref5">12, 11, 23, 34, 5, 29</xref>
        ]. From the perspective of SciSci, identifying fundamental
drivers of science and developing predictive models to capture its evolution are
instrumental for successful science. SciSci reveals that the previously discovered
knowledge mainly inspires new scienti c ideas [
        <xref ref-type="bibr" rid="ref47">47</xref>
        ], and citation [
        <xref ref-type="bibr" rid="ref24">24</xref>
        ] is a
relatively good re ection of this cumulative nature of scienti c research. Citation
count, which has been used to evaluate the quality and in uence of scienti c
work for a long time, stands out from many quanti cation measure metrics of
scienti c impact [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ]. With the rapid evolution of scienti c research, there is a
massive volume of literature published every year, and this situation is expected
to remain within the foreseeable future. Fig. 1 shows the statistics on the
citation data set used in this paper. The data set is extracted from AMiner [
        <xref ref-type="bibr" rid="ref38">38</xref>
        ],
which is a billion-scale academic search and mining system. Fig. 1(a) visualizes
the explosive increase in the volume of publications in the past years from 1990
to 2015. It shows that the literature quantity assumes the exponential order to
grow.
      </p>
      <p>(a) The volume of literatures.</p>
      <p>(b) The citation Distribution.</p>
      <p>
        Scienti c work is founded on prior research. It is not wise nor possible for
researchers to track all existing related work due to the enormous volume of the
existing publications. In general, researchers follow or cite merely a small
proportion of high-quality publications. SciSci provides several quanti cation methods
for scienti c impact measurement in article-level, author-level, and journal-level.
Much SciSci work has been done on the evaluation metrics for the quality and
in uence of scienti c work [
        <xref ref-type="bibr" rid="ref30 ref31">31, 30</xref>
        ], including citation count [
        <xref ref-type="bibr" rid="ref33">33</xref>
        ], h-index [
        <xref ref-type="bibr" rid="ref18">18</xref>
        ],
and impact factor [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ]. One of the most basic quanti cation measure metrics of
scienti c impact is citation count. It measures the number of received citations
for an article. Many other essential evaluation criteria of authors (e.g., h-index)
and journals [
        <xref ref-type="bibr" rid="ref13 ref15">15, 13</xref>
        ] (e.g., Impact Factor) are calculated based on citation count.
      </p>
      <p>
        A lot of SciSci researchers have focused on the characterization of scienti c
impact [
        <xref ref-type="bibr" rid="ref35">35</xref>
        ], such as the universal citation distributions [
        <xref ref-type="bibr" rid="ref28">28</xref>
        ], the characteristics of
citation networks [
        <xref ref-type="bibr" rid="ref19 ref20 ref26">19, 26, 20</xref>
        ], and the growth pattern of scienti c impact [
        <xref ref-type="bibr" rid="ref10 ref17">10, 17</xref>
        ].
The results reveal the regularity of scienti c progress that a few research papers
attract the vast majority of citations [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ], long-distance interdisciplinarity leads to
higher scienti c impact [
        <xref ref-type="bibr" rid="ref21 ref46">21, 46</xref>
        ]. Fig. 1(b) illustrates the citation distribution (the
number of papers vs. citation counts) of about two million papers. The citation
distribution follows the power-law distribution [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ]. It is natural to nd that
not all publications attract equal attention in academia. A few research papers
accumulate the vast majority of citations, and most of the other papers attract
only a few citations [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ]. A small number of scholarly outcomes are more likely to
attract scientists' attention than others accounting for a vast majority. For the
ever-growing literature quantity, it is signi cant to forecast which paper is more
likely to attract more attention in the scienti c community. Zhu et al. [
        <xref ref-type="bibr" rid="ref48">48</xref>
        ] present
several machine learning methods and one multiple linear regression strategy
to predict a paper's future citation. Ali et al. [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ] propose a novel method for
predicting long-term citations of a paper based on the number of its citations
in the rst few years after publication. Daniel et al. [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ] present GNN-based
architecture that predicts the top set of papers at the time of publication.
      </p>
      <p>
        The fact is that the current citation count and the derived metrics can only
capture the past accomplishment. They lack the predictive power to quantify
the future impact [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]. Predicting an individual paper's citation count over time
is signi cant, but (arguably) very di cult. To predict the citation count of
individual items within a complex evolving system, current models are falling into
two main paradigms. One formulates the citation count over time as time
series and then makes predictions by either exploiting temporal correlations [
        <xref ref-type="bibr" rid="ref37">37</xref>
        ]
or tting these time series with certain classes of designed functions [
        <xref ref-type="bibr" rid="ref25 ref3">25, 3</xref>
        ],
including the regression models [
        <xref ref-type="bibr" rid="ref45">45</xref>
        ], the counting process [
        <xref ref-type="bibr" rid="ref39">39</xref>
        ], the point process,
the Poisson process [
        <xref ref-type="bibr" rid="ref40 ref41">40, 41</xref>
        ], Reinforced Poisson Process (RPP) [
        <xref ref-type="bibr" rid="ref32">32</xref>
        ], self-excited
Hawkes Process [
        <xref ref-type="bibr" rid="ref27">27</xref>
        ], RPP with self-excited Hawkes Process [
        <xref ref-type="bibr" rid="ref42">42</xref>
        ]. The designed
functions consider various factors.
      </p>
      <p>
        The other prevalent line utilizes Deep Neural Network (DNN) based
models to solve the scienti c impact prediction problem. Recently, Convolutional
Neural Network (CNN) and Recurrent Neural Network (RNN) have received
considerable attention from academia and the industry. RNN has been proven
to perform particularly well on temporal data series [
        <xref ref-type="bibr" rid="ref36">36</xref>
        ]. Due to the
vanishing gradient problem, RNN always fails to handle the temporal contingencies
present in the input/output sequences spanning long intervals [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ]. The networks
with loops in RNN allow information to persist for a long time. Long short-term
memory (LSTM) is proven to be capable of learning long-term dependencies.
RNN with LSTM units performs rather well in handling long-term temporal
data series [
        <xref ref-type="bibr" rid="ref43">43</xref>
        ].
      </p>
      <p>All the existing methods try to tune the citation distribution precisely as the
original power-law distribution. However, this paper argues that the e ectiveness
of quantifying long-term scienti c impact is fundamentally limited in this routine
thinking. This paper proposes to put more attention on some speci c items,
such as highly cited papers. The authors validate the proposed method on a real
large-scale citation data set. Extensive experiment results demonstrate that the
proposed method possesses remarkable power at predicting long-term scienti c
citation. The most important contribution is that this paper changes the line
of thinking in quantifying the long-term scienti c impact. Instead of simulating
the original power-law distribution, researchers need to emphasize the limited
attention to better stand on giants' shoulders.
2</p>
    </sec>
    <sec id="sec-2">
      <title>Problem Formulation</title>
      <p>The primary evaluation metric for scienti c impact is citation count. The
received citation count of an individual paper d during time period [0; T ] is
characterized by a time-stamped sequence fntdgtT=0, where ntd represents the number
of citation counts received by paper d at time t, ntd is an integer greater than or
equal to zero. In giving the historical citation records, the goal is to model the
future citation count and predict it over an arbitrary time.</p>
      <p>Given the literature corpus D, card(D) = M means nd M papers from D.
In this paper, we believe that the scienti c impact of a literature article is equal
to the number of papers which cite it. The scienti c impact of a literature article
d 2 D at time t is de ned as its citation counts ntd:
citingdt = fd 2 D; d~ 6= d : d~t cites dg; ntd
~
= card(citingdt):
(1)</p>
      <p>The underlying assumption of the citation count here is the accumulated
citations, which make it possible to quantify citations for di erent items at di erent
times. The long-term scienti c impact of the individual item d can be formalized
as the following time series fn0d; ; ntd; ; ndT g. Without loss of generality, the
number of accumulated citation counts increases over time. And then, we have
0 = n0d ntd ndT = Nd.</p>
      <p>In the scienti c impact prediction problem, the input X~ is fn0d; ; ntd; g,
for every paper d 2 D, where ntd is the citation counts of paper d at time t. The
goal of the scienti c impact prediction problem is to learn a predictive function f
to predict the citation counts of an article d after a given period time t. Formally,
we have:</p>
      <p>f (djX~ ; t) ! n^td;
where n^td is the predicted citation count and ntd is the actual one. Based on the
learned prediction function, we can predict the citation count of a paper for the
next years. For example, the citation count of paper d at time t is given by
f (djX~ ; t).
(2)
3</p>
    </sec>
    <sec id="sec-3">
      <title>Scienti c Impact Prediction</title>
      <p>As the most e cient scienti c impact prediction method found so far, RNN
has already achieved compelling performance in predicting the scienti c impact.
This paper embeds the RNN with LSTM units as a baseline and then
emphasizes highly cited articles in the proposed attention mechanism. Although many
other elds have used the attention mechanism, the proposed method gives new
insight into long-term quantifying scienti c impact. Instead of adapting citation
distribution to a power-law distribution, this paper's ndings provide a new line
of thinking for the SciSci research.
(a) The overview.</p>
      <p>(b) The LSTM units.</p>
      <p>
        (c) The attention model.
Given a time-stamped sequence fndgt=0, a K-dimensional feature vector X~ =
t T
fx0d; ; xtd; ; xdT g needs to be designed as input. The input space of every
item with popularity records f(x0; n0); ; (xt; nt); ; (xT ; nT )g re ects the
intrinsic quality of the item. Fig. 2(a) gives an overview of the model architecture.
There are two critical components in the architecture: the RNN with LSTM units
and the attention model. As illustrated in Fig. 2(b), it arranges the LSTM units
in the form of RNN with L layers. In the deep neural network, the parameter
L depends on the input scale. RNN is famous for its popularity and well-known
capability for e cient time series learning [
        <xref ref-type="bibr" rid="ref43">43</xref>
        ]. The LSTM units capture the
long-range dependency in long-term scienti c impact quanti cation.
      </p>
      <p>
        The RNN with LSTM Units. The LSTM units are arranged in the form of
RNN, as illustrated in Fig. 2(b). There are four major components in a standard
LSTM unit, including a memory cell, a forget gate f , an input gate i, and an
output gate o. The gates are responsible for information processing and storage
over arbitrary time intervals. Usually, the outputs of these gates are between 0
and 1. A new study gives suggestions to push the output values of the gates
towards 0 or 1. By doing so, the gates are mostly open or closed instead of in a
middle state [
        <xref ref-type="bibr" rid="ref22">22</xref>
        ]. This paper arranges the LSTM units in the form of RNN. In
this way, introducing the memory cell will solve the vanishing gradient problem.
Thus, it can store information for either short or long periods in the LSTM unit.
      </p>
      <p>Intuitively, the input gate controls the extent to which a new value ows into
the memory cell. The input gate's function passes through the input gate and is
added to the cell state to update it. The following formula for the input gate is
used:
it =</p>
      <p>Wi ht 1; xt + bi ;
(3)
where matrix Wi collects the weights of the input and recurrent connections.
The symbol represents the Sigmoid function. The values of the vector it are
between 0 and 1. If one of the values of it is 0 (or close to 0), it means that this
input gate is closed, and no new information is allowed into the memory cell at
time t. If one of the values is 1, the input gate is open for a new coming value
at time t. Otherwise, the gate is in the state of half-open half-clearance.</p>
      <p>The forget gate controls the extent to which a value remains in the memory
cell. It provides a way to get rid of the previously stored memory value. Here is
the formulation of the forget gate:
ft =</p>
      <p>Wf ht 1; xt + bf ;
(4)
where Wf is the weight matrix that governs the behavior of the forget gate.
Similar to it, ft is also a vector of values between 0 and 1. If one of the values
of ft is 0 (or close to 0), it means that the memory cell should remove that piece
of information in the corresponding component in the cell. If one of the values
is 1, the corresponding information will be kept.</p>
      <p>Remembering information for long periods is practically the default behavior
of LSTM. The long-term accumulative in uence is formulated as follows:
That is, the information in the memory cell consists of two parts: the retained
old information ft ct 1 (controlled by the forget gate) and the new coming
information it c~t (controlled by the input gate).</p>
      <p>The output gate controls the extent to which the value in the cell is used to
compute the output activation of the LSTM unit. The following output function
is used:
ot =</p>
      <p>Wo ht 1; xt + bo :
The weight matrices and bias vector parameters are needed to be learned during
training. This paper updates the current working state as the following formula:
ht =
t tanh ct :
o
ct =
t
f
ct 1 + it</p>
      <p>c~t;
c~t = tanh Wc ht 1; xt + bc :
where denotes the Hadamard product (the element-wise multiplication of
matrices), c~t is calculated as follows:
(5)
(6)
(7)
The items stored in the current working state have an advantage in reading
over those stored in long-term memory. In the time series modeling of scienti c
impact, the recent items stored in the short-term working state have an
advantage over those stored in the long-term memory. The next step introduces the
attention mechanism based on ht.</p>
      <p>The Attention Model. The arti cial attention mechanism, inspired by the
attention behavior in neuroscience, has been applied in deep learning for speech
recognition, translation, and visual identi cation of objects.Broadly, attention
mechanisms are components of prediction systems that allow the system to focus
on di erent subsets of the input sequentially. It aims to capture the critical points
and focuses on the relevant parts more than the remote parts as a human does.
More speci cally, content-based attention generates attention distribution. Only
part of a subset of the input information is focused. The attention function needs
to be di erentiable, so that everywhere of the input is focused, just to di erent
extents.</p>
      <p>The deep learning attention mechanism used in this paper works as follows:
given an input X~ = fx0d; ; xtd; ; xdT g, the aforementioned LSTM units
generate ~h = fh1; ; ht; ; hT g to represent the hidden patterns of the input. The
output is the summary of the ht focusing on information linked to the input.
In this formulation, attention produces a xed-length embedding of the input
sequence by computing an adaptive weighted average of the state sequence ~h.</p>
      <p>The graphical representation of the attention model is shown in Fig. 2(c).
The input X~ and the hidden layer ~h of the LSTM network (an RNN composed
of LSTM units) are the input of the attention model. Then, it computes the
following formula:
at = tanh Wa xt; ht
;
where Wa is the weight matrix. An important remark here is that each at is
computed independently without looking at the other xt0 for t0 6= t. Then, each
at is linked to a Softmax layer, which function is given by:
(9)
(10)
(11)
t =</p>
      <p>eat
P eat ; for t = 1;
t</p>
      <p>; T
O = X</p>
      <p>txt:
t
where Pt t = 1, the t is the softmax of the at projected on a learned direction.
The output is a weighted arithmetic mean of the input, and the weights re ect
the relevance of ~h and the input. It is calculated as the following formula:
Finally, the popularity of item d at time t is given by the prediction f (djX~ ; t) =
O.
3.2</p>
      <sec id="sec-3-1">
        <title>Key Factor in Quantifying Long-term Impact</title>
        <p>
          As widely acknowledged, the citation distribution follows the power-law
distribution. This nding leads the way for research in this domain. Researchers try
to simulate the citation distribution as the power-law distribution. This paper
changes the line of thinking. Although the number of research papers has
exploded, the reading time of scientists has not. At the same time, the attention
shifts toward the top 1% over time [
          <xref ref-type="bibr" rid="ref4">4</xref>
          ]. Even though the citation distribution
follows the power-law distribution, attention is vital in quantifying the long-term
scienti c impact.
        </p>
        <p>
          In the fact of limited attention, the Matthew e ect dominates in quantifying
the long-term scienti c impact. The experiments will con rm it. The citation
count captures the inherent di erences between papers, accounting for the
perceived novelty and the importance of a paper. The "rich-get-richer" phenomenon
summarizes the Matthew e ect of accumulated advantage, i.e., previously
accumulated attention triggers more subsequent attention [
          <xref ref-type="bibr" rid="ref8">8</xref>
          ] than others. In fact,
the highly popular items are more visible and more likely to be viewed than
others. The proposed model emphasizes highly cited papers under limited
attention. The memory cell in the LSTM unit considers the long-term dependencies.
As shown in Eq. (5), previously accumulated attention stored in the long-term
memory triggers more subsequent attention. What is more, the attention model,
which focuses on the most popular part of the time series as Eq. (11) does, also
emphasizes the Matthew e ect.
4
        </p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>Experiments</title>
      <p>4.1</p>
      <sec id="sec-4-1">
        <title>Dataset</title>
        <p>This section demonstrates the e ectiveness of putting particular emphasis on
the vital factor in quantifying the long-term scienti c impact.</p>
        <p>The authors extract the data from an academic search and mining platform
called AMiner and construct a real large-scale scholarly dataset{Academic
Social Network1. The citation network's full graph in this dataset has about 2
million vertices (papers) and 8 million edges (citations). In detail, the dataset
is composed of 2; 092; 356 digitalized papers spanning from 1936 to 2016 (for
more than 80 years) and 8; 024; 869 citations between them. By convention, the
authors eliminate those papers with less than 5 citations during the rst 5 years
after publication and only retain the remaining papers as the training data. As
a result, 143; 902 papers published from 1956 to 2015 are retained.
4.2</p>
      </sec>
      <sec id="sec-4-2">
        <title>Baseline Models and Evaluation Metrics</title>
        <p>
          To compare the predictive performance of the proposed attention model against
other models, we introduce several published models that have been used to
predict scienti c impact. Speci cally, the experiments' comparison methods are LR,
CART, SVR (the three basic machine learning methods used in [
          <xref ref-type="bibr" rid="ref45">45</xref>
          ]), RPP [40,
1 https://www.aminer.cn/data
32], and RNN [
          <xref ref-type="bibr" rid="ref43">43</xref>
          ]. The advantage of deep learning is the utilization of various
features. For fairness, the authors only use the citation count records and the
same feature used in [
          <xref ref-type="bibr" rid="ref32 ref40">40, 32</xref>
          ]. For the fair comparison among di erent kinds of
models, all models use the same vector features, which are the citation records
of the rst ve years after publication. Statistics show that for most papers, the
rst ve years after publication can well re ect their in uence on the current
research stage.
        </p>
        <p>This paper uses two basic scienti c impact evaluation metrics: Mean Absolute
Percentage Error (MAPE) and Accuracy (ACC). Let ntd be the observed citations
of paper d up to time t, and n^td be the predicted one. The MAPE measures
the average deviation between the predicted and observed citations over all the
papers. For a dataset of M papers, the MAPE is given by:
1 XM n^td
M
d=1
ntd
ntd :
(12)</p>
        <p>ACC measures the fraction of papers correctly predicted under a given error
tolerance . Speci cally, the accuracy of citation prediction over M papers is
de ned as:</p>
        <p>
          ACC =
where I[ ] is an indicator function which returns 1 if the statement is true,
otherwise returns 0. We nd that our method always outperforms regardless of
's value. In this paper, we set = 0:3.
The experiment results show that the longer the duration of the training set, the
better the long-term prediction performance. According to our experiment, this
paper sets the training period as 5 years and then predicts the citation counts for
each paper from 1st to 5th after the training period. For example, t = 1 means
that the rst observation year after the training period. In the experiments, the
features with positive contributions are the citation history, the current h-index
of the paper author, and the publication journal level [
          <xref ref-type="bibr" rid="ref44">44</xref>
          ]. For the convenience
of performance comparison, the input feature used here is the citation history
for every sub-window of length 10 years. The value of the parameter L is 2. The
loss function used here is MAPE. Adadelta is the gradient descent optimization
algorithm. The attention layer is fully connected and uses tanh activation. And
the code is available on github 2.
As shown in Table. 1, the proposed model exhibits the best performance in
terms of ACC in all the situations of t = 1, 2, 3, 4, and 5. It means that the
2 https://github.com/AIOpenData/attention
DLAM consistently achieves higher accuracy than other models across di erent
observation times. What is more, the proposed model also exhibits the best
performance in terms of MAPE in all the situations mentioned above. That is, the
proposed model achieves higher accuracy and lowers error rates simultaneously.
In the experiments, all the models used for comparison achieve acceptable low
error rates, except RPP. RPP can avoid this problem with prior [
          <xref ref-type="bibr" rid="ref32">32</xref>
          ], which
incorporates conjugate prior for the tness parameter. However, the RPP with
prior does not improve the ACC performance. Overall, the proposed model also
outperforms RPP with prior.
Compared to the other methods in terms of ACC and MAPE, the proposed
model increases with the number of years after the training period. Compare to
RNN (the most e cient method certi ed in recent works), the proposed model
achieves a few performance improvements, about 1:65% in terms of MAPE and
about 2:13% in terms of ACC when t = 1. However, in the situation of t = 5, the
proposed model achieves signi cant performance improvement of about 24:31%
in terms of MAPE and about 16:7% in terms of ACC. In other words, the
proposed model shows much superiority over other models in scienti c impact
prediction, especially in the long-term situation.
5
        </p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>Further Exploration</title>
      <sec id="sec-5-1">
        <title>E ectiveness of the attention mechanism. The authors remove the</title>
        <p>attention module of the proposed model to verify the e ectiveness of the
attention mechanism. The remainder is RNN with LSTM units (labeled as LT-CCP),
proven to be useful in long-term citation count prediction. In the next step, we
add the attention mechanism in two di erent ways. Firstly, we add the
attention module before the RNN module, labeled as ATT-B-LT (attention before
LT-CCP). In a second way, we add the attention module after the RNN
module, labeled as ATT-A-LT (attention after LT-CCP). As shown in Fig. 3(b) and
Fig. 3(a), the ACC is increased, and the corresponding MAPE is decreased. Both
ATT-B-LT and ATT-A-LT perform better than LT-CCP in terms of MAPE and
ACC. Introducing the attention module improves the ability of scienti c impact
prediction. The e ectiveness of the attention mechanism is veri ed.
(a) ACC comparision.</p>
        <p>(b) MAPE comparision.
(c) LT-CCP (RNN
LSTM).</p>
        <p>with
(d) ATT-B-LT.</p>
        <p>(e) ATT-A-LT (DLAM).</p>
        <p>In addition, we can see that the ATT-A-LT performs better than
ATT-BLT. When the attention model is applied after a deep learning model, it is more
e ective than the reverse combination. It indicates that the deep learning model
can learn the implicit features underlying the citation records, which boots the
performance.</p>
      </sec>
      <sec id="sec-5-2">
        <title>Analysis of the citation distribution. We illustrate the actual and the</title>
        <p>predicted citation distribution of LT-CCP (RNN with LSTM), ATT-B-LT, and
ATT-A-LT (DLAM) when t = 5 in Fig. 3(c), Fig. 3(d), and Fig. 3(e),
respectively. The LT-CCP (RNN with LSTM) illustrated in Fig. 3(c) shows the best
simulation of the power-law distribution. But the ATT-B-LT shown in Fig. 3(d)
and the ATT-A-LT (DLAM) shown in Fig. 3(e) present bad simulation of the
power-law distribution. The results show that LT-CCP (RNN with LSTM)
matches very well with that of real citations, but the ATT-B-LT and the
ATTA-LT (DLAM) don't. Usually, it is believed that the more similar the power-law
distribution, the whole result is better. At rst glance, it seems that LT-CCP
(RNN with LSTM) performs the best.</p>
        <p>However, the rst thought is wrong. As veri ed in Fig. 3(a) and Fig. 3(b), the
LT-CCP (RNN with LSTM) performs the worst. In fact, the LT-CCP only has a
better tting e ect on the papers with little citation counts. On the contrary, the
ATT-B-LT and ATT-A-LT (DLAM) have a better tting e ect on the highly
cited papers. The methods with attention mechanisms achieve better overall
performance than others. It is more accordant with practical prediction
requirements that a few papers occupy a vast number of citations. It further proves the
e ectiveness of the attention model. The experimental results indicate that we
need to change the xed pattern of thinking in quantifying long-term scienti c
impact.
6</p>
      </sec>
    </sec>
    <sec id="sec-6">
      <title>Conclusion</title>
      <p>Scienti c impact evaluation is always a critical point in decision-making
concerning recruitment and funding in the scienti c community. The rapid evolution of
scienti c research has been creating a massive volume of publications every year.
Among the many quanti cation measures of scienti c impact, citation count
stands out for its frequent use in the research community. Although the
peerreview process is the mainly reliable way of predicting a paper's future impact,
the ability to foresee the lasting impact based on citation records is increasingly
essential in the scienti c impact analysis in the era of big data.</p>
      <p>SciSci provides a quantitative understanding of the scienti c impact based on
big data empirical analysis. In this paper, the authors develop an attention
mechanism in long-term scienti c impact prediction and verify its e ectiveness. More
importantly, this paper provides us great insights into understanding the critical
factor in quantifying the long-term scienti c impact. Usually, researchers try to
make the predicted citation distribution similar to the original one. However,
the experimental results in the paper question this solution. In future research
work, we need to change the xed pattern of thinking in quantifying the
longterm scienti c impact and emphasize limited attention to better stand on giants'
shoulders.</p>
    </sec>
    <sec id="sec-7">
      <title>Acknowledgement</title>
      <p>The work is supported by the National Natural Science Foundation of China
(NSFC) under Grant No. 61806111 and NSFC for Distinguished Young Scholar
under Grant No. 61825602.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Abrishami</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Aliakbary</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          :
          <article-title>Predicting citation counts based on deep neural network learning techniques</article-title>
          .
          <source>Journal of Informetrics</source>
          <volume>13</volume>
          (
          <issue>2</issue>
          ),
          <volume>485</volume>
          {
          <fpage>499</fpage>
          (
          <year>2019</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Acuna</surname>
            ,
            <given-names>D.E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Allesina</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kording</surname>
            ,
            <given-names>K.P.</given-names>
          </string-name>
          :
          <article-title>Future impact: Predicting scienti c success</article-title>
          .
          <source>Nature</source>
          <volume>489</volume>
          (
          <issue>7415</issue>
          ),
          <volume>201</volume>
          (
          <year>2012</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Bao</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Shen</surname>
            ,
            <given-names>H.W.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Huang</surname>
            , J., Cheng,
            <given-names>X.Q.</given-names>
          </string-name>
          :
          <article-title>Popularity prediction in microblogging network: a case study on sina weibo</article-title>
          .
          <source>In: International Conference on World Wide Web</source>
          . pp.
          <volume>177</volume>
          {
          <issue>178</issue>
          (
          <year>2013</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Barabasi</surname>
            ,
            <given-names>A.L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Song</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Wang</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          : Publishing:
          <article-title>Handful of papers dominates citation</article-title>
          .
          <source>Nature</source>
          <volume>491</volume>
          (
          <issue>7422</issue>
          ),
          <volume>40</volume>
          (
          <year>2012</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Barbosa</surname>
            ,
            <given-names>S.D.J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Silveira</surname>
            ,
            <given-names>M.S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gasparini</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          :
          <article-title>What publications metadata tell us about the evolution of a scienti c community: the case of the brazilian human{ computer interaction conference series</article-title>
          .
          <source>Scientometrics</source>
          <volume>110</volume>
          (
          <issue>1</issue>
          ),
          <volume>275</volume>
          {
          <fpage>300</fpage>
          (
          <year>2017</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Bengio</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Simard</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Frasconi</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          :
          <article-title>Learning long-term dependencies with gradient descent is di cult</article-title>
          .
          <source>IEEE transactions on neural networks 5(2)</source>
          ,
          <volume>157</volume>
          {
          <fpage>166</fpage>
          (
          <year>1994</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>Cao</surname>
            ,
            <given-names>X.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chen</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Liu</surname>
            ,
            <given-names>K.R.:</given-names>
          </string-name>
          <article-title>A data analytic approach to quantifying scienti c impact</article-title>
          .
          <source>Journal of Informetrics</source>
          <volume>10</volume>
          (
          <issue>2</issue>
          ),
          <volume>471</volume>
          {
          <fpage>484</fpage>
          (
          <year>2016</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>Crane</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sornette</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          :
          <article-title>Robust dynamic classes revealed by measuring the response function of a social system</article-title>
          .
          <source>Proceedings of the National Academy of Sciences of the United States of America</source>
          <volume>105</volume>
          (
          <issue>41</issue>
          ),
          <volume>15649</volume>
          {
          <fpage>53</fpage>
          (
          <year>2008</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <surname>Cummings</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Nassar</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          :
          <article-title>Structured citation trend prediction using graph neural networks</article-title>
          .
          <source>In: ICASSP 2020-2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)</source>
          . pp.
          <volume>3897</volume>
          {
          <fpage>3901</fpage>
          .
          <string-name>
            <surname>IEEE</surname>
          </string-name>
          (
          <year>2020</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <surname>Dong</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Johnson</surname>
            ,
            <given-names>R.A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chawla</surname>
            ,
            <given-names>N.V.</given-names>
          </string-name>
          :
          <article-title>Will this paper increase your h-index?: Scienti c impact prediction</article-title>
          .
          <source>In: Proceedings of the eighth ACM international conference on web search and data mining</source>
          . pp.
          <volume>149</volume>
          {
          <issue>158</issue>
          (
          <year>2015</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <string-name>
            <surname>Dong</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ma</surname>
          </string-name>
          , H.,
          <string-name>
            <surname>Shen</surname>
            ,
            <given-names>Z.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Wang</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          :
          <article-title>A century of science: Globalization of scienti c collaborations, citations, and innovations</article-title>
          .
          <source>In: Proceedings of the 23rd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining</source>
          . pp.
          <volume>1437</volume>
          {
          <fpage>1446</fpage>
          .
          <string-name>
            <surname>ACM</surname>
          </string-name>
          (
          <year>2017</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          12.
          <string-name>
            <surname>Fortunato</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bergstrom</surname>
            ,
            <given-names>C.T.</given-names>
          </string-name>
          , Borner,
          <string-name>
            <given-names>K.</given-names>
            ,
            <surname>Evans</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.A.</given-names>
            ,
            <surname>Helbing</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            ,
            <surname>Milojevic</surname>
          </string-name>
          ,
          <string-name>
            <surname>S.</surname>
          </string-name>
          , et al.:
          <source>Science of science. Science</source>
          <volume>359</volume>
          (
          <issue>6379</issue>
          ) (
          <year>2018</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          13.
          <string-name>
            <surname>Franceschet</surname>
            ,
            <given-names>M.:</given-names>
          </string-name>
          <article-title>The role of conference publications in cs</article-title>
          .
          <source>Communications of the ACM</source>
          <volume>53</volume>
          (
          <issue>12</issue>
          ),
          <volume>129</volume>
          {
          <fpage>132</fpage>
          (
          <year>2010</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          14.
          <string-name>
            <surname>Frank</surname>
            ,
            <given-names>M.R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Wang</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Cebrian</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rahwan</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          :
          <article-title>The evolution of citation graphs in arti cial intelligence research</article-title>
          .
          <source>Nature Machine Intelligence</source>
          <volume>1</volume>
          (
          <issue>2</issue>
          ),
          <volume>79</volume>
          (
          <year>2019</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          15.
          <string-name>
            <surname>Freyne</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Coyle</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Smyth</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Cunningham</surname>
            ,
            <given-names>P.:</given-names>
          </string-name>
          <article-title>Relative status of journal and conference publications in computer science</article-title>
          .
          <source>Communications of the ACM</source>
          <volume>53</volume>
          (
          <issue>11</issue>
          ),
          <volume>124</volume>
          {
          <fpage>132</fpage>
          (
          <year>2010</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          16. Gar eld, E.:
          <article-title>Impact factors, and why they won't go away</article-title>
          .
          <source>Nature</source>
          <volume>411</volume>
          (
          <issue>6837</issue>
          ),
          <volume>522</volume>
          (
          <year>2001</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          17.
          <string-name>
            <surname>Greene</surname>
            ,
            <given-names>M.:</given-names>
          </string-name>
          <article-title>The demise of the lone author</article-title>
          .
          <source>Nature</source>
          <volume>450</volume>
          (
          <issue>7173</issue>
          ),
          <volume>1165</volume>
          (
          <year>2007</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          18.
          <string-name>
            <surname>Hirsch</surname>
            ,
            <given-names>J.E.</given-names>
          </string-name>
          :
          <article-title>An index to quantify an individual's scienti c research output</article-title>
          .
          <source>Proceedings of the National Academy of Sciences of the United States of America</source>
          <volume>102</volume>
          (
          <issue>46</issue>
          ),
          <volume>16569</volume>
          {
          <fpage>16572</fpage>
          (
          <year>2005</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          19.
          <string-name>
            <surname>Hunter</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Smyth</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Vu</surname>
            ,
            <given-names>D.Q.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Asuncion</surname>
            ,
            <given-names>A.U.</given-names>
          </string-name>
          :
          <article-title>Dynamic egocentric models for citation networks</article-title>
          .
          <source>In: Proceedings of the 28th International Conference on Machine Learning (ICML-11)</source>
          . pp.
          <volume>857</volume>
          {
          <issue>864</issue>
          (
          <year>2011</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          20.
          <string-name>
            <surname>Kuhn</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Perc</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Helbing</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          :
          <article-title>Inheritance patterns in citation networks reveal scienti c memes</article-title>
          .
          <source>Physical Review X</source>
          <volume>4</volume>
          (
          <issue>4</issue>
          ),
          <volume>041036</volume>
          (
          <year>2014</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          21.
          <string-name>
            <surname>Lariviere</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Haustein</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          , Borner,
          <string-name>
            <surname>K.</surname>
          </string-name>
          :
          <article-title>Long-distance interdisciplinarity leads to higher scienti c impact</article-title>
          .
          <source>Plos one 10(3)</source>
          (
          <year>2015</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          22.
          <string-name>
            <surname>Li</surname>
            ,
            <given-names>Z.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>He</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Tian</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chen</surname>
            ,
            <given-names>W.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Qin</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Wang</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Liu</surname>
          </string-name>
          , T.Y.:
          <article-title>Towards binaryvalued gates for robust lstm training</article-title>
          .
          <source>In: International Conference on Machine Learning</source>
          . pp.
          <volume>3001</volume>
          {
          <issue>3010</issue>
          (
          <year>2018</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          23.
          <string-name>
            <surname>Ma</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Uzzi</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          :
          <article-title>Scienti c prize network predicts who pushes the boundaries of science</article-title>
          .
          <source>Proceedings of the National Academy of Sciences</source>
          <volume>115</volume>
          (
          <issue>50</issue>
          ),
          <volume>12608</volume>
          {
          <fpage>12615</fpage>
          (
          <year>2018</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          24.
          <string-name>
            <surname>Margolis</surname>
          </string-name>
          , J.:
          <article-title>Citation indexing and evaluation of scienti c papers</article-title>
          .
          <source>Science</source>
          <volume>155</volume>
          (
          <issue>3767</issue>
          ),
          <volume>1213</volume>
          {
          <fpage>1219</fpage>
          (
          <year>1967</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref25">
        <mixed-citation>
          25.
          <string-name>
            <surname>Matsubara</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sakurai</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Prakash</surname>
            ,
            <given-names>B.A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Li</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Faloutsos</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          :
          <article-title>Rise and fall patterns of information di usion: model and implications</article-title>
          .
          <source>In: ACM SIGKDD International Conference on Knowledge Discovery and Data Mining</source>
          . pp.
          <volume>6</volume>
          {
          <issue>14</issue>
          (
          <year>2012</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref26">
        <mixed-citation>
          26.
          <string-name>
            <surname>Pan</surname>
            ,
            <given-names>R.K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kaski</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Fortunato</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          :
          <article-title>World citation and collaboration networks: uncovering the role of geography in science</article-title>
          .
          <source>Scienti c reports 2</source>
          ,
          <issue>902</issue>
          (
          <year>2012</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref27">
        <mixed-citation>
          27.
          <string-name>
            <surname>Peng</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          :
          <article-title>Modeling and predicting popularity dynamics via an in uence-based selfexcited hawkes process</article-title>
          .
          <source>In: ACM International on Conference on Information and Knowledge Management</source>
          . pp.
          <year>1897</year>
          {
          <year>1900</year>
          (
          <year>2016</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref28">
        <mixed-citation>
          28.
          <string-name>
            <surname>Radicchi</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Fortunato</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Castellano</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          :
          <article-title>Universality of citation distributions: Toward an objective measure of scienti c impact</article-title>
          .
          <source>Proceedings of the National Academy of Sciences</source>
          <volume>105</volume>
          (
          <issue>45</issue>
          ),
          <volume>17268</volume>
          {
          <fpage>17272</fpage>
          (
          <year>2008</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref29">
        <mixed-citation>
          29.
          <string-name>
            <surname>Rzhetsky</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Foster</surname>
            ,
            <given-names>J.G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Foster</surname>
            ,
            <given-names>I.T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Evans</surname>
            ,
            <given-names>J.A.</given-names>
          </string-name>
          :
          <article-title>Choosing experiments to accelerate collective discovery</article-title>
          .
          <source>Proceedings of the National Academy of Sciences</source>
          <volume>112</volume>
          (
          <issue>47</issue>
          ),
          <volume>14569</volume>
          {
          <fpage>14574</fpage>
          (
          <year>2015</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref30">
        <mixed-citation>
          30.
          <string-name>
            <surname>Sekercioglu</surname>
            ,
            <given-names>C.H.</given-names>
          </string-name>
          :
          <article-title>Quantifying coauthor contributions</article-title>
          .
          <source>Science</source>
          <volume>322</volume>
          (
          <issue>5900</issue>
          ),
          <volume>371</volume>
          {
          <fpage>371</fpage>
          (
          <year>2008</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref31">
        <mixed-citation>
          31.
          <string-name>
            <surname>Shen</surname>
            ,
            <given-names>H.W.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Barabasi</surname>
            ,
            <given-names>A.L.</given-names>
          </string-name>
          :
          <article-title>Collective credit allocation in science</article-title>
          .
          <source>Proceedings of the National Academy of Sciences</source>
          <volume>111</volume>
          (
          <issue>34</issue>
          ),
          <volume>12325</volume>
          {
          <fpage>12330</fpage>
          (
          <year>2014</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref32">
        <mixed-citation>
          32.
          <string-name>
            <surname>Shen</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Wang</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Song</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Barabasi</surname>
            ,
            <given-names>A.L.</given-names>
          </string-name>
          :
          <article-title>Modeling and predicting popularity dynamics via reinforced poisson processes</article-title>
          .
          <source>In: Twenty-eighth AAAI conference on arti cial intelligence</source>
          (
          <year>2014</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref33">
        <mixed-citation>
          33.
          <string-name>
            <surname>Shen</surname>
            ,
            <given-names>Z.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Yang</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Wu</surname>
          </string-name>
          , J.:
          <article-title>Lognormal distribution of citation counts is the reason for the relation between impact factors and citation success index</article-title>
          .
          <source>Journal of Informetrics</source>
          <volume>12</volume>
          (
          <issue>1</issue>
          ),
          <volume>153</volume>
          {
          <fpage>157</fpage>
          (
          <year>2018</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref34">
        <mixed-citation>
          34.
          <string-name>
            <surname>Sinatra</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Deville</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Szell</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Wang</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Barabasi</surname>
            ,
            <given-names>A.L.:</given-names>
          </string-name>
          <article-title>A century of physics</article-title>
          .
          <source>Nature Physics</source>
          <volume>11</volume>
          (
          <issue>10</issue>
          ),
          <volume>791</volume>
          (
          <year>2015</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref35">
        <mixed-citation>
          35.
          <string-name>
            <surname>Sinatra</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Wang</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Deville</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Song</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Barabasi</surname>
            ,
            <given-names>A.L.</given-names>
          </string-name>
          :
          <article-title>Quantifying the evolution of individual scienti c impact</article-title>
          .
          <source>Science</source>
          <volume>354</volume>
          (
          <issue>6312</issue>
          ),
          <year>aaf5239</year>
          (
          <year>2016</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref36">
        <mixed-citation>
          36.
          <string-name>
            <surname>Sutskever</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Vinyals</surname>
            ,
            <given-names>O.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Le</surname>
            ,
            <given-names>Q.V.</given-names>
          </string-name>
          :
          <article-title>Sequence to sequence learning with neural networks</article-title>
          .
          <source>In: Advances in Neural Information Processing Systems</source>
          . pp.
          <volume>3104</volume>
          {
          <issue>3112</issue>
          (
          <year>2014</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref37">
        <mixed-citation>
          37.
          <string-name>
            <surname>Szabo</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Huberman</surname>
            ,
            <given-names>B.A.</given-names>
          </string-name>
          :
          <article-title>Predicting the popularity of online content</article-title>
          , vol.
          <volume>53</volume>
          .
          <article-title>Communications of the ACM (</article-title>
          <year>2010</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref38">
        <mixed-citation>
          38.
          <string-name>
            <surname>Tang</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          , Zhang, J.,
          <string-name>
            <surname>Yao</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Li</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          , Zhang,
          <string-name>
            <given-names>L.</given-names>
            ,
            <surname>Su</surname>
          </string-name>
          ,
          <string-name>
            <surname>Z.</surname>
          </string-name>
          :
          <article-title>Arnetminer:extraction and mining of academic social networks</article-title>
          .
          <source>In: ACM SIGKDD International Conference on Knowledge Discovery and Data Mining</source>
          . pp.
          <volume>990</volume>
          {
          <issue>998</issue>
          (
          <year>2008</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref39">
        <mixed-citation>
          39.
          <string-name>
            <surname>Vu</surname>
            ,
            <given-names>D.Q.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Asuncion</surname>
            ,
            <given-names>A.U.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hunter</surname>
            ,
            <given-names>D.R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Smyth</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          :
          <article-title>Dynamic egocentric models for citation networks</article-title>
          .
          <source>In: International Conference on International Conference on Machine Learning</source>
          . pp.
          <volume>857</volume>
          {
          <issue>864</issue>
          (
          <year>2011</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref40">
        <mixed-citation>
          40.
          <string-name>
            <surname>Wang</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Song</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Barabasi</surname>
            ,
            <given-names>A.L.</given-names>
          </string-name>
          :
          <article-title>Quantifying long-term scienti c impact</article-title>
          .
          <source>Science</source>
          <volume>342</volume>
          (
          <issue>6154</issue>
          ),
          <volume>127</volume>
          {
          <fpage>132</fpage>
          (
          <year>2013</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref41">
        <mixed-citation>
          41.
          <string-name>
            <surname>Wang</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Mei</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hicks</surname>
            ,
            <given-names>D.:</given-names>
          </string-name>
          <article-title>Comment on \quantifying long-term scienti c impact"</article-title>
          .
          <source>Science</source>
          <volume>345</volume>
          (
          <issue>6193</issue>
          ),
          <volume>149</volume>
          {
          <fpage>149</fpage>
          (
          <year>2014</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref42">
        <mixed-citation>
          42.
          <string-name>
            <surname>Xiao</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Yan</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Li</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Jin</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          :
          <article-title>On modeling and predicting individual paper citation count over time</article-title>
          .
          <source>In: Twenty-Fifth International Joint Conference on Arti cial Intelligence (IJCAI-16)</source>
          . pp.
          <volume>2676</volume>
          {
          <issue>2682</issue>
          (
          <year>2016</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref43">
        <mixed-citation>
          43.
          <string-name>
            <surname>Xiao</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Yan</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Yang</surname>
            ,
            <given-names>X.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zha</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chu</surname>
            ,
            <given-names>S.M.</given-names>
          </string-name>
          :
          <article-title>Modeling the intensity function of point process via recurrent neural networks</article-title>
          .
          <source>In: Thirty-First AAAI Conference on Arti cial Intelligence</source>
          . vol.
          <volume>17</volume>
          , pp.
          <volume>1597</volume>
          {
          <issue>1603</issue>
          (
          <year>2017</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref44">
        <mixed-citation>
          44.
          <string-name>
            <surname>Yan</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Huang</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Tang</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          , Zhang,
          <string-name>
            <given-names>Y.</given-names>
            ,
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <surname>X.</surname>
          </string-name>
          :
          <article-title>To better stand on the shoulder of giants</article-title>
          .
          <source>In: Proceedings of the 12th ACM/IEEE-CS joint conference on Digital Libraries</source>
          . pp.
          <volume>51</volume>
          {
          <issue>60</issue>
          (
          <year>2012</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref45">
        <mixed-citation>
          45.
          <string-name>
            <surname>Yan</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Tang</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Liu</surname>
            ,
            <given-names>X.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Shan</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Li</surname>
            ,
            <given-names>X.</given-names>
          </string-name>
          :
          <article-title>Citation count prediction: learning to estimate future citations for literature</article-title>
          .
          <source>In: ACM International Conference on Information and Knowledge Management</source>
          . pp.
          <volume>1247</volume>
          {
          <issue>1252</issue>
          (
          <year>2011</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref46">
        <mixed-citation>
          46.
          <string-name>
            <surname>Yegros-Yegros</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rafols</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>D'Este</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          :
          <article-title>Does interdisciplinary research lead to higher citation impact? the di erent e ect of proximal and distal interdisciplinarity</article-title>
          .
          <source>PloS one 10(8)</source>
          (
          <year>2015</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref47">
        <mixed-citation>
          47.
          <string-name>
            <surname>Yin</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Wang</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          :
          <article-title>The time dimension of science: Connecting the past to the future</article-title>
          .
          <source>Journal of Informetrics</source>
          <volume>11</volume>
          (
          <issue>2</issue>
          ),
          <volume>608</volume>
          {
          <fpage>621</fpage>
          (
          <year>2017</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref48">
        <mixed-citation>
          48.
          <string-name>
            <surname>Zhu</surname>
            ,
            <given-names>X.P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ban</surname>
            ,
            <given-names>Z.</given-names>
          </string-name>
          :
          <article-title>Citation count prediction based on academic network features</article-title>
          .
          <source>In: 2018 IEEE 32nd International Conference on Advanced Information Networking and Applications (AINA)</source>
          . pp.
          <volume>534</volume>
          {
          <fpage>541</fpage>
          .
          <string-name>
            <surname>IEEE</surname>
          </string-name>
          (
          <year>2018</year>
          )
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>