<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Can we do better than Co-Citations? - Bringing Citation Proximity Analysis from idea to practice in research article recommendation</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Petr Knoth</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Anita Khadka</string-name>
          <email>anita.khadkag@open.ac.uk</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>The Open University</institution>
          ,
          <country country="UK">UK</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>In this paper, we build on the idea of Citation Proximity Analysis (CPA), originally introduced in [1], by developing a step by step scalable approach for building CPA-based recommender systems. As part of this approach, we introduce three new proximity functions, extending the basic assumption of co-citation analysis (stating that the more often two articles are co-cited in a document, the more likely they are related) to take the distance between the co-cited documents into account. Asking the question of whether CPA can outperform co-citation analysis in recommender systems, we have built a CPA based recommender system from a corpus of 368,385 full-texts articles and conducted a user survey to perform an initial evaluation. Two of our three proximity functions used within CPA outperform co-citations on our evaluation dataset.</p>
      </abstract>
      <kwd-group>
        <kwd>Citation Proximity Analysis</kwd>
        <kwd>Co-Citation Analysis</kwd>
        <kwd>Recommender System</kwd>
        <kwd>Information Retrieval</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>
        The number of scholarly articles is increasing exponentially every year, according
to Gipp and Beel [
        <xref ref-type="bibr" rid="ref1 ref2">1,2</xref>
        ] by 3.7%, Bornmann and Mutz claim 8%- 9% (up-to 2010)
[
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]. This creates challenges for researchers to stay in touch with new relevant
articles in their domain.
      </p>
      <p>
        West et al. [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ] state that searching for a particular paper knowing that it
exists, has become trivial except for the pay-wall. Searching for (unknown) but
relevant papers is a challenging task that is at the very centre of the research
process (for example, the task of reviewing the state-of-the-art in a particular
domain). To researchers, recommender systems can help them to stay in touch
with the latest relevant papers in their eld. To authors, recommender systems
can help cater their papers to the relevant audiences resulting in an increased
number of reads and therefore more e ective dissemination of knowledge.
      </p>
      <p>
        In the past, many academic recommender systems unless developed by
publishers or corporations that negotiated access to scienti c literature, faced
several limitations. For instance, limitations of machine access to the full texts of
papers and sometimes even the citation information. Consequently, full-text
features have been so far relatively unexplored. Most of the recommender systems
Recommender systems suggest relevant and useful information to match the need
of its users. These systems are popular in both commerce and academia. Over the
years, various metrics and approaches such as Collaborative Filtering (CF) [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ],
Content Based Filtering (CBF) [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ] [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ], Graph based recommendations [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ][
        <xref ref-type="bibr" rid="ref12">12</xref>
        ]
      </p>
    </sec>
    <sec id="sec-2">
      <title>Document A</title>
      <p>Lorem ipsum dolor sit amet, consectetuer
adipiscing elit. In hac habitas e platea dictumst.
In leo ante, venenatis eu, volutpat ut, imperdiet
auctor, enim. Aenean scelerisque metus eget
sem. Nul a sed lacus. Mauris tempus diam.
Curabitur accumsan felis in erat. Integer porta.
Praesent scelerisque. Quisque pretium rutrum
ligula. Praesent scelerisque. Duis sem velit,
ultrices et, fermentum auctor, rhoncus ut, ligula.
Curabitur risus urna, placerat et, luctus pulvinar,
auctor vel, orci. Praesent semper, neque vel
condimentum hendrerit, lectus elit pretium
ligula, nec consequat nisl velit at dui. Nam
mas a turpis, nonummy et, consectetuer id,
placerat ac, ante. Sed at turpis vitae velit
euismod aliquet. Maecenas justo. Donec ut urna.
Aenean luctus vulputate turpis.</p>
      <p>Pel entesque condimentum felis a sem. Etiam
pharetra lacus sed velit imperdiet bibendum.
Cras ac enim vel dui vestibulum suscipit. Aenean
turpis ipsum, rhoncus vitae, posuere vitae,
euismod sed, ligula. Cras gravida. Sed
elementum, felis quis port itor sol icitudin, augue</p>
      <p>Lorem ipsum dolor sit amet,
consectetur adipiscing elit. Vestibulum
interdum</p>
      <p>a augue accumsan suscipit.</p>
      <p>Phasel us id erat neque. Aliquam
eu
pretium
enim.</p>
      <p>Maecenas
semper
pel entesque</p>
      <p>
        mi ac rhoncus. Donec
[
        <xref ref-type="bibr" rid="ref1">1</xref>
        ] Lorem ipsum
dolor sit [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ].amet,
consectetur adipiscing
elit verbatim
Proin in
      </p>
      <p>mat is elit. Fusce faucibus
mauris ut arcu consequat ultrices.</p>
      <p>
        Donec eu lectus ultrices justo
luctus placerat. Phasel us scelerisque
sapien non blandit lobortis ipsum
Lorem
ipsum
dolor
sit
[
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]
amet,
consectetur adipiscing elit verbatim
[
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]
[
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]
[
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]
.
.
      </p>
      <p>Document C
Document A
Document B
.</p>
      <p>.</p>
    </sec>
    <sec id="sec-3">
      <title>Document B</title>
      <sec id="sec-3-1">
        <title>Citation Proximity</title>
      </sec>
      <sec id="sec-3-2">
        <title>Analysis (CPA) conceptualised by Gipp and Beel (b)</title>
      </sec>
      <sec id="sec-3-3">
        <title>Citation Proximity</title>
        <p>
          Analysis (CPA). Length
of
solid
underlined Red text
signi
[
          <xref ref-type="bibr" rid="ref1">1</xref>
          ]
es
d12 and Length
of dashed underlined Green text
signi
es d23. (Best
viewed in
colour)
and
di
        </p>
        <p>erent citation based concepts have b een implemented and evaluated for
scholarly
recommendation,</p>
        <p>either using full-texts or with meta-data. While
is the state-of-the-art approach for recommending items in commerce, it is
CF
also
prone to some limitations such as the cold
start problem, i.e. the need to have
go o d coverage of ratings. In the domain
of scholarly pap ers, it is
typically
di
cult to obtain ratings. Agarwal et</p>
        <p>
          al. [
          <xref ref-type="bibr" rid="ref13">13</xref>
          ] argue that the research pap er domain
has
relatively
less
users
compared to the large numb er
of online
research
pap ers. This
intro duces
high
dimensionality
and sparsity, which p erforms p o orly
when
algorithms such as k nearest neighb our and CF are
applied. To combat
this issue, their Subspace Clustering Algorithm (SCuBA) approach reduces the
dimensionality
of the subspace [
          <xref ref-type="bibr" rid="ref13">13</xref>
          ]. Sugiyama and Kan in
[
          <xref ref-type="bibr" rid="ref14">14</xref>
          ] used CF to
discover the p otential citation pap ers for a do cument by creating user pro
les from
the
researcher's
list
of
publication.
        </p>
        <p>They
also showed that \Conclusion"
section weight more
than other
section for computing e
ectiveness
of the pap er.</p>
        <sec id="sec-3-3-1">
          <title>However, Nascimento et</title>
          <p>
            al. [
            <xref ref-type="bibr" rid="ref9">9</xref>
            ] argued that title
of any do cument weighs more.
          </p>
          <p>Another p opular approach is</p>
          <p>
            citation based approach, such as, Co-Citation
which was prop osed by Small [
            <xref ref-type="bibr" rid="ref15">15</xref>
            ] and Marshakova [
            <xref ref-type="bibr" rid="ref16">16</xref>
            ]
separately
in
the 70s.
          </p>
          <p>This approach is a well-established in
scholarly
recommendation system by now.</p>
          <p>West
et</p>
          <p>
            al. [
            <xref ref-type="bibr" rid="ref4">4</xref>
            ] have prop osed an automation system which has remarkable
success over the system \Co-download" which uses collab orative
ltering and the
\Co-Citation" system. Although, results are remarkable it may di
er in
di
erent
databases and result could
b e
di
erent in
cross-disciplinary database.
Furthermore, Tran et
          </p>
          <p>
            al. [
            <xref ref-type="bibr" rid="ref17">17</xref>
            ] followed the concept of CPA and used graph based
similarity
measure to demonstrate that do cuments are
more
related in
          </p>
        </sec>
        <sec id="sec-3-3-2">
          <title>Sentence Level</title>
          <p>
            Co-Citation than Paper Level Co-Citation. Gipp et al. [
            <xref ref-type="bibr" rid="ref6">6</xref>
            ] introduced a hybrid
research paper recommender system by using the concept of in-text Impact
factor (ICFA) and in-text citation distance analysis (ICDA). Similarly, Schwarzer
et al. [
            <xref ref-type="bibr" rid="ref18">18</xref>
            ] used the concept of CPA for articles recommendation using links
(SeeAlso links from wikipedia articles) instead of citations for article
recommendations. They claimed that citation based approaches have di erent strength to
text-based approach like More Like This (MLT) and suggested that combining
them could supersede one's caveat with other's advantages. And, Gipp et al.
[
            <xref ref-type="bibr" rid="ref19">19</xref>
            ] evaluated and analysed citation based approach and compared with
character based approach to detect plagiarism and showed citation based approach
provided preferable results than character based approach.
          </p>
          <p>Researchers have tried and tested citation-based approaches which are
proving in some situations, such as for expert users, more powerful than purely text
based approaches. Consequently, it is worthwhile to further explore the role of
citation proximity in recommender systems based on co-citations.
3</p>
          <p>Method
The design of our CPA-based recommender system consists of ve modules
(Extracting Citation Information, Citation Information Normalisation, Sparse
Matrix, Citation Proximity Analysis, Recommendation) depicted in Figure 2. We
start by extracting citation information, including the positions of citation
anchors in the body of the full-texts. We then normalise the extracted reference
strings (typically found at the end of each paper) trying to detect and merge
those referring to the same canonical document. The output of this process is
a sparse square matrix where rows and columns correspond to unique
references found in the full-texts of research papers in the original collection. These
unique references form the set of recommendable items. Each cell of the matrix
contains all the co-occurrences of the corresponding references in any of the
research papers in the original collection, including their character positions. The
information stored in each cell is passed to the CPA component which applies a
proximity function to produce a CPA value.</p>
          <p>To produce recommendations for a given paper reference, it is necessary to
look up a corresponding row (or column), calculate CPA values for each
nonempty cell and select n references with the highest CPA scores as the
recommendations. Next sections describe the process in more detail.
3.1</p>
          <p>
            Extracting citation information
The aim of this component is to:
{ extract reference strings, typically appearing at the end of research papers,
{ identify and parse the reference structure, such as article title, authors,
publication year or DOI, of each reference and
{ detect and extract the character o set of each citation occurrence (citance)
on the body of the research paper.
Vestibulum ante ipsum primis in fauc
us orci luctus et ultrices posuere a cu
cubil a [
            <xref ref-type="bibr" rid="ref1">1</xref>
            ] Curae Donec velit neque,
porta vel ultricies ligula et. [
            <xref ref-type="bibr" rid="ref2">2</xref>
            ]
auctor sit amet aliquam vel ul amcorper sit
amet ligula. Curabitur non
nul a sit amet nisl tempus conval is quis ac
lectus. [
            <xref ref-type="bibr" rid="ref1">1</xref>
            ] Praesent sapien mas a, conval is a
pel entesque.
          </p>
          <p>Mo del
dia
gram
of
the
pro
cess
carried
out
for
recommending
research
articles
The
input
of
this
comp onent
is
the
full
text
of
a
research
pap er
and
the
output
is</p>
          <p>a tuple:
(referenceId,
title,
authors
[],
characterO
sets
[],
yearPublished,
sourceId)
(1)
er
is
of
an
each
array
the
is
where
referenceId
is
an</p>
          <p>identi
of
the
referenced
article,
authors
reference
strin
g,
title
is
the</p>
          <p>title
of
author
name
sets
is
an
array
of
o
sets
of
citances
in
the
b o dy
of
given
reference,
yearPublished
the
year
of
publication
unique
identi
er
of
the
full-text
research
pap er
from
which
strings,
the
full
chara</p>
          <p>ctext
and
this
sourceID
reference
Citation
information</p>
          <p>normalisation
citation
normalisation
comp onent takes
as an
input,
a
dataset
of
the
tuples
sparse
de</p>
          <p>ned
matrix.</p>
          <p>in
As
the</p>
          <p>Equation
1,</p>
          <p>deduplicates
deduplication
is
not
the
key
them
fo cus
naive
deduplication
metho d
that
targets
precision
at
and</p>
          <p>represents
of
the
this
pap er,
we</p>
          <p>use
exp ense
of
recall.</p>
          <p>Firstly,
uments
we
removed the sp ecial
characters
and spaces
from
the
title
of
each
and
group ed
records
with
the
same
title,
publication
date
and
at
do
cleast
one matching author and singled out only one record from the group. By doing
this, we ended up only unique documents in the dataset.</p>
          <p>After normalisation, we represent the output as a Citation-Positions matrix.
Each cell Vi;j in this matrix contains character o set information of all the
citances where a normalised reference i co-occurs with a normalised reference
j in a given source document. This is all that is needed as input for the CPA
module which then calculates the CPA value based on a given proximity function.
3.3</p>
          <p>Proximity functions
The CPA proximity function takes as an input, a set of character o set distances
and produces a single value. The higher the value the higher the relevance. In a
research paper, a reference can be cited multiple times leading to a set of pair
distances for the co-cited pair. Additionally, references can be co-cited in multiple
source documents. The intuition behind the proximity metric is that higher
number of co-citations as well as closer proximity should lead to an increased
relevance.</p>
          <p>In this work, we have de ned and experimented with the following proximity
measures: MinProx, SumProx and MeanProx described below. Our baseline
cocitation method, which does not use any proximity information, can be in this
framework de ned as:
cocitbaabseline = jDocja2Doc^b2Doc;
(2)</p>
          <p>This means that the number of co-citations is de ned by the number of source
documents where references a and b co-occur.
3.3.1 MinProx: Uses only the distance of the closest co-occurrence in the
denominator. During distance computation, one of the hypothesis is the distance
between co-cited documents are never Zero or One. For example, the extreme
case of citations being cited together will be like [X; Y ] and this will always have
a separator character between them. If reference \X" has character o set 102
and reference \Y" has character o set 104 then the distance will be 104-102=2.</p>
          <p>As, in the following proximity functions (3, 4, 5), logarithm is applied for
smoothing large distances. The nominator is equal to the baseline co-citation
measure.</p>
          <p>proxaMbin =</p>
          <p>jDocja2Doc^b2Doc
log(minfd1ab; :::; danbg)
where dab denotes the rst distance between the co-cited references a and b and
1
danb denotes the last distance between them.
3.3.2 SumProx: Uses the sum of the logs of all the co-cited distances in the
denominator.</p>
          <p>(3)
proxaSbum = jDocja2Doc^b2Doc</p>
          <p>Pn
i=1 log(diab)
where di is the ith distance between the co-cited documents a and b
3.3.3 MeanProx: Uses the log of the mean of all the co-cited distances in
the denominator.</p>
          <p>proxaMbean =</p>
          <p>jDocja2Doc^b2Doc
log(meanfd1ab; :::; danbg)
(4)
(5)
where dn is the last distance between the co-cited documents a and b.
4</p>
          <p>
            Experiments
To evaluate CPA against the co-citations baseline, we have developed and \trained"
a recommender system, as described in Section 3, on a sample collection
scienti c documents from CORE [
            <xref ref-type="bibr" rid="ref20">20</xref>
            ] 1. The evaluation dataset consisted of a set of
recommendations retrieved by each evaluated variation of the recommender in
response to di erent queries (research papers for which recommendations should
be produced). Several human judges were asked to provide binary relevance
judgments which form the evaluation ground truth. We will now provide more details
of the experimental setup, evaluation dataset and results.
4.1
          </p>
          <p>
            Experimental system
We used GeneRation Of BIbliographic Data (GROBID) [
            <xref ref-type="bibr" rid="ref21">21</xref>
            ] to convert research
papers in the PDF format into the Text Encoding Initiatives (TEI) format from
which we extracted the required citation information as speci ed in Section 3.1.
The processing of this collection took about an hour and a half on a quad core
system with 20 GB of memory.
          </p>
          <p>As expected, GROBID could not successfully process all the documents (due
to PDFs that were scans, badly encoded PDF les or citation extraction failing
on valid PDFs). We were con dent of 368,385 documents which yielded
citations information along with their positions and used these in our experiment.
The resulting set consisted of 6,609,147 references. This means we have obtained
on average 18 references per document. Figure 3(a) shows the probability
distribution of the number of citation mentions (citances) in a document while
Figure 3(b) shows the probability distribution of the number of references in a
document.</p>
          <p>For an e cient implementation of the subsequent components, i.e.
normalisation and CPA calculation, we made use of the MapReduce paradigm and
implemented the solution using the data ow language Apache Pig. The
normalization step took approximately 2 minutes with 100 parallel processes and
1 https://core.ac.uk/services#datasets
0.030
t
n
e
cum0.025
o
d
r
e
p
se0.020
ffrcenR0.015
e
e
o
n
o
it
u
ib0.010
itrsd
y
lit
ib0.005
a
b
o
r</p>
          <p>P0.0000
20 # of extracted citation mentions 80
40 60
100
20
4#0of References
60
80
100
resulted in a citation-positions sparse matrix containing 142,157,561 co-cited
pairs. Figure 4 shows the distribution plots of relevance score produced by each
proximity functions, we have de ned along the baseline standard's results.
4.2</p>
          <p>
            Evaluation dataset
Evaluating recommender system is generally a challenging task due to the variety
of criteria and goals recommendation methods can be developed with, such as
recency, serendipity, relevance, etc. Consequently, some evaluation datasets can
work better with one metric or task while other's may not [
            <xref ref-type="bibr" rid="ref22 ref23">22,23</xref>
            ]. We created an
evaluation dataset on the basis of the relevance of recommendations to the target
evaluator's expertise. As CPA is still in its relative infancy, our goal was to create
a small-scale pilot evaluation initially. If encouraging results are produced by the
tested method, this will be a signal for us to extend this study and re-evaluate
on a larger dataset.
          </p>
          <p>For the evaluation, we randomly selected 6 sample documents from the
dataset from the area of \Computer Science" (specially \Data mining" and
\Information Retrieval") with which the annotators were familiar with. Ten
annotators (survey participants) from computer science department (working on \Data
mining" and \Information Retrieval") were asked to provide binary relevance
judgments on each recommendation o ered by each evaluated system. As, we
had 4 evaluated metrics, 5 recommendations for each sample document and ten
participants, this yields 6 5 4 10 = 1; 200 individual relevance judgments.
4.3</p>
          <p>Results
We have calculated precision at 3 di erent precision levels as shown in Table 1.
Our experimental results indicate that proximity information helps in producing
(a) proximities using minimum co-(b) proximities using summation of
cited Distance metric co-cited Distances metric
(c) proximities using mean of co-(d) proximities using frequency of
cited Distances metric co-citation metric
better recommendations than the baseline co-citation approach. More speci
cally, out of the three proximity functions two, SumP rox and M eanP rox,
outperform the Baseline. The improvement over the tested dataset over the baseline
for P@5 (from 0:27 to 0:34) corresponds to a more than 25% improvement.</p>
          <p>To assess the subjectivity of the task, we have also calculated inter-rater
reliability statistic to weight the agreement between the contributors. To do so,
we have used Fleiss's as follows:
=</p>
          <p>P
1</p>
          <p>Pe</p>
          <p>Pe
where, Pe denotes the observed agreement and P denotes the probability of
chance agreement. Hence, (1 P ) is the degree of agreement which is obtainable
by chance and P Pe gives the degree of agreement which is actually obtained.
For all the sample data and its recommendations, we have observed = 0:25
suggesting a fair agreement.
(6)</p>
          <p>
            Discussion and future work
There are several things we would like to address in our future work as this is an
initial practical observation of the CPA concept. Firstly, our current proximity
functions are based on working with absolute character o sets, i.e. our distance
measure is a character distance of the citances. In the future, we would like
to extend our work by experimenting with distance measures that re ect the
lexico-syntactic structure of language. For example, we could use information on
whether two papers have been co-cited in the same sentence clause, sentence,
paragraph, section, table, etc. Additionally, we would also like to compute the
CPI values as conceptualised by Beel and Gipp in [
            <xref ref-type="bibr" rid="ref1">1</xref>
            ] and compare the results
by extending the combinations of the metrics like multiplication of CPA and
Co-Citation.
          </p>
          <p>Secondly, our current method treats all co-citations equal. However, it would
be interesting to explore how the concept of \authority" might be applied in this
problem. For example, if one paper is co-cited (or even compared/contrasted)
with another highly signi cant work (e.g. famous authors, policy document or
a high value determined by any particular scientometric method) or if the
cocitation is present in a highly signi cant work, this information should in uence
the strength of this co-citation evidence and be e ectively used in the
recommender.</p>
          <p>Thirdly, more work is needed to understand the impact of document length
on co-citation analysis approaches as long documents, like theses or books,
produce signi cantly higher numbers of co-citations than shorter documents (short,
long papers, demo), hence they have a higher impact on the results of the
recommender.</p>
          <p>Finally, we would also like to use machine learning algorithms using position
of the citation as one of the features to improve the weighing process. However,
to do so, we will be needing big dataset with the ground truth of recommendation
results for training the system so collecting such dataset will be one of the major
hurdles.
6</p>
          <p>Conclusion
In this paper, we performed an experiment to convert the concept of CPA into
practice and benchmark this approach against co-citation Analysis. We
introduced three di erent proximity functions used within the CPA method and
developed a highly scalable CPA implementation that runs on a cluster
using the MapReduce paradigm. Our initial results suggest that CPA can provide
better performance in recommender systems than the co-citation method. More
speci cally, two of our proximity functions outperformed the baseline co-citation
approach on our dataset, the SumProx function by a margin of more than 25%
for precision@5. However, a larger evaluation dataset is needed to con rm these
results.</p>
        </sec>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <given-names>Bela</given-names>
            <surname>Gipp</surname>
          </string-name>
          and Joran Beel.
          <article-title>Citation proximity analysis (cpa)-a new approach for identifying related work based on co-citation analysis</article-title>
          .
          <source>In Proceedings of the 12th International Conference on Scientometrics and Informetrics (ISSI'09)</source>
          , volume
          <volume>2</volume>
          , pages
          <fpage>571</fpage>
          {
          <fpage>575</fpage>
          . Rio de Janeiro (Brazil):
          <source>International Society for Scientometrics and Informetrics</source>
          ,
          <year>2009</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <given-names>Joeran</given-names>
            <surname>Beel</surname>
          </string-name>
          , Bela Gipp, Stefan Langer, and Corinna Breitinger.
          <article-title>Research-paper recommender systems: a literature survey</article-title>
          .
          <source>International Journal on Digital Libraries</source>
          ,
          <volume>17</volume>
          (
          <issue>4</issue>
          ):
          <volume>305</volume>
          {338, jul
          <year>2015</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <given-names>Lutz</given-names>
            <surname>Bornmann</surname>
          </string-name>
          and
          <article-title>Rudiger Mutz. Growth rates of modern science: A bibliometric analysis based on the number of publications and cited references</article-title>
          .
          <source>Journal of the Association for Information Science and Technology</source>
          ,
          <volume>66</volume>
          (
          <issue>11</issue>
          ):
          <volume>2215</volume>
          {
          <fpage>2222</fpage>
          ,
          <year>2015</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Jevin</surname>
            <given-names>D.</given-names>
          </string-name>
          <string-name>
            <surname>West</surname>
          </string-name>
          , Ian
          <string-name>
            <surname>Wesley-Smith</surname>
          </string-name>
          , and
          <string-name>
            <surname>Carl</surname>
            <given-names>T.</given-names>
          </string-name>
          <string-name>
            <surname>Bergstrom</surname>
          </string-name>
          .
          <article-title>A recommendation system based on hierarchical clustering of an article-level citation network</article-title>
          .
          <source>IEEE Transactions on Big Data</source>
          ,
          <volume>2</volume>
          (
          <issue>2</issue>
          ):
          <volume>113</volume>
          {123, jun
          <year>2016</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <given-names>Norman</given-names>
            <surname>Meuschke</surname>
          </string-name>
          , Bela Gipp, and
          <string-name>
            <given-names>Mario</given-names>
            <surname>Lipinsk</surname>
          </string-name>
          .
          <article-title>Citrec: An evaluation framework for citation-based similarity measures based on trec genomics and pubmed central</article-title>
          .
          <source>iConference 2015 Proceedings</source>
          ,
          <year>2015</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <given-names>Bela</given-names>
            <surname>Gipp</surname>
          </string-name>
          , Joran Beel, and
          <string-name>
            <given-names>Christian</given-names>
            <surname>Hentschel</surname>
          </string-name>
          .
          <article-title>Scienstein : A research paper recommender system</article-title>
          .
          <source>In Proceedings of the International Conference on Emerging Trends in Computing</source>
          , pages
          <volume>309</volume>
          {
          <fpage>315</fpage>
          ,
          <year>2009</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <given-names>Bela</given-names>
            <surname>Gipp</surname>
          </string-name>
          , Adriana Taylor, and Joran Beel.
          <article-title>Link Proximity Analysis-Clustering Websites by Examining Link Proximity</article-title>
          .
          <source>Proceedings of the 14th European Conference on Digital Libraries (ECDL'10)</source>
          ,
          <volume>6273</volume>
          (September):
          <volume>449</volume>
          {
          <fpage>452</fpage>
          ,
          <year>2010</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <given-names>Chong</given-names>
            <surname>Wang</surname>
          </string-name>
          and
          <string-name>
            <given-names>David M.</given-names>
            <surname>Blei</surname>
          </string-name>
          .
          <article-title>Collaborative topic modeling for recommending scienti c articles</article-title>
          .
          <source>In Proceedings of the 17th ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, KDD '11</source>
          , pages
          <fpage>448</fpage>
          {
          <fpage>456</fpage>
          , New York, NY, USA,
          <year>2011</year>
          . ACM.
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <given-names>Cristiano</given-names>
            <surname>Nascimento</surname>
          </string-name>
          ,
          <string-name>
            <surname>Alberto H.F. Laender</surname>
          </string-name>
          , Altigran S. da Silva, and
          <article-title>Marcos Andre Goncalves. A source independent framework for research paper recommendation</article-title>
          .
          <source>In Proceedings of the 11th Annual International ACM/IEEE Joint Conference on Digital Libraries, JCDL '11</source>
          , pages
          <fpage>297</fpage>
          {
          <fpage>306</fpage>
          , New York, NY, USA,
          <year>2011</year>
          . ACM.
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <surname>Qi</surname>
            <given-names>He</given-names>
          </string-name>
          , Jian Pei, Daniel Kifer, Prasenjit Mitra, and
          <string-name>
            <given-names>Lee</given-names>
            <surname>Giles</surname>
          </string-name>
          .
          <article-title>Context-aware citation recommendation</article-title>
          .
          <source>In Proceedings of the 19th International Conference on World Wide Web, WWW '10</source>
          , pages
          <fpage>421</fpage>
          {
          <fpage>430</fpage>
          , New York, NY, USA,
          <year>2010</year>
          . ACM.
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <string-name>
            <surname>Yicong</surname>
            <given-names>Liang</given-names>
          </string-name>
          ,
          <string-name>
            <given-names>Qing</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>and Tieyun</given-names>
            <surname>Qian</surname>
          </string-name>
          .
          <source>Finding Relevant Papers Based on Citation Relations</source>
          , pages
          <volume>403</volume>
          {
          <fpage>414</fpage>
          . Springer Berlin Heidelberg, Berlin, Heidelberg,
          <year>2011</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          12.
          <string-name>
            <surname>Ding</surname>
            <given-names>Zhou</given-names>
          </string-name>
          , Shenghuo Zhu, Kai Yu, Xiaodan Song, Belle L. Tseng, Hongyuan Zha, and
          <string-name>
            <given-names>C. Lee</given-names>
            <surname>Giles</surname>
          </string-name>
          .
          <article-title>Learning multiple graphs for document recommendations</article-title>
          .
          <source>In Proceedings of the 17th International Conference on World Wide Web, WWW '08</source>
          , pages
          <fpage>141</fpage>
          {
          <fpage>150</fpage>
          , New York, NY, USA,
          <year>2008</year>
          . ACM.
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          13.
          <string-name>
            <surname>Nitin</surname>
            <given-names>Agarwal</given-names>
          </string-name>
          , Ehtesham Haque, Huan Liu, and Lance Parsons.
          <article-title>Research paper recommender systems: A subspace clustering approach</article-title>
          .
          <source>In Proceedings of the 6th International Conference on Advances in Web-Age Information Management, WAIM'05</source>
          , pages
          <fpage>475</fpage>
          {
          <fpage>491</fpage>
          , Berlin, Heidelberg,
          <year>2005</year>
          . Springer-Verlag.
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          14.
          <string-name>
            <given-names>Kazunari</given-names>
            <surname>Sugiyama</surname>
          </string-name>
          and
          <string-name>
            <surname>Min-Yen Kan</surname>
          </string-name>
          .
          <article-title>Exploiting potential citation papers in scholarly paper recommendation</article-title>
          .
          <source>In Proceedings of the 13th ACM/IEEE-CS Joint Conference on Digital Libraries, JCDL '13</source>
          , pages
          <fpage>153</fpage>
          {
          <fpage>162</fpage>
          , New York, NY, USA,
          <year>2013</year>
          . ACM.
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          15.
          <string-name>
            <given-names>Henry</given-names>
            <surname>Small</surname>
          </string-name>
          .
          <article-title>Co-citation in the scienti c literature: A new measure of the relationship between two documents</article-title>
          .
          <source>Journal of the Association for Information Science and Technology</source>
          ,
          <volume>24</volume>
          (
          <issue>4</issue>
          ):
          <volume>265</volume>
          {
          <fpage>269</fpage>
          ,
          <year>1973</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          16.
          <string-name>
            <surname>Irina</surname>
            <given-names>V</given-names>
          </string-name>
          <string-name>
            <surname>Marshakova</surname>
          </string-name>
          .
          <article-title>System of document connections based on references</article-title>
          .
          <source>Nauchno-Tekhnicheskaya Informatsiya Seriya 2-Informatsionnye Protsessy I Sistemy, (6):3{8</source>
          ,
          <year>1973</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          17.
          <string-name>
            <surname>Nam</surname>
            <given-names>Tran</given-names>
          </string-name>
          , Pedro Alves, Shuangge Ma, and
          <string-name>
            <given-names>Michael</given-names>
            <surname>Krauthammer</surname>
          </string-name>
          .
          <article-title>Enriching pubmed related article search with sentence level co-citations</article-title>
          .
          <source>In AMIA Annual Symposium Proceedings</source>
          , volume
          <year>2009</year>
          , page
          <volume>650</volume>
          ,
          <year>2009</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          18.
          <string-name>
            <surname>Malte</surname>
            <given-names>Schwarzer</given-names>
          </string-name>
          , Moritz Schubotz, Norman Meuschke, Corinna Breitinger, Volker Markl, and
          <string-name>
            <given-names>Bela</given-names>
            <surname>Gipp</surname>
          </string-name>
          .
          <article-title>Evaluating link-based recommendations for wikipedia</article-title>
          .
          <source>In Digital Libraries (JCDL)</source>
          ,
          <year>2016</year>
          IEEE/ACM Joint Conference on, pages
          <volume>191</volume>
          {
          <fpage>200</fpage>
          . IEEE,
          <year>2016</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          19.
          <string-name>
            <surname>Bela</surname>
            <given-names>Gipp</given-names>
          </string-name>
          , Norman Meuschke, and
          <string-name>
            <given-names>Corinna</given-names>
            <surname>Breitinger</surname>
          </string-name>
          .
          <article-title>Citation-based plagiarism detection: Practicability on a large-scale scienti c corpus</article-title>
          .
          <source>Journal of the Association for Information Science and Technology</source>
          ,
          <volume>65</volume>
          (
          <issue>8</issue>
          ):
          <volume>1527</volume>
          {
          <fpage>1540</fpage>
          ,
          <year>2014</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          20.
          <string-name>
            <given-names>Petr</given-names>
            <surname>Knoth</surname>
          </string-name>
          and
          <string-name>
            <given-names>Zdenek</given-names>
            <surname>Zgrahal</surname>
          </string-name>
          . Core:
          <article-title>Three access levels to underpin open access</article-title>
          .
          <string-name>
            <surname>D-Lib</surname>
            <given-names>Magazine</given-names>
          </string-name>
          ,
          <volume>18</volume>
          (
          <issue>11</issue>
          /12), nov
          <year>2012</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          21.
          <string-name>
            <given-names>Patrice</given-names>
            <surname>Lopez</surname>
          </string-name>
          . GROBID:
          <article-title>Combining Automatic Bibliographic Data Recognition and Term Extraction for Scholarship Publications</article-title>
          , pages
          <volume>473</volume>
          {
          <fpage>474</fpage>
          . Springer Berlin Heidelberg, Berlin, Heidelberg,
          <year>2009</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          22.
          <string-name>
            <given-names>Paul</given-names>
            <surname>Resnick</surname>
          </string-name>
          and Hal R Varian.
          <source>Recommender systems. Communications of the ACM</source>
          ,
          <volume>40</volume>
          (
          <issue>3</issue>
          ):
          <volume>56</volume>
          {
          <fpage>58</fpage>
          ,
          <year>1997</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          23.
          <string-name>
            <surname>Jonathan L. Herlocker</surname>
          </string-name>
          , Joseph A.
          <string-name>
            <surname>Konstan</surname>
          </string-name>
          , Loren G. Terveen, and John T. Riedl.
          <article-title>Evaluating collaborative ltering recommender systems</article-title>
          .
          <source>ACM Trans. Inf</source>
          . Syst.,
          <volume>22</volume>
          (
          <issue>1</issue>
          ):5{
          <fpage>53</fpage>
          ,
          <year>January 2004</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>