<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Collecting the Seminal Scienti c Abstracts with Topic Modelling, Snowball Sampling and Citation Analysis</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Hennadii Dobrovolskyi</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Nataliya Keberle</string-name>
          <email>nkeberle@gmail.com</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Department of Computer Science, Zaporizhzhya National University</institution>
          ,
          <addr-line>Zhukovskogo st. 66, 69600, Zaporizhzhya</addr-line>
          ,
          <country country="UA">Ukraine</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>This paper presents a complete information technology for collecting and analysis of a citation network of scienti c publications aimed at detecting of seminal papers in a selected domain of research. The technology consists of the seed paper selection, plain snowball sampling, probabilistic topic modeling, greedy restricted snowball sampling, and analysis of the collected citation network. The topic model is built on the base of word-word co-occurrence probability with combination of sparse symmetric nonnegative matrix factorization and principal component approximation. Experiments with the collection of High Energy Physics abstracts show that the number of topics in the model is determined in natural way and the KullbackLeibler divergence correlates with cosine similarity calculated from keywords provided by publication authors. The citation networks on \critical thinking" and \automatic pronunciation assessment" domains are collected and analyzed. The analysis shows that both networks are \small worlds" and therefore the observed saturation of the restricted snowball sampling can provide the complete set of publications in domains of interest. Multiple runs of the sampling con rm the hypothesis that the set of seminal publications is stable with respect to variations of the seed papers. The modi ed main path analysis allows to distinguish the seminal papers including new publications following main stream of research.</p>
      </abstract>
      <kwd-group>
        <kwd>text mining</kwd>
        <kwd>short text document</kwd>
        <kwd>topic modelling</kwd>
        <kwd>principal component analysis</kwd>
        <kwd>sparse symmetric nonnegative matrix factorization</kwd>
        <kwd>citation network</kwd>
        <kwd>main path analysis</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>getting answers to both questions. The essence of the technology is the collecting
and analysis of a citation network of scienti c publications aimed at detecting
of seminal papers in a selected domain of research.</p>
      <p>To our best knowledge, while the separate parts of the method are developed,
the entire procedure that takes the small set of papers on some scienti c domain
and produces perfect list of references is not known.</p>
      <p>
        The objectives of the presented work are:
{ to present the complete technology that takes a manually selected seed
papers and produces the short list of interconnected scienti c publications that
re ects evolution of main ideas of the selected scienti c domain. The
technology contains both restricted snowball sampling method and citation network
analysis.
{ to test all initial assumptions, namely
if the proposed snowball restriction method [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ] provides adequate
semantic distance between publications;
if the obtained restricted snowball forms the scale-free network [22];
if the restricted snowball provides saturation of the publication dataset;
if the best age of the seed papers is 5{10 years;
if the biased seed papers can produce unbiased citation network.
      </p>
      <p>The distinctive features of the presented method are application of the
probabilistic topic model to perform restricted snowball sampling and collect citation
network, then the main path analysis is applied to point both the most in
uential publications and the main path of scienti c knowledge evolution. The main
path allows detecting the newest publications that follow the mainstream and
the outliers that potentially contain the completely new ideas.</p>
      <p>The structure of the paper is following. Section 2 overviews the publications
related to the presented technology, Section 3 contains description of the crucial
steps of the algorithm, Section 4 states the experiment pre-conditions and Section
5 discusses the results. Conclusion summarises the main results and discusses
future work.
2</p>
    </sec>
    <sec id="sec-2">
      <title>Related Work</title>
      <p>
        Publications on a domain of knowledge can be collected from conference
proceedings [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ], the study of the co-authorship [
        <xref ref-type="bibr" rid="ref15 ref16">15, 16</xref>
        ], elaboration of keywords
and topics [
        <xref ref-type="bibr" rid="ref14">14, 17</xref>
        ], querying academic search engines1 with a set of keywords
or snowball sampling [
        <xref ref-type="bibr" rid="ref1 ref10 ref6">10, 1, 6</xref>
        ]. However, not for every research domain there is
a corresponding conference, as well as one author can write papers on di erent
topics. Building maps and ontologies of large scienti c domains does not provide
the list of references rather the set of interconnected concepts. Querying with
      </p>
      <sec id="sec-2-1">
        <title>1 Google Scholar, https://scholar.google.com,</title>
        <p>Semantic Scholar, https://www.semanticscholar.org/,</p>
        <p>Microsoft Academic, https://academic.microsoft.com/
a set of keywords produces the biased set of publications [19] because di erent
researchers use slightly di erent terms to report their results.</p>
        <p>
          The most appropriate way to collect scienti c papers is snowball sampling
[
          <xref ref-type="bibr" rid="ref1 ref10">10, 1</xref>
          ], when each publication from the current queue is considered then all
referenced papers and all papers referencing to the publication are added to the
next level queue. The snowball sampling allows collecting publications on the
narrow research topic and connect them in the citation network [22]. The high
quality of citation-based search algorithm is provided with phenomena of \small
world" which is a proved property of scale-free networks [
          <xref ref-type="bibr" rid="ref2">2, 22</xref>
          ]. Newman [
          <xref ref-type="bibr" rid="ref15">15</xref>
          ]
has shown that in the most of the cases it is enough to do three iterations.
However, the statistical properties of global citation network including all scienti c
papers is not known, that is why we need to test if the small world assumption is
true for the collected subset of citation network and if the three iterations allows
collecting most of the papers.
        </p>
        <p>
          Another point of snowball sampling is dependence on the initial queue called
a seed collection. The general advice [
          <xref ref-type="bibr" rid="ref10">10</xref>
          ] recommends that the seed papers
should be the seminal papers of the knowledge domain pointed by experts or
the papers selected by the researcher. Valid seed papers should be 5{10 years
old and have to be widely cited. The best seeds are the reviews, foundational
or framing articles on the topic of interest. However, the advice also should
be checked. Moreover, we need to test if the biased seed papers can produce
unbiased citation network.
        </p>
        <p>The snowball sampling cannot be applied directly to publication crawling
because the list of references can contain the items that are not directly related to
the investigated domain. Therefore the straightforward implementation of
snowball publication sampling causes in nite collection in ation and some restrictions
should be introduced to accept or reject the candidate publication. It should be
noted that the introduced restrictions can violate the small world property and
we need to check if the restricted snowball result is scale-free citation network.</p>
        <p>
          To lter out the most relevant papers while sampling Ahad et al. [
          <xref ref-type="bibr" rid="ref1">1</xref>
          ] in their
approach use vector document model and cosine similarity, however the
document vector model relies on word spelling rather than meaning that causes
precision loss when the short texts are considered. Lecy et al. [
          <xref ref-type="bibr" rid="ref10">10</xref>
          ] apply
PageRank calculated by Google Scholar as a measure of paper signi cance. However,
PageRank is a property of a global citation network including all topics of
knowledge, so it cannot be calculated from its small subset. One of the most promising
approaches is the probabilistic topic model (PTM).
        </p>
        <p>
          Probabilistic topic models [24] use a large collection of documents and
statistical approach to model words and documents as vectors in a high-dimensional
semantic space Rn, where n is much less than number of words and number of
documents. The base idea of PTM is to construct few topis which are groups
of tightly connected words. Then document words are represented as a result
of two-stage random sampling. The most known method of topic modelling is
Latent Dirichlet Allocations (LDA) [
          <xref ref-type="bibr" rid="ref5">5</xref>
          ] which is successful and simple enough.
A general introduction and survey of the topic modelling can be found in [24]
along with a novel approach, called Additive Regularization of Topic Models.
        </p>
        <p>However, in most of the scienti c databases, full texts are often protected
by copyright. Therefore the only information we can use are paper title, paper
abstract, and sometimes the database-speci c keywords and topics. So the
documents that we analyse are short and common PTMs based on document-word
statistics lose their precision. This shortcoming is overpassed with approaches
utilizing word co-occurrence statistics in Biterm Topic Model (BTM) [25] and
Word Network Topic Model (WNTM) [26] instead of counting document-word
pairs. Another method of word embedding, called GloVe, is proposed in [18]. It is
based on word-word co-occurrence matrix and uses global matrix factorization,
so it is close to BTM [25] and WNTM [26] statistical topic modelling.</p>
        <p>
          Also, the vague part of common PTMs is that number of topics cannot be
determined with document analysis. To overcome this weakness, handling the
word-word co-occurences with principal component analysis (PCA) and sparse
symmetric nonnegative matrix factorization (Sparse SNMF) was proposed [
          <xref ref-type="bibr" rid="ref6">6</xref>
          ].
        </p>
        <p>
          The collected citation network [22] can be analyzed using citation count and
other simple statistics [
          <xref ref-type="bibr" rid="ref12">12</xref>
          ], PageRank [
          <xref ref-type="bibr" rid="ref11">11</xref>
          ], information retrieval techniques [
          <xref ref-type="bibr" rid="ref1">1</xref>
          ],
knowledge graph [17], combined supervised machine learning approaches [23] or
Main Path analysis [
          <xref ref-type="bibr" rid="ref12">12</xref>
          ].
        </p>
        <p>The most appropriate way to highlight the seminal papers of the small
scienti c domain is main path analysis because the method deals only with the
collected citation network and allows to increase the precision of sampled dataset.
On the contrary, the citation count and other statistics, supervised machine
learning applied by Valenzuela, Ha and Etzioni [23] and PageRank cannot point
out the tightly interconnected subset of the citation network. Klink-2 [17] and
similar algorithms aim to build the map of knowledge domain but do not seek
the most in uenced publications.
3</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>Information Technology Overview</title>
      <p>
        The general work ow of the restricted snowball sampling is introduced in [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ]. It
contains the following steps:
1. Collect a set of seed papers and put them in the initial, 0-th, queue.
2. Run several iterations of the unrestricted snowball sampling to pickup
baseline documents. For n 2 0; 1; 2; 3
1 get a portion of papers from the n-th queue;
2 download the papers referenced by the portion;
3 download the papers referencing the portion;
4 add all the downloaded papers to the (n + 1)-th queue.
3. Create the PTM using baseline documents:
1 extract title and abstract from each document of the collection;
2 split all the titles and abstracts into sentences;
3 create the dictionary containing all the nouns and adjectives that occur
in the sentences;
4 combine all terms from the reduced dictionary occurring in the same
sentence into pairs and build the joint probability matrix;
5 detect the collection speci c stop-words and exclude them from the
reduced dictionary;
6 perform Sparse SNMF to create PTM;
7 map each of the seed papers to a vector of topic probabilities.
4. Perform the batch restricted snowball sampling: for n 2 0; 1; 2; 3
1 get a portion of papers from the n-th queue;
2 download the papers referenced by the portion;
3 download the papers referencing the portion;
4 extract bag of stemmed words from each of downloaded papers;
5 map each of the downloaded papers to a vector of topic probabilities;
6 calculate distance from each downloaded paper to the seed papers;
7 add to the next level queue only those of downloaded papers which are
close to the seed papers.
5. Analyse the citation network.
      </p>
      <p>
        The details of the restricted snowball sampling and probabilistic topic model
construction are discussed in [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ].
4
4.1
      </p>
    </sec>
    <sec id="sec-4">
      <title>Citation Network Analysis</title>
      <sec id="sec-4-1">
        <title>Cycles Elimination</title>
        <p>
          The correctly built citation network must be an acyclic directed graph. However,
the publication database errors accidentally can cause cycles. The problem with
cycles is that if there is a cycle in a network then there is also an in nite number of
paths between some vertices. Since a citation network is usually almost acyclic
to transform it into an acyclic network we use the \preprint" transformation
described by Batagelj [
          <xref ref-type="bibr" rid="ref3">3</xref>
          ]. First, we identify cycles and then each paper from a
cycle is duplicated with its \preprint" version and the papers inside cycle cite
\preprints".
4.2
        </p>
      </sec>
      <sec id="sec-4-2">
        <title>Simple Citation Path Count</title>
        <p>
          Our approach is similar to Search Path Count (SPC) algorithm [
          <xref ref-type="bibr" rid="ref12 ref3">12, 3</xref>
          ]. We
introduce two pseudo-vertices { source and target. A vertex that does not reference
any other publication vertex, gets an edge to the target vertex. A vertex that is
not referenced by any publication vertex, gets an edge from the source vertex,
so the graph becomes connected. Next step is to calculate all simple paths from
the source to the target using Python library NetworkX2. The algorithm uses a
modi ed depth- rst search to generate the paths [
          <xref ref-type="bibr" rid="ref9">9</xref>
          ]. As the result we obtain a
set of paths, each of which is a sequence of vertices. Each pair of direct neighbour
vertices in such a sequence is an edge in the citation graph. For each edge, its
2 NetworkX, https://networkx.github.io
frequency is calculated against all paths { a number of paths through it, simple
path count. Next we calculate edge resistance as inverse proportional to edge
simple path count { this allows diminishing the di erence among the most cited
and least cited papers [
          <xref ref-type="bibr" rid="ref8">21, 8</xref>
          ]. The path resistance is then calculated as the sum
of its edge resistances. Finally, we set an order over the paths using the path
resistances. Using path resistances is a distinguishing feature of the proposed
algorithm.
        </p>
        <p>
          The di erence of the applied algorithm from SPC algorithm is the
preservation of the citation graph connectivity. In the known algorithm [
          <xref ref-type="bibr" rid="ref4">4</xref>
          ], as soon as
the edge SPC scores are calculated the edges having low scores are removed and
the citation network can become a disconnected graph.
4.3
        </p>
      </sec>
      <sec id="sec-4-3">
        <title>Chasing New Ideas in Publications</title>
        <p>Path resistance allows detecting new publications in the eld, not referenced yet
by any other authors but existing in a mainstream of the domain. We can
separate all papers into mainstream research and probably new research elds or
directions. The smallest (up to a certain threshold) path resistances correspond
to the mainstream, whereas the biggest path resistances correspond to the
publications that are either brand new, bad written or published in a low impact
journal/conference proceedings. We assume those publications are the source of
potentially new ideas and topics.
5
5.1</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>Analysis of Experimental Results</title>
      <sec id="sec-5-1">
        <title>Experimental Settings</title>
        <p>We took three di erent corpora: \high energy physics"(HEP), \critical
thinking"(CT) and \pronunciation quality assessment"(PQA).</p>
        <p>HEP publications [20] are available from the European Laboratory for
Nuclear Research. The hep-ex partition of the HEP collection is composed of 2802
abstracts related to experimental high-energy physics that are indexed with 1093
main keywords (the categories), the hep-astroph partition contains 2716
abstracts from astrophysics section and 18114 abstracts on theoretical physics in
hep-th metadata. Each publication is manually annotated with keywords.</p>
        <p>CT corpus is gathered with our snowball sampling software. The CT domain
is characterized with a large noisy publications corpus tightly entangled with
publications on psychology, didactics, pedagogy and phylosophy. The size of CT
corpus is 24040 publication abstracts.</p>
        <p>PQA domain is very speci c and narrow, with a moderate-size corpus
containing 8339 scienti c abstracts collected by our snowball sampling software.</p>
        <p>The sampling was run with following parameters: percentage of stop words to
exclude { 2%; percentage of rare words to exclude { 5%; number of components
in PCA which is maximal number of topics { 200; threshold KL-divergence {
0.18; sparsity parameter { 0.05; number of top citation paths { 50; minimal
number of citations { 3.
5.2</p>
      </sec>
      <sec id="sec-5-2">
        <title>Seminal Publications for PQA domain</title>
        <p>On the part (a) of Figure 1 we can see that the mainstream of pronunciation
assessment contains the publications:</p>
        <p>The mainstream evolution of pronunciation assessment starts from
application of automatic speech recognition, pays some attention to pedagogical aspects,
goes to simple machine learning approaches, then to neural networks and to deep
learning. Some of the seminal publications are the reviews containing discussions
of the feature selection, methods comparison and combination. The part (b) of
Figure 1 shows that the more top paths we keep, the more detailed knowledge
map we obtain.
5.3</p>
      </sec>
      <sec id="sec-5-3">
        <title>Assumptions Checking</title>
      </sec>
      <sec id="sec-5-4">
        <title>PTM as a restriction criteria for the proposed restricted snowball</title>
        <p>sampling method provides adequate semantic distance between
publications. To check the statement we took HEP collection, built PTM for it,
and measure similarity using keywords annotating each publication from HEP
collection. For each pair of HEP publications were calculated both symmetric
KL-divergence and cosine similarity. The results are shown in Figure 2.</p>
        <p>
          As we can see, the PTM-based symmetric Kullback-Leibler divergence [
          <xref ref-type="bibr" rid="ref13">13</xref>
          ]
provides reliable upper bound for the keyword based cosine similarity. The reason
is that the cosine similarity uses only word spelling and PTM uses Rn word
embedding taking into account the meaning of terms.
        </p>
        <p>The restricted snowball sampling forms the scale-free network. Derek
de Solla Price showed in 1965 [22] that the number of references to a paper
(node degree) in a citation network had a heavy-tailed distribution following the
power law and thus that the citation network is scale-free. One of our initial
assumptions was that the restricted snowball sampling results in a scale-free
networks so we need a few number of snowball iterations to achieve the high
recall. Figure 3 was calculated on the base of PQA corpus and shows that for
the small node degrees the logarithm of node number has linear dependency on
the logarithm of node degree and for the large node degrees the dependency
has heavy tail. That means, the restricted snowball sampling produces
scalefree citation network as well as classical snowball. So we can be sure that a
few iterations of the restricted snowball sampling allow collecting most of the
relevant publications.</p>
        <p>Saturation of the restricted snowball sampling. The restricted snowball
sampling can be modelled as Poisson process when the publications appear
sequentially and we can either (a) accept n-th publication and add it to the
snowball or (b) don't accept. So we can calculate the con dence interval of Poisson
distribution of event (a) and compare its upper bound with some pre-de ned
acceptance probability. Figure 4 shows 0.95 con dence interval of Poisson
distribution of paper acceptance as a colored strip and acceptance probability
threshold 0.05 as a straight line. The con dence interval was calculated on the base of
10 snowball runs for CT collection starting from random subsets of seed paper
collection. After some number of tested abstracts the upper bound of con dence
interval becomes lower than the threshold so the restricted snowball sampling
guarantees the saturation.
The Citations Age To study the in uence of a publication age on the
probability of the publication citation we have attributed each edge of the PQA
citation network with age calculated as di erence between years of referencing
publication and referenced one. The number of the edges as a function of edge
age is shown on Figure 5. We can see that the maximal number of the references
is observed for the publications that are 2{8 years old. Such publications are
still regarded as new ones but at the same time are old enough to be read and
estimated by many reserchers.</p>
        <p>The biased seed papers can produce unbiased citation network. To
estimate the stability of the restricted snowball sampling with respect to the
seed papers variation we run the sampling starting from the full set of the PQA
seed papers and mark the relevant papers with main path analysis. Then we run
the sampling again starting from 10 random subsets of the PQA seed papers
and count the number of the runs where each seminal paper occurs. In our
experiments the random subsets contain 50% of the seed papers and 66% of
relevant papers are detected every time, 14% { in 80% of runs, 14% { in 60% of
runs, 6% { at least once. So we can conclude that the PQA citation network is
stable with respect to large seed paper variations being input for the restricted
snowball sampling and the result of the sampling is unbiased.</p>
      </sec>
    </sec>
    <sec id="sec-6">
      <title>Conclusions and Future Studies</title>
      <p>The main objective of the paper was to present the complete information
technology that obtains a set of publications on some scienti c topic as input and
produces a list of seminal publications for that topic. It provides data for future
detailed analysis and serves as a good point to begin investigation in a new
domain. Additionally, we tested several initial assumptions regarding the results of
the technology application and show that:
{ PTM as a restriction criteria for the restricted snowball sampling method
provides adequate semantic distance between publications.
{ The restricted snowball sampling guarantees the saturation.
{ The maximal number of the references is observed for the publications that
are 2{8 years old. Such publications are still regarded as new ones but at the
same time are old enough to be read and estimated by many reserchers.
{ The biased seed papers produce unbiased citation network.
{ The collected citation network is stable with respect to large seed paper
variations being input for the restricted snowball sampling and the result of
the sampling is unbiased.</p>
      <p>The presented technology is implemented as sequence of Python scripts3.</p>
      <sec id="sec-6-1">
        <title>3 https://github.com/gendobr/snowball</title>
        <p>17. Osborne, F., Motta, E.: Klink-2: integrating multiple web sources to generate
semantic topic networks. In: International Semantic Web Conference. pp. 408{424.</p>
        <p>Springer (2015)
18. Pennington, J., Socher, R., Manning, C.D.: Glove: Global vectors for word
representation. In: EMNLP. vol. 14, pp. 1532{1543 (2014)
19. Petticrew, M., Gilbody, S.: Planning and conducting systematic reviews. Health
psychology in practice pp. 150{179 (2004)
20. Raez, A.M., Lopez, L.A.U., Steinberger, R.: Adaptive selection of base classi ers in
one-against-all learning for large multi-labeled collections. In: Advances in Natural
Language Processing, pp. 1{12. Springer (2004)
21. Salganik, M.J., Heckathorn, D.D.: Sampling and estimation in hidden populations
using respondent-driven sampling. Sociological methodology 34(1), 193{240 (2004)
22. de Solla Price, D.J.: Networks of scienti c papers. Science 149(3683), 510{515
(1965)
23. Valenzuela, M., Ha, V., Etzioni, O.: Identifying meaningful citations. In: AAAI</p>
        <p>Workshop: Scholarly Big Data (2015)
24. Vorontsov, K., Potapenko, A.: Tutorial on probabilistic topic modeling: Additive
regularization for stochastic matrix factorization. In: International Conference on
Analysis of Images, Social Networks and Texts x000D . pp. 29{46. Springer (2014)
25. Yan, X., Guo, J., Lan, Y., Cheng, X.: A biterm topic model for short texts. In:
Proceedings of the 22nd international conference on World Wide Web. pp. 1445{
1456. ACM (2013)
26. Zuo, Y., Zhao, J., Xu, K.: Word network topic model: a simple but general solution
for short and imbalanced texts. Knowledge and Information Systems 48(2), 379{
398 (2016)</p>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Ahad</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Fayaz</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Shah</surname>
            ,
            <given-names>A.S.:</given-names>
          </string-name>
          <article-title>Navigation through citation network based on content similarity using cosine similarity algorithm</article-title>
          .
          <source>International Journal of Database Theory and Application</source>
          <volume>9</volume>
          (
          <issue>5</issue>
          ),
          <volume>9</volume>
          {
          <fpage>20</fpage>
          (
          <year>2016</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Barabasi</surname>
            ,
            <given-names>A.L.</given-names>
          </string-name>
          :
          <article-title>Scale-free networks: a decade and beyond</article-title>
          .
          <source>Science</source>
          <volume>325</volume>
          (
          <issue>5939</issue>
          ),
          <volume>412</volume>
          {
          <fpage>413</fpage>
          (
          <year>2009</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Batagelj</surname>
          </string-name>
          , V.:
          <article-title>E cient algorithms for citation network analysis</article-title>
          .
          <source>arXiv preprint cs/0309023</source>
          (
          <year>2003</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Batagelj</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Mrvar</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          :
          <article-title>Pajek-program for large network analysis</article-title>
          .
          <source>Connections</source>
          <volume>21</volume>
          (
          <issue>2</issue>
          ),
          <volume>47</volume>
          {
          <fpage>57</fpage>
          (
          <year>1998</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Blei</surname>
            ,
            <given-names>D.M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ng</surname>
            ,
            <given-names>A.Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Jordan</surname>
            ,
            <given-names>M.I.</given-names>
          </string-name>
          :
          <article-title>Latent dirichlet allocation</article-title>
          .
          <source>Journal of machine Learning research 3(Jan)</source>
          ,
          <volume>993</volume>
          {
          <fpage>1022</fpage>
          (
          <year>2003</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Dobrovolskyi</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Keberle</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Todoriko</surname>
            ,
            <given-names>O.</given-names>
          </string-name>
          :
          <article-title>Probabilistic topic modelling for controlled snowball sampling in citation network collection</article-title>
          .
          <source>In: International Conference on Knowledge Engineering and the Semantic Web</source>
          . pp.
          <volume>85</volume>
          {
          <fpage>100</fpage>
          . Springer (
          <year>2017</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>Ermolayev</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Batsakis</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Keberle</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Tatarintseva</surname>
            ,
            <given-names>O.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Antoniou</surname>
          </string-name>
          , G.:
          <article-title>Ontologies of time: Review and trends</article-title>
          .
          <source>International Journal of Computer Science &amp; Applications</source>
          <volume>11</volume>
          (
          <issue>3</issue>
          ) (
          <year>2014</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>Even</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          :
          <article-title>Graph algorithms</article-title>
          . Cambridge University Press (
          <year>2011</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <surname>Golumbic</surname>
            ,
            <given-names>M.C.</given-names>
          </string-name>
          :
          <article-title>Algorithmic graph theory and perfect graphs</article-title>
          , vol.
          <volume>57</volume>
          .
          <string-name>
            <surname>Elsevier</surname>
          </string-name>
          (
          <year>2004</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <surname>Lecy</surname>
            ,
            <given-names>J.D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Beatty</surname>
            ,
            <given-names>K.E.</given-names>
          </string-name>
          :
          <article-title>Representative literature reviews using constrained snowball sampling and citation network analysis (</article-title>
          <year>2012</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <string-name>
            <surname>Leskovec</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rajaraman</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ullman</surname>
            ,
            <given-names>J.D.</given-names>
          </string-name>
          :
          <article-title>Mining of massive datasets</article-title>
          . Cambridge university press (
          <year>2014</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          12.
          <string-name>
            <surname>Lucio-Arias</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Leydesdor</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          :
          <article-title>Main-path analysis and path-dependent transitions in histcite-based historiograms</article-title>
          .
          <source>Journal of the Association for Information Science and Technology</source>
          <volume>59</volume>
          (
          <issue>12</issue>
          ),
          <year>1948</year>
          {
          <year>1962</year>
          (
          <year>2008</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          13.
          <string-name>
            <surname>MacKay</surname>
          </string-name>
          , D.J.:
          <article-title>Information theory, inference and learning algorithms</article-title>
          . Cambridge university press (
          <year>2003</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          14.
          <string-name>
            <surname>Moya-Anegon</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Vargas-Quesada</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Herrero-Solana</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chinchilla-Rodr guez</surname>
            ,
            <given-names>Z.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Corera-Alvarez</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Munoz-Fernandez</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          :
          <article-title>A new technique for building maps of large scienti c domains based on the cocitation of classes and categories</article-title>
          .
          <source>Scientometrics</source>
          <volume>61</volume>
          (
          <issue>1</issue>
          ),
          <volume>129</volume>
          {
          <fpage>145</fpage>
          (
          <year>2004</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          15.
          <string-name>
            <surname>Newman</surname>
            ,
            <given-names>M.E.</given-names>
          </string-name>
          :
          <article-title>The structure of scienti c collaboration networks</article-title>
          .
          <source>Proceedings of the National Academy of Sciences</source>
          <volume>98</volume>
          (
          <issue>2</issue>
          ),
          <volume>404</volume>
          {
          <fpage>409</fpage>
          (
          <year>2001</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          16.
          <string-name>
            <surname>Newman</surname>
            ,
            <given-names>M.E.</given-names>
          </string-name>
          :
          <article-title>Coauthorship networks and patterns of scienti c collaboration</article-title>
          .
          <source>Proceedings of the national academy of sciences 101(suppl 1)</source>
          ,
          <volume>5200</volume>
          {
          <fpage>5205</fpage>
          (
          <year>2004</year>
          )
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>