<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Exploring Trends of Cancer Research Based on Topic Model</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Mingliang Cui</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Yanchun Liang</string-name>
          <email>ycliang@jlu.edu.cn</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Yuping Li</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Renchu Guan</string-name>
          <email>guanrenchu@jlu.edu.cn</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>College of Computer Science and Technology, Jilin University</institution>
          ,
          <addr-line>Changchun, 130012</addr-line>
          ,
          <country country="CN">China</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Yanchun Liang</institution>
        </aff>
      </contrib-group>
      <fpage>7</fpage>
      <lpage>18</lpage>
      <abstract>
        <p>Cancer research is of great importance in life science and medicine and attracts research funds of thousands of millions dollars each year. With the explosion of biomedical research papers, it becomes more and more necessary to show the research trend in this spotlight area. In this paper, to provide a straightforward research atlas for the top killer cancers, Latent Dirichlet Allocation (LDA) is performed on the massive quantities of biomedical literatures. Moreover, Gibbs Sampling is used to make assessment on the parameters of the LDA model. The proposed evaluation carried out under multiple conditions with different Ks (the number of topics) for the top five cancers in recent five years. Additionally, a biomedical topic model was generated with the LDA model and delicate analysis was performed on the basis of that in order to explore the trending topic in cancer research. It can help the biology and medicine doctors quickly catch the frontiers of the cancer study, improve and expand their research programme, especially in today's era of “big data”.</p>
      </abstract>
      <kwd-group>
        <kwd>Topic Model</kwd>
        <kwd>LDA</kwd>
        <kwd>Gibbs Sampling</kwd>
        <kwd>Topic Analysis</kwd>
        <kwd>Cancer Research Trend</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>1.1</p>
    </sec>
    <sec id="sec-2">
      <title>Introduction</title>
      <sec id="sec-2-1">
        <title>Significance</title>
        <p>
          Cancer is of great threat to human health, scientists and Medicians have been
continuously looking for effective treatment to conquer it, however, cancer research is
a very challenging field in today's life science. Over the past 5 years cancer research
has diverged enormously, partly based on the quickly development of biotechnology
and bioinformatics. It is not easy to summarize recent trends for different cancers’
study, and identify where the new findings are and therapies might come from [
          <xref ref-type="bibr" rid="ref1">1</xref>
          ].
Under this circumstance, we choose an appropriate machine learning method―topic
model―to explore the trend in cancer research, specifically, the topics in cancer
research papers.
        </p>
        <p>
          The topic model is a probabilistic model of text mining appeared in recent years [
          <xref ref-type="bibr" rid="ref2">2</xref>
          ].
It is an algorithm that can discover the topic structure hidden in large-scale data. In
topic model, the vocabulary items is visible while the topic structure hidden. In order
to reduce the dimensions of the feature vector space, texts are usually mapped into the
topic space via the topic model. It is different from the traditional vector space model,
which just simply considers each document as a sample and each word as a feature.
Instead, it maps a high dimensional frequency space to a lower dimensional topic
space. Moreover, the topic model can capture the semantic information, which can
reveal that the latent relations among documents. It also can effectively solve the
polysemy, synonym and other problems, which has the vital significance in document
feature extraction and content analysis. However, using a topic model (i.e. Latent
Dirichlet Allocation) to analyze the trends of cancers has not been reported.
        </p>
        <p>The rest of the article is organized as follows: we start in section 2 with a brief
review of a topic model named as Latent Dirichlet Allocation and its related work. In
section 3, the algorithm of Gibbs Sampling is introduced, and the framework of our
exploration is described in detail. It presents the experimental methodology and
results on Medline dataset in Section 4. At last, conclusions and future work are
depicted in Section 5.
2</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>Background</title>
      <p>Hyper parameters are subject to Dirichlet distribution. We explain the symbol used
to describe the model in Table.1. Now we have a text representation of the probability:
    
p(wm , zm ,θ m , Φ |α ,β )</p>
      <p>Nm     
=p(wm,n ∏ |ϕ zm,n ) p(zm,n |θ m ) ⋅ p(θ m |α ) ⋅ p(Φ | β )
n=1
(1)</p>
      <p>The convergence of LDA is a more critical issue. We use the convergence function
to solve it. Under the given model conditions, we choose the appearance probability
of samples as the evaluation criterion of the model. The performance of the LDA
model:
(2)
(3)
(4)
 M  
p(w | M) = ∏ p(w m |θ m , Φ)</p>
      <p>m=1
=∑K( ∏M ∏Nm nkt + β t ⋅
m=1 n=1 k=1 ∑ Vt =+β 1 nkt t</p>
      <p>nmk + α k
∑ kK =+α 1 nmk k
) .</p>
      <p>
        For the convenience of the calculation, we use log transform to the equation (1),
denoted as . Along with the iterative constantly, is used to determine the model
convergence [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ].
      </p>
      <p>Lˆ = - ∑M ∑Nm log2 ( ∑K ( nkt + β t ⋅
m =n=1 1 k =∑Vt 1 =+β 1 nkt t</p>
      <p>nmk +α k
∑ kK =+α 1 nmk k
))
3
3.1</p>
    </sec>
    <sec id="sec-4">
      <title>Main Work</title>
      <sec id="sec-4-1">
        <title>Core Method</title>
        <p>
          In fact, LDA is one of the probabilistic. Data are generally divided into two parts,
visible variables and latent variables. It is believed in the topic model that the data is
produced by the generation process, which defines the joint probability distribution of
visible random variables and latent random variables. For many modern probabilistic
models, including Bayesian statistics, the priori probability calculation is extremely
difficult. So the core research objective of modern probability modelling is to do
everything possible to obtain an approximate solution. Random sampling is a kind of
methods for solving the approximate solution with good performance. This article
describes a method commonly used sampling MCMC (Markov Chain Monte Carlo)
[
          <xref ref-type="bibr" rid="ref14">14</xref>
          ] and Gibbs Sampling algorithm [
          <xref ref-type="bibr" rid="ref15">15</xref>
          ], Gibbs Sampling algorithm has been widely
used in modern Bayesian analysis.
        </p>
        <p>
          MCMC methods have Gibbs Sampling algorithm and Metropolis-Hastings (MH)
[
          <xref ref-type="bibr" rid="ref16">16</xref>
          ] algorithm commonly used sampling methods. Gibbs Sampling algorithm is a
special case of MH algorithm, Gibbs Sampling from a high-dimensional space are
sampled separately for each dimension, and gradually get higher dimensional
sampling points, making sampling difficult to reduce.
        </p>
        <p>N-dimensional Gibbs Sampling:
1. Random initialization { xi : i =1, 2, ⋅ ⋅ ⋅, n }
2. For t</p>
        <p>=0,1, 2, ⋅ ⋅ ⋅ loop in sampling
x 1(t +1) ～
p(x1 | xt2 , xt3 , ⋅ ⋅ ⋅, xtn )</p>
      </sec>
      <sec id="sec-4-2">
        <title>Complete Process</title>
        <p>
          We choose the top 5 cancer from the top 10 deadliest cancer published by the
LiveScience [
          <xref ref-type="bibr" rid="ref17">17</xref>
          ], which are Breast Cancer, Lung Cancer, Pancreatic Cancer, Prostate
Cancer and Colon Cancer. We search and download the research paper related to
these cancers from NCBI PubMed from 2010 to 2014 separately [
          <xref ref-type="bibr" rid="ref18">18</xref>
          ].
        </p>
        <p>
          Then we make a pre-process to extract the title and abstract for each paper and
make some hyphens for the entity name for a better segmentation. After that we use
WordNet [
          <xref ref-type="bibr" rid="ref19">19</xref>
          ] to stem the text in order to get an exact description for the topic. Before
the document input into model, we will do the pre-processing for the document,
thereby obtaining the document term matrix. It can be seen by term matrix that how
many documents in corpus, how many word terms and how frequently each term
appears in a document. The inputs of the model are the document collection , the
number of topics , and the hyper parameters .The topic of a number K we need
to specify its value according to the experience, we want to take advantages of K
solution, the need for repeated experiments, and then to carry on the value according
to the different K value under the situation of convergence. After repeated
experiments of LDA convergence on the data, we get a reasonable value of K=100.
The hyper parameter is the Dirichlet of the prior distribution. In fact they have
smooth effect on data. Because there is no supervision information too much, we
assign the hyper parameters an empirical value , , tending to take
symmetry value [
          <xref ref-type="bibr" rid="ref20">20</xref>
          ].
        </p>
        <p>Take the dataset of Pancreatic Cancer in 2010 as an example, the number of the
abstract is 2088, includes 54004 words. The number of topic K is set to 100, 200, 300,
400, 500, considered for model convergence condition.</p>
        <p>According to the evaluation function of the convergence mentioned in Section 2,
which shows the cost of compression of the text using the model, the smaller the
better. Figure.2 below respectively under different K values for the iterative
convergence condition of iteration times of 400 and 1000.</p>
        <p>
          We measure the similarity and differences between the 5 cancer research by their
word vector cosine coefficients similarity [
          <xref ref-type="bibr" rid="ref22">22</xref>
          ] (formula 6).
(5)
(6)
        </p>
        <p>
          We further analysis the output of the LDA Gibbs Sampling, particularly in the
topic words file (this file contains topic words most likely words of each topic). We
first transform the topic-word matrix to a topic word vector representing the selection
and the frequency of topic words to describe one cancer research. Then we use feature
scaling normalization [
          <xref ref-type="bibr" rid="ref21">21</xref>
          ] (formula 5) to deal with the word vector, in order to
compare them between different years and different cancers.
        </p>
        <p>Unlike the cosine coefficients which give a numerical description on the trend and
topic of the 5 cancers, the common words are easier for readers to have an intuition on
the 5 cancers topics. We do the both analysis of the results so that we may get a
comprehensive answer.
4</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>Experiments and Discussion</title>
      <p>
        To examine the behavior and the performance of LDA, the experiments are illustrated
on the widely used Medline dataset. In order to straightforwardly compare the trend
and topic resulted from the LDA model, the visual results to get the entire recognition
of the topic words of different cancers are shown, which uses the Word Cloud tool
supported by Tagul.com [
        <xref ref-type="bibr" rid="ref24">24</xref>
        ]. And we also draw the trend of cancer research in
different years, supported by Plot.ly [
        <xref ref-type="bibr" rid="ref25">25</xref>
        ].
      </p>
      <sec id="sec-5-1">
        <title>Experimental Setup</title>
        <p>The publicly available Medline dataset is provided by NCBI PubMed, which contains
a title and an abstract for each paper on the five fatal cancers. In order to get a better
visual understanding of the data, the 5 cancers with different colors (in details, breast
cancer blue, colon cancer orange, lung cancer green, pancreatic cancer red, prostate
cancer purple) are listed in Table 2. An outline of the selected dataset is shown below
(actually we collect the 2014 data before October 2014, that is why it seems all the
2014 paper decrease from the previous years):
After introducing the Gibb Sampling and transforming the topic-word matrix to the
vector, we draw all the five cancer topic words in the recent five years. A snapshot of
the collection of the topic words clouds are shown in Figure.3.</p>
        <p>Filtering the topic words for each cancer into the common words collection, we can
easily discover the main topic for each cancer and new findings and focus for each
year. Take breast cancer as an example shown below:</p>
        <p>In the Figure.4 above, we can find out the big words in it like cell, breast cancer,
tumor, patient, woman and so on, these words are simply the most common used to
elaborate breast cancer, which means if you talk about breast cancer and you just
can’t avoid mentioning them. They are obviously from the different topics, therapy,
factor, risk and other aspects of the breast cancer. The common words are defined as
the words which appear every year for certain cancer research during 2010-2014.
They are shown in Figure.5 below, and the shape of the words represent their
frequency.
4.3</p>
      </sec>
      <sec id="sec-5-2">
        <title>General Comparison and Discussions</title>
        <p>The Figure.6 above shows the contrast result of the vector cosine coefficient between
two cancers in the same year. The higher coefficient, the more similar. We can easily
find out that all the similarities contrast with colon cancer (see colon&amp;lung, colon &amp;
pancreatic, breast&amp;colon, colon&amp;prostate) are higher than the other cancer pairs. It
indicates colon cancer research is more related with the others. It is because that: lung
cancer may easily metastasize to colon; colon cancer and pancreatic cancer are both
belong to lower digestive cancers; lack of exercise and sitting for long time are the
common causes of breast and colon cancer; long-term androgen deprivation therapy
for prostate cancer may increase the risk of colon cancer.</p>
        <p>And we find it interesting that almost all the five cancer research diverse from each
other in 2012, because they have a concave at the time in the figure. Fig.7 shows the
topic words changes for each cancer research during the recent five years. We can
learn from the figure that breast cancer, lung cancer, prostate cancer, these 3 cancer
research only change a little, and pancreatic cancer research changes a lot. It is
because that the pancreatic incidence rate increase fast recently and it is named as
“the king of cancer” because the lowest survival rate. Many scientists engage in the
research of pancreatic cancer.
In this paper, we first apply LDA Gibbs Sampling model on the analysis of the top 5
deadliest cancer research trends, which is extended from cosine coefficient using the
vector space model after transforming topic-word matrix into topic word vector. Then
we generate the common topic word collection for each cancer research, in order to
get the trending topic words. We further explore the trend by comparing and
contrasting topic words for each cancer and their cosine coefficients. It is found that
the trending topic words for the 5 cancers research from 2010 to 2014, which are
depicted in the words clouds. Moreover, it is found that numerical trends of the four
cancers research are as follows: breast cancer, lung cancer, and prostate cancer are of
little change, but pancreatic cancer is changing a lot in the recent five years. But what
cause all five cancer research diverse in 2012, how to visualize the allocation of the
topics for each cancer research, and how to make a computational evaluation for the
trend results, are all what need to be explored in our future work.</p>
      </sec>
    </sec>
    <sec id="sec-6">
      <title>Acknowledgments.</title>
      <p>The authors are grateful to the support of the NSFC (61272207, 61472158, 61103092)
and the Science Technology Development Project from Jilin Province
(20130522106JH, 20140520070JH）</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Cao</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          ,
          <string-name>
            <given-names>DePinho</given-names>
            , R.,
            <surname>Ernst</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            ,
            <surname>Vousden</surname>
          </string-name>
          ,
          <string-name>
            <surname>K.</surname>
          </string-name>
          :
          <article-title>Cancer research: past, present and future</article-title>
          .
          <source>Nature Reviews Cancer</source>
          .
          <volume>11</volume>
          ,
          <fpage>749</fpage>
          -
          <lpage>754</lpage>
          (
          <year>2011</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Rosen-Zvi</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Griffiths</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Steyvers</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>Smyth</surname>
            ,
            <given-names>P.:</given-names>
          </string-name>
          <article-title>The author-topic model for authors and documents</article-title>
          .
          <source>Proceedings of the 20th conference on Uncertainty in artificial intelligence</source>
          .
          <fpage>487</fpage>
          -
          <lpage>494</lpage>
          (
          <year>2004</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Papadimitriou</surname>
            ,
            <given-names>C. H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Tamaki</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Raghavan</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>Vempala</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          :
          <article-title>Latent semantic indexing: A probabilistic analysis</article-title>
          .
          <source>Proceedings of the seventeenth ACM SIGACT-SIGMOD-SIGART symposium on Principles of database systems</source>
          .
          <volume>159</volume>
          -
          <fpage>168</fpage>
          (
          <year>1998</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Hofmann</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          :
          <article-title>Probabilistic latent semantic indexing</article-title>
          .
          <source>Proceedings of the 22nd annual international ACM SIGIR conference on Research and development in information retrieval</source>
          .
          <volume>50</volume>
          -
          <fpage>57</fpage>
          (
          <year>1999</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Blei</surname>
            ,
            <given-names>D. M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ng</surname>
            ,
            <given-names>A. Y.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>Jordan</surname>
            ,
            <given-names>M. I.</given-names>
          </string-name>
          :
          <article-title>Latent dirichlet allocation</article-title>
          .
          <source>the Journal of machine Learning research</source>
          ,
          <volume>3</volume>
          ,
          <fpage>993</fpage>
          -
          <lpage>1022</lpage>
          (
          <year>2003</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Teh</surname>
            ,
            <given-names>Y. W.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Jordan</surname>
            ,
            <given-names>M. I.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Beal</surname>
            ,
            <given-names>M. J.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>Blei</surname>
            ,
            <given-names>D. M.</given-names>
          </string-name>
          :
          <article-title>Hierarchical dirichlet processes</article-title>
          .
          <source>Journal of the american statistical association 101</source>
          .476 (
          <year>2006</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>Mcauliffe</surname>
            ,
            <given-names>Jon D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>David</surname>
            <given-names>M.</given-names>
          </string-name>
          <string-name>
            <surname>Blei</surname>
          </string-name>
          .:
          <article-title>Supervised topic models</article-title>
          .
          <source>Advances in neural information processing systems</source>
          .
          <volume>121</volume>
          -
          <fpage>128</fpage>
          (
          <year>2008</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>Petinot</surname>
          </string-name>
          , Yves,
          <string-name>
            <surname>Kathleen McKeown</surname>
            ,
            <given-names>and Kapil</given-names>
          </string-name>
          <string-name>
            <surname>Thadani</surname>
          </string-name>
          .
          <article-title>: A hierarchical model of web summaries</article-title>
          .
          <source>In: Proceedings of the 49th Annual Meeting of the Association for Computational Linguistics: Human Language Technologies: short papers-Volume</source>
          <volume>2</volume>
          .
          <fpage>670</fpage>
          -
          <lpage>675</lpage>
          . Association for Computational Linguistics (
          <year>2011</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <surname>Diebolt</surname>
            , Jean,
            <given-names>Christian P.</given-names>
          </string-name>
          <string-name>
            <surname>Robert</surname>
          </string-name>
          .:
          <article-title>Estimation of finite mixture distributions through Bayesian sampling</article-title>
          .
          <source>Journal of the Royal Statistical Society. Series B (Methodological)</source>
          .
          <fpage>363</fpage>
          -
          <lpage>375</lpage>
          (
          <year>1994</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <given-names>S.</given-names>
            <surname>Kotz</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Balakrishnan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N. L.</given-names>
            <surname>Johnson</surname>
          </string-name>
          <article-title>: Dirichlet and Inverted Dirichlet Distributions</article-title>
          .
          <source>Continuous Multivariate Distributions, Models and Applications</source>
          . New York: Wiley,
          <year>2000</year>
          ., United
          <string-name>
            <surname>States</surname>
          </string-name>
          (
          <year>2004</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <string-name>
            <surname>Lindley</surname>
          </string-name>
          , Dennis V.:
          <article-title>The use of prior probability distributions in statistical inference and decision</article-title>
          .
          <source>Proc. 4th Berkeley Symp. on Math. Stat. and Prob</source>
          .
          <volume>453</volume>
          -
          <fpage>468</fpage>
          (
          <year>1961</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          12.
          <string-name>
            <surname>Beal</surname>
          </string-name>
          , Matthew James.:
          <article-title>Variational algorithms for approximate Bayesian inference</article-title>
          .
          <source>PhD diss</source>
          . University of London (
          <year>2003</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          13.
          <string-name>
            <surname>Heinrich</surname>
          </string-name>
          , Gregor.:
          <article-title>Parameter estimation for text analysis</article-title>
          .
          <source>Technical report</source>
          . (
          <year>2005</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          14.
          <string-name>
            <surname>Gilks</surname>
          </string-name>
          , Walter R.:
          <article-title>Markov chain monte carlo</article-title>
          . John Wiley &amp; Sons, Ltd. (
          <year>2005</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          15.
          <string-name>
            <surname>Porteous</surname>
            , Ian, David Newman,
            <given-names>Alexander</given-names>
          </string-name>
          <string-name>
            <surname>Ihler</surname>
          </string-name>
          , Arthur Asuncion, Padhraic Smyth, Max Welling.:
          <article-title>Fast collapsed gibbs sampling for latent dirichlet allocation</article-title>
          .
          <source>Proceedings of the 14th ACM SIGKDD international conference on Knowledge discovery and data mining</source>
          .
          <volume>569</volume>
          -
          <fpage>577</fpage>
          (
          <year>2008</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          16.
          <string-name>
            <surname>Chib</surname>
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Greenberg</surname>
            <given-names>E.</given-names>
          </string-name>
          :
          <article-title>Understanding the metropolis-hastings algorithm</article-title>
          .
          <article-title>The american statistician</article-title>
          .
          <volume>49</volume>
          (
          <issue>4</issue>
          ):
          <fpage>327</fpage>
          -
          <lpage>335</lpage>
          (
          <year>1995</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          17.
          <string-name>
            <surname>Chan</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          :
          <article-title>The 10 Deadliest Cancers and Why There's No Cure</article-title>
          , http://www.livescience.com/11041-10
          <article-title>-deadliest-cancers-cure</article-title>
          .html.
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          18.
          <string-name>
            <surname>Ncbi</surname>
          </string-name>
          .nlm.nih.gov,: Home - PubMed - NCBI, http://www.ncbi.nlm.nih.gov/pubmed.
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          19. University, P.:
          <string-name>
            <surname>About WordNet - WordNet - About WordNet</surname>
          </string-name>
          , http://wordnet.princeton.edu/.
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          20.
          <string-name>
            <surname>Griffiths</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <article-title>Steyvers: Finding scientific topics</article-title>
          .
          <source>Proceedings of the National Academy of Sciences. 101</source>
          ,
          <fpage>5228</fpage>
          -
          <lpage>5235</lpage>
          (
          <year>2004</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          21.
          <string-name>
            <surname>Wikipedia</surname>
          </string-name>
          ,: Normalization, http://en.wikipedia.org/wiki/Normalization.
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          22.
          <string-name>
            <surname>Wikipedia</surname>
          </string-name>
          ,: Cosine similarity, http://en.wikipedia.org/wiki/Cosine_similarity.
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          23.
          <string-name>
            <given-names>O</given-names>
            <surname>'Rourke</surname>
          </string-name>
          , Norm,
          <string-name>
            <given-names>R.</given-names>
            <surname>Psych</surname>
          </string-name>
          , Larry Hatcher.
          <article-title>: A step-by-step approach to using SAS for factor analysis and structural equation modeling</article-title>
          .
          <source>Sas Institute</source>
          . (
          <year>2013</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          24.
          <string-name>
            <surname>Tagul</surname>
          </string-name>
          .com,: Tagul - Gorgeous word clouds, http://tagul.com.
        </mixed-citation>
      </ref>
      <ref id="ref25">
        <mixed-citation>
          25.
          <string-name>
            <surname>Plot</surname>
          </string-name>
          .ly,: Plotly, https://plot.ly/.
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>