<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>A research topic evolution prediction approach based on multiplex-graph representation learning ⋆</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Yang Zheng</string-name>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Kaiwen Shi</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Yuhang Dong</string-name>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Xiaoguang Wang</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Hongyu Wang</string-name>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>School of Information Engineering, Zhongnan University of Economics and Law</institution>
          ,
          <addr-line>Wuhan</addr-line>
          ,
          <country country="CN">China</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>School of Information Management, Wuhan University</institution>
          ,
          <addr-line>Wuhan</addr-line>
          ,
          <country country="CN">China</country>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>School of Management, Wuhan University of Technology</institution>
          ,
          <addr-line>Wuhan</addr-line>
          ,
          <country country="CN">China</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>The intensification of international technological innovation competition and the evolution of scientific research paradigms have led to a continuous expansion of scientific literature, making information analysis increasingly complex and diversified. To address the challenges of accurately assessing topic evolution within the context of vast literature big data, traditional methods of expert evaluation or visualization analysis based on scientific knowledge networks are inadequate. From the perspectives of artificial intelligence and big data, this paper proposes a universal method for automated and intelligent discrimination and prediction of research topic evolution hotness. This method involves integrating content and structural features of keywords to track the evolution of keyword frequency strength over time in research topic networks characterized by keywords. This study conducts a case analysis in the field of information science. The results demonstrate that the prediction of keyword strength is improved after integrating content and structural features, which has significant reference value for tasks such as future research topic evolution trend discrimination, research direction, and policy planning.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;topic evolution</kwd>
        <kwd>keyword citation network</kwd>
        <kwd>text mining</kwd>
        <kwd>graph representation learning 1</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>With the intensification of international technological
innovation competition and the evolution of the
fourth paradigm of scientific research driven by big
data development, the growing volume of scientific
literature, shifting scholarly interests, and the
emergence of new research topics pose significant
challenges to traditional methods of research topic
analysis[1,2,3]. How to comprehensively and finely
reveal the research topics and their characteristic
dimensions: content, structure, and strength[4,5].
Utilizing the powerful representation and feature
learning capabilities of deep representation learning
algorithms such as word embedding and graph
embedding[6], it is possible to model the complex
nonlinear relationships between entities represented
by keywords, the smallest units of knowledge[7,8].
Tracking changes in topic strength, as indicated by
keyword frequency over time, can reflect the evolving
trends of topics[9,10,11].</p>
      <p>Therefore, this paper utilizes a multiplex-graph
representation learning method combined with
interactions in topic keyword content and structure to
assess changes in topic strength, achieving prediction
of topic evolution hotness. And select the field of
"Information Science" for case analysis, aiming to
address the following two scientific questions:
1. How to reveal the evolution process of topics
at the micro level, thereby tracking the
evolutionary trends of research topics?
2. How to effectively and comprehensively
model and integrate the multi-dimensional
Step1:Data Preparation
Field-specific literature is selected from databases
such as Web of Science and Scopus, with titles,
keywords, abstracts, and references extracted as basic
data. After cleaning and filtering the data, an original
dataset  is constructed, which specifically includes
the keyword citation relationship dataset  , the
keyword frequency dataset  , and the integrated
dataset  containing titles, abstracts, and keywords.
-- LDaantagsueatg:eW: OEnSglish Data cleaning
- Collection: SCI &amp; SSCI
- Query : ‘Information Science’ Hump naming
- Year : 2010 – 2023
- Amount: 58,119
- Fields : TI, AB, DE, CR,
PY, DOI
Step3: Evolution prediction</p>
      <p>MAE</p>
      <p>MSE</p>
      <p>Evaluation and testing
Forecasting task MAE MSE
1.93 13.41
1.91 12.63
1.90 13.26
1.88 12.24</p>
      <p>Null-data validation
data process</p>
      <p>Word Frequency</p>
      <p>Ai 16
BigData 84
Gcn 26
Cat 9
... ...</p>
      <p>Bert 33</p>
      <p>Gpt 124
Example of result</p>
      <p>Data extraction
Year: 2010 . 2022 2023</p>
      <p>Datasets for
2010-2023
+
+
+ +
Representation fusion
experimental group</p>
      <p>MLP</p>
      <p>forecast
Word frequency</p>
      <p>Distance
calculation
Fusion of representations
features of research topic evolution
representations on knowledge networks?
After tracking these multi-dimensional
evolutionary features, will the assessment of
topic evolution trends become more
accurate?</p>
    </sec>
    <sec id="sec-2">
      <title>2. Methodology</title>
      <p>This study aims to predict the strength of keyword
frequency, using changes in keyword frequency
strength to reflect variation in topic evolution hotness.</p>
      <p>The research process is divided into 3 steps: Step
1 involves retrieving and cleaning the source data to
obtain all the data needed for subsequent experiment.
Step 2 involves obtaining content, structure, and
strength representations of keywords. Step 3 utilizes
deep learning models to integrate multi-dimensional
data representations and conduct prediction of topic
evolution hotness with the integrated experimental
data. Figure 1 shows the detailed research process.
Step2:Multidimensional data representation</p>
      <p>Dataset of titles, abstracts, and keywords
Dataset of keyword citation relationships</p>
      <p>Dataset of keyword frequencies</p>
      <p>Year:2010 . 2022 2023 YeGar:lo20V1e0, .N2u0m22py2,02P3andas and oYtehaer:r2m01o0d.el2s0a22nd20t2o3olkits
Representation of content Representation of reference structure Representation of strength
GCN
GAT</p>
      <p>Hotness
Citation
Semantic</p>
      <p>Represent
data
MAE</p>
      <p>Back-Propagation</p>
      <sec id="sec-2-1">
        <title>2.2. Multi-dimensional feature extraction</title>
        <p>To prepare for multi-dimensional feature integration,
this study will perform feature extraction on keyword
data across three representational dimensions:
content, structure, and strength of keywords.</p>
        <p>（1）Content feature extraction
In the process of extracting keyword content features,
this study opts to use the  static word
embedding method to capture the semantic
relationships of keywords in their global context[12],
facilitating the embedding of keywords, as it offers
greater stability and requires less computational
resources[13]. The principle is as shown in 2-1 and
2The generated word vector is then used to
calculate the cosine similarity between words using
formula 2-3. This process results in obtaining the
semantic distance matrix 
keyword content feature extraction.</p>
        <p>=   = (  , ),  =  , for
 = [ 1,  2, ,   ], B = [ 1,  2, ,   ]
 
=
∑

 =1   ×  
√∑ =1   2 × √∑ =1   2
（2）Structure feature extraction
Some scholars have proposed using a
"keywordcitation-keyword"</p>
        <p>
          method to construct keyword
citation networks[14], meaning when literature  1
cites literature  2, there exists a "Keyword-Cartesian
product mapping" citation relationship between the
keywords of the two literatures. The specific principle
is shown in Figure 2. Based on this theory, this study
constructs a keyword citation network with citation
frequency as edge weight, resulting in a keyword

citation matrix   =    = (  , ),  =  .
(2-1)
(
          <xref ref-type="bibr" rid="ref2">2-2</xref>
          )
(
          <xref ref-type="bibr" rid="ref2 ref3">2-3</xref>
          )
dataset  , constructing a "frequency-year" frequency


1
2
3
4

∑
=  =1
        </p>
        <p>|  −  ̂ |

=  =1   −  ̂ 2
∑
2, where   is the number of times word  appears in
the context of word  ,   is the total number of words
appearing in the context of word  , and   is the
probability of word  appeare in the context of word  .</p>
        <p>
          = ∑  
  =   |  =  

 
(
          <xref ref-type="bibr" rid="ref2 ref3 ref4">2-4</xref>
          )
(
          <xref ref-type="bibr" rid="ref2 ref3 ref4 ref5">2-5</xref>
          )
matrix 
features,
=   =  ℎ . While extracting strength
this
also
generates
the
strength
representation  of keywords.
        </p>
      </sec>
      <sec id="sec-2-2">
        <title>2.3. Model construction and prediction</title>
        <p>Graph Attention Network(GAT)based on the attention
mechanism, can effectively capture complex semantic
dependencies
between
keywords,
while</p>
        <sec id="sec-2-2-1">
          <title>Graph</title>
          <p>Convolutional Network(GCN) can efficiently process
the strcture information of the graph structure itself
by aggregating the features of neighboring nodes.
Therefore, this study chooses to use GAT and GCN
graph
neural network
models to capture the
relationships between nodes in graph-structured data
from content and structure perspectives, respectively,
and
employs
an</p>
        </sec>
        <sec id="sec-2-2-2">
          <title>Multilayer</title>
        </sec>
        <sec id="sec-2-2-3">
          <title>Perceptron(MLP)</title>
          <p>regression
model to integrate
multi-dimensional
features of keywords for strength prediction.</p>
          <p>This study constructs an ablation experiment
group, as shown in Table 1, for predicting the hotness
of topic evolution. And details of the model settings
are shown in Figure 1. It's worth noting that set an
initial identity matrix allows the neural network to
gradually adjust and optimize feature representations
during
the
learning
process.</p>
          <p>Therefore,
after
obtaining the content matrix 
and the structure
matrix</p>
          <p>, GAT and GCN are used to perform
convolution operations on these two matrices on a
predefined 50-dimensional identity
matrix. After
obtaining content representation  and structure
representation

of
keywords,
the
three
representation data are concatenated directly for
integration, and prediction is made based on MLP[15].</p>
          <p>Above,   represents the ith element of  , and  is
the number of elements.</p>
        </sec>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>3. Experiment</title>
      <p>The detailed process of data acquisition can be found
in the appendix under  .   .</p>
      <sec id="sec-3-1">
        <title>3.1. Experiment Preparation</title>
        <p>When extracting features from experimental data, this
study choose to work with four sets of yearly data:
2017-2019, 2018-2020, 2019-2021, and 2020-2022.
The first three sets are used as the training groups,
and the last set as the test group. Specifically, the data
from 2019 to 2022 serve as the basis for operations
(all operations on yearly data will follow this
fouryear standard). Taking 2019 as an example, semantic
content matrix  , citation structure matrix  , and
frequency strength matrix  are constructed for
that year's keyword data.</p>
      </sec>
      <sec id="sec-3-2">
        <title>3.2. Model training and prediction</title>
        <p>
          The objective is to determine the optimal parameters
for various model groups in order to ensure accurate
predictions. This study has set four hyperparameters:
learning rate (1e-2, 1e-3, 1e-4), number of training
epochs (10, 50, 100, 200), hidden layer dimensions
(10, 30, 50), and stopping steps (
          <xref ref-type="bibr" rid="ref5">5, 10, 20</xref>
          ), using the
training data from 2019 to 2021 to train models
across four experimental groups, and evaluating the
final MAE and MSE results on the training set to
determine the most suitable hyperparameters for
each model. The optimal parameter settings for
different models are shown in Table 2.
        </p>
      </sec>
      <sec id="sec-3-3">
        <title>3.3. Results and Discusstion</title>
        <p>The prediction results are evaluated using two
indicators: MAE and MSE, with the evaluation results
listed in Table 3. From the table, it can be observed
that using only the strength representation  for
predicting topic hotness, the MAE and MSE between
the predicted and actual values are 1.93484 and
13.41658, respectively. However, after integrating the
content representation  or structure representation
 , the values of MAE and MSE both decrease. The best
result for predicting topic hotness are achieved by
integrating all three types of representations,
resulting in the lowest values of MAE and MSE.</p>
        <p>The results indicate that predicting the evolution
hotness of research topics by integrating
multidimensional features such as content and structure
through multiplex-graph representation learning is
more accurate than traditional prediction methods.</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>4. Conclusion</title>
      <p>This study proposes a novel approach based on
multiplex-graph representation learning to predict
the evolution of research topics. And the
contributions are follows: First, in feature modeling,
GCN and GAT graph neural network models are used
to perform convolution operations on content and
structure features on unit matrices of specified
dimensions, adaptively aligning data across different
dimensions and time windows to ensure
comparability. Second, this study integrates semantic
content features, citation structure features, and
frequency strength features of keywords for research
topic hotness prediction, showcasing the interaction
between knowledge structures and cognitive
structures from a multidimensional perspective,
offering a deeper insight into predicting research
topic evolution hotness. Third, after integrating
content and structure features, a domain case analysis
is conducted, and the result indicates that combining
these two types of features indeed makes the
prediction of research topic evolution hotness more
accurate.</p>
      <p>Owing to the desire to directly validate whether
integrating multiple representations of topic
evolution enhances the accuracy of topic evolution
analysis, this paper choose to predict the future
frequency of topic keywords, which has certain
limitations. Subsequent tasks such as research topic
trend discrimination, research direction, and policy
planning can be developed based on the effective
analysis results of this study.</p>
    </sec>
    <sec id="sec-5">
      <title>Acknowledgements</title>
      <p>This work was funded by the National Natural Science
Fund of China (No. 71874129), the Open-end Fund of
Information Engineering Lab of ISTIC and the
Independent Innovation Foundation of Wuhan
University of Technology (No. 233103002).
keywords in building an initial reading list of
research papers in scientific paper retrieval and
recommender systems." Information Processing
&amp; Management 53.3 (2017): 577-594.
[10] Yoon, Young Seog, et al. "Exploring the dynamic
knowledge structure of studies on the Internet
of things: Keyword analysis." ETRI Journal 40.6
(2018): 745-758.
[11] Ohniwa, Ryosuke L., and Aiko Hibino.
"Generating process of emerging topics in the
life sciences." Scientometrics 121.3 (2019):
1549-1561.
[12] Pennington, Jeffrey, Richard Socher, and
Christopher D. Manning. "Glove: Global vectors
for word representation." Proceedings of the
2014 conference on empirical methods in
natural language processing (EMNLP). 2014.
[13] Wang, Yuxuan, et al. "From static to dynamic
word representations: a survey." International
Journal of Machine Learning and Cybernetics 11
(2020): 1611-1630.
[14] Q. Chen, J. Wang, and W. Lu. "Discovering
Domain Vocabularies Based on Citation Co-word
Network" Data Analysis and Knowledge
Discovery 3.6 (2019): 57-65. (in Chinese)
[15] Liu, Weijia, et al. "Category-universal witness
discovery with attention mechanism in social
network." Information Processing &amp;
Management 59.4 (2022): 102947.
[16] Yan, Yuwei, et al. "Data mining of customer
choice behavior in internet of things within
relationship network." International Journal of
Information Management 50 (2020): 566-574.
[17] Gandhudi, Manoranjan, et al. "Causal aware
parameterized quantum stochastic gradient
descent for analyzing marketing advertisements
and sales forecasting." Information Processing &amp;
Management 60.5 (2023): 103473.</p>
    </sec>
    <sec id="sec-6">
      <title>A. Online Resources</title>
      <p>The resources of this article can be downloaded at
https://github.com/Hipkevin/EEKE-hotness.</p>
    </sec>
    <sec id="sec-7">
      <title>B. Data resources</title>
      <p>This study uses "Information Science" as a case study
topic, selecting the SCI and SSCI core databases in
WOS. Conducting literature searches in the
"Information Science &amp; Library Science" field with the
search query "Document Types: Article or Review
Article; Languages: English," ultimately selecting
literature from 2010 to 2023, totaling 58,119 articles,
as experimental data.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <surname>Zhu</surname>
          </string-name>
          ,
          <string-name>
            <surname>Hengmin</surname>
          </string-name>
          , et al.
          <article-title>"Evolution analysis of online topics based on 'word-topic'coupling network</article-title>
          .
          <source>" Scientometrics 127.7</source>
          (
          <year>2022</year>
          ):
          <fpage>3767</fpage>
          -
          <lpage>3792</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <surname>Hu</surname>
          </string-name>
          ,
          <string-name>
            <surname>Kai</surname>
          </string-name>
          , et al.
          <article-title>"Understanding the topic evolution of scientific literatures like an evolving city: Using Google Word2Vec model and spatial autocorrelation analysis</article-title>
          .
          <source>" Information Processing &amp; Management 56.4</source>
          (
          <year>2019</year>
          ):
          <fpage>1185</fpage>
          -
          <lpage>1203</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <surname>Huo</surname>
          </string-name>
          , Chaoguang, Shutian Ma, and Xiaozhong Liu.
          <article-title>"Hotness prediction of scientific topics based on a bibliographic knowledge graph."</article-title>
          <source>Information Processing &amp; Management 59.4</source>
          (
          <year>2022</year>
          ):
          <fpage>102980</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>Z.</given-names>
            <surname>Liu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Wang</surname>
          </string-name>
          , and
          <string-name>
            <given-names>R.</given-names>
            <surname>Bai</surname>
          </string-name>
          .
          <article-title>"Research on Visualization Analysis Method of Discipline Topics Evolution from the Perspective of Multi Dimensions:A Case Study of the Big Data in the Field of Library and Information Science in China."</article-title>
          <source>Journal of Library Science in China 42.6</source>
          (
          <year>2016</year>
          ):
          <fpage>67</fpage>
          -
          <lpage>84</lpage>
          . (in Chinese)
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>K.</given-names>
            <surname>Cui</surname>
          </string-name>
          .
          <source>The Research and lmplementation of Topic Evolution Based on LDA. Diss. National University of Defense Technology</source>
          <year>2010</year>
          .
          <article-title>(in Chinese)</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <surname>Zhou</surname>
          </string-name>
          ,
          <string-name>
            <surname>Yuan</surname>
          </string-name>
          , et al.
          <article-title>"A deep learning framework to early identify emerging technologies in largescale outlier patents: An empirical study of CNC machine tool</article-title>
          .
          <source>" Scientometrics</source>
          <volume>126</volume>
          (
          <year>2021</year>
          ):
          <fpage>969</fpage>
          -
          <lpage>994</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <surname>Şenel</surname>
            ,
            <given-names>Lütfi</given-names>
          </string-name>
          <string-name>
            <surname>Kerem</surname>
          </string-name>
          , et al.
          <article-title>"Learning interpretable word embeddings via bidirectional alignment of dimensions with semantic concepts</article-title>
          .
          <source>" Information Processing &amp; Management 59.3</source>
          (
          <year>2022</year>
          ):
          <fpage>102925</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <surname>Shi</surname>
          </string-name>
          ,
          <string-name>
            <surname>Bin</surname>
          </string-name>
          , et al.
          <article-title>"RelaGraph: Improving embedding on small-scale sparse knowledge graphs by neighborhood relations."</article-title>
          <source>Information Processing &amp; Management 60.5</source>
          (
          <year>2023</year>
          ):
          <fpage>103447</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <surname>Raamkumar</surname>
            ,
            <given-names>Aravind</given-names>
          </string-name>
          <string-name>
            <surname>Sesagiri</surname>
            , Schubert Foo, and
            <given-names>Natalie</given-names>
          </string-name>
          <string-name>
            <surname>Pang</surname>
          </string-name>
          .
          <article-title>"Using author-specified</article-title>
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>