<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Topic Evolution Path and Semantic Relationship Discovery Based on ∗ Patent Entity Relationship</article-title>
      </title-group>
      <contrib-group>
        <aff id="aff0">
          <label>0</label>
          <institution>Department of Information Management, School of Economics &amp; Management Nanjing University of Science and Technology Nanjing</institution>
          <country country="CN">China</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Jinzhu Zhang</institution>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>Linqi Jiang Department of Information Management, School of Economics &amp; Management Nanjing University of Science and Technology Nanjing</institution>
          <country country="CN">China</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2021</year>
      </pub-date>
      <fpage>77</fpage>
      <lpage>79</lpage>
      <abstract>
        <p>Topic evolution analysis describes the emergence, transition, and extinction of a topic in a technical field, which can help researchers understand the history and current situation of the research field. Current studies is mainly patent text-based methods, which often uses relationships among keywords to construct co-occurrence network and analyses evolution using topic clustering algorithms. However, it didn't consider all the words in the patent and the semantic relationship between them. In addition, the relationships among topics should be more concrete, we should not only find the evolution relationship, but also need to reveal the semantic relationships among topics. Therefore, this paper uses representation learning method to get the semantic representation of each entity/word, and computes the semantic similarity among them to find out pairs of words which are different but with the same meaning in a special context. Moreover, we define multiple semantic relationships among topics, and design a method to use patent entity relationships to obtain the semantic relationships among topics. Experiments in the technical field of UAV transportation have confirmed that the method in this paper can effectively identify the evolutionary relationship between topics and the semantic relationship between topic, Make the evolutionary relationship between topics more abundant and Interpretable. And provide a reference for further enriching and improving the topic evolution analysis method.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;Topic Evolution Path</kwd>
        <kwd>semantic relationship between topic</kwd>
        <kwd>Patent Entity Relationship</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>1</p>
    </sec>
    <sec id="sec-2">
      <title>Introduction</title>
      <p>
        Topic evolution analysis describes the emergence, transition, and
extinction of a topic in a technical field, which can help
researchers understand the history and current situation of the
research field. The result can quickly identify research hotspots,
trends and gaps, which is essential to scientific and technological
innovation
        <xref ref-type="bibr" rid="ref1">(Liu H,2020)</xref>
        .
      </p>
      <p>In the study of topic evolution analysis, topic evolution path and
relationship discovery play important role in related research.</p>
      <p>
        Current studies could be classified into two classes, including
patent citation analysis-based and patent text-based methods
        <xref ref-type="bibr" rid="ref2">(Yu
D,2020)</xref>
        . This paper focuses on the latter method, which often
uses relationships among keywords to construct co-occurrence
network and analyses evolution using topic clustering algorithms
        <xref ref-type="bibr" rid="ref3">(No H. J,2015)</xref>
        . Then the evolution path and relationship are
discovered through comparisons of common keywords among
topics in different time series.
      </p>
      <p>However, these common keywords cannot cover the pair of words
which are different but with the same meaning in a special context.</p>
      <p>In addition, the relationships among topics should be more
concrete, for example, we should not only find the evolution
relationship (i.e., emergence, transition, and extinction), but also
need to reveal the semantic relationships among topics (i.e.,
function-realization or function-area).</p>
      <p>
        Therefore, this paper uses representation learning method
        <xref ref-type="bibr" rid="ref4">(Birunda,2021)</xref>
        to get the semantic representation of each
entity/word, and computes the semantic similarity among them to
find out pairs of words which are different but with the same
meaning in a special context. Moreover, we define multiple
semantic relationships among topics, and design a method to use
patent entity relationships to obtain the semantic relationships
among topics.
Firstly, a manual labelled dataset is made and a neural network
model is trained to extract all patent entities. Then the topic is
identified by the clustering method and the evolution path is
Copyright 2021 for this paper by its authors. Use permitted under Creative Commons License Attribution 4.0 International (CC BY 4.0).
detected based on semantic similarity among. Finally, a neural
network model is trained to extract the relationship between
patent entities, and the semantic relationship among topics is
discovered through relationships between patent entities
      </p>
    </sec>
    <sec id="sec-3">
      <title>2.1 Data collection</title>
      <p>This paper uses the Derwent Innovations Index database as the
patent data retrieval platform and the data is retrieved on July
2019. The patent search expression is “IP = B64* AND TI=
(((unmanned OR automatic OR autonomous OR remotely piloted OR
nonhuman) AND (aircraft OR "aerial vehicle" OR airship* OR
drone OR plane OR aircraft* OR airplane OR aerobat* OR
aerostat*)) OR "UAV"”, the time interval is from 2008 to 2020. A
total of 4507 patents with title, abstract, patent application date
and other features are retrieved and processed as the data source.
It is divided into different time series for evolution analysis
considering the number of patents in each time period.
2.2 Discovery of Topic Evolution Path Based on</p>
      <p>
        Semantic Similarity Among Entities
It has following four steps for semantic similarity among entities.
Firstly, a subset about training and testing set is made where the
entities are manually labelled. Secondly, a BiLSTM-CRF
        <xref ref-type="bibr" rid="ref5">(Lample,
G.2016)</xref>
        model is trained on this dataset and evaluated through
quantitative indicators. After 100 training iterations, the accuracy
of the model exceeded 90% and became stable. Similarly, the loss
dropped to below 5.8 and stabilized. Thirdly, this model is used to
detect entities on all patens of each patent. Finally, K-Means is
applied for clustering topics and get entities of each topic. Patent
documents contain a lot of long professional vocabulary.
Compared with commonly used LDA, the identified patent entity
will not lose professional information. Fourthly, a word
representation learning method is applied and the semantic
similarity among entities of each topic could be calculated.
Based on semantic similarity among entities, a similarity
threshold is determined, in which two different entities could be
treated as the same meaning if the similarity is higher than
threshold. We define five topic evolution patterns, including
development, division, integration, extinction, and emerging.
They are shown in Table 1, in which the size of the circle
represents the number of entities under each topic. Moreover,
T1(t1) and T2(t2) means the topic T1 and T2 at time t1 and t2
respectively.
      </p>
      <p>T2(t2) is treated as an emerging topic which
has no evolutionary relationship with
previous topic.
2.3 Discovery of Semantic Relationship Among</p>
      <p>
        Topics
Firstly, we predefine five types of semantic relationships among
patent entities, which are shown in Table 2. Secondly, we
manually label a small dataset for training and testing with
predefined relationships. Thirdly, we train a OpenNRE
        <xref ref-type="bibr" rid="ref6">(Han, X.,
2019)</xref>
        model on this dataset and evaluate it through quantitative
indicators. Fourthly, the model is used to predict all relationships
among entities. Finally, the semantic relationship between two
topics is determined based on the semantic relationship among all
pairs of entities.
3
      </p>
    </sec>
    <sec id="sec-4">
      <title>Result</title>
      <p>We set semantic similarity to 0.7 where a pair of entities with
similarity higher than it is considered to have the same meaning.
In the discovery of topic evolution path, the results are obtained
from 2015-2016 as an example. There are six topics in 2015 and
eight topics in 2016, where the evolution probability is shown in
Table 3.</p>
    </sec>
    <sec id="sec-5">
      <title>Conclusion</title>
      <p>This paper proposes a method for discovery of Topic Evolution
Path and Semantic Relationship among topics Based on Patent
Entity Relationship. The result could prove the effectiveness of
this method and could enrich and improve the topic evolution
analysis method. In the next step, we would like to apply other
neural network models that may do better in patent entity
relationship extraction and compared with baseline method for
deep analysis.</p>
    </sec>
    <sec id="sec-6">
      <title>ACKNOWLEDGMENTS</title>
      <p>This work is supported by the National Natural Science
Foundation of China (No. 71974095).</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          <string-name>
            <surname>Liu</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chen</surname>
            ,
            <given-names>Z.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Tang</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          et al.
          <article-title>Mapping the technology evolution path: a novel model for dynamic topic detection and tracking</article-title>
          .
          <source>Scientometrics 125</source>
          ,
          <fpage>2043</fpage>
          -
          <lpage>2090</lpage>
          (
          <year>2020</year>
          ). https://doi.org/10.1007/s11192-020-03700-5
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          <string-name>
            <surname>Yu</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Xu</surname>
            ,
            <given-names>Z.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>Wang</surname>
            ,
            <given-names>X.</given-names>
          </string-name>
          (
          <year>2020</year>
          ).
          <article-title>Bibliometric analysis of support vector machines research trend: a case study in China</article-title>
          .
          <source>International Journal of Machine Learning and Cybernetics</source>
          ,
          <volume>11</volume>
          (
          <issue>3</issue>
          ),
          <fpage>715</fpage>
          -
          <lpage>728</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          <string-name>
            <surname>No</surname>
            ,
            <given-names>H. J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>An</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>Park</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          (
          <year>2015</year>
          ).
          <article-title>A structured approach to explore knowledge flows through technology-based business methods by integrating patent citation analysis and text mining</article-title>
          .
          <source>Technological Forecasting and Social Change</source>
          ,
          <volume>97</volume>
          ,
          <fpage>181</fpage>
          -
          <lpage>192</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          <string-name>
            <surname>Birunda</surname>
            ,
            <given-names>S. S.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>Devi</surname>
            ,
            <given-names>R. K.</given-names>
          </string-name>
          (
          <year>2021</year>
          ).
          <article-title>A Review on Word Embedding Techniques for Text Classification</article-title>
          .
          <source>In Innovative Data Communication Technologies and Application</source>
          (pp.
          <fpage>267</fpage>
          -
          <lpage>281</lpage>
          ). Springer, Singapore.
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          <string-name>
            <surname>Lample</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ballesteros</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Subramanian</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kawakami</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>Dyer</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          (
          <year>2016</year>
          ).
          <article-title>Neural architectures for named entity recognition</article-title>
          .
          <source>arXiv preprint arXiv:1603</source>
          .
          <fpage>01360</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          <string-name>
            <surname>Han</surname>
            ,
            <given-names>X.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gao</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Yao</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ye</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Liu</surname>
            ,
            <given-names>Z.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>Sun</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          (
          <year>2019</year>
          ).
          <article-title>OpenNRE: An open and extensible toolkit for neural relation extraction</article-title>
          . arXiv preprint arXiv:
          <year>1909</year>
          .
          <fpage>13078</fpage>
          ..
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>