<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Use of Pseudo Relevance Feedback for Patent Clustering with Fuzzy C-means</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Noushin Fadaei</string-name>
          <email>fadaei@uni-hildesheim.de</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Thomas Mandl</string-name>
          <email>mandl@uni-hildesheim.de</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Hildesheim University</institution>
          ,
          <addr-line>Hildesheim</addr-line>
          ,
          <country country="DE">Germany</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2005</year>
      </pub-date>
      <abstract>
        <p>Patent databases are meaningful resources for technology trend detection as they collect information on the recent key innovations; however the importance of wordings in patents and use of complex content are remarkable challenges in key word extraction in the text mining phase. Moreover patents share information by nature and depending on the criterion of classi cation such as materials or uses, one may belong to multiple classes. For clustering patents, this work proposes an updating fuzzy c-means clustering which employs pseudo relevance feedback originating from information retrieval in order to improve features extracted from the patent collection following feedbacks. The results show a noticeable improvement in clustering after applying pseudo relevance feedback clustering patents under topic contact lenses.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>The increasing competition in research conducted at
universities and other research institutes as well as in
industry, further intensi ed by the increasing
globalization, reinforces the importance of identifying new
trends at an early stage. According to a study by
Thomson Reuters 70% to 90% of the information
covered in patents depending on the research area is
not published anywhere else [Cor07]. This quality of
patents makes such databases vital resources for
nding new trends in the industry. Yet patent databases
Copyright © CIKM 2018 for the individual papers by the papers'
authors. Copyright © CIKM 2018 for the volume as a collection
by its editors. This volume and its papers are published under
the Creative Commons License Attribution 4.0 International (CC
BY 4.0).
are growing very fast and required to be more
organized for such purposes. According to European
Patent O ce (EPO), there is a steady increase in this
huge information resource in terms of led patents
since 2010b [O 18]. In 2016, EPO recorded the
highest number of granted patents which went up to 10.1%
in 2017 [O 18]. To manage the detection of subtopics
and their potentiality of being a trend in such big data,
further grouping of patents is inevitable.</p>
      <p>Patent analysis approaches are either qualitative or
quantitative [Hon09]. The focus of this work lies on
qualitative patent analysis, i.e., it uses text-mining
techniques regarding the content of patents unlike
quantitative approaches which employ metadata. The
process typically aims to transform documents to
vectors using weighting systems (e.g., TFIDF) for key
terms; Clustering methods then exploit similarity
measures to examine related vectors (e.g., Euclidean
distance) and group them into clusters. The e ciency
of clustering is highly dependent on the selected key
terms [TJC04]. While the key words extraction
requires the contribution of experts in various domains,
automatic methods restrict the number of selected
terms by applying thresholds for term frequency in
documents (TF) or document frequency (DF) [TLL07]
or consider building key phrases for example by
creating a co-relation matrix of high frequent terms in
documents where the co-occurrence of terms in the dataset
is reserved [THTL06]. This work adopts pseudo
relevance feedback [MRS08b] to obtain the key terms that
tend to form the core concept of clusters.</p>
      <p>One of the most commonly used clustering
algorithms in text-mining is k-means; however k-means
is not the best clustering method for patent
analysis. Patents are very rich documents in terms of
professional information and they may cover a range of
technologies, applications or use of various materials.
Therefore exhaustive clustering con nes patents that
potentially belong to di erent classes of a certain
criterion such as technology or material [FMS+15].
Moreover k-means requires the number of clusters and
cannot cope with outliers [CHPT05].</p>
      <p>To avoid low accuracy, fuzzy non-exhaustive
clustering methods are exploited for patent clustering, e.g.,
by [THTL06] and [DD09]. The fuzziness of the
clustering method provides a likelihood of possession (or
membership to cluster) instead of a rigid distinction
and the overlapping characteristic of clustering allows
one patent re ecting a number of claims contributes
in multiple clusters. Setting a threshold for
membership controls to what extent of similarity patents may
show up in a cluster. Fuzzy c-means (FCM) [Bez81] is
the most popular fuzzy clustering algorithms and here
we use it to softly partition our patent collection on
certain topics. Also, the membership matrix enables
FCM to cope with the issue of the outliers as all
membership values of one document to all clusters ought to
add up to 1; thus, an unrelated document to all
clusters receives insigni cant membership values for each
and every cluster and can be ignored.</p>
      <p>The ultimate goal of this study is to investigate the
main idea of partitioning a patent dataset by means of
a fuzzy clustering which allows for updating through
relevance feedback. Practically this study will run
with help of patent expert users however for
experimental purposes and analysing the validity of the
approach, we organize it with FCM and pseudo
relevance feedback. For the evaluation, we developed a
benchmark based on World Intellectual Property
Organization's patent reports on recent technology trends
which were edited by experts in the domains. The
reports describe technologies and trends within a
domain, e.g. contact lenses or robotic arms. These
reports are published at the website of the World
Intellecutal Property Organization (WIPO)1. The selected
reports are provided by Gridlogics Technologies Pvt.
Ltd and were partially generated by the use of Patent
iNSIGHT Pro. The developed benchmark can be used
for experiments in classi cation, clustering, trend
analysis or other intelligent patent processing systems.
2</p>
    </sec>
    <sec id="sec-2">
      <title>Problem Statement</title>
      <p>Given the query of the expert user which determines
the scope of the topic or more generally given the
International Patent Classi cation (IPC) 2 of the topic of
interest, we retrieve the patents of the domain from the
patent dataset. The set of n patents is then supposed
1http://www.wipo.int/patentscope/en/programs/patent
landscapes/plrdb.html</p>
      <p>2According to World Intellectual Property
Organization(WIPO), The International Patent Classi cation (IPC),
established by the Strasbourg Agreement 1971, provides for a
hierarchical system of language independent symbols for the
classi cation of patents.
to be represented as vectors (X = (x1; x2; : : : ; xn)2
c n). Each vector xk 2 Rs is built up by s features
where the features are the selected key terms from the
dataset. The goal is to organize X into groups that
may share vectors for further patent analysis purposes
like identifying trends in the given topic [HXW+12].
2.1</p>
      <sec id="sec-2-1">
        <title>Citations In Text</title>
        <p>Citations within the text should indicate the author's
last name and year[?]. Reference style[?] should follow
the style that you are used to using, as long as the
citation style is consistent.</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>Our Approach 3</title>
      <p>3.1</p>
      <sec id="sec-3-1">
        <title>Fuzzy C-Means Algorithm (FCM) [Bez81]</title>
        <p>Fuzzy c-means algorithm (FCM) [Bez81] is based on
the fuzzy membership matrix. The membership
describes the likelihood of each vector (document) xk
being a part of cluster (subtopic) ci where 1 k n
and 1 i c with c being the number of clusters. The
overall membership of each vector is normalized to 1.
The algorithm starts o with initializing c vectors as
centroids of the clusters. Then the membership (wik)
is calculated through the Euclidean distances (dik) of
each vector (xk) to the centroid (Pi) of the cluster it
belongs to and to the centroids of other clusters.</p>
        <p>Pi =</p>
        <p>Pn</p>
        <p>k=1 (wik)mxk
Pn</p>
        <p>k=1 (wik)m
c
Wi(kb) = X
1</p>
        <p>2
j=1 [( ((ddjikk))((bb)) ) m 1 ]
(1)
(2)
(3)</p>
        <p>The FCM membership function is calculated as
[HXW+12]:</p>
        <p>c
(i;j) = [X ( jjxj
t=1 jjxj
vijjA ) m2 1 ]
vtjjA
1
(i;j) represents the membership value of jth patent
of the dataset and ith topic whose centroid is vi. jjjjA
stands for norm function.</p>
        <p>The memberships updates centroids vectors until
the overall distance of the updated centroids is less
than " compared to the last set of centroids. The
steps of FCM algorithm are as follows:
1. Set the number of clusters to be found (c)
2. Set an Euclidean normalization and fuzziness (m)
3. Initialize of cluster prototype P 0, set the iterative
counter (b)
4. Obtain membership matrix using Equation 2
above
5. Update the centroids using Equation 1 above
6. Repeat starting from step 2 until jjP (b)</p>
        <p>P (b 1)jj &lt;</p>
        <p>The algorithm description and formulations are
inspired by [NNB15].</p>
        <p>Relevance feedback is a common concept in
information retrieval that involves the user knowledge to
identify and fetch more similar results to the query.
Pseudo relevance feedback resembles the same
concept but the entire process is carried out
automatically. The goal is to identify the relevant key terms
regarding the query and to expand them in order to
update the query and make it in line with more
relevant documents. The procedure starts with retrieving
relevant documents and sorting them based on their
similarity score to the query. Then the top k
documents are considered as the feedback for best results.
Pseudo relevance feedback considers these documents
the source of relevant key terms to the query.</p>
        <p>We make use of this method in FCM so that we
nd the patents that are more similar to the core
concept of the cluster or centroids. The hypothesis is the
key terms provided by these patents re ect the main
idea of the cluster and their corresponding vectors can
represent a more manageable dataset for FCM. The
procedure is shown in gure 1. It generates an
overlapping clustering based on the membership matrix of
FCM. As long as a relevance feedback is requested the
following procedure runs:
1. Rank patents based on their membership to the
cluster
2. Obtain the top relevant patents for building the
features for vectors (k% of the size of the
corresponding cluster)
3. Update X = (x1; x2; : : : ; xn)2
vectors) with new feature
c
n (dataset of
4. Pass X to FCM so it starts updating the
membership function
5. Stop if after clustering result of FCM no relevance
feedback (RF) is requested
4
4.1</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>Experiment</title>
      <sec id="sec-4-1">
        <title>Datasets and Experiments Settings</title>
        <p>We used freely available queries that are provided
by the World Intellectual Property Organization
(WIPO) to form a gold standard. The result sets
were gained through the European Patent Fulltext
(EPFULL) repository. For this work, we used a set,
namely contact lenses to be clustered and another
set, robotic arms to show an insight of the built gold
standard. The code was implemented in Python using
Natural Language Toolkit (NLTK) for tokenization
and stemming but for removing the stop words,
we used our list of stop words for patent analysis.
For Fuzzy c-means clustering we adopted the code
provided in Github 3. The overlapping is controlled
by allowing the patents that are at least 99% similar
to a member of the cluster inside the cluster as a
member. The fuzzi er is set on 1.2, the error on 0.001
and the number of iterations on 200. We run the
method for c=3, c=13 and c=23 and for visualization
in 2D, we have used Principal Component analysis
(PCA).</p>
        <p>3https://github.com/holtwashere/PossibilisticCMeans
4.2
Evaluation measures used for granulated clusters (e.
g. k-means clusters) such as Dunn, Mutual
Information (MI), F-measure, Rand Index and Jaccard are not
quite useful for measuring overlapping clusters; for
instance Dunn gives a higher score to the clustering
systems that assign the data points to more distant
clusters while the contents of each cluster are pretty close.
Considering a criterion like uses, we know that one
patent might have several usages and such measures
are not revealing any required information. Purity
exposes the very nature of a cluster: the degree of
consistency. The higher is the purity the less is random
clustering. This is one important characteristic of clusters
in the patent clustering task, nevertheless we cannot
neglect the drawbacks of Purity: it is highly dependent
on the number of clusters. Like Purity, MI is also
inuenced by the number of clusters, while Normalized
Mutual Information (NMI) enables us to compare the
clusters with each other (it ranges between 0 and 1)
and it is not a ected greatly by the inaccurate number
of clusters. For assessing the quality of this clustering,
we have used Normalized Mutual Information (NMI)
by Fred and Jain [FJ03] and Purity [MRS08a].</p>
      </sec>
      <sec id="sec-4-2">
        <title>Gold Standard based on World Intellectual Property Organization (WIPO) Reports</title>
        <p>There are a number of technological reports available
by WIPO under the universal title of public health/life
science. These reports usually cover two types of
queries; one results in the main class (usually shares
the title with the topic that is reported) and one is
breaking the main query into pieces; thus produces
subclasses. Some reports categorize the foregoing
subclasses and provide fewer yet more general subclasses.
The following example depicts search strings provided
by WIPO report which results to a main class, namely
robotic arms followed by a further query that along
with the main query leads into a subclass of robotic
arms:
1. Query to the main class robotic arm :
(FT=(robot* or (arti cial w/2 intelligence) or
android or cyborg or humanoid*)) or (TAC=
(manipulator* or manipulater* or actuator*
or actuater* or drives or joint or joints or
actuation or ("end e ector" or "end e ecter")
or ((pneumatic* or air) w/2 muscle*))) and
((IC= B25J9/02 or B25J9/04 or B25J9/06 or
B25J13/02 or B25J13/08 or B25J17 or B25J18)
or (UC=901/2 or 901/14 or 901/19 or 901/27
or 901/31 or 901/39 or 700/245 or 700/248 or
700/261))</p>
      </sec>
      <sec id="sec-4-3">
        <title>2. Query to its subclass</title>
        <p>Robot :
Anthropomorphic
(TAC) contains (humanoid or android or
anthropomorphic* or anthropomor c*)
Query guide: FT-Full Text, TAC- Title Abstract
Claim, IC- International Class, UC- US Class.
w/2 shows the maximum number of intervening
unmatched positions doesn't exceed 2.</p>
        <p>Depending on the authors of the reports, queries are
described in di erent languages and formats.
Therefore the queries had to be adjusted to match Json
data type. Using Elasticsearch's API, we collected
data through the EPFULL database and obtained
the main class and subsequently the corresponding
subclasses. For instance under the topic robotic arms,
519 patents were retrieved, out of which 511 patents
were covered by subclasses from which 293 hits belong
to the subclasses of the types criterion, 474 hits t
in the subclasses of applications criterion and the
subclasses of parts criterion cover 701 hits. Patents
may share di erent subclasses and some available
patents in EPFULL may not be covered by any
subclass. The results of the proposed clustering method
have been examined against these retrieved subclasses.
5</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>Results</title>
      <p>The results of clustering under the topic contact lenses
shows using pseudo relevance feedback can de nitely
help patents picking the better cluster. According to
table 1, while the Purity results remain more or less
the same the normalized mutual information (NMI)
improved 4% to 11% for all three number of clusters.
Purity is signi cantly dependent on the number of
clusters as when the data themes are noticeably fewer
than the number of clusters the chance of having more
consistent clusters would raise. However The results of
Purity under the topic Contact Lenses do not change
diversely for di erent number of clusters and it remain
relatively high. This re ects the acceptable ability of
FCM in this clustering. Moreover, visualizing
clustering at 13 clusters with FCM ( gure 2) and FCM with
in uence of RF ( gure 3) using PCA, we observe the
data points are less scattered after using RF.
6</p>
    </sec>
    <sec id="sec-6">
      <title>Future work</title>
      <p>The main purpose of this study is to test whether the
use of relevance feedback in clustering can play an
important role in grouping patents. In future we would
use explicit relevance feedback of the expert users to
modify the dimension of the vectors with more useful
features. We would also make use of lexical resources
on top of the feedbacks. In future we would also survey
the position of feedback vectors as they are selected
by users and may not be necessary appear around the
centroid.
[Bez81]
[Cor07]
[DD09]
[FJ03]</p>
      <sec id="sec-6-1">
        <title>The Thomson Corporation. Global patent sources: An overview of international patents. 2007.</title>
      </sec>
      <sec id="sec-6-2">
        <title>Turkay Dereli and Alptekin Durmusoglu. Classifying technology patents to identify trends: Applying a fuzzy-based clustering approach in the turkish textile industry.</title>
        <p>Technology in Society, 31(3):263 { 272,
2009.</p>
      </sec>
      <sec id="sec-6-3">
        <title>Ana L N Fred and Anil K. Jain. Ro</title>
        <p>bust data clustering. In Proceedings of
the IEEE Computer Society Conference
on Computer Vision and Pattern
Recognition, volume 2, 2003.
[FMS+15] Noushin Fadaei, Thomas Mandl, Michael
Schwantner, Mustafa Sofean, Julia M.</p>
        <p>Stru , Katrin Werner, and Christa
Womser-Hacker. Patent analysis and
patent clustering for technology trend
mining. In HIER workshop, Hildesheim,</p>
        <p>Germany, July 2015.
[Hon09]</p>
      </sec>
      <sec id="sec-6-4">
        <title>Soonwoo Hong. The magic of patent infor</title>
        <p>mation. World Intellectual Property
Organization (WIPO). Available via
DIA</p>
        <p>LOG., December 2009.
[HXW+12] Ming Huang, Zhixun Xia, Hongbo Wang,
Qinghua Zeng, and Qian Wang. The range
of the value for the fuzzi er of the fuzzy
c-means algorithm. Pattern Recogn. Lett.,
33(16):2280{2284, December 2012.
[MRS08a] Christopher D. Manning, Prabhakar
Raghavan, and Hinrich Schutze.
Introduction to Information Retrieval.
Cambridge University Press, New York, NY,</p>
        <p>USA, 2008.
[MRS08b] Christopher D. Manning, Prabhakar
Raghavan, and Hinrich Schutze.
Introduction to Information Retrieval. Cambridge</p>
        <p>University Press, Cambridge, UK, 2008.
[NNB15]</p>
      </sec>
      <sec id="sec-6-5">
        <title>Janmenjoy Nayak, Bighnaraj Naik, and</title>
        <p>HS Behera. Computational intelligence in
[O 18]</p>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          <source>data mining</source>
          . Springer, New Delhi, India,
          <year>2015</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          <string-name>
            <surname>European</surname>
          </string-name>
          Patent O ce.
          <source>Epo quality report 2017</source>
          .
          <year>2018</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          <string-name>
            <surname>[THTL06] Amy J. C. Trappey</surname>
          </string-name>
          ,
          <string-name>
            <surname>Fu-Chiang</surname>
            <given-names>Hsu</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Charles</surname>
            <given-names>V.</given-names>
          </string-name>
          <string-name>
            <surname>Trappey</surname>
          </string-name>
          , and
          <string-name>
            <surname>Chia-I Lin</surname>
          </string-name>
          .
          <article-title>Development of a patent document classi - cation and search platform using a backpropagation network</article-title>
          .
          <source>Expert Syst. Appl.</source>
          ,
          <volume>31</volume>
          :
          <fpage>755</fpage>
          {
          <fpage>765</fpage>
          ,
          <year>2006</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [TJC04] [TLL07]
          <string-name>
            <given-names>Y. H.</given-names>
            <surname>Tseng</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D. W.</given-names>
            <surname>Juang</surname>
          </string-name>
          , and
          <string-name>
            <surname>S. H.</surname>
          </string-name>
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          <string-name>
            <surname>Chen</surname>
          </string-name>
          .
          <article-title>Global and local term expansion for text retrieval</article-title>
          .
          <source>In Proceedings of the fourth NTCIR workshop on evaluation of information retrieval, automatic text summarization and question answering</source>
          , Tokyo, Japan,
          <year>June 2004</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          <string-name>
            <surname>Yuen-Hsien</surname>
            <given-names>Tseng</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chi-Jen Lin</surname>
          </string-name>
          , and
          <string-name>
            <surname>Yu-I Lin</surname>
          </string-name>
          .
          <article-title>Text mining techniques for patent analysis</article-title>
          .
          <source>Inf. Process. Manage.</source>
          ,
          <volume>43</volume>
          (
          <issue>5</issue>
          ):
          <volume>1216</volume>
          {
          <fpage>1247</fpage>
          ,
          <year>September 2007</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>