<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Exploring Keyphrase Extraction and IPC Classification Vectors for Prior Art Search</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Manisha Verma</string-name>
          <email>manisha.verma@research.iiit.ac.in</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Vasudeva Varma</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Search And Information Extraction Lab, Language Technologies Research Centre, International Institute of Information Technology</institution>
          ,
          <addr-line>Hyderabad</addr-line>
          ,
          <country country="IN">India</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>In this paper we describe experiments conducted for CLEFIP 2011 Prior Art Retrieval track. We examined the impact of 1) using key phrase extraction to generate queries from input patent and 2) the use of citation network and (International Patent Classification) IPC class vector in ranking patents. Variations of a popular key phrase extraction technique were explored for extracting and scoring terms of query patent. These terms are used as queries to retrieve similar patents. In the second approach, we use a two stage retrieval model to find similar patents. Each patent is represented as an IPC class vector. Citation network of patents is used to propagate these vectors from a node (patent) to its neighbors (cited patents). Similar patents are found by comparing query vector with vectors of patents in the corpus. Text based search is used to re-rank this solution set to improve precision. Two-stage system is used to retrieve and rank patents. Finally, we also extract and add citations present within the text of a query patent to the result set. Adding these citations (present in query patent text) to the results shows significant improvement in Mean Average Precision (MAP).</p>
      </abstract>
      <kwd-group>
        <kwd>Prior Art Retrieval</kwd>
        <kwd>Patent Retrieval</kwd>
        <kwd>Key phrase extraction</kwd>
        <kwd>CLEF-IP track</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>We participated in the CLEF-IP 2011 Prior Art Retrieval track to evaluate the
performance of existing approaches and a new representation for patents on a
large collection of documents and queries. Our goal was to use and evaluate key
phrase extraction for constructing queries from input patents, impact of IPC
class information and citation network of patents in the corpus on recall and
contribution of citation mining in enriching initial set of search results.</p>
      <p>
        IPC class information can be useful in filtering or re-ordering the search
results. In CLEF-IP task, BiTeM group [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ] have used IPC codes to filter patents
which do not share at least one IPC code with the query patent. Key phrase
extraction has been previously used to construct queries from input patents. In [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]
candidate n-grams are selected using a classifier. The authors manually
annotate potential keywords to train the classifier. Extraction of citation
information present in the the patent text has also improved (Mean Average Precision)
MAP in previous CLEF-IP tracks [
        <xref ref-type="bibr" rid="ref2 ref6">2, 6</xref>
        ]. In CLEF 2011 we try and evaluate three
ways to improve prior art retrieval. Firstly, we evaluate variations of a popular
key phrase extraction technique (TextRank) for extracting and scoring terms of
query patent. These terms are used as queries to retrieve similar patents.
Secondly, we use a novel representation of patents and a two-stage retrieval approach
to improve both precision and recall. Finally, we add citations extracted from
the patent text to the search results, as it has improved MAP scores previously.
      </p>
      <p>In Section 2, we explain briefly the approaches used for retrieving and
reranking patents. The experiments, result and analysis are explained in Section 3
and Section 4 respectively. Conclusion and future work are discussed in Section 5.
2
2.1</p>
    </sec>
    <sec id="sec-2">
      <title>Our Approach</title>
      <sec id="sec-2-1">
        <title>Key Phrase Extraction</title>
        <p>
          Reducing the input patent text to a query which can fetch related prior art
was our first objective. In [
          <xref ref-type="bibr" rid="ref3 ref5">5, 3</xref>
          ] we explored several supervised and unsupervised
key phrase extraction approaches to extract candidate terms to form a query.
Since the training set of queries was small, we decided to use only unsupervised
methods of term extraction from patents. In unsupervised approaches, TextRank
outperformed tf-idf in selection of terms from a patent. TextRank uses the
information around a word to calculate its importance whereas tf or tf − idf scores
do not reflect this information. Hence, we use TextRank to extract terms from
a query patent and use following to assign weight to top 20 terms in the query.
        </p>
        <p>TextRank: Since patents contain significant amount of text, co-occurrence
information present in it can be safely used to determine key terms for a patent.
Weight of each term wi in the query is the score given by TextRank algorithm.</p>
        <p>TextRank*idf: TextRank extracts words which are central/important for a
patent. However, patents either use general or totally new terms to explain new
concepts. TextRank extracts both these types of terms with efficiency (with the
help of co-occurrence information) but a word’s score does not reflect its rarity.
Hence to capture rarer terms central to the query patent we use modified version
of TextRank score to weigh each word in the query. Note that words are still
selected on the basis of their TextRank score, only their weight in the query is
determined by the following:
Weight of each term wi in the query is the score T extRank(wi) ∗ idf (wi) where
T extRank(wi) is the TextRank score of wi,
idf (wi) =</p>
        <p>N
log( 1.0+df(wi) )
log(2)
(1)
N is the number of documents in the collection and df (wi) is the number of
documents containing wi.
2.2</p>
        <p>
          IPC-Vector based Retrieval
Patents contain meta-data other than text which can be leveraged to improve
retrieval accuracy. A patent has manually assigned classification code, defining
broad area of the invention. It also cites other patents to discuss similar
inventions in the past. Our approach is to combine both the classification and citation
information to represent a patent. Each patent is manually assigned one or more
International Patent Classification (IPC) codes. We use this information to
represent each patent as an IPC class vector. Citation network of patents is used
to propagate these vectors from a node (patent) to its neighbors (cited patents).
Thus, each patent vector is a weighted combination of its neighbor IPC
information and its own. Vector formation and propagation are explained in [
          <xref ref-type="bibr" rid="ref4">4</xref>
          ]. Two
stages of the system are :
1. Stage 1 : Converting the query and corpus patents into vectors using IPC
codes and citation network. For the runs submitted in CLEF-IP, we use only
cosine similarity to retrieve similar vectors.
2. Stage 2 : Re-Ranking top K documents using text of the query patent. We
use tf-idf, TextRank and T extRank ∗ idf score to select and weigh top 20
words in the query.
2.3
        </p>
        <p>Citation extraction from queries
Some query patents contain cited patent numbers within the text of their
description. These patent numbers were not filtered out of the text of the query
patent, which can be added to the search results. Adding this information to the
experimental results is reported to demonstrate the impact of using this kind of
information.</p>
        <p>For the large topic collection containing 3973 query patents, citations were
extracted from 1419 patent topics and found to be IDs of patents in the indexed
collection. Other extracted citations that do not exist in the collection were
discarded.
3
3.1</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>Evaluation</title>
      <sec id="sec-3-1">
        <title>Data</title>
        <p>The data collections are extracts of the MAREC1 dataset, containing over 2.6
million patent documents pertaining to 1.3 million patents from the European</p>
        <sec id="sec-3-1-1">
          <title>1 http://www.ir-facility.org/prototypes/marec</title>
          <p>Patent Office with content in English, German and French, and extended by
documents from the WIPO. The queries have been translated to English from
German and French with the help of Google Translator2. Only English translations
of original patents are used for making queries. The data has been indexed using
Lemur3 toolkit. All the fields of a patent (title, abstract, description, claims,
citations and IPC class information) have been indexed. Of 1.3 million patents,
0.8 million patents cite at least one patent in the corpus and 0.64 million patents
are cited by at least one patent. Dimension of concatenated IPC class vector for
this dataset is 79963, of which level 1 has 875, level 2 has 8631 and level 3 has
70457 dimensions respectively. 3973 query patents were provided including 1351
English and 2622 German and French patents.
3.2</p>
          <p>Evaluation Method
In the CLEF-IP Workshop, we use the mean average precision (MAP), Recall at
100 (R@100), R@200 and R@1000 as evaluation measures. For CLEF-IP Prior
Art Search task we compare the following methods:</p>
          <p>Base: Simple Text Retrieval, 20 words, from the query patent, with high
tf-idf values are used to form a weighted query. The weight of each word is its
tf-idf score.</p>
          <p>TextRank: 20 words, from the query patent, with high TextRank values
are used to form a weighted query. The weight of each word is its TextRank
score.</p>
          <p>TextRank∗idf: 20 words, from the query patent, with high TextRank
values are used to form a weighted query. The weight of each word wi in the query
is T extRank(wi) ∗ idf (wi).</p>
          <p>Since limited number of runs could be submitted for the task, it was found
on the training data that T extRank ∗ idf performed the best. Hence, we did not
submit the results of Base and TextRank.</p>
          <p>COS: Cosine similarity (COS) has been used to measure similarity between
a patent and query IPC vectors. The process for generating IPC vectors for
patents in the corpus is explained in 2.2.</p>
          <p>COS, tf-idf: IPC information present in the patent is used to make the
vector. Cosine is used to calculate similarity between a patent and query. For
a query patent top 1000 similar patents are retrieved. These patents are
reranked using query generated by TextRank method. It does not contain citations
extracted from the patents.</p>
          <p>COS, TR: IPC information present in the patent is used to make the vector.
Cosine Similarity is used to calculate similarity between a patent and query. For
a query patent top 1000 similar patents are retrieved and re-ranked using queries
generated by TextRank method mentioned above.</p>
        </sec>
        <sec id="sec-3-1-2">
          <title>2 http://translate.google.com</title>
          <p>3 http://www.lemurproject.org/</p>
          <p>For the runs (COS,tf-idf), (COS,TR) and (TextRank∗idf) we also add the
extracted citations from query patents .
4</p>
        </sec>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>Results And Discussion</title>
      <p>
        The results for the submitted runs are shown in Table 1. In the runs without
citations, methods using vector representation of patents perform well in terms of
Recall. Importance of re-ranking documents is evident from COS method (using
only vector representation to find similar patents) results, as the MAP is still
low. However, re-ranking the top 1000 documents does not result in significant
change in MAP value either. This is primarily due to the approach used for
reranking top documents. After the submission of the results, it was found linear
combination of the COS and Text (Re-rank using queries generated by tf − idf ,
TextRank etc) score resulted in higher MAP values [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ]. The high recall is due
to the representation of a patent as vectors and propagation of vectors in the
citation graph.
      </p>
      <p>The TextRank∗idf method has low recall as compared to vector based
methods. This is primarily due to limited coverage of queries. The queries created by
using only the patent text cannot be used for retrieving documents which share
meta information such as IPC Class information and citations. Such queries may
not ensure very high recall while still managing high precision.</p>
      <p>Citation addition to the initial set of results improves performance of
TextRank∗idf significantly but pushes down the MAP for COS methods. The
lowering of MAP may be due to the relevance judgments given for the queries.
The relevance judgements contain more than one level of relevance. It may be
the case that documents with higher priority in relevance judgments which were
ranked higher by COS methods had been pushed to lower ranks due to citation
addition which inturn resulted in low MAP values.
TextRank*idf 0.055 0.200 0.26</p>
      <p>COS 0.049 0.255 0.354
COS,tf-idf 0.061 0.288 0.385</p>
      <p>COS,TR 0.057 0.281 0.378
TextRank*idf + Citation 0.097 0.244 0.297</p>
      <p>COS,tf-idf + Citation 0.055 0.269 0.362
COS,TR + Citation 0.052 0.262 0.355
0.399
0.600
0.601
0.595
0.423
0.599
0.594
5</p>
    </sec>
    <sec id="sec-5">
      <title>Conclusion and Future Work</title>
      <p>In CLEF-IP 2011 we experimented with some approaches for query formation
from an input patent. We also explored a two-stage approach to find related
prior art. First, a vector based representation which uses IPC information of
patent and its neighbors to retrieve similar patents. This representation proves
effective in increasing the recall. Then, re-ranking top 1000 documents in second
stage is used to improve precision. We also used extracted citations from the
query patent text to improve the results. Vector based representation proved
to be effective in increasing recall, however improvement in precision was not
achieved with simple re-ranking of documents. Approaches like TextRank which
use co-occurence information in the text to find out key terms were better than
frequency based measures of selecting words from the text. Citation extraction
and addition certainly proved instrumental in increasing mean average precision.
However, ways of citation addition to the result set needs further investigation.
An extension to this work for future participation would be to use a
learningto-rank approach to re-rank top documents. It would be interesting to observe
effects of combining both vector representation with patent text to avoid
reranking.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <given-names>J.</given-names>
            <surname>Gobeill</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Pasche</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Teodoro</surname>
          </string-name>
          , and
          <string-name>
            <given-names>P.</given-names>
            <surname>Ruch</surname>
          </string-name>
          .
          <article-title>Simple pre and post processing strategies for patent searching in clef intellectual property track 2009</article-title>
          .
          <article-title>In Proceedings of the 10th cross-language evaluation forum conference on Multilingual information access evaluation: text retrieval experiments</article-title>
          ,
          <source>CLEF'09</source>
          , pages
          <fpage>444</fpage>
          -
          <lpage>451</lpage>
          , Berlin, Heidelberg,
          <year>2009</year>
          . Springer-Verlag.
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <given-names>P.</given-names>
            <surname>Lopez</surname>
          </string-name>
          and
          <string-name>
            <given-names>L.</given-names>
            <surname>Romary</surname>
          </string-name>
          .
          <article-title>Experiments with citation mining and key-term extraction for Prior Art Search</article-title>
          .
          <source>In CLEF 2010 - Conference on Multilingual and Multimodal Information Access Evaluation</source>
          , Padua Italie,
          <year>2010</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <given-names>M.</given-names>
            <surname>Verma</surname>
          </string-name>
          and
          <string-name>
            <given-names>V.</given-names>
            <surname>Varma</surname>
          </string-name>
          .
          <article-title>Applying key phrase extraction to aid invalidity search</article-title>
          .
          <source>In Proceedings of the 13th International Conference on Artificial Intelligence and Law</source>
          ,
          <source>ICAIL '11</source>
          , pages
          <fpage>249</fpage>
          -
          <lpage>255</lpage>
          , New York, NY, USA,
          <year>2011</year>
          . ACM.
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <given-names>M.</given-names>
            <surname>Verma</surname>
          </string-name>
          and
          <string-name>
            <given-names>V.</given-names>
            <surname>Varma</surname>
          </string-name>
          .
          <article-title>Patent search using IPC Class vectors</article-title>
          . To Appear
          <source>In Proceedings of the 4th international workshop on Patent information retrieval</source>
          ,
          <source>PaIR '11.</source>
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <given-names>R.</given-names>
            <surname>Mihalcea</surname>
          </string-name>
          and
          <string-name>
            <given-names>P.</given-names>
            <surname>Tarau</surname>
          </string-name>
          . TextRank:
          <article-title>Bringing order into texts</article-title>
          .
          <source>In Proc. of EMNLP</source>
          ,
          <year>2004</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <given-names>W.</given-names>
            <surname>Magdy</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Leveling</surname>
          </string-name>
          , and
          <string-name>
            <given-names>G. J. F.</given-names>
            <surname>Jones</surname>
          </string-name>
          . DCU @
          <article-title>CLEF-IP 2009: Exploring standard IR techniques on patent retrieval</article-title>
          .
          <source>In 10th Workshop of the Cross-Language Evaluation Forum</source>
          ,
          <string-name>
            <surname>CLEF</surname>
          </string-name>
          <year>2009</year>
          , Corfu, Greece,
          <source>September 30-October 2, Revised Selected Papers, Lecture Notes in Computer Science (LNCS)</source>
          . Springer,
          <year>2010</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>