<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Automatically Generating Queries for Prior Art Search</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Erik Graf</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Leif Azzopardi</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Keith van Rijsbergen</string-name>
          <email>keithg@dcs.gla.ac.uk</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>General Terms</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="editor">
          <string-name>Patent Retrieval, Prior Art Search, Automatic Query Formulation</string-name>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Measurement</institution>
          ,
          <addr-line>Performance, Experimentation</addr-line>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>University of Glasgow</institution>
        </aff>
      </contrib-group>
      <abstract>
        <p>This report outlines our participation in CLEF-IP's 2009 prior art search task. In the task's initial year our focus lay on the automatic generation of e ective queries. To this aim we conducted a preliminary analysis of the distribution of terms common to topics and their relevant documents, with respect to term frequency and document frequency. Based on the results of this analysis we applied two methods to extract queries. Finally we tested the e ectiveness of the generated queries on two state of the art retrieval models.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Introduction</title>
      <p>The formulation of queries forms a crucial step in the work ow of many patent related retrieval
tasks. This is speci cally true within the process of Prior Art search, which forms the main task
of the CLEF-IP 09 track. Performed both, by applicants and the examiners at patent o ces, it
is one of the most common search types in the patent domain, and a fundamental element of the
patent system. The goal of such a search lies in determining the patentability (See Section B IV
1/1.1 in [3] for a more detailed coverage of this criterion in the European patent system) of an
application by uncovering relevant material published prior to the ling date of the application.
Such material may then be used to limit the scope of patentability or completely deny the inherent
claim of novelty of an invention. As a consequence of the judicial and economic consequences linked
to the obtained results, and the complex technical nature of the content, the formulation of prior
art search queries requires extensive e ort. The state of the art approach consists of laborious
manual construction of queries, and commonly requires several days of work dedicated to the
manual identi cation of e ective keywords. The great amount of manual e ort, in conjunction
with the importance of such a search, forms a strong motivation for the exploration of techniques
aimed at the automatic extraction of viable query terms. Throughout the remainder of these
working notes we will provide details of the approach we have taken to address this challenge. In
the subsequent section we will provide an overview of prior research related to the task of prior
art search. Section 3 covers the details of our experimental setup. In section 4 we report on the
o cial results and perform an analysis. Finally in the last section we provide a conclusion and
future outlook.
2</p>
    </sec>
    <sec id="sec-2">
      <title>Prior Research</title>
      <p>The remainder of this section aims at providing an overview of prior research concerning retrieval
tasks related to the CLEF-IP 09 task. As the majority of relevant retrieval research in the patent
domain has been pioneered by the NTCIR series of evaluation workshops [1], additionally a brief
overview of relevant collections and the associated tasks is provided. Further we review a variety
of successful techniques applied by participating groups of relevant NTCIR tasks.</p>
      <p>
        First introduced in the third NTCIR workshop [
        <xref ref-type="bibr" rid="ref6">9</xref>
        ], the patent task has led to the release of
several patent test collections. Details of these collections are provided in Table 1. From the
listing in Table 1 we can see that the utilized collections are comparative in size to the CLEF-IP
09 collection, and that the main di erences consist of a more limited time period and a much
smaller amount of topics speci cally for the earlier collections.
      </p>
      <sec id="sec-2-1">
        <title>Workshop</title>
        <p>NTCIR-3
NTCIR-4
NTCIR-5
NTCIR-6</p>
      </sec>
      <sec id="sec-2-2">
        <title>Document Type</title>
        <p>Patent JPO(J)</p>
        <p>Abstracts(E/J)
Patent JPO(J), Abstracts(E)
Patent JPO(J), Abstracts(E)</p>
        <p>Patent USPTO(E)</p>
        <p>Time Period
1998-1999
1995-1999
1993-1997
1993-2002
1993-2002
# of Docs.</p>
        <p>
          697,262
ca. 1,7 million
1,700,000
3,496,252
1,315,470
# of Topics
31
31
103
1223
3221
Based on these collections the NTCIR patent track has covered a variety of di erent tasks,
ranging from cross-language and cross-genre retrieval (NTCIR 3 [
          <xref ref-type="bibr" rid="ref6">9</xref>
          ]) to patent classi cation
(NTCIR 5 [
          <xref ref-type="bibr" rid="ref4">7</xref>
          ] and 6 [
          <xref ref-type="bibr" rid="ref5">8</xref>
          ]). A task related to the Prior Art search task is presented by the invalidity
search run at NTCIR 4 [
          <xref ref-type="bibr" rid="ref1">4</xref>
          ],5 [
          <xref ref-type="bibr" rid="ref2">5</xref>
          ], and 6 [
          <xref ref-type="bibr" rid="ref3">6</xref>
          ]). Invalidity searches are exercised in order to render
speci c claims of a patent, or the complete patent itself, invalid by identifying relevant prior art
published before the ling date of the patent in question. As such, this kind of search, that can be
utilized as a means of defense upon being charged with infringement, is related to prior art search.
Likewise the starting point of the task is given by a patent document, and a viable corpus may
consist of a collection of patent documents. In course of the NTCIR evaluations, for each search
topic (i.e. a claim), participants were required to submit a list of retrieved patents and passages
associated with the topic. Relevant matter was de ned as patents that can invalidate a topic
claim by themselves (1), or in combination with other patents (2). In light of these similarities,
the following listing provides a brief overview of techniques applied by participating groups of the
invalidity task at NTCIR 4-6:
        </p>
        <p>Claim Structure Based Techniques: Since the underlying topic consisted of the text
of a claim, the analysis of its structure has been one of the commonly applied techniques.
More precisely the di erentiation between premise and invention parts of a claim and the
application of term weighting methods with respect to these parts has been shown to yield
successful results.</p>
      </sec>
      <sec id="sec-2-3">
        <title>Document Section Analysis Based Techniques: Further one of the e ectively applied</title>
        <p>
          assumptions has been, that certain sections of a patent document are more likely to contain
useful query terms. For example it has been shown that from the 'detailed descriptions
corresponding to the input claims, e ective and concrete query terms can be extracted' NTCIR
4 [
          <xref ref-type="bibr" rid="ref1">4</xref>
          ].
        </p>
      </sec>
      <sec id="sec-2-4">
        <title>Merged Passage and Document Scoring Based Techniques: Further grounded on</title>
        <p>the comparatively long length of patent documents, the assumption was formed that the
occurrence of query terms in close vicinity can be interpreted as a stronger indicator of
relevance. Based on this insight, a technique based on merging passage and document scores
has been successfully introduced.</p>
        <p>
          Bibliographical Data Based Techniques: Finally the usage of bibliographical data
associated with a patent document has been applied both for ltering and re-ranking of retrieved
documents. Particularly the usage of the hierarchical structure of the IPC classes and
applicant identities have been shown to be extremely helpful. The NTCIR 5 proceedings [
          <xref ref-type="bibr" rid="ref2">5</xref>
          ] cover
the e ect of applying this technique in great detail and note that, 'by comparing the MAP
values of Same' (where Same denotes the same IPC class) 'and Di in either of Applicant
or IPC, one can see that for each run the MAP for Same is signi cantly greater than the
MAP for Di . This suggests that to evaluate contributions of methods which do not use
applicant and IPC information, the cases of Di need to be further investigated.' [
          <xref ref-type="bibr" rid="ref2">5</xref>
          ]. The
great e ectiveness is illustrated by the fact that for the mandatory runs of NTCIR the best
reported MAP score for 'Same' was 0,3342 MAP whereas the best score for 'Di ' was 0,916
MAP.
        </p>
        <p>As stated before our experiments focused on devising a methodology for the identi cation
of e ective query terms. Therefore in this initial participation, we did not integrate the above
mentioned techniques in our approach. In the following section the experimental setup and details
of the applied query extraction process will be supplied.
3</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>Experimental Setup</title>
      <p>The corpus of the CLEF-IP track consists of 1,9 million patent documents published by the
European Patent O ce (EPO). This corresponds to approximately 1 million individual patents
led between 1985 and 2000. As a consequence of the statutes of the EPO, the documents of
the collection are written in English, French and German. While most of the early published
patent documents are mono-lingual, most documents published after 2000 feature title, claim, and
abstract sections in each of these three languages. The underlying document format is based on
an innovative XML schema 1 developed at Matrixware2.</p>
      <p>Indexing of the collection took place utilizing the Indri3 and Lemur retrieval system4. To this
purpose the collection was wrapped in TREC format. The table below provides details of the
created indices:</p>
      <p>Index Name</p>
      <sec id="sec-3-1">
        <title>Retrieval System</title>
      </sec>
      <sec id="sec-3-2">
        <title>Stemming</title>
        <p>Lem-Stop
Indri-Stop</p>
      </sec>
      <sec id="sec-3-3">
        <title>Lemur Indri none none</title>
      </sec>
      <sec id="sec-3-4">
        <title>Stop-Worded</title>
      </sec>
      <sec id="sec-3-5">
        <title>Stop-worded Stop-worded UTF-8 No</title>
        <p>Yes
1http://www.ir-facility.org/pdf/clef/patent-document.dtd
2http://www.matrixware.com/
3http://www.lemurproject.org/indri/
4http://www.lemurproject.org/lemur/
on the English language was applied to all indices. A minimalistic stop-word list was applied in
order to mitigate potential side e ects. The challenges associated with stop-wording in the patent
domain are described in more detail by Blanchard [2]. No stop-wording for French and German
was performed. The creation of the Indri-Stop index was made necessary in order to allow for
experiments based on the ltering terms by language. Lemur based indices do not support UTF-8
encoding and therefore did not allow for ltering of German or French terms by use of constructed
dictionaries.
3.1</p>
        <p>E ective Query Term Identi cation
As stated before the main aim of our approach lies in the extraction of e ective query terms from
a given patent document. The underlying assumption of our subsequently described method is,
that such terms can be extracted based on an analysis of the distribution of terms common to a
patent application and its referenced prior art.</p>
        <p>The task of query extraction therefore took place in two phases: In the rst phase we contrasted the
distribution of terms shared by source documents and referenced documents with the distribution
of terms shared by randomly chosen patent document pairs. Based on these results the second
phase consisted of the extraction of queries and their evaluation based on the available CLEF-IP
09 training data. In the following subsections both steps are discussed in more detail.
3.1.1</p>
        <sec id="sec-3-5-1">
          <title>Analysing the common term distribution</title>
          <p>The main aim of this phase lies in the identi cation of term related features whose distribution
varies among source-reference pairs and randomly chosen pairs of patent documents. As stated
before the underlying assumption is, that such variations can be utilized in order to extract query
terms whose occurrences are characteristic for relevant document pairs. To this extent we evaluated
the distribution of the following features:
1. The corpus wide term frequency (tf)
2. The corpus wide document frequency (df)</p>
          <p>In order to uncover such variations the following procedure was applied: For a given number
of n source-reference pairs an equal number of randomly chosen document pairs was generated.
Secondly the terms common to document pairs in both groups were identi ed. Finally an analysis
with respect to the above listed features was conducted.</p>
          <p>As a result of this approach gure 1 depicts the number of common terms for source-reference
pairs and randomly chosen pairs with respect to the corpus wide term frequency. In the graph,
the x-axis denotes the collection wide term frequency, while on the y-axis the total number of
occurrences of common terms with respect to this frequency is provided. Evident from the graph
are several high-level distinctive variations: The rst thing that can be observed is that the total
number of shared terms of source-reference pairs is higher than for those of random pairs. Further
the distribution of shared terms in random pairs, shown in blue, resembles a straight line on the
log-log scale. Assuming that the distribution of terms in patent documents follows a Zipf like
distribution this can be interpreted as an expected outcome. In contrast to this, the distribution
of shared terms in source-reference pairs, depicted in red, varies signi cantly. This is most evident
in the low frequency range of approximately 2-10000.</p>
          <p>Given our initial goal of identifying characteristic di erences in the distribution of terms shared
within relevant pairs, this distinctive pattern can be utilized as a starting point of the query
extraction process. Therefore, as will be evident in more detail in the subsequent section, we
based our query extraction process on this observation.
Based on the characteristic variations of the distribution of terms common to source-reference pairs
our query term extraction process uses the document frequency as selection criterion. The applied
process hereby consisted of two main steps. Based on the identi cation of the very low document
frequency range as most characteristic for source-reference pairs, we created sets of queries with
respect to the df of terms (1) for each training topic. These queries were then submitted to a
retrieval model, and their performance was evaluated by use of the available training relevance
assessments (2).</p>
          <p>Following this approach two series of potential queries were created via the introduction of two
thresholds.</p>
        </sec>
        <sec id="sec-3-5-2">
          <title>Document Frequency (df ) Based Percentage Threshold: Based on this threshold,</title>
          <p>queries are generated by including only terms whose df lies below an incrementally increased
limit. To allow for easier interpretation the incrementally increased limit is expressed as
dNf 100, where N denotes the total number of documents in the collection. A percentage
threshold of 0.5% therefore denotes, that we include only terms in the query that appear in
less than 0.5% of the documents in the collection.</p>
          <p>Query Length Threshold: A second set of queries was created by utilization of an
incrementally increased query length as underlying threshold. In this case for a given maximum
query length n, a query was generated by including the n terms with the lowest document
frequency present in the topic document. The introduction of this threshold was triggered
by the observation that the amount of term occurrences with very low df varies signi cantly
for the topic documents. As a consequence of this a low df threshold of 1000 can yield a lot
of query terms for some topics, and in the extreme case no query terms for other topics.</p>
          <p>We generated queries based on a percentage threshold ranging from 0.25%-3% with an
increment of 0.25, and with respect to query lengths ranging from 10-300 with an increment of 10. The
performance of both query sets was then evaluated by utilization of the large training set of the
main task with the BM25 and Cosine retrieval models.
In the following a description of the submitted runs and an analysis of their performance will be
conducted. In total our group submitted ve runs. While the performance of one of our runs is
in line the with the observations based on the training set, the performance of the other four runs
resulted in a completely di erent and order of magnitudes lower results. Unfortunately this was
induced by a bug occurring in the particular retrieval setup utilized for their creation. The amount
of analysis that can be drawn from the o cial results is therefore very limited. After obtaining
the o cial qrels we re-evaluated the baseline run of these four runs in order to verify the observed
tendencies of the training phase.
4.1</p>
          <p>Description of Submitted Runs and Results
We participated in the Main task of this track with four runs for the Medium set that contains
1,000 topics in di erent languages. All four runs where based on the BM25 retrieval model using
standard parameter values (b = 0.75, k1 = 1.2, k3 =1000), and utilized a percentage threshold of
3.0. These runs are listed below:</p>
          <p>BM25medStandard: No ltering of query terms by language was applied. Query terms where
selected solely considering their df.</p>
          <p>BM25EnglishTerms: German and French terms were ltered out.</p>
          <p>BM25FrenchTerms: English and German terms were ltered out.</p>
          <p>BM25GermanTerms: English and French terms were ltered out.</p>
          <p>Additionally we submitted a run for the XL set consisting of 10000 topics. This run also utilized
a threshold of 3.0, used the Cosine retrieval model, and ltered out French and German terms
via the utilization of dictionaries that were constructed based on the documents in the Clef-IP 09
corpus. Table 4 lists the o cial results of the above described runs.
4.2</p>
          <p>Analysis
While the CosEnglishTerms run showed comparable performance to the observations during the
training phase outlined in Figure 3, it can be seen from the results that the performance of the
BM25 based runs was signi cantly lower than the observed results in Figure 2 . Therefore rst
(a) MAP for varying percentage thresholds with the BM25 Model
(b) Number of retrieved rel. documents for varying percentage thresholds with the BM25 Model</p>
          <p>(a) MAP performance for varying query length with the Cosine model
(b) Number of retrieved rel. documents for varying query length with the Cosine model
of all, it is not possible for us to draw any conclusions towards the e ect of the applied ltering
by language from these results. In retrospective analysis we identi ed that this almost complete
failure in terms of performance was linked to applying the BM25 model to the Indri indices created
to allow for language ltering. While this problem has not yet been resolved and we were therefore
not able to re-evaluate the language ltering based runs, we re-evaluated the BM25medstandard
run using a Lemur based index and the released o cial qrels. This resulted in the below listed
performance, that is in the same range of what we witnessed during the BM25 training phase. It
con rms our observation that BM25 seems to be more e ective than the Cosine model.
run id
BM25medstandard</p>
          <p>P5
0.1248</p>
          <p>P10
0.0836</p>
          <p>P100
0.0188</p>
          <p>R
0.511</p>
          <p>MAP
0.1064</p>
        </sec>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>Conclusion and Future Outlook</title>
      <p>Based on one of the submitted runs and our training results this preliminary set of experiments has
shown that our proposed method of automatic query formulation may be interpreted as a
promising start towards e ective automatic query formulation. As such a technique may signi cantly
facilitate the process of prior art search through the automatic suggestion of e cient keywords,
it is planned to extend our experimentation in several directions. These extensions include the
consideration of a patent document's structure (i.e. title, description, claims) in the selection
process, and the introduction of a mechanism that will allow the weighted inclusion of term related
features in addition to the document frequency.
6</p>
    </sec>
    <sec id="sec-5">
      <title>Acknowledgements</title>
      <p>We would like to express our gratitude for being allowed to use the Large Data Collider (LDC)
computing resource provided and supported by the Information Retrieval Facility5. We speci cally
also want to thank Matrixware6 for having co-funded this research.
[1] National institue of informatics</p>
      <p>http://research.nii.ac.jp/ntcir/.
[2] Antoine Blanchard. Understanding and customizing stopword lists for enhanced patent
mapping. World Patent Information, 29(4):308 { 316, 2007.
[3] European Patent O ce (EPO). Guidelines for Examination in the European Patent O
December 2007.
ce,
5http://www.ir-facility.org/
6http://www.matrixware.com/
test
collection
for
ir
systems
(ntcir).</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>Atsushi</given-names>
            <surname>Fujii</surname>
          </string-name>
          , Makoto Iwayama, and
          <string-name>
            <given-names>Noriko</given-names>
            <surname>Kando</surname>
          </string-name>
          .
          <article-title>Overview of patent retrieval task at ntcir-4</article-title>
          .
          <source>In Proceedings of NTCIR-4 Workshop Meeting</source>
          ,
          <year>2004</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>Atsushi</given-names>
            <surname>Fujii</surname>
          </string-name>
          , Makoto Iwayama, and
          <string-name>
            <given-names>Noriko</given-names>
            <surname>Kando</surname>
          </string-name>
          .
          <article-title>Overview of patent retrieval task at ntcir-5</article-title>
          .
          <source>In Proceedings of NTCIR-5 Workshop Meeting</source>
          ,
          <year>2005</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>Atsushi</given-names>
            <surname>Fujii</surname>
          </string-name>
          , Makoto Iwayama, and
          <string-name>
            <given-names>Noriko</given-names>
            <surname>Kando</surname>
          </string-name>
          .
          <article-title>Overview of the patent retrieval task at the ntcir-6 workshop</article-title>
          .
          <source>In Proceedings of NTCIR-6 Workshop Meeting</source>
          , pages
          <volume>359</volume>
          {
          <fpage>365</fpage>
          ,
          <year>2007</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>Makoto</given-names>
            <surname>Iwayama</surname>
          </string-name>
          , Atsushi Fujii, and
          <string-name>
            <given-names>Noriko</given-names>
            <surname>Kando</surname>
          </string-name>
          .
          <article-title>Overview of classi cation subtask at ntcir-5 patent retrieval task</article-title>
          .
          <source>In Proceedings of NTCIR-5 Workshop Meeting</source>
          ,
          <year>2005</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>Makoto</given-names>
            <surname>Iwayama</surname>
          </string-name>
          , Atsushi Fujii, and
          <string-name>
            <given-names>Noriko</given-names>
            <surname>Kando</surname>
          </string-name>
          .
          <article-title>Overview of classi cation subtask at ntcir-6 patent retrieval task</article-title>
          .
          <source>In Proceedings of NTCIR-6 Workshop Meeting</source>
          , pages
          <volume>366</volume>
          {
          <fpage>372</fpage>
          ,
          <year>2007</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>Makoto</given-names>
            <surname>Iwayama</surname>
          </string-name>
          , Atsushi Fujii, Noriko Kando, and
          <string-name>
            <given-names>Akihiko</given-names>
            <surname>Takano</surname>
          </string-name>
          .
          <article-title>Overview of patent retrieval task at ntcir-3</article-title>
          .
          <source>In Proceedings of NTCIR-3 Workshop Meeting</source>
          ,
          <year>2002</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>