<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Patent Retrieval Experiments in the Context of the CLEF IP Track 2009</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Daniela Becks</string-name>
          <email>becks@uni-hildesheim.de</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Christa Womser-Hacker</string-name>
          <email>womser@uni-hildesheim.de</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Thomas Mandl</string-name>
          <email>mandl@uni-hildesheim.de</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Ralph Kölle</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Information Science, University of Hildesheim</institution>
          ,
          <addr-line>Marienburger Platz 22 D-31141 Hildesheim</addr-line>
          ,
          <country country="DE">Germany</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Intellectual Property</institution>
          ,
          <addr-line>Evaluation</addr-line>
          ,
          <country>Patent Retrieval System</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>At CLEF 2009 the University of Hildesheim focused on the main task of the Intellectual Property Track which aims at finding prior art for a specified patent [cf. Information Retrieval Facility 2009]. The experiments of the University of Hildesheim concentrated on a baseline approach including stopword elimination, stemming and simple term queries. Furthermore only title and claim were included into the index as especially the second one is considered to be the most important patent part during a prior art search [cf. Graf/Azzopardi 2008: 64].</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1 Introduction</title>
      <p>•
•
•</p>
      <sec id="sec-1-1">
        <title>Bibliographic data (e.g. inventor)</title>
        <p>Disclosure (e.g. title and detailed description)</p>
        <p>Claims
A look at the patent documents of the collection as well as the topic set reveals a similar structure. Furthermore
one can see that the description is the longest part of the document. It is followed by the claim section which in
general consists of a number of claims [cf. e.g. Patent number EP-1114924-B1].</p>
        <p>As we wanted to investigate whether simple statistical methods work well in such a special domain we adopted a
baseline approach including stopword elimination, stemming and simple term queries. In the future, our goal is
to compare this approach with a more sophisticated linguistic indexing and to apply a bag of terminological
resources.</p>
      </sec>
    </sec>
    <sec id="sec-2">
      <title>2 Indexing approach for patent documents</title>
      <p>As already said before, the University of Hildesheim adopted a baseline approach. A first analysis of the test
collection as well as the topic set revealed that in each patent document one could find a German, English and
French title. The same went for the claims. That´s why we decided to perform different monolingual runs – for
English and German - including the above mentioned parts of the patent document. It should be said that we
didn´t concentrate on French title and claims because German and English can be find more frequently in the
patent domain.</p>
      <p>All of the experiments were done using a simple retrieval system. More information about it will be given in
section 2.2. In particular syntax and patent specific terms caused some difficulties during the system setup. Some
of these problems will be discussed in more detail in the next section (2.1).</p>
    </sec>
    <sec id="sec-3">
      <title>2.1 Patent Terminology</title>
      <p>Patent documents are as special as the whole domain. Many scientists have already figured out the differences
between patents and other kinds of documents. A very important fact is described in Schamlu 1985. Here the
author points out that there are special rules which a patentee has to follow. Because of these rules some
sentence constructions return in each patent text [cf. Schamlu 1985: 63 ff]. This fact is confirmed by Ahmad and
Al-Thubaity [cf. 2003: 48].</p>
      <p>In the test collection as well as in the topic set this is the case for constructions like “as set forth in claim” [cf.
e.g. Patent number EP-1114924-B1]. The same goes for words like comprises or claim. As these words appear
frequently we decided to add them to the English stopword list. In German titles and claims the words umfasst or
Anspruch return in most patents. These ones were included into the German stopword list, too. Besides these text
patterns patent documents contain a lot of technical or general terms [cf. Graf/Azzopardi 2008: 64] which
strongly affected the retrieval experiments.</p>
      <p>In particular the huge amount of technical terms made the parsing process more difficult. Parsing errors occurred
for example because of German terms like AGR-System [cf. Patent number EP-1114924-B1]. As can be seen this
word contains a hyphen which had to be removed before. The same went for numbers which frequently appeared
in the claims. Besides parsing difficulties caused by the technical language the rather general vocabulary
influenced the retrieval process. As many patentees use vague and general expressions [cf. Graf/Azzopardi 2008:
64] like “Verfahren und Vorrichtung zur” [cf. e.g. Patent number EP-1117189-B1] many relevant documents are
returned if only simple term queries are conducted.</p>
      <p>An overview over the retrieval system used for the experiments of the University of Hildesheim will be given in
the next section (2.2).</p>
    </sec>
    <sec id="sec-4">
      <title>2.2 System Setup</title>
      <p>The experiments of the University of Hildesheim were done using a simple retrieval system based on Apache
Lucene1. This framework provides classes for stemming, indexing and searching. We followed the traditional
retrieval process which is listed below.</p>
      <p>As the collection consisted of patents in XML format a parser was necessary to first read out the content of the
documents. We finally decided to integrate the SAX Parser2 because using the DOM Parser seemed to take a lot
of time. The major reason for this might be that patents are much longer than other document types [cf.
Graf/Azzzopardi 2008: 64]. Following Iwayama et al. patent documents are even 24 times longer than
newspaper articles [cf. 2003: 254].</p>
      <sec id="sec-4-1">
        <title>1 http://lucene.apache.org/java/</title>
        <p>2 http://www.saxproject.org/
•
•
•</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>Indexing</title>
      <sec id="sec-5-1">
        <title>UCID</title>
        <p>CLAIM-TEXT</p>
        <p>INVENTION-TITLE</p>
      </sec>
    </sec>
    <sec id="sec-6">
      <title>Search process</title>
    </sec>
    <sec id="sec-7">
      <title>Stopword removal and stemming</title>
      <p>Because we performed monolingual runs with English and German terms we integrated one specific stopword
list3 for each language. As described in section 2.1 we added some patent specific terms to these lists.
After having removed all the stopwords we employed a stemmer to the text. While the Lucene Standard
Analyzer was chosen for English the German text was stemmed using the German Analyzer provided with
Lucene.</p>
      <p>Index field</p>
      <p>Part of patent</p>
      <sec id="sec-7-1">
        <title>Patent number</title>
        <p>Claims (including all claims available in a patent)
Title of the invention
To avoid having one huge file a separate index per language was created. Finally, a German as well as an
English index existed and could be used as the basis of the actual search process. The following fields were
included into the index file (cf. Table 1).</p>
        <p>Prior art search is performed to determine whether an invention or part of it already existed [cf. Graf/Azzopardi
2008: 64]. Any document that states prior art is relevant to the query. In the context of the Intellectual Property
Track a query is said to be a given patent.</p>
        <p>After having parsed the topic files which existed in XML format stopwords were removed. Again only title and
claims were considered for query formulation. Finally, we employed the same stemmer as during the indexing
process. The remaining terms were used as simple term queries.
3</p>
      </sec>
    </sec>
    <sec id="sec-8">
      <title>Results and Analysis</title>
      <p>The experiments of the University of Hildesheim concentrated on the main task of the Intellectual Property
Track 2009. Because of the above mentioned parsing problems as well as the time that was needed to adapt the
system to the specialties of the domain we actually submitted only one run. Furthermore to investigate the
influence of the single parameters some post runs were performed. A detailed description of the post runs can be
found in section 3.2.</p>
    </sec>
    <sec id="sec-9">
      <title>3.1 Submitted runs</title>
      <p>The University of Hildesheim submitted one run within the main task. For this run we concentrated on German
titles and claims only. The search process was based on the German index file. Equivalent to the described
German approach a second run considering only English titles and claims was performed. We didn´t submit this
run, but analyzed the results based on the provided relevance assessments.</p>
      <p>Unfortunately the results of our Hildesheim_MethodeA_Main_S run as well as the equivalent run with English
terms did not satisfy our expectations.</p>
      <p>One problem might be the huge amount of results returned by the system. This caused that among the submitted
result list a lot of relevant patents were missing. As we examined our results in more detail it was shown that
some relevant documents have been actually found by the system, but didn’t appear among the first 1000 patents
of the ranking list. Furthermore the results revealed that the retrieval results seemed to be better if the topic was</p>
      <sec id="sec-9-1">
        <title>3 http://members.unine.ch/jacques.savoy/clef/index.html</title>
        <p>of type “B1” meaning that the patent has already been granted [cf. Graf/Azzopardi 2008: 67]. For example in
case of topic number EP1169314 nine patent documents were said to be relevant. The system of the University
of Hildesheim returned seven of them.</p>
        <p>Unfortunately the performance as a whole remained relatively bad. To find out if the results of the experiments
can be improved in some way further runs were performed.</p>
      </sec>
    </sec>
    <sec id="sec-10">
      <title>3.2 Post runs</title>
      <p>As described before the results of the submitted run did not satisfy our expectations. During an intensive analysis
we tried to figure out which approach could lead to the improvement of the system´s performance. Finally, the
following post runs were performed. It should be said that these experiments had been run using a smaller topic
set of the first 50 XML documents only.
•
•
•</p>
      <p>De_ohne_Snowball</p>
      <sec id="sec-10-1">
        <title>En_ohne_Snowball</title>
        <p>In this case we concentrated on German terms and removed stopwords. In contrast to the submitted run the
Snowball Stemmer4 was included instead of the German Analyzer.</p>
        <p>This post run is the English equivalent to the first one. In this case the Standard Analyzer has been replaced
by the Snowball Stemmer.</p>
        <p>De_mit_Snowball</p>
        <p>In case of this run again German terms were used, but we didn´t remove stopwords.</p>
        <p>With the help of the first two post runs we wanted to investigate whether the retrieval results can be influenced
by the choice of analyzer. The experiments made clear that the performance of the retrieval system strongly
depends on the included analyzer. Figure 1 shows the results for the German runs.
were found. Instead during the post run (right side) the system only returned about 50 relevant documents. To
summarize the observations we have to state that in case of our German runs the Snowball Analyzer didn´t work
well.</p>
        <p>In contrast the results of the English run improved as the Snowball Analyzer had been included. This can be
clearly seen in figure 2. Although the number of documents which were not found by the system is relatively
high the number of relevant patents found obviously increased.</p>
        <p>It should be said that these results only refer to a relatively small topic set of 50 XML documents. Maybe this
would change if one would take into account the whole topic set. As a single run took us a lot of time we decided
in a first attempt to use only part of the given topic set.
The first two post runs already revealed a relation between the implemented analyzer and the number of relevant
documents returned by the retrieval system. Furthermore the influence of stopwords was investigated.
Our submitted run focussed on the basic retrieval approach which included stopword removal, but as the patent
domain differs in many ways this might not be the best approach. That´s why a further post run without
removing stopwords was performed. The results are illustrated in figure 3.</p>
      </sec>
    </sec>
    <sec id="sec-11">
      <title>Outlook</title>
      <p>The patent domain is quite different from other domains. Especially the linguistic features like terminology and
text structure bear a lot of difficulties. Because of these problems the University of Hildesheim only submitted
one run with German terms. With the help of further post runs we were able to figure out that the analyzer used
during the stemming process strongly influences the retrieval results. Furthermore there seems to be little
evidence that in the patent domain stopwords should not be removed before stemming, but this needs further
investigation.</p>
      <p>If we have a look at the results as a whole they are not quite good. In the future we will have to adopt our
retrieval system to the specialties of the patent domain. This includes the implementation of a better analyzer as
well as a robust parser. We will also have to investigate whether simple term queries are sufficient in the area of
patent retrieval.</p>
      <p>Overall we need to implement a more sophisticated search strategy. The experiments of this year provide a good
baseline for our further experiments.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <surname>Ahmad</surname>
          </string-name>
          , Khurshid; Al-Thubaity,
          <source>AbdulMosen</source>
          (
          <year>2003</year>
          )
          <article-title>: Can Text Analysis Tell us Something about Technology Progress?</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <surname>Chu</surname>
            , Aaron; Sakurai, Shigeyuki; Cardenas,
            <given-names>F.</given-names>
          </string-name>
          <string-name>
            <surname>Alfonso</surname>
          </string-name>
          (
          <year>2008</year>
          )
          <article-title>: Automatic Detection of Treatment Relationships for Patent Retrieval</article-title>
          .
          <source>In: Proceedings of the PaIR</source>
          <year>2008</year>
          , Octobre 30,
          <year>2008</year>
          ,
          <string-name>
            <given-names>Napa</given-names>
            <surname>Valley</surname>
          </string-name>
          , USA, pp.
          <fpage>9</fpage>
          -
          <lpage>14</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <surname>Graf</surname>
          </string-name>
          , Erik; Azzopardi,
          <string-name>
            <surname>Leif</surname>
          </string-name>
          (
          <year>2008</year>
          )
          <article-title>: A methodology for building a test collection for prior art search</article-title>
          .
          <source>In: Proceedings of the 2nd International Workshop on Evaluating Information Access (EVIA)</source>
          ,
          <source>December</source>
          <volume>16</volume>
          ,
          <year>2008</year>
          , Tokyo, Japan, pp.
          <fpage>60</fpage>
          -
          <lpage>71</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>Information</given-names>
            <surname>Retrieval Facility</surname>
          </string-name>
          (
          <year>2009</year>
          )
          <article-title>: CLEF-IP09 Track</article-title>
          . &lt;http://www.ir-facility.org/the_irf/clef-ip09-track&gt; (
          <volume>19</volume>
          .
          <fpage>08</fpage>
          .
          <year>2009</year>
          ,
          <volume>17</volume>
          :20)
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <surname>Iwayama</surname>
          </string-name>
          , Makoto; Fujii, Atsushi; Kando, Noriko; Marukawa,
          <string-name>
            <surname>Yuzo</surname>
          </string-name>
          (
          <year>2003</year>
          )
          <article-title>: An Empirical Study on Retrieval Models for Different Document Genres: Patents and Newspaper Articles</article-title>
          .
          <source>In: Proceedings of the ACM SIGIR 2003, July 28 - August 1</source>
          ,
          <year>2003</year>
          , Toronto, Canada, pp.
          <fpage>251</fpage>
          -
          <lpage>258</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <surname>Kando</surname>
          </string-name>
          ,
          <string-name>
            <surname>Noriko</surname>
          </string-name>
          (
          <year>2000</year>
          )
          <article-title>: What Shall We Evaluate? - Preliminary Discussion for the NTCIR Patent IR Challenge (PIC) Based on the Brainstorming with the Specialized Intermediaries in Patent Searching and Patent Attorneys</article-title>
          .
          <source>In: ACMSIGIR Workshop on Patent Retrieval, July</source>
          <volume>28</volume>
          ,
          <year>2000</year>
          , Athens, Greece, pp.
          <fpage>37</fpage>
          -
          <lpage>42</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <surname>Schamlu</surname>
          </string-name>
          ,
          <string-name>
            <surname>Mariam</surname>
          </string-name>
          (
          <year>1985</year>
          )
          <article-title>: Patentschriften - Patentwesen. Eine argumentationstheoretische Analyse der Textsorte Patent am Beispiel der Patentschriften zu Lehrmitteln</article-title>
          . Indicium-Verlag,
          <year>München</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>