<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Running CLEF-IP experiments using a graphical query builder</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>W. Alink</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>R. Cornacchia</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>A. P. de Vries Centrum Wiskunde en Informatica</string-name>
        </contrib>
      </contrib-group>
      <abstract>
        <p>The CWI submission for the CLEF-IP track shows results for out-of-the-box querying using a graphical strategy design interface. This domain-independent search platform has been enriched with patent-speci c information, which was then readily available to the query interface. The search strategies for the 4 runs submitted have been constructed by simple drag&amp;drop operations in this graphical interface, subsequently compiled into probabilistic relational algebra (PRA) [3] and SQL, and then executed on a relational high-performance database system [2]. The four search strategies compare boolean search, ranked retrieval, and categorybased re-ranking. The main lesson learned is that using selection on category only yields a high recall.</p>
      </abstract>
      <kwd-group>
        <kwd>Question answering</kwd>
        <kwd>Questions beyond factoids</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>di erent retrieval strategies without the need of re-programming or re-con guring it. More details
on the runs are provided in Section 4.
2</p>
    </sec>
    <sec id="sec-2">
      <title>Ob jectives</title>
      <p>The main objectives of the submitted runs are to show exibility in expressing di erent strategies
for patent-document retrieval using a seamless combination of information retrieval and database
technologies: possibility to mix freely probabilistic (IR) and exact (DB) match criteria, and a neat
separation of the retrieval speci cations (IR) and physical data management (DB).</p>
      <p>Tasks performed by intellectual property specialists are often ad hoc, and continuously require
new approaches to search a collection of documents. Our objective is therefore to focus on the ease
of expressing new search strategies for IP search experts who not necessarily have high IR/database
expertise. Intellectual property specialists also need a high degree of control over the searches that
are performed. Our objective is therefore also to be able to combine exact match operators with
ranking operators, and provide the intellectual property specialist with an intuitive overview of
the search steps used in his strategy, so that results can be explained and veri ed.</p>
      <p>By targeting the XL experiments, our aim also includes to build a scalable solution, although
due to time restrictions, and the early stage of the system used, this issue has been not been
thoroughly addressed.
3</p>
    </sec>
    <sec id="sec-3">
      <title>Approach</title>
      <p>
        The CWI submission for the CLEF-IP track was powered by the LHM project, a joint project
with Apriorie [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ] and a leading IP search provider company. The main aim of the project is to
build an integrated system composed of:
Strategy Builder: a graphical user interface that enables patent experts to create complex search
strategies in a drag&amp;drop fashion.
      </p>
      <p>HySpirit: a software framework for probabilistic reasoning on relational and object-relational
data, developed by Apriorie.</p>
      <p>MonetDB: an open source high-performance database management system developed by CWI.</p>
      <p>The hypothesis is that such an integrated system enables users of the Strategy Builder to
formulate and execute expressive queries and retrieval strategies e ciently on the large-scale patent
corpus of the CLEF-IP track. This goal is achieved by implementing several automatic translation
steps. First, the graphical, user-oriented, strategy is composed of building blocks which are
internally expressed in terms of the HySpirit Probabilistic Relational Algebra (PRA). This guarantees
the search strategy to be grounded on a solid theoretical framework, that properly propagates
relevance probabilities throughout the whole search process, while hiding explicit management of
such probabilities. Second, the PRA speci cation is translated to a database query, and executed
on the high-performance database engine MonetDB, using the standard SQL query language.</p>
      <p>Fig. 1 shows the category -run strategy and excerpts of the intermediate compiled strategy
representations. A graphical representation of the complete category -run strategy is shown in
Fig. 1a. Fig. 1b zooms in on a single building block of that strategy. The corresponding probabilistic
(PRA) that is attached to this building block is shown in Fig. 1c. The compilation of the PRA
snippet yields the SQL code depicted in Fig. 1d which can be directly executed on a database
engine.
4</p>
    </sec>
    <sec id="sec-4">
      <title>Tasks performed</title>
      <p>In total 4 runs have been submitted to the CLEF-IP track. A short explanation of each of the
runs:
(a) Strategy Builder - a complete strategy</p>
      <p>(b) Strategy Builder - a building block
BLOCK_SOURCE(docID) = INPUT1_result
BLOCK_SUBJECTS(neID) = INPUT2_result</p>
      <sec id="sec-4-1">
        <title>BLOCK_prd(neID,docID)</title>
        <p>= PROJECT ALL [neID, docID] (
SELECT[predicate="%SELECTION%"] (</p>
      </sec>
      <sec id="sec-4-2">
        <title>INPUT1_ne_doc ) );</title>
      </sec>
      <sec id="sec-4-3">
        <title>BLOCK_nes(neID, docID)</title>
        <p>= PROJECT ALL [neID, docID] (
JOIN INDEPENDENT [docID=docID] (</p>
      </sec>
      <sec id="sec-4-4">
        <title>BLOCK_SOURCE_result,</title>
      </sec>
      <sec id="sec-4-5">
        <title>BLOCK_prd ) );</title>
      </sec>
      <sec id="sec-4-6">
        <title>BLOCK_result(docID)</title>
        <p>= PROJECT DISTINCT[docID] (
JOIN INDEPENDENT [neID=neID] (</p>
      </sec>
      <sec id="sec-4-7">
        <title>BLOCK_SUBJECTS_result,</title>
      </sec>
      <sec id="sec-4-8">
        <title>BLOCK_nes ) );</title>
        <p>(c) PRA query for the building block in Fig. 1b</p>
      </sec>
      <sec id="sec-4-9">
        <title>CREATE VIEW BLOCK_prd_1 AS</title>
      </sec>
      <sec id="sec-4-10">
        <title>SELECT neID AS a1, predicate AS a2, docID AS a3, prob</title>
      </sec>
      <sec id="sec-4-11">
        <title>FROM INPUT1_ne_doc WHERE predicate='%SELECTION%';</title>
      </sec>
      <sec id="sec-4-12">
        <title>CREATE VIEW BLOCK_prd AS</title>
      </sec>
      <sec id="sec-4-13">
        <title>SELECT a1, a3 AS a2, prob</title>
      </sec>
      <sec id="sec-4-14">
        <title>FROM BLOCK_prd_1;</title>
      </sec>
      <sec id="sec-4-15">
        <title>CREATE VIEW BLOCK_nes_1 AS</title>
      </sec>
      <sec id="sec-4-16">
        <title>SELECT INPUT1.a1 AS a1,</title>
      </sec>
      <sec id="sec-4-17">
        <title>BLOCK_prd.a1 AS a2, BLOCK_prd.a2 AS a3,</title>
      </sec>
      <sec id="sec-4-18">
        <title>INPUT1.prob * BLOCK_prd.prob AS prob</title>
      </sec>
      <sec id="sec-4-19">
        <title>FROM INPUT1, BLOCK_prd WHERE INPUT1.a1= BLOCK_prd.a2;</title>
      </sec>
      <sec id="sec-4-20">
        <title>CREATE VIEW BLOCK_nes AS</title>
      </sec>
      <sec id="sec-4-21">
        <title>SELECT a1, a3 AS a2, prob</title>
      </sec>
      <sec id="sec-4-22">
        <title>FROM BLOCK_nes_1;</title>
      </sec>
      <sec id="sec-4-23">
        <title>CREATE VIEW BLOCK_result_1 AS</title>
      </sec>
      <sec id="sec-4-24">
        <title>SELECT INPUT2.a1 AS a1,</title>
      </sec>
      <sec id="sec-4-25">
        <title>BLOCK_nes.a1 AS a2, BLOCK_nes.a2 AS a3,</title>
      </sec>
      <sec id="sec-4-26">
        <title>INPUT2.prob * BLOCK_nes.prob AS prob</title>
      </sec>
      <sec id="sec-4-27">
        <title>FROM INPUT2, BLOCK_nes WHERE INPUT2.a1=BLOCK_nes.a1;</title>
      </sec>
      <sec id="sec-4-28">
        <title>CREATE VIEW BLOCK_result AS</title>
      </sec>
      <sec id="sec-4-29">
        <title>SELECT a3 AS a1, 1-prod(1-prob) AS prob</title>
      </sec>
      <sec id="sec-4-30">
        <title>FROM BLOCK_result_1 GROUP BY a3;</title>
        <p>(d) SQL query for the building block in Fig. 1b</p>
        <p>
          The reason we think this somewhat resembles boolean retrieval is that from various patent
search experts we have heard that often the initial phase of a search is done by selecting key
terms from the patent under inspection, and generating such a query with those words so
that an amount of results is retrieved of which it is feasible to read all the abstracts.
bm25 -run: A well-known and often applied strategy in information retrieval is the BM25 relevance
model [
          <xref ref-type="bibr" rid="ref7">7</xref>
          ]. Our bm25 -run uses 15 keywords taken from the topic patent, and searches the
patent-document collection using the BM25 formula.
        </p>
        <p>The keywords taken from the topic document are weighted based on tf-idf. The initial weight
of the keywords is taken into account when ranking the documents.
category -run: In the category -run patent-documents are selected that matched one or more IPCR-categories
of the topic-patent. The IPCR-categories are weighted based on idf. The patent-documents
are ranked by the sum of matching category scores.
category-bm25 -run: uses the category strategy and applies the bm25 strategy to the results of this strategy.</p>
        <p>This run combines the bm25 -run and the category -run. First the patent-documents are
selected that match one or more of the categories in the topic-patent, and afterwards this
set of documents is searched using keywords extracted from the topic patent. Scores are
propagated at each step, so BM25 gets as input a list of weighted documents. The same
keywords have been used as were used in the bm25 -run</p>
        <p>For the boolean, bm25, and category-bm25 runs, text search has been performed on all the
`textual' sections of a document (title, abstract, description, and claims), and no speci c sections
have been queried.
5</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>Experimental Setup</title>
      <p>In our approach the process of creating indices for the data is separated from querying the data.
The same indices are used for each of the submitted runs. The only di erences between the 4 runs
are changes in the strategy.</p>
      <p>
        The schema used is comparable to RDF [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ] triple schema; all entries are subject, predicate,
object tuples. The most noticeable di erence of our schema compared to RDF is that probabilities
are attached to each tuple. The CLEF-IP corpus has been provided as a set of XML documents in
a custom XML format. All data has been loaded as XML using MonetDB/XQuery [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. Keyword
indices have been created using the PF/TIJAH indexer [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ]. Structural relations between patents
are obtained by using a domain speci c knowledge that was expressed as a set of XQuery [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ]
queries.
      </p>
      <p>The strategies are expressed in the strategy builder's `building blocks'. Each building block
contains snippets of PRA code, and the combined code is then compiled into a full PRA expression,
Subsequently the PRA is compiled into SQL statements using the PRA2SQL conversion of the
HySpirit engine. This compiler is a result of the LHM project. The nal script is then executed
on a MonetDB/SQL database engine.</p>
      <p>The strategy-builder had initially been build to run a single query with a given set of
parameters. To allow the system to execute a full run of topics at a time, it was changed in such a
way that it would compile the query-template once, and then would substitute the keywords and
categories for each topic.</p>
      <p>The main software components that are used:</p>
      <p>LHM Strategy Builder v0.2, con gured with a speci c workspace `CLEF-IP'
run
boolean
bm25
category
categorybm25
other1
other2
MonetDB Feb2009 SP2 release, patched so the MonetDB/TIJAH indexer results can be used
in MonetDB/SQL.</p>
      <p>The data is physically distributed over 4 di erent databases, each holding the indices of 500k
patent-documents. Creating the indices for the documents took little over 10 hours. For the
experiments on the 2 million patent documents we have been allowed to use some of the IR
Facility resources; the LDC, an Altix-4700 machine which has ample resources to use (80 cores
and 360GB of shared memory). Both during indexing and querying (only) 4 cores in parallel are
used. During querying 3 GB of memory per database are needed. During indexing much more
memory is needed, but due to the fact that the LDC has ample memory available (360GB) no
problems occurred.
6</p>
    </sec>
    <sec id="sec-6">
      <title>CLEF-IP Results</title>
      <p>
        Results of the CLEF-IP runs have been made available in [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ], and are summarised in [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ]. To
describe the overall result of the CWI submission for the CLEF-IP: in terms of scores none of the
runs have real good results, at least when compared to other participants. The results over the
S, M and XL bundles seem similar. The results for the M bundle (Table 9 in [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ]) are used here
for analysis, as this is the largest bundle for which all our 4 runs submitted results. The most
interesting item using only category information of a patent can yield high recall In Fig. 3a and
Fig. 3b the precision and recall scores are presented for the 4 submitted runs. For comparison,
the results of other participants that obtained the highest scores are also shown.
      </p>
      <p>There are a few observations to be made when looking at each of the runs individually. The
boolean-run provides poor retrieval quality. The reason why it has relatively high precision is
probably due to the fact that not always 1000 results are returned, and often much less. The
category -run has high recall, but MAP2 is very low.</p>
      <p>The bm25 -run resembles the most classic method of ranking documents. Results are (slightly)
lower than other participants methods. This is perhaps due to the fact that poor parameters have
been used. The category-bm25 -run does somewhat improve precision and recall over the bm25
run, but MAP is lower. Compared to the category -run, the category-bm25 -run does somewhat
improve precision and MAP, but recall is lower.</p>
      <p>A part of the explanation for the results could be the aggregation of patent-documents to
patents. For each of the runs the same aggregation method is used. The selection of the patents
has in each of the runs been the nal operation of the strategy: in the category-bm25 -run the
intermediate results between the category and bm25 part are patent-documents and not patents.</p>
      <p>The execution time for a single topic in the category-, bm25- and boolean-runs is roughly 3 to
6 seconds. The execution time for a single topic in the bm25-category -run is roughly 30 seconds.
Why the bm25-category -run is much slower than the other runs has not been analysed in detail,
but it may possibly be due to query plan optimisation for this (more complex) query.
0,2
0,15
0,1
0,05
0</p>
      <p>P5</p>
      <p>Boolean
BM25
Category
CategoryBM25
other1
other2
0,8
0,7
0,6
0,5
0,4
0,3
0,2
0,1
0</p>
      <p>Boolean
BM25
Category
CategoryBM25
other1
other2
P10</p>
      <p>P100</p>
      <p>P</p>
      <p>R5</p>
      <p>R10</p>
      <p>R100</p>
      <p>R
(a) Precision at 5, 10, 100 and overall preci- (b) Recall at 5, 10, 100 and overall recall for
sion for bundle M, task Main.In gray are highest bundle M, task Main.In gray are highest scores
scores by other participants. by other participants.
Participating in CLEF-IP 2009 has been an interesting experience. The rst of our objectives,
exibility and ease of use, is reached: constructing strategies in a graphical interface using
highlevel abstract concepts worked well and proved to be exible enough to express the retrieval
strategies for the CLEF-IP 2009 submission, without any collection-speci c additional coding. In
particular, combining exact and ranked matches required no e ort: this distinction and the proper
propagation of scores (probabilities) are totally transparent in the graphical user interface. The
integration with a general purpose database engine as a back-end worked smoothly as well, with
all the physical details abstracted away from the query interface. It would be interesting to see
whether other retrieval strategies used in CLEF-IP 2009 could be easily formulated in our Strategy
Builder.</p>
      <p>The second of our objectives, scalability, is only partly reached: we were able to handle the
2M patents of the CLEF-IP 2009 corpus, but this is still an order of magnitude less than all
patent-documents digitally available.</p>
      <p>Quality of retrieval results is not excellent, and should be improved. The main interesting
results: the category -run exhibits good recall, but poor MAP. This could mean that the IPCR
classi cation is good for selecting relevant patent-documents, but seems a poor criteria for ranking
in the way it is used in our strategy. More investigation is needed to determine whether category
information can be e ectively used for ranking patents.</p>
      <p>The main points to be improved, or at least to be further investigated:</p>
      <p>The parameters used in the BM25 ranking formula were not chosen carefully, also the
`patentdocument to patent' aggregation may be an interesting point of research, and could be
improved upon.</p>
      <p>For intermediate steps inside a composed strategy, it might be more useful to regard all
patent-documents of a retrieved patent again, instead of the individual retrieved
patentdocuments. This has not been analysed.</p>
      <p>The strategies used are oblivious of the language in which the patent has been written.
Better retrieval should be possible if language is taken into account
IDF of top topic-terms is computed over the XL bundle of topic documents, rather than over
the patent corpus. It is currently unknown whether this has a high e ect on the results.</p>
      <p>Finally, we would like to thank the IRF for providing access to the LDC, which made the
experiments much easier to perform.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>Peter</given-names>
            <surname>Boncz</surname>
          </string-name>
          , Torsten Grust, Maurice van Keulen,
          <string-name>
            <surname>Stefan Manegold</surname>
            , Jan Rittinger, and
            <given-names>Jens</given-names>
          </string-name>
          <string-name>
            <surname>Teubner</surname>
          </string-name>
          .
          <article-title>Monetdb/xquery: a fast xquery processor powered by a relational engine</article-title>
          .
          <source>In SIGMOD '06: Proceedings of the 2006 ACM SIGMOD international conference on Management of data</source>
          , pages
          <volume>479</volume>
          {
          <fpage>490</fpage>
          , New York, NY, USA,
          <year>2006</year>
          . ACM.
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <surname>CWI.</surname>
          </string-name>
          <article-title>MonetDB website</article-title>
          . http://www.monetdb.nl/.
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>Norbert</given-names>
            <surname>Fuhr</surname>
          </string-name>
          and
          <article-title>Thomas Rolleke. A probabilistic relational algebra for the integration of information retrieval and database systems</article-title>
          .
          <source>ACM Trans. Inf</source>
          . Syst.,
          <volume>15</volume>
          (
          <issue>1</issue>
          ):
          <volume>32</volume>
          {
          <fpage>66</fpage>
          ,
          <year>1997</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>Djoerd</given-names>
            <surname>Hiemstra</surname>
          </string-name>
          , Henning Rode, Roel van Os,
          <article-title>and Jan Flokstra. PF/Tijah: text search in an XML database system</article-title>
          . In Second International Workshop on Open Source Information Retrieval, Seattle, USA,
          <year>2006</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>Apriorie</given-names>
            <surname>LTD</surname>
          </string-name>
          .
          <article-title>Apriorie website</article-title>
          . http://www.apriorie.co.uk/.
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>Florina</given-names>
            <surname>Piroi</surname>
          </string-name>
          , Giovanna Roda, and
          <string-name>
            <given-names>Veronika</given-names>
            <surname>Zenz. CLEF-IP 2009 Evaluation Summary</surname>
          </string-name>
          ,
          <year>2009</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <surname>Stephen</surname>
            <given-names>E.</given-names>
          </string-name>
          <string-name>
            <surname>Robertson</surname>
            , Steve Walker, Susan Jones, Micheline Hancock-Beaulieu, and
            <given-names>Mike</given-names>
          </string-name>
          <string-name>
            <surname>Gatford</surname>
          </string-name>
          .
          <article-title>Okapi at TREC-3</article-title>
          . In Third Text REtrieval Conference (TREC
          <year>1994</year>
          ),
          <year>1994</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>Giovanna</given-names>
            <surname>Roda</surname>
          </string-name>
          , John Tait, Florina Piroi, and
          <string-name>
            <given-names>Veronika</given-names>
            <surname>Zenz</surname>
          </string-name>
          .
          <article-title>CLEF-IP 2009: retrieval experiments in the Intellectual Property domain</article-title>
          .
          <source>In CLEF working notes 2009</source>
          , Corfu, Greece,
          <year>2009</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <fpage>W3C</fpage>
          .
          <article-title>Resource description framework</article-title>
          . http://www.w3.org/RDF/.
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          <source>[10] W3C. XQuery 1</source>
          .0:
          <string-name>
            <surname>An</surname>
            <given-names>XML</given-names>
          </string-name>
          <string-name>
            <surname>Query</surname>
          </string-name>
          <article-title>Language</article-title>
          . http://www.w3.org/TR/xquery/.
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>