<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Report on the CLEF-IP 2013 Experiments: Multilayer Collection Selection on Topically Organized Patents</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Anastasia Giachanou</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Michail Salampasis</string-name>
          <email>salampasis@ifs.tuwien.ac.at</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Maya Satratzemi</string-name>
          <email>maya@uom.gr</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Nikolaos Samaras</string-name>
          <email>samaras@uom.gr</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>University of Macedonia, Department of Applied Informatics</institution>
          ,
          <addr-line>Thessaloniki</addr-line>
          ,
          <country country="GR">Greece</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Vienna University of Technology, Institute of Software Technology and Interactive Systems</institution>
          ,
          <addr-line>1040, Vienna</addr-line>
          ,
          <country country="AT">Austria</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>This technical report presents the work which has been carried out using Distributed Information Retrieval methods for federated search of patent documents for the passage retrieval starting from claims (patentability or novelty search) task. Patent documents produced worldwide have manually-assigned classification codes which in our work are used to cluster, distribute and index patents through hundreds or thousands of sub-collections. For source selection, we tested CORI and a new collection selection method, the Multilayer method. We also tested CORI and SSL results merging algorithms. We run experiments using different combinations of the number of collections requested and documents retrieved from each collection. One of the aims of the experiments was to test older DIR methods that characterize different collections using collection statistics like term frequencies and how they perform in patent search and in suggesting relevant collections. Also to experiment with Multilayer, a new collection selection method that follows a multilayer, multi-evidence process to suggest collections taking advantage of the special hierarchical classification of patent documents. We submitted 8 runs. According to PRES @100 our best DIR approach ranked 6th across 21 submitted results.</p>
      </abstract>
      <kwd-group>
        <kwd>Patent Search</kwd>
        <kwd>IPC</kwd>
        <kwd>Source Selection</kwd>
        <kwd>Federated Search</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>This technical report presents the participation of the University of Macedonia in
collaboration with the Vienna University of Technology in the passage retrieval
(patentability or novelty search) task. Our experiments aim to explore an important issue,
the thematic organization of patent documents using the subdivision of patent data by
International Patent Classification (IPC) codes, and if this organization can be used to
improve patent search effectiveness using DIR methods in comparison to centralized
index approaches. We have also developed and tested a new collection selection
method that follows a multilayer, multi-evidence process to suggest collections taking
advantage of the special hierarchical classification of patent documents.</p>
      <p>Patent documents produced worldwide have manually-assigned classification
codes which in our experiments are used to topically organize, distribute and index
patents through hundreds or thousands of sub-collections. Our system automatically
selects the best collections for each query submitted to the system, something which
very precisely and naturally resembles the way patents professionals do various types
of patents searches, especially patent examiners doing invalidity search.</p>
      <p>In the experiments which are reported in this paper, we divided the CLEF-IP
collection using the subclass (Split-3), the main group (Split-4) and the subgroup level
(Split-5). The patents have been allocated to sub-collections based on the IPC codes
specified in them. In the experiments we report here, we allocated a patent to each
sub-collection specified by at least one of its IPC code, i.e. a sub-collection might
overlap with others in terms of the patents it contains.</p>
      <p>Topics in the patentability or novelty search task are sets of claims extracted from
actual patent application documents. Participants are asked to return passages that are
relevant to the topic claims. The topics contain also a pointer to the original patent
application file. Our participation was limited only at the document level. We didn’t
perform the claims to passage task because the main objective of our method is to
identify relevant IPCs. We submitted 8 runs. According to PRES @100 our best
submitted DIR approach ranked 6th across 21 submitted results.</p>
      <p>This paper is organized as follows. In Section 2 we present in detail how patents
are topically organized in our work using their IPC code. In Section 3 we describe the
DIR techniques that were tested on patent documents for our study and the new
methodology for collection selection proposed in this paper. In Section 4 we describe the
details of our experimental setup and the results. We follow with a discussion of the
rationale of our approach in Section 5 and future work and conclusions in Section 6.
2</p>
    </sec>
    <sec id="sec-2">
      <title>Topically Organised Patents for DIR</title>
      <p>
        The experiments which are reported in this paper extend our previous work of
applying DIR methods to topically organized patents
        <xref ref-type="bibr" rid="ref21 ref8 ref9">(Salampasis et al. 2012)</xref>
        . We
propose a new collection selection method that surpasses previous source/IPC selection
methods for topically organised patents. Another collection selection study involving
topically organized patents is reported in the literature
        <xref ref-type="bibr" rid="ref10">(Larkey et al. 2000)</xref>
        , however
this study was conducted many years ago with a different (USPTO) patent dataset.
Also, our approach of dividing patents is different and closer to the actual way of
patent examiners conducting patent searches, as we divide patents into a much larger
number of sub-collections. Additionally, we apply CORI in multiple layers and
evaluate its performance.
      </p>
      <p>
        All patents have manually assigned IPC codes
        <xref ref-type="bibr" rid="ref5">(Chen &amp; Chiu 2011)</xref>
        . IPC is an
internationally accepted standard taxonomy for classifying, sorting, organizing,
disseminating, and searching patents. It is officially administered by World Intellectual
Property Organization (WIPO). The IPC provides a hierarchical system of language
independent symbols for the classification of patents according to the different areas
of technology to which they pertain. IPC has currently about 71,000 nodes which are
organized into a five-level hierarchical system which is also extended in greater levels
of granularity. IPC codes are assigned to patent documents manually by technical
specialists.
      </p>
      <p>Patents can be classified by a number of different classification schemes. European
Classification (ECLA) and U.S. Patent Classification System (USPTO) are the most
known classification schemes used by EPO and USPTO respectively. Recently, EPO
and USPTO signed a joint agreement to develop a common classification scheme
known as Cooperative Patent Classification (CPC). The CPC that has been developed
as an extension of the IPC contains over 260,000 individual codes. For this study,
patents were organized based on IPC codes because this was the available
classification scheme in the test collection CLEF-IP.</p>
      <p>
        Although IPC codes are used to topically cluster patents into sub-collections,
something which is a prominent prerequisite for DIR, there are some important
differences which motivated us to re-examine and adapt existing DIR techniques in patent
search. Firstly, IPC are assigned by humans in a very detailed and purposeful
assignment process, something which is very different by the creation of sub-collections
using automated clustering algorithms. Also, patents are published electronically
using a strict technical form and structure
        <xref ref-type="bibr" rid="ref1">(Adams 2010)</xref>
        . This characteristic is another
reason to reassess existing DIR techniques because these have been mainly developed
for structureless and short documents such as newspapers or poorly structured web
documents. Another important difference is that patent search is recall oriented
because very high recall is required in most searchers
        <xref ref-type="bibr" rid="ref12">(Lupu et al. 2011)</xref>
        , i.e. a single
missed patent in a patentability search can invalidate a newly granted patent. This
contrasts with web search where high precision of initially returned results is the
requirement and about which DIR algorithms were mostly concentrated and evaluated
        <xref ref-type="bibr" rid="ref15 ref16">(Paltoglou et al. 2008)</xref>
        .
      </p>
      <p>Before we describe our study further we should explain IPC which determines how
we created the sub-collections in our experiments. Top-level IPC nodes consist of
eight sections which are: human necessities, performing operations, chemistry,
textiles, fixed constructions, mechanical engineering, physics, and electricity. A section
is divided into classes which are subdivided into subclasses. Subclass is divided into
main groups which are further subdivided into subgroups. In total, the current IPC has
8 sections, 129 classes, 632 subclasses, 7.530 main groups and approximately 63,800
subgroups.</p>
      <p>Table 1 shows a part of IPC. Section symbols use uppercase letters A through H. A
class symbol consists of a section symbol followed by two-digit numbers like F01,
F02 etc. A subclass symbol is a class symbol followed by an uppercase letter like
F01B. A main group symbol consists of a subclass symbol followed by one to
threedigit numbers followed by a slash followed by 00 such as F01B7/00. A subgroup
symbol replaces the last 00 in a main group symbol with two-digit numbers except for
00 such as F01B7/02. Each IPC node is attached with a noun phrase description
which specifies some technical fields relevant to that IPC code. Note that a subgroup
may have more refined subgroups (i.e. defining 6th, 7th level etc). Hierarchies among
subgroups are indicated not by subgroup symbols but by the number of dot symbols
preceding the node descriptions as shown in Table 1.</p>
    </sec>
    <sec id="sec-3">
      <title>Distributed IR on Patent Search</title>
      <sec id="sec-3-1">
        <title>Prior Work on Collection Selection</title>
        <p>
          Distributed Information Retrieval (DIR), also known as federated search
          <xref ref-type="bibr" rid="ref23 ref24">(Si &amp; J.
Callan 2003a)</xref>
          , offers users the capability of simultaneously searching multiple online
remote information sources through a single point of search. The DIR process can be
perceived as three separate but interleaved sub-processes: Source representation, in
which surrogates of the available remote collections are created
          <xref ref-type="bibr" rid="ref3">(Callan &amp; Connell
2001)</xref>
          . Source selection, in which a subset of the available information collections is
chosen to process the query
          <xref ref-type="bibr" rid="ref12 ref18">(Paltoglou et al. 2011)</xref>
          and results merging, in which the
separate results returned from remote collections are combined into a single merged
result list which is returned to the user for examination
          <xref ref-type="bibr" rid="ref15 ref16">(Paltoglou et al. 2008)</xref>
          .
        </p>
        <p>
          There are a number of Source Selection approaches including CORI
          <xref ref-type="bibr" rid="ref2 ref4">(Callan et al.
1995)</xref>
          , gGlOSS
          <xref ref-type="bibr" rid="ref6">(French et al. 1999)</xref>
          , and others
          <xref ref-type="bibr" rid="ref22">(Si et al. 2002)</xref>
          , that characterize
different collections using collection statistics like term frequencies. These statistics,
which are used to select or rank the available collections’ relevance to a query, are
usually assumed to be available from cooperative search providers. Alternatively,
statistics can be approximated by sampling uncooperative providers with a set of
queries
          <xref ref-type="bibr" rid="ref3">(Callan &amp; Connell 2001)</xref>
          .
        </p>
        <p>
          The Decision-Theoretic framework (DTF) presented by Fuhr
          <xref ref-type="bibr" rid="ref7">(Fuhr 1999)</xref>
          is one of
the first attempts to approach the problem of source selection from a theoretical point
of view. The Decision-Theoretic framework (DTF) produces a ranking of collections
with the goal of minimizing the occurring costs, under the assumption that retrieving
irrelevant documents is more expensive than retrieving relevant ones.
        </p>
        <p>
          In more recent years, there has been a shift of focus in research on source selection,
from estimating the relevancy of each remote collection to explicitly estimating the
number of relevant documents in each. ReDDE
          <xref ref-type="bibr" rid="ref23 ref24">(Si &amp; Callan 2003b)</xref>
          focuses at
exactly that purpose. It is based on utilizing a centralized sample index, comprised of all
the documents that are sampled in the query-sampling phase and ranks the collections
based on the number of documents that appear in the top ranks of the centralized
sample index. Its performance is similar to CORI at testbeds with collections of
similar size and better when the sizes vary significantly. Other methods see source
selection as a voting method where the available collections are candidates and the
documents that are retrieved from the set of sampled documents are voters
          <xref ref-type="bibr" rid="ref17">(Paltoglou et al.
2009)</xref>
          . Different voting mechanism can be used (e.g. BordaFuse, ReciRank,
Compsum) mainly inspired by data fusion techniques.
        </p>
        <p>There is a major difference between CORI and the other collection selection
algorithms presented in the paragraph above. CORI builds a hyperdocument representing
the sub-collection while using the other methods the collection selection or not is
based on the retrieval or not of individual documents from the single centralized
sample index. Due to this characteristic CORI may not work well in environments
containing a mix of “small” and “very large” document databases. On the other hand for
homogenous sub-collections as the ones produced from patents belonging to a single
IPC, representative hyperdocument in CORI should normally encompass a strong
discriminating power, something useful for effective and robust collection selection.</p>
        <p>In the experiments which we report in this paper we use both CORI and our
Multilayer method which adapts the way any selection method (CORI in the experiments
presented here) can be applied in patent domain.
3.2</p>
      </sec>
      <sec id="sec-3-2">
        <title>Multilayer Collection Selection</title>
        <p>We exploit the hierarchical organization of the IPC classification scheme and the
idea of topically organized patents to propose a new multiple-evidence Multilayer
collection selection method. The new method ranks collections/IPCs not only based
on the subdivision of patents in a specific IPC layer, but additionally utilizes the
ranking of their ancestors, if the same selection process (query) had been applied at a
higher level. This method can effectively suggest relevant collections at any
professional search system where high value documents exist that are organized
hierarchically according to an appropriate classification scheme.</p>
        <p>The motivation behind the Multilayer method is to select as many as possible
relevant collections at lower IPC levels (level 4, level 5 etc). IPC code selection when
applied at low levels can effectively help patent examiners to identify quickly the
subgroups they should focus and this can become a real time saver. In a recent field
survey we conducted, patent examiners expressed the problem of spending time
exploring IPC codes (sub-groups) that discover later they are not relevant. That happens
more often in smaller patent offices where patent examiners are usually asked to
examine patents in areas which they are relatively knowledgeable but not experts.
Especially in such conditions collection/IPC selection methods and tools could be very
useful for patents examiners while searching relevant patents.</p>
        <p>
          The proposed method is based on collections selected by CORI. Previous studies
showed that CORI performs better than other collection selection methods
(BordaFuse, Reciprocal Rank) when applied at the patent domain
          <xref ref-type="bibr" rid="ref21 ref8 ref8 ref9 ref9">(Salampasis et al.
2012; Giachanou et al. 2012)</xref>
          . We believe the reason is that CORI is based on a
content-based representation of sub-collections using a hyperdocument approach, while
the other methods use individual retrieved documents from a sub-collection to
estimate the relevance of a sub-collection. However, CORI tends to produce poorer
results at low IPC levels (level 4 or level 5). One reason is that the technological area of
patents belonging to a sub-collection is more accurately represented in higher IPC
levels (e.g. subclass) because it consists of less sub-collections. At higher IPC levels,
documents in one sub-collection are relatively homogeneous and better distinguished
from patents in other IPCs, something that is more difficult to capture in lower levels
of classification. For example, sub-collections of level 4 that contains about ten times
more sub-collections than level 3, are less easier differentiated between each other
using a hyperdocument approach, resulting in a decreased CORI performance. To
depict this differentiation more clearly, patents that represent methods for oral or
dental hygiene can be more easily differentiated from radiation therapy patents at level 3
while patents represent dental machines for boring may not be so easily differentiated
from patents that represent dental tools at level 4.
        </p>
        <p>
          In other words our method to apply source selection introduces a normalisation
procedure which takes into account the source selection results at several
classification levels. Of course, the proposed method can utilise multiple evidence, if the
documents are organized in at least two different levels. In this paper, we focus on
level 3 (subclass), level 4 (main group) and level 5 (subgroup). We used the CORI
collection selection algorithm to retrieve the relevant collections as it has been proven
more effective than other collection selection algorithms (e.g. BordaFuse, RR) that we
tested before
          <xref ref-type="bibr" rid="ref21 ref8 ref8 ref9 ref9">(Salampasis et al. 2012; Giachanou et al. 2012)</xref>
          .
        </p>
        <p>The lists returned from leveli and leveli+1 can be represented by two plots using the
collection and the score:
{(CollA, scoreA), (CollB, scoreB), ..., (CollN, scoreN)}
{(CollA.1, scoreA.1), (CollA.2, scoreA.2), ..., (CollA.M, scoreA.M),(CollB.1, scoreB.1),..., (CollN.1,
scoreN.1),..., (CollN.M’, scoreN.M’)}
where N is the number of sub-collections suggested at leveli, M is the number of
subcollections at leveli+1 that are children of collectionA and M’ is the number of
subcollections at leveli+1 that are children of collectionN.</p>
        <p>The new collection selection algorithm combines the information gathered from
the two levels to produce a new list of relevant collections. The new algorithm
evaluates the new scores for collections at leveli+1 according to the following equation:
scorey.z=a*scorey+(1-a)*scorey.z
(1)
where y is a sub-collection at leveli and z is a sub-collection at leveli+1 which is child
of the CollY . Parameter α determines the weight that each level will take to decide the
final score of a sub-collection (IPC).</p>
        <p>Another parameter of our method is the collection window which represents the
number of sub-collections that will be re-ranked after taking evidence from a higher
level. For example if the aim is to produce a list of N suggested IPCs at level 5, the
method should define how many IPCs in the initial rank produced by running CORI
in level 5, initially positioned after position N, will be reconsidered in the second
round re-ranking process. This is the window parameter and this decision can be
based either on a fixed threshold such as 100 or on a number relative to the number of
IPCs that should be suggested (i.e. 2 * N, 3 * N etc). Another parameter of our
method is influence factor, i.e. how many IPCs from a higher level (level 4 in our
example) should be used to re-rank the collection window IPCs in the lower level
(level 5). For example, if we want to re-rank 2*N IPCs in level 5, a parameter in our
method is how many IPCs from level 4 we will use to make the re-ranking.</p>
        <p>
          For the experiments in this study, we decided to use the parameters that optimized
the performance of Multilayer in a previous study
          <xref ref-type="bibr" rid="ref8 ref9">(Giachanou et al. 2012)</xref>
          . Based on
the results of that study, we decided the parameter a to be assigned with the value of
0.8. Finally, at split-4 the collection window and the influence factor were assigned
with the values of 20 and 200 respectively while at split-5 those parameters were
assigned with the values of 200 and 2000. The decision was based on a previous study
that was performed on CLEF 2012.
4
        </p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>Experimental Setup</title>
      <p>The data collection which was used in the study is CLEF-IP 2013 where patents
are extracts of the MAREC dataset, containing over 2.6 million patent documents
pertaining to 1.3 million patents from the EPO with content in English, German and
French, and extended by documents from the WIPO. We indexed the collection with
the Lemur toolkit. The fields which have been indexed are: title, abstract, description
(first 500 words), claims, inventor, applicant and IPC class information. Patent
documents have been pre-processed to produce a single (virtual) document representing a
patent. Our pre-processing involves also stop-word removal and stemming using the
Porter stemmer. In our study, we use the Inquery algorithm implementation of Lemur.</p>
      <p>We have divided the CLEF-IP collection using the subclass (split3), the main
group (split4) and the sub-group level (split5). This decision is driven by the way that
patent examiners work when doing patent searches who basically try to incrementally
focus into a narrower sub-collection of documents. In the present system, we allocate
a patent to each sub-collection specified by at least one of its IPC codes, i.e. a
subcollection might overlap with others in terms of the patents it contains. This is the
reason why the column #patents presents a number larger than the 1.3 million patents
that constitute the CLEF-IP 2011 collection.</p>
      <p>
        To test our system, we used a subset of the official queries provided in CLEF-IP
2013 dataset. The queries generated using the title, the abstract, the description and
the claims. Topics in French and German were first translated in English using the
WIPO Translation Assistant1. We tested CORI and Multilayer source selection
methods at split3, split4 and split5. For results merging, we applied CORI results merging
algorithm
        <xref ref-type="bibr" rid="ref2 ref4">(Callan et al. 1995)</xref>
        that is based on a heuristic weighted scores merging
1 https://www3.wipo.int/patentscope/translate/translate.jsf
algorithm and SSL. We also performed a run with the centralized index. The
multilayer method was tested at split4 and split5. To test the Multilayer method, we used
the collections selected by CORI at split3, split4 and split5.
5
      </p>
    </sec>
    <sec id="sec-5">
      <title>Results and Discussion</title>
      <p>As it is shown in Table 3 Multilayer performs better than the CORI at the main
group (Split4) level. To obtain a more complete picture of the results, we calculated a
Run description
centralised.EN
10-100.CORI.SSL.split5.EN
10-100.CORI.CORI.split3.EN
20-50.CORI.CORI.split5.EN
10-100.Multilayer.CORI.split4.EN
20-50.Multilayer.CORI.split5.EN
Median (us NOT including)- EN
Average (us NOT including)- EN
10-100.Multilayer.CORI.split5.EN</p>
      <p>10-100.CORI.SSL.split4.EN
Average (us NOT including) - all</p>
      <p>10-100.CORI.CORI.split3.all
10-100.Multilayer.CORI.split4.all</p>
      <p>centralised.all
20-50.CORI.CORI.split5.all
Median (us NOT including) - all</p>
      <p>
        10-100.CORI.SSL.split5.all
20-50.Multilayer.CORI.split5.all
10-100.Multilayer.CORI.split5.all
10-100.CORI.SSL.split4.all
recall measure Rn which is used to compare the performance of source selection
algorithms
        <xref ref-type="bibr" rid="ref11 ref14 ref2 ref4">(Callan et al. 1995; Nottelmann &amp; Fuhr 2003; Larson 2003)</xref>
        .
      </p>
      <p>Table 4 shows the results produced from the source selection algorithms ranked
according to Rk @10, Rk @20 and Rk @50 at split4 and split5. The best performing
algorithm at split4 is the Multilayer method where the first 50 suggested collections
contain about 70% of all relevant documents while CORI managed to identify about
45%. This is a very encouraging result that strongly suggests that source selection
algorithms can be effectively used to suggest sub-collections as starting points for
information seekers to search.</p>
      <p>Another important finding is that the best runs are those requesting fewer
subcollections (10 collections) and more documents from each selected sub-collection.
This fact is probably the result of the small number of relevant documents which exist
for each topic. To validate these observations we did a post-run analysis of the topics
and how their relevant documents are allocated to sub-collections in each split (Table
5). Table 5 reveals useful information which shows that to some extend relevant IPC
codes can be effectively identified if IPC classification codes are already assigned to a
topic. This is a feature that we didn’t use in our experiments and can be used as a
heuristic that could substantially increase the performance of source selection.</p>
      <p>In addition to the comments already discussed, perhaps the most interesting and
important finding for this study is that DIR approaches managed to perform similar or
better than the centralized index approaches. It is also very interesting that the
performance remains relatively high at subgroup level (split-5), the level that patent
examiners focus on. This is a very interesting finding which shows that DIR approaches
can be used to suggest collections at low levels while being effective and efficient.</p>
      <p>
        It seems that in patent domain the cluster-based approaches to information retrieval
        <xref ref-type="bibr" rid="ref25">(Willett 1988)</xref>
        <xref ref-type="bibr" rid="ref8">(Fuhr et al. 2012)</xref>
        which utilize document clusters (sub-collections),
could be utilized so efficiency or effectiveness can be improved. As for efficiency,
searching and browsing on sub-collections rather than the complete collection of
documents could significantly reduce the retrieval time of the system and more
significantly the information seeking time of users. In relation to effectiveness, the potential
of DIR retrieval stems from the cluster hypothesis
        <xref ref-type="bibr" rid="ref20">(Van Rijsbergen 1979)</xref>
        which
states that related documents residing in the same cluster (sub-collection) tend to
satisfy same information needs. The cluster hypothesis has been utilized in various
settings for information retrieval such as for example cluster-based retrieval, extensions
of IR models with clusters, latent semantic indexing. The expectation in the context of
source selection, which is of primarily importance for this study, is that if the correct
sub-collections are selected then it will be easier for relevant documents to be
retrieved from the smaller set of available documents and more focused searches can be
performed.
      </p>
      <p>
        The field of DIR has been explored in the last decade mostly as a response to
technical challenges such as the prohibitive size and exploding rate of growth of the web
which make it impossible to be indexed completely
        <xref ref-type="bibr" rid="ref19">(Raghavan &amp; Garcia-Molina
2001)</xref>
        . Also there is a large number of online sources (web sites), collectively known
as invisible web which are either not reachable by search engines because they sit
behind pay-to-use turnstiles, or for other reasons do not allow their content to be
indexed by web crawlers, offering their own search capabilities
        <xref ref-type="bibr" rid="ref13">(Miller 2007)</xref>
        . As the
main focus of this paper is patent search, we should mention this is especially true in
the patent domain as nearly all authoritative online patent sources (e.g. EPO’s
espacenet) are not indexable and therefore not accessible by general purpose search
engines.
6
      </p>
    </sec>
    <sec id="sec-6">
      <title>Conclusion and Future Work</title>
      <p>In this paper we presented the work which has been carried out using Distributed
Information Retrieval methods for federated search of patent documents for the
CLEF-IP 2013 passage retrieval starting from claims (patentability or novelty search)
task. We tested CORI and Multilayer methods for source selection and CORI and
SSL for results merging. We have divided the CLEF-IP collection using the subclass
(Split-3), the main group (Split-4) and the subgroup (Split-5) level to experiment with
different levels and depth of topical organization.</p>
      <p>We submitted 8 runs. According to PRES @100 our best DIR approach ranked 6th
across 21 submitted results. The methods we apply performed similar or better than
the centralised approach.</p>
      <p>We plan to explore further this line of work with exploring modifications to the
Multilayer and to make it more effective for patent search. We believe that the
discussion and the experiment presented in this paper are also useful to the designers of
patent search systems which are based on DIR methods.</p>
      <p>ACKNOWLEDGMENT. The second author is supported by a Marie Curie
fellowship from the IEF project PerFedPat (www.perfedpat.eu).
7</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Adams</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <year>2010</year>
          .
          <article-title>The text, the full text and nothing but the text: Part 1 - Standards for creating textual information in patent documents and general search implications</article-title>
          .
          <source>World Patent Information</source>
          ,
          <volume>32</volume>
          (
          <issue>1</issue>
          ), pp.
          <fpage>22</fpage>
          -
          <lpage>29</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Callan</surname>
            ,
            <given-names>J P</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lu</surname>
            ,
            <given-names>Z</given-names>
          </string-name>
          &amp;
          <string-name>
            <surname>Croft</surname>
            ,
            <given-names>W B</given-names>
          </string-name>
          ,
          <year>1995</year>
          .
          <article-title>Searching distributed collections with inference networks</article-title>
          .
          <source>In Proceedings of the 18th annual international ACM SIGIR conference on Research and development in information retrieval. ACM</source>
          , pp.
          <fpage>21</fpage>
          -
          <lpage>28</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Callan</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          &amp;
          <string-name>
            <surname>Connell</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <year>2001</year>
          .
          <article-title>Query-based sampling of text databases</article-title>
          .
          <source>ACM Transactions on Information Systems</source>
          ,
          <volume>19</volume>
          (
          <issue>2</issue>
          ), pp.
          <fpage>97</fpage>
          -
          <lpage>130</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Callan</surname>
          </string-name>
          ,
          <string-name>
            <surname>James</surname>
            <given-names>P</given-names>
          </string-name>
          , Lu, Zhihong &amp; Croft,
          <string-name>
            <given-names>W</given-names>
            <surname>Bruce</surname>
          </string-name>
          ,
          <year>1995</year>
          .
          <article-title>Searching distributed collections with inference networks</article-title>
          .
          <source>In Proceedings of the 18th annual international ACM SIGIR conference on Research and development in information retrieval - SIGIR '95</source>
          . Seattle, Washington: ACM New York, NY, USA, pp.
          <fpage>21</fpage>
          -
          <lpage>28</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Chen</surname>
            ,
            <given-names>Y.-L.</given-names>
          </string-name>
          &amp;
          <string-name>
            <surname>Chiu</surname>
          </string-name>
          , Y.-T.,
          <year>2011</year>
          .
          <article-title>An IPC-based vector space model for patent retrieval</article-title>
          .
          <source>Information Processing &amp; Management</source>
          ,
          <volume>47</volume>
          (
          <issue>3</issue>
          ), pp.
          <fpage>309</fpage>
          -
          <lpage>322</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>French</surname>
            ,
            <given-names>J.C.</given-names>
          </string-name>
          et al.,
          <year>1999</year>
          .
          <article-title>Comparing the Performance of Database Selection Algorithms</article-title>
          .
          <source>Proceedings of the 22nd annual international ACM SIGIR conference on Research and development in information retrieval SIGIR 99</source>
          , pp.
          <fpage>238</fpage>
          -
          <lpage>245</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>Fuhr</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <year>1999</year>
          .
          <article-title>A decision-theoretic approach to database selection in networked IR</article-title>
          .
          <source>ACM Transactions on Information Systems</source>
          ,
          <volume>17</volume>
          (
          <issue>3</issue>
          ), pp.
          <fpage>229</fpage>
          -
          <lpage>249</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>Fuhr</surname>
          </string-name>
          , N. et al.,
          <year>2012</year>
          .
          <article-title>The optimum clustering framework: implementing the cluster hypothesis</article-title>
          .
          <source>Information Retrieval</source>
          ,
          <volume>15</volume>
          (
          <issue>2</issue>
          ), pp.
          <fpage>93</fpage>
          -
          <lpage>115</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <surname>Giachanou</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Salampasis</surname>
            ,
            <given-names>M</given-names>
          </string-name>
          &amp; Paltoglou,
          <string-name>
            <surname>G</surname>
          </string-name>
          ,
          <year>2012</year>
          .
          <article-title>Multilayer Collection Selection and Search of Topically Organized Patents</article-title>
          .
          <source>In ceur-ws.org.</source>
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <surname>Larkey</surname>
            ,
            <given-names>L.S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Connell</surname>
            ,
            <given-names>M.E.</given-names>
          </string-name>
          &amp;
          <string-name>
            <surname>Callan</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <year>2000</year>
          .
          <article-title>Collection selection and results merging with topically organized U.S. patents and TREC data</article-title>
          .
          <source>In Proceedings of the ninth international conference on Information and knowledge management - CIKM '00</source>
          .
          <string-name>
            <surname>McLean</surname>
          </string-name>
          , Virginia, USA: ACM New York, NY, USA, pp.
          <fpage>282</fpage>
          -
          <lpage>289</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <string-name>
            <surname>Larson</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <year>2003</year>
          .
          <article-title>Distributed IR for digital libraries T</article-title>
          . Koch &amp; I. Sølvberg, eds.
          <source>… and Advanced Technology for Digital Libraries</source>
          ,
          <volume>2769</volume>
          , pp.
          <fpage>487</fpage>
          -
          <lpage>498</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          12.
          <string-name>
            <surname>Lupu</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          et al.,
          <year>2011</year>
          . Current Challenges in
          <string-name>
            <given-names>Patent</given-names>
            <surname>Information Retrieval M. Lupu</surname>
          </string-name>
          et al., eds., Springer.
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          13.
          <string-name>
            <surname>Miller</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <year>2007</year>
          .
          <article-title>Most fed data is un-googleable</article-title>
          .
          <source>Federal ComputerWeek.</source>
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          14.
          <string-name>
            <surname>Nottelmann</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          &amp;
          <string-name>
            <surname>Fuhr</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <year>2003</year>
          .
          <article-title>Evaluating different methods of estimating retrieval quality for resource selection</article-title>
          .
          <source>In Proceedings of the 26th annual international ACM SIGIR conference on Research and development in informaion retrieval - SIGIR '03</source>
          . Toronto, Canada: ACM New York, NY, USA, pp.
          <fpage>290</fpage>
          -
          <lpage>297</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          15.
          <string-name>
            <surname>Paltoglou</surname>
            ,
            <given-names>G</given-names>
          </string-name>
          , Salampasis,
          <string-name>
            <given-names>M</given-names>
            &amp;
            <surname>Satratzemi</surname>
          </string-name>
          ,
          <string-name>
            <surname>M,</surname>
          </string-name>
          <year>2008</year>
          .
          <article-title>A results merging algorithm for distributed information retrieval environments that combines regression methodologies with a selective download phase</article-title>
          .
          <source>Information Processing &amp; Management</source>
          ,
          <volume>44</volume>
          (
          <issue>4</issue>
          ), pp.
          <fpage>1580</fpage>
          -
          <lpage>1599</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          16.
          <string-name>
            <surname>Paltoglou</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Salampasis</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          &amp;
          <string-name>
            <surname>Satratzemi</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <year>2008</year>
          .
          <article-title>A results merging algorithm for distributed information retrieval environments that combines regression methodologies with a selective download phase</article-title>
          .
          <source>Information Processing &amp; Management</source>
          ,
          <volume>44</volume>
          (
          <issue>4</issue>
          ), pp.
          <fpage>1580</fpage>
          -
          <lpage>1599</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          17.
          <string-name>
            <surname>Paltoglou</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Salampasis</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          &amp;
          <string-name>
            <surname>Satratzemi</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <year>2009</year>
          . Advances in Information Retrieval M.
          <string-name>
            <surname>Boughanem</surname>
          </string-name>
          et al., eds., Berlin, Heidelberg: Springer Berlin Heidelberg.
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          18.
          <string-name>
            <surname>Paltoglou</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          , Salampasis.,
          <string-name>
            <given-names>M</given-names>
            &amp;
            <surname>Satratzemi</surname>
          </string-name>
          ,
          <string-name>
            <surname>M.</surname>
          </string-name>
          ,
          <year>2011</year>
          .
          <article-title>Modeling information sources as integrals for effective and efficient source selection</article-title>
          .
          <source>Information Processing &amp; Management</source>
          ,
          <volume>47</volume>
          (
          <issue>1</issue>
          ), pp.
          <fpage>18</fpage>
          -
          <lpage>36</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          19.
          <string-name>
            <surname>Raghavan</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          &amp;
          <string-name>
            <surname>Garcia-Molina</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <year>2001</year>
          .
          <article-title>Crawling the Hidden Web</article-title>
          .
          <source>In Proceedings of the International Conference on Very Large Data Bases. Citeseer</source>
          , pp.
          <fpage>129</fpage>
          -
          <lpage>138</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          20.
          <string-name>
            <surname>Van Rijsbergen</surname>
            ,
            <given-names>C.J.</given-names>
          </string-name>
          ,
          <year>1979</year>
          . Information Retrieval, Butterworths.
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          21.
          <string-name>
            <surname>Salampasis</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Paltoglou</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          &amp;
          <string-name>
            <surname>Giahanou</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <year>2012</year>
          .
          <article-title>Report on the CLEF-IP 2012 Experiments: Search of Topically Organized Patents</article-title>
          . In P. Forner,
          <string-name>
            <given-names>J.</given-names>
            <surname>Karlgren</surname>
          </string-name>
          , &amp;
          <string-name>
            <surname>C.</surname>
          </string-name>
          Womser-Hacker, eds.
          <source>CLEF</source>
          (Online Working Notes/Labs/Workshop).
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          22.
          <string-name>
            <surname>Si</surname>
          </string-name>
          , L et al.,
          <year>2002</year>
          .
          <article-title>A language modeling framework for resource selection and results merging</article-title>
          .
          <source>In ACM CIKM 02</source>
          . ACM Press, pp.
          <fpage>391</fpage>
          -
          <lpage>397</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          23.
          <string-name>
            <surname>Si</surname>
            , Luo &amp; Callan,
            <given-names>J.</given-names>
          </string-name>
          ,
          <year>2003a</year>
          .
          <article-title>A semisupervised learning method to merge search engine results</article-title>
          .
          <source>ACM Transactions on Information Systems</source>
          ,
          <volume>21</volume>
          (
          <issue>4</issue>
          ), pp.
          <fpage>457</fpage>
          -
          <lpage>491</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          24.
          <string-name>
            <surname>Si</surname>
            , Luo &amp; Callan,
            <given-names>J.</given-names>
          </string-name>
          ,
          <year>2003b</year>
          .
          <article-title>Relevant document distribution estimation method for resource selection</article-title>
          .
          <source>In Proceedings of the 26th annual international ACM SIGIR conference on Research and development in informaion retrieval - SIGIR '03</source>
          . Toronto, Canada: ACM New York, NY, USA, pp.
          <fpage>298</fpage>
          -
          <lpage>305</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref25">
        <mixed-citation>
          25.
          <string-name>
            <surname>Willett</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <year>1988</year>
          .
          <article-title>Recent trends in hierarchic document clustering: A critical review</article-title>
          .
          <source>Information Processing &amp; Management</source>
          ,
          <volume>24</volume>
          (
          <issue>5</issue>
          ), pp.
          <fpage>577</fpage>
          -
          <lpage>597</lpage>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>