<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>CLEF-IP 2011: Retrieval in the Intellectual Property Domain</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Florina Piroi</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Mihai Lupu</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Allan Hanbury</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Veronika Zenz</string-name>
          <email>veronika.zenz@max-recall.com</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Vienna University of Technology Institute for Software Technology and Interactive Systems Favoritenstr.</institution>
          <addr-line>911/188, 1040 Vienna</addr-line>
          ,
          <country country="AT">Austria</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>max-recall Information Systems OG Pulverturmgasse 17/3</institution>
          ,
          <addr-line>1090 Vienna</addr-line>
          ,
          <country country="AT">Austria</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>The patent system is designed to encourage disclosure of new technologies and novel ideas by granting exclusive rights on the use of inventions to their inventors, for a limited period of time. Before a patent can be granted, patent oces around the world perform thorough searches to ensure that no previous similar disclosures were made. In the intellectual property terminology, such kind of searches are called prior art searches. In some industries, the number of granted patents a company owns has a high impact on the market value of the company. This underlines the importance of well-performed prior art searches. Together with the TrecChem track [5], also organized by our institution, the ClefIp eort comes to complete the work that is being done in the series of Ntcir workshops (see for example [ 4]). The rst ClefIp track ran within Clef 20091. The purpose of the track was twofold: to encourage and facilitate research in the area of patent retrieval by providing a large clean data set for experimentation; to create a large test collection of patents in the three main European languages for the evaluation of crosslingual information access. The ClefIp data set includes documents published by the European Patent Oce (Epo) which contain a mixture of English, German and French content. The track focused on the task of prior art search. In 2010 and 2011, the ClefIp track was organized as a benchmarking activity (lab) in the Clef conference. In these years, the main goal of the ClefIp eort remained the same to foster research in the patent retrieval area, and provide a large clean data set. To this end, the number of tasks in the track was increased and the data set was enlarged. Recognizing the importance of patent classications in the daily activity of an intellectual property professional, in 2010 the ClefIp benchmarking activity included a patent classication task. The participants were asked to classify 1 http://www.clef-campaign.org</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Introduction</title>
      <p>
        given patent application documents according to the International Patent
Classication ( Ipc) system [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ].
      </p>
      <p>This year (2011), in addition to the (now classic) prior art search task and
the patent classication task two patent image related tasks were oered. Often,
patent applications contain images that clarify details about the invention they
describe. Images in patents may be drawn by hand, by computer, or both, may
contain text, and are generally black-and-white (i.e. not even monochrome).
Depending on the technological area of a patent, images may be technical drawings
of a mechanical component, or an electric component, owcharts if the patent
describes, for example, a workow, chemical structures, tables, etc. When a
patent expert browses through a list of search results given by a search engine,
he or she can very quickly dismiss irrelevant patents to the patent application by
just glancing at the images in the retrieved patents. The number of documents to
be looked at in more detail is thus greatly reduced. With the Image-based
Document Retrieval and Image-based Classication tasks in ClefIp we try to make
this aspect of an IP expert’s daily work familiar to the research communities
12 international teams have participated this year, we present here an overview
their work and research results. The paper is structured as follows: Section 2
describes the test collection used this year, section 3 presents the participating
teams and gives an overview of the methods the teams involved. In the same
section we also present the main measurements done in this track.
2
2.1</p>
    </sec>
    <sec id="sec-2">
      <title>The 2011 ClefIp</title>
    </sec>
    <sec id="sec-3">
      <title>Collection</title>
      <sec id="sec-3-1">
        <title>The Documents in the Collection Corpus</title>
        <p>
          The ClefIp collection contains patents, physically stored as a collection of
Xml les encoding patent documents. A patent document may be an application
document, a search report, or a granted patent document. Each patent document
is assigned a kind code, which appears as a sux to the patent identier (e.g.
EPnnnnnnn-A1, WO-nnnnnnnnnn-A2). In the case of the Epo, patent application
documents that include a search report carry the code A1, patent application
search reports carry the code A3, granted patent documents carry the code B1,
etc.2 For a description of key terms and steps in a patent’s lifecycle see [
          <xref ref-type="bibr" rid="ref6">6</xref>
          ].
        </p>
        <p>
          An important tool in organizing the large amount of patent data which patent
oces regulate is the classication system . A patent classication system ‘sorts’
the patents according to the technical area they belong to, and it is a basis for a
quick investigation of the state of the art in a eld [
          <xref ref-type="bibr" rid="ref1 ref6">1,6</xref>
          ]. The mostly used patent
classication systems are the International Patent Classication system ( Ipc),
the European Classication System ( Ecla), the US Classication System. In the
ClefIp lab, the patent classication tasks make use only of the Ipc system.
2 A list of kind Epo kind codes is listed at https://register.epo.org/espacenet/
help?topic=kindcodes.
        </p>
        <p>Kind codes used by the Wipo are listed at http://www.wipo.int/patentscope/en/
wo_publication_information/kind_codes.html</p>
        <p>The 2011 ClefIp data collection is based on the 2010 data, and is extracted
from the Marec3 data corpus. The ClefIp collection contains mainly patent
documents published by the Epo.</p>
        <p>
          Two important additions were made to the 2010 collection corpus. The rst
one was to include in the distributed corpus certain patents published by the
World Intellectual Patent Organization ( Wipo). A high percentage of the Epo
patents contained in the ClefIp corpus are patent applications internationally
led under the Patent Cooperation Treaty ( Pct [
          <xref ref-type="bibr" rid="ref2">2</xref>
          ]) in which case, the Epo
does not republish the whole patent application, but only a bibliographic entry
referring to the original application. For these patents we added their Wipo
equivalent to the ClefIp collection, in order to provide the participants a
collection that is both larger and more realistic. The second addition to the
ClefIp collection regards one of the new patent imagebased tasks, the Patent
imagebased retrieval task. For this task we have added to the ClefIp target
collection the patent images for three Ipc subclasses: A43B, A61B, and H01L.
        </p>
        <p>In number of documents, adding the Wipo patent documents to the
collection corpus increased with 1.2 million patent documents, to a nal number of
approximately 3.5 million Xml documents, referring to approximately 1.5
million patents. The images corresponding to the 47 thousands Xml documents in
the three Ipc subclasses in the ImgPac task occupy 5.4 Gb for 290880 ti les.</p>
        <p>The same as in the previous years, the test collection corpus was delivered
to the participants as is, without merging the documents related to the same
patent into one document. Each patent is identied by a unique patent number
a string identifying the publishing oce (EP for the Epo and WO for the
Wipo) followed by a series of digits. Corresponding to each patent is a directory
containing the Xml documents representing the patent documents related to
that patent. For the EP patents, the layout is 00000n/nn/nn/nn/*.xml, where
the sequence of digits in the directory layout corresponds to the one in the patent
number. For the WO patents, the layout is 00nnnn/nn/nn/nn/*.xml, where the
rst 4 digits (after ‘00’) represent the document publication year.</p>
        <p>For example, to patent EP 0981201 corresponds the directory containing les
EP-0981201-A2.xml, EP-0981201-A3.xml, and EP-0981201-B1.xml:
&gt; pwd
EP/000000/98/12/01
&gt; ls
EP-0981201-A2.xml EP-0981201-A3.xml EP-0981201-B1.xml
To patent WO 1994030029 corresponds the directory containing the le
WO1994030029-A1.xml:
&gt; pwd
WO/001994/03/00/29
&gt; ls
WO-1994030029-A1.xml
3 The Marec data corpus is a collection of over 19 million patent documents, in Xml
format, made available by the Irf for research purposes.</p>
        <p>The patent image les are stored as tif les in one separate folder, the
correspondence between the image le and its Xml le is established with a small set
of rules.</p>
        <p>All textual documents in the ClefIp collection contain the following main
Xml elds: bibliographic data, abstract, description, and claims. Not all
documents actually have content in these elds. The content of the various Xml elds
can be English, German, or French. Some elds may occur more than once, each
time with a dierent language. The Xml patent le have also a document
language (English, German or French), this not excluding that its subelds occur
with a dierent language attribute than the document language. For example,
granted EP patent documents (EP-nnnnnnn-Bn.xml) must contain claims in
three languages (English, German and French).
2.2</p>
        <p>Tasks and Topics
5 tasks were proposed to the participants: Prior Art Candidates Search ( Pac),
Patent Classication ( Cls1), Rened Patent Classication ( Cls2) Patent
Imagebased Retrieval (ImgPac ), and Patent Image-based Classication ( ImgCls ).
The topics for each of the tasks were chosen from the same topic pool we have
used in 2010. We will detail each of the proposed tasks in the following.
Prior Art Candidates Search. The rst task in this track ( Pac) consisted
in nding patent documents in the target collection that may invalidate a given
patent application. The participants were provided with one set of 3973 topics.
The topics were formulated as ‘Find all patents in the collection that potentially
invalidate patent application EP-nnnnnnn-An.’, where the Xml le storing the
patent application document EP-nnnnnnn-An (which we call the topic le or
topic patent ) was given in an attached archive. A third of the topic les’
document language was English, another third was German, and the last third was
French. The task did not restrict the language used for retrieving the
documents, but participants were encouraged to use the multilingual characteristic
of the collection (namely, that claims in granted patent documents are provided
in three languages). A small set of 300 training topics was also provided, and
participants were allowed to use the ClefIp 2010 topics sets in their system
training.</p>
        <p>Patent Classication. The second task in the ClefIp track (Cls1) required
to classify a given patent document according to the Ipc system. the topics
were formulated as ‘Classify patent document EP-nnnnnnn-An according to the
Ipc system.’, where the Xml le storing the patent application document
EPnnnnnnn-An was given in an attachment. The Ipc system is hierarchically
organized into sections, classes, subclasses, groups and subgroups. The classication
was to be given at the subclass level. The set of classication topics contained
3000 patent documents, a dierent set than the one used in the Pac task. Again,
the task didn’t restrict the language used for classication, but the topic language
was English for one third of the topics, German for the next third, and French
for the last third of the topics. Participants could use the ClefIp collection
corpus as a training set.</p>
        <p>Rened Patent Classication. This task is, at least in formulation, very
similar to the Cls1 task. It required the participants to classify given patent
documents according to the Ipc system. The subclass was given, the participants
had to return the group/subgroup classication levels.</p>
        <p>Patent Image-based Prior Art Search. This task (ImgPac ) was
introduced as a pilot task. It has the same aim as the Pac task, except that the
images and Xml les corresponding to the patents were available and were to be
used. Only the three Ipc subclasses listed above were used, as for these classes,
patent searchers often rely on visual comparison of images in the patents to nd
relevant prior art. The queries consisted of the text and complete set of images
of 211 patents, with the topics formulated in the same way as for the Pac task.
Patent Image-based Classication. The aim of the image classication task
(ImgCls ) was to automatically classify patent images based on visual content.
For the ImgCls task, only images extracted from patents, not full patents, were
provided. Participants to this task did not need the full ClefIp corpus. The
classication was into nine classes: drawing, chemical structure, program
listing, gene sequence, ow chart, graph, mathematics, table, and symbol. Training
data with between 300 and 6,000 training images for each of these classes was
provided, and only these data were to be used to train image type classication
techniques. The task was to train a classier using the provided training data,
and test the resulting classier on a set of 1,000 patent images.
2.3</p>
        <p>
          Relevance Assessments
The relevance assessments used to evaluate the Pac and ImgPac submissions
were obtained automatically from the patent citations stored in the collection
documents. Since the average number of citations per patent in the ClefIp
collection is low (below 4), we have looked for methods to extend the set of
relevant documents per topic. For this we used an extended list of citations,
where to the patents listed in the patent’s search report (the direct citations), we
added also the patent citations listed in the family members of the topic patent,
as well as the family members of the cited patents. For detailed explanations of
the citation extraction procedure, we point the reader to the overview article [
          <xref ref-type="bibr" rid="ref7">7</xref>
          ].
        </p>
        <p>The relevance assessments used to evaluate the Cls1 and Cls2 submissions
were also obtained automatically from the documents that originated the Cls
topics. We have extracted the Ipc codes, restricted to the subclass level, and
group/subgroup level respectively, from the patent documents.</p>
        <p>The relevance assessments for the ImgCls topics were done manually by
us.</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>Submissions and Results</title>
      <p>
        With the exception of the ImgCls task, a submission consisted of a single text
le with at most 1,000 answers per topic, in the standard format used for the
Trec submissions. The submissions to the ImgCls task consisted of a single
text le with one entry per topic, each entry containing the topic id and nine
space separated values one for each class used in the classication. 12
participating groups have submitted a total of 77 runs, (unequally) distributed over the
ve proposed task. Table 1 shows the list of participating groups, marking the
tasks where runs were submitted. The submissions were uploaded to the Direct
system [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]4.
ca sl
ca sl1 sl2 gP gC
      </p>
      <p>P C C Im Im
DE x
DE x
RU x
IN x
AT
NL x
AT
AT x
CH x
NL x x x
KR x x x
FR
This section is based on the descriptions provided by the participants. We present
here which Xml elds were used in document processing, what kind of pre and
postprocessing was done, the retrieval and ranking system that was used to
4 The Direct system is currently developed under the Promise project
obtain the results, crosslanguage techniques involved, as well as any other
relevant details.
⋆ The hildesheim and chemnitz groups collaborated in try to identify how
patent IR is aected by the query length and the use of linguistics in the
retrieval process. Their method was developed on top of the Xtrieval framework,
which provides a common interface to several search engines. In preprocessing,
they used a specic stopword list, especially created for the patent domain.
Subsequently, each language was indexed separately. In addition to the text, the
Ipc classes were also added to the index. Like other groups this year, they
extracted dierent types of phrases in order to improve the precision of the results.
However, their method was not purely statistical, but also used a rule based
dependency parser.</p>
      <p>The search process considered three types of queries: term-based,
phrasesbased and a combination of the two. They also compared very long queries with
short (and precise) queries. Their results would indicate that using long queries
is better than using small precise ones. It is arguable however, how indicative of
the content of the given patent application a short query can be. Furthermore,
they show that using linguistic phrases did not increase the eectiveness of the
retrieval system.
⋆ At the time of writing this text, we have no information on how the experiments
submitted by the hprussia participant were obtained.
⋆ The team from Hyderabad combines three methods in order to retrieve and
rank prior art. At the base, their method relies on the Lemur toolkit and the
translation of queries into English. First, they apply a key phrase extraction
method in order to create queries from the topic patent documents. Then, they
identify references to other patens within the text of the document and use them
in two ways: rst, to create a document vector based on the Ipc codes assigned
to the current patent application and to all the referenced patents. Second, to
add them directly to the result list.</p>
      <p>Without using this citation list, the best results appear to be those
using both the Ipc information and the text search. When using citations, the
Ipcinformation seems to reduce the quality of the results.
⋆ The joanneum group participated in the ImgCls task. They use features
such as Local Binary Patterns (LBP), MPEG-7 features as well as OCR’d text
extracted from the images. Support Vector Machines are used for the
classication. Runs either use a single type of feature, or combine the classication results
of dierent types of features using late fusion. The best run is the one using the
LBP only.
⋆ The lugano and tuwien-2 teams took up on the Pac task by focusing on
generating patent summaries that improve the query formulation. Topic patent
documents are summarized by the PatTextTiling technique, which is an
adaptation of the TextTiling summarization algorithm. The topics’ Ipc codes are
used to dene a relevance set (corpus documents that share Ipc codes with the
topic) which is used to get closer to a better relevance model. Query models
can be built summary-based or description-based (from the description
document sections). After removing stop-words and stemming, the patent document
corpus was indexed using Terrier. The ranking model used is BM25. The
experiments are ltered by the topic’s Ipc code and patent citations extracted from
the topic’s text are added to the results. Only English topics were considered,
for this reason the plot values for the German and French languages are missing
in Figure 3.
⋆ In its approach to tackle the Patent Classication tasks, the nijgmenen team
use the Linguistic Classication System ( Lcs) to implement three classiers:
Naive Bayes, Winnow, and SVM lihgt. Various experimental settings were
compared to gradually improve classication results. Among such settings are the use
of dierent document sections, of patent citations, dierent document
representations, and dierent training data sets. Parameter tuning during training also
contributed to better classication results. In the training phase only English
abstracts and descriptions were used. As metadata the Ipc codes, and the
applicant, inventor, and address elds were extracted. It turned out that applicant,
inventor and address information did not contribute much to the classication
results. To the abstracts and (rst 400 words of) the description thus extracted
the Aegir dependency parser was applied, its results being added to the
document representations. Patent citations in the topic les were used to rerank
the classication results.</p>
      <p>For the Prior Art Candidates search task, the nijgmenen group teamed up
with the spinque group, focusing on using bag-of-words approach enriched with
syntacticsemantic information. Only the English content of the titles, abstracts,
claims and descriptions was used. (For this reason, Figure 7 doesn’t have plot
entries for the nijgmenen experiments.) In a separate processing step, the Ipc
information and the rst 400 words of the description were extracted. The
extracted English content was cleaned up by removing image references and claim
headers, sentenced, and parsed with Aegir. Both the corpus and the topics were
processed in this way. Selecting the query terms was done based on their
relevance to Ipc classes, relevance computed with the ( Lcs) software. Finally, the
retrieval was done using with the Spinque framework, which allows the denition
of search strategies via a graphical user interface.
⋆ The wisenut group participated both in the Prior Art Candidate search and
Classication tasks, although the Pac participation was due to implementing a
KNN-like classication using Pac search results. This solution was chosen based
on the experience that training and classifying documents suer from having
few training documents, not enough memory for the (usually) large
classication model, and from using too much processing time. To obtain the Pac search
results, the Xml les were processed in the following way: weighted keywords
were selected out of the title, abstract, description, and claims eld. Then, after
Pos tagging, functional words, stop words and noncontent words are removed
and cooccurrence terms are added. Terms and text content in German and
French were translated into English using the MyMemory translation service 5.
The system used to index, search and classify is based on Lucene, to which a
simple Java-based interface was implemented. The interface included also an
application for classication that applied a improved weighting scheme in the KNN
classication to obtain the Cls1 and Cls2 results.
⋆ The xerox-sas group participated in both the ImgPac and ImgCls tasks.
Images for both tasks are represented using Fisher vectors. For the ImgCls
task, linear classiers are used. The best results are obtained for the run in
which the images were articially rotated and added to the training set, to take
into account that the images are sometimes rotated. For the ImgPac task,
different strategies are investigated to compare one set of images to another (as
patents consisting of a group of images, not single images, are to be retrieved).
Runs are also created in which the predicted ImgCls image classes are used.
For the retrieval of the patent text, dierent sections of the patent are weighted
dierently. Similarities are also calculated based on shared Ipc categories and
similarities of the patent citation graph, with late fusion used to combine the
similarities. A weighted late fusion strategy was again used to combine the text and
image ranking, with the text rankings weighted higher than the image rankings.
While the visual retrieval performs poorly, when combined with text retrieval it
outperforms the text-only retrieval.
3.2</p>
      <p>Evaluation Results
We have evaluated the submitted experiments using the most common metrics in
IR. Before we ran the evaluation software, some simple cleanup of the data was
done. This included replacing whitespace sequences with only one blank space,
ltering out experiment entries which did not belong to a given topic patent
documents.</p>
      <p>For each submitted Pac experiment we computed the following measures:
For each submitted Cls experiment we computed the following measures:</p>
      <p>All computations were done using the trec_eval 9.0 software provided by
Nist. Figures 1 through 8 show some of the calculated measures. Detailed values
for each of the mentioned measures were sent to the lab participants and are soon
to be published into a technical report.
5 MyMemory, http://mymemory.translated.net
0.1
0.09
0.08
0.07
0.06
0.05
0.04
0.03
0.02
0.01
0
map</p>
      <p>P
P at 100</p>
      <p>The gures below use a shortened version of the original run names, whose
length would have made the pictures less legible. The mapping between the
original run name and the shortened versions is shown in the appendix. Note
that, although run names might be the same, the experiment les are dierent
between tasks. For each submitted ImgCls experiment, we computed Equal
.y5H .3hC .h2C .u1L .Lu2 .Lu3 .Lu4 .h7C .h1C i.3N i.1N .4hC .y3H i.7W i.6W i.5W i.8W i.4W i.3W i.2W i.1W .y4H .y6H .y2H i.2N .y1H .h6C i.4N .h5C .1pH
Error Rate (EER) and Area Under Curve (AUC) of a ROC curve, and True
Positive Rate (TPR) per class averaged over all classes, as well as confusion
matrices. This was done using a custom-written script running in Octave 6. The
results of all runs are summarized in Table 2.
4</p>
    </sec>
    <sec id="sec-5">
      <title>Final Observations</title>
      <p>We have given an account of the benchmarking activities done in the frame
of the ClefIp 2011. Compared to the last year, collaborations between
research groups has intensied. Another positive observation is that participants
are drawing on the research results obtain in the previous years to improve their
retrieval methods. The coagulation of research groups leads to a consolidation
of the methods used for patent retrieval and allows them to reach a maturity</p>
      <sec id="sec-5-1">
        <title>6 http://www.gnu.org/software/octave/</title>
        <p>Rat100
all
en
de
fr
0.7
0.6
0.5
0.4
0.3
0.2
0.1
0
0.11
0.1
0.09
0.08
0.07
0.06
0.05
0.04
0.03
0.02
0.01
0</p>
        <p>C H W W W W W W W W H H H H C C H</p>
        <p>Fig. 3. Map measures per languages for the Pac runs
Pat1
Pat5
Rat5
F1at5
0.9
0.8
0.7
0.6
0.5
0.4
0.3
0.2
0.1</p>
        <p>0
0.9
0.8
0.7
0.6</p>
      </sec>
      <sec id="sec-5-2">
        <title>Fig.5. Measuresforthe Cls1runs</title>
        <p>0.25</p>
      </sec>
      <sec id="sec-5-3">
        <title>Fig.7. Measuresbylanguageforthe Cls1runs</title>
        <p>all
en
de
fr
0.35
0.3
0.25
0.2
0.15
0.1
0.05
joanneum.alphacentauri 0.15 0.91 0.66
joanneum.arcturus 0.24 0.81 0.50
joanneum.betelgeuse 0.18 0.90 0.62
joanneum.canopus 0.16 0.91 0.65
joanneum.procyon 0.37 0.67 0.27
joanneum.rigel 0.16 0.90 0.63
joanneum.sirius 0.16 0.91 0.64
joanneum.vega 0.32 0.72 0.28
xerox-sas.RUNORH 0.06 0.98 0.85
xerox-sas.RUNORH_ROTRAIN 0.04 0.99 0.91
xerox-sas.FV_ORH_SP 0.08 0.92 0.85
xerox-sas.MEAN_ALL 0.08 0.91 0.85</p>
        <p>Table 2. Summary of the ImgCls run results.
level which would make them a candidate for commercial exploitation. However,
it also reduces the diversity of the submissions. The number of participating
groups was lower this year and at the workshop we need to explore the ways in
which the evaluation of patent retrieval tools needs to go ahead.</p>
        <p>One such was was identied as of last year. Image retrieval is extremely
important for many technologies patented. However, participation in the image
tasks was low in the rst. The ImgPac pilot task is challenging due to the
multimodal nature of the retrieval task, the large amount of data and the full
patent retrieval (containing a set of images) as opposed to single image retrieval.
It is hoped that this task can lead to more collaboration between image and text
retrieval groups in the next years. The ImgCls task was planned so as to have
a lower threshold of entry for groups with image classication expertise. While
6 groups registered to obtain the data, only 2 participated.</p>
        <p>Acknowledgements This work was partly supported by the EU Network
of Excellence PROMISE 7 (FP7-258191) and the Austrian Research Promotion
Agency (FFG) FIT-IT project IMPEx 8 (No. 825846).</p>
      </sec>
      <sec id="sec-5-4">
        <title>7 http://www.promise-noe.eu 8 http://www.joanneum.at/?id=3922</title>
        <p>Original id
NIJMEGEN.RUN_ADMWCIT
NIJMEGEN.RUN_ADMWOWCIT
NIJMEGEN.RUN_ADMWOW
NIJMEGEN.RUN_ADMW
NIJMEGEN.RUN_ADWOWCIT
NIJMEGEN.RUN_ADWOW
NIJMEGEN.RUN_ADWTCIT
NIJMEGEN.RUN_ADWT
WISENUT.WISENUT_R1_BASE
WISENUT.WISENUT_R2_BASE_10
WISENUT.WISENUT_R3_BASE_20
WISENUT.WISENUT_R4_BASE_30
WISENUT.WISENUT_R5_CO
WISENUT.WISENUT_R6_CO_10
WISENUT.WISENUT_R7_CO_20
WISENUT.WISENUT_R8_CO_30</p>
      </sec>
      <sec id="sec-5-5">
        <title>Original id Short</title>
        <p>id
NIJMEGEN.RUN_WINNOW_WORDS Ni.1
WISENUT.WISENUT_R1_BASE Wi.1
WISENUT.WISENUT_R2_BASE_10 Wi.2
WISENUT.WISENUT_R3_BASE_20 Wi.3
WISENUT.WISENUT_R4_BASE_30 Wi.4
WISENUT.WISENUT_R5_CO Wi.5
WISENUT.WISENUT_R6_CO_10 Wi.6
WISENUT.WISENUT_R7_CO_20 Wi.7
WISENUT.WISENUT_R8_CO_30 Wi.8</p>
      </sec>
      <sec id="sec-5-6">
        <title>Original id Short</title>
        <p>id
XEROX-SAS.3MAX3MEAN Xe.1
XEROX-SAS.3MAX3MEAN_LATEMONO Xe.2
XEROX-SAS.3MAX3MEAN_MT_CIT Xe.3
XEROX-SAS.3MAX_LATEMONO Xe.4
XEROX-SAS.3MAX_MT Xe.5
XEROX-SAS.FVORH_3MAX Xe.6
XEROX-SAS.FVORH_3MAX3MEAN Xe.7
XEROX-SAS.MAXMEANMODAD Xe.8
XEROX-SAS.MAXMEANMODAD_MT Xe.9 Table 6. Original and short run
XEROX-SAS.MAX_MT_CIT Xe.10 names for the Pac task</p>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <given-names>International</given-names>
            <surname>Patent</surname>
          </string-name>
          <article-title>Classication (IPC)</article-title>
          . http://www.wipo.int/ classifications/ipc/en/.
          <source>last retrieved: August</source>
          ,
          <year>2011</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <given-names>Patent</given-names>
            <surname>Cooperation Treaty (PCT).</surname>
          </string-name>
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <given-names>Marco</given-names>
            <surname>Dussin</surname>
          </string-name>
          and
          <string-name>
            <given-names>Nicola</given-names>
            <surname>Ferro</surname>
          </string-name>
          .
          <article-title>DIRECT: Applying the DIKW Hierarchy to LargeScale Evaluation Campaigns</article-title>
          .
          <source>Bulletin of IEEE Technical Committee on Digital Libraries (IEEE-TCDL)</source>
          ,
          <volume>5</volume>
          (
          <issue>1</issue>
          ),
          <year>2009</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <given-names>Atsushi</given-names>
            <surname>Fujii</surname>
          </string-name>
          , Makoto Iwayama, and
          <string-name>
            <given-names>Noriko</given-names>
            <surname>Kando</surname>
          </string-name>
          .
          <article-title>Overview of the Patent Retrieval Task at the NTCIR-6 Workshop</article-title>
          . In Noriko Kando and David Kirk Evans, editors,
          <source>Proceedings of the Sixth NTCIR Workshop Meeting on Evaluation of Information Access Technologies: Information Retrieval</source>
          , Question Answering, and CrossLingual Information Access , pages
          <fpage>359365</fpage>
          ,
          <fpage>2</fpage>
          -1-2 Hitotsubashi, Chiyoda-ku, Tokyo 101-8430, Japan, May
          <year>2007</year>
          . National Institute of Informatics.
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <given-names>M.</given-names>
            <surname>Lupu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Huang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Zhu</surname>
          </string-name>
          , and
          <string-name>
            <given-names>J.</given-names>
            <surname>Tait</surname>
          </string-name>
          .
          <article-title>TREC-CHEM: large scale chemical information retrieval evaluation at TREC</article-title>
          .
          <source>SIGIR Forum</source>
          ,
          <volume>43</volume>
          (
          <issue>2</issue>
          ),
          <year>December 2009</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <given-names>F.</given-names>
            <surname>Piroi</surname>
          </string-name>
          and
          <string-name>
            <surname>J. Tait. CLEF-IP</surname>
          </string-name>
          <year>2010</year>
          :
          <article-title>Retrieval experiments in the intellectual property domain</article-title>
          .
          <source>Technical Report IRFTR201000005</source>
          , Information Retrieval Facility, Vienna,
          <year>September 2010</year>
          .
          <article-title>Also available as a Notebook Paper of the CLEF 2010 Informal Proceedings</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <given-names>G.</given-names>
            <surname>Roda</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Tait</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Piroi</surname>
          </string-name>
          , and
          <string-name>
            <given-names>V.</given-names>
            <surname>Zenz.</surname>
          </string-name>
          CLEF-IP
          <year>2009</year>
          :
          <article-title>Retrieval Experiments in the Intellectual Property Domain</article-title>
          . To appear.
          <source>In Proc. of CLEF, Revised Selected Papers</source>
          . Springer,
          <year>2010</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>