<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Improving Patent Classification using AI-Generated Summaries</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Naoya Yoshikawa</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Ralf Krestel</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Japan Patent Ofice</institution>
          ,
          <addr-line>3-4-3 Kasumigaseki, Chiyoda-ku, 100-8915, Tokyo</addr-line>
          ,
          <country country="JP">Japan</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>ZBW - Leibniz Information Centre for Economics &amp; Kiel University</institution>
          ,
          <addr-line>Düsternbrooker Weg 120, 24105, Kiel</addr-line>
          ,
          <country country="DE">Germany</country>
        </aff>
      </contrib-group>
      <fpage>8</fpage>
      <lpage>17</lpage>
      <abstract>
        <p>This study investigates the efectiveness of using summaries generated by large language models (AI-generated summaries) to improve the performance of automatic patent classification. We propose a novel approach to use AI-generated summaries of patent text fields (abstract, claims and detailed description) as training data for classification models: using two patent datasets, USPTO-70k and CLEF-IP 2011, we perform experiments focused on both subclass-level multi-label classification and subgroup-level multi-class classification tasks. The results show that models trained on AI-generated summaries of claims and detailed descriptions achieve significantly higher scores than models trained on the original text. This result suggests that AI-generated summaries efectively extract information relevant to patent classification and contributes to the development of automatic patent classification technology.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;Patent Classification</kwd>
        <kwd>Large Language Model</kwd>
        <kwd>Patent Summarization</kwd>
        <kwd>Subgroup Classification</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>
        ifcation task, and supervised machine learning models
using large amounts of classified patent text data
pubPatent classifications play an important role in the accu- lished by WIPO and national patent ofices have been
rate and eficient management of patent information. In widely adopted. One of the elements that can improve
addition to the International Patent Classification (IPC) the performance of these models is the availability of
maintained by the World Intellectual Property Organiza- high-quality and suficient training data. Patent
docution (WIPO), patent ofices around the world maintain ments have a variety of text fields, such as title, abstract,
their own patent classification systems, including the claims, and detailed description, which contain rich
texCooperative Patent Classification (CPC) of the European tual information, but also much information that is not
Patent Ofice (EPO) and the United States Patent and relevant to the classification task. Since it is dificult to
Trademark Ofice (USPTO), and the File Index (FI) and manually extract high-quality textual information from
File Forming Term (F-term) of the Japan Patent Ofice a large number of patent documents, training data is
(JPO). The IPC is an internationally uniform classifica- usually created by mechanically extracting textual
intion with a hierarchical structure of sections (e.g. "A"), formation from the text fields of patent documents. In
classes (e.g. "A61"), subclasses (e.g. "A61B"), main groups previous studies, textual information such as title and
(e.g. "A61B17") and subgroups (e.g. "A61B17/29"). Patent abstract, claims, and the first few hundred words of the
professionals assign IPC codes to patent applications at detailed description have been used [
        <xref ref-type="bibr" rid="ref2 ref3 ref4 ref5">2, 3, 4, 5</xref>
        ].
the subgroup-level according to IPC guidelines [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. This Moreover, mamy previous studies have focused on
assignment process requires expertise and is complex and classification tasks at subclass-level, rather than at the
costly. Recent advances in natural language processing subgroup-level for the IPC and CPC. Because subgroups
and deep learning have shown excellent results in vari- are subdivided into approximately 70,000 labels for IPC
ous classification tasks. However, the diversity of patent and 240,000 labels for CPC, and patent classification is
classifications, evolving technology levels, inconsistent a multi-label classification task where multiple labels
ifeld-specific terms in patent documents, and the multi- can be assigned, accurate automatic classification at the
label nature in which multiple labels can be assigned to subgroup-level is very dificult. Subgroups are
charactera single application still make it challenging to automate ized by specific parts of the subject matter covered by the
patent classification accurately. upper main group or subclass [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. As a result, the title
There are a number of approaches to the patent classi- and abstract of a patent document, for example, may not
contain suficient information about the subgroups.
5th Workshop on Patent Text Mining and Semantic Technologies Since the introduction of the transformer architecture,
(PatentSemTech) 2024 one of the main approaches has been fine-tuning
prer$kry@oisnhfiokramwaat-ikn.auonyia-k@iejpl.doe.g(oR.j.pK(rNe.stYeol)shikawa); trained models using task-specific training data. Many
0000-0002-5036-8589 (R. Krestel) previous studies have shown good performance using
© 2024 Copyright for this paper by its authors. Use permitted under Creative Commons License transformer-based models. However, these studies used
Attribution 4.0 International (CC BY 4.0).
training data that was simply extracted, such as the first pre-trained language models to investigate multi-label
few hundred words of each text field, and the quality patent classification performance for patent text. The
of the training data is debatable. Ideally, the input text first 128 words of each text field were compared as input
should be no more than a few hundred words, containing text, and the combination of title and abstract showed
important information relevant to the assigned classifi- good performance. They also noted that longer sentences
cation. should be considered when using detailed patent
descrip
      </p>
      <p>
        Large Language Models (LLMs), which are pre-trained tions and claims. Pujari et al. [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ] addressed the issue of
on large amounts of text data, have shown high perfor- limited input text size for neural models based on
premance in a variety of natural language processing tasks. trained transformers. They proposed a new approach
Using LLMs to summarize text has the potential to ex- to efectively integrate information obtained from
multract information that is highly relevant to classification tiple text fields. They published the USPTO-70k dataset
tasks. extended to include claims, detailed description,
brief
      </p>
      <p>
        The objective of this study is to compare the perfor- summary, and figure description, and in particular, they
mance of models trained on summaries of patent docu- found that the brief-summary text field in the US patent
ments generated by LLMs with the performance of mod- document is the most useful for CPC subclass-level
clasels trained on the original patent text. We also evalu- sification. Yadrintsev et al. [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ] compared KNN (k-nearest
ate the performance of subclass-level and subgroup-level neighbors) and fastText as IPC subclass-level classifiers
classification tasks; text summarization with LLMs can be using CLEF-IP competition data. In the CLEF English test
a promising approach to improve the quality of training sample (1000 documents), the micro-averaged F1-score
data for patent classification tasks. By generating con- were 71.0 for KNN and 70.4 for fastText.
cise summaries from lengthy text of patent documents All of the above studies were conducted at the IPC or
that contain important information relevant to classifi- CPC subclass-level; there are not many reports of
studcation, the training eficiency and performance of the ies on automatic classification at the IPC subgroup-level.
model can be improved. Furthermore, the summarized Hoshino et al. [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ] proposed a new decoder architecture
text would be of a suitable length to serve as input for the that takes into account the hierarchical structure of IPC
transformer-based model, thus capturing the meaning of and a model that considers the content of all claims by
the entire text while reducing computational cost. extracting important information from the claims. The
model showed a significant improvement in accuracy
compared to previous methods, especially in
subgroup2. Related Work level prediction. They extracted nouns and their
proportions from the claims as input text, but suggest that other
Our goal is to classify patents based on automatically methods of information compression should be
considgenerated summaries (AI-generated summaries). There- ered. Zuo et al. [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ] compared diferent approaches to
fore, we explore in the following related work on patent automatically classify French patent documents at the
classification on the one hand, and patent summarization IPC main group and subgroup-levels. Their experiments
on the other hand. showed the need for more sophisticated techniques such
as data augmentation, clustering, and negative sampling
2.1. Patent Classification at deeper levels such as subgroups. Chen and Chang [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ]
proposed a three-step classification (TPC) algorithm that
achieved 36.07% accuracy at the IPC subgroup-level
classification. D’hondt et al. [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ] demonstrated the
efectiveness of combining words and PoS-filtered skipgrams.
      </p>
      <sec id="sec-1-1">
        <title>They showed that extending the textual representation</title>
        <p>from traditional word-only based features to more
finegrained phrase-based features significantly improves the
performance of automatic classification at the
subgrouplevel.</p>
        <p>
          Recent research results that are highly relevant to this
study are presented. In recent years, deep learning
methods have attracted a great deal of attention in patent
classification tasks. Li et al. [
          <xref ref-type="bibr" rid="ref4">4</xref>
          ] proposed a deep learning
algorithm DeepPatent, which combines word embedding
and convolutional neural networks, and achieved 73.88%
precision at the IPC subclass-level using the USPTO-2M
dataset and 83.98% using the CLEF-IP 2011 dataset. They
compared various combinations of text fields as training
data and concluded that using the first 100 words of the
title and abstract was optimal. Lee and Hsiang [
          <xref ref-type="bibr" rid="ref3">3</xref>
          ]
finetuned a pre-trained BERT model for patent classification
and reported that they achieved better performance than
        </p>
      </sec>
      <sec id="sec-1-2">
        <title>DeepPatent. They also showed that using only claims as training data instead of title and abstract produced comparable performance. Roudsari et al. [2] fine-tuned</title>
        <sec id="sec-1-2-1">
          <title>2.2. Patent Document Summarization and</title>
        </sec>
        <sec id="sec-1-2-2">
          <title>LLM-based Summary Generation</title>
        </sec>
      </sec>
      <sec id="sec-1-3">
        <title>Sharma et al. [11] proposed BIGPATENT, a large dataset containing 1.3 million U.S. patent documents and abstract and coherent summaries written by humans. Experiments with BIGPATENT suggest that summarization</title>
        <p>
          tasks for specialized texts such as patent documents re- Table 1
quire deeper understanding and abstraction than simply Description of datasets
extracting phrases from the original text. Ding et al. [
          <xref ref-type="bibr" rid="ref12">12</xref>
          ]
generated summaries for 1630 patent document abstract
and claims combinations using several text summariza- USPTO-70k
tion models. They concluded that the GPT-3.5-turbo
model can summarize patent documents better than other
models, and they stated that prompting strategy is the
key to the success of patent document summarization.
        </p>
        <p>
          Yang et al. [
          <xref ref-type="bibr" rid="ref13">13</xref>
          ] evaluated ChatGPT’s performance on
various text summarization tasks using a benchmark dataset
and found that the summaries generated by ChatGPT
were comparable to those generated by traditional
methods, suggesting that ChatGPT is a promising powerful
tool for text summarization. As in the above study, there
have been reports of using LLMs to generate summaries
of long texts, including patent documents, and
evaluating their quality. However, to the best of my knowledge,
there are no reported cases of using LLMs to generate a
summary of a patent document and using the generated
summary to perform a patent classification task.
        </p>
        <p>Subclass Multi-label
# of documents
Avg. labels per patent
Total labels
Subclass Multi-label
# of documents
Avg. labels per patent
Total labels
Subgroup Multi-class
# of documents
Total labels
CLEF-IP 2011 subset</p>
        <p>Train</p>
        <p>Valid</p>
        <p>Test</p>
      </sec>
    </sec>
    <sec id="sec-2">
      <title>3. Datasets</title>
      <sec id="sec-2-1">
        <title>For our experiments, we used two popular patent</title>
        <p>
          datasets: one from USPTO and one from CLEF.
USPTO-70k Dataset. The USPTO-70k [
          <xref ref-type="bibr" rid="ref14">14</xref>
          ] is a dataset
for CPC subclass-level patent classification tasks shown
in the upper part of Table 1. It consists of training data
from 50250 USPTO patent documents from 2006 to 2017,
validation data from 10,000 documents in 2018, and
testing data from 10,000 documents in 2019. This dataset
contains various text fields of patent documents such as
title, abstract, claims, and detailed description.
        </p>
        <p>
          Since the USPTO-70k does not contain subgroup-level
classification information, we augmented the dataset
with subgroups of the main IPC (the IPC that best
represents the technical field to which the patent belongs)
by referring to USPTO Bulk Data Storage System.1 As in
the previous studies [
          <xref ref-type="bibr" rid="ref10 ref9">9, 10</xref>
          ], only subgroup labels with
at least seven training documents were retained for the
subgroup-level classification task. After data cleaning,
the patent documents shown in the middle part of Table 1
were obtained.
        </p>
      </sec>
      <sec id="sec-2-2">
        <title>CLEF-IP 2011 Subset. The CLEF-IP 2011 dataset [15]</title>
        <p>
          consists of over 2.6 million patent documents from the
EPO and 0.4 million patent documents from the WIPO,
ifled between 1978 and 2009. We extracted EPO English
patent documents from 2000 to 2009 that contain IPC
labels, title, abstract, claims, and detailed description from
the CLEF-IP 2011 dataset. Training data and validation
data were collected according to the data size of
USPTO70k dataset, as shown in the lower part of Table 1. We
used the CLEF-IP test sample (1000 documents) as test
data. The word-piece token distributions in the diferent
text fields of the CLEF-IP 2011 subset shown in Figure 1
show a similar distribution to the token distribution for
the USPTO-70k dataset presented by Pujari et al. [
          <xref ref-type="bibr" rid="ref14">14</xref>
          ]
        </p>
      </sec>
      <sec id="sec-2-3">
        <title>While the USPTO has adopted the concept of "main</title>
      </sec>
      <sec id="sec-2-4">
        <title>IPC" and uniquely determines the IPC corresponding</title>
        <p>to every patent document, the EPO does not follow the
main IPC rule. Therefore, the subgroup-level multi-class
classification task was performed only on the USPTO-70k
dataset.</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>4. Experimental Setup</title>
      <p>In this study, the following procedure was used for the
experiments:
1. Summary Generation. Using each text field
(abstract, claims, and detailed description) of the
patent document as input, the LLM generated a
summary for each text field.
2. Patent Classification. AI-generated summaries
were used to fine-tune the pre-trained models to
adapt them to a multi-label or multi-class
classification task. Multi-label classification is a problem
where each sample may belong to more than one
label, and multi-class classification is a problem
where each sample belongs to one of the classes.</p>
      <sec id="sec-3-1">
        <title>The fine-tuned model predicts the classification</title>
        <p>using the generated summary as input.</p>
      </sec>
      <sec id="sec-3-2">
        <title>3. Evaluation. To evaluate the performance of the</title>
        <p>ifne-tuned model, the prediction results were
an(a) Abstract
(b) Claims
(c) Detailed description
You are a patent expert. Summarize the patent document given by the user in 100 words focusing on
an addition to the state of the art. "Addition to the state of the art" means the diference between the
subject matter in a patent document and the collection of all technical subject matter that has already
been placed within public knowledge. Output only the summary.
alyzed using hierarchical precision, recall, and</p>
      </sec>
      <sec id="sec-3-3">
        <title>F1-score (See section 4.3.). For comparison, the</title>
        <p>same evaluation was performed on the results
of model training and classification prediction
with the original texts instead of AI-generated
summaries.</p>
        <p>Each procedure is further described in detail in the
following subsections.</p>
        <sec id="sec-3-3-1">
          <title>4.1. Summary Generation</title>
          <p>the system will do its best to sample deterministically,
and repeated requests with the same seed and parameters
should return the same results. Note, however, that this
is currently a beta feature and does not always produce
exactly the same output.3</p>
          <p>Prompts are created so that the sum of system and
user messages does not exceed the maximum context
window size (16,385 tokens). If the maximum context
window size is exceeded, the first portion of text up to the
maximum context window size is used and the remainder
is truncated.</p>
        </sec>
        <sec id="sec-3-3-2">
          <title>4.2. Patent Classification</title>
        </sec>
      </sec>
      <sec id="sec-3-4">
        <title>We used the gpt-3.5-turbo-0125 model to generate sum</title>
        <p>
          maries from patent documents. The prompts given to
the model consisted of one of the three types of system
role messages shown in Table 2 and a user role message
containing the patent text. The "Elaborate" prompt in
Table 2 is based on the description in VIII. PRINCIPLES
OF THE CLASSIFICATION of the IPC guidelines [
          <xref ref-type="bibr" rid="ref1">1</xref>
          ].2
        </p>
      </sec>
      <sec id="sec-3-5">
        <title>To ensure as much reproducibility of the experiment</title>
        <p>
          as possible, the temperature value of the model was set to
0.0, and a seed was specified (seed=42). With the seed set,
In this study, we applied the RoBERTa (Robustly
Optimized BERT Pretraining Approach) model [16] to
multilabel/multi-class classification tasks for patent
documents. RoBERTa has been reported to require less
time for fine-tuning than other transformer-based
models [
          <xref ref-type="bibr" rid="ref2">2</xref>
          ]. We used the pre-trained RoBERTa-Base model
and adapted it to the patent classification task by adding
a dropout layer and a linear layer of size equal to the
number of classification classes. The model was
opti2T"choelle"Ectliaobnoroaftael"lpterochmnpictarlesquubirjeecstamtiamtteesrtaalmrepadays tihnethbeaspisubfolirc tdhoe- mized with the Adam optimizer using the
hyperparamemain" to be followed faithfully. However, since it would be unfair ters shown in Table 3.
to give metadata for individual patent documents only to the "Elab- Furthermore, based on the results of preliminary
experorate" prompt and not to the other prompts ("Patent", "Simple"), we iments conducted with reference to Merchant et al. [17],
decided not to give their time stamps. Detailed descriptions usually
include a description of the background technology, so it would be
possible to infer "state of the art" based on that information. 3https://platform.openai.com/docs/api-reference/chat/create
we found that freezing all layers except the linear layer,
the pooler layer, and the last three layers of RoBERTa had
little efect on performance. Therefore, we unfroze these
layers and performed fine-tuning for training eficiency.
        </p>
        <p>For the multi-label classification task, we trained the
model with the prediction threshold set to 0.5, based on
the report by Giczy et al. [18]. After training, the
prediction threshold of the model was varied from 0.1 to
0.9 in 0.1 increments, label predictions were made on
the validation data, and the prediction threshold with
the highest micro-average hierarchical F1-score (See
section 4.3.) was selected as the final model. As shown in
Figure 2, the maximum micro-averaged F1-score was
obtained with a prediction threshold of 0.2 and 0.3 when
using the USPTO-70k dataset and the CLEF-IP 2011
subset, respectively.</p>
        <sec id="sec-3-5-1">
          <title>4.3. Evaluation</title>
          <p>(a) USPTO-70k dataset
Evaluating AI-generated Summaries. In this study,
two automatic evaluation metrics, ROUGE [19] and</p>
        </sec>
      </sec>
      <sec id="sec-3-6">
        <title>BERTScore [20], were used to evaluate the AI-generated</title>
        <p>summaries. For both indicators, scores were calculated (b) CLEF-IP 2011 subset
using the original text of the patent document abstract Figure 2: Comparision of various threshold values for
as the reference summary. RoBERTa model (subclass-level classification, abstract original</p>
        <p>The ROUGE score is a measure that evaluates the de- text)
gree of n-gram overlap between the generated summary
and the reference summary. ROUGE-1 and ROUGE-L
were used in this study; ROUGE-1 measures unigram label classification. For each patent document  in the test
overlap, while ROUGE-L evaluates summary similarity data, the predicted label set  consists of all predicted
based on the longest common subsequence. labels and their ancestors. Similarly, the true label set</p>
        <p>The BERTScore is a measure that evaluates the simi- consists of true labels and their ancestors. These
evalualarity between generated and reference summaries using tion metrics consider the degree of agreement between
the pre-trained language model BERT. AI-generated sum- predicted and true labels in hierarchically structured
clasmaries are likely to use diferent vocabulary than the sifications, even at higher levels of the hierarchy. For
original abstracts, but by using BERTScore, word seman- example, if the true label is "G06F", the predicted label
tic similarity can be taken into account. of Model A is "G06K", and the predicted label of Model</p>
      </sec>
      <sec id="sec-3-7">
        <title>B is "H04L", the prediction of Model A, which is con</title>
        <p>
          Evaluating Patent Classification Performance. Fol- sistent with the class-level "G06", is rated better than
lowing Pujari et al. [
          <xref ref-type="bibr" rid="ref14 ref6">14, 6</xref>
          ], we used hierarchical preci- that of Model B. The subgroup-level multi-class
classision, recall, and F1-score defined as ℎ = ∑︀∑|︀|∩|| , fication was evaluated using accuracy, following Chen
ℎ = ∑︀ |∩| and ℎ 1 = 2ℎ· ℎ+· ℎℎ and proposed by and Chang [
          <xref ref-type="bibr" rid="ref9">9</xref>
          ] and D’hondt et al. [
          <xref ref-type="bibr" rid="ref10">10</xref>
          ]. Since this is a
Kiritchen∑k︀o|e|t al. [21] to evaluate subclass-level multi- multi-class classification task, the accuracy is consistent
with the micro-averaged F1-score.
Claims
Original text
(first 100 words)
Detailed description
Original text
(first 100 words)
Claims
AI-generated
summary
Detailed description
AI-generated
summary
        </p>
        <p>Text
"An exterior mirror assembly including an attachment member for supporting an approach light.</p>
        <p>The attachment member interconnects a mirror housing to a vehicle and includes an opening for
receiving a lens. Light projects through the lens from the attachment member in order to illuminate
a predetermined area in proximity to the vehicle. A light source may be housed within the support
member or, alternatively, a light source may be housed interior to the vehicle and a light path
transport light from the light source to the lens for projection from the support member."
"1. An exterior rear view mirror assembly comprising: a housing; a reflective mirror supported by
the housing and arranged in a rearwardly facing direction; a support member having an outboard
end attached to the housing and an inboard end secured to an outer surface of a motor vehicle, the
support member including an opening; and a light transmitting lens disposed in the opening of
the support member, the lens enabling light to project from an interior of the support member to
an exterior of the support member to illuminate a predetermined area in proximity to the motor
vehicle. 2. The"
"DESCRIPTION OF THE PREFERRED EMBODIMENTS Referring now to the figures, and in
particular FIGS. 1–3, a mirror assembly10includes a housing12which houses and supports a reflective
element14, such as mirror glass. Mirror assembly10of the present invention will be described
with respect to an exterior driver side rear view mirror which enables the vehicle operator to view
an area beside and rearward with respect to the vehicle. Housing12connects to an attachment
member16including a sail portion18and an integral, laterally projecting support arm or bracket
portion20. The interconnection between housing12and attachment member16enables housing12to
pivot with respect to the support arm portion20. Particularly, housing12pivots rearward and"
"The patent describes an exterior mirror assembly with an attachment member supporting an
approach light. The attachment member connects the mirror housing to a vehicle and has an opening
for a lens. Light passes through the lens to illuminate an area near the vehicle. The light source can
be within the attachment member or inside the vehicle, with a light path transporting light from
the source to the lens for projection."
"The patent document describes an exterior rear view mirror assembly for a motor vehicle. It
includes a housing, a reflective mirror, a support member with a light transmitting lens for illuminating
an area near the vehicle, and optional features like a light assembly with adjustable light source,
reflective elements, environmental seal, sail for securing the support member, and a cover for the
lens. The system can also include a timer, speed sensor, switch, and control circuit for activating
the light source based on vehicle speed or various sensors."
"The patent document describes a mirror assembly for vehicles with an integrated approach light
system. The mirror assembly includes a housing supporting a reflective element and an attachment
member with a sail portion and a support arm portion. The approach light can be located in
the support arm portion or sail portion, providing illumination for areas adjacent to the vehicle.</p>
        <p>Various configurations for the approach light placement, activation, and control are discussed,
including adjustable reflectors, variable timeout functions, and integration with vehicle systems.</p>
        <p>The invention aims to improve visibility, safety, and serviceability of the approach light system in
vehicle mirror assemblies."</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>5. Results and Discussion</title>
      <p>and the AI-generated summary of the claims specifically
describe the components of the invention. This is
be5.1. Summary Generation cause patent claims must clearly indicate the scope of the
patent right. The original text of the detailed description
Table 4 shows examples of original patent document text describes a specific embodiment of the invention. The
AIand AI-generated summaries on the USPTO-70k dataset. generated summary of the detailed description describes
Both the original text and the AI-generated summary of the most comprehensive content of the invention, such
the abstract describe the basic content of the invention. as its structure, function, and purpose, compared to the
The original text of the claims is written in a manner spe- other texts.
cific to patent claims. In addition, both the original text Table 5 shows the metrics for the summary generated</p>
      <p>AB: Abstract, CL: Claims, DD: Detailed description, AI: AI-generated summary
Table 7 detailed description. This means that summaries
generPerformance comparison on USPTO-70k dataset for subgroup- ated from claims and detailed description contain more
level multi-class classification words and meanings that are diferent from the original
Text field hAcc Acc Macro-avg. tcelaxitmosf ttohedeatbasilterdacdte. sAcrsipt htieo nte,xt htemdoevceredafsreoimn BabEsRtTraScctotroe
hF1 F1 was smaller than the decrease in ROUGE score. This
AB (OT) 52.2 22.3 16.0 11.0 indicates that although matching at the word level is
CL (OT) 52.0 21.9 16.0 11.1 decreasing, semantic similarity is relatively maintained.
DD (OT) 41.2** 14.3** 9.8** 6.5** This indicates that the gpt-3.5-turbo-0125 model tends to
AB (AI) 52.0 22.0 16.4 11.3 generate contextually relevant summaries for longer and
CL (AI) 53.9** 23.9** 17.7** 12.9** more complex texts, such as claims and detailed
descripDD (AI) 53.9** 23.5** 18.2** 13.1** tion.</p>
      <p>hAcc: hierarchical Accuracy, Acc: Accuracy, AB: Abstract, CL:
Claims, DD: Detailed description, OT: Original text, AI:
AIgenerated summary
*p&lt;0.05, **p&lt;0.01
from each text field in the patent document. The ROUGE
score represents the score of the respective AI-generated
summary of the abstract, claims, and detailed description
when the original text of the abstract is used as the
reference; the BERT score is similar. In both datasets, each
score decreased as text moved from abstract to claims to</p>
      <sec id="sec-4-1">
        <title>5.2. Patent Classification</title>
        <sec id="sec-4-1-1">
          <title>USPTO-70k Dataset. Tables 6 and 7 show the results</title>
          <p>
            of the patent classification task using AI-generated
summaries as training data for the "Patent" prompt (shown in
Table 2) on the USPTO-70k dataset. According to Welch’s
t-test with a sample size of 5, scores were significantly
higher for both subclass-level multi-label classification
and subgroup-level multi-class classification tasks when
using AI-generated summaries of claims or detailed
deText field
hR
Table 10 the original text of the abstract. This may be because
Performance comparison of AI-generated summaries with the detailed description on the USPTO-70k dataset often
three diferent prompts on USPTO-70k dataset for subgroup- begins with figure captions or notes on the scope of the
level multi-class classification invention’s disclosure, and in many cases, the technical
Prompt hAcc Acc Macro-avg. wfeoatrudrse.s of the invention are not included in the first 100
hF1 F1 Comparing the results with Pujari et al.’s THMM [
            <xref ref-type="bibr" rid="ref6">6</xref>
            ],
Patent 54.2 23.6 18.2 13.1 the combination of RoBERTa and the AI-generated
sumSimple 54.0 23.5 18.0 12.9 mary was superior to the combination of THMM and
Elaborate 54.3 24.4* 18.5 13.4 the original text when claims or detailed descriptions
hAcc: hierarchical Accuracy, Acc: Accuracy, hF1: hierarchical F1- were used. In particular, the micro-averaged hierarchical
score, F1: F1-score, DD: Detailed description, AI: AI-generated F1-score was 5.0 points higher, and the macro-averaged
summary hierarchical F1-score was 3.6 points higher when detailed
*p&lt;0.05, **p&lt;0.01 descriptions were used, indicating that AI-generated
summaries efectively extract important information from
scriptions than when using the original abstract text. long texts such as detailed descriptions.
          </p>
        </sec>
        <sec id="sec-4-1-2">
          <title>However, using AI-generated summaries of abstracts did</title>
          <p>not result in significant diferences in scores. This sug- CLEF-IP 2011 Subset. Table 8 shows the results on
gests that the generative model can provide efective the CLEF-IP 2011 subset. each score was significantly
text for classification tasks by summarizing important higher when using the AI-generated summary of claims
information from claims or detailed description. and detailed descriptions than when using the original</p>
          <p>The scores were significantly lower when using the abstract text.
original text of the detailed description than when using In particular, the micro-averaged hierarchical F1-score</p>
        </sec>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>6. Conclusion</title>
      <p>improved by 4.4 points when using the AI-generated
summaries of the detailed description compared to when
using the original abstract text. This indicates that the AI- In this study, we investigate the efect of using summaries
generated summaries eefctively extracts classification- generated by LLMs to improve the performance of
aurelevant information from the detailed description. tomatic patent classification. Our experiments on the</p>
      <p>The scores were also significantly higher when the USPTO-70k dataset and the CLEF-IP 2011 subset
demonoriginal text of the detailed description was used. This strated that models trained on AI-generated summaries
may be because EPO patent documents often include the of claims and detailed descriptions achieve significantly
technical field to which the invention belongs (e.g., "The higher scores compared to those trained on original
abpresent invention relates to a sound and heatinsulating stract text in both subclass-level multi-label classification
material.) at the beginning of the detailed description. and subgroup-level multi-class classification tasks. These</p>
      <p>
        In all cases, the results are below Yadrintsev et al.’s results suggest that AI-generated summaries adequately
KNN [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ]. This is likely due to the fact that this study uses capture information relevant to patent classification.
only a single text field and about one-tenth the size of The proposed approach improves automatic patent
the CLEF-IP 2011 dataset for training data. classification techniques by utilizing LLMs to generate
      </p>
      <p>
        The best (non-hierarchical) F1-score for this method high-quality summaries. We believe that this research
is 63.1 points when using AI-generated summaries of builds new possibilities for improving the accuracy and
detailed description, which is lower than the scores from eficiency of patent classification, which is important for
previous studies such as those from DeepPatent by Li et managing the ever-increasing amount of patent
informaal. [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ] (83.98 points) and KNN by Yadrintsev et al. [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ] (71.0 tion.
points). This is likely due to the fact that this method Future research directions include exploring the
optiuses only a single text field and only about one-tenth the mal prompts for generating summaries of patent
docusize of the CLEF-IP 2011 dataset as training data. ments by LLMs and investigating the applicability of the
proposed approach to other patent classification models.
      </p>
      <sec id="sec-5-1">
        <title>Although this study experimented with a flat approach</title>
        <p>that does not consider the hierarchical structure of patent
classification, the performance of an automatic patent
classification system could be further enhanced by
incorporating the interdependence of each hierarchical label
into the automatic classification process.</p>
      </sec>
      <sec id="sec-5-2">
        <title>Prompt Comparison. Tables 9 and 10 shows a comparison of classification performance when using AIgenerated summaries by each of the three prompts (Patent, Simple, and Elaborate).</title>
        <p>For the subclass-level classification task, the
performance with the "Patent" prompt and with the "Simple"
prompt is competitive on the USPTO-70k dataset, while
the performance with the "Simple" prompt is slightly
higher on the CLEF-IP 2011 subset. On the other hand,
performance with the "Elaborate" prompt tends to be
slightly poorer on both datasets. These results suggest
that overly detailed summaries are not necessarily
effective for subclass-level patent classification tasks but
rather that even summaries generated with the "Simple"
prompt are efective enough.</p>
      </sec>
      <sec id="sec-5-3">
        <title>For the subgroup-level classification task, the perfor</title>
        <p>mance of the "Elaborate" prompt was slightly better than
that of the "Patent" and "Simple" prompts. The results
suggest that the summary generated by the "Elaborate"
prompt may be useful for more detailed classification
tasks.</p>
        <p>From the above, we believe that in the patent
classification task, it is important to use prompts with a
reasonable level of detail depending on the
classification level, although the improvement in performance
obtained by optimizing the prompts is limited.
However, since only three prompts were compared in this
experiment, a comprehensive investigation of the efects
of the various prompts on performance on the patent
classification task is a topic for future work.</p>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>World</given-names>
            <surname>Intellectual Property Organization</surname>
          </string-name>
          ,
          <article-title>Guide to the international patent classification (</article-title>
          <year>2023</year>
          ), https: //www.wipo.int/edocs/pubdocs/en/wipo
          <article-title>-guide-i pc-2023-en-guide-to-the-international-patent-cla ssif ication-2023</article-title>
          .pdf ,
          <year>2023</year>
          . Accessed:
          <fpage>2024</fpage>
          -4-16.
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>A.</given-names>
            <surname>Haghighian Roudsari</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Afshar</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W.</given-names>
            <surname>Lee</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Lee</surname>
          </string-name>
          ,
          <article-title>Patentnet: multi-label classification of patent documents using deep learning based language understanding</article-title>
          ,
          <source>Scientometrics</source>
          <volume>127</volume>
          (
          <year>2022</year>
          )
          <fpage>207</fpage>
          -
          <lpage>231</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>J.-S.</given-names>
            <surname>Lee</surname>
          </string-name>
          ,
          <string-name>
            <surname>J. Hsiang,</surname>
          </string-name>
          <article-title>PatentBERT: Patent classification with fine-tuning a pre-trained BERT model</article-title>
          ,
          <source>World Patent Information</source>
          <volume>61</volume>
          (
          <year>2020</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>S.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Hu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Cui</surname>
          </string-name>
          , J. Hu,
          <article-title>DeepPatent: patent classification with convolutional neural networks and word embedding</article-title>
          ,
          <source>Scientometrics</source>
          <volume>117</volume>
          (
          <year>2018</year>
          )
          <fpage>721</fpage>
          -
          <lpage>744</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>Y.</given-names>
            <surname>Zuo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Sagot</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Gerdes</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Mouzoun</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S. Ghamri</given-names>
            <surname>Doudane</surname>
          </string-name>
          ,
          <article-title>Exploring data-centric strategies for French patent classification: A baseline and comparisons</article-title>
          , in: Actes de CORIA-TALN
          <year>2023</year>
          .
          <article-title>Actes de la 30e Conférence sur le Traitement Automatique des Langues Naturelles (TALN)</article-title>
          ,
          <source>volume Proceedings, CEUR-WS.org</source>
          ,
          <year>2011</year>
          . 1 : travaux de recherche originaux - articles longs, [16]
          <string-name>
            <given-names>Y.</given-names>
            <surname>Liu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Ott</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Goyal</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Du</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Joshi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Chen</surname>
          </string-name>
          ,
          <string-name>
            <surname>ATALA</surname>
          </string-name>
          ,
          <year>2023</year>
          , pp.
          <fpage>349</fpage>
          -
          <lpage>365</lpage>
          .
          <string-name>
            <given-names>O.</given-names>
            <surname>Levy</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Lewis</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Zettlemoyer</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Stoyanov</surname>
          </string-name>
          ,
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>S. C.</given-names>
            <surname>Pujari</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Mantiuk</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Giereth</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Strötgen</surname>
          </string-name>
          ,
          <article-title>Roberta: A robustly optimized bert pretraining apA. Friedrich, Evaluating neural multi-field doc-</article-title>
          proach,
          <year>2019</year>
          . arXiv:
          <year>1907</year>
          .11692.
          <article-title>ument representations for patent classification</article-title>
          , [17]
          <string-name>
            <given-names>A.</given-names>
            <surname>Merchant</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Rahimtoroghi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Pavlick</surname>
          </string-name>
          , I. Tenney, in
          <source>: BIR</source>
          <year>2022</year>
          :
          <article-title>12th International Workshop on What happens to BERT embeddings during fineBibliometric-enhanced Information Retrieval at tuning?</article-title>
          ,
          <source>in: Proceedings of the Third BlackboxNLP ECIR</source>
          <year>2022</year>
          , April 10,
          <year>2022</year>
          , hybrid,
          <year>2022</year>
          . Workshop on Analyzing and Interpreting Neural
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>V.</given-names>
            <surname>Yadrintsev</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Bakarov</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Suvorov</surname>
          </string-name>
          ,
          <string-name>
            <surname>I. Sochenkov</surname>
          </string-name>
          ,
          <article-title>Networks for NLP, Association for Computational Fast and accurate patent classification in search Linguistics</article-title>
          ,
          <year>2020</year>
          , pp.
          <fpage>33</fpage>
          -
          <lpage>44</lpage>
          . doi:
          <volume>10</volume>
          .18653/v1/20 engines,
          <source>Journal of Physics: Conference Series 1117 20.blackboxnlp-1</source>
          .
          <fpage>4</fpage>
          . (
          <year>2018</year>
          )
          <article-title>012004</article-title>
          . doi:
          <volume>10</volume>
          .1088/
          <fpage>1742</fpage>
          -6596/1117/1 [18]
          <string-name>
            <given-names>A. V.</given-names>
            <surname>Giczy</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N. A.</given-names>
            <surname>Pairolero</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A. A.</given-names>
            <surname>Toole</surname>
          </string-name>
          , Identify/012004.
          <article-title>ing artificial intelligence (ai) invention: A novel ai</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>Y.</given-names>
            <surname>Hoshino</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Utsumi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Matsuda</surname>
          </string-name>
          , Y. Tanaka, patent dataset,
          <source>The Journal of Technology Transfer K. Nakata, IPC prediction of patent documents 47</source>
          (
          <year>2022</year>
          )
          <fpage>476</fpage>
          -
          <lpage>505</lpage>
          .
          <article-title>using neural network with attention for hierar-</article-title>
          [19]
          <string-name>
            <surname>C.-Y. Lin</surname>
            ,
            <given-names>ROUGE:</given-names>
          </string-name>
          <article-title>A package for automatic evalchical structure</article-title>
          ,
          <source>PLoS One</source>
          <volume>18</volume>
          (
          <year>2023</year>
          )
          <article-title>e0282361. uation of summaries</article-title>
          , in: Text Summarization doi:https://doi.org/10.1371/journal.po Branches Out, Association for Computational Linne.
          <volume>0282361</volume>
          . guistics,
          <year>2004</year>
          , pp.
          <fpage>74</fpage>
          -
          <lpage>81</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>Y.-L.</given-names>
            <surname>Chen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.-C.</given-names>
            <surname>Chang</surname>
          </string-name>
          ,
          <article-title>A three-phase method</article-title>
          [20]
          <string-name>
            <given-names>T.</given-names>
            <surname>Zhang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Kishore</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Wu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K. Q.</given-names>
            <surname>Weinberger</surname>
          </string-name>
          , for patent classification, Information Processing Y. Artzi, Bertscore:
          <article-title>Evaluating text generation with</article-title>
          &amp;
          <source>Management</source>
          <volume>48</volume>
          (
          <year>2012</year>
          )
          <fpage>1017</fpage>
          -
          <lpage>1030</lpage>
          . doi:https: bert, in: International Conference on Learning Rep//doi.org/10.1016/j.ipm.
          <year>2011</year>
          .
          <volume>11</volume>
          .001. resentations,
          <year>2020</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <surname>E. D'hondt</surname>
            , S. Verberne,
            <given-names>N.</given-names>
          </string-name>
          <string-name>
            <surname>Oostdijk</surname>
            , L. Boves, [21]
            <given-names>S.</given-names>
          </string-name>
          <string-name>
            <surname>Kiritchenko</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          <string-name>
            <surname>Matwin</surname>
            ,
            <given-names>A. F.</given-names>
          </string-name>
          <string-name>
            <surname>Famili</surname>
          </string-name>
          , et al.,
          <article-title>FuncPatent Classification on Subgroup Level Using Bal- tional annotation of genes using hierarchical text anced</article-title>
          <source>Winnow</source>
          , Springer Berlin Heidelberg,
          <year>2017</year>
          , categorization,
          <source>in: Proc. of the ACL Workshop</source>
          pp.
          <fpage>299</fpage>
          -
          <lpage>324</lpage>
          . doi:
          <volume>10</volume>
          .1007/978-3-
          <fpage>662</fpage>
          -53817-3 on Linking Biological Literature, Ontologies and _
          <fpage>11</fpage>
          . Databases: Mining Biological Semantics,
          <year>2005</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>E.</given-names>
            <surname>Sharma</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <surname>BIGPATENT:</surname>
          </string-name>
          <article-title>A largescale dataset for abstractive and coherent summarization, in: Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics</article-title>
          , Association for Computational Linguistics,
          <year>2019</year>
          , pp.
          <fpage>2204</fpage>
          -
          <lpage>2213</lpage>
          . doi:
          <volume>10</volume>
          .18653/v1/
          <fpage>P19</fpage>
          -121
          <fpage>2</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>J.</given-names>
            <surname>Ding</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Chen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Kolapudi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Pobbathi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Nguyen</surname>
          </string-name>
          ,
          <article-title>Quality evaluation of summarization models for patent documents</article-title>
          ,
          <source>in: 2023 IEEE 23rd International Conference on Software Quality, Reliability, and Security (QRS)</source>
          ,
          <year>2023</year>
          , pp.
          <fpage>250</fpage>
          -
          <lpage>259</lpage>
          . doi:
          <volume>10</volume>
          .1109/QRS60937.
          <year>2023</year>
          .
          <volume>00033</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <given-names>X.</given-names>
            <surname>Yang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Zhang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Chen</surname>
          </string-name>
          , W. Cheng,
          <article-title>Exploring the limits of chatgpt for query or aspect-based text summarization</article-title>
          ,
          <year>2023</year>
          . arXiv:
          <volume>2302</volume>
          .
          <fpage>08081</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <given-names>S. C.</given-names>
            <surname>Pujari</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Friedrich</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Strötgen</surname>
          </string-name>
          ,
          <article-title>A multi-task approach to neural multi-label hierarchical patent classification using transformers</article-title>
          ,
          <source>in: Proceedings of the 43rd EUROPEAN CONFERENCE ON INFORMATION RETRIEVAL, Online</source>
          ,
          <year>2021</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <given-names>F.</given-names>
            <surname>Piroi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Lupu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Hanbury</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Zenz</surname>
          </string-name>
          ,
          <string-name>
            <surname>CLEFIP</surname>
          </string-name>
          <year>2011</year>
          :
          <article-title>Retrieval in the intellectual property domain</article-title>
          ,
          <source>in: CLEF 2011 Labs and Workshop</source>
          , Notebook Papers,
          <fpage>19</fpage>
          -22
          <source>September</source>
          <year>2011</year>
          , Amsterdam, The Netherlands, volume
          <volume>1177</volume>
          of CEUR Workshop
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>