<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Automatic Identification of Patent Claim Types: Enhancing Eficiency in Patent Analysis</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Rima Dessí</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Hidir Aras</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Mark Prince</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>René Hackl-Sommer</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>CAS - Chemical Abstracts Service</institution>
          ,
          <addr-line>Columbus, Ohio</addr-line>
          ,
          <country country="US">USA</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>FIZ Karlsruhe - Leibniz Institute for Information Infrastructure</institution>
          ,
          <country country="DE">Germany</country>
        </aff>
      </contrib-group>
      <fpage>67</fpage>
      <lpage>71</lpage>
      <abstract>
        <p>Patents are an important source of technological innovation. Due to the large number of patents published each year, it has become increasingly dificult to find precise information on specific patent inventions, which requires not only professional search systems but also many years of patent expertise. Patent claims are the backbone of inventions and define the scope of legal protection. They can be classified in diferent types in terms of what they claim, e.g. for a physical entity we can speak of "product claims", while for an activity we can refer to as a "process claim". Manual identification of these claim types for a large set of documents is labor-intensive and time-consuming. To address this challenge, we developed a Patent Claim Type Recognition (PCTR) model based on Deep Learning (DL), which is able to automatically identify pre-defined types of patent claims. Further, we also built a rule-based heuristic approach, to generate training data to be used by the PCTR model. The proposed model was evaluated by using a dataset labeled by Subject Matter Experts (SMEs). Our experimental results demonstrate that the PCTR model accurately identifies the type of given patent claims, ofering a promising approach to streamline patent analysis and evaluation processes.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>
        1. Introduction dependent claims, together building a hierarchy1.
Typically, independent claims contain core inventive
inforPatents enable inventors to disclose their inventions and mation whereas dependent claims specify improvements
protect them legally by preventing others from using, sell- or variations. Such variations can be rather miniscule or
ing, and producing the invention without permission [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. they can be quite substantive. However, the legal jargon
Therefore, patents encourage further inventions by grant- in patent claims can make it dificult to understand what
ing exclusive rights to inventors and thus foster more exactly a text is about. Further, the claims often relate
research and development. However, the ever-increasing to a particular subject matter, such as apparatus,
comnumber of available patents and their complex charac- position, process, or a combination thereof. Therefore,
teristics in nature poses several challenges to scientists, it is crucial to precisely identify the type of a claim for
lawyers and information professionals. These documents accurate patent analysis. However, manual
identificaare diverse, encompassing text, formulas, drawings, ta- tion of claim types in larger result sets is expensive and
bles, and more, while also being lengthy and filled with time-consuming.
domain-specific vocabulary that is tailored to the target Several studies have been proposed to eficiently and
ifeld. The so-called full text of a patent document often efectively analyze and understand patents as well as
consists of sections, namely, title, abstract, claims, and their claims. Most of these methods employ Machine
description. Learning (ML) and Natural Language Processing (NLP)
      </p>
      <p>
        The claims are a crucial component of the patent doc- techniques to automate patent classification, claim
clasuments, they define the legal scope of protection of an sification, and claim type identification. However, often
invention. Essentially, they specify the subject matter they require manually labeled large amounts of training
that is sought to be protected and refer to the core in- data. Further, approaches focus on claims designed to
ventive information of a patent. These claims serve as measure diferent aspects of patents such as comparison
the foundation of what aspects of an invention should be of patent claims and economic growth and social
welprotected from infringement. Therefore, an accurate anal- fare [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ] or to measure technological patent scope with
ysis and understanding of them are vital for inventors, semantic analysis of patent claims [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ].
examiners, and scientists. There can be independent and In this paper, we propose a Patent Claim Type
Recognition (PCTR) model based on Deep Learning (DL)
techniques. As mentioned before, patent claims refer to a
particular subject matter, i.e., types such as apparatus,
composition, system, etc. The main goal of PCTR is to
automatically identify this type information for each given
patent claim. Additionally, we design a rule-based heuris- method is based on a deep neural network that includes
tic model to generate training data for the PCTR model. a transformer layer, it does not require any manually
The training data consists of claims paired with their re- labeled data, instead, we define a heuristic model to label
spective types, determined through the heuristic model. large amounts of data eficiently and efectively without
Subsequently, this curated data is used to train the PCTR requiring any manual efort. Consequently, this dataset
model, finally, the trained model is able to assign a claim is then used to train the proposed PCTR model.
type to a given patent claim.
      </p>
      <p>Overall the main contributions of the paper are as
follows: 3. Patent Claim Type Recognition
• Introducing a rule-based heuristic model that as- In this section, we give a definition for the regarded
probsigns types to claims, facilitating the generation lem and describe the predefined claim types for the model
of training data. prediction.
• A transformer-based deep neural network archi- Problem Definition: Given a claim text and a
predetecture designed to automatically identify patent ifned type list, the task of the PCTR model is to assign
claim types. the most relevant type from the predefined type list.
• A comprehensive evaluation of the PCTR model Predefined Types: The specific claim types listed below,
using data labeled by three diferent Subject Mat- which are defined by name and example in the WIPO 3
ter Experts (SMEs). publicly available documentation, focusing on the
subset of claim types that are directed at the nature of the
invention.</p>
    </sec>
    <sec id="sec-2">
      <title>2. Related Work</title>
      <p>
        Due to the significant importance of patent documents
for individuals, enterprises, and industries, there has been
a considerable amount of study and research dedicated to
this domain. These studies cover a wide range of topics
such as patent classification [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ], patent landscaping [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ],
prior art searches [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ], and more. Further, since claims
are the core part of patent documents that define the
legal boundaries to protect inventions, there has been
a concerted efort to utilize them efectively to propose
scientific solutions.
      </p>
      <p>
        As mentioned in the Introduction section patent claims
can have diferent types, [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ] analyzes the occurrence of
process claims in large US patent corpus and reports that To automatically classify a given claim into the
abovesubstantial increase over the last century in such type of described types we used a (rule-based) heuristic method
claims. Further, the authors also developed a patent claim (see section 4.1) to generate training data. This data is
classification tool 2 that recognizes three types of claims then used to train the Patent Claim Type Recognition
namely, process claims, product claims, and product-by- (PCTR) model which is based on a deep neural network
process claims. [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ] investigates the relation between the for automatic claim type identification. Figure 1
illuspatent examination process and the patent’s scope. The trates the architectural design of the PCTR model.
proposed method relies on the claim length and count. PCTR is a transformer-based multi-class classification
Another interesting study performed by [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ], in which model that is capable of assigning the most relevant type
authors first collect patents from diferent countries to to a given patent accurately. Figure 1 illustrates the claim
assess the country’s technological capability by compar- type recognition model, i.e., the deep neural network
ing the number of patents and claims. It turns out that model that has been designed for this study. It consists
patent claims are much more reliable than the number of a transformer block which is integrated as a layer,
of patents to reflect the country’s technological advance- followed by a pooling layer, a dense layer, and a final
ment. softmax layer. The input to the model comprises a claim
      </p>
      <p>In contrast to these approaches, our proposed ap- paired with the document sections such as the title and
proach difers in two main aspects. First, the focus of abstract. Then the output is the type of the given claim,
our work is to build a comprehensive model to identify represented as  ( = |), where  denotes the patent
types of pre-defined patent claims. Second, although our
• Method: recites a sequence of steps that
com</p>
      <p>plete a task or accomplish a result
• Use: depicts intended/inventive application of</p>
      <p>novelty
• Composition: invention pertains to the
chemi</p>
      <p>cal nature of materials/components used.
• Process: claims define a process of manufacture,
it should be noted that WIPO documentation
labels it as “product-by-process”.
• Apparatus: protects an apparatus or device
• System: an assemblage or combination of things
or parts forming a unitary whole</p>
      <sec id="sec-2-1">
        <title>2https://zenodo.org/records/6395308</title>
      </sec>
      <sec id="sec-2-2">
        <title>3https://www.wipo.int/edocs/mdocs/aspac/en/wipo_ip_phl_16/</title>
        <p>wipo_ip_phl_16_t5.pdf</p>
        <sec id="sec-2-2-1">
          <title>Claim Type</title>
        </sec>
      </sec>
      <sec id="sec-2-3">
        <title>System</title>
        <p>Process</p>
        <p>Method
Composition
Apparatus
Use
#Sample</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>4. Experimental Results</title>
      <sec id="sec-3-1">
        <title>This section gives an overview of the generation of the</title>
        <p>training data, test data annotated by subject matter
exclaim type (e.g., apparatus, composition, system, etc.). perts (SMEs), and the experimental outcomes obtained
The model aims to classify the claims based on the pro- from the PCTR models.
vided input text. Figure 1 illustrates an example input of
a patent section and its associated claim. Initially, the text 4.1. Training Data Generation
is tokenized, and the token embeddings serve as input
to the transformer block, with these embeddings being To train the PCTR models which are based on the DL
randomly initialized. architecture, it is crucial to have a suficient amount of</p>
        <p>Classifying patent claims according to given types (as representative training data. Manually, generating
mildefined above) is a challenging task. It should be noted lions of labeled data is expensive and time-consuming.
that in this study, we distinguish claim types based on For the training data generation only simple claims (cf.
their level of complexity as simple claim types and com- Section3), which contain the type information within the
plex claim types. Simple claim types, as the name sug- text are considered. We designed a rule-based heuristic
gests, are straightforward and include claim-type infor- model that is able to label given simple claim texts based
mation within the claim text. On the other hand, complex on a defined regular expression (regex) rule. On the other
claim types either do not include explicit claim-type in- hand, the test set contains only complex claims (cf.
Secformation in the claim text or refer to more than one tion4.2), which do not include the claim type within the
type, making it challenging to identify them. The focus text.
of the developed Patent Claim Type Recognition (PCTR) The designed regex baseline relies on start and end
approach is to automatically identify the type of com- markers. In between such markers are the targets.
plex claims by designing and developing the deep neural Following is an example of the beginning of a patent
network. claim:
3.1. Feature Selection</p>
      </sec>
      <sec id="sec-3-2">
        <title>To train the PCTR model, we primarily utilized textual</title>
        <p>features, including the claim, title, and abstract. Other
potential features, such as CPC/IPC4 information and
additional contextual or structural elements from the
patent document, remain for our future analysis.</p>
        <p>To train the PCTR model with diferent feature
combinations we designed the following versions of it:
PCTR_V1 utilizes claim, PCTR_V2 utilizes claim + title,
PCTR_V3 utilizes claim + title + abstract as input to
perform the claim type prediction. Essentially, the version of
Example 1. "68. A method according to claim 67,"
with "68. A" being a start marker, "according to" an end
marker, and "method" the target.</p>
        <p>End markers are diferent depending on whether they
are dependent or independent claims. Therefore, such
information has to be provided, e.g. by using diferent
tools designed for this purpose. For that, the internally
developed tool has been employed. Three types of
targets are identified. High-probability targets where start
and end markers only capture a single word and that
word matches one of the known claim types
(appara4https://www.epo.org/en/searching-for-patents/helpful-resources/ tus, compound, composition, device, method, process,
ifrst-time-here/classification system, use). Medium-probability targets where several</p>
        <sec id="sec-3-2-1">
          <title>Claim Type</title>
          <p>System
Process</p>
          <p>Method
Composition
Apparatus</p>
          <p>Use
#Sample
words a captured, but one of them is a known claim type.
And lastly, low-probability targets, where more than one
known claim type is captured.</p>
          <p>After applying the heuristic model to the claim texts of
internal sources which encompasses patents from two
diferent sources, namely, the World Intellectual
Property Organization (WIPO), and the US Patent Ofice. 2
million patent samples were selected employing random
sampling and subjected to preprocessing. In other words,
a subset of data points has been selected as a training
set from a larger dataset. The preprocessing involved
removing duplicates as well as invalid claim types,
abstracts, and titles. Finally, we had 2,092,385 samples with
high probability targets, i.e., claim types. All the claims
with low-probability or medium-probability targets are
ifltered. The statistics of the training data are shown in
Table 1.</p>
          <p>It should be noted that to avoid biases towards the
rulebased method for labeling training samples and allow for
the generalization of the trained models, we removed the
claim type information from each patent claim before
feeding them into the PCTR models.
4.2. Test Data Generation</p>
        </sec>
      </sec>
      <sec id="sec-3-3">
        <title>To accurately assess the efectiveness of the PCTR model,</title>
        <p>it is essential to have representative test data that the
model is expected to encounter in real-world scenarios.
To this end, the test data was generated and labeled by
overall 4 SMEs following an iterative process in order
to achieve the required level of agreement. Initially, 300
samples were randomly selected from the two internal
data sources to be labeled by the SMEs. Additionally, 200
samples were also selected from an external database to
balance the test set, ensuring an equal number of
expertlabeled samples for each type. The statistics of the test
set are presented in Table2. For determining the final
claim type, we adopted a majority vote approach.
4.3. Evaluation of PCTR Models</p>
      </sec>
      <sec id="sec-3-4">
        <title>We have trained 3 diferent PCTR models namely:</title>
        <p>PCTR_V1, PCTR_V2, PCTR_V3. Table 3 presents the
accuracy results of the models, calculated as the ratio
of correctly classified data to the total test data. Upon
analyzing these results and consulting with SMEs for
improvement suggestions, we received valuable feedback.</p>
        <p>The analysis by SMEs led to a revised interpretation of
the PCTR model’s performance. Some samples initially
deemed as falsely identified types were actually correctly
identified, according to the feedback of the SMEs. The
essential feedback that we applied to our evaluation
process to improve the performance of the PCTR models is
as follows:
• "Product-by-process" claims should be regarded
as "process" claims.
• "Use" claims can be seen as a sub-category of
"method" claims.</p>
        <p>• Claims can have multiple (2) types.</p>
        <p>Table 4 presents the improved accuracy of the PCTR
models after considering this feedback and re-computing
the accuracy.</p>
        <p>Another aspect has been considered to improve the
models’ performance. The generated dataset through
random sampling is quite unbalanced and generally, it
is a good practice to have a balanced data set for any
machine learning model. In our eforts to achieve a
balanced dataset, we had to reduce the size of the training
data by downsampling the dataset as some claim types
were underrepresented. However, this reduction in the
training data resulted in a drop in accuracy, as there
were fewer samples for each claim type hence a smaller
dataset. Therefore, we decided to use the originally
randomly sampled training set to train the models. It should
be noted that ensuring a balanced training dataset while
randomly sampling is a time-consuming process that
requires expert assistance to ensure an equal number of
training samples for each class or claim type.</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>5. Conclusion</title>
      <sec id="sec-4-1">
        <title>In this paper, we presented an approach for the automatic</title>
        <p>identification of predefined claim types based on a Deep
Learning model. Looking at the performance of the PCTR
model’s performance focusing on both dataset
characteristics and feature selection aspects, we can report the
following results of a deeper analysis.</p>
        <p>Data from diferent Patent Authorities : As claim type
definitions have been discovered to vary by jurisdiction,
it is needed to customize the developed models per patent
ofice (or collections of patent ofices) accordingly. As
stated earlier, the claim type definitions are based on
WIPO documentation.</p>
        <p>The importance of feature selection: The experiments
indicate that including more features, specifically the title,
and abstract, can significantly improve the accuracy of
the PCTR models for claim type identification. In general,
providing more context that can be used to describe and
distinguish each claim type from each other is helpful.
Further, our preliminary experiments with a very small
set of datasets suggest that including CPC information,
improved the accuracy. Nevertheless, including CPC
as a feature requires a systematic evaluation and data
sampling considering the various domains and extraction
of balanced data of suficient size for each claim type. We
leave this as our future work.</p>
        <p>The role of SMEs: The results also demonstrate that
incorporating the feedback and insights from SMEs can
greatly enhance the accuracy of the PCTR models. This
highlights the importance of involving subject matter
experts in the development process. Herewith, starting
from investigated examples in our data analysis it was,
for example, confirmed that a claim can be assigned
several types. Besides that, it turned out that it is viable
to consider several probabilities for the correct target
label (type), as for example in the case of
product-byprocess claims which were predicted widely as process.
With this valuable feedback, we plan to further refine
the PCTR implementation to allow for assigning multiple
type information in future iterations.</p>
        <p>Finally, the developed PCTR model can be applied in
various real-world scenarios: (1) as a stand-alone model
targeting only complex type claims, or (2) as part of a
hybrid system combining a heuristic model (cf. Section 4.1)
with the PCTR model.</p>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>R.</given-names>
            <surname>Dessi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Aras</surname>
          </string-name>
          ,
          <string-name>
            <surname>M.</surname>
          </string-name>
          <article-title>Alam, Exploring the impact of negative sampling on patent citation recommendation</article-title>
          ,
          <source>PatentSemTech</source>
          (
          <year>2023</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>S.</given-names>
            <surname>Niwa</surname>
          </string-name>
          ,
          <article-title>Patent claims and economic growth</article-title>
          ,
          <source>Economic Modelling</source>
          <volume>54</volume>
          (
          <year>2016</year>
          )
          <fpage>377</fpage>
          -
          <lpage>381</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>S.</given-names>
            <surname>Wittfoth</surname>
          </string-name>
          ,
          <article-title>Measuring technological patent scope by semantic analysis of patent claims-an indicator for valuating patents</article-title>
          ,
          <source>World Patent Information</source>
          <volume>58</volume>
          (
          <year>2019</year>
          )
          <fpage>101906</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>C. J.</given-names>
            <surname>Fall</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Törcsvári</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Benzineb</surname>
          </string-name>
          , G. Karetka,
          <article-title>Automated categorization in the international patent classification</article-title>
          ,
          <source>in: Acm Sigir Forum</source>
          , volume
          <volume>37</volume>
          , ACM New York, NY, USA,
          <year>2003</year>
          , pp.
          <fpage>10</fpage>
          -
          <lpage>25</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>S.</given-names>
            <surname>Choi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Lee</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Park</surname>
          </string-name>
          , S. Choi,
          <article-title>Deep learning for patent landscaping using transformer and graph embedding</article-title>
          ,
          <source>Technological Forecasting and Social Change</source>
          <volume>175</volume>
          (
          <year>2022</year>
          )
          <fpage>121413</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>S.</given-names>
            <surname>Bashir</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Rauber</surname>
          </string-name>
          ,
          <article-title>Improving retrievability of patents in prior-art search</article-title>
          ,
          <source>in: European Conference on Information Retrieval</source>
          , Springer,
          <year>2010</year>
          , pp.
          <fpage>457</fpage>
          -
          <lpage>470</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>B.</given-names>
            <surname>Ganglmair</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W. K.</given-names>
            <surname>Robinson</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Seeligson</surname>
          </string-name>
          ,
          <article-title>The rise of process claims: Evidence from a century of us patents, ZEW-Centre for European Economic Research Discussion Paper (</article-title>
          <year>2022</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>A. C.</given-names>
            <surname>Marco</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. D.</given-names>
            <surname>Sarnof</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Charles</surname>
          </string-name>
          ,
          <article-title>Patent claims</article-title>
          and patent scope,
          <source>Research Policy</source>
          <volume>48</volume>
          (
          <year>2019</year>
          )
          <fpage>103790</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>X.</given-names>
            <surname>Tong</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. D.</given-names>
            <surname>Frame</surname>
          </string-name>
          ,
          <article-title>Measuring national technological performance with patent claims data</article-title>
          ,
          <source>Research policy 23</source>
          (
          <year>1994</year>
          )
          <fpage>133</fpage>
          -
          <lpage>141</lpage>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>