<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>PATExpert: Semantic Processing of Patent Documentation</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Leo Wanner</string-name>
          <email>leo.wanner@upf.edu</email>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Sören Brügmann</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Barrou Diallo</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Mark Giereth</string-name>
          <email>mark.giereth@vis.uni-stuttgart.de</email>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Yiannis Kompatsiaris</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Emanuele Pianta</string-name>
          <email>pianta@itc.it</email>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Gautam Rao</string-name>
          <email>rao@iale.es</email>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Pia Schoester</string-name>
          <email>pia.schoester@pst.fhg.de</email>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Vasiliki Zervaki</string-name>
        </contrib>
      </contrib-group>
      <abstract>
        <p>- PATExpert is a recently started “Specific Targeted Research Project” funded by the EC in FP 6, IST priority. PATExpert's goal is to change the paradigm currently followed for patent processing from textual to semantic. We are about to develop a semantic multimedia content representation based on Semantic Web technologies for selected technology areas and to investigate some central topics from the semantic representation angle: patent retrieval and classification, content extraction, generation of multilingual user comprehensible patent information, visualization of and navigation in patent content spaces, and patent valuing and technology area assessment.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>Index Terms— semantic representation, (multimedia) ontology,
OWL-DL, SUMO, PULO, patent processing techniques.</p>
    </sec>
    <sec id="sec-2">
      <title>I. INTRODUCTION</title>
      <p>Patents belong to the few types of public information that
have a big impact on the European economy, and whose proper
monitoring, retrieval, representation, interpretation, and
assessment so clearly depend on the access to its content, and, thus,
on advances in semantics-based techniques. However, research
and development in the area of patent processing still focuses
on selected traditional tasks such as text retrieval,
classification, and shallow linguistic analysis. Recent initiatives that
target the automatic access to content of patents attempt to cover
ALL knowledge areas. This forces them to rely on term
frequency, term co-occurrence and grammatical term categories.
I.e., despite the use of a Semantic Web-based formalism, the
resulting representation is not a real content representation. As
a consequence, tasks that ultimately require knowledge-based
multimedia techniques (content-oriented search, assessment,
abstracting, etc.) are still, to a major extent, carried out
manually. The overall goal of the PATExpert project, which
started 01.02. 2006 and is funded by the EC (FP6,
IST028116, http://www.patexpert.org), is to change the paradigm
currently followed for patent processing from textual (viewing
patents as text blocks enriched by “canned” picture material,
sequences of morpho-syntactic tokens, or collections of partial
syntactic structures) to semantic (viewing patents as
multimedia knowledge objects) processing. PATExpert is about to
develop a multimedia content representation formalism based
on Semantic Web technologies for selected technology areas
and to investigate some central topics from the semantic
representation angle: patent retrieval and classification, content
extraction, generation of multilingual user comprehensible
patent information, visualization of and navigation in patent
content spaces, and patent valuing and technology area
assessment, taking into account the information needs of all user
types as defined in a user typology. PATExpert’s technological
goal is to develop a showcase that demonstrates the viability
of PATExpert’s approach to content representation for real
applications.</p>
    </sec>
    <sec id="sec-3">
      <title>II. SEMANTIC REPRESENTATION OF PATENTS</title>
      <p>The semantic representation of patent documentation must
cover, on the one hand, propositional, multimedia and
metadata information and, on the other hand, lingustic
knowledge—first of all the characteristic text structures
encountered in patent documentation and the lexical information.
We developed an initial working schema of the knowledge
representation (KR) in PATExpert, with OWL-DL as the KR
language.</p>
      <p>As a rule, the content in patent documentation makes
reference to knowledge of three levels of abstraction: (a)
common sense knowledge, (b) patent-specific knowledge and
terminology, and (c) domain-specific knowledge that refers to
technology area details. PATExpert focuses on the ontologies
of two technology areas: optical recording media and
mechanical engineering tools.</p>
      <p>
        Traditionally, common sense knowledge representation is
dealt with by core ontologies and domain-specific knowledge
representation by domain-ontologies. As core ontology, we
use SUMO [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]. Following a recent trend in semantic web
representation technologies (see, e.g., the Midlevel Ontology,
      </p>
      <sec id="sec-3-1">
        <title>MILO by Teknowledge Knowledge Systems Group [3]), we</title>
        <p>capture patent-specific knowledge and terminology by a
midlevel ontology, called Patent Upper Level Ontology, PULO.
PULO aims to bridge the gap between the high level concept
descriptions in SUMO and the detailed domain ontologies.</p>
        <p>
          Multiple media (in particular images such as photographs,
diagrams, flow charts, drawings, etc.) come into play in the
representation of patent documentation and must thus be
modelled by a multimedia ontology. For linguistic knowledge
representation, we foresee a patent document structure
ontology and lexical (word level) ontologies. As lexical ontologies,
we use WordNets [
          <xref ref-type="bibr" rid="ref1">1</xref>
          ]. Patent-related meta information (such
as the patent holder, inventor and his current affiliation, etc.)
are captured in a separate metadata ontology. Figure 1 shows
the initial working schema of the knowledge representation in
PATExpert. It will be revised as needed when the work on the
tasks that make use of the semantic representation progresses.
        </p>
        <p>Core Ontology: Holistic Representation - SUMO
Patent Upper Level Ontology - PULO
Multimedia Ontologies</p>
        <p>Domain Ontologies</p>
        <p>Linguistic Ontologies</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>III. PROCESSING PATENT DOCUMENTS</title>
      <p>The state of the art in a number of central patent processing
areas suffers from the lack of an adequate representation of the
content and content structure of patent documentation. With
the KR-schema presented above at hand, PATExpert addresses
these areas. The most central of them are listed below. In the
poster presentation, more details will be given on each topic.</p>
      <sec id="sec-4-1">
        <title>A. Information Extraction from Patents</title>
        <p>Extraction of content information (e.g., composition and
function of the invention) and meta information (e.g., the
productivity of an inventor) from patent documentation is one
of the burning issues in patent processing. In PATExpert, this
topic is approached from two angles: as a stand alone task,
and as a way to populate the knowledge base. Strategies using
partial syntactic and semantic analysis, information extraction
techniques, inference mechanisms are being explored.</p>
      </sec>
      <sec id="sec-4-2">
        <title>B. Patent Retrieval and Classification</title>
        <p>The (multi) media content representation of patent
documentation will allow us to develop patent retrieval
strategies that go considerably beyond the state of the art patent
retrieval techniques. In particular, it will facilitate semantic
retrieval, i.e., search for patents that describe inventions with
specific content features, image-based retrieval and document
similarity-based retrieval. It will also allow for a classification
and clustering of patent documents along semantic criteria.</p>
      </sec>
      <sec id="sec-4-3">
        <title>C. Production of Multilingual Patent Information</title>
        <p>The language style in patent documentation is very complex
and repetitive. It is thus hard to comprehend by human readers.
Our goal is to provide the reader with a comprehensible variant
of text passages chosen by him in the language of his choice.
Two topics are addressed: (a) paraphrasing of patent passages
and (b) generation of multilingual gists of given passages. For
both, shallow techniques and deep techniques that draw on
content representation of the text passages in question (and that
implement thus text generation proper) are being developed.</p>
      </sec>
      <sec id="sec-4-4">
        <title>D. Visualization and Navigation in Patent Knowledge Spaces</title>
        <p>Given the complexity of patent knowledge spaces, the
availability of techniques for visualization of patent content
material that is retrieved from the patent KB or selected
by the user while browsing the patent KB is crucial. We
develop techniques that make the complex content structures
transparent (as, e.g., the IS-A, PART-OF, CAUSE, OPERATE,
etc. relations between objects, semantic similarity / entailment
links, etc.) and help navigate through such structures. The
navigation techniques combine browsing mechanisms with
advanced strategies that guide navigation taking the user’s focus,
the context and the discourse relations between knowledge
objects into account.</p>
      </sec>
      <sec id="sec-4-5">
        <title>E. Patent Valuing and Technology Area Assessment</title>
        <p>Currently, high quality valuing of patents and patent
applications and the assessment of technology areas with respect
to their potential to give rise to patent applications is done
mainly manually—which is very costly and time consuming.
We are developing techniques that use statistical and semantic
information from patent (applications/) as well as user based
data for market aspects to prognosticate the value of a patent
(application).</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>IV. THE STATE OF AFFAIRS</title>
      <p>After having developed the initial schema of the knowledge
representation, we work on the topics sketched in Section III.
The first prototypical implementations of the techniques are
planned to be operational by the end of May 2007, some of
them (e.g., the gist generation) already by the end of January
2007.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>C.</given-names>
            <surname>Fellbaum</surname>
          </string-name>
          (ed.),
          <source>WordNet. An Electronic Lexical Database</source>
          . Cambridge, MA: The MIT Press,
          <year>1998</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>I.</given-names>
            <surname>Niles</surname>
          </string-name>
          and
          <string-name>
            <given-names>A.</given-names>
            <surname>Pease</surname>
          </string-name>
          , “
          <article-title>Towards a Standard Upper Ontology,”</article-title>
          <source>in Proceedings of the 2nd International Conference on Formal Ontology in Information Systems (FOIS-</source>
          <year>2001</year>
          ),
          <string-name>
            <given-names>C.</given-names>
            <surname>Welty</surname>
          </string-name>
          and
          <string-name>
            <given-names>B.</given-names>
            <surname>Smith</surname>
          </string-name>
          , Eds., Ogunquit, MA,
          <year>2001</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>I.</given-names>
            <surname>Niles</surname>
          </string-name>
          and
          <string-name>
            <given-names>A.</given-names>
            <surname>Terry</surname>
          </string-name>
          , “
          <article-title>The MILO: A General-Purpose,</article-title>
          <string-name>
            <surname>Mid-Level</surname>
            <given-names>Ontology</given-names>
          </string-name>
          ,”
          <source>in Proceedings of the 2004 International Conference on Information and Knowledge Engineering</source>
          , Las Vegas,
          <string-name>
            <surname>NE</surname>
          </string-name>
          ,
          <year>2004</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>