<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>NLEL-MAAT at CLEF-IP</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Santiago Correa</string-name>
          <email>scorrea@dsic.upv.es</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Davide Buscaldi</string-name>
          <email>dbuscaldi@dsic.upv.es</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Paolo Rosso. NLE Lab</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>ELiRF Research Group</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Universidad Politécnica de Valencia</institution>
          ,
          <country country="ES">Spain</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>This report presents the work carried out at NLE Lab for the IP@CLEF-2009 competition. We adapted the JIRS passage retrieval system for this task, with the objective to exploit the stylistic characteristics of the patents. Since JIRS was developed for the Question Answering task and this is the first time its model was used to compare entire documents, we had to carry out some transformations on the patent documents. The obtained results are not good and show that the modifications adopted in order to use JIRS represented a wrong choice, compromising the performance of the retrieval system.</p>
      </abstract>
      <kwd-group>
        <kwd>Passage Retrieval</kwd>
        <kwd>Intellectual Property</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Introduction</title>
      <p>relating to 1.022.388 patents. The provided documents are encoded in XML format, emphasizing these sections:
title, language, summary and description.</p>
      <p>A total of 500 patents are analyzed using the supplied corpus to determine their prior art; for each one of them
the systems must return a list of 1000 documents with their score ranking.
3</p>
    </sec>
    <sec id="sec-2">
      <title>The passage retrieval engine JIRS</title>
      <p>The passage retrieval system JIRS is a based on n-grams (an n-gram is a sequence of n adjacent words). JIRS has
the ability to find structures of words sequences in a large collection of documents quickly and efficiently,
through the use of different n-grams models. In order to do this, JIRS searches for all possible n-grams of the
words sequence in the collection and it gives them a weight in relation to the n-grams quantity and weight that
appear in these passages. E.g.: suppose to search a document collection, using the JIRS system, in order to find
articles related to the phrase: “anti-lock braking system”; The system could retrieve the following two passages:
“…braking system consists of disk brakes …” and “…anti-lock braking system developed by…”. In a standard
IR engine the first passage would obtain a higher weight due to the repetitions of the words with the “brak” stem.
In JIRS the second passage is ranked higher because of the presence of the 3-gram “anti-lock braking system”. In
order to calculate the n-grams weight of each passage, first of all it is necessary to identify the most relevant
ngram and assign to it a weight equal to the sum of all term weights. The weight of each term is set to:
  = 1 − log ⁡(  )</p>
      <p>1+log ⁡( )
    ,</p>
      <p>=  =1  ∈ ℎ  ,</p>
      <p>=1  ∈ ℎ  , 
ℎ  ,  
=
0

 =1</p>
      <p>∈  
  ℎ</p>
      <p>Where nk is the number of passages in which the term appears and N is the total number of passages in the
system.</p>
      <p>The target is to establish a measure of similarity between a passage (d) and a text (q).</p>
      <p>
        The function ℎ  ,   , in the equation (
        <xref ref-type="bibr" rid="ref2">2</xref>
        ), returns a weight for the j-gram x with respect to the set of j-grams
(  ) in the passage and is defined by:
      </p>
      <p>A more detailed description of the system JIRS can be found in [3].
4</p>
    </sec>
    <sec id="sec-3">
      <title>Approach used</title>
      <p>
        The objective was to use the JIRS PR system to detect plagiarism between a candidate patent and any other
invention described in the prior art. We suppose that a high similarity value between the candidate patent and
another patent in the collection corresponds to the fact that the candidate patent does not represent an original
invention. A problem in carrying out this comparison is that JIRS was designed for the QA task, where the input
is a question: the JIRS model was not developed to compare a full document to another one but only a sentence
(the question) to documents (the passages). Therefore, it was necessary to determine a strategy to summarize the
candidate patent in a sequence of words that could be used as a query for JIRS. The summarization technique is
based on the random walks method proposed by Hassan et al. [4]: the query is composed by the title of the patent
followed by the most relevant n-grams composed by the heaviest terms, according to the weights assigned using
the random walks method, assuming a window size of 2 words.
(
        <xref ref-type="bibr" rid="ref1">1</xref>
        )
(
        <xref ref-type="bibr" rid="ref2">2</xref>
        )
(
        <xref ref-type="bibr" rid="ref3">3</xref>
        )
NLEL-MAAT at CLEF-IP 3
For instance, consider the patent EP-1445166 “Foldable baby carriage”, having the following abstract:
“A folding baby carriage (20) comprises a pair of seating surface supporting side bars (25)
extending back and forth along both sides of a seating surface in order to support the seating
surface from beneath. Each seating surface supporting side bar (25) has a rigid inward
extending portion (25a) extending toward the inside so as to support the seating surface from
beneath, at a rear portion thereof. The inward extending portion (25a) is formed by bending a
rear end portion of the seating surface supporting side bar (25) toward the inside.”
The random walks method extracts the relevant n-gram seating surface from the patent document, composed by
the heaviest terms occurring in the document. The resulting query is “Foldable baby carriage, seating surface”.
      </p>
      <p>Another problem was to transform the patents in documents that could be indexed by JIRS. In order to do so,
we decided to eliminate all the irrelevant information, extracting from each document its title and the description
in the original language in which it was submitted. Each patent has also an identification number, but often the
identification number is used to indicate that the present document is a revision of a previously submitted
document: in this case it is necessary to examine all documents that are part of a same patent and remove them
from the collection. With these transformations we obtained a database that was indexed by the search engine
JIRS, in which each of the patents was represented by a single passage.
5</p>
    </sec>
    <sec id="sec-4">
      <title>Results</title>
      <p>We submitted one run for the task size S (500 topics), obtaining the following results:</p>
    </sec>
    <sec id="sec-5">
      <title>Conclusions</title>
      <p>The obtained results were not satisfactory, possibly due to the reduction process carried out on the provided
corpus; however we believe that the assumptions made in the approximation still constitute a valid approach,
capable of returning appropriate results; in the future, we will attempt to study how to reduce the database size in
order to delete as less amount of relevant information as possible.</p>
      <p>The development of the queries regarding each of the patents is one of the weaknesses which must be taken
into account for future participations: it will be necessary to refine or improve the summarization process and to
compare this model to other summarization models and other standard similarity measures between documents.</p>
    </sec>
    <sec id="sec-6">
      <title>Acknowledgements</title>
      <p>The work of the first author has been possible thanks to a scholarship funded by Maat Knowledge in the context
of the joint project with the Universidad Politécnica de Valencia “Módulo de servicios semánticos de la
plataforma G”. We also thank the TEXT-ENTERPRISE 2.0 TIN2009-13391-C04-03 research project.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <surname>Callan</surname>
            ,
            <given-names>J. P.</given-names>
          </string-name>
          <year>1994</year>
          .
          <article-title>Passage-level evidence in document retrieval</article-title>
          .
          <source>In Proceedings of the 17th Annual international ACM SIGIR Conference on Research and Development in information Retrieval</source>
          (Dublin, Ireland,
          <source>July 03 - 06</source>
          ,
          <year>1994</year>
          ). W. B. Croft and C. J. van Rijsbergen, Eds.
          <source>Annual ACM Conference on Research and Development in Information Retrieval</source>
          . Springer-Verlag New York, New York, NY,
          <fpage>302</fpage>
          -
          <lpage>310</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <surname>Buscaldi</surname>
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sanchis</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rosso</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          <article-title>N-gram vs</article-title>
          .
          <source>Keyword-based Passage Retrieval for Question Answering. Lecture Notes in Computer Science</source>
          . Vol.
          <volume>4730</volume>
          pp.
          <fpage>377</fpage>
          -
          <lpage>384</lpage>
          ,
          <year>2007</year>
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <surname>Buscaldi</surname>
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Rosso</surname>
          </string-name>
          , JM Gomez,
          <string-name>
            <surname>E.</surname>
          </string-name>
          <article-title>Sanchis Answering Questions with an n-gram based Passage Retrieval Engine</article-title>
          .
          <source>Journal of Intelligent Information Systems (82)</source>
          ,
          <year>2009</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <surname>Hassan</surname>
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Mihalcea</surname>
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Banea</surname>
            <given-names>C.</given-names>
          </string-name>
          ,
          <article-title>Random-Walk TermWeighting for Improved Text Classification</article-title>
          . Department of Computer Science University of North Texas.
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>