<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Applying Logic Forms and Statistical Methods to CL-SR Performance</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>R. M. Terol</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>P. Mart</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>nez-Barco</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>M. Palomar</string-name>
          <email>mpalomarg@dlsi.ua.es</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="editor">
          <string-name>Speech Retrieval, Information Retrieval, Logic Forms</string-name>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Departamento de Lenguajes y Sistemas Inform</institution>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>aticos Universidad de Alicante Carretera de San Vicente del Raspeig - Alicante - Spain Tel</institution>
        </aff>
      </contrib-group>
      <abstract>
        <p>This paper describes in detail the combination of NLP methods applied to the treatment of logic forms in the topic processing and statistical methods applied to the search engine in the frame of the CL-SR performance. The method that infers the logic form of a topic is based on dependency analysis between the words of the topic. These dependencies between the words of the topic are calculated using the MINIPAR parser. Di®erent combinations of the topic, description and narrative ¯elds are used in the runs to perform the retrieval process. The based on logic forms method processes the description and narrative ¯elds of the topics. This processing task consists on the removal of several terms according to the logic structure of the processed ¯eld in the logic form. On the other hand, the statistical processing applied to the search engine consists on using IR-n system. IR-n system is a passage retrieval system that manages overlapping of variable passages that are composed by a number of sentences. Di®erent statistical similarity measures are managed by IR-n system to acquire the topic terms weight. The removal of several topic terms according to the logic structure of the topic origines that the rest of the topic terms acquire a better relevance.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Introduction</title>
      <p>Di®erent combinations of the topic, description and narrative ¯elds can be applied to perform the
information retrieval process. This fact implies that all the terms of these ¯elds are used by the
search engine to accomplish its goal. The search engine usually removes many terms that can be
considered as stop-words (prepositions, articles and so on). If we have a look to the structure
of the description and narrative ¯elds of a topic (see table 1), we can deduce that there exists
many terms that would not be as relevant as other terms in the information retrieval process. Our
system processes the topics according to an NLP based approach. The topic processing basically
consists on removing several terms of the description and narrative ¯elds of the topic. Obviously,
these removed terms are consider as not relevant and then they will not be processed by the
information retrieval engine.</p>
      <p>In this new Cross-Language Speech Retrieval (CL-SR) Track, our research e®ort has been
focused on combining the use of NLP and statistical methods in the CL-SR performance. Concretely,</p>
      <sec id="sec-1-1">
        <title>Child survivors in Sweden</title>
      </sec>
      <sec id="sec-1-2">
        <title>Description</title>
        <sec id="sec-1-2-1">
          <title>Describe survival mechanisms</title>
          <p>of children born in 1930-1933
who spend the war in
concentration camps or in ...</p>
        </sec>
      </sec>
      <sec id="sec-1-3">
        <title>Narrative</title>
        <sec id="sec-1-3-1">
          <title>The relevant material should</title>
          <p>describe the circumstances and
inner resources of the</p>
          <p>
            surviving children
our main research goal has been centered on demonstrating that the use of NLP methods by way
of processing the topics according to the logic structure of their associated logic form increases
the results obtained by the statistical search engine. We applied the new version of IR-n system
[
            <xref ref-type="bibr" rid="ref1">1</xref>
            ] as statistical search engine.
          </p>
          <p>The following section shows the topic processing by way of applying NLP rules based on the
logic structure of their associated logic forms. Finally, we describe the submitted runs, the obtained
results in these submitted runs, and discuss the application of NLP methods to the statistical IR-n
system.
2</p>
        </sec>
      </sec>
    </sec>
    <sec id="sec-2">
      <title>System Description</title>
      <p>
        This section presents the topic processing applying NLP rules based on logic forms [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]. The format
of the applied logic forms is based on the format of the logic form de¯ned by eXtended
WordNet [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]. For example, the associated logic form of the topic \The liberation of Buchenwald and
Dachau" is instantiated as \liberation:NN(x4) of:IN(x4, x2) buchenwald:NN(x3) and:CC(x2, x3,
x1) dachau:NN(x1)".
      </p>
      <p>The topic processing by way of logic forms consists on removing many terms of the topic
according to the logic structure of its logic form. A combination of the text, description and
narrative ¯elds of the topic has been employed to perform the information retrieval process according
to the submitted runs. The rules are only applied to the description and narrative ¯elds of the
topic because these ¯elds contains a lot of information (see table 1) that would be previously
¯ltered before to be processed by the information retrieval process. These rules are
independently applied to the description and narrative ¯elds of the topic when the number of words of
these ¯elds are upper to 10 words. In other case, there is not necessary to remove any word (term).</p>
      <p>These rules consist on the removal of the ¯rsts words until a preposition (predicate type IN),
or a main verb (predicate type VB or VBE), or a compositional structure (predicate type CC),
both included. If a preposition or a verb are the ¯rsts words in the sentence, we removed them
and then the processing continues until ¯nding another preposition, main verb or compositional
structure. If in this search process the system detects a noun (predicate type NN) coinciding with
the nouns of the topic ¯eld then the search process is aborted until this noun. Table 2 shows how
an example of the application of these rules.</p>
      <p>
        Once these rules are applied, the next process consists on performing the search in the document
collection according to combination of updated ¯elds by the application of these rules. The
statistical IR-n system [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ] accomplishes this goal.
3
      </p>
    </sec>
    <sec id="sec-3">
      <title>Submitted Runs</title>
      <p>This section describes the submitted runs in which our system has participated in. The di®erences
between these ¯ve submitted runs are basically based on the combination of the topic ¯elds and</p>
      <p>Field
Describe survival mechanisms
of children born in 1930-1933</p>
      <p>who spend the war in
concentration camps or in ...</p>
      <p>The relevant material should
describe the circumstances and
inner resources of the
surviving children ...</p>
      <sec id="sec-3-1">
        <title>Logic structure</title>
        <p>describe:VB(e2, x11, e1) survival:NN(x1)</p>
        <p>NNC(x8, x1, x9) mechanism:NN(x9)</p>
        <p>of:IN(x8, x2) child:NN(x2)
bear:VB(e1, x8, x10) in:IN(e1, x4) ...</p>
        <p>relevant:JJ(x1) material:NN(x1)
describe:VB(e1, x1, x5) circumstance:NN(x6)</p>
        <p>and:CC(x5, x6, x3) inner:NN(x2)</p>
        <p>NNC(x3, x2, x4) resource:NN(x4) ...
on the indexation of a combination of di®erent segment ¯elds from the document collection. In all
submitted runs we use the indexing and searching processes developed by our IR-n system using
the English as query language. There is not used any kind of thesaurus terms as keywords in
the indexing and in the searching processes. Following subsections show the features of these ¯ve
submitted runs according to the judgment pool priority order:
² UA TDN FL ASR06BA1A2 Run. In this run IR-n system indexes the combination of
the ASRTEXT2006B, AUTOKEYWORD2004A1 and AUTOKEYWORD2004A2
segment ¯elds of the document collection. The English title, description and narrative topic
¯elds are used in the construction of the queries. This was the unique submitted run in
which we apply the rules based on the topic processing by way of logic forms described in
previous section.
² UA TDN ASR06BA1A2 Run. In this run, as previous submitted run, IR-n system
indexes the combination of the ASRTEXT2006B, AUTOKEYWORD2004A1 and
AUTOKEYWORD2004A2 segment ¯elds of the document collection. The English title,
description and narrative topic ¯elds are used in the construction of the queries.
² UA TD ASR06BA2 Run. In this run IR-n system indexes the combination of the
ASRTEXT2006B and AUTOKEYWORD2004A2 segment ¯elds of the document collection.</p>
        <p>Only the title and description topic ¯elds are used in the construction of the queries.
² UA TDN ASR06BA2 Run. In this run, as in previous run, IR-n system indexes the
combination of the ASRTEXT2006B and AUTOKEYWORD2004A2 segment ¯elds
of the document collection. The English title, description and narrative topic ¯elds are used
in the construction of the queries.
² UA TD ASR06B Run. In this required run, IR-n system only indexes the ASRTEXT2006B
segment ¯eld of the document collection. Only the title and description topic ¯elds are used
in the construction of the queries.
4</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>Results 5</title>
    </sec>
    <sec id="sec-5">
      <title>Conclusion</title>
      <p>In this research we have demonstrated that the previous preprocessing of the topics according to
NLP methods produces an improvement in the statistical retrieval process. The NLP methods are
based on the logic structure of the narrative and description topic ¯elds.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>Fernando</given-names>
            <surname>Llopis</surname>
          </string-name>
          and
          <string-name>
            <given-names>Elisa</given-names>
            <surname>Noguera</surname>
          </string-name>
          .
          <article-title>Combining Passages in Monolingual Experiments with IR-n system</article-title>
          .
          <source>In Workshop of Cross-Language Evaluation Forum (CLEF</source>
          <year>2005</year>
          ), in this volume, Vienna, Austria.
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <surname>Rafael</surname>
            <given-names>M.</given-names>
          </string-name>
          <string-name>
            <surname>Terol</surname>
          </string-name>
          , Patricio Mart¶
          <article-title>³nez-Barco and Manuel Palomar. Applying Logic Forms to Biomedical Q-A</article-title>
          . In
          <source>International Symposium on Innovations in Intelligent Systems and Applications (INISTA</source>
          <year>2005</year>
          ), pages
          <fpage>29</fpage>
          {
          <fpage>32</fpage>
          ,
          <string-name>
            <surname>Istambul</surname>
          </string-name>
          , Turkey,
          <year>Juny 2004</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>S.</given-names>
            <surname>Harabagiu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.A.</given-names>
            <surname>Miller</surname>
          </string-name>
          ,
          <string-name>
            <surname>and D.I. Moldovan.</surname>
          </string-name>
          <article-title>WordNet 2 - A Morphologically and Semantically Enhanced Resource</article-title>
          .
          <source>In Proceedings of ACL-SIGLEX99: Standardizing Lexical Resources</source>
          , Maryland,
          <year>June 1999</year>
          , pp.
          <fpage>1</fpage>
          -
          <lpage>8</lpage>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>