<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>CACAO PROJECT AT THE TEL@CLEF 2009 TASK</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Alessio Bosca</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Luca Dini</string-name>
          <email>dini@celi.it</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>General Terms</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="editor">
          <string-name>Cross-Language Information Retrieval, Query Expansion, Translations Disambiguation, Digital
Libraries</string-name>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Measurement</institution>
          ,
          <addr-line>Performance, Experimentation</addr-line>
        </aff>
      </contrib-group>
      <abstract>
        <p>This paper presents the participation of the CACAO prototype to the TEL@CLEF 2009 task, an evaluation track focusing on multilingual document retrieval over a collection of library catalogues. CACAO (Cross-language Access to Catalogues And On-line libraries) is an EU project devoted to enabling cross-language access to the contents of a federation of digital libraries with a set of software tools for harvesting, indexing and serching over such data. CACAO project consortium participated both to the monolingual and the bilingual subtasks from TEL@CLEF since they constitute a perfect opportunity in order to test the CACAO system prototype and obtain feedbacks for its enhancement. The prototype showed good performances with respect to the other participants, resulting among the best 5 in the French Monolingual subtrack, and quite a signi cative enhancement in comparison with the past year participation.</p>
      </abstract>
      <kwd-group>
        <kwd>H</kwd>
        <kwd>3 [Information Storage and Retrieval]</kwd>
        <kwd>H</kwd>
        <kwd>3</kwd>
        <kwd>1 Content Analysis and Indexing</kwd>
        <kwd>H</kwd>
        <kwd>3</kwd>
        <kwd>3 Information Search and Retrieval</kwd>
        <kwd>H</kwd>
        <kwd>3</kwd>
        <kwd>4 Systems and Software</kwd>
        <kwd>H</kwd>
        <kwd>3</kwd>
        <kwd>7 Digital Libraries</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>prototype and obtain feedbacks for its enhancement. The prototype showed good performances
with respect to the other participants, resulting among the best 5 in the French Monolingual
subtrack, and quite a signi cative enhancement in comparison with the past year participation.</p>
      <p>This paper is organized as follows. We present the architecture of our system in 2, in 3 we
describe our experiments, the evaluation measures and the evaluation results, and nally conclude
in 4.
2</p>
    </sec>
    <sec id="sec-2">
      <title>CACAO Pro ject</title>
      <p>CACAO (Cross-language Access to Catalogues And On-line libraries) is an EU project funded
under the eContentplus program and proposes an innovative approach for accessing, understanding
and navigating multilingual textual content in digital libraries and OPACs, enabling European
users to better exploit the available European electronic content.</p>
      <p>By coupling sound Natural Language Processing techniques with available information retrieval
systems the project aims at the delivery of a non-intrusive infrastructure to be integrated with
current OPAC and digital libraries. The result of such integration will be the possibility for the user
to type in queries in his/her own language and retrieve volumes and documents in any available
language. CACAO aims at o ering cross-lingual and cross-border access to the content of classical
and digital libraries and enabling users to nd digital content irrespective of the language. In fact,
in a context of interlaced cross-border libraries, such as the ones proposed by META OPAC, the
absence of a cross-language perspective is likely to cause a substantial impasse: if a user wanted
to access a META OPAC including the National Libraries of France, Germany, Italy, Poland and
Hungary, s/he would have to type ve queries in ve di erent languages. Much of the advantage
of having a unique access point is thus lost.</p>
      <p>CACAO project proposes a system based on the assumptions that users look more and more
at library contents using free keyword queries (as those used with a web search engine) rather
than more traditional library-oriented access (e.g. via Subject Heading); therefore, the only way
to face the cross-language issue is by translating the query into all languages covered by the
library/collection (rather than, for instance, translating subject headings, as in the MACS approach,
https://macs.vub.ac.be/pub/). The system will then yield results in all desired languages.</p>
      <p>Validation is another important aspect in the project: all CACAO core technologies are indeed
sound, but they have never been massively deployed in the eld of digital libraries. CACAO aims
at crossing the chasm between sound innovation and adoption by library institutions for real life
purposes. CACAO proposes the development of an infrastructure for multilingual access to digital
content, including an information retrieval system able to search for books and texts in all the
available languages. The core of the search engine takes advantage of information contained in
existing catalogues and texts of the digital libraries that is enriched by means of NLP techniques
such as word sense disambiguation and named entities recognition. The goal of such integration
is to avoid confusing the user by providing irrelevant results due to bad translations and thus
enabling a better access to the digital content.
2.1</p>
      <sec id="sec-2-1">
        <title>Architecture Overview</title>
        <p>The general architecture of the Cacao system could be summarized as the result of the interactions
of few functional subsystems, coordinated by a central manager and reacting to external stimuli
represented by end users queries:</p>
        <p>Harvesting subsystem is in charge of collecting data from digital libraries, abstracting from
the multiplicity of standards and protocols, and storing them into a repository.
Corpus Analysis subsystem performs speci c analysis on the data collected from libraries
and infers new information used to support query processing and resource retrieval (e.g.
query expansion, terms disambiguation,..).
Web Services subsystem represents third party software providing speci c services (e.g.
linguistic analysis, translations,..).</p>
        <p>Query Processing subsystem: a set of components is devoted to process the original
monolingual user query, transforming and enriching it by means of translations and expansions.
3
3.1</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>Experiments</title>
      <sec id="sec-3-1">
        <title>Indexing and Enhancing MetaData</title>
        <p>
          In order to acquire the collections metadata into the CACAO system a speci c harvester module
for importing the XML corpus documents has been deployed. The textual information contained
in the dc:subject, dc:title and dc:dexcription are lemmatised using the XIP incremental parser
from XEROX (see [
          <xref ref-type="bibr" rid="ref1">1</xref>
          ])) and all the data is then indexed using the Lucene open source engine (see
[
          <xref ref-type="bibr" rid="ref3">3</xref>
          ])). By means of lexical semantics technologies (we exploited Random Indexing approach, see
[
          <xref ref-type="bibr" rid="ref2">2</xref>
          ]) a corpus based word space model has been created for each of the TEL@CLEF collections;
these word space resources have been used by the CACAO system as a means to disambiguate
the candidate translations and for query expansion purposes.
3.2
        </p>
      </sec>
      <sec id="sec-3-2">
        <title>Topics Processing</title>
        <p>The approach adopted by CACAO system for dealing with user queries is based on the free
keywords search; therefore while the title eld of TEL topics already tted this model, the description
eld has been processed in order to extract a set of relevant keywords from the sentence. For this
purpose a simple keyword extractor module has been used for each of the main languages present
in the corpus (English, French and German).</p>
        <p>The keywords retrieved in this process are lemmatised and the system assigns di erent weight
to them according to their frequency in both of the topic elds (title and description). In the
lemmatisation process named entities (i.e. person and geographic names) were also identi ed,
and they were treated di erently from common keywords with respect to translation to target
languages.</p>
        <p>According to the subtasks (monolingual or bilingual) the keywords were translated to the target
language or directly submitted to the Lucene search engine.</p>
        <p>The translation process exploits internal resources (inter-lingual indexes or bilingual
dictionaries) and online dictionaries as Ergane; the obtained translation candidates are disambiguated using
the corpus based semantic vectors, computed by the CACAO system on the collections metadata
2.1 and according the following approach:</p>
        <p>As a rst step the system automatically groups keywords in sets of semantically related
terms by comparing their similarity, de ned as the cosine of the angle between the vector
representations of the terms; this process allows the system to group together all the keywords
bearing a common meaning.</p>
        <p>
          Then the translation candidates of each keywords group are analysed in order to prune away
all the elements with a low similarity with the center of the translation group, computed as
the sum of the vector representation of terms (a variation of the algorithm proposed by [
          <xref ref-type="bibr" rid="ref4">4</xref>
          ]).
        </p>
        <p>Experiments involving query expansion enriched the keywords groups (either in the original or
in the target language) by means of di erent strategies :</p>
        <p>Exploiting the corpus based semantic vectors, by adding the N nearest neighbours of each
group center where; the actual value of N depends on the cardinality of the keyword group.
Extracting terms from the titles of the top 10 documents restrieved for each topic
By expanding geographic names items with the list of the geographic entity they are
contained in (i.e. Turin is expanded as Piedmont, Italy, Europe)
By translating the terms in the title eld of TEL topics to other languases then the
default one of the collection in order to capture the multilinguality of the data (i.e. in the
monolingual English task, keywords from the title led are also translated to french and
german)
3.3</p>
      </sec>
      <sec id="sec-3-3">
        <title>Submitted Runs</title>
        <p>The results of the top result for each subtask are provided in the following table:</p>
        <p>Run ID
SEMVECT EXPANDED
DE-EN BASE
ML EXPANDED
EN-FR ML EXPANDED
ALL EXPANDED
EN-DE BASE</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>Conclusions</title>
      <p>The prototype showed good performances with respect to the other participants, resulting among
the best 5 in the French Monolingual subtrack, and quite a signi cative enhancement in comparison
with the past year participation.
5</p>
    </sec>
    <sec id="sec-5">
      <title>Acknowledgements</title>
      <p>This work has been supported and founded by CACAO EU project (ECP 2006 DILI 510035).</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <surname>At-Mokhtar</surname>
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chanod</surname>
            <given-names>J-P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Roux</surname>
            <given-names>C</given-names>
          </string-name>
          .
          <article-title>Robustness beyond shallowness: incremental dependency parsing</article-title>
          <source>NLE Journal</source>
          ,
          <year>2002</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <surname>Sahlgren</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          <article-title>An Introduction to Random Indexing</article-title>
          .
          <source>Proceedings of the Methods and Applications of Semantic Indexing Workshop at the 7th International Conference on Terminology and Knowledge Engineering</source>
          ,
          <string-name>
            <surname>TKE</surname>
          </string-name>
          <year>2005</year>
          ,
          <year>August</year>
          16, Copenhagen, Denmark.
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <surname>Lucene</surname>
          </string-name>
          .
          <article-title>The Lucene search engine</article-title>
          . URL: http://jakarta.apache.org/lucene/.
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>A.</given-names>
            <surname>Bosca</surname>
          </string-name>
          and
          <string-name>
            <given-names>L.</given-names>
            <surname>Dini</surname>
          </string-name>
          .
          <article-title>Query expansion via library classication systems</article-title>
          .
          <source>LNCS proceedings on CLEF@TEL</source>
          ,
          <year>2008</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>