<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>UAIC: Participation in TEL@CLEF task</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Adrian Iftene</string-name>
          <email>adiftene@info.uaic.ro</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Alina-Elena Mihăilă</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Ingride-Paula Epure</string-name>
          <email>paula.epure@info.uaic.ro</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>General Terms Experimentation</institution>
          ,
          <addr-line>Performance, Measurement, Algorithms</addr-line>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>UAIC: Faculty of Computer Science, “Alexandru Ioan Cuza” University</institution>
          ,
          <country country="RO">Romania</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2009</year>
      </pub-date>
      <abstract>
        <p>In 2009 was first time when we built a system in order to participate in the TEL@CLEF competition. In this competition the aim is to build retrieval algorithms on multilingual collections of catalog records from TEL collections. Our system has four main components: the pre-processing and indexing component, the component responsible with applying rules, the translation component and the searching component. This paper presents how we implement these components and how interacts them in order to achieve the desired purpose.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1 Introduction</title>
      <sec id="sec-1-1">
        <title>The aim of the 2009 Ad Hoc1 track was to improve the last year's experience, with the</title>
        <p>same three tasks: Tel@CLEF, Persian@CLEF, and Robust-WSD. The general aim of
the task was to create good reusable test collections for each of them. The main task
offers monolingual and cross-language search on library catalog records in English,</p>
      </sec>
      <sec id="sec-1-2">
        <title>French, and German, organized in collaboration with The European Library (TEL2) (Agirre et al., 2008).</title>
      </sec>
      <sec id="sec-1-3">
        <title>In 2009, TEL@CLEF evaluated retrieval algorithms on multilingual collections of catalog records. As in 2008, the collections were derived from the English, French and German archives of The European Library. The task is to search and retrieve relevant items from collections of library catalog cards.</title>
      </sec>
      <sec id="sec-1-4">
        <title>Data from this year was very different from the news corpora previously used in</title>
        <p>the CLEF ad hoc track, consisting of bibliographic data (document surrogates).</p>
      </sec>
      <sec id="sec-1-5">
        <title>Whereas in the traditional ad hoc task, the user searches for document containing</title>
        <p>information of interest, here the user will be searching to identify which publications
are of potential interest – according to the information provided by the catalog card.</p>
      </sec>
      <sec id="sec-1-6">
        <title>The question the user is asking is “Is the publication described by the bibliographic record relevant to my information need?”.</title>
      </sec>
      <sec id="sec-1-7">
        <title>Three target collections were provided: TEL Catalog records in English (Copyright</title>
      </sec>
      <sec id="sec-1-8">
        <title>British Library (BL)), TEL Catalog records in French (Copyright Bibliothèque</title>
        <p>nationale de France (BnF)), and TEL Catalog records in German (Copyright Austrian</p>
      </sec>
      <sec id="sec-1-9">
        <title>National Library (ONB)). All three collections are to some extent multilingual and</title>
        <p>contain documents (catalog records) in many additional languages.</p>
      </sec>
      <sec id="sec-1-10">
        <title>The way in which we built the system for TEL track and it components are</title>
        <p>presented in Section 2, while Section 3 presents the runs submitted details. Last</p>
      </sec>
      <sec id="sec-1-11">
        <title>Section presents conclusions regarding our participation in TEL 2009 track.</title>
        <p>2</p>
      </sec>
      <sec id="sec-1-12">
        <title>Our system performs the following operations: pre-processing, indexing, applying rules, translation and searching. The Figure 1 presents the system architecture.</title>
      </sec>
      <sec id="sec-1-13">
        <title>English</title>
      </sec>
      <sec id="sec-1-14">
        <title>TEL data</title>
        <sec id="sec-1-14-1">
          <title>Digester</title>
        </sec>
      </sec>
      <sec id="sec-1-15">
        <title>Filtered</title>
      </sec>
      <sec id="sec-1-16">
        <title>TEL data</title>
        <sec id="sec-1-16-1">
          <title>Lucene</title>
        </sec>
      </sec>
      <sec id="sec-1-17">
        <title>Lucene</title>
      </sec>
      <sec id="sec-1-18">
        <title>Index</title>
      </sec>
      <sec id="sec-1-19">
        <title>Final</title>
      </sec>
      <sec id="sec-1-20">
        <title>Result</title>
      </sec>
      <sec id="sec-1-21">
        <title>Initial</title>
      </sec>
      <sec id="sec-1-22">
        <title>Queries</title>
        <sec id="sec-1-22-1">
          <title>Rules</title>
        </sec>
      </sec>
      <sec id="sec-1-23">
        <title>Lucene</title>
      </sec>
      <sec id="sec-1-24">
        <title>Queries 1</title>
        <p>Google
Translate</p>
      </sec>
      <sec id="sec-1-25">
        <title>Lucene</title>
      </sec>
      <sec id="sec-1-26">
        <title>Queries 2</title>
      </sec>
      <sec id="sec-1-27">
        <title>Details about the main system components are presented below.</title>
        <p>2.1 Pre-processing and Indexing</p>
      </sec>
      <sec id="sec-1-28">
        <title>Pre-processing step help us in selection of relevant tags from XML files. We used</title>
      </sec>
      <sec id="sec-1-29">
        <title>Digester3 for selecting just the attributes we are interested in (title and subject). We</title>
        <p>have an xml configure file named "playRules.xml" for Digester, where we put the
rules that Digester must respect when we parse the xml files (from the library).</p>
      </sec>
      <sec id="sec-1-30">
        <title>At this step we use Lucene4 (Hatcher and Gospodnetic, 2005). We used Lucene in</title>
        <p>order to index the XML files filtered with Digester and after that to search in this
index the relevant documents.</p>
        <p>First we created an object of Document type in Lucene terminology, which will be
assigned to each article that we want to index. We considered the fields: “docno”,
“title” and “subject” and we filled them with the corresponding values for the article
that we want to index. Pairs (field, value) represents the terms of the current
document. The only condition is that the name has to be a String. Add method for the
document will take a Field object type that we build using one of static methods in
class Field. In the end, we need to consider an IndexWriter for indexing of the</p>
      </sec>
      <sec id="sec-1-31">
        <title>Document.</title>
      </sec>
      <sec id="sec-1-32">
        <title>The Document class has a get() method that can be used to extract the information</title>
        <p>that was stored in the index. For example, to get the author from the Document we
would code doc.get("author").</p>
      </sec>
      <sec id="sec-1-33">
        <title>Since we added the article itself as Field.UnStored, attempting to get it will return</title>
        <p>null. However, since we added the URL of the article to the index, we can get the</p>
      </sec>
      <sec id="sec-1-34">
        <title>URL and display it to the user in our result list.</title>
        <p>2.2 Applying Rules
At this step, in original query in English we identify for every word the lemma and
after that accordingly with the result, we build a new query using the following rules:
• If there are named entities in the query, these entities will be included in title
field and also in the subject field, having, of course, a greater relevance in
comparison with other words (we used boost factor 2, instead of default
boost factor which is 1);
• Else every element of the query is added in the title field, as well as in the
subject field;
• Also, any element of the query (non-named entity) is searched in the Lemma
file and thus obtaining the lemmas corresponding to the word. These lemmas
will be included in the Lucene query having a smaller relevance (boost factor
0.75) than the original word.</p>
      </sec>
      <sec id="sec-1-35">
        <title>In the end, a complex Lucene query is obtained, which is to be sent to the next</title>
        <p>module (Translation), so that the Lucene query can be translated.</p>
      </sec>
      <sec id="sec-1-36">
        <title>3 Digester: http://commons.apache.org/digester/</title>
      </sec>
      <sec id="sec-1-37">
        <title>4 Lucene: http://lucene.apache.org/</title>
        <p>For example the query will look like:
&lt;query&gt;
&lt;identifier&gt;10.2452/701-AH&lt;/identifier&gt;
&lt;lang&gt;en&lt;/lang&gt;
&lt;text&gt;subject:document^0.7 subject:documenting^0.7
((subject:documents subject:species)^2.0)
(subject:fauna^2.0 subject:about^2.0 subject:arctic^2.0)
(title:arctic title:animals)&lt;/text&gt;
&lt;/query&gt;
And the result will be:
10.2452/701-AHQ0 010786904 0 0.15470393 runDefault1
10.2452/701-AHQ0 010786905 1 0.15470393 runDefault1
10.2452/701-AHQ0 010786906 2 0.15470393 runDefault1
10.2452/701-AHQ0 011249616 3 0.15470393 runDefault1
10.2452/701-AHQ0 011249617 4 0.15470393 runDefault1
……………………………………………………………
2.3 Translation</p>
      </sec>
      <sec id="sec-1-38">
        <title>We used the Google Java Language API5 to perform the translation of Lucene query from English to French and German.</title>
      </sec>
      <sec id="sec-1-39">
        <title>The Lucene queries were received in an XML document structured as: the root</title>
        <p>element we have tag query, and the children are id, text (that was the query itself) and
the language (en in our case). Every query had an id assigned which helped us keep
the track of the way the translation was performed and the final result was generated
in the result XML.</p>
      </sec>
      <sec id="sec-1-40">
        <title>We created a DOM parser to obtain every Lucene query, then we processed it</title>
        <p>using java.util.regex API for pattern matching with regular expressions, so we could
separate the words that needed translation from the ones that don’t.</p>
      </sec>
      <sec id="sec-1-41">
        <title>Finally, after the translation was performed we merged the three queries: the original one, the one in French and the one in German, and created the output file that had basically the same structure. These queries were put in the result.xml file and after that they are sending to the index module.</title>
        <p>An example of translate query looks like:
&lt;query&gt;
&lt;identifier&gt;10.2452/701-AH&lt;/identifier&gt;
&lt;text&gt;
subject:document^0.7 subject:documenting^0.7
((subject:documents subject:species)^2.0) (subject:fauna^2.0
subject:about^2.0 subject:arctic^2.0)
(title:arctic title:animals) subject:document^0.7
subject:documentation^0.7 ((subject:documents
5 Google Java Language API: http://code.google.com/p/google-api-translate-java/
subject:espÃ¨ces)^2.0) (subject:faune^2.0 subject:Ã propos
de^2.0 subject:Arctique^2.0) (title:Arctique title:animaux)
subject:Dokument^0.7 subject:Dokumenting^0.7
((subject:Dokuments subject:Arten)^2.0) (subject:Fauna^2.0
subject:Ã¼ber^2.0 subject:Arktis^2.0) (title:Arktis
title:Tiere)
&lt;/text&gt;
&lt;/query&gt;
2.4 Searching</p>
      </sec>
      <sec id="sec-1-42">
        <title>At this part we used again Lucene and we created an IndexSearcher in order to extract</title>
        <p>relevant information from previous created Lucene index.
For searching we considered relevant the following values:
• Measure of documents relevance to a query,
• Factors:
o tf: factor of term frequency in document,
o idf: factor of documents with term in index,
o boost: field-level boost,
o coord: factor-based # of query terms in document,
o queryNorm: normalization for query weights.</p>
      </sec>
      <sec id="sec-1-43">
        <title>For example when we are looking for “Arctic Animals” with description “Find</title>
        <p>documents about arctic fauna species” the query will look like:
title:arctic animals AND "arctic fauna species"
with meaning: search for arctic animals in field title from Lucene index and search
for arctic fauna species in field subject which is default searching field in Lucene
index.</p>
      </sec>
      <sec id="sec-1-44">
        <title>As we can see the search will search simultaneously in two fields of the index</title>
        <p>created: title and subject.</p>
      </sec>
    </sec>
    <sec id="sec-2">
      <title>3 Submitted Runs</title>
      <sec id="sec-2-1">
        <title>We submitted three runs with the following characteristics:</title>
        <p>Run 1:
Run 2:
subject has boost 2 and the title has boost 1(default)
[entities have boost 2 and lemmas have boost 0.7]
[subject and title were given in l languages English, French and German]
subject has boost default and the title has boost 2
[entities have boost 2 and lemmas have boost 0.7]
Run 3:
[subject and title were given in l languages English, French and German]
subject and title have the same boost
[entities have boost 2 and lemmas have boost 0.7]
[subject and title were given in l languages English, French and German]</p>
      </sec>
      <sec id="sec-2-2">
        <title>The best result was obtained for Run 1. Statistics for this run are presented in below statistics.</title>
      </sec>
      <sec id="sec-2-3">
        <title>The presented system has four components: (1) pre-processing and indexing, (2)</title>
        <p>applying rules, (3) translation and (4) searching. First component after selection of
relevant tags with Digester, uses Lucene libraries in order to create a Lucene index.</p>
      </sec>
      <sec id="sec-2-4">
        <title>Second component transforms the initial query from natural language to Lucene query</title>
        <p>and it adds to Lucene query boost factors. Third component translates the Lucene
query from English to French and German and obtained a more complex Lucene
query with elements in all three languages. The last component searches in Lucene
index using Lucene queries and offered like output the system result.</p>
      </sec>
      <sec id="sec-2-5">
        <title>From three runs submitted the first run was the better.</title>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>Acknowledgements</title>
      <sec id="sec-3-1">
        <title>The authors would like to thank to the students Ionuţ-Alexandru Bujdei, George</title>
      </sec>
      <sec id="sec-3-2">
        <title>Leonard Chetreanu, Alexandra Apopoaiei and their colleagues from group 2B, second year, for their help and support at different stages of system development.</title>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Agirre</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Di Nunzio</surname>
            ,
            <given-names>G. M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ferro</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Mandl</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          and
          <string-name>
            <surname>Peters</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          : CLEF 2008:
          <article-title>Ad Hoc Track Overview</article-title>
          .
          <source>In Proceedings of the CLEF 2008 Workshop</source>
          . 17-
          <fpage>19</fpage>
          September. Aarhus, Denmark. (
          <year>2008</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Hatcher</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          and
          <string-name>
            <surname>Gospodnetic</surname>
            ,
            <given-names>O.</given-names>
          </string-name>
          :
          <article-title>Lucene in action</article-title>
          .
          <source>Manning Publications Co</source>
          .
          <article-title>(</article-title>
          <year>2005</year>
          )
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>