<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>UNT at ImageCLEF 2010: CLIR for Wikipedia Images</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Miguel E. Ruiz</string-name>
          <email>Miguel.Ruiz@unt.edu</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Jiangping Chen</string-name>
          <email>Jiangping.Chen@unt.edu</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Karthikeyan Pasupathy</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Pok Chin</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Ryan Knudson</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>University of North Texas, College on Information, Department of Library and Information Sciences</institution>
          ,
          <addr-line>1155 Union Circle 311068 Denton, Texas 76203-1068</addr-line>
          ,
          <country country="US">USA</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>This paper presents the results of the team of the University of North Texas in the Wikipedia image retrieval track of Image-CLEF-2010. Our approach is based on performing translation of the French and German image captions to English and using of Language Models for generating our runs. We also explore the use of complex queries by asking two users to manually build queries based on the original topics distributed. Our results indicate that the approach of translating the image captions is feasible and yields results that are quite competitive with other teams that participated in the same track.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1 Introduction</title>
      <p>This paper presents the results of the UNT team participation in the Wikipedia
retrieval task. Traditionally, the most common approach to solve the cross language
retrieval problem is to perform automatic translation of the user queries into the
language of the document to be retrieved. However, in the presence of short queries
the automatic translation might not have enough context to generate an appropriate
translation. Our main goal was to explore the efficacy of using the captions associated
with the Wikipedia images and providing automatic translations of them in English.
We also address the effectiveness of using this approach using automatic queries as
well as manual queries constructed by real users.</p>
      <p>Section 2 of this paper presents a short background of the CLIR retrieval problem
in image retrieval. Section 3 presents the methods used to conduct our experiments.
Section 4 presents our results and preliminary analysis of results. The last section of
this paper presents our conclusion and plans for future work.</p>
    </sec>
    <sec id="sec-2">
      <title>2 Background</title>
      <p>
        Retrieval of images in multilingual collections is a task that has been studied in CLEF
since 2003
        <xref ref-type="bibr" rid="ref4">(Peters, 2009)</xref>
        . Previous research in CLEF addressing this problem have
explored the use of different resources for translation and for most part concentrated
on combining visual and textual features automatically extracted fro
        <xref ref-type="bibr" rid="ref3">m images
(Müller, et al., 2009</xref>
        ). There has been a lot on emphasis on trying to improve the
current Content-Based Image Retrieval (CBIR) using automatically extracted visual
features to match sample images given in the official topics. However, our own
research as well as the results from other participants has shown that the most
successful approach to solve the image retrieval problem relies on high quality text
retrieval (
        <xref ref-type="bibr" rid="ref3">Müller, et al., 2009</xref>
        ;
        <xref ref-type="bibr" rid="ref5 ref6">Ruiz M. E., 2006</xref>
        ;
        <xref ref-type="bibr" rid="ref6">Ruiz &amp; Névéol, 2007</xref>
        ). The
combination of visual and textual features has proven to contribute to small
improvements in mean average precision (MAP) which is the standard measure that is
used in CLEF to compare system performance.
      </p>
      <p>One of the key issues that needs to be explored is to find approaches that can
contribute more to solve the retrieval problem when the given data collection contains
annotations generated in multiple languages. For the current Wikipedia retrieval task
this is an important issue since the collection has a relatively even distribution of
annotations in three languages (English, French and German). The CLIR problem has
been solved using several methods but the most commonly used approach consists of
translating the text of the user query to the language of the document, performing
monolingual retrieval in each language and then combining the results of several
monolingual runs. However, this approach has two main potential challenges:
1. Use of machine translation on short queries can be difficult due to the loss
of context and finding appropriate disambiguation for automatic
translation.
2. Finding optimal parameters to adjust the mechanism to merge results from
multiple monolingual runs is challenging. Moreover, these optimal
parameters can change from one collection to another making it hard to
find a general optimal set of parameters.</p>
      <p>We decided to explore a solution that translates all the documents to a single
language and perform the retrieval in that language only. We recognize that this is an
expensive solution that might not work for a general CLIR problem. However, for
image retrieval it is a viable option due to the relatively short length of image captions
(compared to the full text associated with the images in an article in Wikipedia). Also,
image captions usually contain enough contextual information to allow appropriate
translation disambiguation. The translation of captions can be achieved relatively fast
using MT translation systems that are freely available on the Internet such as Google
Translation. This also reduces the CLIR problem to a simple monolingual translation
for which the technology is more stable, and there is no need to deal with merging the
results that come from different collections.
3</p>
    </sec>
    <sec id="sec-3">
      <title>Methodology</title>
      <p>
        We used the Indri/Lemur Retrieval system to index our collection using standard
Krovetz stemming and the standard Language Model implemented in Lemur
        <xref ref-type="bibr" rid="ref2">( Lemur
Project, 2001-2008)</xref>
        .
      </p>
      <sec id="sec-3-1">
        <title>Data Collection preparation:</title>
        <p>For our experiments we translated all the French and German texts in the captions
associated with the images into English using the Google Translation service. This
translation was added to the caption using a new field that was indexed together with
the original English caption (if it was available).</p>
      </sec>
      <sec id="sec-3-2">
        <title>Topics preparation:</title>
        <p>For our runs we used just one language at a time from the three provided in the
original ImageCLEFwiki Topics and built topics automatically using a simple
strategy that converted all the words to a “#combine” statement in lemur. We used
first the English topics as our base line. A second run with French topics that were
translated into English was created to measure the effect on query translation. We also
asked two of the members of our group to use the Indri Query Language and create
manual queries that could take advantage of the advanced option of the more
advanced operators in Lemur. For this purpose we made available for these users the
Indri web search engine (which is based on Lemur) and asked them to conduct
searches with the system until they were satisfied with the results that were retrieved.
Each user learned the syntax of the Indri Query Language and then created queries
that tried to use the capabilities of the query language. For example, for our first user
the procedure followed to build the query is described below:</p>
        <p>All query statements used to perform manual image retrieval were built based on
the Indri Query Language. The user developed all seventy manual query statements
using the following methods:
1. The user tried different combinations of keywords to retrieve images from a
sample database.
2. Based on the returned images, the user refined the query statements based on the
following criteria:
a. Incorporate those observable objects within images that can match the
question topics into the query keywords using Indri Query Language
operators such as #combine and #filreq.
b. Reject those observable objects within images that cannot match the
question topics using Indri Query Language operators such as #filrej.
c. The user reviewed the first 50 images returned and reiterated the two
steps mentioned above until the precision of the first 50 images
reached at least 80%.</p>
        <p>For example, for topic number seven. In order to find most images representing
“striking lighting in the sky”, the user tried the method mentioned in 2a to incorporate
all potential keywords that could imply “striking lighting in sky” such as lightning,
day, night, strike, struck, striking, and sky. The user also rejected the keyword
“fighter” using method mentioned in 2b so that the aircraft fighter “lighting” would
not be selected for this question. The final query submitted in the official run for this
topic is:</p>
        <p>#filrej(fighter #filreq(lightning #combine(day night strike struck striking sky)))</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>Results and Analysis</title>
      <p>We submitted three official runs and have an unofficial manual run
 untaTxEn: This run uses automatic query construction using the portion of the
original topic.
 untaTxFr: This run uses automatic query construction using the original
French portion of the text which was translated to English using Google
Translation. (This can be considered a standard CLIR scenario)
 untMan1En: This run uses the final version of the queries created by our first
user for each of the 70 topics.
 untMan2En: This run uses the final version of the queries created by our
second user for each of the 70 topics.</p>
      <p>Comparing the runs using a standard Recall-Precision graph gives a better picture
of the performance of the manual and automatic runs (see Figure 1). The manual runs
in general perform better on the early R-P levels (0-0.2) while the automatic runs
perform better on the higher levels of recall. This seems to be correlated to the
amount of images retrieved which basically indicate that the manual queries are
optimized to generate high precision but low recall (this can also be appreciated in the
total number of retrieved documents in Table 1).</p>
      <p>Regarding the CLIR run that uses French as original language and English as target
(or collection language) we can see that the performance is very close to approach
that translates the documents instead of the queries.
We conclude that the approach of translating the captions instead of the queries is
feasible for image collection and competitive with other more complex approaches
that use more complex algorithms for performing CLIR.</p>
      <p>The manually created queries allowed us to explore potential strategies that could
help in future research that can take advantage of the complex query language
available in Lemur. We still have to do more analysis on the official runs as well as
other unofficial runs that will be included for the extended paper of the proceedings.
We also plan to explore the use of visual features with these queries but need to get a
better understanding of the way users would interact with the system using an
appropriate interface that combines CBIR and CLIR.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          <string-name>
            <surname>Gonzalo</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Clough</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>Karlgren</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          (
          <year>2008</year>
          ).
          <source>Overview of iCLEF</source>
          <year>2008</year>
          :
          <article-title>Search Log Analysis for Multilingual Image Retrieval</article-title>
          .
          <source>Working Notes for the CLEF 2008 Workshop</source>
          . Aarhus, Denmark.
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          <string-name>
            <given-names>Lemur</given-names>
            <surname>Project</surname>
          </string-name>
          (
          <year>2001</year>
          -
          <fpage>2008</fpage>
          ).
          <article-title>The Lemur Project</article-title>
          . ( University of Massachusetts and Carnegie Mellon University)
          <source>Retrieved August 15</source>
          ,
          <year>2010</year>
          , from http://www.lemurproject.org
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          <string-name>
            <surname>Müller</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kalpathy-Crame</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Eggel</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bedrick</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Radhouani</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bakke</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          , et al. (
          <year>2009</year>
          ).
          <article-title>Overview of the CLEF 2009 medical image retrieval track</article-title>
          ,
          <source>CLEF working notes 2009</source>
          . Corfu, Greece.
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          <string-name>
            <surname>Peters</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          (
          <year>2009</year>
          ).
          <article-title>What happened in CLEF 2009: Introduction to the Working Notes</article-title>
          .
          <source>Working Notes for the CLEF 2009 Workshop</source>
          . Corfu, Greece.
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          <string-name>
            <surname>Ruiz</surname>
            ,
            <given-names>M. E.</given-names>
          </string-name>
          (
          <year>2006</year>
          ).
          <article-title>Combining Image Features, Case Descriptions and UMLS Concepts to Improve Retrieval of Medical Images</article-title>
          .
          <source>Proceedings of the Symposium of the American medical Informatics Association</source>
          , (pp.
          <fpage>674</fpage>
          -
          <lpage>8</lpage>
          ). Washington, D.C.
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          <string-name>
            <surname>Ruiz</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>Névéol</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          (
          <year>2007</year>
          ).
          <article-title>Evaluation of Automatically Assigned MeSH Terms for Retrieval of Medical Images</article-title>
          .
          <source>In Revised and selected papers of the 2006 Cross Language Evaluation Forum</source>
          . Springer.
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>