<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Information retrieval of visual descriptions with IR-n system based on passages</article-title>
      </title-group>
      <contrib-group>
        <aff id="aff0">
          <label>0</label>
          <institution>Sergio Navarro, Fernando Llopis, Rafael Mun ̃oz Guillena, Elisa Noguera Grupo de Investigacio ́n en Procesamiento del Lenguaje Natural y Sistemas de Informacio ́n Departamento de Lenguajes y Sistemas Informa ́ticos University of Alicante</institution>
          ,
          <country country="ES">Spain</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>This paper describes an approach made to the development of a textual image retrieval system, by the university of Alicante using IR-n, a Information Retrieval (IR) textbased system. With only a minimal quantity of adaptations to the features of this task, our system has obtained precision results over the mean average of participants at ImageCLEF07: for English (0.1604 vs 0.1388) and for Spanish (0.1482 vs 0.1450). For German, our results were under the mean (0.0991 vs 0.1331), it could be due to our system does not incorporate a splitter for the treatment of this agglutinative language. We obtain these results, without incorporate specific adaptations of the dominion of the recovery of images. This leads us to believe that we start from a good base point from which to work to obtain better results.</p>
      </abstract>
      <kwd-group>
        <kwd>Question answering</kwd>
        <kwd>Questions beyond factoids</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>
        For our task of the ImageCLEF we have used the IR, IR-n [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ]. It’s a information retraival
system that using statistical techniques, has given good results in flat text based tasks. The
objective to use a system of these characteristics is to contrast, in the scope of the recovery of
images, the results of a statistical system, with others which includes NLP techniques.
      </p>
      <p>This paper is structured as follows: Firstly, presents the main characteristics of the IR-n
system; the following section explains the task with which we have evaluated the system and the
carried out training; finally, in the last section we present the results and the conclusions.
2</p>
    </sec>
    <sec id="sec-2">
      <title>IR-n System</title>
      <p>
        For this approach we have used IR-n, that is a IR based on passages. The RP systems, consider
each document like a set of passages, where a passage defines as a portion or contiguous text
block. This kind of systems opposite to those that are based on documents, allow to consider the
proximity of appearance of the words in a document, to value their relevance [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ].
      </p>
      <p>
        IR-n like RP system difference of other systems pertaining to the same category, in the method
proposed for the definition of the passages. The unit that uses the IR-n system to define the
passages is the phrase. Thus, the passages are defined by a number of consecutive phrases of the
document. [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ].
      </p>
      <p>In this section the main characteristics of the IR-n system are described, and the techniques
used for ImageCLEF 2007 are detailed.
2.1</p>
      <sec id="sec-2-1">
        <title>Resources: stemmers and stopword list</title>
        <p>Stemmers and stopword list are used with the objective to discriminate the information that is
going to be used for retrieval. Thus, the stopword list of each language contains those words that,
in spite of appearing in the query, does not consider important its appearance in the document
to determine if this one is relevant or no. As far as stemers, these, obtain the root of a word
eliminating the suffixes and prefixes of the same ones, for their indexing and search.</p>
        <p>IR-n uses stemmers and the stopword list available in the web www.unine.ch/info/clef.
2.2</p>
      </sec>
      <sec id="sec-2-2">
        <title>Weighting models</title>
        <p>
          The weighting models allow to quantify the similarity between a text (a complete document or
a passage) and the query. These measures are based fundamentally on the terms that share the
text and the query, as well as on the discriminatory importance of each term. IR-n uses several
weighting models. For this competition we have used dfr [
          <xref ref-type="bibr" rid="ref3">3</xref>
          ] and okapi [
          <xref ref-type="bibr" rid="ref6">6</xref>
          ]. The document ranking
produced by each weighting model is obtained using the same general expression, this one is
defined as the product of the weight of a term in the document by the weight of the term in the
query.
        </p>
        <p>sim(q, d) =</p>
        <p>X wt,p · wt,q
t∈q∧p
(1)
Variables List Here is described the list of variables used in the following formulas. 2, 3:
• ft,p is the frequency of the term t on the passage p,
• ft,q is the frequency of the term t on the query q,
• n is the number of documents in the collection,
• nt is the number of documents in which it appears t,
• c, k1, b and k3 are constant values,
• ld is the length of the document,
• avgld is the average of the length of the documents
Okapi Using the okapi model, the weight of a passage p for a query q is given by:
wt,q =
wt,p =
(k1 + 1) · ft,p</p>
        <p>K · ft,p
(k3 + 1) · ft,q · wt</p>
        <p>k3 · ft,q
wt = log2
K = (1 − b) + b · ld</p>
        <p>avrld
n − nt + 0.5</p>
        <p>Using this model, the weight of a passage p for query q is given by:
wt,p = (log2(1 + wt) + wt′,p · log2( 1 + wt )) ·</p>
        <p>wt
′
wt,p = ft,p · log2(1 +
wt,q = ft,q
ft + 1
nt · (wt′,p + 1)
c · avrld )</p>
        <p>ld
wt =
ft
n
2.3</p>
      </sec>
      <sec id="sec-2-3">
        <title>Query expansion</title>
        <p>
          Most IR systems use query expansion techniques [
          <xref ref-type="bibr" rid="ref4">4</xref>
          ] based on adding the most frequent terms
contained in the most relevant documents to the original query. The IR-n architecture allows us
to use query expansion based on either the most relevant passages or the most relevant documents.
In previous researches, we obtained better results using the most relevant passages.
3
        </p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>Training</title>
      <p>IR-n is a parametrizable system, that allows to adapt it to the concrete characteristics of the task
to make. The parameters for his configuration are: the number of phrases that form a passage,
the weighting model to use, the type of expansion, the number of documents/passages on which
the expansion is based, and the average number of words by document.</p>
      <p>This section describes the training process which has been carried out in order to obtain the
best features to improve the performance of the system. Firstly, the collections and resources are
described, and in following section, the specific experiments carried out.
3.1</p>
      <sec id="sec-3-1">
        <title>Data Collections</title>
        <p>We have participated in the following monolingual tasks of ImageCLEF 2007: English, German
and Spanish. For its training with English and German we have been based on corpus of previous
years. Table 3.1 shows the characteristics of the language collections.</p>
        <p>• WDAvg: is the average of words by document.
• NrQue: is the number of queries that are used in the experiments on each collection.
(2)
(3)</p>
        <p>Colection
St Andrews(ImageCLEF2004)
IAPR TC-12 (ImageCLEF2006)
IAPR TC-12 (ImageCLEF2006)</p>
        <p>To the Spanish we did not have previous corpus, that’s the reason why we have used as guide
the results obtained for the English and German.</p>
        <p>As we can see at Table, the difference between corpus of the year 2004 and corpus of the year
2006 is substantial, since the one of the 2004 presents greater amount of information (greater
number of documents and greater average of words by document). Furthermore, at corpus of 2006
only a 70% of the documents has all the information, the 10% hasn’t description, another 10%
has only location and date, and finally last 10% hasn’t any annotation. All of this, increases the
possibilities of success in a task of textual information retrieval over the corpus of 2004.</p>
        <p>The collections of data has a semi structured format. We took advantage of it selecting the
information used as input of the IR-n system.</p>
        <p>The information that is not interesting for a textual recovery is rejected. Thus, we used as input
for IR-n only the fields that correspond with the title (TITLE), the description (DESCRIPTION),
notes (NOTES), the place (LOCATION) and the date of the photo (DATES).</p>
        <p>The queries associated to each corpus (60 queries by corpus) also has a semi structured format
as is possible to see in Figure 1, and in addition the text contains images with similar contents to
the target to retrieve.
&lt;num&gt; Number: 1 &lt;/num&gt;
&lt;title&gt; accommodation with swimming pool &lt;/title&gt;
&lt;narr&gt; Relevant images will show the building of an accommodation facility
(e.g. hotels, hostels, etc.) with a swimming pool.</p>
        <p>Pictures without swimming pools or without buildings are not relevant.
&lt;/narr&gt;
&lt;image&gt; images/03/3793.jpg &lt;/image&gt;
&lt;image&gt; images/06/6321.jpg &lt;/image&gt;
&lt;image&gt; images/06/6395.jpg &lt;/image&gt;
&lt;/top&gt;</p>
        <p>Only the queries in English of ImageCLEF 2006 and ImageCLEF 2004 contain narrative
(NARR) that accompanies the query, the rest of sets of queries of ImageCLEF 2006 in other
languages only contains title (TITLE) as textual information.</p>
        <p>Furthermore from a set of 60 queries in ImageCLEF 2006, only 30 were responded with textual
information, 10 with visual information and 20 indifferently with textual or visual information.
3.2</p>
      </sec>
      <sec id="sec-3-2">
        <title>Experiments</title>
        <p>The aim of the experiment phase is set up the optimum value of the input parameters for each
collection. Next, we describe the input parameter of the system:
• Size of the Passage (sp): Number of phrases that form the passage.
• Weight model (wm): We used two weighting models : okapi y dfr.
• Opaki parameters: k1, b and avgld (k3 is fixed as 1000).
• Dfr parameters: c and avgld.
• Query expansion parameters: If exp has value 1, this denotes we use relevance feedback
based on passages in this experiment. But, if exp has value 2, the relevance feedback is
based on documents. Moreover, num denote the number of passages or documents that the
expansion will use, and term indicates the k terms extracted from the best ranked passages
or documents from the original query
• Evaluation Measure: MAP Mean average precision (avgP) is the evaluation measure
used in order to evaluate the experiments.
3.2.1</p>
        <p>English
As we can see for corpus of 2004 at Table 2, expansion based on passages and dfr obtains better
results. It is important to stand out that, better results are obtained establishing the average
length (in bytes) by document to values greater than corpus has.</p>
        <p>For 2006 English corpus we can see that the precision is reduced considerably. This is justified
by the fact that corpus has minor amount of textual information that in the edition of the 2004,
as reflects Table 3.1. And too because corpus of the 2005 has a 30% of incomplete documents.
Also we observed, that the best results are obtained with dfr and techniques of expansion based on
documents. For this corpus we see that the average length by document, has been standardized,
corresponding with numbers nearer the real ones.
exp
num</p>
        <p>term
exp
num</p>
        <p>term
1
1
2
1
10
5
5
10
The results in German are lower as we can see at Table 4.</p>
        <p>It could be a combination of two causes: on the one hand the fact that IR-n does not
incorporate a mechanism for the treatment of compound words (in spite of being very common in
this language), on the other hand the circumstance that the queries in German do not contain
narrative, which reduce the success possibilities.</p>
        <p>For German the weighting model that better results gives back is okapi, with expansion by
documents.
b
exp
num</p>
        <p>term</p>
        <p>These results were obtained with the best configuration for each corpus and language. For
our participation at ImageCLEF07 we used the same configuration than the best one used with
ImageCLEF06 corpus. The decission is justified by the fact that the corpus for ImageCLEF07 is
the same that we used at ImageCLEF06, and it will improve our success rate. Despite the fact
that we have discarded the 2004 results for the training process, they have helped us to meassure
how the features of the corpus can affect the results.
4</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>Results at ImageCLEFPhoto-2007</title>
    </sec>
    <sec id="sec-5">
      <title>Conclusion and future work</title>
      <p>In our first ImageCLEF participation, we used an IR text-based system. We used it with a minimal
quantity of adaptations to the features of this task. We highlight that its precision results are
above average for English and Spanish. The lower results in German could be due to IR-n not
incorporate a mechanism for the treatment of compound words (in spite of being very common in
this language). We will work to solve this in future projects.</p>
      <p>In addition, analysing the results of training with the corpus of previous years, makes us
meassure how the no completeness of the corpus and the absence of theme tags (like in 2004
corpus) result in decrease of precision values, and moreover the reduction of the length of the
corpus has a direct effect over the passage size.</p>
      <p>Finally, the fact to obtain these results, without to have incorporated specific adaptations of
the dominion of the recovery of images, makes us to think, that we start from a good base point,
from which work, to obtain better results.</p>
      <p>
        To continue improving the system there are several ways that can be taken in account. One
of them is to consider to do an approach to the resolution of a multilingual corpus using the
calculation of the documentary relevance in two steps [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ]. Another work to be developed is
to improve the system with the incorporation of a system CBIR, that complements the textual
information retrieval that IR-n carries out. Therefore, as future work we will explore the possibility
of shape extraction from images, establishing a relationship between them and associated terms.
And finally, we will try to add NLP techniques to the local query expansion.
6
      </p>
    </sec>
    <sec id="sec-6">
      <title>Acknowledgements</title>
      <p>This research has been partially funded by the Spanish Government under project TEXT-MESS
(TIN-2006-15265-C06-01), and by European Union (EU) under QALL-ME project
(FP6-IST033860), and by the Valencia Government under project number GV06-161.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>[1] http://ir.shef.ac.uk/imageclef.</mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>[2] http://www.clef-campaign.org.</mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>G.</given-names>
            <surname>Amati</surname>
          </string-name>
          and
          <string-name>
            <given-names>C. J. Van</given-names>
            <surname>Rijsbergen</surname>
          </string-name>
          .
          <article-title>Probabilistic Models of information retrieval based on measuring the divergence from randomness</article-title>
          .
          <source>ACM TOIS</source>
          ,
          <volume>20</volume>
          (
          <issue>4</issue>
          ):
          <fpage>357</fpage>
          -
          <lpage>389</lpage>
          ,
          <year>2002</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>Aitao</given-names>
            <surname>Chen and Fredric C.</surname>
          </string-name>
          <article-title>Gey. Combining Query Translation and Document Translation in Cross-Language Retrieval</article-title>
          . In Carol Peters, Julio Gonzalo,
          <string-name>
            <given-names>Martin</given-names>
            <surname>Braschler</surname>
          </string-name>
          , and et al., editors,
          <source>4th Workshop of the Cross-Language Evaluation Forum, CLEF 2003, Lecture notes in Computer Science, Lecture notes in Computer Science</source>
          , Trondheim, Norway,
          <year>2003</year>
          . SpringerVerlag.
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>Michael</given-names>
            <surname>Grubinger</surname>
          </string-name>
          , Paul Clough, Allan Hanbury, and
          <article-title>Henning Mu¨ller. Overview of the ImageCLEFphoto 2007 photographic retrieval task</article-title>
          .
          <source>In Working Notes of the 2007 CLEF Workshop</source>
          , Budapest, Hungary,
          <year>September 2007</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <surname>Savoy</surname>
            <given-names>J.</given-names>
          </string-name>
          <article-title>Fusion of Probabilistic Models for Effective Monolingual Retrieval</article-title>
          . In Carol Peters, Julio Gonzalo,
          <string-name>
            <given-names>Martin</given-names>
            <surname>Braschler</surname>
          </string-name>
          , and et al., editors,
          <source>4th Workshop of the Cross-Language Evaluation Forum, CLEF 2003, Lecture notes in Computer Science</source>
          , Trondheim, Norway,
          <year>2003</year>
          . Springer-Verlag.
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>Fernando</given-names>
            <surname>Llopis.</surname>
          </string-name>
          IR-n: Un Sistema de Recuperacio´n de Informacio´
          <article-title>n Basado en Pasajes</article-title>
          . In
          <source>PhD thesis</source>
          ,
          <year>2003</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>Fernando</given-names>
            <surname>Mart</surname>
          </string-name>
          <article-title>´ınez Santiago. El problema de la fusio´n de colecciones en la recuperacio´n de informacio´n multilingu¨e y distribuida: ca´lculo de la relevancia documental en dos pasos</article-title>
          . In
          <source>PhD thesis</source>
          ,
          <year>2004</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>