<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Targeting Diversity in Photographic Retrieval Task with Commonsense Knowledge</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Supheakmungkol Sarin</string-name>
          <email>mungkol@fuji.waseda.jp</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Wataru Kameyama</string-name>
          <email>wataru@waseda.jp</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Commonsense knowledge</institution>
          ,
          <addr-line>Image Retrieval, Diversity, AnalogySpace, Query/Document Expansion, Rerank</addr-line>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Graduate School of Global Information and Telecommunication Studies, Waseda University 1011 Okuboyama, Nishi-Tomida</institution>
          ,
          <addr-line>Honjo-shi, Saitama-ken 367-0035</addr-line>
          ,
          <country country="JP">Japan</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>Image search engines have a very limited usefulness since it is still dicult to provide dierent users with what they are searching for. This is because most research eorts to date have only been concentrating on relevancy rather than diversity which is also a quite important factor, given that the search engine knows nothing about the user's context. In this paper, we describe our approach for ImageCLEF 2008 photographic retrieval task. The novelty of our technique is the use of AnalogySpace [3], the reasoning technique over commonsense knowledge for document and query expansion, which aims to increase the diversity of the results. Our proposed technique combines AnalogySpace mapping with other two mappings namely, location and full-text. We then re-rank the resulting images from the mapping by trying to eliminate duplicate and near duplicate results in the top 20. We present our preliminary experiments and the results conducted using the IAPR TC-12 photographic collection with 20,000 natural still photographs. The results show that our integrated method with AnalogySpace yields slightly better performance in terms of cluster recall and the number of relevant photographs retrieved . We nally identify the weakness in our approach and ways on how the system could be optimized and improved.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Introduction</title>
      <p>This paper describes our participation for the photographic retrieval task of ImageCLEF 2008.
ImageCLEF 2008 is a track running as part of the CLEF (Cross Language Evaluation Forum) campaign. It
comprises ve tasks related to image retrieval and annotation techniques, namely, photographic retrieval,
medical retrieval, general photographic concept detection, medical automatic image annotation, and
image retrieval task from a collection of Wikipedia images. We present our development and contributions
to the rst task of which the goal is to promote diversity in the top ranked list of resulting images.
2</p>
    </sec>
    <sec id="sec-2">
      <title>Approach and Implementation</title>
      <p>Using surrounding text of the images or annotation as a means to interpret them is a classic research
methodology. To date, however, most research eorts have only been concentrating on relevancy than
diversity. The latter is also a quite important factor since the search engine usually knows nothing about
the user. Furthermore, most of the time, people solve the problem through selecting some keywords and
features of images to represent the photograph rather than trying to understand the semantic nature of
annotation and the query. In this paper, we approach these problems as follows:</p>
      <p>To enable diversity, we use commonsense knowledge as a tool for term expansion. We consider
ConceptNet [8] as our commonsense knowledge database. ConceptNet is made up of a network
of everyday concepts that have been automatically generated from English sentences of the Open
Mind Common Sense corpus. The corpus has been handcrafted by the general public since 2000
[9]. Those concepts are connected by one of about twenty relationships such as IsA, PartOf ,
locationAt, Desires, CapableOf , UsedFor, etc. We use ConceptNet for diversity purposes
because a term can be expanded to its contextually related concepts that are not necessarily its
synonyms. Furthermore, those related concepts reect the commonsense way of people’s thinking
and how they relate concepts since they are input by human beings with a specic purpose. For
instance, drink coee relates to wake up, yawn, read newspaper, etc. However, diversity
should not come as a compensation of relevancy. Therefore, we also try to maintain the level of
precision by combining the former with both full-text and location matching.</p>
      <p>
        Re-ranking technique is performed afterward to re-rank the results of the previous step by trying
to eliminate duplicate and near duplicate results.
As shown in Figure 1, we introduce three kinds of matching between query and annotation of the image,
namely location, AnalogySpace, and full-text.
We begin by parsing the annotation to get location named entities. GATE is used for this purpose
[
        <xref ref-type="bibr" rid="ref6">10</xref>
        ]. Then, we establish a location hierarchy from the annotation before we perform the matching. For
instance, Lima is expanded to Lima &gt;&gt; Peru &gt;&gt; South America . Location names found in image
annotations and query topics are expressed as sets with prepositions found in the query as a matching
condition. To do this, we simply create two sets of prepositions namely, include set and exclude set.
Prepositions in include set are such as ’in’, ’of’, ’along’, ’on’, ’near’, ’by’, ’in’, etc., while the other set
includes prepositions such as ’out of’, ’outside’, etc. For example, in the query Sport stadium outside
Australia, outside serves as an excluding condition.
2.1.2
      </p>
      <sec id="sec-2-1">
        <title>AnalogySpace matching</title>
        <p>
          AnalogySpace is a vector space representation of commonsense knowledge built on the top of ConceptNet
using Principal Component Analysis [
          <xref ref-type="bibr" rid="ref2">3</xref>
          ]. This representation can be used as a reasoning tool as it reveals
large-scale patterns in the data while smoothing over noise. In our case, we use an implementation of
AnalogySpace called Divisi [
          <xref ref-type="bibr" rid="ref7">11</xref>
          ] to create ad-hoc category for each annotation and query. We then match
the query against the annotation. The degree of similarity between the two ad-hoc categories is the dot
product of matrices of the shared similar concepts and features.
        </p>
        <p>Since ConceptNet depends on sentences contributed from human, it does not contain all the terms a
dictionary has. To cope up with unknown terms, we use their synonym and hypernym. We create the
set of expanded term for the unknown term using its Wordnet’s synsets and hypernym regardless of its
part of speech. However, we only choose one term as our replacement for the unknown term. The best
term is the term that is most uniform to other terms of the annotation. This is achieved via dot product
of the matrix of an ad-hoc category created from a combination of other terms of the annotation, against
the ad-hoc categories created from each term from the expanded set if it exists in ConceptNet. We chose
the term that has the highest similarity score. Figure 2 shows the process.
2.1.3</p>
        <p>full-text matching
Vector Space Model is used to represent the annotations and query topics. Term frequency is used for
our vector space model. Each document is represented as a vector, where each dimension corresponds to
the frequency of a given term. In our case, terms are reduced to their stems.</p>
        <p>Some terms from query topics might not be found in the index of the annotation documents. To cope up
with this, we expand unknown query terms with their synsets and hypernym from WordNet. We select
top three terms among the set of synonyms found. AnalogySpace is used to compute the similarity score
between the unknown term and its synonyms.</p>
        <p>The similarity distance between a document vector and a query vector is expressed as cosine distance.
Figure 3 illustrates the technique.</p>
        <p>Finally, we normalize each matching score according to its maximum and minimum value. The total
matching score is expressed as the product of all the three matching scores. This is the simplest way to
combine the scores and yet make the large dierences count for even more.
2.2</p>
        <sec id="sec-2-1-1">
          <title>Re-ranking</title>
          <p>In this step, the results from the rst step are re-ranked according to their semantic similarity by giving
penalty to the ones with high similarity between each other.
We calculate full-text and location similarity. Same as in the matching process between query topic and
photograph annotation, boolean logic is used for location similarity calculation, while vector space model
is used for full-text similarity calculation. We compute the total pair distance of images as the product
of both distance scores.
2.2.2</p>
        </sec>
      </sec>
      <sec id="sec-2-2">
        <title>Re-rank</title>
        <p>The similarity distance score obtained can be used to lter and re-rank the preliminary results. We use
a method called Hill Climbing to nd a threshold of similarity distance that can help optimize both the
precision and diversity. We introduce a loop where Hill Climbing starts with a random threshold and
looks for the set of solutions which are better from its neighbors. The loop goes on until we obtain the
best compromise.
3
3.1</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>EVALUATION</title>
      <sec id="sec-3-1">
        <title>Process</title>
        <p>Organizers of ImageCLEF 2008 provide participants with a collection of annotated images, together with
query topics. Participants use these resources with their retrieval systems and submit to the organizers
the identiers of the relevant documents for each query topic. Then, the organizers evaluate the result
set of each submission from every participant and rank submissions according to standard evaluation
measures.
3.2</p>
      </sec>
      <sec id="sec-3-2">
        <title>Dataset 3.3</title>
      </sec>
      <sec id="sec-3-3">
        <title>Query</title>
        <p>
          The collection of images used for ImageCLEF 2008 is the IAPR TC-12 photo collection consisting of 20,000
natural images taken from locations around the world [
          <xref ref-type="bibr" rid="ref1">2</xref>
          ]. The collection includes images of various sports
and actions, photos of people, animals, cities, landscapes and many other aspects of contemporary life.
Each image is also associated with an alphanumeric caption stored in a semi-structured format. These
captions include the title of the image, its creation date, the location at which the photograph was taken,
a semantic description of the contents of the image by the photographer and some additional notes. Table
1 shows the example of a photograph and its metadata. In our system, we use only the title, description,
and location parts of the metadata.
        </p>
        <p>There are a total of 39 queries used in this study ranging from the very specic to the very abstract ones
with dierent levels of diculty. Here are some of the query topics: "animal swimming", "destinations
in Venezuela", "church with more than two towers", "sunset over water", etc. Query topics are provided
as a structured information. It is composed of the query title, cluster, narration of how relevant images
should be, and some examples of relevant image les. Table 2 shows the example of a query topic. In
our system, we use only the topic title.
3.4</p>
      </sec>
      <sec id="sec-3-4">
        <title>Measurement techniques</title>
        <p>
          To ensure both relevancy and diversity, the evaluation is based principally on two measures, namely,
precision at 20, and instance recall at rank 20 [
          <xref ref-type="bibr" rid="ref3">4</xref>
          ]. The technique is a relatively new evaluation
methodology that considers results of a query as interdependence rather than a standalone. A good engine will
produce results that maximize the two measurements.
4
        </p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>Results and Discussions</title>
      <p>We present below the results of the four runs that we submitted to ImageCLEF photographic task 2008.
1. AnalogySpace: In this run, we combine location matching and AnalogySpace.
2. Full-text: In this run, we simply use location matching and full-text search.
3. Full-text (no query expansion) + AnalogySpace: In this run, we combine location matching,
fulltext matching, and AnalogySpace matching.
4. Full-text (with query expansion) + AnalogySpace: The same as the previous one, we combine
the three matching. We further expand the terms of query topics in full-text matching with their
synsets and hypernym.</p>
      <sec id="sec-4-1">
        <title>Runs</title>
      </sec>
      <sec id="sec-4-2">
        <title>AnalogySpace Full-text Full-text (no query expansion) + AnalogySpace Full-text (with query expansion) + AnalogySpace</title>
        <p>P5</p>
        <p>We still believe that ConceptNet could help enriching diversity in the resulting images. To our
understanding, the reason why we could not achieve better results is because of the fact that there are lots of
terms that ConceptNet does not cover. When we try to expand those unknown terms using WordNet,
we only introduce noise. That is because WordNet’s synsets contain all the synonyms of the word from
all its possible senses. We did not implement any sense disambiguation. We did not even check the part
of speech. Therefore, most of the time, the replacement only twists the meaning of the original word
since we do not select the most appropriate sense of the word. Moreover, we limit the number of selected
synonym to only one in AnalogySpace term expansion, and only up to three in our full-text query
expansion. This reduces the coverage of the meanings. Due to limited time, content-based technology was not
taken into consideration. Should we have incorporated another content-based pair similarity distance in
the re-ranking step, we might be able to get better resulting images. Hence, we are planning to tackle
these issues in our future works.</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>Related Works</title>
      <p>
        Image search at the major search engines today largely relies on looking at words that are used around
images on the pages that host them, in image le names, and in ALT text associated with them. No
real image recognition is done by any of the majors. Datta et al. have recently produced a complete
survey of the current image related techniques [
        <xref ref-type="bibr" rid="ref5">6</xref>
        ]. Hsu et al [1] have used ConceptNet as tool for
query and document expansion in image retrieval task. Nevertheless, in doing this, the authors only use
spatial relationship function to nd the concepts that co-exist in space of the real world. Google recently
introduced VisualRank a method that guesses how the images would be linked together, with those
being most similar having more virtual links to each other. As a result, the most "linked to" images are
calculated to rank rst [7].
6
      </p>
    </sec>
    <sec id="sec-6">
      <title>Conclusion</title>
      <p>User satisfaction is not solely a function of relevancy. When nothing is known about the user, diversity
plays an important role in getting the results that user would like to see. We present a novel approach
to enable rich diversity in the results by incorporating commonsense knowledge expansion and result
re-ranking through elimination of duplicate and near duplicate results. The presented results are just our
preliminary ones. Even they are not conclusive yet, they pave the way to help us to improve our current
system. We are now working to address the weak points that we have discussed earlier.
[1] Ming-Hung Hsu and Hsin-Hsi Chen, 2006. Information retrieval with commonsense knowledge.
Proceedings of the 29th ACM SIGIR ’06, 651652.
[7] Yushi Jing and Shumeet Bajula. PageRank for Product Image Search. In Proceedings of the World</p>
      <p>Wide Web Conference. ACM WWW ’08.
[8] Havasi, C., Speer, R. &amp; Alonso, J. (2007) ConceptNet 3: a Flexible, Multilingual Semantic Network
for Common Sense Knowledge. Proceedings of Recent Advances in Natural Languages Processing
2007
[9] Singh P, Lin T, Mueller E T, Lim G, Perkins T and Zhu W L: ‘Open mind commonsense: knowledge
acquisition from the general public’, Proceedings of the First International Conference on Ontologies,
Databases, and Applications of Semantics for Large Scale Information Systems, Lecture Notes in
Computer Science No 2519 Heidelberg, Springer (2002).
over
semantic</p>
      <sec id="sec-6-1">
        <title>Website:</title>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [2]
          <string-name>
            <surname>Grubinger</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Clough</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Mller</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          <article-title>and</article-title>
          <string-name>
            <surname>Deselaers</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <year>2006</year>
          .
          <article-title>The IAPR TC-12 Benchmark: A New Evaluation Resource for Visual Information Systems</article-title>
          , In Proceedings of International Workshop OntoImage'
          <year>2006</year>
          <article-title>Language Resources for Content-Based Image Retrieval, held in conjunction with LREC'06</article-title>
          , pages
          <fpage>13</fpage>
          -
          <lpage>23</lpage>
          , Genoa, Italy, 22 May
          <year>2006</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>Robert</given-names>
            <surname>Speer</surname>
          </string-name>
          , Catherine Havasi, and Henry Lieberman.
          <article-title>AnalogySpace: Reducing the Dimensionality of Common Sense Knowledge</article-title>
          , Chicago, Illinois, AAAI 2008
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [4]
          <string-name>
            <surname>Zhai</surname>
            ,
            <given-names>C. X.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Cohen</surname>
            ,
            <given-names>W. W.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Laerty</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          <year>2003</year>
          .
          <article-title>Beyond independent relevance: methods and evaluation metrics for subtopic retrieval</article-title>
          .
          <source>In Proceedings of the 26th Annual international ACM SIGIR Conference on Research and Development in Information Retrieval</source>
          (Toronto, Canada,
          <source>July 28 - August 01</source>
          ,
          <year>2003</year>
          ). SIGIR '
          <volume>03</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>Catherine</given-names>
            <surname>Havasi</surname>
          </string-name>
          , Robert Speer, and Jason Alonso.
          <article-title>ConceptNet 3: a Flexible, Multilingual Semantic Network for Common Sense Knowledge, RANLP 2007</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>R.</given-names>
            <surname>Datta</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Joshi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>and J. Z.</given-names>
            <surname>Wang</surname>
          </string-name>
          .
          <article-title>Image Retrieval: Ideas, Inuence, and Trends of the New Age</article-title>
          .
          <source>ACM Computing Surveys</source>
          , vol.
          <volume>40</volume>
          , no.
          <issue>2, article</issue>
          5, 60 pages,
          <year>2008</year>
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>H.</given-names>
            <surname>Cunningham</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Maynard</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Bontcheva</surname>
          </string-name>
          , and
          <string-name>
            <given-names>V.</given-names>
            <surname>Tablan</surname>
          </string-name>
          ,
          <article-title>GATE: a framework and graphical development environment for robust NLP tools and applications</article-title>
          ,
          <source>in Proceedings of the 40th Anniversary Meeting of the Association for Computational Linguistics (ACL '02)</source>
          , Philadelphia, Pa, USA,
          <year>July 2002</year>
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [11]
          <article-title>Divisi: a general-purpose tool for reasoning http://divisi</article-title>
          .media.mit.edu/ (Last visit:
          <source>August</source>
          <volume>11</volume>
          ,
          <year>2008</year>
          )
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>