<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Building a Diversity Featured Search System by Fusing Existing Tools</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>General Terms</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Jiayu Tang, Thomas Arni, Mark Sanderson, Paul Clough Department of Information Studies, University of Sheffield</institution>
          ,
          <country country="UK">UK</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Measurement</institution>
          ,
          <addr-line>Performance, Experimentation</addr-line>
        </aff>
      </contrib-group>
      <abstract>
        <p>This paper describes our diversity featured retrieval system which are built for the task of ImageCLEFPhoto 2008. Two existing tools are used: Solr and Carrot2. We have experimented with different settings of the system to see how the performance changes. The results suggest that the system can indeed increase diversity of the retrieved results and keep the precision about the same.</p>
      </abstract>
      <kwd-group>
        <kwd>H</kwd>
        <kwd>3 [Information Storage and Retrieval]</kwd>
        <kwd>H</kwd>
        <kwd>3</kwd>
        <kwd>1 Content Analysis and Indexing</kwd>
        <kwd>H</kwd>
        <kwd>3</kwd>
        <kwd>3 Information Search and Retrieval</kwd>
        <kwd>H</kwd>
        <kwd>3</kwd>
        <kwd>4 Systems and Software</kwd>
        <kwd>H</kwd>
        <kwd>3</kwd>
        <kwd>7 Digital Libraries</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Introduction</title>
      <p>In this paper, we describe how to quickly set up a diversity featured search system by combining
existing tools that are available to the public, namely Solr [2] and Carrot2 [3]. We further describe
how to tune the system for better performance in the context of ImageCLEFPhoto 2008.
We use Solr for text search and Carrot2 for increasing diversity of results.</p>
      <p>Solr is a text search server based on the popular search library Lucene [1]. Solr provides some
very useful and convenient features. The web-services like API allows users to index documents
via XML over HTTP, query it via HTTP GET and receive results in XML format. The fields and
field types of documents can be easily defined in the schema.xml file. In this file, users can also
specify the Solr “out of the box” tokenizers and token filters to use for indexing and query. In
addition, the HTML administration interface gives users comprehensive insight into the system.</p>
      <p>Carrot2 is a open source search results clustering engine. It provides five different algorithms
for automatically organising search results into thematic categories. Carrot2 works as a pipeline
of three kinds of components: input components, filter components and visualisation components.
Input components obtain search results from a source of choice (e.g. YahooAPI, GoogleAPI,
Lucene, Solr, etc.), then filter components apply clustering algorithms to the search results, and
finally visualisation components render the clustered results to the user.
3</p>
    </sec>
    <sec id="sec-2">
      <title>System Setup</title>
      <p>In ImageCLEFPhoto 2008, each image comes with 9 fields of data: DOCNO, TITLE,
DESCRIPTION, NOTES, LOCATION, DATE, IMAGE, THUMBNAIL and TOPIC. We decided that
TITLE, DESCRIPTION, NOTES, LOCATION are the fields that would provide useful information
for text based image retrieval. Therefore, we constructed a new field named TEXT by combining
the text from the four fields, and specified it as the default search field in Solr. Solr’s default
configuration of tokenizers and token filters is used in our experiments, namely WhitespaceTokenizer,
SynonymFilter, StopFilter, WordDelimiterFilter, LowerCaseFilter, EnglishPorterFilter and
RemoveDuplicatesTokenFilter. More details on each tokenizer and token filter can be found on [2].
After feeding the Solr server with all the 20,000 documents in XML format, we have a running
search engine that is able to return a ranked list of documents based on the query submitted by
the user. Note that all the fields are indexed by Solr.</p>
      <p>
        In Carrot2, we construct an input component for acquiring results from the Solr server. Then,
a filter component is used for clustering the results, and a visualisation component is used for
displaying the clusters and their members. Since ImageCLEFPhoto 2008 assesses the S-recall [
        <xref ref-type="bibr" rid="ref2">5</xref>
        ]
of the top 20 results of the list to be submitted, we use the following procedure to find the best
20 results from the input/ranked list generated by Solr:
1: repeat
2: for each doc in the input list do
3: if membership of the current doc not exist in the list of cluster ID then
4: add the current doc to output list
5: remove the current doc from input list
6: add its membership to the list of cluster ID
7: if length of output list equals 20 then
8: break
9: end if
10: if length of cluster ID list equals number of clusters found by Carrot2 then
11: clear the list of cluster ID
12: break
13: end if
14: end if
15: end for
16: until length of output list equals 20
      </p>
      <p>Basically, the above pseudo code chooses documents based on two criteria: 1. appear as early
as possible in the input/ranked list, 2. cover as many different clusters as possible. The 20
documents chosen by the above procedure form the list to be submitted for assessment.
4</p>
    </sec>
    <sec id="sec-3">
      <title>Experiments</title>
      <p>We have submitted 32 runs, all of which are EN-EN-AUTO-TXT, meaning that all the submissions
are English-English monolingual runs using fully automated text clustering methods. Among
them, there are 8 groups, each of which includes 4 runs with the same system configuration except
that different number of documents are used for clustering. In other words, for each group, we
vary the number of documents (40, 60, 80 and 100) that are used from the top of the input list
for clustering by Carrot2, in order to get different output lists. For example, a run with top 40
documents applies clustering on the 40 documents and chooses 20 for submission based on the
procedure in Section 3.</p>
      <p>We increasingly change another 3 kinds of settings of the system to examine how the
performances would change. Firstly, for topics whose CLUSTER field is Country, City, State or
Location, we specify Carrot2 to cluster the documents by the LOCATION filed, otherwise by the
TEXT field. Secondly, we change the parameters of the clustering algorithm used in Carrot2.
Finally, we apply expansion to the indexing and query stage. In the following, we describe each
group of runs.
4.1</p>
      <sec id="sec-3-1">
        <title>Baseline</title>
        <p>
          This group uses the default settings of Solr and Carrot2. In terms of text retrieval, expansion
is not applied during indexing or query. Clustering is applied to the TEXT field for all topics.
Carrot2’s default clustering algorithm (Lingo [
          <xref ref-type="bibr" rid="ref1">4</xref>
          ]) and default parameters (0.150, 0.775) are used.
This group is used as a comparison baseline for other groups.
4.2
        </p>
      </sec>
      <sec id="sec-3-2">
        <title>Clustering by Location</title>
        <p>It seems that for topics that have been specified to be clustered by Country, City, State and
Location, the LOCATION filed in each document contains the essential information for clustering.
Therefore, such topics are clustered based on the LOCATION field. This group is different from
the Baseline group in that clustering on some topics are based on the LOCATION field, while
others based on the TEXT field.
4.3</p>
      </sec>
      <sec id="sec-3-3">
        <title>Parameters of Lingo set to (0.05, 0.95)</title>
        <p>
          Lingo [
          <xref ref-type="bibr" rid="ref1">4</xref>
          ] is a singular value decomposition based clustering algorithm that has been implemented
in Carrot2. The first parameter 0.05 is the Cluster Assignment Threshold, determining how precise
the assignment of documents to clusters should be. Lower threshold assign more documents to
clusters and less to “Other Topics”, which contains unclassified documents. With a low threshold,
more irrelevant documents are also assigned to the clusters. The second parameter 0.95 is the
Candidate Cluster Threshold, determining how many clusters Lingo will try to create. Higher
values give more clusters. This group is based on group 4.2, but uses (0.05, 0.95) as the parameters.
4.4
        </p>
      </sec>
      <sec id="sec-3-4">
        <title>Parameters of Lingo set to (0.10, 0.90)</title>
        <p>This group is the same as group 4.3 except that it uses (0.10, 0.90).
4.5</p>
      </sec>
      <sec id="sec-3-5">
        <title>Parameters of Lingo set to (0.15, 0.85)</title>
        <p>This group is the same as group 4.3 except that it uses (0.15, 0.85).
4.6</p>
      </sec>
      <sec id="sec-3-6">
        <title>Indexing and Query Expansion with (0.05, 0.95)</title>
        <p>This group is based on group 4.3, but applies indexing and query expansion in Solr. Specific domain
ontologies have been a popular choice for expansion. Cyclopedia websites such as Wikipedia have
also been adopted for expansion. Due to limited time, we used neither of the approaches. Instead,
we examined the topics and data-set, and then manually built an indexing expansion list and
query expansion list, as shown in Appendix. These lists are used as the synonym lists in Solr for
indexing and query. In the expansion lists, for lines containing “&gt;=”, any of the words before
“&gt;=” are replaced by words after “&gt;=” during expansion. For example, “ship, ships =&gt; ship,
vehicle” will replace any “ship” or “ships” by “ship, vehicle”. For lines without “&gt;=”, any word
from the line will be replaced by the whole set of words from the same line. For example, “USA”
or “United States of America” or “US” will be replaced by “USA, United States of America, US”.</p>
        <p>Precision@20
0.4
0.35
0.3
0.25
0.2
0.15
0.05
0</p>
      </sec>
      <sec id="sec-3-7">
        <title>Indexing and Query Expansion with (0.10, 0.90)</title>
        <p>This group is the same as group 4.6 except that it uses (0.10, 0.90) in Lingo.
4.8</p>
      </sec>
      <sec id="sec-3-8">
        <title>Indexing and Query Expansion with (0.15, 0.85)</title>
        <p>This group is the same as group 4.6 except that it uses (0.15, 0.85) in Lingo.
5</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>Results and Discussions</title>
      <p>As shown in Figure 1, group 6, 7 and 8 have clearly higher precisions than the other groups.
This is due to indexing and query expansion using the expansion lists in Appendix. As have been
mentioned, the expansion lists are built based on an examination of the data-set. For example,
in Topic 48 “vehicle in South Korea”, “South Korea” normally means the country, so there is
not much ambiguity. However, “vehicle” can mean many things, e.g. car, bus, boat. Intuitively,
expanding names of different types of vehicle with the word “vehicle” during indexing will boost
the precision, because many images are only annotated with specific vehicle names rather than
the word “vehicle”. Therefore, after indexing expansion, “car” becomes “car vehicle”. Similarly,
we expanded specific animal names with the word “animal”, so “fish” becomes “fish animal”.</p>
      <p>In terms of cluster recall, we can see in Figure 2 that different parameters of the clustering
algorithm (Lingo) have led to different performances. It is a little surprising that group 2 (Section
4.2) performed worse than group 1 (Section 4.1), the reason of which needs to be examined. In
addition, it seems that low Cluster Assignment Threshold (i.e. more documents are clustered)
and high Candidate Cluster Threshold (i.e. more clusters are created) give better cluster recall.
In our experiments, (0.05, 0.95) gives the best results: group 3 is better than group 4 and 5; group
6 is better than group 7 and 8.
0.45
0.4
0.35
0.3
0.25
0.2
0.15
0.1
0.05
0
40
60
80
100
1
2
3
4
5
6
7
8</p>
      <p>Group</p>
      <p>By comparing the two figures, it can also be noticed that groups with the same settings of Solr
have very similar precisions, no matter what settings of Carrot2 were used. For example, with
different parameters of Lingo, the precisions of group 6, 7 and 8 are relative stable, but the values
of cluster recall vary. This can be seen as an evidence that the diversity featured retrieval system
can make the results more diverse while maintaining the precision.
6</p>
    </sec>
    <sec id="sec-5">
      <title>Conclusions and Future Work</title>
      <p>This paper has described our submissions to ImageCLEFPhoto 2008. We have changed 4 kinds
of settings: the field used for clustering, the number of images used for clustering, indexing and
query expansion, and parameters of the clustering algorithm. The results suggest that indexing
and query expansion can fairly improve precision. Appropriately chosen clustering method can
increase diversity of the results while keeping precision almost the same.</p>
      <p>As we have mentioned, the performance of group 2 is a little against our initial anticipation. It
would be interesting to find out why. On the other hand, we plan to build an automatic expansion
approach using resources such as ontologies, rather than using the manually built expansion lists.
7</p>
    </sec>
    <sec id="sec-6">
      <title>Acknowledgements</title>
      <p>Work undertaken in this paper is supported by the EU-funded TrebleCLEF project (Grant
agreement: 215231) and by Tripod (Contract No. 045335).
[1] Apache lucene project. http://lucene.apache.org. [Visited 23/07/08].
[2] Apache solr project. http://lucene.apache.org/solr/index.html. [Visited 23/07/08].
[3] Carrot2 project. http://project.carrot2.org/. [Visited 23/07/08].</p>
      <p>A
A.1</p>
    </sec>
    <sec id="sec-7">
      <title>Appendix</title>
      <p>Indexing Expansion List
# vehicles
ship, ships =&gt; ship, vehicle
cutter, cutters =&gt; cutter, vehicle
train, trains =&gt; train, vehicle
rail, rails =&gt; rail, vehicle
locomotives, locomotive =&gt; locomotive, vehicle
wagon, wagons =&gt; wagon, vehicle
tractor, tractors =&gt; tractor, vehicle
bus, buses =&gt; bus, vehicle
car, cars =&gt; car, vehicle
forklift, forklifts =&gt; forklift, vehicle
boat, boats =&gt; boat, vehicle
Penguin, Penguins =&gt; penguin, animal
condors, condor =&gt; condor, animal
monkey, monkeys =&gt; monkey, animal
bird, birds =&gt; bird, animal
iguana, iguanas =&gt; iguana, animal
snail, snails =&gt; snail, animal
toucan, toucans =&gt; toucan, animal
lion, lions =&gt; lion, animal
llama, llamas =&gt; llama, animal
snake, snakes =&gt; snakes, animal
Tortoise, Tortoises =&gt; Tortoise, animal
parrot, parrots =&gt; parrot, animal
Booby, Boobies =&gt; Booby, animal
horse, horses =&gt; horse, animal
sea lion, sea lions =&gt; seal, animal</p>
      <p># sports
football =&gt; football, sport
surf =&gt; surf, sport
motorcycle =&gt; motorcycle, sport
race =&gt; race, sport</p>
      <p># people
fans, fan =&gt; fan, people</p>
      <p># water
river =&gt; river, water
lake =&gt; lake, water</p>
      <p># stone, rock
rock =&gt; rock, stone
brick, bricks =&gt; brick, stone
A.2</p>
      <sec id="sec-7-1">
        <title>Query Expansion List</title>
        <p># countries
USA, United States of America, US</p>
        <p># synonym
church, churches, cathedral, cathedrals
oxidised, rusty
observing, watch, see
football, soccer
match, game
accommodation, room</p>
        <p># common sense
swimming =&gt; swimming in the water
drawings, drawing =&gt; drawing, line, petroglyph
prize, prizes =&gt; prize, medal
water =&gt; water, sea</p>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>S.</given-names>
            <surname>Osinski</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Stefanowski</surname>
          </string-name>
          , and
          <string-name>
            <given-names>D.</given-names>
            <surname>Weiss</surname>
          </string-name>
          . Lingo:
          <article-title>Search results clustering algorithm based on singular value decomposition</article-title>
          .
          <source>In Proceedings of the International Conference on Intelligent Information Systems</source>
          , pages
          <fpage>359</fpage>
          -
          <lpage>368</lpage>
          , Zakopane, Poland,
          <year>2004</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>Cheng</given-names>
            <surname>Xiang Zhai</surname>
          </string-name>
          , William W. Cohen,
          <string-name>
            <given-names>and John</given-names>
            <surname>Lafferty</surname>
          </string-name>
          .
          <article-title>Beyond independent relevance: methods and evaluation metrics for subtopic retrieval</article-title>
          .
          <source>In Proceedings of the 26th annual international ACM SIGIR conference on Research and development in informaion retrieval</source>
          , pages
          <fpage>10</fpage>
          -
          <lpage>17</lpage>
          , Toronto, Canada,
          <year>2003</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>