<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>MMIS at ImageCLEF 2008: Experiments combining different evidence sources</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Simon Overell</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
          <xref ref-type="aff" rid="aff4">4</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Ainhoa Llorente</string-name>
          <email>a.llorente@open.ac.uk</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
          <xref ref-type="aff" rid="aff4">4</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Haiming Liu</string-name>
          <email>h.liu@open.ac.uk</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
          <xref ref-type="aff" rid="aff4">4</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Rui Hu</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
          <xref ref-type="aff" rid="aff4">4</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Adam Rae</string-name>
          <email>a.rae@open.ac.uk</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
          <xref ref-type="aff" rid="aff4">4</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Jianhan Zhu</string-name>
          <email>jianhanzhu@gmail.com</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
          <xref ref-type="aff" rid="aff4">4</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Dawei Song</string-name>
          <email>d.song@open.ac.uk</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
          <xref ref-type="aff" rid="aff4">4</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Stefan Ru¨ger</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
          <xref ref-type="aff" rid="aff4">4</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>General Terms</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
          <xref ref-type="aff" rid="aff4">4</xref>
        </contrib>
        <contrib contrib-type="editor">
          <string-name>Measurement, Performance, Experimentation</string-name>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Department of Computing, Imperial College London</institution>
          ,
          <addr-line>SW7 2AZ</addr-line>
          ,
          <country country="UK">UK</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>E-48170 Zamudio</institution>
          ,
          <addr-line>Bizkaia</addr-line>
          ,
          <country country="ES">Spain</country>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>Knowledge Media Institute</institution>
        </aff>
        <aff id="aff3">
          <label>3</label>
          <institution>ROBOTIKER-TECNALIA</institution>
          ,
          <addr-line>Parque Tecnol ́ogico, Edificio 202</addr-line>
        </aff>
        <aff id="aff4">
          <label>4</label>
          <institution>The Open University</institution>
          ,
          <addr-line>Milton Keynes, MK7 6AA</addr-line>
          ,
          <country country="UK">UK</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>This paper presents the work of the MMIS group at ImageCLEF 2008. The results for three tasks are presented: Visual Concept Detection Task (VCDT), ImageCLEFphoto and ImageCLEFwiki. We combine image annotations, CBIR, textual relevance and a geographic filter using our generic data fusion method. We also compare methods for BRF and clustering. Our top performing method in the VCDT enhances supervised learning by modifying probabilities based on a matrix that shows how terms appear together. Although it occurred in the top quartile of submitted runs, the enhancement did not provide a statistically significant improvement. In the ImageCLEFphoto task we demonstrate that evidence from image retrieval can provide a contribution to retrieval; however we are yet to find a way of combining text and image evidence in a way to provide an improvement over the baseline. Due to the relative performances of difference evidences in ImageCLEFwiki and our failure to improve over a baseline we conclude that text is the dominant feature in this collection.</p>
      </abstract>
      <kwd-group>
        <kwd>H</kwd>
        <kwd>3 [Information Storage and Retrieval]</kwd>
        <kwd>H</kwd>
        <kwd>3</kwd>
        <kwd>1 Content Analysis and Indexing</kwd>
        <kwd>H</kwd>
        <kwd>3</kwd>
        <kwd>3 Information Search and Retrieval</kwd>
        <kwd>H</kwd>
        <kwd>3</kwd>
        <kwd>4 Systems and Software</kwd>
        <kwd>H</kwd>
        <kwd>3</kwd>
        <kwd>7 Digital Libraries</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Introduction</title>
      <p>In this paper we describe the experiments of the MMIS group at ImageCLEF’08. We participated
in three tasks: Visual Concept Detection Task (VCDT), ImageCLEFphoto and ImageCLEFwiki.</p>
      <p>All experiments were performed in a single framework of independently testable and tuneable
modules. The framework is described in detail in Section 2. The experiments conducted and
individual runs submitted for each task are described in Sections 3, 4 and 5. Finally we present
our conclusions in Section 6.
2</p>
    </sec>
    <sec id="sec-2">
      <title>System</title>
      <p>Our system framework is shown in Figure 1. The text elements of the corpus are indexed as
a bag-of-words, analysed geographically and stored in a geographic index. Texture and colour
features are extracted from images to form feature indexes, these features are further analysed to
automatically annotate the images. This allows us to compare query images to our index using
both semantic annotations and low-level features.</p>
      <p>
        Blind Relevance Feedback (BRF) is employed across media types similarly to [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ]. In training
experiments we found the text results to have the highest precision of all individual media types.
Because of this we use the top text results as feedback for the Image Retrieval engine to provide
an additional Image BRF rank.
      </p>
      <p>Our intermediate format is the standard TREC format. This allows us to evaluate and tune
each module independently. The results of all the independent modules are combined in the data
fusion module. The data fusion module combines both the ranks provided by the image and
text query engine, and the filters provided by the geographic and annotation query engines. The
difference between a rank and a filter is all elements in a filter are considered of equal relevance.</p>
      <p>In the ImageCLEFphoto task we are evaluated on the novelty of our top results and provided
with a subject which this novelty will be judged with respect to. We have clustered our top results
using the geographic index, image annotations and text index.</p>
      <p>Further details on the individual modules and tuning are described in the following sections.
2.1</p>
      <sec id="sec-2-1">
        <title>Image Feature Extractor</title>
        <p>Content-Based Image Retrieval (CBIR) provides a way to browse or search images from large image
collections based on visual similarity. CBIR is normally performed by computing the dissimilarity
between the object images and query images based on their multidimensional representations in
content feature spaces, for example, colour, texture and structure. In this section, we are going to
introduce the key issues of the image search applied to our three tasks.
2.1.1</p>
        <sec id="sec-2-1-1">
          <title>Feature Extraction.</title>
          <p>
            ImageCLEFphoto Task. Colour feature is the most commonly used visual feature e.g. HSV,
RGB [
            <xref ref-type="bibr" rid="ref14">14</xref>
            ]. In [
            <xref ref-type="bibr" rid="ref8">8</xref>
            ] Colour feature HSV outperforms the texture feature Gabor and structure feature
Konvolution on CBIR . Thus we use HSV in our CBIR system. HSV is a cylindrical colour space.
Its representation appears to be more intuitive to humans than the hardware representation of
RGB. The hue coordinate H is angular and represents the colour, the saturation S represents the
pureness of the colour and is the radial distance, finally the brightness V is the vertical distance.
We extract the HSV feature globally for every image in the query and test sets.
ImageCLEFwiki Task. We used the Gabor feature for the ImageCLEFwiki task. Gabor is a
texture feature generated using Gabor wavelets [
            <xref ref-type="bibr" rid="ref7">7</xref>
            ]. Here we decompose each image into two scales
and four directions.
          </p>
          <p>Figure 1: Our Application Framework
VCDT. The features used in the VCDT experiments are a combination of colour feature feature,
CIELAB, and texture feature, Tamura.</p>
          <p>
            CIE L ∗ a ∗ b∗ (CIELAB) [
            <xref ref-type="bibr" rid="ref4">4</xref>
            ] is the most complete colour space specified by the International
Commission on Illumination (CIE). Its three coordinates represent the lightness of the colour (L∗),
its position between red/magenta and green (a∗) and its position between yellow and blue (b∗).
          </p>
          <p>
            The Tamura texture feature is computed using three main texture features called “contrast”,
“coarseness”, and “directionality”. Contrast aims to capture the dynamic range of grey levels in
an image. Coarseness has a direct relationship to scale and repetition rates and it was considered
by Tamura et al. [
            <xref ref-type="bibr" rid="ref16">16</xref>
            ] as the most fundamental texture feature and finally, directionality is a global
property over a region.
          </p>
          <p>The process for extracting each feature is as follows, each image is divided into nine equal
rectangular tiles, the mean and second central moment feature per channel are calculated in each
tile. The resulting feature vector is obtained after concatenating all the vectors extracted in each
tile.
2.1.2</p>
        </sec>
        <sec id="sec-2-1-2">
          <title>Dissimilarity Measure.</title>
          <p>ImageCLEFphoto Task. The χ2 statistic is a statistical measure that compares two objects
in a distributed manner and basically assumes that the feature vector elements are samples. The
dissimilarity measure is given by
dχ2 (A, B) =</p>
          <p>Xn (ai − mi)2</p>
          <p>
            ,
mi
mi =
i=1
ai + bi
2
where A = (a1, a2, ..., an) and B = (b1, b2, ..., bn) are the query vector and test object vector
respectively. It measures the difference of the query vector (observed distribution) from the mean
of both vectors (expected distribution)[
            <xref ref-type="bibr" rid="ref20">20</xref>
            ]. The χ2 statistic was chosen as it was one of the
consistently best performing dissimilarity measures in our former research [
            <xref ref-type="bibr" rid="ref8">8</xref>
            ].
ImageCLEFwiki Task. We use the City Block distance for the ImageCLEFwiki task, to
compute the distance between a query image and each test image. The City Block distance belongs
to the Minkowski family, which is given by:
when the parameter p equals one.
2.1.3
          </p>
        </sec>
        <sec id="sec-2-1-3">
          <title>Search Method.</title>
          <p>n
dp(A, B) = (X |ai − bi|p) p1 ,</p>
          <p>i=1
ImageCLEFphoto Task. In the ImageCLEFphoto task, every query topic includes three
independent example images a, b, c. The probability of an object image x being a relevant match
for a query topic is determined by the joint probability of the relevance between the images a, b,
c in the query topic and the object image x. According to probability theory, the joint result is
given by</p>
          <p>D(abc, x) = d(a, x) × d(b, x) × d(c, x),
where D(abc, x) is the distance between a query topic including images a, b and c with an image
x. d(a, x), d(b, x) and d(c, x) are distances between the three images a, b, c and the image x,
respectively.
(1)
(2)
(3)
(4)
2.1.4</p>
        </sec>
        <sec id="sec-2-1-4">
          <title>Blind Relevance Feedback.</title>
          <p>
            We employed Blind Relevance Feedback (BRF) for CBIR in a similar fashion to [
            <xref ref-type="bibr" rid="ref9">9</xref>
            ]. Two different
BRF method were applied in the ImageCLEFphoto task.
          </p>
          <p>• The first BRF method takes the top seven results from the text retrieval results as new
query examples, and the seven ranked result are combined by the search method introduced
in Section 2.1.3.
• The second BRF method was employed in both the ImageCLEFphoto and ImageCLEFwiki
tasks. The final ranked result of this BRF is the sum of the ranks of the top five examples
from the text result with equal weights. The detail of this combining method are described
in Section 2.6.
2.2</p>
        </sec>
      </sec>
      <sec id="sec-2-2">
        <title>Image Annotator</title>
        <p>The Image Annotator is the core part of the Visual Concept Detection Task (VCDT) which aims
to create a model able to detect the presence or absence of 17 visual concepts in the images of
the collection. The input is a training set of 1825 images that have already been annotated with
words coming from a vocabulary of 17 visual concepts. The output is the annotations.</p>
        <p>
          We use as a baseline for this module the framework developed by Yavlinsky et al. [
          <xref ref-type="bibr" rid="ref19">19</xref>
          ] who
used global features together with a non-parametric density estimation.
        </p>
        <p>The process can be described as follows. First, images are segmented into nine equal tiles, and
then, low-level features are extracted. The features used to model the visual concept densities are
a combination of colour CIELAB and texture Tamura, as explained in Section 2.1.</p>
        <p>The next step is to extract the same feature information from an unseen picture in order
to compare it with all the previously created models (one for each concept). The result of this
comparison yields a probability value of each concept being present in each image.</p>
        <p>Then we modify some of these probabilities using additional knowledge from the image context
in order to improve the accuracy of the final annotations.</p>
        <p>The context of the images is computed using a co-occurrence matrix where each cell represents
the number of times two visual concepts appear together annotating an image of the training set.</p>
        <p>The underlying idea of this algorithm is to detect incoherence between words with the help
of this correlation matrix. Once incoherence between words has been detected, the probability of
the word associated to the lowest probability will be lowered, as well as all the words which are
semantically similar.</p>
        <p>
          Among the many uses of the concept “semantic similarity,” we refer to the definition by Miller
and Charles [
          <xref ref-type="bibr" rid="ref10">10</xref>
          ] who consider it as the degree of contextual interchangeability or the degree
to which one word can be replaced by another in a certain context. Consequently, two words
are similar if they refer to entities that are likely to co-occur together like “mountains” and
“vegetation”, “beach” and “water”, “buildings” and “road”, etc. As shown in Figure 2, our
vocabulary of 17 visual concepts adopt a hierarchical structure. In the first level we find two
general concepts like “indoor” and “outdoor” which are mutually exclusive while in lower levels
of the hierarchy we find more specific concepts that are subclasses of the previous ones. Some
concepts can belong to more than one class, for instance, a “person” can be part of an “indoor” or
“outdoor” scene but others are mutually exclusive, a scene can not represent “day” and “night”
at the same time.
        </p>
        <p>Thus, by modifying the probability values of some concepts, annotations are produced by
selecting the concepts with the highest probability. The output of this module, the annotations,
will be the input for the Annotation Query Engine and the Document Cluster Generator.
2.2.1</p>
        <sec id="sec-2-2-1">
          <title>Annotation Query Engine</title>
          <p>The input for this module is textual queries which are the concepts of our vocabulary and the
annotations produced by the Image Annotator module. Given a query term, we represent the top
n images annotated by it following the standard TREC format. This constitutes one of the inputs
for the Data Fusion module. Retrieval performance is evaluated with the mean-average precision
(MAP) on the whole vocabulary of terms, which is the average precision, over all queries, at the
ranks where recall changes (where relevant items occur).
2.3</p>
        </sec>
      </sec>
      <sec id="sec-2-3">
        <title>Text Indexer</title>
        <p>
          Our text retrieval system is based on Apache Lucene [
          <xref ref-type="bibr" rid="ref15">15</xref>
          ]. Text fields are pre-processed by a
customised analyser similar to Lucene’s default analyser: text is split at white space into tokens,
the tokens are then converted to lower case, stop words discarded and stemmed with the “Snowball
Stemmer”. The processed tokens are held in Lucene’s inverted index.
        </p>
        <p>We use only the English meta-data and queries (monolingual retrieval). In ImageCLEFphoto
both the title and location fields are searched as text. We do not use the notes field as in previous
training experiments we have found this gives worse results.
2.4</p>
      </sec>
      <sec id="sec-2-4">
        <title>Geographic Indexer</title>
        <p>
          We process the text fields with Sheffield University’s General Architecture for Text
Engineering (GATE) [
          <xref ref-type="bibr" rid="ref1">1</xref>
          ]. The bundled Information Extraction Engine, ANNIE, performs named entity
recognition, extracting named entities and tagging them as locations. Our disambiguation system
matches these placenames to unique locations in the Getty Thesaurus of Geographical Names
(TGN) [
          <xref ref-type="bibr" rid="ref5">5</xref>
          ].
        </p>
        <p>
          In ImageCLEFwiki we compare two different disambiguation methods, both based on a
geographic co-occurrence model mined from Wikipedia [
          <xref ref-type="bibr" rid="ref13">13</xref>
          ]. The first method (MR), builds a default
gazetteer based on statistics from Wikipedia on which locations are most referred to by each
placename. It is not context aware and disambiguates every placename with the same name to
the same location. For example every reference to Cambridge will be matched to Cambridge,
Massachusettes regardless of context.
        </p>
        <p>The second method (Neighbourhoods), builds neighbourhoods of trigger words from Wikipedia.
Depending which trigger words occur in the vicinity of an ambiguous placename dictates which
location it will be disambiguated as. For example if Oxford occurs in the context of Cambridge,
Cambridge will be disambiguated as Cambridge, Cambridgeshire as Oxford is a trigger word of
Cambridge, Cambridgeshire.</p>
        <p>As the ImageCLEFphoto corpus contains minimal referrences to ambiguous placenames we
only use the MR method.</p>
        <p>To query our geographic index we extract locations from the query and based on the topological
data contained in the TGN return a filter of all the documents mentioning either the query locations
or locations within the query locations. For example a query location of the “United States” will
return all the documents mentioning the United States (or synonyms such as “USA”, “the States”
etc.) and all documents mentioning states, counties, cities and towns within the United States.
2.5</p>
      </sec>
      <sec id="sec-2-5">
        <title>Document Cluster</title>
        <p>
          We only employ clustering in ImageCLEFphoto. We propose a simple method of re-ordering
the top of our rank based on document annotations. We consider three sources of annotations:
Automated annotations assigned to images, words matched to WordNet and locations extracted
from text (described in Section 2.4). WordNet is a freely available semantic lexicon of 155,287
words mapping to 117,659 semantic senses [
          <xref ref-type="bibr" rid="ref17">17</xref>
          ]. In our experiments we compare two sources of
annotations: automated image annotations (Image clustering) and words matched to WordNet
(WordNet clustering).
        </p>
        <p>
          In Image clustering all the images have been annotated with concepts coming from the Corel
ontology. This ontology was created using SUMO (Suggested Upper Merged Ontology) [
          <xref ref-type="bibr" rid="ref12">12</xref>
          ] and
enriching it with a taxonomy of animals created by the University of Michigan [
          <xref ref-type="bibr" rid="ref11">11</xref>
          ]. After that, the
ontology was populated with the vocabulary of 374 terms used for annotating the Corel dataset.
Among many categories, we can find animal, vehicle and weather.
        </p>
        <p>For example, if the cluster topic is “animal” we will split the ranked list of results into sub
ranked lists, one corresponding to every type of animal and an additional uncategorised rank.
These ranks will then be merged to form a new rank, where all the documents at rank 1 in a sub
list appear first, followed by the documents at rank 2, followed by rank 3 etc. The documents
of equal rank in the sublists are ranked amongst themselves based on their rank in the original
list. This way the document at rank 1 in the original list remains at rank 1 in the re-ranked
list. We only maximise the diversity of the top 20 documents, after the 20th document the other
documents maintain their ordering in the original list.</p>
        <p>Similarly in WordNet clustering we build clusters of sub-categories of animal, bird, sport,
vehicle and weather. We match these to the bag-of-words for each document contained in the text
index. For example if the cluster topic is sport we will have a sub-ranked list containing every
document that mentions “tennis”. If for image clustering the cluster topic is not animal , vehicle
or weather, or for WordNet clustering the topic is not animal, bird, sport, vehicle or weather, we
default to location clustering. In this case we have a sublist for every different location a document
is annotated with.
2.6</p>
      </sec>
      <sec id="sec-2-6">
        <title>Data Fusion</title>
        <p>With multiple sources of evidence being provided by the different query engines, a method of
combining these data was required. Each query engine produced an output rank of data set
images ordered with respect to their relevance to the input query. Not all engines gave values for
the entire data set—some gave ranks of relevant sub-sets of the main data set.</p>
        <p>
          We adopted a group consensus function based on the Borda Count method [
          <xref ref-type="bibr" rid="ref2 ref6">6, 2</xref>
          ] as it was
simple to implement and fair to the input data, but it was extended by introducing per-rank
scaling parameters.
        </p>
        <p>Data were split into two categories which were treated differently; ranks and filters. Ranks
were processed without adjustment and directly compared to each other during the combination
process. Filters were used to filter out non-relevant results from an intermediate combined rank
stage. This was used, for example, with the output of the geographic query engine, where it was
judged more appropriate to ‘filter in’ explicitly determined relevant results than to treat it like a
rank.</p>
        <p>The final output of the data fusion algorithm was the combined result of the multiple input
ranks and filters.
where n is the number of pure ranks to combine.</p>
        <p>This function then defined our parameter space within which we had to search to find the more
appropriate weights for the rank combination process. We performed a brute force search of the
entire parameter space to find the optimum weights. We used past years’ data for training and
optimum sets of weights were considered those that gave the maximum mean average precision
(MAP). The final weights are given in Table 1 for the two tasks that used rank combination; the
ImageCLEFwiki and the ImageCLEFphoto tasks.</p>
        <p>The parameter values used by the filter stage were also important. Unlike the pure rank weights
that were used to multiply an entire rank, the filters were used by penalising those values that
were not present in the filter. These penalisation weights pi, making up weight vector P , were
subject to the same constraints and were derived in a similar way to the pure rank weights by
exhaustively searching all possible values up to a limit which was defined as when the variation in
the parameter yielded no change in the final rank’s MAP value greater than 0.0001.</p>
        <p>The overall process is described here as a five stage algorithm:
1. Read in rank data</p>
        <p>Data was provided by the query engines in ordered ranked lists with those elements that
were judged to be most relevant by the engine appearing first, with each rank denoted Ri.
Each of the m elements was assigned a value based on its position in the list, so the first
element was given ‘1’, and increased consecutively down the list until the last element had
a value of m.
2. Multiply by rank weights</p>
        <p>Each query engine output list was scaled by multiplying every rank value in the list by its
parameter value wi.
3. Sum ranks</p>
        <p>The scaled ranks were then combined by producing a new rank R0 where every element r in
each Ri was replaced by value r0. Stages 2 and 3 can be summarised thus:
(5)
(6)
r0 =
m
X wiri
i=1
where m is the number of ranks to combine and wi is an element of the weight vector W .
4. Filter rank</p>
        <p>The newly combined intermediary rank R0 was then subjected to the output of each filter
engine to produce rank R00. For each element in the rank R0, any element that was not found
in the filter data had its rank value penalised by that rank’s parameter value as described
by filter function f (r). This pushed less relevant elements further down the ordered rank.
f (r0i) =
(
r0i r0i present in filter
r0ipi r0i not present in filter
(7)
5. Sort rank</p>
        <p>The filtered rank R00 was then sorted to ensure that the rank values were in ascending order.</p>
        <p>The sorted rank is then the output of the combination stage of the overall system.
3
3.1</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>VCDT</title>
      <sec id="sec-3-1">
        <title>Experiments</title>
        <p>The objective of the Visual Concept Detection Task (VCDT) is to detect the presence or absence
of 17 visual concepts in the 1000 images that constitute the test set. In addition to that, some
confidence scores are provided once an object is detected. The higher the value the greater the
confidence of the presence of the object in the image.</p>
        <p>For the VCDT task, we submitted four different algorithms, all of them correspond to
automatic runs dealing with visual information. The second run uses statistical information about
visual concepts co-occurring together in addition to the visual information, and the final run is a
combination of the other three.
3.1.1</p>
        <sec id="sec-3-1-1">
          <title>Automated Image Annotation Algorithm</title>
          <p>
            This algorithm corresponds to the work carried out by Yavlinsky et al. [
            <xref ref-type="bibr" rid="ref19">19</xref>
            ]. Their method is
based on a supervised learning model that uses a Bayesian approach together with image analysis
techniques. The algorithm exploits simple global features together with robust non-parametric
density estimation using the technique of kernel smoothing in order to estimate the probability of
the words belonging to a vocabulary being present in each one of the pictures of the test set. This
algorithm was previously tested with the Corel dataset and the Getty collection.
3.1.2
          </p>
        </sec>
        <sec id="sec-3-1-2">
          <title>Enhanced Automated Image Annotation Algorithm</title>
          <p>This second algorithm is described in detail in Section 2.2. The submitted run is based on an
enhanced version of the algorithm described in the previous section. The input is the annotations
achieved by the algorithm developed by Yavlinsky et al. together with a matrix that represents
the probabilities of all the words of the vocabulary being present in the images. This algorithm
was also tested on the Corel collection of 5,000 pictures and a vocabulary of 374 words obtaining
statistical significant results (5%).
3.1.3</p>
        </sec>
        <sec id="sec-3-1-3">
          <title>Dissimilarity Measures Algorithm</title>
          <p>The third algorithm follows a simple approach based on dissimilarity measures. Given one test
image, we compute its global distance (Cityblock distance) to the mean of the training images
which share one common keyword. The smaller the distance value the higher the probability for
a test image to be annotated by the keyword of that category. Our submitted result is based
on the combination of two results from two single feature spaces. One is the CIELAB colour
feature. For each pixel, we compute the CIELAB colour values. Within each tile the mean and
the second moment are computed for each channel. The other is the Tamura texture feature. The
combination here is a simple additive combination of the probability for each test image for each
category.</p>
          <p>Algorithm
Enhanced automated image annotation
Automated image annotation
Combined algorithm
Dissimilarity measures algorithm
This submitted run is based on the combination of all the other runs submitted by this team, in
addition to one extra combined result formed from the output of a feature extraction algorithm for
the Tamura and CIELAB features. This was a simple additive combination which, through testing
was shown to be useful when used in combination with the other algorithm outputs. By testing
the output of each individual algorithm on subsets of the training data they were produced with,
a rough indication of performance of the algorithm per concept was derived. These figures then
allowed an automated system to pick for each concept the best performing algorithm’s output and
combine it into a new result set. The individual algorithms’ outputs were not adjusted or scaled
in any way (other than the pre-combined result set mentioned above).</p>
          <p>
            In our testing routines based on splitting the available training data into training and
evaluation sets, the combination performed marginally better than the individual component algorithm
outputs. This was due to selecting those algorithms which performed better at certain concepts
and classes of concepts. Further work will be carried out to more robustly take advantage of
concept classification when combining algorithm results.
The evaluation metric followed by the ImageCLEF organisation is based on ROC curves. Initially,
a receiver operating characteristic (ROC) curve was used in signal detection theory to plot the
sensitivity versus (1 - specificity) for a binary classifier as its discrimination threshold is varied.
Later on, ROC curves [
            <xref ref-type="bibr" rid="ref3">3</xref>
            ] were applied to information retrieval in order to represent the fraction
of true positives (TP) against the fraction of false positives (FP) in a binary classifier. The Equal
Error Rate (EER) is the error rate at the threshold where FP=FN. The area under the ROC
curve, AUC, is equal to the probability that a classifier will rank a randomly chosen positive
instance higher than a randomly chosen negative one. The results obtained by the four algorithms
developed by our group are represented in Table 2. Our best result corresponds to the “Enhanced
Automated Image Annotation” algorithm as seen in Figure 3.
3.3
          </p>
        </sec>
      </sec>
      <sec id="sec-3-2">
        <title>Analysis</title>
        <p>
          Our best algorithm was previously tested with the Corel dataset, a collection of 5,000 images
and 374 terms obtaining a statistically significant (5%) improvement over the baseline approach
followed by Yavlinksy et al. [
          <xref ref-type="bibr" rid="ref19">19</xref>
          ]. However, with the IAPR collection and the vocabulary of 17
terms used in VCDT, the results although better were not statistically significant.
4
4.1
        </p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>ImageCLEFphoto</title>
      <sec id="sec-4-1">
        <title>Experiments</title>
        <p>In this section, we describe the nine submitted runs.
Txt. Our text only run uses Lucene as a base with TF•IDF term weighting and comparing
queries to documents using the vector space model. Stemming is performed using the snow ball
stemmer. Stop words are removed and text is lowercased and stripped of diacritics.
Img. This run is a pure CBIR. The dissimilarity value between an image example and an object
image is computed by χ2 Statistics measure based on their HSV colour feature space. Probability
theory is adapted for combining the three dissimilarity values of the image examples in each topic
with an object image. The final result is the Mean Average Precision (MAP) of 39 query topics.
ImgBrfWeightingIterative. This run is a combination of image evidence and text evidence.
BRF is employed in this run. The top seven examples from ranked text results of iteration one
were taken as new image examples for each query topic. The same search method as with “Img”
is used in this run. The final result is the combination of the result ranks of this run and the ranks
of “Img” using weight 0.6 and 0.4.</p>
        <p>ImgBrfWeighting. The differences between this run and “ImgBrfWeightingIterative” is that
the BRF rank for this run was produced by summing the ranks of the first five results of the text
engine’s results when the original query was used as input.</p>
        <p>ImgTxt / TxtGeo / ImgTxtGeo. These three runs are the combination of image evidence
and text evidence, text evidence and geographic evidence, image evidence and text evidence and
geographic evidence, respectively. These evidences were merged by the data fusion process as
described in Section 2.6. The ranks’ and filters’ weights are derived through exhaustive search of
their parameter spaces. The final combination is made with optimal weights (see Table 1).
ImgTxtGeo-ImageCluster. This run combines the text, image and geographic data as with
the ImgTxtGeo run and then clusters the result using the Image Clustering method described in
Section 2.5.</p>
        <p>Run ID MAP
Txt 0.0923
Img 0.0256
ImgBrfWeightingIterative 0.0326
ImgBrfWeighting 0.0217
ImgTxt 0.0715
TxtGeo 0.0696
ImgTxtGeo 0.078
ImgTxtGeo-ImageCluster 0.0465
ImgTxtGeo-MetaCluster 0.0461
ImgTxtGeo-MetaCluster. This run combines the text, image and geographic data as with
the ImgTxtGeo run and then clusters the result using the WordNet Clustering method described
in Section 2.5.</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>5 ImageCLEFwiki</title>
      <p>
        The ImageCLEFwiki task aims to investigate retrieval approaches in the context of a larger scale
and heterogeneous collection of images. The dataset of this task is a wikipedia image collection,
which contains 151,518 images created and employed by the INEX Multimedia track in 2006-2007
[
        <xref ref-type="bibr" rid="ref18">18</xref>
        ]. 75 topics are considered.
5.1
      </p>
      <sec id="sec-5-1">
        <title>Experiments</title>
        <p>In this section, we describe the six submitted runs.</p>
        <p>SimpleText. This run is our baseline – pure text based search. The detailed description can be
found in Section 2.3.</p>
        <p>TextGeoNoContext. This run is a combination of text evidence and geographic evidence. It
uses the MR method or placename disambiguation described in Section 2.4, which is not context
aware. The geographic filter and text rank were combined by multiplying the rank of each
document in the text rank not appearing in the geographic filter by a penalisation value of 3.0 (detailed
in Section 2.6).</p>
        <p>Run ID MAP
SimpleText 0.1918
TextGeoNoContext 0.1896
TextGeoContext 0.1896
Image 0.0037
ImageText 0.1225
ImageTextBRF 0.1225
TextGeoContext. This run is a combination of text evidence and geographic evidence. It uses
the context aware Neighbourhood method or placename disambiguation described in Section 2.4.
Text and geographic evidence are combined in the same way as the TextGeoNoContext run.
Image. This run is a pure content based image search. Gabor texture features were extracted
from the test collection and used to form a high dimensional space. The query images were
compared to the corpus images using the Cityblock distance.</p>
        <p>ImageText. This run is a combination of text evidence and image evidence. The two were
combined using a convex combination of ranks (Section 2.6) using weights 0.3 and 0.7.
ImageTextBRF. This run is a combination of text evidence and image evidence. The same
text and image retrieval system are used as in the previous section. We use the top five results
from the text retrieval results as query images for our BRF system. The three were combined
using a convex combination of ranks using weights 0.25, 0.1 and 0.65 for Image, BRF and Text
relevance respectively.
5.2</p>
      </sec>
      <sec id="sec-5-2">
        <title>Results</title>
        <p>Our conclusions from ImageCLEF’08 are limited as none of our experiments achieved a statistically
significant improvement over a baseline. Discussions of the results are provided below.
VCDT. The enhanced automated image annotation method performed well appearing in the
top quartile of all methods submitted, however it failed to provide significant improvement over
the automated image annotation method. An explanation for this can be found in the small
number of terms of the vocabulary that hinders the functioning of the algorithm and another in
the nature of the vocabulary itself, where instead of incoherence we have mutually exclusive terms
and almost no semantically similar terms.</p>
        <p>ImageCLEFphoto. As mentioned in Section 4.2, despite the combination of Text and Image
evidence performing worse than the text baseline with respect to MAP, the superior GMAP shows
that Image evidence is improving performance in the worst case. In future work we would like
to explore further data fusion methods and find a way to take full advantage of this additional
evidence without undermining text retrieval where it is performing well.</p>
        <p>ImageCLEFwiki. Our text baseline outperformed all other retrieval methods. From this we
conclude that text is by far the dominant feature on retrieval for this heterogeneous collection. The
fact that blind relevance feedback made negligible positive or negative difference to image retrieval
combined with the very low retrieval results for image retrieval alone, implies simple features can
contribute little to retrieval on this collection. Geographic retrieval offered some improvement in
some measures (p@10 and R-prec.), but not statistically significant. We can only conclude that
a geographically aware system could provide some improvement, but due to the short length of
the documents, context based placename disambiguation will be unlikely to provide a significant
improvement.</p>
        <p>WordNet,
online
lexical</p>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>H</given-names>
            <surname>Cunningham</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D</given-names>
            <surname>Maynard</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V</given-names>
            <surname>Tablan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C</given-names>
            <surname>Ursu</surname>
          </string-name>
          , and
          <string-name>
            <given-names>K</given-names>
            <surname>Bontcheva</surname>
          </string-name>
          .
          <article-title>Developing language processing components with GATE</article-title>
          .
          <source>Technical report</source>
          , University of Sheffield,
          <year>2001</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>M</given-names>
            <surname>Van Erp</surname>
          </string-name>
          and
          <string-name>
            <given-names>L</given-names>
            <surname>Schomaker</surname>
          </string-name>
          .
          <article-title>Variants of the Borda Count method for combining ranked classifier hypotheses</article-title>
          . In B. Zhang,
          <string-name>
            <given-names>D.</given-names>
            <surname>Ding</surname>
          </string-name>
          , and L. Zhang, editors,
          <source>International Workshop on Frontiers in Handwriting Recognition</source>
          ,
          <year>2000</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>Tom</given-names>
            <surname>Fawcett</surname>
          </string-name>
          .
          <article-title>An introduction to ROC analysis</article-title>
          .
          <source>Pattern Recognition Letters</source>
          ,
          <volume>27</volume>
          (
          <issue>8</issue>
          ):
          <fpage>861</fpage>
          -
          <lpage>874</lpage>
          ,
          <year>2006</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>A</given-names>
            <surname>Hanbury</surname>
          </string-name>
          and
          <string-name>
            <given-names>J</given-names>
            <surname>Serra</surname>
          </string-name>
          .
          <article-title>Mathematical morphology in the CIELAB space</article-title>
          .
          <source>Image Analysis &amp; Stereology</source>
          ,
          <volume>21</volume>
          :
          <fpage>201</fpage>
          -
          <lpage>206</lpage>
          ,
          <year>2002</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>P</given-names>
            <surname>Harping.</surname>
          </string-name>
          <article-title>User's Guide to the TGN Data Releases</article-title>
          .
          <source>The Getty Vocabulary Program, 2.0 edition</source>
          ,
          <year>2000</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>T K</given-names>
            <surname>Ho</surname>
            , J J Hull
          </string-name>
          , and
          <string-name>
            <given-names>S N</given-names>
            <surname>Srihari</surname>
          </string-name>
          .
          <article-title>Decision combination in multiple classifier systems</article-title>
          .
          <source>Pattern Analysis and Machine Intelligence</source>
          ,
          <volume>16</volume>
          (
          <issue>1</issue>
          ):
          <fpage>66</fpage>
          -
          <lpage>75</lpage>
          ,
          <year>1994</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>P</given-names>
            <surname>Howarth and S Ru</surname>
          </string-name>
          <article-title>¨ger. Robust texture features for still-image retrieval</article-title>
          .
          <source>Vision</source>
          , Image and
          <string-name>
            <given-names>Signal</given-names>
            <surname>Processing</surname>
          </string-name>
          ,
          <volume>6</volume>
          (
          <issue>152</issue>
          (
          <issue>6</issue>
          )):
          <fpage>868</fpage>
          -
          <lpage>874</lpage>
          ,
          <year>2005</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>H</given-names>
            <surname>Liu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D</given-names>
            <surname>Song</surname>
          </string-name>
          , S Ru¨ger, R Hu, and
          <string-name>
            <given-names>V</given-names>
            <surname>Uren</surname>
          </string-name>
          .
          <article-title>Comparing dissimilarity measures for contentbased image retrieval</article-title>
          .
          <source>In Asian Information Retrieval Symposium</source>
          , pages
          <fpage>44</fpage>
          -
          <lpage>50</lpage>
          ,
          <year>2008</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>N</given-names>
            <surname>Maillot</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J</given-names>
            <surname>Chevallet</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V</given-names>
            <surname>Valea</surname>
          </string-name>
          ,
          <string-name>
            <given-names>and J H</given-names>
            <surname>Lim. IPAL</surname>
          </string-name>
          inter
          <article-title>-media pseudo-relevance feedback approach to imageCLEF 2006 photo retrieval</article-title>
          .
          <source>In CLEF 2006 Workshop</source>
          , Working notes,
          <year>2006</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>G A</given-names>
            <surname>Miller and W G Charles</surname>
          </string-name>
          .
          <article-title>Contextual correlates of semantic similarity</article-title>
          .
          <source>Journal of Language and Cognitive Processes</source>
          ,
          <volume>6</volume>
          :
          <fpage>1</fpage>
          -
          <lpage>28</lpage>
          ,
          <year>1991</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>P</given-names>
            <surname>Myers</surname>
            , R Espinosa
          </string-name>
          ,
          <string-name>
            <given-names>C S</given-names>
            <surname>Parr</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T</given-names>
            <surname>Jones</surname>
          </string-name>
          , G S Hammond, and
          <string-name>
            <given-names>T A</given-names>
            <surname>Dewey.</surname>
          </string-name>
          <article-title>The animal diversity web (online)</article-title>
          . http://animaldiversity.org.
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>I</given-names>
            <surname>Niles</surname>
          </string-name>
          and
          <string-name>
            <given-names>A</given-names>
            <surname>Pease</surname>
          </string-name>
          .
          <article-title>Towards a standard upper ontology</article-title>
          .
          <source>In International Conference on Formal Ontology in Information Systems</source>
          , pages
          <fpage>2</fpage>
          -
          <lpage>9</lpage>
          , New York, NY, USA,
          <year>2001</year>
          . ACM Press.
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <given-names>S</given-names>
            <surname>Overell</surname>
          </string-name>
          and
          <string-name>
            <given-names>S</given-names>
            <surname>Ru</surname>
          </string-name>
          <article-title>¨ger. Geographic co-occurrence as a tool for GIR</article-title>
          .
          <source>In CIKM Workshop on Geographic Information Retrieval</source>
          ,
          <year>2007</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <given-names>M J</given-names>
            <surname>Pickering and S Ru</surname>
          </string-name>
          <article-title>¨ger. Evaluation of key frame based retrieval techniques for video</article-title>
          .
          <source>Computer Vision</source>
          and Image Understanding,
          <volume>92</volume>
          (
          <issue>2</issue>
          ):
          <fpage>217</fpage>
          -
          <lpage>235</lpage>
          ,
          <year>2003</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <given-names>Apache</given-names>
            <surname>Lucene</surname>
          </string-name>
          <article-title>Project</article-title>
          . http://lucene.apache.org/java/docs/.
          <source>Accessed 1 August</source>
          <year>2007</year>
          ,
          <year>2007</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [16]
          <string-name>
            <given-names>H</given-names>
            <surname>Tamura</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T</given-names>
            <surname>Mori</surname>
          </string-name>
          , and
          <string-name>
            <given-names>T</given-names>
            <surname>Yamawaki</surname>
          </string-name>
          .
          <article-title>Textural features corresponding to visual perception</article-title>
          .
          <source>Systems, Man and Cybernetics</source>
          ,
          <volume>8</volume>
          (
          <issue>6</issue>
          ):
          <fpage>460</fpage>
          -
          <lpage>473</lpage>
          ,
          <year>1978</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>[17] Princeton University. http://www.cogsci.princeton.edu/∼wn/.</mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          [18]
          <string-name>
            <given-names>T</given-names>
            <surname>Westerveld and R van Zwol</surname>
          </string-name>
          .
          <article-title>The inex 2006 multimedia track</article-title>
          . In N. Fuhr,
          <string-name>
            <given-names>M.</given-names>
            <surname>Lalmas</surname>
          </string-name>
          ,
          <article-title>and</article-title>
          <string-name>
            <surname>A</surname>
          </string-name>
          . Trotman, editors,
          <source>Advances in XML Information Retrieval: INEX 2006</source>
          . Springer-Verlag,
          <year>2007</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          [19]
          <string-name>
            <given-names>A</given-names>
            <surname>Yavlinsky</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E</given-names>
            <surname>Schofield</surname>
          </string-name>
          , and
          <string-name>
            <given-names>S</given-names>
            <surname>Ru</surname>
          </string-name>
          <article-title>¨ger. Automated image annotation using global features and robust non-parametric density estimation</article-title>
          .
          <source>In International ACM Conference on Image and Video Retrieval</source>
          , pages
          <fpage>507</fpage>
          -
          <lpage>517</lpage>
          ,
          <year>2005</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          [20]
          <string-name>
            <given-names>D</given-names>
            <surname>Zhang</surname>
          </string-name>
          and
          <string-name>
            <given-names>G</given-names>
            <surname>Lu</surname>
          </string-name>
          .
          <article-title>Evaluation of similarity measurement for image retrieval</article-title>
          .
          <source>In International Conference on Neural Networks &amp; Signal Processing</source>
          , pages
          <fpage>928</fpage>
          -
          <lpage>931</lpage>
          ,
          <year>2003</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>