<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Scene of Crime Information System: Playing at St. Andrews</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Bogdan Vrusias</string-name>
          <email>b.vrusias@surrey.ac.uk</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Mariam Tariq</string-name>
          <email>m.tariq@surrey.ac.uk</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Lee Gillam</string-name>
          <email>l.gillam@surrey.ac.uk</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Department of Computing University of Surrey</institution>
          ,
          <addr-line>England</addr-line>
        </aff>
      </contrib-group>
      <pub-date>
        <year>1955</year>
      </pub-date>
      <volume>35</volume>
      <fpage>3</fpage>
      <lpage>14</lpage>
      <abstract>
        <p>This paper discusses the adaptation of the Scene of Crime Information System developed within an EPSRC-funded project, to the collection of data within the ImageCLEF track of the Cross Language Evaluation Forum 2003. The adaptations necessary to participate in this activity are detailed, and initial results are briefly presented. ImageCLEF is concerned with the retrieval of images from a specific collection by the captions associated to those images, and is running in relation to an EPSRC-funded project at Sheffield University (Eurovision, GR/R56778/01). The image collection consists of around 28,133 images from the photographic collection provided by St Andrews University Library (Clough et al. 2003). The 28133 images are each referred to and annotated by a single text file, and the full set of annotations are contained within one SGML-based document1. Each annotation comprises identifiers to the text file and the image files (DOCNO, SMALL_IMG, LARGE_IMG), the caption of the image (HEADLINE), a set of categories that have been assigned to this image (CATEGORIES), a database record identifier (RECORD_ID) and an unlabelled chunk of text describing the image, denoted below in italics.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. ImageCLEF Collection</title>
      <p>From the above example, it is apparent that some of the categories assigned to the images may not be wholly reliable.
While some of the associations are clear:
jetty
Fife
rowing (boat)
fishing (rods)
piers and landing stages
Fife all views
rowing boats, rowing, fishing vessels
angling, fresh water fishing, fishing
equipment
others could be associated to information that appears, but is not in the correct context – the combination of “Open
Championship” and “St Andrews” being candidates for explaining the golfing categories – while the assignment of a
“battlefields” category is less easily obvious.</p>
    </sec>
    <sec id="sec-2">
      <title>2. Task 1: Automatic Ad Hoc Retrieval</title>
      <p>The automatic ad hoc retrieval task aims at the ranked-retrieval of upto 1000 images from the Eurovision collection. The
images are to be retrieved in response to a set of pre-formulated queries. The queries themselves comprise of 50 topics.
Each topic has an English query, plus narrative description of the expected result of the query, and the English query has
been translated into 5 other languages, French, German, Spanish, Italian and Dutch. Some queries have more than one
translation for a given language.</p>
      <p>The retrieval results are to be assessed by personnel from the University of Sheffield such that they can be evaluated
using the trec_eval program with recall and precision metrics. Similar to TREC, the results will subsequently be
published.</p>
      <p>An example topic encoded in XML2 is shown below:
&lt;top&gt;
&lt;num&gt;Number: 25&lt;/num&gt;
&lt;EN-title n="1"&gt;Golf course bunkers&lt;/EN-title&gt;
&lt;EN-narr&gt;A relevant image will show a picture of a golf course in which a bunker can be clearly identified. The picture
must be a photograph or a postcard, but not a drawing, e.g. a plan of the golf course. A bunker is a sandy hollow
formed by wearing away of the turf, or nowadays an artificial sand-hole with a built-up face. An example relevant
document is [stand03_1714/stand03_7020].&lt;/EN-narr&gt;
&lt;/top&gt;
&lt;top&gt;
&lt;num&gt;Number: 25&lt;/num&gt;
&lt;DE-title n="1"&gt;Golfplatz Bunker&lt;/DE-title&gt;
&lt;/top&gt;
&lt;top&gt;
&lt;num&gt;Number: 25&lt;/num&gt;
&lt;FR-title n="1"&gt;Bunkers de terrain de golfe&lt;/FR-title&gt;
&lt;/top&gt;
&lt;top&gt;
&lt;num&gt;Number: 25&lt;/num&gt;
&lt;IT-title n="1"&gt;Un bunker in un percorso di golf&lt;/IT-title&gt;
&lt;IT-title n="2"&gt;bunkers in un campo di golf&lt;/IT-title&gt;
&lt;/top&gt;
&lt;top&gt;
&lt;num&gt;Number: 25&lt;/num&gt;
&lt;ES-title n="1"&gt;B&amp;#x00FA;nkers en un campo de golf&lt;/ES-title&gt;
&lt;ES-title n="2"&gt;Pista de golf&lt;/ES-title&gt;
&lt;/top&gt;
&lt;top&gt;
&lt;num&gt;Number: 25&lt;/num&gt;
&lt;NL-title n="1"&gt;Bunkers op een golfbaan&lt;/NL-title&gt;
&lt;/top&gt;</p>
      <p>The example shown is for Topic 25, for which Golf course bunkers has been translated once into each of German,
French and Dutch, and twice each for Spanish and Italian. With multiple translations for some languages for the 50
topics, we have the following number of queries for the various languages:
2 Similar character issues as reported previously were also fixed for this collection.</p>
      <p>These 421 queries are to be made against the 28,133 annotations to retrieve images from the collection.</p>
    </sec>
    <sec id="sec-3">
      <title>3. The SoCIS Archetype</title>
      <p>The EPSRC-funded Scene of Crime Information System (SoCIS) project was run from October 1999 to March 2003.
The aim of the project was to study the link between images and texts within a specialist domain context. A method has
been outlined for developing an intelligent content-based image retrieval (CBIR) system, which can store and retrieve
images based on the linguistic descriptions of the images. The corpus-based method uses the lexical and semantic
properties of specialist texts for extracting key terms and for discovering the ontological organisation of the terms.</p>
      <p>
        A prototype CBIR system was developed in the Java programming language for demonstrating the efficacy of the
corpus-based method. The system, which is based on a 3-tier architecture of client, server, and database, can be accessed
via a local intranet. SoCIS is an intelligent CBIR system that automatically: (a) labels (and indexes) images by keywords
as well as relational facts extracted from the descriptions provided by domain experts; (b) extracts physical features of an
image; (c) populates a database comprising domain-specific terminology, together with the semantic relationships
between terms, starting from a random selection of collateral texts of the domain; and (d) learns to link image and text by
using neural networks
        <xref ref-type="bibr" rid="ref1 ref10">(Ahmad et al., 2002)</xref>
        . SoCIS has integrated modules from (a) System Quirk
        <xref ref-type="bibr" rid="ref2">(Ahmad &amp; Rogers,
2001)</xref>
        - a set of tools for building and managing multilingual term bases with the use of powerful text analysis
techniques, and (b) GATE
        <xref ref-type="bibr" rid="ref9">(Cunningham et al., 2002)</xref>
        - a framework and graphical development environment comprising
robust NLP tools. The main advantages that SoCIS can be said to have over other text-based and CBIR systems is its
ability to extract information from both texts and images, to encode this information for indexing, and to build thesauri,
all automatically.
      </p>
      <p>
        The SoCIS prototype3 was evaluated using images normally used for the training of Scene of Crime Officers (SoCOs)
together with a description provided by the SoCOs as well as other collateral texts like crime scene reports and forensic
science research papers and manuals. The question of (inter) indexer-variability, the variances in the output of different
indexers for the same image, has been explored in the project
        <xref ref-type="bibr" rid="ref11 ref3 ref4">(Handy &amp; Ahmad, 2003)</xref>
        . This study further reinforced the
need for automatic thesauri construction to aid in query expansion
        <xref ref-type="bibr" rid="ref11 ref3 ref4">(Ahmad et al., 2003a)</xref>
        .
      </p>
    </sec>
    <sec id="sec-4">
      <title>4. Adapting SoCIS</title>
      <p>
        SoCIS was specifically targeted at the use of specialist languages – or Languages for Special Purposes (LSP)
        <xref ref-type="bibr" rid="ref12 ref2">(Harris,
1988, Ahmad &amp; Rogers, 2001)</xref>
        . The system has been built based on the knowledge gathered from Scene of Crime
experts, from the testing and evaluation sessions performed with them, and from a domain-specific text corpus. The
system had to be adapted to deal with multilinguality as well as structured data from a more general domain for the
ImageCLEF collection. SoCIS does not have a translation tool so the translation of the queries from the other languages
to English had to be carried out offline as discussed in section 4.1. A parser had to be written to extract the various fields
containing textual information (in English) about the images from the provided XML document that could be used for
indexing purposes. The indexing module was used to extract single and compound terms from the output of the parser.
The main difficulty we encountered (see section 4.2) was the creation of a terminology dictionary and thesaurus related
to the general domain, which is needed for the automatic indexing and query expansion modules. We decided to use
Wordnet for query expansion purposes but the indexing had to be carried out without using a terminology dictionary to
filter out invalid terms. A new relevance ranking mechanism, which is briefly described in section 4.3, was adopted to
handle the expanded terms retrieved from Wordnet.
4.1.
      </p>
      <sec id="sec-4-1">
        <title>Handling Multilinguality</title>
        <p>The first step necessary was the translation of the various queries to English. Without in-house software, we relied upon
translation engines as found on the Internet. Some work was done in an attempt to exploit Google’s translation tools for
this purpose, however there were difficulties encountered in this. Eventually, Altavista’s Babelfish was selected as the
principal translation engine (http://babelfish.altavista.com/), however since this system does not translate Dutch,
FreeTranslation.com (http://www.freetranslation.com/) was also used.</p>
        <p>
          To translate the queries, Java code was used to wrap definitions of the query syntax used by these sites (with the
HTTP POST command being used in both cases). Each query was posted to the site with its requested translation
language pair, and the HTML result was retrieved. Using the Java JTidy utility, the resulting HTML was converted to
XML
          <xref ref-type="bibr" rid="ref6">(Bray et al, 2000)</xref>
          , and XSLT
          <xref ref-type="bibr" rid="ref7">(Clark, 1999)</xref>
          employed to strip out the end result of the translation.
        </p>
        <sec id="sec-4-1-1">
          <title>3 http://www.surrey.ac.uk/socis</title>
          <p>The results of translating the various languages for topic number 25 (Golf course bunkers) are shown in the table
below:</p>
          <p>German Golf course shelter
French Bunkers of ground of gulf
Italian (1) A bunker in a distance of golf
Italian (2) bunkers in a golf course
Spanish (1) B??nkers in a golf course
Spanish (2) Track of golf</p>
          <p>Dutch Bunkers on a wave job</p>
          <p>Immediately, certain of these translations will cause problems with the retrieval. The topic identifies the image
stand03_1714/stand03_7020 as being relevant. In the run, this was located only for English, Italian (2), and Dutch at
ranks 798, 798 and 45 respectively. The quality of returned translation will therefore have a significant impact on the
results being returned.
4.2.</p>
        </sec>
      </sec>
      <sec id="sec-4-2">
        <title>Synonymy and Morphology</title>
        <p>
          The thesaurus construction module of SoCIS was developed to provide a query expansion facility for the system. There
are general-purpose thesauri or lexicons available such as Wordnet4, which could be used but are inadequate in specialist
domains due to a deficiency in specialized terminology. For example, the two key compound terms ‘forensic science’
and ‘crime scene’ are not present in Wordnet. The method we developed was based on the analysis of a representative
domain-specific text corpus to automatically extract key terms and relationships, which were then used to build the
thesaurus
          <xref ref-type="bibr" rid="ref11 ref20 ref3 ref4">(Ahmad et al., 2003a, Tariq et al., 2003)</xref>
          . Since the ImageCLEF collection comprised of a wide range of
mainly general topics such as buildings, golfers, animals, boats and so on, to apply our method we would have had to
construct and analyze a corpus representing most of general knowledge, a clearly difficult and unpractical task. We
decided that Wordnet could be a possible resource to use for query expansion since its coverage is based on a general
English dictionary.
        </p>
        <p>A program was written to query a Wordnet database to provide a set of synonyms and hyponyms for each of the
query terms. In Wordnet, English nouns, verbs, adjectives and adverbs are ordered into synonym sets (synsets). Each
synset can be said to contain the words that represent a specific concept. The synsets are then linked to each other based
on semantic relations such as antonymy, hyponymy and meronymy. Given a query term, the program returns all the
words in the synset that the particular term is an element of, as well as all the hyponyms of each synset element to a
specified level in the hierarchy. Initially we planned to go down 2 levels in the hierarchy but ended up using just the
synonyms due to system performance issues related to the large number of expanded terms returned, which is discussed
in section 5. Taking the query “Boats on Loch Lomond” as an example, the term ‘boat’ returned 53 expanded words
going down one level in the hierarchy. Some synonyms returned were: travel on water, sauceboat, gravy boat; some
hyponyms returned included motorboat, mail boat, mailboat gondola, propel by oars, propel by paddles, yacht, and so
on. ‘Loch’ returned one synonym lough while ‘Lomond’ was not present since it is a proper noun. The very common
term ‘man’ had 131 expanded words going down one level and 344 expanded words going down two levels with words
such as private, make swollen, belly out, candy striper, Homo erectus, clothes horse, ridicule with a satire, and
gentleman.</p>
        <p>Some basic morphological analysis was also carried out for each query term to account for the use of variants such as
singular or plural terms as well as the verb or adjective forms. The morphology module uses standard rules (for example
if a word ends with ‘ss’ or ‘h’ then the plural form is usually derived by adding an ‘es’) as well as some common
exceptions (for example the plural of ~man will be ~men). This was also important for the query expansion part since
Wordnet only has singular forms of words as part of the synsets so a plural word used as the query term will return no
results.
4.3.</p>
      </sec>
      <sec id="sec-4-3">
        <title>Relevance Ranking</title>
        <p>Each keyword carried a proportion of its frequency in an annotation divided by the total number of terms allocated to this
annotation. The original keyword was then multiplied with weight 1, each expanded term (synonyms) returned by
WordNet with weight 0.9, and words containing substrings of the original keywords with weight 0.1. The total ranking
was then given by:</p>
        <p>Rank = ∑  td × w 
 f</p>
        <p>t 
 N d </p>
        <p>Where ftd is the term frequency of term t in document d, wt is the weight of a term t as described previously, and Nd
is the total number of words in document d.</p>
        <sec id="sec-4-3-1">
          <title>4 http://www.cogsci.princeton.edu</title>
        </sec>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>5. Performance Issues</title>
      <p>The main factor to have an effect on the performance of SoCIS was that the system has been designed for the analysis
of free text in specialist domains whereas with the ImageCLEF collection we were dealing with structured texts in a
general domain. This resulted in difficulties for SoCIS when indexing the images – the indices produced were relatively
unreliable due to the different syntactic structure of the ImageCLEF text when compared to free text, which also affected
the ranking. One example here is that the system considered all the category terms given by the ImageCLEF description
in the XML document (since they where enclosed in square brackets) as a single compound term. Also due to the fact
that we used Wordnet for query expansion, we encountered problems associated with polysemous words as well as
different word forms (see the example of boat and man in section 4.2). Due to the amount of time it was taking to
process the expanded queries (some times reaching up to 300 words, see section 4.2) we had to limit the expansion to just
synonyms of the original query terms. Even so we had six computers running in parallel to finish the processing, which
was taking approximately 8 hours per language.</p>
    </sec>
    <sec id="sec-6">
      <title>6. Results and Evaluation</title>
      <p>Although the combination of features outlined above would require significant efforts to develop as a usable real-world
system (parallelisation and optimisation issues at least), the combination of technologies and techniques presented did
enable participation in the ImageCLEF track. A system that in principle would allow a user to query a collection of
images that have been annotated in English, using a query in one of six languages has been prototyped from this
combination. According to the abstract from the Eurovision project, such a system had not been implemented or
researched. Though far from perfect, the evaluation of the results obtained at this stage is important.</p>
      <p>Across all languages, the following sets of results were obtained (missing topics and quantities for that topic are given
in the third column):</p>
      <sec id="sec-6-1">
        <title>Spanish</title>
        <p>For this selection of 6 topics, the exemplar image is only found for English for topic 21. This is an initially
disappointing result. We consider, first, the top image being retrieved for each of these topics.
7 (En) stand03_1749/stan</p>
        <p>d03_22144
14 (En) stand03_1502/stan</p>
        <p>d03_16737
21 (En) stand03_1675/stan</p>
        <p>d03_22740
28 (En) stand03_1714/stan</p>
        <p>d03_7540
35 (En) stand03_1851/stan</p>
        <p>d03_7899
42 (En) stand03_1590/stan</p>
        <p>d03_28349
7 (Fr)
14 (Fr)
21 (Fr)
28 (Fr)
35 (Fr)
42 (Fr)</p>
        <p>No results
stand03_1853/stan
d03_12134
stand03_1675/stan
d03_22740
stand03_2046/stan
d03_13818
stand03_1851/stan
d03_7899
stand03_1590/stan
d03_28349
7 (De)</p>
        <p>No results
14 (De) stand03_1857/stan</p>
        <p>d03_9586
21 (De) stand03_1675/stan</p>
        <p>d03_22740
28 (De) stand03_2046/stan</p>
        <p>d03_13818
35 (De) stand03_1851/stan</p>
        <p>d03_7899
42 (De) stand03_2054/stan</p>
        <p>Littlehampton. The</p>
      </sec>
      <sec id="sec-6-2">
        <title>Parade.</title>
        <p>The Castle, Loch an
Eilein
Engraving of a painting
of a Biblical scene,
[Noah, family and the
Ark at Mount Ararat].</p>
        <p>Old Tom Morris, golfer,
St Andrews. (ca 1900)
Trossachs. Loch
Achray, Trossachs
Church and Ben An or
Binnein (Ben A 'an).</p>
        <p>Samuel Messieux,
refugee from Paris and
teacher of French at
Madras College [South
Street], St Andrews.</p>
        <p>Boat of Garten.</p>
        <p>Engraving of a painting
of a Biblical scene,
[Noah, family and the
Ark at Mount Ararat].</p>
        <p>Kingsbarns. Old Grave
Stone, Kingsbarns
Churchyard.</p>
        <p>Trossachs. Loch
Achray, Trossachs
Church and Ben An or
Binnein (Ben A 'an).</p>
        <p>Samuel Messieux,
refugee from Paris and
teacher of French at
Madras College [South
Street], St Andrews.</p>
        <p>View of ship at sea.</p>
        <p>Engraving of a painting
of a Biblical scene,
[Noah, family and the
Ark at Mount Ararat].</p>
        <p>Kingsbarns. Old Grave
Stone, Kingsbarns
Churchyard.</p>
        <p>Trossachs. Loch
Achray, Trossachs
Church and Ben An or
Binnein (Ben A 'an).</p>
        <p>Motherwell. Town Hall.</p>
        <p>7 (It)
14 (It)
21 (It)
28 (It)
35 (It)
42 (It)
d03_18895
stand03_1587/stan
d03_28525
stand03_1502/stan
d03_16737
stand03_1675/stan
d03_22740
stand03_2046/stan
d03_13818
stand03_1778/stan
d03_4502
stand03_2054/stan
d03_18895
7 (Es) stand03_1587/stan</p>
        <p>d03_7524
14 (Es) stand03_1853/stan</p>
        <p>d03_12134
21 (Es) stand03_1675/stan</p>
        <p>d03_22740
28 (Es) stand03_2046/stan</p>
        <p>d03_13818
35 (Es) stand03_2092/stan</p>
        <p>d03_14170
42 (Es) stand03_1590/stan</p>
        <p>d03_28349
7 (Nl) stand03_1587/stan</p>
        <p>d03_7524
14 (Nl) stand03_1502/stan</p>
        <p>d03_16737
21 (Nl) stand03_1974/stan</p>
        <p>d03_11773
28 (Nl) stand03_2046/stan</p>
        <p>d03_13818
35 (Nl) stand03_1853/stan</p>
        <p>d03_21295
42 (Nl) stand03_2054/stan
d03_18895
[Walker family?]
Untitled portrait of a
man.</p>
        <p>The Castle, Loch an
Eilein
Engraving of a painting
of a Biblical scene,
[Noah, family and the
Ark at Mount Ararat].</p>
        <p>Kingsbarns. Old Grave
Stone, Kingsbarns
Churchyard.</p>
        <p>Launch X.</p>
        <p>Motherwell. Town Hall.</p>
        <p>Man in theatrical
costume. [St Andrews ?].</p>
        <p>Boat of Garten.</p>
        <p>Engraving of a painting
of a Biblical scene,
[Noah, family and the
Ark at Mount Ararat].</p>
        <p>Kingsbarns. Old Grave
Stone, Kingsbarns
Churchyard.</p>
        <p>Lochgilphead. Crinan
Canal at
Samuel Messieux,
refugee from Paris and
teacher of French at
Madras College [South
Street], St Andrews.</p>
        <p>No results
The Castle, Loch an
Eilein
Brompton Oratory. Altar
of Our Lady of Good
Counsel.</p>
        <p>Kingsbarns. Old Grave
Stone, Kingsbarns
Churchyard.</p>
        <p>Fettercairn. Cairn o'
Mount and Clatterin'
Brig
Motherwell. Town Hall.</p>
        <p>These tables of results show some interesting features. For Topic 7, 3 of the queries returned no results, while those
that did have a different first result. For topic 14, 5 of the 6 results refer to just 2 images. For topic 21, all but then Dutch
result refer to the same image. For topic 28, all but the English result refer to the same image, however judging by the
caption, the English result is the best. For topic 35, one image is referred to in 3 results. For topic 42, 2 images are
equally referred to.</p>
        <p>For Topic 14, the top 5 results have been taken once for each language, and the similarity matrix between these
results is as follows:</p>
        <p>The top 5 results show degrees of similarity between the English, Italian and Dutch results, with German and Spanish
showing similarities, and French showing the most marked behavioural difference. This top 5 have captions as follows:
The Castle, Loch an Eilein
Boat of Garten.</p>
        <p>Dunkeld. Loch of Craiglush and Creag nam Mial (Creagnam Hill).</p>
        <p>Linlithgow Palace and Loch, from the air.</p>
        <p>Bearsden. St Germain's Loch.</p>
        <p>It would appear that a number of Lochs, apart from Loch Lomond with any boats on have been discovered in
response to this query! Indeed, none of the 16 results above make mention of Lomond.</p>
        <p>From this, it is apparent that although similar behaviour is achieved for certain language translations, the end result of
retrieval is not correctly weighted. The initial concern that translation would have a significant bearing on retrieval is
perhaps now not so relevant as the retrieval itself.</p>
        <p>Taking a list of the exemplar images for retrieval, the ranking (where it exists) of that image within the 1000 results
for each language was considered. For each language, if the exemplar image was retrieved within the first 1000, this was
counted. If it was retrieved within the top 100 results, this was also noted. The following table presents the results
obtained.</p>
      </sec>
      <sec id="sec-6-3">
        <title>High Low Ave</title>
        <p>Nl
De
En
Fr
It
Es
20
21
28
33
42
51
50
50
50
51
103
117
17
11
8
11
1
7
971
973
798
995
967
884
319.55
309.05
257.96
337.52
353.14
314.52</p>
        <p>In top
100
7
11
12
11
14
19</p>
        <p>For two queries in Italian, both for Topic 19, the exemplar image was retrieved in first place. This is certainly a result
of interest given the analysis of other results in this paper. In the above table, the first column represents the language
code, the second the amount of exemplar images retrieved in the 1000 results, the third is the amount of queries, the
fourth and fifth show the highest and lowest ranking of the exemplars, with the sixth column showing the average
ranking. Column 7 shows the quantity of exemplars occurring in the first 100 retrieved results. This set of results tends
to indicate that there is some value to the approach taken here, but how that compares to other approaches remains to be
seen.</p>
      </sec>
    </sec>
    <sec id="sec-7">
      <title>7. A Note on Text and Image Retrieval</title>
      <p>
        Increasingly, images are being indexed and retrieved by both their visual content and by related texts such as captions
that describe the image
        <xref ref-type="bibr" rid="ref13 ref18 ref19 ref5">(Srihari, 1995, Srihari et al., 2000, Paek et al., 1999, Barnard &amp; Forsyth, 2001)</xref>
        . Image
descriptors extracted directly from image data (colour, texture and shape) tend to capture little of an image’s semantic
content
        <xref ref-type="bibr" rid="ref17">(Squire et al., 2000, Eakins, 2002)</xref>
        – hence there is a need to extract information about the image content from
collateral texts
        <xref ref-type="bibr" rid="ref1 ref10 ref11 ref15 ref16 ref3 ref4">(Smeulders et al., 2000, Gillam et al., 2002, Salway &amp; Frehen, 2002, Ahmad et al., 2003a)</xref>
        .
      </p>
    </sec>
    <sec id="sec-8">
      <title>8. Future Work</title>
      <p>Numerous improvements suggest themselves, for example if the system could be grid-enabled then the different
processing modules, as well as instances of the same module, could be run as a service, in parallel, which would
significantly improve the processing time. The ranking mechanism needs to be further refined and tuned by carrying out
more trial runs. To improve the query expansion one suggestion could be to use part-of-speech information from the
query sentence to filter out some of the irrelevant expanded terms returned from Wordnet – for example in the query
“Boats on loch Lomond”, the term boat is being used in the noun and not the verb form so the synonyms (propel by oars,
propel by paddles) related to the verb form of boat would not be retrieved. Otherwise, an attempt could be made to
analyze the British National Corpus5, which might perhaps yield only the more frequently used term associations.</p>
      <p>
        Since the system deals with image retrieval, we are investigating methods of effectively combining text-based with
image-based retrieval techniques. The physical features of an image such as colour, texture, and shape can be extracted
and used in combination with the text features. This technique when incorporated into a system that learns how to index,
would result in a significant improvement in performance
        <xref ref-type="bibr" rid="ref11 ref3 ref4">(Ahmad et al., 2003b)</xref>
        . We are also investigating the creation
of multimedia thesauri, based on Picard’s initial work
        <xref ref-type="bibr" rid="ref14">(Picard, 1995)</xref>
        . The premise here is that since specialist texts can
be said to be a reflection of the ontological commitment of domain experts, specialist images may also reflect some form
of ontological commitment on the part of the expert. Also, objects depicted in specialist images often represent the same
concepts that are represented by lexical units in texts. The method discussed in
        <xref ref-type="bibr" rid="ref1">Ahmad et al. (2002)</xref>
        could help in
establishing the link between an image and text.
      </p>
    </sec>
    <sec id="sec-9">
      <title>9. Acknowledgements</title>
      <p>This work was partially funded by the European Union through the Generic Information-based Decision Assistant
(GIDA: IST-2000-31123) Projects and the EPSRC through the Scene of Crime Information System (SOCIS:
GR/M89041/01) project.</p>
      <sec id="sec-9-1">
        <title>5 http://www.hcu.ox.ac.uk/BNC</title>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          <string-name>
            <surname>Ahmad</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          <string-name>
            <surname>Vrusias</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Tariq</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          , (
          <year>2002</year>
          )
          <article-title>“Co-operative neural networks and integrated classification”</article-title>
          .
          <source>In Proceedings of the International Joint Conference on Neural Networks. Hawaii</source>
          , USA May
          <year>2002</year>
          . IEEE Press.
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          <string-name>
            <surname>Ahmad</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          and
          <string-name>
            <surname>Rogers</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          , (
          <year>2001</year>
          ) “
          <article-title>Corpus linguistics and terminology extraction”</article-title>
          . In Wright S.E. and
          <string-name>
            <surname>Budin</surname>
            <given-names>G</given-names>
          </string-name>
          . (eds.)
          <source>Handbook of Terminology Management</source>
          , Vol.
          <volume>2</volume>
          , Amsterdam/Philadelphia: Benjamins, pp.
          <fpage>725</fpage>
          -
          <lpage>760</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          <string-name>
            <surname>Ahmad</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Tariq</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Vrusias</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          and
          <string-name>
            <surname>Handy</surname>
            <given-names>C.</given-names>
          </string-name>
          (
          <year>2003a</year>
          ) “
          <article-title>Corpus-Based Thesaurus Construction for Image Retrieval in Specialist Domains”</article-title>
          . In F. Sebastiani (ed.)
          <source>: Proceedings of the 25th European Conference on Information Retrieval Research</source>
          , ECIR-
          <volume>03</volume>
          , Pisa,
          <source>Italy LNCS-2633</source>
          . Heidelberg: Springer Verlag. pp
          <fpage>502</fpage>
          -
          <lpage>510</lpage>
          ,.
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          <string-name>
            <surname>Ahmad</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Casey</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Vrusias</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Saragiotis</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          (
          <year>2003b</year>
          ) “
          <article-title>Combining Multiple Modes of Information Using Unsupervised Neural Classifiers”</article-title>
          . In: Windeatt,
          <string-name>
            <given-names>T.</given-names>
            and
            <surname>Roli</surname>
          </string-name>
          ,
          <string-name>
            <surname>F</surname>
          </string-name>
          . (eds.),
          <source>Proceedings of Multiple Classifier Systems 4th Int. Workshop</source>
          , Guildford,
          <string-name>
            <surname>UK</surname>
          </string-name>
          , June 11-13,
          <year>2003</year>
          , LNCS 2709. Heidelberg: Springer-Verlag, pp.
          <fpage>236</fpage>
          -
          <lpage>245</lpage>
          ,
          <year>2003b</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          <string-name>
            <surname>Barnard</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Forsyth</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          (
          <year>2001</year>
          )
          <article-title>“Learning the Semantics of Words and Pictures”</article-title>
          .
          <source>International Conference on Computer Vision</source>
          , Vol
          <volume>2</volume>
          , pp
          <fpage>408</fpage>
          -
          <lpage>415</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          <string-name>
            <surname>Bray</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Paoli</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sperberg-McQueen</surname>
            ,
            <given-names>C.M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Maler</surname>
          </string-name>
          , E. (eds.), (
          <year>2000</year>
          ).
          <article-title>“Extensible Markup Language (XML)”</article-title>
          , Version 2.0.
          <string-name>
            <given-names>W3C</given-names>
            <surname>Recommendation</surname>
          </string-name>
          . http://www.w3.org/TR/REC-xml
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          <string-name>
            <surname>Clark</surname>
            ,
            <given-names>J</given-names>
          </string-name>
          . (ed.), (
          <year>1999</year>
          ).
          <article-title>“XSL Transformations (XSLT)”</article-title>
          , Version 1.0.
          <string-name>
            <given-names>W3C</given-names>
            <surname>Recommendation</surname>
          </string-name>
          . http://www.w3.org/TR/xslt
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          <string-name>
            <surname>Clough</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sanderson</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Reid</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          (
          <year>2003</year>
          ).
          <article-title>“The Eurovision St Andrews Photographic Collection (ESTA)”</article-title>
          . http://ir.shef.ac.uk/imageclef/guide.pdf (
          <year>February 2003</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          <string-name>
            <surname>Cunningham</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          <string-name>
            <surname>Maynard</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          <string-name>
            <surname>Bontcheva</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          <string-name>
            <surname>Tablan</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          (
          <year>2002</year>
          )
          <article-title>“GATE: A framework and graphical development environment for robust NLP tools and applications”</article-title>
          .
          <source>In Proceedings of the 40th Anniversary Meeting of the Association for Computational Linguistics</source>
          ,
          <year>2002</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          <string-name>
            <surname>Gillam</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ahmad</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          and
          <string-name>
            <surname>Salway</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          (
          <year>2002</year>
          )
          <article-title>Digital Heritage and the use of Terminology</article-title>
          .
          <source>In Proceedings of 6th International Conference Terminology and Knowledge Engineering (TKE) 2002 ISBN 2-</source>
          <fpage>7261</fpage>
          -1217-X
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          <string-name>
            <surname>Handy</surname>
            ,
            <given-names>C.J.</given-names>
          </string-name>
          and
          <string-name>
            <surname>Ahmad</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          (
          <year>2003</year>
          )
          <article-title>“Indexer Variability in Visual Domains”</article-title>
          . To appear
          <source>in: Proceedings of the 13th LSP Conference.</source>
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          <string-name>
            <surname>Harris</surname>
            ,
            <given-names>Z.S.</given-names>
          </string-name>
          , (
          <year>1998</year>
          )
          <article-title>“Language and Information”</article-title>
          . In: Nevin,
          <string-name>
            <surname>B</surname>
          </string-name>
          . (ed.)
          <source>Computational Linguistics</source>
          , Vol.
          <volume>14</volume>
          , No.4, Columbia University Press, New York, pp.
          <fpage>87</fpage>
          -
          <lpage>90</lpage>
          ,
          <year>1988</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          <string-name>
            <surname>Paek</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sable</surname>
            ,
            <given-names>C. L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hatzivassiloglou</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Jaimes</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Schiffman</surname>
            ,
            <given-names>B.H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chang</surname>
            ,
            <given-names>S-F.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>McKeown</surname>
            ,
            <given-names>K. R.</given-names>
          </string-name>
          (
          <year>1999</year>
          )
          <article-title>“Integration of visual and text based approaches for the content labelling and classification of</article-title>
          <source>Photographs” ACM SIGIR'99 Workshop on Multimedia Indexing and Retrieval</source>
          , Berkeley, California, USA.
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          <string-name>
            <surname>Picard</surname>
            ,
            <given-names>R.W.</given-names>
          </string-name>
          , (
          <year>1995</year>
          ) “
          <article-title>Towards a Visual Thesaurus”</article-title>
          . In: Ian Ruthven (ed.) Springer Verlag Workshops in Computing, MIRO
          <volume>95</volume>
          ,
          <string-name>
            <surname>Glasgow</surname>
          </string-name>
          , Scotland.
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          <string-name>
            <surname>Salway</surname>
          </string-name>
          and
          <string-name>
            <surname>Frehen</surname>
          </string-name>
          (
          <year>2002</year>
          ), “
          <article-title>Words for Pictures: analysing a corpus of art texts”</article-title>
          .
          <source>In Proceedings of 6th International Conference Terminology and Knowledge Engineering (TKE) 2002 ISBN 2-</source>
          <fpage>7261</fpage>
          -1217-X
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          <string-name>
            <surname>Smeulders</surname>
            ,
            <given-names>A.W.M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Worring</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Santini</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gupta</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Jain</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          (
          <year>2000</year>
          ) “
          <article-title>Content-Based Image Retrieval at the End of the early Years”</article-title>
          .
          <source>IEEE Transactions on Pattern Analysis and Machine Intelligence</source>
          , Vol.
          <volume>22</volume>
          , No. 12, IEEE Press, pp.
          <fpage>1349</fpage>
          -
          <lpage>1380</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          <string-name>
            <surname>Squire</surname>
          </string-name>
          ,
          <string-name>
            <surname>McG</surname>
            .
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Muller</surname>
            ,
            <given-names>W.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Muller</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Pun</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          (
          <year>2000</year>
          ) “
          <article-title>Content-Based Query of Image databases: Inspirations from Text Retrieval”</article-title>
          .
          <source>Pattern Recognition Letters</source>
          , Vol.
          <volume>21</volume>
          . No.
          <volume>13</volume>
          -
          <fpage>14</fpage>
          . Elsevier Science, Netherlands, pp.
          <fpage>1193</fpage>
          -
          <lpage>1198</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          <string-name>
            <surname>Srihari R.K.</surname>
          </string-name>
          , (
          <year>1995</year>
          ) “
          <article-title>Use of Collateral Text in Understanding Photos”</article-title>
          .
          <source>Artificial Intelligence Review (Special Issue on Integrating Language and Vision)</source>
          , Vol.
          <volume>8</volume>
          , pp.
          <fpage>409</fpage>
          -
          <lpage>430</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          <string-name>
            <surname>Srihari</surname>
            ,
            <given-names>R.K.</given-names>
          </string-name>
          and
          <string-name>
            <surname>Zhang</surname>
            ,
            <given-names>Z.</given-names>
          </string-name>
          (
          <year>2000</year>
          )
          <article-title>“Show&amp;Tell: a Semi-Automated Image Annotation System”</article-title>
          .
          <source>IEEE Multimedia</source>
          , Vol.
          <volume>7</volume>
          , No.
          <issue>3</issue>
          , pp.
          <fpage>61</fpage>
          -
          <lpage>71</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          <string-name>
            <surname>Tariq</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Manumaisupat</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Al-Sayed</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          and
          <string-name>
            <surname>Ahmad</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          , (
          <year>2003</year>
          )
          <article-title>“Experiments in Ontology Construction from Specialist Texts”</article-title>
          . To appear
          <source>in: Proceedings of EUROLAN Workshop: Ontologies and Information Extraction</source>
          , Bucharest, Romania,
          <source>July 28 -August</source>
          <volume>08</volume>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>