<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Overview of the ImageCLEF 2007 Medical Retrieval and Annotation Tasks</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Henning Mu¨ller</string-name>
          <email>henning.mueller@sim.hcuge.ch</email>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Thomas Deselaers</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Eugene Kim</string-name>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Jayashree Kalpathy-Cramer</string-name>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Thomas M. Deserno</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>William Hersh</string-name>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Computer Science Dep., RWTH Aachen University</institution>
          ,
          <country country="DE">Germany</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Medical Informatics, RWTH Aachen University</institution>
          ,
          <country country="DE">Germany</country>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>Medical Informatics, University and Hospitals of Geneva</institution>
          ,
          <country country="CH">Switzerland</country>
        </aff>
        <aff id="aff3">
          <label>3</label>
          <institution>Oregon Health and Science University (OHSU)</institution>
          ,
          <addr-line>Portland, OR</addr-line>
          ,
          <country country="US">USA</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>This paper describes the medical image retrieval and medical image annotation tasks of ImageCLEF 2007. Separate sections describe each of the two tasks, with the participation and an evaluation of major findings from the results of each given. A total of 13 groups participated in the medical retrieval task and 10 in the medical annotation task. The medical retrieval task added two news data sets for a total of over 66'000 images. Tasks were derived from a log file of the Pubmed biomedical literature search system, creating realistic information needs with a clear user model in mind. The medical annotation task was in 2007 organised in a new format as a hierarchical classification had to be performed and classification could be stopped at any confidence level. This required algorithms to change significantly and to integrate a confidence level into their decisions to be able to judge where to stop classification to avoid making mistakes in the hierarchy. Scoring took into account errors and unclassified parts.</p>
      </abstract>
      <kwd-group>
        <kwd>Image Retrieval</kwd>
        <kwd>Performance Evaluation</kwd>
        <kwd>Image Classification</kwd>
        <kwd>Medical Imaging</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>
        ImageCLEF1 [
        <xref ref-type="bibr" rid="ref2 ref3">3, 2</xref>
        ] started within CLEF2 (Cross Language Evaluation Forum [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ]) in 2003 with
the goal to benchmark image retrieval in multilingual document collections. A medical image
1http://ir.shef.ac.uk/imageclef/
2http://www.clef-campaign.org/
retrieval task was added in 2004 to explore domain–specific multilingual information retrieval and
also multi-modal retrieval by combining visual and textual features for retrieval. Since 2005, a
medical retrieval and a medical image annotation task were both part of ImageCLEF [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ].
      </p>
      <p>The enthusiastic participation in CLEF and particularly for ImageCLEF has shown the need
for benchmarks and their usefulness to the research community. Again in 2007, a total of 48 groups
registered for ImageCLEF to get access to the data sets and tasks. Among these, 13 participated
in the medical retrieval task and 10 in the medical automatic annotation task.</p>
      <p>
        Other important benchmarks in the field of visual information retrieval include TRECVID3
on the evaluation of video retrieval systems [
        <xref ref-type="bibr" rid="ref18">18</xref>
        ], ImagEval4, mainly on visual retrieval of images
and image classification, and INEX5 (INiative for the Evaluation of XML retrieval) concentrating
on retrieval of multimedia based on structured data. Close contact exists with these initiatives to
develop complementary evaluation strategies.
      </p>
      <p>
        This article focuses on the two medical tasks of ImageCLEF 2007, whereas two other papers
[
        <xref ref-type="bibr" rid="ref4 ref7">7, 4</xref>
        ] describe the new object classification task and the new photographic retrieval task. More
detailed information can also be found on the task web pages for ImageCLEFmed6 and the medical
annotation task7. A detailed analysis of the 2005 medical image retrieval task and its outcomes
is also available in [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ].
2
      </p>
    </sec>
    <sec id="sec-2">
      <title>The Medical Image Retrieval Task</title>
      <p>The medical image retrieval task has been run for four consecutive years. In 2007, two new
databases were added for a total of more than 66’000 images in the collection. For the generation
of realistic topics or information needs, log files of the medical literature search system Pubmed
were used.
2.1</p>
      <sec id="sec-2-1">
        <title>General Overview</title>
        <p>Again and as in previous years, the medical retrieval task showed to be popular among many
research groups registering for CLEF. In total 31 groups from all continents and 25 countries
registered. A total of 13 groups submitted 149 runs that were used for the pooling required for
the relevance judgments.
2.2</p>
      </sec>
      <sec id="sec-2-2">
        <title>Databases</title>
        <p>
          In 2007, the same four datasets were used as in 2005 and 2006 and two new datasets were added.
The Casimage8 dataset was made available to participants [
          <xref ref-type="bibr" rid="ref13">13</xref>
          ], containing almost 9’000 images
of 2’000 cases [
          <xref ref-type="bibr" rid="ref14">14</xref>
          ]. Images present in Casimage include mostly radiology modalities, but also
photographs, PowerPoint slides and illustrations. Cases are mainly in French, with around 20%
being in English and 5% without annotation. We also used the PEIR9 (Pathology Education
Instructional Resource) database with annotation based on the HEAL10 project (Health Education
Assets Library, mainly Pathology images [
          <xref ref-type="bibr" rid="ref1">1</xref>
          ]). This dataset contains over 33’000 images with
English annotations, with the annotation being on a per image and not a per case basis as in
Casimage. The nuclear medicine database of MIR, the Mallinkrodt Institute of Radiology11 [
          <xref ref-type="bibr" rid="ref22">22</xref>
          ],
was also made available to us for ImageCLEFmed. This dataset contains over 2’000 images mainly
from nuclear medicine with annotations provided per case and in English. Finally, the PathoPic12
3http://www-nlpir.nist.gov/projects/t01v/
4http://www.imageval.org/
5http://inex.is.informatik.uni-duisburg.de/2006/
6http://ir.ohsu.edu/image
7http://www-i6.informatik.rwth-aachen.de/~deselaers/imageclef07/medicalaat.html
8http://www.casimage.com/
9http://peir.path.uab.edu/
10http://www.healcentral.com/
11http://gamma.wustl.edu/home.html
12http://alf3.urz.unibas.ch/pathopic/intro.htm
collection (Pathology images [
          <xref ref-type="bibr" rid="ref6">6</xref>
          ]) was included into our dataset. It contains 9’000 images with
extensive annotation on a per image basis in German. A short part of the German annotation is
translated into English.
        </p>
        <p>In 2007, we added two new datasets. The first was the myPACS 13 dataset of 15’140 images and
3’577 cases, all in English and containing mainly radiology images. The second was the Clinical
Outcomes Research Initiative (CORI 14) Endoscopic image database contains 1’496 images with
an English annotation per image and not per case. This database extends the spectrum of the
total dataset as so far there were only few endoscopic images in the dataset. An overview of the
datasets can be seen in Table 1</p>
        <p>As such, we were able to use more than 66’000 images, with annotations in three different
languages. Through an agreement with the copyright holders, we were able to distribute these
images to the participating research groups. The myPACS database required an additional copyright
agreement making the process slightly more complex than in previous years.
2.3</p>
      </sec>
      <sec id="sec-2-3">
        <title>Registration and Participation</title>
        <p>In 2007, 31 groups from all 6 continents and 25 countries registered for the ImageCLEFmed
retrieval task, underlining the strong interest in this evaluation campaign. As in previous years,
only about half of the registered groups finally submitted results, often blaming a lack of time for
this. The feedback of these groups remains positive as they say to use the data for their research
as a very useful resource.</p>
        <p>The following groups finally also submitted results for the medical image retrieval task:
• CINDI group, Concordia University, Montreal, Canada;
• Dokuz Eylul University, Izmir, Turkey;
• IPAL/CNRS joint lab, Singapore, Singapore;
• IRIT–Toulouse, Toulouse, France;
• MedGIFT group, University and Hospitals of Geneva, Switzerland;
• Microsoft Research Asia, Beijing, China;
• MIRACLE, Spanish University Consortium, Madrid, Spain;
13http://www.mypacs.net/
14http://www.cori.org</p>
        <sec id="sec-2-3-1">
          <title>Ultrasound with rectangular sensor.</title>
          <p>Ultraschallbild mit rechteckigem Sensor.</p>
          <p>Ultrason avec capteur rectangulaire.
• MRIM–LIG, Grenoble, France;
• OHSU, Oregon Health &amp; Science University, Portland, OR, USA;
• RWTH Aachen Pattern Recognition group. Aachen, Germany;
• SINAI group, University of Jaen Intelligent Systems, Jaen, Spain;
• State University New York (SUNY) at Buffalo, NY, USA;
• UNAL group, Universidad Nacional Colombia, Bogot`a, Colombia;
In total, 149 runs were submitted, with the maximum being 36 of a single group and the minimum
a single run per group. Several runs had incorrect formats. These runs were corrected by the
organisers whenever possible but a few runs were finally omitted from the pooling process and
the final evaluation because trec eval could not parse the results even after our modifications. All
groups have the possibility to describe further runs in their working notes papers after the format
corrections as the qrels files were made available to all.
2.4</p>
        </sec>
      </sec>
      <sec id="sec-2-4">
        <title>Query Topics</title>
        <p>Query topics for 2007 were generated based on a log file of Pubmed15. The log file of 24 hours
contained a total of 77’895 queries. In general, the search terms were fairly vague and did not
contain many image–related topics, so we filtered out words such as image, video, and terms
relating to modalities such as x–ray, CT, MRI, endoscopy etc. We also aimed for the resulting
terms to cover at least two or more of the axes: modality, anatomic region, pathology, and visual
observation (e.g., enlarged heart).</p>
        <p>A total of 50 candidate topics were taken from these and sometimes an additional axis such as
modality was added. From these topics we checked whether at least a few relevant images are in
the database and once this was finished, 30 topics were selected.</p>
        <p>All topics were categorised with respect to the retrieval approach expected to perform best:
visual topics, textual (semantic) topics and mixed topics. This was performed by an experienced
image retrieval system developer. For each of the three retrieval approach groups, ten topics were
selected for a total of 30 query topics that were distributed among the participants. Each topic
consisted of the query itself in three languages (English, German, French) and 2–3 example images
for the visual part of the topic. Topic images were searched for on the Internet and were not part
of the database. This made visual retrieval significantly harder as most images were taken with
different collections compared to those in the database and had changes in the grey level or colour
values.</p>
        <p>Figure 1 shows a visual topic, Figure 2 a topic that should be retrieved well with a mixed
approach and Figure 3 a topics with very different images in the results sets that should be
well-suited for textual retrieval, only.</p>
        <sec id="sec-2-4-1">
          <title>Lung xray tuberculosis. R¨ontgenbild Lunge Tuberkulose. Radio pulmonal de tuberculose. Figure 2: Example for a mixed topic.</title>
        </sec>
        <sec id="sec-2-4-2">
          <title>Pulmonary embolism all modalities. Lungenembolie alle Modalita¨ten. Embolie pulmonaire, toutes les formes. Figure 3: Example for a semantic topic.</title>
          <p>Relevance judgments in ImageCLEFmed were performed by physicians and other studennts in the
OHSU biomedical informatics graduate program. All were paid an hourly rate for their work. The
pools for relevance judging were created by selecting the top ranking images from all submitted
runs. The actual number selected from each run has varied by year. In 2007, it was 35 images
per run, with the goal of having pools of about 800-1200 images in size for judging. The average
pool size in 2007 was 890 images. Judges were instructed to rate images in the pools are definitely
relevant (DR), partially relevant (PR), or not relevant (NR). Judges were instructed to use the
partially relevant desingation only in case they could not determine whether the image in question
was relevant.</p>
          <p>One of the problems was that all judges were English speakers but that the collection had a
fairly large number of French and German documents. If the judgment required reading the text,
judges had more difficulty ascertaining relevance. This could create a bias towards relevance for
documents with English annotation.
2.6</p>
        </sec>
      </sec>
      <sec id="sec-2-5">
        <title>Submissions and Techniques</title>
        <p>This section quickly summarises the main techniques used by the participants for retrieval and the
sort of runs that they submitted. We had for the first time several problems with the submissions
although we sent out a script to check runs for correctness before submission. In 2006, this script
was part of the submission web site, but performance problems had us change this setup. The
unit for retrieval and relevance was the image and not the case but several groups submitted case
IDs that we had to replace with the first image of the case. Other problems include the change of
upper/lower case for the image IDs and the change of the database names that also changed the
image IDs. Some groups reused the 2006 datasets that were corrected before 2007 and also ended
up with invalid IDs.</p>
        <p>The CINDI group submitted a total of 4 valid runs, two feedback runs and two automatic runs,
each time one with mixed media and a purely visual run. Text retrieval uses a simple tf/idf
weighting model and uses English, only. For visual retrieval a fusion model of a variety of features
and image representations is used. The mixed media run simply combine the two outcomes in a
linear fashion.</p>
        <p>Dokuz Eylul University submitted 5 runs, 4 visual and one textual run. The text runs is a
simple bag of words approach and for visual retrieval several strategies were used containing color
layout, color structure, dominant color and an edge histogram. Each run contained only one single
technique.
2.6.1
2.6.2</p>
        <p>DEU
2.6.3</p>
        <p>IPAL
2.6.7</p>
        <p>LIG
2.6.8</p>
        <p>OHSU
2.6.9</p>
        <p>RWTH
IPAL submitted 6 runs, all of them text retrieval runs. After having had the best performance
for two years, the results are now only in the middle of the performance scale.
2.6.4</p>
        <p>IRIT
The IRIT group submitted a single valid run, which was a text retrieval run.
2.6.5</p>
        <p>MedGIFT
The MedGIFT group submitted a total of 13 runs. For visual retrieval the GIFT (GNU Image
Finding Tool) was used to create a sort of baseline run, as this system had been used in the
same configuration since the beginning of ImageCLEF. Multilingual text retrieval was performed
with EasyIR and a mapping of the text in the three languages towards MeSH (Medical Subject
Headings) to search in semantic terms and avoid language problems.
2.6.6</p>
        <p>MIRACLE
MIRACLE submitted 36 runs in total and thus most runs of all groups. The text retrieval runs
were among the best, whereas visual retrieval was in the midfield. The combined runs were worse
than text alone and also only in the midfield.</p>
        <p>MRIM–LIG submitted 6 runs, all of them textual runs. Besides the best textual results, this was
also the best overall result in 2007.</p>
        <p>The OHSU group submitted 10 textual and mixed runs, using Fire as a visual system. Their
mixed runs had good performance as well as the best early precision.</p>
        <p>The Human language technology and pattern recognition group from the RWTH Aachen
University in Aachen, Germany submitted 10 runs using the FIRE image retrieval system. The runs are
based on a wide variety of 8 visual descriptors including image thumbnails, patch histograms, and
different texture features. For the runs using textual information, a text retrieval system is used
in the same way as in the last years. The weights for the features are trained with the maximum
entropy training method using the qrels of the 2005 and 2006 queries.</p>
        <p>The SINAI group submitted 30 runs in total, all of them textual or mixed. For text retrieval, the
terms of the query are mapped onto MeSH, and then, the query is expanded with these MeSH
terms.</p>
        <p>SUNY submitted 7 runs, all of which are mixed runs using Fire as visual system. One of the runs
is among the best mixed runs.</p>
        <p>The UNAL group submitted 8 runs, all of which are visual. The runs use a single visual feature,
only and range towards the lower end of the performance spectrum.</p>
        <p>The combination of runs from RWTH, OHSU, MedGIFT resulted in 13 submissions, all of which
were automatic and all used visual and textual information. The combinations were linear and
surprisingly the results are significantly worse than the results of single techniques of the
participants.
For the first time in 2007, the best overall system used only text for the retrieval. Up until now
the best systems always used a mix of visual and textual information. Nothing can really be said
on the outcome of manual and relevance feedback submissions as there were too few submitted
runs.</p>
        <p>It became clear that most research groups participating had a single specialty, usually either
visual or textual retrieval. By supplying visual and textual results as example, we gave groups
the possibility to work on multi-modal retrieval as well.
2.7.1</p>
        <p>Automatic Retrieval
As always, the vast majority of results were automatic and without any interaction. There were
146 runs in this category, with 27 visual runs, 80 mixed runs and 39 textual submissions, making
automatic mixed media runs the most popular category. The results shown in the following tables
are averaged over all 30 topics, thus hiding much information about which technique performed
well for what kind of tasks.</p>
        <p>Visual Retrieval Purely visual retrieval was performed in 27 runs and by six groups. Results
from GIFT and FIRE (Flexible Image Retrieval Engine) were made available for research groups
not having access to a visual retrieval engine themselves.</p>
        <p>To make the tables shorter and to not bias results shown towards groups with many
submissions, only the best two and the worst two runs of every group are shown in the results tables
of each category. Table 2 shows the results for the visual runs. Most runs had an extremely low
MAP (&lt;3% MAP), which had been the case during the previous years as well. The overall results
were lower than in preceding years, indiacting that tasks might have become harder. On the other
hand, two runs had good results and rivaled, at least for early precision, the best textual results.
These two runs actually used data from 2005 and 2006 that was somewhat similar to the tasks
in 2007 to train the system for optimal feature selection. This showed that an optimised feature
weighting may result in a large improvement!
textual retrieval A total of 39 submissions were purely textual and came from nine research
groups.</p>
        <p>Table 3 shows the best and worst two results of every group for purely textual retrieval. The
best overall runs were from LIG and were purely textual, which happened for the first time in
ImageCLEF. (LIG participated in ImageCLEF this year for the first time. Early precision (P10)
was only slightly better than the best purely visual runs and the best mixed runs had a very high
early precision whereas the highest P10 was actually a purely textual system where the MAP was
situated significantly lower. (Despite its name, MAP is more of a recall-oriented measure.)
mixed retrieval Mixed automatic retrieval had the highest number of submissions of all
categories. There were 80 runs submitted by 8 participating groups.</p>
        <p>Table 4 summarises the best two and the worst two mixed runs of every group. For some
groups the results for mixed runs were better than the best text runs but for others this was not
the case. This underlines the fact that combinations between visual and textual features have
to be done with care. Another interesting fact is that some systems with only a mediocre MAP
performed extremely well with respect to early precision.
2.8</p>
      </sec>
      <sec id="sec-2-6">
        <title>Manual and Interactive retrieval</title>
        <p>Only three runs in 2007 were in the manual or interactive sections, making any real comparison
impossible. Table 5 lists these runs and their performance</p>
        <p>Although information retrieval with relevance feedback or manual query modifications are seen
as a very important area to improve retrieval performance, research groups in ImageCLEF 2007
did not make use of these categories.
Visual retrieval without learning had very low results for MAP and even for early precision
(although with a smaller difference from text retrieval). Visual topics still perform well using visual
techniques. Extensive learning of feature selection and weighting can have enormous gain in
performance as shown by the FIRE runs.</p>
        <p>Purely textual runs had the best overall results for the first time and text retrieval was shown
to work well for most topics. Mixed–media runs were the most popular category and are often
better in performance than text or visual features alone. Still, in many cases the mixed media
runs did not perform as well as text alone, showing that care needs to be taken to combine media.</p>
        <p>Interactive and manual queries were almost absent from the evaluation and this remains an
important problem. ImageCLEFmed has to put these domains more into the focus of the researchers
although this requires more resources to perform the evaluation. System–oriented evaluation is
an important part but only interactive retrieval can show how well a system can really help the
users.</p>
        <p>With respect to performance measures, there was less correlation between the measures than
in previous years. The runs with the beast early precision (P10) were not close in MAP to the
best overall systems. This needs to be investigated as MAP is indeed a good indicator for overall
system performance but early precision might be much more what real users are looking for.
3</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>The Medical Automatic Annotation Task</title>
      <p>
        Over the last two years, automatic medical image annotation has been evolved from a simple
classification task with about 60 classes to a task with about 120 classes. From the very start
however, it was clear that the number of classes cannot be scaled indefinitely, and that the number
of classes that are desirable to be recognised in medical applications is far to big to assemble
sufficient training data to create suitable classifiers. To address this issue, a hierarchical class
structure such as the IRMA code [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ] can be a solution which allows to create a set of classifiers
for subproblems.
      </p>
      <p>The classes in the last years were based on the IRMA code where created by grouping similar
codes in one class. This year, the task has changed and the objective is to predict complete IRMA
codes instead of simple classes.</p>
      <p>This year’s medical automatic annotation task builds on top of last year: 1,000 new images
were collected and are used as test data, the training and the test data of last year was used as
training and development data respectively.
3.1</p>
      <sec id="sec-3-1">
        <title>Database &amp; Task Description</title>
        <p>The complete database consists of 12’000 fully classified medical radiographs taken randomly from
medical routine at the RWTH Aachen University Hospital. 10’000 of these were release together
with their classification as training data, another 1’000 were also published with their classification
as validation data to allow for tuning classifiers in a standardised manner. One thousand additional
images were released at a later date without classification as test data. These 1’000 images had
to be classified using the 11’000 images (10’000 training + 1’000 validation) as training data.</p>
        <p>Each of the 12’000 images is annotated with its complete IRMA code (see Sec. 3.1.1). In
total, 116 different IRMA codes occur in the database, the codes are not uniformly distributed,
but some codes have a significant larger share among the data than others. The least frequent
codes however, are represented at least 10 times in the training data to allow for learning suitable
models.</p>
        <p>Example images from the database together with textual labels and their complete code are
given in Figure 4.</p>
        <p>1121-120-200-700
T: x-ray, plain radiography, analog, overview image
D: coronal, anteroposterior (AP, coronal), unspecified
A: cranium, unspecified, unspecified
B: musculosceletal system, unspecified, unspecified</p>
        <p>T: x-ray, plain radiography, analog, overview image
D: coronal, anteroposterior (AP, coronal), unspecified
A: spine, cervical spine, unspecified</p>
        <p>
          B: musculosceletal system, unspecified, unspecified
1121-127-700-500
TD:: xco-rraoyn,apl,laainnterraodpioogstrearpihoyr,(aAnPa,locgo,roonvaerl)v,ieswupiimneage
A: abdomen, unspecified, unspecified
B: uropoietic system, unspecified, unspecified
1123-211-500-000
T: x-ray, plain radiography, analog, high beam energy
D: sagittal, lateral, right-left, inspiration
A: chest, unspecified, unspecified
B: unspecified, unspecified, unspecified
3.1.1 IRMA Code
Existing medical terminologies such as the MeSH thesaurus are poly-hierarchical, i.e., a code
entity can be reached over several paths. However, in the field of content-based image retrieval,
we frequently find class-subclass relations. The mono-hierarchical multi-axial IRMA code strictly
relies on such part-of hierarchies and, therefore, avoids ambiguities in textual classification [
          <xref ref-type="bibr" rid="ref9">9</xref>
          ].
In particular, the IRMA code is composed from four axes having three to four positions, each in
{0, . . . 9, a, . . . z}, where ”‘0”’ denotes ”‘not further specified”’. More precisely,
• the technical code (T) describes the imaging modality;
• the directional code (D) models body orientations;
• the anatomical code (A) refers to the body region examined; and
• the biological code (B) describes the biological system examined.
        </p>
        <p>This results in a string of 13 characters (IRMA: TTTT – DDD – AAA – BBB). For instance, the
body region (anatomy, three code positions) is defined as follows:
AAA
000 not further specified
...
400 upper extrimity (arm)
410 upper extrimity (arm); hand
411 upper extrimity (arm); hand; finger
412 upper extrimity (arm); hand; middle hand
413 upper extrimity (arm); hand; carpal bones
420 upper extrimity (arm); radio carpal joint
430 upper extrimity (arm); forearm
431 upper extrimity (arm); forearm; distal forearm
432 upper extrimity (arm); forearm; proximal forearm
440 upper extrimity (arm); ellbow
...</p>
        <p>The IRMA code can be easily extended by introducing characters in a certain code position,
e.g., if new imaging modalities are introduced. Based on the hierarchy, the more code position
differ from ”‘0”’, the more detailed is the description.
3.1.2</p>
        <p>Hierarchical Classification
To define a evaluation scheme for hierarchical classification, we can consider the 4 axes to be
independent, such that we can consider the axes independently and just sum up the errors for
each axis independently.</p>
        <p>
          Hierarchical classification is a well-known topic in different field. For example the classification
of documents often is done using a ontology based class hierarchy [
          <xref ref-type="bibr" rid="ref20">20</xref>
          ] and in information extraction
similar techniques are applied [
          <xref ref-type="bibr" rid="ref11">11</xref>
          ]. In our case, however we developed a novel evaluation scheme
to account for the particularities of the IRMA code which considers errors that are made early in
a hierarchy to be worse than errors that are made at a very fine level, and it is explicitly possible
to predict a code partially, i.e. to predict a code up to a certain position and put wild-cards for
the remaining positions, which is penalised but only with half the penalty a misclassification is
penalised.
        </p>
        <p>Our evaluation scheme is described in the following, where we only consider one axis. The
same scheme is applied to each axis individually.</p>
        <p>Let l1I = l1, l2, . . . , li, . . . , lI be the correct code (for one axis) of an image, i.e. if a classifier
predicts this code for an image, the classification is perfect. Further, let lˆ1I = lˆ1, lˆ2, . . . , lˆi, . . . , lˆI
be the predicted code (for one axis) of an image.</p>
        <p>The correct code is specified completely: li is specified for each position. The classifiers however,
are allowed to specify codes only up to a certain level, and predict “don’t know ” (encoded by *)
for the remaining levels of this axis.</p>
        <p>Given an incorrect classification at position lˆi we consider all succeeding decisions to be wrong
and given a not specified position, we consider all succeeding decisions to be not specified.</p>
        <p>We want to penalise wrong decisions that are easy (fewer possible choices at that node) over
wrong decisions that are difficult (many possible choices at that node), we can say, a decision at
position li is correct by chance with a probability of b1i if bi is the number of possible labels for
position i. This assumes equal priors for each class at each position.</p>
        <p>Furthermore, we want to penalise wrong decisions at an early stage in the code (higher up in
the hierarchy) over wrong decisions at a later stage in the code (lower down on the hierarchy) (i.e.
li is more important than li+1).</p>
        <p>Assembling the ideas from above in a straight forward way leads to the following equation:
with
where the parts of the equation account for</p>
        <p>XI 1 1 δ(li, lˆi)
i=1 |{bzi} |{iz} | ({cz) }</p>
        <p>(a) (b)
0 if lj = lˆj ∀j ≤ i
δ(li, lˆi) = 0.5 if lj = * ∃j ≤ i
1</p>
        <p>if lj 6= lˆj ∃j ≤ i
(a) accounts for difficulty of the decision at position i (branching factor)
(b) accounts for the level in the hierarchy (position in the string)
(c) correct/not specified/wrong, respectively.</p>
        <p>In addition, for every code, the maximal possible error is calculated and the errors are normed
such that a completely wrong decision (i.e. all positions wrong) gets an error count of 1.0 and a
completely correctly classified image has an error of 0.0.</p>
        <p>Table 7 shows examples for a correct code with different predicted codes. Predicting the
completely correct code leads to an error measure of 0.0, predicting all positions incorrectly leads
to an error measure of 1.0. The examples demonstrate that a classification error in a position at
the back of the code results in a lower error measure than a position in one of the first positions.
The last column of the table show the effect of the branching factor. In this column we assumed
the branching factor of the code is 2 in each node of the hierarchy. It can be observed that the
errors for the later positions have more weight compared to the real errors in the real hierarchy.
3.2</p>
      </sec>
      <sec id="sec-3-2">
        <title>Participating Groups &amp;</title>
      </sec>
      <sec id="sec-3-3">
        <title>Methods</title>
        <p>In the medical automatic annotation task, 29 groups registered of which 10 groups participated,
submitting a total of 68 runs. The group with the highest number of submissions had 30 runs in
total.</p>
        <p>In the following, groups are listed alphabetically and their methods are described shortly.
3.2.1</p>
        <p>
          BIOMOD: University of Liege, Belgium
The Bioinformatics and Modelling group from the University Liege16 in Belgium submitted four
runs. The approach is based on an object recognition framework using extremely randomised trees
and randomly extracted sub-windows [
          <xref ref-type="bibr" rid="ref10">10</xref>
          ]. The runs all use the same technique and differ how
the code is assembled. One run predicts the full code, one run predicts each axis independently
and the other two runs are combinations of the first ones.
3.2.2
        </p>
        <p>
          BLOOM: IDIAP, Switzerland
The Blanceflor-om2-toMed group from IDIAP in Martigny, Switzerland submitted 7 runs. All
runs use support vector machines (either in one-against-one or one-against-the-rest manner).
Features used are downscaled versions of the images, SIFT features extracted from sub-images, and
combinations of these [
          <xref ref-type="bibr" rid="ref21">21</xref>
          ].
        </p>
        <p>16http://www.montefiore.ulg.ac.be/services/stochastic/biomod
3.2.3</p>
        <p>
          Geneva: medGIFT Group, Switzerland
The medGIFT group17 from Geneva, Switzerland submitted 3 runs, each of the runs uses the
GIFT image retrieval system. The runs differ in the way, the IRMA-codes of the top-ranked
images are combined [
          <xref ref-type="bibr" rid="ref23">23</xref>
          ].
3.2.4
        </p>
        <p>CYU: Information Management AI lab, Taiwan
The Information Management AI lab from the Ching Yun University of Jung-Li, Taiwan submitted
one run using a nearest neighbour classifier using different global and local image features which
are particularly robust with respect to lighting changes.
3.2.5</p>
        <p>MIRACLE: Madrid, Spain
The Miracle group from Madrid Spain18 submitted 30 runs. The classification was done using a
10-nearest neighbour classifier and the features used are gray-value histograms, Tamura texture
features, global texture features, and Gabor features, which were extracted using FIRE. The runs
differ which features were used, how the prediction was done (predicting the full code, axis-wise
prediction, different subsets of axes jointly), and whether the features were normalised or not.
3.2.6</p>
        <p>Oregon Health State University, Portland, OR, USA
The Department of Medical Informatics and Clinical Epidemiology19 of the Oregon Health and
Science University in Portland, Oregon submitted two runs using neural networks and GIST
descriptors. One of the runs uses a support vector machine as a second level classifier to help
discriminating the two most difficult classes.
3.2.7</p>
        <p>
          RWTHi6: RWTH Aachen University, Aachen, Germany
The Human Language Technology and Pattern Recognition group20 of the RWTH Aachen
University in Aachen, Germany submitted 6 runs, all are based on sparse histograms of image patches
which were obtained by extracting patches at each position in the image. The histograms have
65536 or 4096 bins [
          <xref ref-type="bibr" rid="ref5">5</xref>
          ]. The runs differ in the resolution of the images. One run is a combination
of 4 normal runs, and one run does the classification axis-wise, the other runs, directly predict the
full code.
3.2.8
        </p>
        <p>IRMA: RWTH Aachen University, Medical Informatics, Aachen, Germany
The IRMA group from the RWTH Aachen University Hospital21, in Aachen Germany submitted
three baseline runs using weighted combinations of nearest neighbour classifiers using texture
histograms, image cross correlations, and the image deformation model. The parameters used are
exactly the same as used in previous years. The runs differ in the way in which the codes of the
five nearest neighbours are used to assemble the final predicted code.
3.2.9</p>
        <p>
          UFR: University of Freiburg, Computer Science Dep., Freiburg, Germany
The Pattern Recognition and Image Processing group from the University Freiburg22, Germany,
submitted four runs using relational features calculated around interest points which are later
combined to form cluster cooccurrence matrices [
          <xref ref-type="bibr" rid="ref17">17</xref>
          ]. Three different classification methods were
used: a flat classification scheme using all of the 116 classes , an axiswise-flat classification scheme
17http://www.sim.hcuge.ch/medgift/
18http://www.mat.upm.es/miracle/introduction.html
19http://www.ohsu.edu/dmice/
20http://www-i6.informatik.rwth-aachen.de
21http://www.irma-project.org
22http://lmb.informatik.uni-freiburg.de/
(i.e. 4 multi-class classifiers), and a binary classification tree (BCT) based scheme. The BCT based
approach is much faster to train and classify, but this comes at a slight performance penalty. The
tree was generated as described in [
          <xref ref-type="bibr" rid="ref16">16</xref>
          ].
        </p>
        <p>
          UNIBAS: University of Basel, Switzerland
The Databases and Information Systems group from the University Basel23, Switzerland submitted
14 runs using a pseudo two-dimensional hidden Markov model to model image deformation in the
images which were scaled down keeping the aspect ratio such that the longer side has a length of
32 pixels [
          <xref ref-type="bibr" rid="ref19">19</xref>
          ]. The runs differ in the features (pixels, Sobel features) that were used to determine
the deformation and in the k-parameter for the k-nearest neighbour.
3.3
        </p>
      </sec>
      <sec id="sec-3-4">
        <title>Results</title>
        <p>The results of the evaluation are given in Table 7. For each run, the run-id, the score as described
above and additionally, the error rate, which was used in the last years to evaluate the submissions
to this task are given.</p>
        <p>The method which had the best result last year is now at rank 8, which gives an impression
how much improvement in this field was achieved over the last year.</p>
        <p>Looking at the results for individual images, we noted, that only one image was classified
correctly by all submitted runs (top left image in Fig. 4). No image was misclassified by all runs.
3.4</p>
      </sec>
      <sec id="sec-3-5">
        <title>Discussion</title>
        <p>Analysing the results, it can be observed that the top-performing runs do not consider the
hierarchical structure of the given task, but rather use each individual code as one class and train a
116 classes classifier. This approach seems to work better given the currently limited amount of
codes, but obviously would not scale up infinitely and would probably lead to a very high demand
for appropriate training data if a much larger amount of classes is to be distinguished. The best
run using the code is on rank 6, builds on top of the other runs from the same group and uses the
hierarchy only in a second stage to combine the four runs.</p>
        <p>Furthermore, it can be seen that a method that is applied once accounting for the hierarchy/axis
structure of the code and once using the straight forward classification into 116 classes approach,
the one which does not know about the hierarchy clearly outperforms the other one (runs on ranks
11 and 13/7 and 14,16).</p>
        <p>Another clear observation is that methods using local image descriptors outperform methods
using global image descriptors. In particular, the top 16 runs are all using either local image
features alone or local image features in combination with a global descriptor.</p>
        <p>It is also observed that images where a large amount of training data is available are more far
more likely to be classified correctly.</p>
        <p>Considering the ranking wrt. to the applied hierarchical measure and the ranking wrt. to the
error rate it can clearly be seen that there are hardly any differences. Most of the differences
are clearly due to use of the code (mostly inserting of wildcard characters) which can lead to an
improvement for the hierarchical evaluation scheme, but will always lead to a deterioration wrt.
to the error rate.
3.5</p>
      </sec>
      <sec id="sec-3-6">
        <title>Conclusion</title>
        <p>The success of the medical automatic annotation task could be continued, the number of
participants is pretty constant, but a clear performance improvement of the best method could be
observed. Although only few groups actively tried to exploit the hierarchical class structure many
of the participants told us that they consider this an important research topic and that a further
investigation is desired.
0.5
0
10</p>
        <p>100
code frequency
1000</p>
        <p>Our goal for future tasks is to motivate more groups to participate and to increase the database
size such that it is necessary to use the hierarchical class structure actively.
4</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>Overall Conclusions</title>
      <p>The two medical tasks of ImageCLEF again attracted a very large number of registrations and
participation. This underlines the importance of such evaluation campaigns giving researchers the
opportunity to evaluate their systems without the tedious task of creating databases and topics.
In domains such as medical retrieval this is particularly important as data access if often difficult.</p>
      <p>In the medical retrieval task, visual retrieval without any learning only obtained good results
for a small subset of topics. With learning this can change strongly and deliver even for purely
visual retrieval fairly good results. Mixed–media retrieval was the most popular category and
results were often better for mixed–media than textual runs of the same groups. This shows that
mixed–media retrieval requires much work and more needs to be learned on such combinations.
Interactive retrieval and manual query modification were only used in 3 out of the 149 submitted
runs. This shows that research groups prefer submitting automatic runs , although interactive
retrieval is important and still must be addressed by researchers.</p>
      <p>For the annotation task, it was observed that techniques that rely heavily on recent
developments in machine learning and build on modern image descriptors clearly outperform other
methods. The class hierarchy that was provided could only lead to improvements for a few
groups. Overall, the runs that use the class hierarchy perform worse than those which consider
every unique code as a unique class which gives the impression that for the current number of 116
unique codes the training data is sufficient to train a joint classifier. As opposed to the retrieval
task, none of the groups used any interaction although this might allow for a big performance
gain.</p>
    </sec>
    <sec id="sec-5">
      <title>Acknowledgements</title>
      <p>We would like to thank the CLEF campaign for supporting the ImageCLEF initiative. We would
like to thank all the organizations who provided images and annotations for this year’s task,
including myPACS.net (Rex Jakobovits) and the OHSU CORI project (Judith Logan).</p>
      <p>This work was partially funded by the DFG (Deutsche Forschungsgemeinschaft) under
contracts Ne-572/6 and Le-1108/4, the Swiss National Science Foundation (FNS) under contract
205321-109304/1, the American National Science Foundation (NSF) with grant ITR–0325160,
and the EU Sixth Framework Program with the SemanticMining project (IST NoE 507505) and
the MUSCLE NoE.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>Chris. S.</given-names>
            <surname>Candler</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Sebastian. H.</given-names>
            <surname>Uijtdehaage</surname>
          </string-name>
          ,
          <article-title>and</article-title>
          <string-name>
            <given-names>Sharon. E.</given-names>
            <surname>Dennis. Introducing</surname>
          </string-name>
          <string-name>
            <surname>HEAL</surname>
          </string-name>
          :
          <article-title>The health education assets library</article-title>
          .
          <source>Academic Medicine</source>
          ,
          <volume>78</volume>
          (
          <issue>3</issue>
          ):
          <fpage>249</fpage>
          -
          <lpage>253</lpage>
          ,
          <year>2003</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>Paul</given-names>
            <surname>Clough</surname>
          </string-name>
          , Henning Mu¨ller, and
          <string-name>
            <given-names>Mark</given-names>
            <surname>Sanderson</surname>
          </string-name>
          .
          <article-title>The CLEF 2004 cross language image retrieval track</article-title>
          . In C. Peters,
          <string-name>
            <given-names>P.</given-names>
            <surname>Clough</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Gonzalo</surname>
          </string-name>
          , G. Jones,
          <string-name>
            <given-names>M.</given-names>
            <surname>Kluck</surname>
          </string-name>
          , and B. Magnini, editors,
          <source>Multilingual Information Access for Text</source>
          ,
          <article-title>Speech and Images: Results of the Fifth CLEF Evaluation Campaign</article-title>
          , pages
          <fpage>597</fpage>
          -
          <lpage>613</lpage>
          . Lecture Notes in Computer Science (LNCS), Springer, Volume
          <volume>3491</volume>
          ,
          <year>2005</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>Paul</given-names>
            <surname>Clough</surname>
          </string-name>
          , Henning Mu¨ller, and
          <string-name>
            <given-names>Mark</given-names>
            <surname>Sanderson</surname>
          </string-name>
          .
          <article-title>Overview of the CLEF cross-language image retrieval track (ImageCLEF) 2004</article-title>
          . In Carol Peters,
          <string-name>
            <given-names>Paul D.</given-names>
            <surname>Clough</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Gareth J. F.</given-names>
            <surname>Jones</surname>
          </string-name>
          , Julio Gonzalo,
          <string-name>
            <given-names>M.</given-names>
            <surname>Kluck</surname>
          </string-name>
          , and B. Magnini, editors,
          <source>Multilingual Information Access for Text</source>
          ,
          <article-title>Speech and Images: Result of the fifth CLEF evaluation campaign</article-title>
          , Lecture Notes in Computer Science, Bath, England,
          <year>2005</year>
          . Springer-Verlag.
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>Thomas</given-names>
            <surname>Deselaers</surname>
          </string-name>
          , Allan Hanbury, and et al.
          <article-title>Overview of the ImageCLEF 2007 object retrieval task</article-title>
          .
          <source>In Working Notes of the 2007 CLEF Workshop</source>
          , Budapest, Hungary,
          <year>September 2007</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>Thomas</given-names>
            <surname>Deselaers</surname>
          </string-name>
          , Andre Hegerath, Daniel Keysers, and Hermann Ney.
          <article-title>Sparse patchhistograms for object classification in cluttered images</article-title>
          .
          <source>In DAGM</source>
          <year>2006</year>
          ,
          <string-name>
            <given-names>Pattern</given-names>
            <surname>Recognition</surname>
          </string-name>
          ,
          <source>26th DAGM Symposium</source>
          , volume
          <volume>4174</volume>
          of Lecture Notes in Computer Science, pages
          <fpage>202</fpage>
          -
          <lpage>211</lpage>
          , Berlin, Germany,
          <year>September 2006</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>K.</given-names>
            <surname>Glatz-Krieger</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Glatz</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Gysel</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Dittler</surname>
          </string-name>
          , and
          <string-name>
            <given-names>M. J.</given-names>
            <surname>Mihatsch</surname>
          </string-name>
          .
          <article-title>Webbasierte Lernwerkzeuge fu¨r die Pathologie - web-based learning tools for pathology</article-title>
          .
          <source>Pathologe</source>
          ,
          <volume>24</volume>
          :
          <fpage>394</fpage>
          -
          <lpage>399</lpage>
          ,
          <year>2003</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>Michael</given-names>
            <surname>Grubinger</surname>
          </string-name>
          , Paul Clough, Allan Hanbury, and
          <article-title>Henning Mu¨ller. Overview of the ImageCLEF 2007 photographic retrieval task</article-title>
          .
          <source>In Working Notes of the 2007 CLEF Workshop</source>
          , Budapest, Hungary,
          <year>September 2007</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>William</given-names>
            <surname>Hersh</surname>
          </string-name>
          , Henning Mu¨ller, Jeffery Jensen,
          <string-name>
            <given-names>Jianji</given-names>
            <surname>Yang</surname>
          </string-name>
          , Paul Gorman, and
          <string-name>
            <given-names>Patrick</given-names>
            <surname>Ruch</surname>
          </string-name>
          .
          <article-title>Imageclefmed: A text collection to advance biomedical image retrieval</article-title>
          .
          <source>Journal of the American Medical Informatics Association</source>
          , September/October,
          <year>2006</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <surname>Thomas</surname>
            <given-names>M.</given-names>
          </string-name>
          <string-name>
            <surname>Lehmann</surname>
          </string-name>
          , Henning Schubert, Daniel Keysers, Michael Kohnen, and
          <string-name>
            <surname>Bertold</surname>
            <given-names>B.</given-names>
          </string-name>
          <string-name>
            <surname>Wein</surname>
          </string-name>
          .
          <article-title>The IRMA code for unique classification of medical images</article-title>
          .
          <source>In SPIE</source>
          <year>2003</year>
          , volume
          <volume>5033</volume>
          , pages
          <fpage>440</fpage>
          -
          <lpage>451</lpage>
          ,
          <year>2003</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <surname>Rapha</surname>
          </string-name>
          <article-title>¨el Mar´ee, Pierre Geurts</article-title>
          , Justus Piater, and
          <string-name>
            <given-names>Louis</given-names>
            <surname>Wehenkel</surname>
          </string-name>
          .
          <article-title>Random subwindows for robust image classification</article-title>
          .
          <source>In Cordelia Schmid</source>
          , Stefano Soatto, and Carlo Tomasi, editors,
          <source>Proceedings of the IEEE International Conference on Computer Vision and Pattern Recognition (CVPR</source>
          <year>2005</year>
          ), volume
          <volume>1</volume>
          , pages
          <fpage>34</fpage>
          -
          <lpage>40</lpage>
          . IEEE,
          <year>June 2005</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <surname>Diana</surname>
            <given-names>Maynard</given-names>
          </string-name>
          , Wim Peters, and
          <string-name>
            <given-names>Yaoyong</given-names>
            <surname>Li</surname>
          </string-name>
          .
          <article-title>Metrics for evaluation of ontology-based information extraction. In Evaluation of Ontologies for the Web (EON</article-title>
          <year>2006</year>
          ), Edinburgh, UK,
          <year>2006</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12] Henning Mu¨ller, Thomas Deselaers,
          <string-name>
            <surname>Thomas</surname>
            <given-names>M.</given-names>
          </string-name>
          <string-name>
            <surname>Lehmann</surname>
            , Paul Clough, and
            <given-names>William</given-names>
          </string-name>
          <string-name>
            <surname>Hersh</surname>
          </string-name>
          .
          <article-title>Overview of the imageclefmed 2006 medical retrieval and annotation tasks</article-title>
          .
          <source>In CLEF working notes</source>
          , Alicante, Spain, Sep.
          <year>2006</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13] Henning Mu¨ller, Antoine Rosset, Jean-Paul Vall´ee, Francois Terrier, and
          <string-name>
            <given-names>Antoine</given-names>
            <surname>Geissbuhler</surname>
          </string-name>
          .
          <article-title>A reference data set for the evaluation of medical image retrieval systems</article-title>
          .
          <source>Computerized Medical Imaging and Graphics</source>
          ,
          <volume>28</volume>
          :
          <fpage>295</fpage>
          -
          <lpage>305</lpage>
          ,
          <year>2004</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <surname>Antoine</surname>
            <given-names>Rosset</given-names>
          </string-name>
          , Henning Mu¨ller, Martina Martins, Natalia Dfouni, Jean-Paul Vall´ee, and Osman Ratib.
          <article-title>Casimage project - a digital teaching files authoring environment</article-title>
          .
          <source>Journal of Thoracic Imaging</source>
          ,
          <volume>19</volume>
          (
          <issue>2</issue>
          ):
          <fpage>1</fpage>
          -
          <lpage>6</lpage>
          ,
          <year>2004</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <given-names>Jacques</given-names>
            <surname>Savoy</surname>
          </string-name>
          .
          <source>Report on CLEF-2001 experiments. In Report on the CLEF Conference 2001 (Cross Language Evaluation Forum)</source>
          , pages
          <fpage>27</fpage>
          -
          <lpage>43</lpage>
          , Darmstadt, Germany,
          <year>2002</year>
          . Springer LNCS 2406.
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [16]
          <string-name>
            <given-names>Lokesh</given-names>
            <surname>Setia</surname>
          </string-name>
          and
          <string-name>
            <given-names>Hans</given-names>
            <surname>Burkhardt</surname>
          </string-name>
          .
          <article-title>Learning taxonomies in large image databases</article-title>
          .
          <source>In ACM SIGIR Workshop on Multimedia Information Retrieval</source>
          , Amsterdam, Holland,
          <year>2007</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [17]
          <string-name>
            <surname>Lokesh</surname>
            <given-names>Setia</given-names>
          </string-name>
          , Alexandra Teynor, Alaa Halawani, and
          <string-name>
            <given-names>Hans</given-names>
            <surname>Burkhardt</surname>
          </string-name>
          .
          <article-title>Image classification using cluster-cooccurrence matrices of local relational features</article-title>
          .
          <source>In Proceedings of the 8th ACM International Workshop on Multimedia Information Retrieval</source>
          , Santa Barbara, CA, USA,
          <year>2006</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          [18]
          <string-name>
            <surname>Alan</surname>
            <given-names>F.</given-names>
          </string-name>
          <string-name>
            <surname>Smeaton</surname>
            , Paul Over, and
            <given-names>W.</given-names>
          </string-name>
          <string-name>
            <surname>Kraaij</surname>
          </string-name>
          .
          <article-title>TRECVID: Evaluating the effectiveness of information retrieval tasks on digital video</article-title>
          .
          <source>In Proceedings of the international ACM conference on Multimedia 2004 (ACM MM</source>
          <year>2004</year>
          ), pages
          <fpage>652</fpage>
          -
          <lpage>655</lpage>
          , New York City, NY, USA,
          <year>October 2004</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          [19]
          <string-name>
            <given-names>Michael</given-names>
            <surname>Springmann</surname>
          </string-name>
          , Andreas Dander, and
          <string-name>
            <given-names>Heiko</given-names>
            <surname>Schuldt</surname>
          </string-name>
          .
          <source>T.b.a. In Working Notes of the 2007 CLEF Workshop</source>
          , Budapest, Hungary,
          <year>September 2007</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          [20]
          <string-name>
            <given-names>Aixin</given-names>
            <surname>Sun</surname>
          </string-name>
          and
          <string-name>
            <surname>Ee-Peng Lim</surname>
          </string-name>
          .
          <article-title>Hierarchical text classification and evaluation</article-title>
          .
          <source>In IEEE International Conference on Data Mining (ICDM</source>
          <year>2001</year>
          ), pages
          <fpage>521</fpage>
          -
          <lpage>528</lpage>
          , San Jose, CA, USA,
          <year>November 2001</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          [21]
          <string-name>
            <surname>Tatiana</surname>
            <given-names>Tommasi</given-names>
          </string-name>
          , Francesco Orabona, and Barbara Caputo.
          <article-title>CLEF2007 Image Annotation Task: an SVM-based Cue Integration Approach</article-title>
          .
          <source>In Working Notes of the 2007 CLEF Workshop</source>
          , Budapest, Hungary,
          <year>September 2007</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          [22]
          <string-name>
            <surname>Jerold</surname>
          </string-name>
          . W. Wallis,
          <string-name>
            <surname>Michelle. M. Miller</surname>
            ,
            <given-names>Tom. R.</given-names>
          </string-name>
          <string-name>
            <surname>Miller</surname>
          </string-name>
          ,
          <article-title>and</article-title>
          <string-name>
            <given-names>Thomas. H.</given-names>
            <surname>Vreeland</surname>
          </string-name>
          .
          <article-title>An internetbased nuclear medicine teaching file</article-title>
          .
          <source>Journal of Nuclear Medicine</source>
          ,
          <volume>36</volume>
          (
          <issue>8</issue>
          ):
          <fpage>1520</fpage>
          -
          <lpage>1527</lpage>
          ,
          <year>1995</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          [23]
          <string-name>
            <surname>Xin</surname>
            <given-names>Zhou</given-names>
          </string-name>
          , Julien Gobeill, Patrick Ruch, and Henning Mu¨ller. University and Hospitals of Geneva at ImageCLEF 2007.
          <source>In Working Notes of the 2007 CLEF Workshop</source>
          , Budapest, Hungary,
          <year>September 2007</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>