<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Overview of the LifeCLEF 2014 Fish Task</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Concetto Spampinato</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Simone Palazzo</string-name>
          <email>simone.palazzog@dieei.unict.it</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Bas Boom</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Robert B. Fisher</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Department of Electrical, Electronics and Computer Engineering University of Catania</institution>
          ,
          <country country="IT">Italy</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>School of Informatics, University of Edinburgh</institution>
          ,
          <country country="UK">UK</country>
        </aff>
      </contrib-group>
      <fpage>616</fpage>
      <lpage>624</lpage>
      <abstract>
        <p>This paper describes the LifeCLEF 2014 sh task, which aimed at benchmarking automatic sh detection and recognition methods by processing underwater visual data. The task consisted of videobased subtasks for sh detection and sh species recognition in videos and one image-based task for sh species classi cation in still images. Our underwater visual datasets consisted of about 2,000 videos taken from the Fish4Knowledge video repository and more than 200,000 annotations automatically obtained and manually validated. About fty teams registered to the sh task, but only two teams submitted runs: the I3S team for subtask 3 and LSIS/DYNI team for subtask 4. The results achieved by both teams are satisfactory.</p>
      </abstract>
      <kwd-group>
        <kwd>Underwater video analysis</kwd>
        <kwd>Object classi cation</kwd>
        <kwd>Fine-grained visual categorisation</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>Underwater video and imaging systems are used increasingly in a range of
monitoring or exploratory applications, in particular for biological (e.g. benthic
community structure, habitat classi cation), sheries (e.g. stock assessment, species
richness), geological (e.g. seabed type, mineral deposits) and physical surveys
(e.g. pipelines, cables, oil industry infrastructure). Their usage has bene tted
from the increasing miniaturisation and cost-e ectiveness of submersible ROVs
(remotely operated vehicles) and advances in underwater digital cameras. These
technologies have revolutionised our ability to capture high-resolution images in
challenging aquatic environments and are also greatly improving our ability to
e ectively manage natural resources, increasing our competitiveness and
reducing operational risks in industries that operate in both marine and freshwater
systems. Despite these advances in data collection technologies, the analysis of
video data usually requires very time-consuming and expensive input by human
observers. This is particularly true for ecological and shery video data, which
often requires laborious visual analysis. This analytical \bottleneck" greatly
restricts the use of these otherwise powerful video technologies and demands
effective methods for automatic content analysis to enable proactive provision of
analytical information.</p>
      <p>
        The development of automatic video analysis tools is, however, particularly
challenging because of the complexities of underwater video recordings in terms
of the variability of scenarios and factors that may degrade the video quality such
as water clarity and/or depth. A recent advance in this direction was developed
as part of the Fish4Knowledge (F4K)3 project (funded by The European Union
7th Framework Programme), where computer vision and machine learning
techniques were developed to extract information about sh density and richness
from videos taken by underwater cameras installed at coral reefs in Taiwan [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ].
      </p>
      <p>However, the underwater is a rather complex environment because of several
factors that make the study of it particularly complex and challenging. In
particular, dynamic or multi-modal backgrounds, abrupt lighting changes (due also to
water caustics), and radical and instant water turbidity changes a ect the ability
to perform visual tasks also for humans. To complicate even more the situation,
the underwater environment shows two almost exclusive characteristics with
respect to other domains: three degrees of freedom and erratic and extremely fast
movements of objects (i.e. sh). This makes sh less predictable than people or
vehicles, as sh may move in all three directions changing suddenly their size and
their shape in the video. As a consequence of this, although in the F4K project
reliable approaches for video-based sh detection and species identi cation were
devised, the problem of automatic analysis of underwater visual data remains
still open.</p>
      <p>In this paper we describe the \Fish Task" organised as part of the LifeCLEF
2014, where video and image based approaches for sh detection and recognition
were tested on underwater video/image datasets achieving very good results.</p>
      <p>The remainder of the paper is as follows: Sect. 2 provides an overview of the
task as well as the underwater video and image datasets used, Sect. 3 presents
the participants to the sh task and their approaches whose results, compared
to our baselines, are given in Sect. 4. Sect. 5 concludes this report.
2</p>
    </sec>
    <sec id="sec-2">
      <title>Fish Task Description</title>
      <p>The LifeCLEF 2014 Fish task aimed at benchmarking automatic sh detection
and recognition methods by processing underwater visual data. It basically
consisted of three video-based subtasks and one image-based task. The video-based
subtasks were:</p>
      <p>Subtask 1 { detecting moving objects in videos by either background
modeling or object detection methods;
Subtask 2 { detecting sh instances in video frames, thus discriminating
sh instances from non- sh ones;
Subtask 3 { detecting sh instances in video frames and recognising their
species;</p>
      <p>The image-based subtask, instead, had the goal:
3 www. sh4knowledge.eu
Subtask 4 { to identify sh species using only still images containing only
one sh instance.</p>
      <p>The participants had to submit at most three runs for a subtask. The run le
must have had the same format as the ground truth xml le (see below), i.e. it
must contain the frame where the sh was detected together with the bounding
box (for all the subtasks 1, 2, and 3), contours (only task 1) and species name
(for subtasks 3 and 4) of the sh.
2.1</p>
      <p>
        Dataset
The underwater video dataset was derived from the Fish4Knowledge video
repository, which contains about 700,000 10-minute video clips that were taken in the
past ve years to monitor Taiwan coral reefs. The Taiwan area is particularly
interesting for studying the marine ecosystem, as it holds one of the largest
sh biodiversities of the world with more than 3,000 di erent sh species whose
taxonomy is available at http:// shdb.sinica.edu.tw. The dataset contains videos
recorded from sunrise to sunset showing several phenomena, e.g. murky water,
algae on camera lens, etc., which makes the sh identi cation task more
complex. Each video has a resolution of either 320x240 or 640x480 with 5 to 8 fps.
As the LifeCLEF 2014 sh task included four subtasks, we employed di erent
datasets for each subtask. In particular, for subtask 1 we used eight videos (four
for training and four for testing) fully labeled (each single sh instance was
annotated) using the tool in [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ] resulting in 21,106 annotations corresponding to 9,852
di erent sh for the training dataset and in 14,829 annotations corresponding
to 899 sh for the test set. For the other three subtasks, we used the (partially
overlapping) following datasets:
{ The D20M dataset, which contains about 20 million underwater images
randomly selected from the whole F4K dataset.
{ The D35K dataset, which is a subset of D20M containing about 35000 sh
images belonging to 10 sh species. Each image of this dataset was manually
annotated with the corresponding species.
{ The D1M dataset, which contains about 1 million automatically-annotated
images by using the method in [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ] and further manually checked. This dataset
is meant to be the subset of the images in D20M which belong to the classes
annotated in D35K.
      </p>
      <p>The D20M dataset represents a randomly-selected fraction of the data
collected within the Fish4Knowledge project, amounting to more than a billion sh
images. The D35K dataset is a subset of the D20M one, containing only (but
not all) sh images belonging to the 10 most common species. Image annotation
was carried out manually and validated by expert marine biologists. Figure 1
shows sample pictures of the chosen 10 species.</p>
      <p>
        The D1M dataset is also a subset of D20M, obtained by the automatic
annotation technique described in [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ] and further manually checked: brie y, images
from D35K are used as queries to a similarity-based search in D20M ; after a
check for false positives, the resulting images are then assigned to the same
species label as the query image. We did not use directly the D35K dataset sh
species classi cation, because it contained several near duplicate images.
      </p>
      <p>The datasets used in the LifeCLEF 2014 sh task (subtasks 2, 3 and 4)
were generated from the D1M dataset. In particular for subtask 2, which was
meant to identify only some sh instances (unlike the subtask 1 which asked
for identifying all sh in a video) in video frames, we used 112,078 annotations
taken from 957 videos for the training set and 15,245 annotations from 89 videos
for the test set.</p>
      <p>For subtasks 3 and 4, we, instead, used 24,441 (belonging to 285 videos)
annotated sh (and their species) in the training set and 6,956 annotations
(belonging to 116 videos) in the test set. Details about the species distribution
in the training and test datasets are shown in Table 1.</p>
      <p>It is important to note that some species were more common than others;
although we tried to make the dataset as uniformly distributed across species as
possible, for some of them (most evidently, Lutjanus fulvus ) it was quite di cult
to nd a large number of adequate images, which resulted in a lower presence in
the dataset.</p>
      <p>The datasets (both the training and the test ones) were provided as zipped
folders containing:
{ A ground truth folder where the ground truth for all subtasks were given as</p>
      <p>XML les (Fig. 2):
{ A video folder where all the videos of the entire dataset were contained
{ A image folder which contains the images corresponding to the bounding
boxes given in the XML les. These images were provided only for subtasks
2, 3, 4 in the training phase and only for subtask 4 in the test phase.
About 50 teams registered to the sh task, but only two teams submitted runs
for the sh task: one, the I3S team, for subtask 3 and one, the LSIS/DYNI team
for subtask 4.</p>
      <p>The strategy employed by the I3S team for sh identi cation and
recognition (subtask 3) consisted of, rst, applying a background modeling approach
based on Mixture of Gaussian for moving object segmentation. SVM learning
using keyframes of species as positive entries and background of current video
as negative entries was used for sh species classi cation.</p>
      <p>
        The LSIS/DYNI team submitted three runs for subtask 4. Each run followed
the strategy proposed in [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ] which, basically, consisted of extracting low level
features, patch encoding, pooling with spatial pyramid for local analysis and
a linear large-scale supervised classication by averaging posterior probabilities
estimated through linear regression of linear SVM's outputs. No image speci c
pre-processing regarding illumination correction or background subtraction was
performed.
      </p>
    </sec>
    <sec id="sec-3">
      <title>Results</title>
      <p>
        Performance evaluation was carried out on the released test sets by computing:
average precision and recall, and precision and recall for each sh species for
subtask 3; average recall and recall for each sh species for subtask 4. The
baselines for subtasks 3 and 4 were, respectively:
{ Subtask 3: The ViBe [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ] background modeling approach for sh detection
combined to VLFeat BoW [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ] for sh species recognition
{ Subtask 4: VLFeat BoW for sh species recognition
      </p>
      <p>The results achieved by the I3S team for subtask 3 are reported in Fig. 3
and compared to our baseline. When computing the performance in terms of
sh detection, we relaxed the constraint on the PASCAL score as the I3S team
often detected correctly a sh but the bounding box' size was much bigger than
the one provided in our dataset (see Fig. 4).</p>
      <p>While the average recall obtained by the I3S team was lower than the
baseline's recall, the precision was improved (see Fig. 5), thus implying that their
sh species classi cation approach was reliable more than the sh detection
approach. Furthermore, we allowed participants to provide a ranked list (top three
species) of sh species for each sh instance in the test set and we treated a
recognition as a true positive if the correct species was in the top three species.
Considering only the most probable class provided in the results submitted by
the I3S team makes the performance drop considerably. The reason behind this
may be found in the size of the bounding boxes extracted by the I3S team. In
fact bigger bounding boxes may contain also background objects and/or other
sh instances, thus a ecting the performance of the sh species classi cation
approach.</p>
      <p>The results obtained by the LSIS/DYNI team for subtask 4 are shown in Fig.
6. In this case, the performance were computed by taking into account only the
most probable sh species (i.e. the rst class in the provided ranked list).</p>
      <p>Please note the image-based recognition task (subtask 4) was easier than
subtask 3 since it does need any sh identi cation module (which is the most
complex part in video-based sh identi cation) and we had only ten sh species
with very distinctive features. However, LSIS/DYNI team did a great job
outperforming our baseline.
5</p>
    </sec>
    <sec id="sec-4">
      <title>Concluding remarks</title>
      <p>In this report, we described the LifeCLEF 2014 sh task, which aimed at
benchmarking machine learning and computer vision methods for sh detection and
recognition in underwater \real-life" video footage. Although the fty teams
registered for the sh task, only two teams submitted runs: the I3S team for subtask
3 and LSIS/DYNI team for subtask 4. The main challenge of our underwater
dataset was to identify sh correctly (as moving objects) by processing the entire
video sequences. This appears evident by looking the results achieved by the I3S
team (Fig. 3 and 5): for instance, the I3S team was not able to detect any of
the Lutjanus fulvus instances, as it is a rather tiny sh which often tends to
hide behind rocks. The subtask 4 was easier (as shown by the results achieved
by our baseline) as the sh were already segmented from the background and
the considered sh species have very distinctive features that make their
identi cation in the feature space almost trivial. Nevertheless, the results achieved
by the LSIS/DYNI team are excellent, outperforming our baseline for almost all
species.</p>
      <p>We have just completed the labelling of other 25 species and we are working
on removing the near duplicates from the dataset to make the sh identi cation
and recognition task more challenging.</p>
      <p>We would like to thank to all teams who participated to the task and the
ImageCLEF and LifeCLEF organisation who made the rst edition of sh task
possible.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Boom</surname>
            ,
            <given-names>B.J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>He</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Palazzo</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Huang</surname>
            ,
            <given-names>P.X.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Beyan</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chou</surname>
            ,
            <given-names>H.M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lin</surname>
            ,
            <given-names>F.P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Spampinato</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Fisher</surname>
          </string-name>
          , R.B.:
          <article-title>A research tool for long-term and continuous analysis of sh assemblage in coral-reefs using underwater camera footage</article-title>
          . Ecological Informatics http://dx.doi.org/10.1016/j.ecoinf.
          <year>2013</year>
          .
          <volume>10</volume>
          .006 (
          <year>2013</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Kavasidis</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Palazzo</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Salvo</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Giordano</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Spampinato</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          :
          <article-title>An innovative web-based collaborative platform for video annotation</article-title>
          .
          <source>Multimedia Tools and Applications</source>
          <volume>70</volume>
          (
          <issue>1</issue>
          ) (
          <year>2014</year>
          )
          <volume>413</volume>
          {
          <fpage>432</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Giordano</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Palazzo</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Spampinato</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          :
          <article-title>Nonparametric label propagation using mutual local similarity in nearest neighbors</article-title>
          . To appear on Computer Vision and Image
          <string-name>
            <surname>Understanding</surname>
          </string-name>
          (
          <year>2014</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4. Paris, S.,
          <string-name>
            <surname>Halkias</surname>
            ,
            <given-names>X.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Glotin</surname>
          </string-name>
          , H.:
          <article-title>Sparse coding for histograms of local binary patterns applied for image categorization: Toward a bag-of-scenes analysis</article-title>
          .
          <source>In: 2012 21st International Conference on Pattern Recognition (ICPR)</source>
          .
          <source>(Nov</source>
          <year>2012</year>
          )
          <volume>2817</volume>
          {
          <fpage>2820</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Barnich</surname>
            ,
            <given-names>O.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Van Droogenbroeck</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          :
          <article-title>Vibe: A universal background subtraction algorithm for video sequences</article-title>
          .
          <source>Image Processing, IEEE Transactions on 20(6)</source>
          (
          <year>June 2011</year>
          )
          <volume>1709</volume>
          {
          <fpage>1724</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Vedaldi</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Fulkerson</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          :
          <article-title>VLFeat - an open and portable library of computer vision algorithms</article-title>
          . In: ACM International Conference on Multimedia. (
          <year>2010</year>
          )
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>