<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Unsupervised individual whales identi cation: spot the di erence in the ocean</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Alexis Joly</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Jean-Christophe Lombardo</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Julien Champ</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Anjara Saloma</string-name>
          <email>anjara@cetamada.com</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Cetamada</institution>
          ,
          <addr-line>NGO</addr-line>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Inria ZENITH team</institution>
          ,
          <addr-line>LIRMM</addr-line>
          ,
          <country country="FR">France</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>Identifying organisms is a key step in accessing information related to the ecology of species. But unfortunately, this is di cult to achieve due to the level of expertise necessary to correctly identify and record living organisms. To try bridging this gap, enormous work has been done on the development of automated species identi cation tools such as image-based plant identi cation or audio recordings-based bird identi cation. Yet, for some groups, it is preferable to monitor the organisms at the individual level rather than at the species level. The automatizing of this problem has received much less attention than species identi cation. In this paper, we address the speci c scenario of discovering humpack whales individuals in a large collections of pictures collected by nature observers. The process is initiated from scratch, without any knowledge on the number of individuals and without any training samples of these individuals. Thus, the problem is entirely unsupervised. To address it, we set up and experimented a scalable ne-grained matching system allowing to discover small rigid visual patterns in highly clutter background. The evaluation was conducted in blind in the context of the LifeCLEF evaluation campaign. Results show that the proposed system provides very promising results with regard to the di culty of the task but that there is still room for improvements to reach higher recall and precision in the future.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>
        Identifying organisms is a key step in accessing information related to the
ecology of species. This is an essential step in recording any specimen on earth to
be used in ecological studies. But unfortunately, this is di cult to achieve due
to the level of expertise necessary to correctly identify and record living
organisms. Watson et al.[
        <xref ref-type="bibr" rid="ref6">6</xref>
        ] discussed in 2004 the potential of automated species
identi cation approaches typically based on machine learning and multimedia
data analysis methods. They suggested that, if the scienti c community is able
to (i) overcome the production of large training datasets, (ii) more precisely
identify and evaluate the error rates, (iii) scale up automated approaches, and
(iv) detect novel species, it will then be possible to initiate the development of a
generic automated species identi cation system that could open up vistas of new
opportunities for pure and applied work in biological and related elds. Since
the question raised by Watson in 2004 ("automated species identi cation: why
not?"), enormous work has been done on the development of e ective methods
such as image-based plant identi cation [
        <xref ref-type="bibr" rid="ref2 ref7 ref8">2,8,7</xref>
        ], bird songs identi cation [
        <xref ref-type="bibr" rid="ref22">22</xref>
        ], sh
species identi cation [
        <xref ref-type="bibr" rid="ref20">20</xref>
        ], etc.
      </p>
      <p>The problem of automatically identifying individual organisms rather than species
has received much less attention (except for humans of course). Yet, for some
groups, it is preferable to monitor the organisms at the individual level rather
than at the species level. This is notably the case of big animals, such as
whales and elephants, whose population are scarcer and who are traveling longer
distances. Monitoring individual animals allow gathering valuable information
about population sizes, migration, health, sexual maturity and behavior
patterns. Tracking devices and tagging technologies are only part of the solution
because of their invasive character, relatively high cost and limited lifetime.
Morphological/biometric approaches are a complementary approach that is less
invasive, more durable and cheaper for nature observers mobilized on a given
spot. Using natural markings to identify individual animals over time is usually
known as photo-identi cation. This research technique is used on many species
of marine mammals. Initially, scientists used arti cial tags to identify individual
whales, but with limited success (most tagged whales were actually lost or died).
In the 1970s, scientists discovered that individuals of many species could be
recognized by their natural markings. These scientists began taking photographs
of individual animals and comparing these photos against each other to identify
individual animal's movements and behavior over time. Since its development,
photo-identi cation has proven to be a useful tool for learning about many
marine mammal species including humpbacks, right whales, nbacks, killer whales,
sperm whales, bottlenose dolphins and other species to a lesser degree.
Nowadays, this process is still mostly done manually making it impossible to get an
accurate count of all the individuals in a given large collection of observations.
Researchers usually survey a portion of the population, and then use statistical
formulae to determine population estimates. To limit the variance and bias of
such an estimator, it is however required to use large-enough samples which still
makes it a very time-consuming process. Automating the photo-identi cation
process could drastically scale-up such surveys and open brave new research
opportunities for the future.</p>
      <p>
        In this paper, we address more particularly the problem of discovering all
individual humpack whales appearing in a large collection of caudal's images in a
fully unsupervised way, i.e. without any knowledge on the number of individuals
and without any training samples of these individuals. This is in essence a di
erent and more challenging problem than the supervised recognition of individual
whales such as the challenge proposed by NOAA Fisheries through the Kaggle
platform4. Such supervised scenario is actually only a ordable when the
individuals are already well known and well illustrated by tens of pictures that were
hardly collected along the years. On the other side, the unsupervised identi
ca4 https://www.kaggle.com/c/noaa-right-whale-recognition
tion scenario targeted in this paper has the great advantage to allow the use of
unlabeled or very labeled collections of observations which is the vast majority
of available data today.
The experiment reported in this paper was part of the 2016-th edition of the
the LifeCLEF international evaluation campaign [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ] (in particular in the scope
of the sea organisms identi cation task). The data shared through this
challenge consisted of 2005 images of humbpack whales caudals collected by the
CetaMada5 NGO between 2009 and 2014 in Madagascar area. Cetamada is a
Malagasy Non-Pro t Association created in May, 2009 whose goal is to protect
marine mammal population and their habitat in Madagascar through
sustainable eco-tourism and scienti c research. There are presently 4 citizen sciences
data collection sites (St. Marys, Majunga, Ifaty and Fort Dauphin) for which
hotel-establishements and their customers have become sentinels for data
collection. This method helps obtain more than 250 photo IDs each year, which
e ectively helps produce a photo catalogue of humpback whales reproducing on
Malagasy coasts.
      </p>
      <p>
        After acquisition, each photograph was manually cropped so as to focus only on
5 https://www.cetamada.org/
the caudal n that is the most discriminant pattern for distinguishing an
individual whale from another. Figure 1 displays six of such cropped images, each
line corresponding to two images of the same individual. As one can see, the
individual whales can be distinguished thanks to their natural markings and/or
the scars that appear along the years. Automatically nding such matches in the
whole dataset and rejecting the false alarms is di cult for three main reasons.
The rst reason is that the number of individuals in the dataset is high, around
1; 200, so that the proportion of true matches is actually very low (around 0:05%
of the total number of potential matches). The second di culty is that distinct
individuals can be very similar at a rst glance as illustrated by the false positive
examples displayed in Figure 2. To discriminate the true matches from such false
positives, it is required to detect very small and ne-grained visual variations
such as in a spot-the-di erence game. The third di culty is that all images have
a similar water background of which the texture generates quantities of local
mismatches.
To di erentiate biomarkers from the mass of other visual patterns without any
supervision, the research line we investigate in this work is to rely on the spatial
ltering of low-level visual correspondences. Our hypothesis is that the
biomarkers are su ciently localized on the n to be considered as non deformable objects
so that two views of the same biomarker in two di erent images are supposed to
be related by epipolar geometry. On the other side, the raw noisy visual
correspondences at the origin of false alarms should be ltered by the use of geometric
rules. The standard solution to perform such epipolar geometry estimation is to
use the RANSAC algorithm [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ]; it consists in generating transformation
hypotheses using a minimal number of low-level visual correspondences and then
evaluating each hypothesis based on the number of inliers among all features
under that hypothesis. The main advantage of the RANSAC algorithm is that it
is robust to the presence of a high number of outliers which makes it suitable to
deal with the large numbers of false alarms that are generated by the raw visual
matching of the local features.
      </p>
      <p>
        As the RANSAC algorithm can be rather slow, an e cient variant, LO-RANSAC,
was proposed by Chum et al. [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ] and has been proved to provide consistent
speedups in many image retrieval frameworks [
        <xref ref-type="bibr" rid="ref1 ref17 ref18 ref19">19,18,17,1</xref>
        ]. It involves generating
hypotheses of an approximate model thanks to the shape information provided
with the a ne-invariant image regions from which the visual features were
extracted. With this method, an hypothesis can be generated with only a single
pair of corresponding features whereas two or three are required when using
only the feature positions. This greatly reduces the number of possible
hypotheses which need to be considered by the RANSAC algorithm and signi cantly
speeds up the spatial veri cation procedure. An even faster strategy [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ],
consists in considering only the shape information of the image regions, without
exploiting the positions of the features at all. A rough approximation of the best
transformation can then actually be estimated by a Hough-like voting strategy
on the quantized di erences of the characteristic orientation and scale of each
visual correspondence. Using this so-called weak geometry method allows
trading quality for time and is the only acceptable solution when dealing with huge
image sets and real-time contexts (e.g. a search engine working on billions of
images).
      </p>
      <p>The spatial veri cation we use in our own system is also a variant of the RANSAC
algorithm making use of weak geometry rules generated from the region shape
characteristics. We however do not use the weak geometry to directly generate
an hypothesis from a single visual correspondence. We rather use it to lter the
exact hypothesis generated by the classical RANSAC algorithm. Concretely, if
we restrict our class of transformations to rotation and scaling, the RANSAC
algorithm can generate an hypothesis from any pair of visual correspondences. To
quickly decide whether this hypothesis is relevant or not, we check its consistency
with regard to the two approximate hypothesis generated from the shape
characteristics of each visual correspondence. If any of the two approximate models
does not t the RANSAC hypothesis, we reject that solution without computing
the costly consensus phase. In practice, up to 99% of the RANSAC hypothesis
can be rejected in that way (leading to a consistent speed-up).</p>
      <p>
        Another major di erence between our method and the ones in [
        <xref ref-type="bibr" rid="ref1 ref17 ref18 ref19">19,18,17,1</xref>
        ]
is that we use the ranking of the visual correspondences to further improve the
matching. Our retrieval framework does actually not rely on the popular
bag-ofwords model to generate the raw visual correspondences but on a more accurate
approximate KNN search algorithm (described in section 4). The main
benet is that the precision of our raw visual matches is already much better than
the ones produced by the bag-of-words model (based on vector quantization).
The RANSAC algorithm therefore works on less correspondences and less false
alarms. Another bene t is that each raw visual correspondence fx; yg is
associated with a rank rx(y)). This allows two things: (i) to restrict the generation of
the hypothesis of the RANSAC algorithm to the best match of each feature x in
the transformed image IY . The number of evaluated hypothesis is consequently
reduced, particularly in the presence of numerous repeated visual patterns (the
burstiness phenomenon [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ]) (ii) the ranking can be used in the computation of
the nal score by weighing the contribution of each inlier according to its rank
in the whole dataset. Closest points are then favored to the detriment of the
farthest ones, independently from the feature space density in the neighborhood
of xq. More formally, given a couple of images IX and IY , represented by sets
of local features X and Y , we de ne the following spatially consistent match
kernel:
where rx(y) : Rd ! N+ is a ranking function that returns the rank of the local
feature y according to its L2-distance to the query feature x (within the whole
dataset). The function '() is a decreasing function allowing to give more weights
to the top ranked features (we used the inverse function in our experiments).
Finally, X;Y (x; y) is an indicator function equal to one if the correspondence
(x; y) is an inlier of the geometric model estimating the transformation between
IX and IY , i.e.:
      </p>
      <p>X;Y (x; y) =</p>
      <p>Px
(A^ Py + B^ ) &lt;
(3)
where (A^ ; B^ ) are the parameters of the best transformation estimated by our
accelerated RANSAC algorithm, Px and Py are the spatial positions of x and
y, is a user de ned threshold de ning the spatial tolerance of the inliers (we
used = 16 pixels in our experiments). The number of probes of the RANSAC
algorithm was set to 10K in all our experiments.
4</p>
    </sec>
    <sec id="sec-2">
      <title>Approximate K-NN search scheme</title>
      <p>In practice, to speed up the computation of the matching, the ranking function
rx(y) : Rd ! N+ is implemented as an approximate nearest neighbors search
algorithm based on hashing and probabilistic accesses in the hash table. It takes
as input a query feature x and an inverted index of all local features z 2 Z
extracted from the image collection. It returns a set of m approximated neighbors
with an approximated rank rxm(y) (we used m = 500 in all our experiments). The
exact ranking function rx(y) is simply replaced by this approximated ranking
function in all equations above. Note that the features y that are not returned in
the top-m approximated nearest neighbors are simply removed from the match
kernel equations conducting to a considerable reduction of the computation time.
Consequently, they are implicitly considered as having a rank-based activation
function '(rx(y)) equal to zero which is a good approximation as their rank is
supposed to be higher than m. The higher the value of m is and the lower the
error compared to the exact match kernel.</p>
      <p>
        Let us now describe more precisely our approximate nearest neighbors indexing
and search method. It rst compresses the original feature vectors z 2 Z into
compact binary hash codes h(z) of length b thanks to the use of a data-dependent
high-dimensional hash function. In our experiments, we used RMMH hash
function [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ] that has the advantage to be easily implemented and to be e ective for
any kind of visual features or data distribution. The distance between any two
features x and z can then be e ciently approximated by the Hamming distance
between h(x) and h(y). In all our experiments we used b = 128 bits.
To avoid scanning the whole dataset, the hash codes h(z) derived from the local
features of the entire set Z are then indexed in a hash table whose keys are the
t-length pre x of the hash codes h(z). At search time, the hash code h(x) of a
query feature x is computed as well as its t-length pre x. We then use a
probabilistic multi-probe search algorithm inspired by the one of [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ] to select the
buckets of the hash table that are the most likely to contain exact nearest
neighbors. This is done by using a probabilistic search model that is trained o ine on
the exact m-nearest neighbors of M sampled features z 2 Z. We however use a
simpler search model than the one of [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ]. We actually use a normal distribution
with independent components parameterized by a single vector that is trained
over the exact nearest neighbors of the training samples. At search time, we also
use a slightly di erent probabilistic multi-probe algorithm trading stability for
time. Instead of probing the buckets by decreasing probabilities, we rather use
a greedy algorithm that computes the probability of neighboring buckets and
select only the ones having a probability greater than a threshold that is xed
over all queries. The value of is trained o ine on M training samples and their
exact nearest neighbors so as to reach on average cumulative probability over
the visited buckets. In our experiments, we always used = 0:95 meaning that
on average we retrieve 95% of the exact nearest neighbors in the original feature
space. Once the most probable buckets have been selected, the re nement step
computes the Hamming distance between h(x) and the h(z)'s belonging to the
selected buckets and keep only the top-m matches thanks to a max heap.
      </p>
    </sec>
    <sec id="sec-3">
      <title>Experiments</title>
      <p>
        As mentioned earlier, the system described above was evaluated in the context
of the Sea task of the LifeCLEF 2016 evaluation campaign [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ]. This means that
the experiment was conducted in blind, i.e. without having access to the ground
truth, and thus, without any possibility of learning or tuning the parameters of
the system.
5.1
      </p>
      <sec id="sec-3-1">
        <title>Task Description</title>
        <p>The task was simply to detect as many true matches as possible from the whole
dataset, in a fully unsupervised way. Each evaluated system had to return a run
le (i.e., a raw text le) containing as much lines as the number of discovered
matches, each match being a triplet of the form:</p>
        <p>&lt; imageX:jpg imageY:jpg score &gt;
where score is a con dence score in [0; 1] (1 for highly con dent matches). The
retrieved matches had to be sorted by decreasing con dence score. A run should
not contain any duplicate match (e.g., &lt; image1:jpg image2:jpg score &gt; and
&lt; image2:jpg image1:jpg score &gt; should not appear in the same run le). The
metric used to evaluate each run is the Average Precision:</p>
        <p>AveP =</p>
        <p>PK
k=1 P (k) rel(k)</p>
        <p>M
where M is the total number of true matches in the groundtruth, k is the rank in
the sequence of returned matches, K is the number of retrieved matches, P (k) is
the precision at cut-o k in the list, and rel(k) is an indicator function equaling
1 if the match at rank k is a relevant match, 0 otherwise. The average is over all
true matches and the true matches not retrieved get a precision score of 0.
5.2</p>
      </sec>
      <sec id="sec-3-2">
        <title>Submitted runs</title>
        <p>
          We submitted a total of 3 run les to the LifeCLEF benchmark corresponding
to three con gurations of our system. Each run was computed by (i) searching
each image of the collection one by one, (ii) computing the score K of the image
pairs according to Equation 1 and (iii) rank all pairs by decreasing value of K.
{ Run ZenithINRIA SiftGeo: In this run we used SIFT local features [
          <xref ref-type="bibr" rid="ref15">15</xref>
          ]
extracted around Harris Hessian regions [
          <xref ref-type="bibr" rid="ref16">16</xref>
          ] (without threshold).
{ Run ZenithINRIA GoogleNet 3layers borda: In this run we used o -the-shelf
local features extracted at three di erent layers of GoogLeNet convolutional
neural network [
          <xref ref-type="bibr" rid="ref21">21</xref>
          ] (layer conv2-3x3 : 3136 local features per image, layer
inception 3b output : 784 local features par image, layer inception 4c output :
196 local features per image). The matches found using the 3 distinct layers
were merged through a late-fusion approach based on Borda.
{ Run ZenithINRIA SiftGeo QueryExpansion: This the last run di ers from
the run ZenithINRIA SiftGeo in that a query expansion strategy was used
to re-issue the regions matched with a su cient degree of con dence as new
queries (using the method described in [
          <xref ref-type="bibr" rid="ref12">12</xref>
          ]).
5.3
        </p>
        <p>O</p>
        <p>
          cial LifeCLEF results
The table in Figure 3 provides the scores achieved by the three con gurations
of our system as well as the scores obtained by the system of the other
competitor and described in [
          <xref ref-type="bibr" rid="ref4">4</xref>
          ] (using a Fisher Vector image representation based on
SIFT features and a GMM visual codebook of 256 visual words). Runs
bmetmit whalerun 2 and bmetmit whalerun 3 di er from bmetmit whalerun 1 in that
segmentation propagation was used beforehand so as to separate the background
(the water) from the whales caudal n.
The main conclusion we can draw from the results of this evaluation is that
the spatial arrangement of the local features is a crucial information for rejecting
the false positives (as proved by the much higher Average Precision of our
system compared to the one of bme mit ). As powerful as aggregation-based methods
such as Fisher Vectors are for ne-grained classi cation, they do not capture the
spatial arrangement of the local features which is a precious information for
rejecting the mismatches without supervision. Another reason explaining the good
performance of the best run ZenithINRIA SiftGeo is that it is based on a ne
invariant local features contrary to ZenithINRIA GoogleNet 3layers borda and
bme mit runs that use grid-based local features. Such features are more sensitive
to small shifts and local a ne deformations even when learned through a
powerful CNN such in our run ZenithINRIA GoogleNet 3layers borda. The comparison
of our two runs ZenithINRIA SiftGeo and ZenithINRIA SiftGeo QueryExpansion
show that query expansion did not succeeded in improving the results. Query
expansion is actually a risky solution in that it is highly sensitive to the decision
threshold used for selecting the re-issued matched regions. It can be considerably
increase recall when the decision threshold is well estimated but at the opposite,
it can also boost the false positives when the threshold is too low.
To get a more practical understanding of the performance achieved by our
system, Figure 4 plots the recall-precision curve of our best run. It shows that the
main strength of our system is that it has a very high precision on the best
found matches (thanks to the spatial ltering). Actually, the top-100 matches
were found with a perfect precision of 1:00 which makes our system already
usable for automatically discovering some matches without any human control. On
the other side, our system fails in reaching high recall values automatically. It
would require further human validation to reach reasonable recalls in the range
of 30 80%. But still, this would be a much more easier process than discovering
the matches from scratch.
Table 1 provides the search processing time of each run. It shows that using
the A ne SIFT features or the o -the-shelf CNN features requires an equivalent
amount of time. The total search time for discovering all the matches in the
image collection was about 24 hours. This is yet not negligible but de nitely
acceptable compared to the di culty of doing that manually. One should also
notice that we used a very high quality approximate nearest neighbors search
(alpha = 95%) to favor quality over time. Much more faster runs could be
obtained using moderate values of alpha (e.g. 80%) without degrading much the
results.
In this paper, we addressed the problem of identifying humpack whales
individuals in a large collections of aerial pictures of caudal ns in a fully unsupervised
way. We therefore designed a scalable ne-grained matching system allowing to
discover small rigid visual patterns in highly clutter background. It was
experimented in the context of a blind system-oriented evaluation in which it performed
the best. The comparison to the other evaluated system show that the spatial
arrangement of the local features is a crucial information to discriminate the
individual whales as well as to lter the potentially huge number of false
positive matches. Overall, the Average Precision of our system is about 49. This is
still not satisfactory for a fully automatic detection scenario but, on the other
side, this might already drastically simplify the manual work of the biologists
through the release of interactive validation tools. In further work, we will
attempt to use localized spatially consistent similarities rather than estimating a
global a ne transformation at the image level. Also, we will explore possible
extensions of convolutional auto-encoders as a way to discover the semi-deformable
bio-markers.
        </p>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Arandjelovic</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zisserman</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          :
          <article-title>Three things everyone should know to improve object retrieval</article-title>
          .
          <source>In: Proc. CVPR</source>
          (
          <year>2012</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Bruno</surname>
            ,
            <given-names>O.M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>de Oliveira</surname>
            <given-names>Plotze</given-names>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            ,
            <surname>Falvo</surname>
          </string-name>
          , M.,
          <string-name>
            <surname>de Castro</surname>
          </string-name>
          , M.:
          <article-title>Fractal dimension applied to plant identi cation</article-title>
          .
          <source>Information Sciences</source>
          <volume>178</volume>
          (
          <issue>12</issue>
          ),
          <volume>2722</volume>
          {
          <fpage>2733</fpage>
          (
          <year>2008</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Chum</surname>
            ,
            <given-names>O.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Matas</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Obdrzalek</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          :
          <article-title>Enhancing ransac by generalized model optimization</article-title>
          .
          <source>In: Proc. of the ACCV</source>
          . vol.
          <volume>2</volume>
          , pp.
          <volume>812</volume>
          {
          <issue>817</issue>
          (
          <year>2004</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4. David Papp,
          <string-name>
            <surname>D.L.</surname>
          </string-name>
          , Szucs, G.:
          <article-title>Object detection, classi cation, tracking and individual recognition for sea images and videos</article-title>
          .
          <source>In: Working notes of CLEF</source>
          (
          <year>2016</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Fischler</surname>
            ,
            <given-names>M.A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bolles</surname>
            ,
            <given-names>R.C.</given-names>
          </string-name>
          :
          <article-title>Random sample consensus: A paradigm for model tting with applications to image analysis and automated cartography</article-title>
          .
          <source>Commun. ACM</source>
          <volume>24</volume>
          (
          <issue>6</issue>
          ),
          <volume>381</volume>
          {
          <fpage>395</fpage>
          (
          <year>1981</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Gaston</surname>
          </string-name>
          ,
          <string-name>
            <surname>K.J.</surname>
            ,
            <given-names>O</given-names>
          </string-name>
          <string-name>
            <surname>'Neill</surname>
            ,
            <given-names>M.A.</given-names>
          </string-name>
          :
          <source>Automated species identi cation: why not? Philosophical Transactions of the Royal Society of London B: Biological Sciences</source>
          <volume>359</volume>
          (
          <issue>1444</issue>
          ),
          <volume>655</volume>
          {
          <fpage>667</fpage>
          (
          <year>2004</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7. Goeau, H.,
          <string-name>
            <surname>Bonnet</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Joly</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bakic</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Barthelemy</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Boujemaa</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Molino</surname>
            ,
            <given-names>J.F.</given-names>
          </string-name>
          :
          <article-title>The imageclef 2013 plant identi cation task</article-title>
          .
          <source>In: CLEF</source>
          (
          <year>2013</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>Hossain</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Amin</surname>
            ,
            <given-names>M.A.</given-names>
          </string-name>
          :
          <article-title>Leaf shape identi cation based plant biometrics</article-title>
          .
          <source>In: Computer and Information Technology (ICCIT)</source>
          ,
          <year>2010</year>
          13th International Conference on. pp.
          <volume>458</volume>
          {
          <fpage>463</fpage>
          .
          <string-name>
            <surname>IEEE</surname>
          </string-name>
          (
          <year>2010</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <surname>Jegou</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Douze</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Schmid</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          :
          <article-title>On the burstiness of visual elements</article-title>
          .
          <source>In: Computer Vision and Pattern Recognition</source>
          ,
          <year>2009</year>
          .
          <article-title>CVPR 2009</article-title>
          . IEEE Conference on. pp.
          <volume>1169</volume>
          {
          <fpage>1176</fpage>
          .
          <string-name>
            <surname>IEEE</surname>
          </string-name>
          (
          <year>2009</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <surname>Jegou</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Douze</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Schmid</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          :
          <article-title>Improving bag-of-features for large scale image search</article-title>
          .
          <source>International Journal of Computer Vision</source>
          <volume>87</volume>
          (
          <issue>3</issue>
          ),
          <volume>316</volume>
          {
          <fpage>336</fpage>
          (
          <year>2010</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <string-name>
            <surname>Joly</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Buisson</surname>
            ,
            <given-names>O.:</given-names>
          </string-name>
          <article-title>A posteriori multi-probe locality sensitive hashing</article-title>
          .
          <source>In: Proceedings of the 16th ACM international conference on Multimedia</source>
          . pp.
          <volume>209</volume>
          {
          <fpage>218</fpage>
          .
          <string-name>
            <surname>ACM</surname>
          </string-name>
          (
          <year>2008</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          12.
          <string-name>
            <surname>Joly</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Buisson</surname>
            ,
            <given-names>O.</given-names>
          </string-name>
          :
          <article-title>Logo retrieval with a contrario visual query expansion</article-title>
          .
          <source>In: Proceedings of the 17th ACM international conference on Multimedia</source>
          . pp.
          <volume>581</volume>
          {
          <fpage>584</fpage>
          .
          <string-name>
            <surname>ACM</surname>
          </string-name>
          (
          <year>2009</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          13.
          <string-name>
            <surname>Joly</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Buisson</surname>
            ,
            <given-names>O.</given-names>
          </string-name>
          :
          <article-title>Random maximum margin hashing</article-title>
          .
          <source>In: Computer Vision and Pattern Recognition (CVPR)</source>
          ,
          <source>2011 IEEE Conference on</source>
          . pp.
          <volume>873</volume>
          {
          <fpage>880</fpage>
          .
          <string-name>
            <surname>IEEE</surname>
          </string-name>
          (
          <year>2011</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          14.
          <string-name>
            <surname>Joly</surname>
          </string-name>
          ,
          <article-title>Alexis and Goeau, Herve and Glotin, Herve and Spampinato, Concetto and Bonnet, Pierre and Vellinga , Willem-Pier and Champ, Julien and Planque, Robert and Palazzo, Simone and Muller, Henning booktitle</article-title>
          =
          <source>Proceedings of CLEF</source>
          <year>2016</year>
          , y.:
          <article-title>Lifeclef 2016: multimedia life species identi cation challenges</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          15.
          <string-name>
            <surname>Lowe</surname>
            ,
            <given-names>D.G.</given-names>
          </string-name>
          :
          <article-title>Object recognition from local scale-invariant features</article-title>
          .
          <source>In: Computer vision</source>
          ,
          <year>1999</year>
          .
          <source>The proceedings of the seventh IEEE international conference on. vol. 2</source>
          , pp.
          <volume>1150</volume>
          {
          <fpage>1157</fpage>
          .
          <string-name>
            <surname>Ieee</surname>
          </string-name>
          (
          <year>1999</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          16.
          <string-name>
            <surname>Mikolajczyk</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Schmid</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          :
          <article-title>Scale &amp; a ne invariant interest point detectors</article-title>
          .
          <source>International journal of computer vision 60(1)</source>
          ,
          <volume>63</volume>
          {
          <fpage>86</fpage>
          (
          <year>2004</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          17.
          <string-name>
            <surname>Perd</surname>
            'och,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chum</surname>
            ,
            <given-names>O.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Matas</surname>
            ,
            <given-names>J.:</given-names>
          </string-name>
          <article-title>E cient representation of local geometry for large scale object retrieval</article-title>
          .
          <source>In: Computer Vision and Pattern Recognition</source>
          ,
          <year>2009</year>
          .
          <article-title>CVPR 2009</article-title>
          . IEEE Conference on. pp.
          <volume>9</volume>
          {
          <fpage>16</fpage>
          .
          <string-name>
            <surname>IEEE</surname>
          </string-name>
          (
          <year>2009</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          18.
          <string-name>
            <surname>Philbin</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chum</surname>
            ,
            <given-names>O.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Isard</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sivic</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zisserman</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          :
          <article-title>Lost in quantization: Improving particular object retrieval in large scale image databases</article-title>
          .
          <source>In: Proceedings of the IEEE Conference on Computer Vision</source>
          and Pattern
          <string-name>
            <surname>Recognition</surname>
          </string-name>
          (
          <year>2008</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          19.
          <string-name>
            <surname>Philbin</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chum</surname>
            ,
            <given-names>O.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Isard</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sivic</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zisserman</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          :
          <article-title>Object retrieval with large vocabularies and fast spatial matching</article-title>
          .
          <source>In: Computer Vision and Pattern Recognition</source>
          ,
          <year>2007</year>
          . CVPR'07. IEEE Conference on. pp.
          <volume>1</volume>
          {
          <issue>8</issue>
          .
          <string-name>
            <surname>IEEE</surname>
          </string-name>
          (
          <year>2007</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          20.
          <string-name>
            <surname>Shortis</surname>
            ,
            <given-names>M.R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ravanbakskh</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Shaifat</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Harvey</surname>
            ,
            <given-names>E.S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Mian</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Seager</surname>
            ,
            <given-names>J.W.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Culverhouse</surname>
            ,
            <given-names>P.F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Cline</surname>
            ,
            <given-names>D.E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Edgington</surname>
            ,
            <given-names>D.R.:</given-names>
          </string-name>
          <article-title>A review of techniques for the identi cation and measurement of sh in underwater stereo-video image sequences</article-title>
          .
          <source>In: SPIE Optical Metrology</source>
          <year>2013</year>
          . pp.
          <source>87910G{87910G. International Society for Optics and Photonics</source>
          (
          <year>2013</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          21.
          <string-name>
            <surname>Szegedy</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Liu</surname>
            ,
            <given-names>W.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Jia</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sermanet</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Reed</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Anguelov</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Erhan</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Vanhoucke</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rabinovich</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          :
          <article-title>Going deeper with convolutions</article-title>
          .
          <source>In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition</source>
          . pp.
          <volume>1</volume>
          {
          <issue>9</issue>
          (
          <year>2015</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          22.
          <string-name>
            <surname>Tyagi</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hegde</surname>
            ,
            <given-names>R.M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Murthy</surname>
            ,
            <given-names>H.A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Prabhakar</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          :
          <article-title>Automatic identi cation of bird calls using spectral ensemble average voice prints</article-title>
          .
          <source>In: Signal Processing Conference</source>
          ,
          <year>2006</year>
          14th European. pp.
          <volume>1</volume>
          {
          <issue>5</issue>
          .
          <string-name>
            <surname>IEEE</surname>
          </string-name>
          (
          <year>2006</year>
          )
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>