<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Shared nearest neighbors match kernel for bird songs identi cation - LifeCLEF 2015 challenge</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Alexis Joly</string-name>
          <email>alexis.joly@inria.fr</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Valentin Leveau</string-name>
          <email>valentin.leveau@inria.fr</email>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Julien Champ</string-name>
          <email>julien.champ@inra.fr</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Olivier Buisson</string-name>
          <email>olivier.buisson@ina.fr</email>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Inra</institution>
          ,
          <addr-line>AMAP, LIRMM, Montpellier</addr-line>
          ,
          <country country="FR">France</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Inria, LIRMM</institution>
          ,
          <addr-line>Montpellier</addr-line>
          ,
          <country country="FR">France</country>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>Institut National de l'Audiovisuel (INA)</institution>
          ,
          <addr-line>Bry-sur-Marne</addr-line>
          ,
          <country country="FR">France</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>This paper presents a new ne-grained audio classi cation technique designed and experimented in the context of the LifeCLEF 2015 bird species identi cation challenge. Inspired by recent works on ne-grained image classi cation, we introduce a new match kernel based on the shared nearest neighbors of the low level audio features extracted at the frame level. To make such strategy scalable to the tens of millions of MFCC features extracted from the tens of thousands audio recordings of the training set, we used high-dimensional hashing techniques coupled with an e cient approximate nearest neighbors search algorithm with controlled quality. Further improvements are obtained by (i) using a sliding window for the temporal pooling of the raw matches (ii) weighting each low level feature according to the semantic coherence of its nearest neighbors. Results show the e ectiveness of the proposed technique which ranked 2nd among the 7 research groups participating to the LifeCLEF bird challenge.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>
        Building accurate knowledge of the identity, the geographic distribution and the
evolution of living species is essential for a sustainable development of
humanity as well as for biodiversity conservation. In this context, using multimedia
identi cation tools is considered as one of the most promising solution to help
bridging the taxonomic, i.e. the di culty for common people to name observed
living organisms and then produce or access to useful knowledge. The LifeCLEF
[
        <xref ref-type="bibr" rid="ref10">10</xref>
        ] lab proposes to evaluate this challenge in the continuity of the image-based
plant identi cation task was run within ImageCLEF the years before but with
a broader scope (considering birds and sh in addition to plants and audio and
video contents in addition to images). This paper particularly reports the
participation of Inria ZENITH research group to the audio-based bird identi cation
task. Inspired by some recent works on ne-grained image classi cation [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ], we
introduce a new match kernel based on the shared nearest neighbors of the low
level audio features extracted at the frame level. Section 2 describes the
preliminary audio processing and features extraction steps. Section 3 then presents
our new match kernel and the resulting explicit representations to be further
classi ed thanks to a linear supervised classi er (section 4). Section 5 and ??
nally reports and discuss the results we obtained within the LifeCLEF 2015
challenge .
2
      </p>
    </sec>
    <sec id="sec-2">
      <title>Pre-processing and features extraction</title>
      <p>The dataset used for this challenge is composed of 33,203 audio recordings
belonging to 999 bird species from Brazil area. As various recording devices are
used, and because it is di cult to capture these sounds as birds are often far
away from the recording devices, many recordings contains a lot of noise. To
overcome this problem, we used SoX, the "Swiss Army knife of sound
processing programs"4. As a rst step, we used the noisered specialised lter, to lter
out noise from the audio, and then we reduce the length of large (i.e. &gt; 0:1s)
silent passages from audio les to 0:1s. In order to obtain audio les with
ideally no more noise but still enough signal, we tried removing as much noise as
possible (using the noisered amount parameter) while guaranteeing that the
resulting audio le was at least 20% the size of the initial audio record. After this
pre-processing step, we used an open source software framework, marsyas5, to
extract MFCC features with parameters based on the provided audio features
in the Birdclef task : MFCC are computed on windows of 11.6 ms, each 3.9
ms, and we additionally derive their speed resulting in 26-dimensional feature
vectors (13+13) for each frame.
3</p>
    </sec>
    <sec id="sec-3">
      <title>Shared Nearest Neighbors Match Kernel</title>
      <p>
        We consider two recordings Ix and Iy represented by sets of 26-dimensional
MFCC features X = fxg and Y = fyg. We then build on the normalized sum
match kernel proposed by [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ] to compare feature sets:
      </p>
      <p>K(X; Y ) = (X)T (Y ) =</p>
      <p>1 X X k(x; y)
jXj jY j x y
where k() is itself a Mercer kernel allowing to compare two individual local
features x and y. In our case, k() is however not de ned as a direct matching
between x and y but rather as the degree of correlation of their matches in a large
training set. Let denote as Z such a training set composed of N 26-dimensional
MFCC feature vectors z. We introduce the following shared nearest neighbors
(SNN) match kernel :</p>
      <p>KS(X; Y ) =</p>
      <p>1 X X X 'x(z):'y(z)
jXj jY j x y z
4 http://sox.sourceforge.net/
5 http://marsyas.info/
(1)
(2)
with 'x(z) a rank-based activation function given by :
'x(z) =
r log(K) log(rx(z))
log(K)
(3)
where rx(z) : Z ! R+ is a ranking function returning the rank of an item
z 2 Z according to its distance to x and K is the maximum number of items
returned by this ranking function. The distance itself could be a L2 metric in the
original feature space but, as we will see in section 3, we use in practice a more
e cient Hamming embedding scheme. Whatever the distance used, the intuition
of the SNN match kernel is that it counts the number of common neighbors in
the neighborhood of x and in the one of y. The product kz(x; y) = 'x(z):'y(z)
is actually equal to one if z is the nearest neighbor of both x and y and close to
zero if z is not in the top neighbors of either x or y.</p>
      <p>
        Using this shared nearest neighbors kernel instead of a more classical distance
in the feature space has several justi cations and advantages. First,
sharedneighbors techniques are known to overcome several shortcomings of traditional
metrics. They are notably less sensitive to the dimensionality curse, more robust
to noisy data and more stable over unusual features distribution [
        <xref ref-type="bibr" rid="ref2 ref5">2, 5</xref>
        ] .
Measuring the similarity between features by the degree to which their neighbourhoods
in the training set resemble one another is actually a form of generative
metric learning. Features belonging to dense clusters are actually more likely to
share neighbors than uniformly distributed and isolated features. So that their
contribution in the global kernel will be enhanced. Secondly, using an indirect
matching rather than a direct one allows to have en explicit formulation of the
embedded space (X). By factorizing equation 3, it is actually easy to show that
KS (X; Y ) = S (X)T S (Y ) with:
      </p>
      <p>S(X) = XN 1</p>
      <p>
        X 'x(zi): !ei
i=1 jXj x
(4)
So that, the explicit N-dimensional feature vector (X) representing each audio
recording in the training set can be computed before training a simple linear
classi er on top of them. This principle of this approach was already introduced
in the intermediate matching kernel of [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ] and further re-used in many methods
including the NBNN kernel of [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ]. Such methods did however rely on the
distance between the features of the candidate object X and the ones of the training
set Z so that they did not bene t from the nice properties of the SNN-kernel.
Last but not least, one of the main advantage of the SNN match kernel is that
is that it can be easily converted to a sparse representation as the ratio of the
number of values close to zeros is very high. Only the features z lying in the top
neighbors of both x and y lead to consistent component values. In practice, it is
therefore su cient to consider only the top-m neighbors of each feature x and
y to get a good approximation of K(X; Y ). This allows using e cient nearest
neighbors search techniques to construct the explicit representations (X) and
to use an e cient sparse encoding when training linear classi ers on top of them.
Temporal pooling of the raw SNN-based representations As elegant as
the explicit representations S (X) is, it does not conduct to good classi cation
results in practice. It's very high dimensionality, equals to the number of features
in the training set (often millions), actually leads to strong over tting even when
using L2 regularizers with high values of the regularization control parameter
. It is therefore required to group the individual matches of the SNN kernel
before deriving an e ective explicit representation. In this work, we do focus
on the temporal pooling of the raw matches rather than aggregating them in
the feature space as done in many popular image representations such as BoW,
Fisher Vectors or VLAD. We consequently loose some generalization capacity in
the feature space compared to these methods but on the other side we strongly
boost the locality, the interpretability and the discrimination of the trained audio
patterns.
      </p>
      <p>Practically, our temporal pooling algorithm rst aggregates the raw matches
within a sliding window (centered around each frame) and then keep the max
score over the whole record. More formally, we can reformulate our explicit
formulation of Equation 4 as:</p>
      <p>M 0
Sw(X) = X
ti+(w=2)</p>
      <p>X
max X 'x(ztm)A :e!m
ti2[1;Tm] t=ti (w=2) x2X
1
(5)
where M is the number of audio recordings in the training set, Tm the number
of frames of the m-th recording and ztm the MFCC feature of the t-th frame of the
m-th recording. The size w of the sliding window was trained by cross-validation
and then xed to w = 1000 frames (resulting in a sliding window of 3.9 seconds).
Approximate K-NN search scheme In practice, to speed up the
computation of our SNN based representations, the ranking function rx(z) : Z ! R+
is implemented as an approximate nearest neighbors search algorithm based on
hashing and probabilistic accesses in the hash table. It takes as input each query
feature x of the audio recording Ix to be described and returns a set of K
approximated neighbors in Z with an approximated rank rx0(z). The exact ranking
function rx(z) is simply replaced by this approximated ranking function in all
equations above. Note that the features z 2 Z that are not returned in the top-K
approximated nearest neighbors are simply removed from the SNN match
kernels equations conducting to a considerable reduction of the computation time.
Consequently, they are implicitly considered as having a rank-based activation
function 'x(z) equal to zero which is a good approximation as their rank is
supposed to be higher than K.</p>
      <p>
        Let us now describe more precisely our approximate nearest neighbors
indexing and search method. It rst compresses the original feature vectors z 2 Z
into compact binary hash codes h(z) of length b. This is done by using RMMH
[
        <xref ref-type="bibr" rid="ref8">8</xref>
        ], a recent data-dependent hash function family, in order to embed the original
feature vectors in compact binary hash codes of b = 128 bits (the parameter M
of RMMH was xed to M = 32). The distance between any two features x and
z can then be e ciently approximated by the Hamming distance between their
respective 128-length hash codes h(z) and h(x). According to our experiments,
this hashing method provides in our context better performances than several
other tested methods, including random projections or hamming embedding [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ]
(orthogonal random projections).
      </p>
      <p>
        To avoid scanning the whole dataset, the hash codes h(z) derived from the
local features of the training set Z are then indexed in a hash table whose keys
are the t-length pre x of the hash codes h(z). At search time, the hash code h(x)
of a query feature x is computed as well as its t-length pre x. We then use a
probabilistic multi-probe search algorithm inspired by the one of [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ] to select the
buckets of the hash table that are the most likely to contain exact nearest
neighbors. This is done by using a probabilistic search model that is trained o ine on
the exact m-nearest neighbors of M sampled features z 2 Z. We however use a
simpler search model than the one of [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ]. We actually use a normal distribution
with independent components parameterized by a single vector that is trained
over the exact nearest neighbors of the training samples. At search time, we also
use a slightly di erent probabilistic multi-probe algorithm trading stability for
time. Instead of probing the buckets by decreasing probabilities, we rather use
a greedy algorithm that computes the probability of neighboring buckets and
select only the ones having a probability greater than a threshold that is xed
over all queries. The value of is trained o ine on M training samples and their
exact nearest neighbors so as to reach on average cumulative probability over
the visited buckets. In our experiments, we always used = 0:80 meaning that
on average we retrieve 80% of the exact nearest neighbors in the original feature
space. Once the most probable buckets have been selected, the re nement step
computes the Hamming distance between h(x) and the h(z)'s belonging to the
selected buckets and keep only the top-m matches thanks to a max heap.
Weak semantic weighting As we are in the case of weakly annotated audio
recordings with multiple classes (primary and secondary species) and highly
cluttered contexts, we suggest improving our SNN match kernel by weighting the
query features according to the semantic coherence of their k nearest neighbours.
We therefore compute a discrimination score f (x) for all MFCC features x 2 X
of a given audio recording IX . A weak label l(x) is rst estimated for each x as
the most represented label within the k-nearest neighbors of x in the training set
(actually the ones computed by the hash-based k-nn search method described
in section 3). The semantic weight f (x) is then computed as the percentage of
the k-nearest neighbors having the same weak label than the feature itself (i.e.
the percentage of k-nearest neighbors whose label is equal to l(x)). Finally, our
representation of a given audio recording IX becomes:
      </p>
      <p>M 0
Sw0(X) = X
max X f (x):'x(ztm)A :e!m
ti2[1;Tm] t=ti (w=2) x2X
(6)
4</p>
    </sec>
    <sec id="sec-4">
      <title>Training and classi cation</title>
      <p>To achieve an e ective supervised classi cation task, we trained a linear
discriminant model on top of our proposed SNN matching-based representations
(cf. Equation 5). This requires rst building the representations of all audio
recordings in the training set and then in learning as many linear classi ers as
the number of species in the training set. The resulting linear classi ers are of
the form:</p>
      <p>h( Sw0(X)) = !T : Sw0(X) + b
so that they interestingly a ect weights !j to each audio recording in the
training set according to its relevance for the targeted class (rather than a ecting
weights to the individual MFCC features as in the raw representation of
Equation 4). In our experiments, we used a linear support vector machine for training
these discriminant linear models. We more precisely used the LibLinear
implementation of the scikit-learn library with a squared hinge loss function and a L2
penalty. The C parameter of the SVM was xed to C = 100:0 weight(class)
where weight(class) is a class-dependent weight that is automatically adjusted
to be inversely proportional to the class frequency. Finally, the scores returned
by the SVM are converted into probabilities using the following p-value test:
P (class) =
where erf is the Gauss error function and (class) and (class) are
respectively the mean and the standard deviation of the SVM score across the
considered class. We will see in the experiments that this conversion provides a
noticeable accuracy improvement.
5
5.1</p>
    </sec>
    <sec id="sec-5">
      <title>Experiments and results</title>
      <sec id="sec-5-1">
        <title>Dataset and task</title>
        <p>
          The LifeCLEF 2015 bird dataset [
          <xref ref-type="bibr" rid="ref3">3</xref>
          ] is built from the Xeno-canto collaborative
database 6 involving at the time of writing more than 140k audio records covering
8700 bird species observed all around the world thanks to the active work of more
than 1400 contributors. The subset of Xeno-canto data used for the 2015 edition
6 http://www.xeno-canto.org/
of the task contains 33,203 audio recordings belonging to 999 bird species in
Brazil area, i.e. the ones having the more recordings in Xeno-canto database. The
dataset contains minimally 14 recordings per species and minimally 10 di erent
recordists per species.Audio records are associated to various metadata such
as the type of sound (call, song, alarm, ight, etc.), the date and localization
of the observations (from which rich statistics on species distribution can be
derived), some textual comments of the authors, multilingual common names
and collaborative quality ratings (more details can be found in [
          <xref ref-type="bibr" rid="ref3">3</xref>
          ]). The task was
evaluated as a bird species retrieval task. A part of the collection was delivered as
a training set available a couple of months before the remaining data is delivered.
The goal was to retrieve the singing species among the top-k returned for each
of the undetermined observation of the test set. Participants were allowed to use
any of the provided metadata complementary to the audio content but we did
not in our own submissions.
We submitted three runs to be evaluated within the LifeCLEF 2015 challenge:
        </p>
        <p>
          INRIA Zenith Run 1: This run was not based on the method described in
this paper, but on our former instance-based classi cation method [
          <xref ref-type="bibr" rid="ref9">9</xref>
          ] evaluated
within the 2014 BirdCLEF challenge [
          <xref ref-type="bibr" rid="ref4">4</xref>
          ]. This allows us measuring progresses
between that former approach and the new one proposed in this paper. It
basically relied on a very similar matching process than the one described in this
paper but it did not train any supervised classi er on top of the resulting
matching score. It actually only computed the top-30 most similar training records of
each query and then used a simple vote on the labels of the retrieve records as
classi er. It however included a pre- ltering of the training set that removed the
less discriminant MFCC features from the training set.
        </p>
        <p>INRIA Zenith Run 2: The new approach described in this paper.</p>
        <p>INRIA Zenith Run 3: The same approach than Run 2 (i.e. the main
contribution of that paper), but without the conversion of the SVM scores into
probabilities (see section 4).</p>
        <p>The results of the whole challenge, including our own results as well as the
results of the other participating research groups, are reported in Figure 1 and
Table 1.
5.3</p>
      </sec>
      <sec id="sec-5-2">
        <title>Discussion and perspectives</title>
        <p>
          Our system globally achieved very good performance and ranked as the second
best one among the 7 participating research groups. Our best run, i.e. the one
based on the method proposed in this paper, achieved a mAP of 0; 334 when
considering only the primary species of each test recording. This is 3 points
better than the state-of-the-art approach of the QMUL research group which
makes use of unsupervised feature learning as described in [
          <xref ref-type="bibr" rid="ref14">14</xref>
          ] whereas we used
classical MFCC features. Also, compared to the mAP of our rst run (equals to
0:265), it shows that training discriminant models using our SNN match kernel
is much more e ective than using our former semantic pruning and
instancebased classi cation approach. The weights learned by the SVM on the pooled
matches actually compensate most of the bias involved by the heterogeneity of
the noise level in the recordings and the heterogeneity of the recordings length.
The intermediate performance of INRIA Zenith Run 3 shows, however, that the
conversion of the SVM scores into probabilities plays an important role in the
performance of Run2. Our interpretation of this phenomenon is related to the
fact that the number of training records per species follows an heavily tailed
distribution (as in most biodiversity data). The SVM scores are consequently
boosted for the most populated species to the detriment of the less populated
ones. Our p-value normalization allows compensating this bias by normalizing
the distribution across all classes.
        </p>
        <p>
          Now, the performance of our approach is still much lower than the best
performing system of MNB TSA which has a mAP equal to 0:453. Note that their
approach is in essence not so far from ours as they also represent the audio
recordings thanks to their matching score in a reference set of audio segments
[
          <xref ref-type="bibr" rid="ref11">11</xref>
          ]. A major di erence however is that they pre-compute a clean set of
relevant audio segments whereas we use all the recordings of the training set as
vocabulary. They notably consider only the audio recordings with the highest
user ratings in the metadata, and, then extract only the segments that are likely
to contain a bird song (thanks to bandwidth considerations). A second di
erence is that their matching is computed at the signal level whereas we are using
MFCC features that might loose some important information. We believe that
integrating these two additional paradigms within our framework could make
it competitive with their approach. Investigating more in depth the semantic
pruning strategy that we introduced in [
          <xref ref-type="bibr" rid="ref9">9</xref>
          ] but in the context of our new SNN
match kernel might for instance be an e ective way of further improving the
quality of the reference set.
        </p>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Boughorbel</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Tarel</surname>
            ,
            <given-names>J.P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Boujemaa</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          :
          <article-title>The intermediate matching kernel for image local features</article-title>
          .
          <source>In: Neural Networks</source>
          ,
          <year>2005</year>
          . IJCNN'
          <fpage>05</fpage>
          .
          <string-name>
            <surname>Proceedings</surname>
          </string-name>
          . 2005 IEEE International Joint Conference on. vol.
          <volume>2</volume>
          , pp.
          <volume>889</volume>
          {
          <fpage>894</fpage>
          .
          <string-name>
            <surname>IEEE</surname>
          </string-name>
          (
          <year>2005</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Ertoz</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Steinbach</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kumar</surname>
          </string-name>
          , V.:
          <article-title>A new shared nearest neighbor clustering algorithm and its applications</article-title>
          .
          <source>In: Workshop on Clustering High Dimensional Data and its Applications at 2nd SIAM International Conference on Data Mining</source>
          . pp.
          <volume>105</volume>
          {
          <issue>115</issue>
          (
          <year>2002</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Glotin</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Joly</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          , Goeau, H.,
          <string-name>
            <surname>Vellinga</surname>
            ,
            <given-names>W.P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rauber</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          :
          <article-title>Lifeclef bird identi - cation task 2015</article-title>
          . In: CLEF working notes
          <year>2015</year>
          (
          <year>2015</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4. Goeau, H.,
          <string-name>
            <surname>Glotin</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Vellinga</surname>
            ,
            <given-names>W.P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rauber</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          :
          <article-title>Lifeclef bird identi cation task 2014</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Jarvis</surname>
            ,
            <given-names>R.A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Patrick</surname>
            ,
            <given-names>E.A.</given-names>
          </string-name>
          :
          <article-title>Clustering using a similarity measure based on shared near neighbors</article-title>
          . Computers, IEEE Transactions on
          <volume>100</volume>
          (
          <issue>11</issue>
          ),
          <volume>1025</volume>
          {
          <fpage>1034</fpage>
          (
          <year>1973</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Jegou</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Douze</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Schmid</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          :
          <article-title>Hamming embedding and weak geometric consistency for large scale image search</article-title>
          .
          <source>Computer Vision{ECCV</source>
          <year>2008</year>
          pp.
          <volume>304</volume>
          {
          <issue>317</issue>
          (
          <year>2008</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>Joly</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Buisson</surname>
            ,
            <given-names>O.:</given-names>
          </string-name>
          <article-title>A posteriori multi-probe locality sensitive hashing</article-title>
          .
          <source>In: Proceedings of the 16th ACM international conference on Multimedia</source>
          . pp.
          <volume>209</volume>
          {
          <fpage>218</fpage>
          .
          <string-name>
            <surname>ACM</surname>
          </string-name>
          (
          <year>2008</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>Joly</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Buisson</surname>
            ,
            <given-names>O.</given-names>
          </string-name>
          :
          <article-title>Random maximum margin hashing</article-title>
          . In: CVPR. IEEE,
          <string-name>
            <surname>Colorado</surname>
            <given-names>springs</given-names>
          </string-name>
          ,
          <source>United States (Jun</source>
          <year>2011</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <surname>Joly</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Champ</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Buisson</surname>
            ,
            <given-names>O.</given-names>
          </string-name>
          :
          <article-title>Instance-based bird species identication with undiscriminant features pruning-lifeclef 2014</article-title>
          . In: CLEF2014 (
          <year>2014</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <surname>Joly</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          , Muller, H., Goeau, H.,
          <string-name>
            <surname>Glotin</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Spampinato</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rauber</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bonnet</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Vellinga</surname>
            ,
            <given-names>W.P.</given-names>
          </string-name>
          , Fisher,
          <string-name>
            <surname>B.</surname>
          </string-name>
          :
          <article-title>Lifeclef 2014: multimedia life species identi cation challenges</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <string-name>
            <surname>Lasseck</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          :
          <article-title>Large-scale identi cation of birds in audio recordings</article-title>
          .
          <source>In: Working notes of CLEF 2014 conference (</source>
          <year>2014</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          12.
          <string-name>
            <surname>Leveau</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Joly</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Buisson</surname>
            ,
            <given-names>O.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Valduriez</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          :
          <article-title>Kernelizing spatially consistent visual matches for ne-grained classi cation</article-title>
          .
          <source>In: ICMR</source>
          <year>2015</year>
          (
          <year>2015</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          13.
          <string-name>
            <surname>Lyu</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          :
          <article-title>Mercer kernels for object recognition with local features</article-title>
          .
          <source>In: Computer Vision and Pattern Recognition</source>
          ,
          <year>2005</year>
          .
          <article-title>CVPR 2005</article-title>
          . IEEE Computer Society Conference on. vol.
          <volume>2</volume>
          , pp.
          <volume>223</volume>
          {
          <fpage>229</fpage>
          .
          <string-name>
            <surname>IEEE</surname>
          </string-name>
          (
          <year>2005</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          14.
          <string-name>
            <surname>Stowell</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Plumbley</surname>
          </string-name>
          , M.D.:
          <article-title>Automatic large-scale classi cation of bird sounds is strongly improved by unsupervised feature learning</article-title>
          .
          <source>PeerJ 2</source>
          ,
          <issue>e488</issue>
          (
          <year>2014</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          15.
          <string-name>
            <surname>Tuytelaars</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Fritz</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Saenko</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Darrell</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          :
          <article-title>The nbnn kernel</article-title>
          .
          <source>In: Computer Vision</source>
          (ICCV),
          <year>2011</year>
          IEEE International Conference on. pp.
          <year>1824</year>
          {
          <year>1831</year>
          . IEEE (
          <year>2011</year>
          )
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>