<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Bag of MFCC-based words for bird identi cation</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Julien Ricard</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Herve Glotin</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>LSIS/DYNI University of Toulon</institution>
          ,
          <country country="FR">France</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>The algorithm used by the authors in the bird identi cation task of LifeCLEF 2016 consists in creating a dictionary of MFCC-based words using k-means clustering, computing histograms of these words over short audio segments and feeding them to a random forest classi er. The o cial score achieved is 0.15 MAP.</p>
      </abstract>
      <kwd-group>
        <kwd>bird identi cation</kwd>
        <kwd>MFCC</kwd>
        <kwd>k-means</kwd>
        <kwd>bag-of-words</kwd>
        <kwd>random forest</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Foreword</title>
      <p>
        The algorithm presented here is quite standard and was initially used on smaller
datasets to improve, in a late fusion scheme, a classi er based on pairs of
spectrogram peaks, described in the context of audio ngerprinting in [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. Because
of some memory issues we have not been able to run this latter algorithm on
the challenge dataset. We propose in the discussion some options to possibly
overcome this issue.
The method is based on the bag-of-words approach, initially used in text
analysis to model long-term distribution of words [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ] and more recently in audio
signal analysis for tasks including spoken language identi cation [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ] or urban
soundscape identi cation [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ]. The di erent steps of the algorithm are:
1. The original 44.1 kHz audio les were split in 0.2s segments with 50% overlap.
2. Only the segments having energy values higher than a relative (to the whole
audio le) value and spectral atness values smaller than an absolute
threshold were kept. This method assumes that a segment containing bird
vocalizations is voiced (i.e. is made of stable time-frequency components, or partials)
and has high energy compared to the environmental noise. Even though not
all bird vocalizations are voiced, this simple technique has proven to give
high precision (close to 1) in a bird presence/absence classi cation1.
1 The recall was not great (about 0.5), but we were more interested in making sure the
detected segments actually contained bird vocalizations than in detecting all these
segments.
The algorithm was implemented in Python, using Numpy, PySoundFile2,
librosa3 and scikit-learn4.
3
      </p>
    </sec>
    <sec id="sec-2">
      <title>Results</title>
      <p>Our method achieved the following results (MAP):
{ with background species: 0.149
{ only species: 0.183
{ soundscape: 0.037
The full list of results is given in http://www.imageclef.org/lifeclef/2016/bird.
4</p>
    </sec>
    <sec id="sec-3">
      <title>Discussion</title>
      <p>The proposed method is quite naive and achieve low performance compared to
the other participants.</p>
      <p>
        The implementation of the random forest algorithm used5 required to be fed
with the whole training dataset, which, because of the large amount of data,
lead to memory issues. We had to limit the number of trees and branches, which
probably decreased the performance of the models. An online (mini-batch)
implementation, such as the ones proposed in [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ] or [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ] could be used to overcome
this problem.
2 https://github.com/bastibe/PySoundFile
3 https://github.com/librosa/librosa
4 http://scikit-learn.org
5 http://scikit-learn.org/stable/modules/generated/sklearn.ensemble.RandomForestClassi er.html
      </p>
      <p>As mentioned in the foreword, the predictions were supposed to be fused with
those of another algorithm. Some preliminary tests on the training set (splitting
the training set in a new 70% training/30% test set) showed that this fusion
improved the performance from 0.13 to 0.2 MAP, using the same random forest
parameters as shown earlier. Combining this with a mini-batch implementation
of random forest to use a larger number of trees and larger trees would possibly
have improved further the performance.</p>
      <p>
        We believe that the weaknesses of our method lie in two main aspects. First,
apart from the MFCC derivatives, which describe only very short-term changes,
we did not encode any temporal information. It has been shown in [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ] that
extracting some signal modulations helped in automatic bird identi cation, and
we assume 6 that the analysis of acoustic sequences, as described for instance
in [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ], might also help. Secondly, while we focused here on the voiced parts of
the signal, some bird vocalizations are unvoiced and should also be taken into
account, by adding some features that could describe them and modifying our
detection of the segments of interest accordingly.
6 We have not yet run any experiment to prove that sequence analysis helped in
automatic bird identi cation, and we have not found any studies showing it. However,
it has been shown that some acoustic sequences in bird songs can be explained by
simple hidden Markov processes [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ], and while it does not prove that it can help in
our task, it proves that some sequencing rules exist, and we assume that they might
be helpful.
      </p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Wang</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          :
          <article-title>An Industrial Strength Audio Search Algorithm</article-title>
          . In: ISMIR, pp.
          <volume>7</volume>
          {
          <issue>13</issue>
          (
          <year>2003</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Sebastiani</surname>
            ,
            <given-names>F..</given-names>
          </string-name>
          <article-title>Machine learning in automated text categorization</article-title>
          .
          <source>In: ACM Computing Surveys 34</source>
          , pp
          <fpage>1</fpage>
          -
          <lpage>47</lpage>
          (
          <year>2002</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Ma</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          and
          <string-name>
            <surname>Haizhou</surname>
            <given-names>L.</given-names>
          </string-name>
          :
          <article-title>Spoken language identi cation using bag-of-sounds (</article-title>
          <year>2005</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Aucouturier</surname>
          </string-name>
          , J.-J.,
          <string-name>
            <surname>Defreville</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          and
          <string-name>
            <surname>Pachet</surname>
            <given-names>F.</given-names>
          </string-name>
          :
          <article-title>The bag-of-frames approach to audio pattern recognition: A su cient model for urban soundscapes but not for polyphonic music</article-title>
          .
          <source>In: Journal of the Acoustical Society of America 122.2</source>
          , pp
          <volume>881</volume>
          {
          <issue>891</issue>
          (
          <year>2007</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <article-title>Sa ari, A</article-title>
          . et al.:
          <article-title>Online Random Forests</article-title>
          .
          <source>In 3rd IEEE ICCV Workshop on On-line Computer Vision</source>
          (
          <year>2009</year>
          ). Code: https://github.com/amirsa ari/online-multiclasslpboost
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Lakshminarayanan</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          et al.:
          <article-title>Mondrian forests: E cient online random forests</article-title>
          .
          <source>In Advances in Neural Information Processing Systems</source>
          (
          <year>2014</year>
          ). Code: https://github.com/balajiln/mondrianforest
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>Stowell</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          and
          <string-name>
            <surname>Plumbley</surname>
          </string-name>
          , M. D.:
          <article-title>Largescale analysis of frequency modulation in birdsong data bases</article-title>
          .
          <source>In Methods in Ecology and Evolution 5</source>
          .9, pp.
          <volume>901</volume>
          {
          <issue>912</issue>
          (
          <year>2014</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>Kershenbaum</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          et al.:
          <article-title>Acoustic sequences in nonhuman animals: a tutorial review and prospectus</article-title>
          .
          <source>In Biological Reviews</source>
          (
          <year>2014</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <surname>Katahira</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          et al.:
          <article-title>Complex sequencing rules of birdsong can be explained by simple hidden Markov processes</article-title>
          .
          <source>In PloS one 6</source>
          .9,
          <issue>e24516</issue>
          (
          <year>2011</year>
          ).
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>