<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Overview of BirdCLEF 2019: Large-Scale Bird Recognition in Soundscapes</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Stefan Kahl</string-name>
          <email>stefan.kahl@informatik.tu-chemnitz.de</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Fabian-Robert Stoter</string-name>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Herve Goeau</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Herve Glotin</string-name>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Robert Planque</string-name>
          <xref ref-type="aff" rid="aff4">4</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Willem-Pier Vellinga</string-name>
          <xref ref-type="aff" rid="aff4">4</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Alexis Joly</string-name>
          <email>alexis.jolyg@inria.fr</email>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>CIRAD, UMR AMAP</institution>
          ,
          <addr-line>Montpellier</addr-line>
          ,
          <country country="FR">France</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Chemnitz University of Technology</institution>
          ,
          <country country="DE">Germany</country>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>Inria/LIRMM ZENITH team</institution>
          ,
          <addr-line>Montpellier</addr-line>
          ,
          <country country="FR">France</country>
        </aff>
        <aff id="aff3">
          <label>3</label>
          <institution>Universite de Toulon, Aix Marseille Univ</institution>
          ,
          <addr-line>CNRS, LIS, DYNI team, Marseille</addr-line>
          ,
          <country country="FR">France</country>
        </aff>
        <aff id="aff4">
          <label>4</label>
          <institution>Xeno-canto Foundation</institution>
          ,
          <country country="NL">The Netherlands</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>The BirdCLEF challenge|as part of the 2019 LifeCLEF Lab [7]|o ers a large-scale proving ground for system-oriented evaluation of bird species identi cation based on audio recordings. The challenge uses data collected through Xeno-canto, the worldwide community of bird sound recordists. This ensures that BirdCLEF is close to the conditions of real-world application, in particular with regard to the number of species in the training set (659). In 2019, the challenge was focused on the di cult task of recognizing all birds vocalizing in omni-directional soundscape recordings. Therefore, the dataset of the previous year was extended with more than 350 hours of manually annotated soundscapes that were recorded using 30 eld recorders in Ithaca (NY, USA). This paper describes the methodology of the conducted evaluation as well as the synthesis of the main results and lessons learned.</p>
      </abstract>
      <kwd-group>
        <kwd>LifeCLEF</kwd>
        <kwd>bird</kwd>
        <kwd>song</kwd>
        <kwd>call</kwd>
        <kwd>species</kwd>
        <kwd>retrieval</kwd>
        <kwd>audio</kwd>
        <kwd>collection</kwd>
        <kwd>identi cation</kwd>
        <kwd>ne-grained classi cation</kwd>
        <kwd>evaluation</kwd>
        <kwd>benchmark</kwd>
        <kwd>bioacoustics</kwd>
        <kwd>ecological monitoring</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>
        Accurate knowledge of the identity, the geographic distribution and the evolution
of bird species is essential for a sustainable development of humanity as well as
for biodiversity conservation. The general public, especially so-called `birders' as
well as professionals such as park rangers, ecological consultants and of course
ornithologists are potential users of an automated bird sound identi cation system
in the context of wider initiatives related to ecological surveillance or biodiversity
conservation. The BirdCLEF challenge |as part of the 2019 LifeCLEF Lab [7]|
evaluates the state-of-the-art of audio-based bird identi cation systems at a very
large scale. Before BirdCLEF started in 2014, three previous initiatives on the
evaluation of acoustic bird species identi cation took place, including two from
the SABIOD6 group [
        <xref ref-type="bibr" rid="ref2 ref4">5,4,2</xref>
        ]. In collaboration with the organizers of these previous
challenges, the BirdCLEF challenges went one step further by (i) signi cantly
increasing the species number by an order of magnitude, (ii) working on
realworld social data built from thousands of recordists, and (iii) moving to a more
usage-driven and system-oriented benchmark by allowing the use of metadata
and de ning information retrieval oriented metrics. Overall, these tasks were
much more di cult than previous benchmarks because of the higher confusion
risk between the classes, the higher background noise and the higher diversity
in the acquisition conditions (di erent recording devices, contexts diversity, etc.).
      </p>
      <p>The main novelty of the 2017 and 2018 editions of the challenge with respect
to the previous years was the inclusion of soundscape recordings containing
timecoded bird species annotations. Usually Xeno-canto recordings focus on a single
foreground species and result from using mono-directional recording devices.
Soundscapes, on the other hand, are generally based on omnidirectional
recording devices that monitor a speci c environment continuously over a long period.
This new kind of recording re ects passive acoustic monitoring scenarios that
could soon augment the number of collected sound recordings by several orders
of magnitude. Despite the technological progress in recent years, the results of
the previous editions on this challenging soundscape task were quite low. We
decided to shift the focus of the 2019 challenge to soundscape analysis only. We
extend the previous dataset with North American bird species for which more
annotated data was available. In particular, we built a dataset of 350 hours of
soundscapes that were recorded and annotated by expert birders of the
Cornell Lab of Ornithology in Ithaca, NY, USA (see Figure 1). This large volume
of data allowed us to share a fully-annotated, three-day validation dataset to
enable participants to thoroughly evaluate their systems.
2</p>
    </sec>
    <sec id="sec-2">
      <title>Task description</title>
      <p>The 2019 BirdCLEF challenge featured the largest, fully-annotated collection of
soundscape recordings. With respect to real-world use cases, labels and metrics
were chosen to re ect the vast diversity of bird vocalizations and high ambient
noise levels in omnidirectional recordings.
2.1</p>
      <sec id="sec-2-1">
        <title>Goal and evaluation protocol</title>
        <p>The goal of the task is to localize and identify all audible birds within the
provided soundscape test set. Each soundscape is divided into segments of 5</p>
        <sec id="sec-2-1-1">
          <title>6 Scaled Acoustic Biodiversity http://sabiod.univ-tln.fr</title>
          <p>seconds, and a list of species associated to probability scores had to be returned
for each segment. The used evaluation metric is the classi cation mean Average
Precision (cmAP ), considering each class c of the ground truth as a query. This
means that for each class c, all predictions with ClassId = c are extracted
from the run le and ranked by decreasing probability in order to compute the
average precision for that class. The mean across all classes is computed as the
main evaluation metric. More formally:
where C is the number of classes (species) in the ground truth and AveP (c) is
the average precision for a given species c computed as:
cmAP =</p>
          <p>PC
c=1 AveP (c)</p>
          <p>C
AveP (c) =</p>
          <p>Pkn=c1 P (k)
nrel(c)
rel(k)
:
where k is the rank of an item in the list of the predicted segments containing c,
nc is the total number of predicted segments containing c, P (k) is the precision
at cut-o k in the list, rel(k) is an indicator function equaling 1 if the segment
at rank k is a relevant one (i.e. is labeled as containing c in the ground truth)
and nrel(c) is the total number of relevant segments for class c.
2.2</p>
        </sec>
      </sec>
      <sec id="sec-2-2">
        <title>Dataset</title>
        <p>The 2019 dataset contains about 350 hours of manually annotated soundscapes|
most of which were recorded using eld recorders between January and June of
2017 in Ithaca, NY, USA. We used SWIFT recording units provided by the
Bioacoustics Research Program7 of the Cornell Lab of Ornithology (Figure 2).
These omnidirectional recorders capture audio over an array of 30 units spanning</p>
        <sec id="sec-2-2-1">
          <title>7 http://www.birds.cornell.edu/brp/</title>
          <p>(a) SWIFT recorder assembly line
(b) SWIFT recorder in the eld
one square mile of diverse vegetation and water bodies. We randomly selected
one le for each hour of the day recorded with one of the 30 recorders to compile
a data collection of 15 days. Each hour-long recording was then annotated by
experts who provided more than 80,000 bounding boxes|one for each audible
bird vocalization. For the sake of comparability with previous editions, these
annotations were condensed into label lists for 5-second segments of audio.</p>
          <p>In addition, we also re-used the soundscape data from the previous years
of BirdCLEF. More speci cally, this concerns about 4,5 hours of soundscapes
recorded in Columbia by Paula Caycedo Rosales, ornithologist from the
Biodiversa Foundation of Colombia and an active member of Xeno-canto. More
details about this soundscape data (locations, authors, etc.) can be found in the
overview working note of BirdCLEF 2018 [6].</p>
          <p>As for training data, we provided an newly composed Xeno-Canto subset
covering 659 species from South and North America (including all species
annotated in the soundscapes). The vast collection of recordings provided by the
Xeno-canto community often features multiple hundreds of recordings per
(common) bird species. Especially North American species are well represented in
the collection. Therefore, we limited the amount of audio les to 100 recordings
per species. This way, we decreased data imbalance and provided a manageable
amount of data. In total, the training data featured 50,153 les with a total run
length of 608 hours. We selected recordings based on their community rating to
preserve a high quality for most species. Each recording contained weak labels
that state the presence of fore and background species.</p>
          <p>Recordings are associated to various metadata such as the type of sound (call,
song, alarm, ight, etc.), the date of recording, the location, textual comments
of the authors, multilingual common names and collaborative quality ratings.
Additionally, we provided eBird.org frequency lists to enable participants to
decide which species are plausible for a given time, date and location. Frequency
estimations of bird species occurrences were compiled using eBird checklist data
for the soundscape recording locations in the US and Colombia provided by the
eBird API 1.1 (which was unfortunately discontinued in March 2019).</p>
          <p>The shift in acoustic domains between mono-species, high quality recordings
and omnidirectional soundscapes with high ambient noise levels is one of the
major challenges in bird sound recognition for avian activity monitoring.
Participants were required to submit at least one run that used training data only.
Aside from that, participants were allowed to use validation data for training
(despite the fact that this would require extensive annotation when switching
recording locations in real-world applications).
3</p>
          <p>Results
103 participants registered for the BirdCLEF 2019 challenge and downloaded
the dataset. Five of them succeeded in submitting runs, but only four teams
published their approaches. Details of the methods and systems used in the
runs are synthesized in the individual working notes of the participants and are
summarized in this section. In Figure 3 we report the performance achieved by
the 25 collected runs, Table 1 provides more detailed insights and additional
scores for each of the two soundscape recording locations.
3.1</p>
        </sec>
      </sec>
      <sec id="sec-2-3">
        <title>MfN [11], Best overall</title>
        <p>Lasseck managed to achieve top scores in most of the past editions in
BirdCLEF. Most notably, his 2018 performance topped all previous results in the
mono-species recording domain [10] and led to the observation that this task
can be considered solved. This year, MfN build upon the results of past
editions and managed to outperform all other participating teams with his very
deep Inception and ResNet architectures that were pre-trained on ImageNet, as
a continuation of [12]. Lasseck used 5-second spectrograms with mel-compressed
frequency and dB amplitude scale. Sophisticated data augmentation methods
lead to consistent improvements and can be considered a major contribution
to the eld of bioacoustics. Additionally, the use of validation data to ne-tune
the pre-trained networks has a signi cant e ect on the overall scores.
Considering this, annotating soundscapes to adapt neural networks to speci c recording
conditions appears to be well worth the costs.
This team also used Inception and ResNet architectures to conduct their
experiments. Again, task-speci c data augmentation was key to achieve higher scores.
ASAS followed the spectrogram extraction approach of the 2018 Baseline
Repository [8]. The results however{although very competitive|do not outperform
the approach of Lasseck despite the similarities in deep neural network design.
This leads to the assumption that sophisticated augmentation strategies are of
particular importance since they provide the needed variance to the input data
distribution which prevents over tting. The authors state that pre-processing of
training data could have a signi cant impact on the overall performance due to
the di culties of weakly labeled data.
3.3</p>
        <p>
          NWPU [
          <xref ref-type="bibr" rid="ref1">1</xref>
          ]
Consistent with all other approaches, these participants used mel-scale
spectrogram extracted from the provided training data to build a Inception-v3 feature
extractor and classi er. The model was pre-trained on ImageNet and a
number data augmentation methods were applied. However, training deep neural
networks is costly and subsequently, the participants were not able to conduct
the amount of experiments needed to achieve higher scores. The results state
once again the most notable observation across all submission: An elaborate
training regime is key to good overall performance. It appears that this applies
independent of the underlying network architecture.
The submission results of this participant con rm this thought. The author states
that he was able to con rm that deeper architectures do not necessarily lead to
better performance, especially when computational constraints limit the choice
of hyperparameters. Despite very low scores across all runs, the observation
that task-speci c training and model layouts matter was consistent with the
submissions of other teams.
        </p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>Conclusion</title>
      <p>In this edition of the BirdCLEF challenge, participants built on established
systems from previous years, all submitted runs featured a CNN classi er trained
on spectrograms|very deep networks once again performed best. Participants
were able to signi cantly improve the detection performance. In fact, we saw an
increase of more than 180% for the best performing runs (2018: 0.193 - 2019:
0.356). This result is probably largely due to the high number of North
American soundscapes that are less complex than their South American counterparts.
However, the recognition performance for South American soundscapes also
increased signi cantly compared to 2018 with a cmAP of 0.293 in 2019 over 0.222
from last year. Participants were allowed to use any publicly available metadata
and even the provided validation data to improve the performance of their
systems. Although expert annotations are not an adequate (or even easy-to-acquire)
addition for the training of a recognition system for unseen habitats, the increase
in overall performance is considerable. The highest scoring run submitted by MfN
achieved a sample-wise mean average precision (our secondary metric) of 0.446
without the use of validation samples and 0.745 when validation data was used
for training. These scores imply that domain adaption to new acoustic
environments (and recorder characteristics) plays a crucial role and should be subject
of investigation in future editions.</p>
      <p>Acknowledgements The organization of the BirdCLEF task is supported
by the Xeno-canto Foundation, the European Union and the European Social
Fund (ESF) for Germany, as well as by the French CNRS project SABIOD.ORG
and EADM GDR CNRS MADICS, BRILAAM STIC-AmSud, and Floris'Tic.
The annotations of some soundscapes were prepared by the wonderful Lucio
Pando of Explorama Lodges, with the support of Pam Bucur, H. Glotin and
Marie Trone. We want to thank all expert birders who annotated North
American soundscapes with incredible e ort: Cullen Hanks, Jay McGowan, Matt
Young, Randy Little, and Sarah Dzielski.
5. Glotin, H., Dufour, O., Bas, Y.: Overview of the 2nd challenge on acoustic bird
classi cation. In: Proc. Neural Information Processing Scaled for Bioacoustics. NIPS
Int. Conf., Ed. Glotin H., LeCun Y., Artieres T., Mallat S., Tchernichovski O.,
Halkias X., USA (2013), http://sabiod.univ-tln.fr/nips4b
6. Goeau, H., Glotin, H., Planque, R., Vellinga, W.P., Kahl, S., Joly, A.: Overview
of birdclef 2018: monophone vs. soundscape bird identi cation. In: CLEF working
notes 2018 (2018)
7. Joly, A., Goeau, H., Botella, C., Kahl, S., Servajean, M., Glotin, H., Bonnet, P.,
Vellinga, W.P., Planque, R., Stoter, F.R., Muller, H.: Overview of lifeclef 2019:
Identi cation of amazonian plants, south &amp; north american birds, and niche
prediction. In: Proceedings of CLEF 2019 (2019)
8. Kahl, S., Wilhelm-Stein, T., Klinck, H., Kowerko, D., Eibl, M.: Recognizing birds
from sound - the 2018 birdclef baseline system. arXiv preprint arXiv:1804.07177
(2018)
9. Koh, C.Y., Chang, J.Y., Tai, C.L., Huang, D.Y., Hsieh, H.H.: Bird sound classi
cation using convolutional neural networks. In: CLEF working notes 2019 (2019)
10. Lasseck, M.: Audio-based bird species identi cation with deep convolutional neural
networks. In: Working Notes of CLEF 2018 (Cross Language Evaluation Forum)
(2018)
11. Lasseck, M.: Bird species identi cation in soundscapes. In: CLEF working notes
2019 (2019)
12. Sevilla, A., Glotin, H.: Audio bird classi cation with inception v4 joint to an
attention mechanism. In: Working Notes of CLEF 2017 (Cross Language Evaluation
Forum) (2017)</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Bai</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Wang</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chen</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Fu</surname>
            ,
            <given-names>Z.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chen</surname>
          </string-name>
          , J.:
          <article-title>Inception-v3 based method of lifeclef 2019 bird recognition</article-title>
          .
          <source>In: CLEF working notes 2019</source>
          (
          <year>2019</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Briggs</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Huang</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Raich</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Eftaxias</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          , et al.,
          <string-name>
            <surname>Z.L.</surname>
          </string-name>
          :
          <article-title>The 9th mlsp competition: New methods for acoustic classi cation of multiple simultaneous bird species in noisy environment</article-title>
          .
          <source>In: IEEE Workshop on Machine Learning for Signal Processing (MLSP)</source>
          . pp.
          <volume>1</volume>
          {
          <issue>8</issue>
          (
          <year>2013</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Costandache</surname>
            ,
            <given-names>M.C.</given-names>
          </string-name>
          :
          <article-title>Bird species identi cation using neural networks</article-title>
          .
          <source>In: CLEF working notes 2019</source>
          (
          <year>2019</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Glotin</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Clark</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>LeCun</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Dugan</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Halkias</surname>
            ,
            <given-names>X.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sueur</surname>
          </string-name>
          , J.:
          <article-title>Bioacoustic challenges in icml4b</article-title>
          .
          <source>In: in Proc. of 1st workshop on Machine Learning for Bioacoustics. No. USA, ISSN 979-10-90821-02-6</source>
          (
          <year>2013</year>
          ), http://sabiod.org/ ICML4B2013_proceedings.pdf
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>