<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Methods for Training Convolutional Neural Networks to Identify Bird Species in Complex Soundscape Recordings</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Konstantin Dmitriev</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Lomonosov Moscow State University</institution>
          ,
          <addr-line>1 Leninskie Gory, Moscow, 119992, Russian Federation</addr-line>
        </aff>
      </contrib-group>
      <abstract>
        <p>The task of bird species identification is very important in ecosystem monitoring. Modern methods based on the use of deep learning will allow such research to be carried out cheaply and on a regular basis. However, creating such algorithms is not an easy task due to the wide variety of birds, their calls, recording conditions, and equipment used. In this paper, some methods are presented for training Convolutional Neural Networks (CNNs) that improve the efectiveness of these models. This includes recording length standardization, data augmentation, mixing, sample selection, and weighting.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;Audio classification</kwd>
        <kwd>sound event detection</kwd>
        <kwd>signal processing</kwd>
        <kwd>convolutional neuron network</kwd>
        <kwd>augmentations</kwd>
        <kwd>spectrogram</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>Bird species diversity and its change in time serve as good indicators of ecosystem state. Traditional
methods of monitoring require the presence of a qualified observer who can manually identify the bird.
It’s quite hard and expensive to conduct such surveys regularly, especially in the case of large areas.
Many birds are small and hard to notice, but they have a loud voice. So, it seems promising to use small
and cheap audio recording devices with omnidirectional microphones instead of human observers and
to process the recordings using modern methods based on deep learning.</p>
      <p>Creating and training such algorithms is a dificult task, however. The first problem is the diversity
of bird species and their recording conditions. There are many birds that can imitate other birds or
even repeat a sound they once heard and liked. Many animals and insects sound like birds. The second
problem is the diference between the available training data and the real recordings to be processed.
Usually, the training data is a set of short bird call recordings made by diferent people at diferent
locations. They try to make the recordings clean and loud, without noise or interference. So, good
equipment is used, including directional microphones, and bad recordings are dropped. The third
problem is that only weak labels are given that identify the bird’s existence in each recording but not
the exact call position.</p>
      <p>
        BirdCLEF 2024 is a competition that is supposed to address the mentioned problems [
        <xref ref-type="bibr" rid="ref1 ref2">1, 2</xref>
        ]. It is a part
of the LifeCLEF 2024 conference [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]. The task is to identify bird calls in a set of recordings made in
the Western Ghats, India. The training dataset consists of recordings from the xeno-canto project [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ].
Each of them has primary and secondary labels. The primary label corresponds to the main bird that
can be heard, and the secondary labels are used to mark additional birds that can accompany the main
bird. 182 bird species, whose existence needs to be predicted, were selected by the organizers. The
predictions must be made for each of the 5-second-long intervals of about 1100 recordings that form
the hidden dataset. An additional constraint is that the task must be completed in 2 hours using CPU
only. The macro-averaged AUC ROC that discards classes with no true labels was used as a metric in
the competition. To prevent overfitting, the full hidden dataset is split into the public and private parts
(approximately 35% and 65% of the data, respectively). The corresponding public and private scores are
calculated independently, and only the public score is known at the competition time.
      </p>
      <p>This article presents the methods that can be used to overcome the aforementioned dificulties and
improve the results of bird call recognition.</p>
    </sec>
    <sec id="sec-2">
      <title>2. Methods</title>
      <sec id="sec-2-1">
        <title>2.1. Model architecture</title>
        <p>
          The used model is based on the model proposed in [
          <xref ref-type="bibr" rid="ref5">5</xref>
          ]. Its scheme is presented in Fig 1. After the signal
is loaded and normalized, the spectrogram transform is performed. Then, it is fed to a backbone CNN,
followed by  multi-head attention blocks. Finally, a log-sum-exp pooling layer is used to extract the
label. The multi-head attention mechanism is described in the paper [
          <xref ref-type="bibr" rid="ref6">6</xref>
          ], and it is implemented in
PyTorch by the torch.nn.MultiheadAttention class. Its main parameters are the number  of
parallel attention heads in each of the blocks and the embedding dimension .
        </p>
        <p>To find the best combination of the model parameters, a number of tests were conducted. Instead of
using the macro-averaged AUC ROC score, which became very close to one after a few epochs, the
accuracy score was used in a 5-fold cross-validation (CV) scheme. In the tests conducted, the backbone
as well as the values , , and  were varied. The results are presented in Table 1.</p>
        <p>
          The simple resnet18 [
          <xref ref-type="bibr" rid="ref7">7</xref>
          ] backbone was used to check diferent parameters. Using it, the best results
were achieved with  = 256,  = 64 and  = 2. This slightly difers from the parameters presented
in [
          <xref ref-type="bibr" rid="ref5">5</xref>
          ], where they were set as  = 768,  = 8 and  = 2. The results significantly improve with
heavier backbones, among which seresnext26t_32x4d [
          <xref ref-type="bibr" rid="ref8">8</xref>
          ] seems the best.
        </p>
      </sec>
      <sec id="sec-2-2">
        <title>2.2. Dealing with diferent lengths of the recordings</title>
        <p>An important problem with the recordings is that they have diferent lengths. The shortest of them
lasts only 0.5 seconds, while the longest is more than 1.5 hours. Diferent lengths don’t allow using
these recordings in batches while training the model. This causes the training to be slow.</p>
        <p>There are several possible ways to overcome this dificulty. Let 0 be the fixed final length of each
recording processed. If the initial recording is shorter, then it is simply padded with zeros. If it is
longer, the “First” approach is to use only the first 0-long interval of the recording. The “First and last”
approach is to use the first 0/2 and last 0/2 intervals stacked together. This seems reasonable because,
as a rule, the recordings were processed by their authors before being uploaded to the website. So, one
can suspect that irrelevant sounds were cut of from the beginning and the end of each recording. These
two approaches are quite popular among the competition solutions. However, a lot of information is
lost. Instead, the third approach is proposed, which is called the “Sum” approach and illustrated in Fig 2.
It consists of the following steps.</p>
        <p>1. Each recording, having length , is padded with z =  − 0int zeros, where int = ⌈/0⌉, i.e.,
the resulting recording contains a whole number of intervals with the length of 0.
2. As an augmentation, a random circular shift is performed.
3. The recording is split into int intervals, and they are summed together. The length of the result
is equal to 0.</p>
        <p>There is no information loss in the third approach. The overlapping of bird calls that may occur doesn’t
seem to be a problem since it corresponds to a situation when many birds vocalize at the same time.</p>
        <p>The same model (seresnext26t_32x4d;  = 768,  = 8,  = 2) was trained using all the described
approaches with diferent 0 values. Every training was repeated three times with diferent random
seeds, and the “best” of them with the highest score on the public dataset was selected. The resulting
public and private scores are presented in Table 2.</p>
        <p>The results produced with diferent approaches are close to each other. However, the scores of the
“First” approach are slightly better with low 0 values. The growth of 0 doesn’t improve the scores
of the “First” and “First and last” approaches, but the scores of the “Sum” approach increase, and it
becomes preferable with large 0. At the same time, increasing 0 makes the model training longer.</p>
      </sec>
      <sec id="sec-2-3">
        <title>2.3. Train data selection</title>
        <p>
          The train dataset suggested at BirdCLEF 2024 competition contains approximately half of all the
recordings from xeno-canto project [
          <xref ref-type="bibr" rid="ref4">4</xref>
          ], with primary labels corresponding to 182 target birds. So, the
obvious step is to download the absent data and create an additional dataset. Merged together, these
two datasets contain about 40 000 recordings.
        </p>
        <p>Using the whole dataset, however, doesn’t improve the model score. It seems strange because the
higher diversity of data usually causes an increase in the generalization ability of the model. So, one
may expect that the additional data corrupts the dataset somehow, making it not correspond to the
hidden dataset.</p>
        <p>To find out the reason for the described behavior, the AUC ROC probing technique is proposed. This
technique is based on splitting the target data into several parts and masking the model predictions so
that only one part of the data is scored each time. In the BirdCLEF competitions, the bird species can be
grouped together. For example, let  = 182 be the total number of bird species. One can select a group
of 0 &lt;  &lt;  bird species. During the submission, the predictions corresponding to the rest  − 
species are set to zero. The constant prediction produces an AUC ROC score of 0.5. So, the resulting
AUC ROC score  is equal to  = ( + 0.5( − ))/ , and the AUC ROC score  of the selected
group is equal to  = ( − 0.5( − ))/. Using this simple formula, it is possible to estimate the
model performance for diferent groups of species. If these scores difer significantly and the size of
each group is large enough, one may assume that the feature used for group selection is important.</p>
        <p>One of the possible features that can afect the model’s performance is the number rec of available
recordings of each bird species, i.e., the frequency of its occurrence in the dataset. Following the
proposed AUR ROC probing technique, the species were split into five groups (Table 3). It can be
assumed that the model will work well for those birds for which the training set contains many records
and poorly in the other case. However, the situation is diferent. Group 5 with at least 200 recordings of
each bird species, has almost as low scores as Group 1 with a maximum of 20 recordings.</p>
        <p>One can conclude that, for some reason, the model has significant dificulties when dealing with
common birds. One of the reasons for this is that common birds may be present in many recordings
in the background while they are not marked, even with secondary labels. Another possible reason is
the geographical distribution of the places at which the recordings were made. Indeed, many common
birds were recorded in Europe, America, or Africa, far away from the region of interest. Birds can have
local dialects. Also, a bird, which is common in Europe, can be rare in India.</p>
        <p>The geographic information can be easily taken into account since GPS coordinates are provided.
The locations of places at which five bird species were recorded are presented in Fig 3. Here, “zitcis1”,
“commoo3” and “barswa” are the primary labels of common birds, with the number of recordings equal
to 500 in the competition dataset. The fourth bird, “revbul” is medium-rare and has 101 recordings, while
the fifth, “maltro1”, is a rare bird with 17 recordings. However, it is noticeable that almost all recordings
of common birds were made outside India, while “revbul” and “maltro1” are endemic birds. As a result,
the number of recordings of all common and rare species that are made in the Indian region is quite
low and significantly less than that of medium-rare birds. This observation explains the diferences in
model scores across diferent groups of species.</p>
        <p>To handle this observation, the algorithm for train data selection is proposed, which consists of the
following steps.</p>
        <p>
          1. Prepare the whole dataset with all the recordings from the xeno-canto project [
          <xref ref-type="bibr" rid="ref4">4</xref>
          ] that contain
the target bird species calls.
2. For each recording, calculate its distance WG to the Western Ghats region. This can be done, for
example, by placing a large number of points in the Western Ghats region and calculating the
minimal distance between these points and the point where the recording was made.
3. Specify the maximum distance max and drop all the recordings with larger distances.
4. Calculate the number of recordings ̃︀rec for each bird species in the resulting dataset.
5. Calculate the distance weight  for each of the recordings. This weight is a decreasing function
of WG. For example,  = 1 + cos( WG/max) may be used.
6. Calculate the class imbalance weight for each of the recordings. This weight is a decreasing
function of ̃︀rec. For example, imb = 1/̃︀rec may be used.
7. The weight of each of the recordings in the final dataset is the product of distant and class
imbalance weights:  =  · imb.
        </p>
        <p>The value of max is important. In the current competition, setting max = 4000 km was a good
choice. From a geographical point of view, it allows to discard European, American, and most African
data while covering Southern Asia. Using the lower max decreases the diversity of species, and with its
higher value, the training set includes irrelevant data. As a result, the public score of the model worsens
in both cases.</p>
      </sec>
      <sec id="sec-2-4">
        <title>2.4. Additional noise sources</title>
        <p>
          As it was mentioned earlier, there is a huge domain shift between the competition data used for training
and for model scoring. The training dataset contains the recordings from the xeno-canto project [
          <xref ref-type="bibr" rid="ref4">4</xref>
          ].
As a rule, these are recordings of high quality. However, the model is supposed to work well with the
data recorded with an omnidirectional microphone in a noisy environment. The unlabeled soundscapes
dataset is provided with recordings similar to those used for model scoring.
        </p>
        <p>Listening to the unlabeled recordings, it is possible to make a list of present noise sources. They
include car trafic and horns, aircraft noises, sirens, human voices, music, frogs, and cicadas, as well
as broadband noises of rain, wind, or even uncertain nature. The example spectrogram of such an
unlabeled sounscape is presented in Fig 4. One has to introduce all these kinds of interferences to the
model to increase its generalization ability and reduce the domain shift. At the same time, the addition
of an extra sound to the recording must not contain bird calls that can disorient the model.</p>
        <p>
          There are several datasets that can help introduce these noises; for example, the Vehicle Type Sound
Dataset [
          <xref ref-type="bibr" rid="ref9">9</xref>
          ], the Noise Audio Data Dataset with short sounds of diferent natures [
          <xref ref-type="bibr" rid="ref10">10</xref>
          ], the Rain Forest
Dataset [
          <xref ref-type="bibr" rid="ref11">11</xref>
          ] with recordings of several frog species, and the Hindi Speech Classification Dataset with
recordings of short phrases [
          <xref ref-type="bibr" rid="ref12">12</xref>
          ]. As an addition, manual selection of recordings not containing bird
calls can be used [
          <xref ref-type="bibr" rid="ref13">13</xref>
          ]. Although the use of a dataset with the regional dialects spoken in the Western
Ghats alongside Hindi may seem more appropriate, it’s quite hard to find a suficient number of such
recordings distributed freely. At the same time, the influence of including these dialects on model
performance seems minor.
        </p>
        <p>The broadband noises are quite hard to add. On the one hand, many existing recordings of rain and
wind noises contain bird calls, which have to be manually filtered. On the other hand, these noises are
nonstationary, so they can’t be precisely modeled with any kind of simple stationary noise. A similar
situation takes place with the sounds produced by cicadas.</p>
        <p>To deal with this situation, it is proposed to use unlabeled soundscapes. These recordings, however,
contain bird calls that must be excluded. The idea of doing it is based on the fact that background noise
as well as cicadas form patterns on the spectrogram that slowly change over time. The patterns of bird
calls and other noises are irregular. The first step of the algorithm is taking the 1D Fourier transform of
the spectrogram along the time axis. On the second step, only the components with the largest absolute
values remain, while the others are put to zero. In the third step, the inverse 1D Fourier transform is
performed, and the result is multiplied by random noise. The described filtering procedure significantly
reduces the amount of information a spectrogram contains, and its “thin structure” disappears, including
bird calls. The example is presented in Fig 5.</p>
      </sec>
      <sec id="sec-2-5">
        <title>2.5. Standard augmentations and post-processing</title>
        <p>The “standard” augmentations can be used in the BirdCLEF competition. They are performed on
spectrograms and include XY masking, random grid shufle, and recording mixing. XY masking selects
a few rectangular areas in the spectrogram and sets the data inside it to a constant. Random grid shufle
splits the spectrogram into a grid and shufles all its cells. This transform can be used only along the
time axis, and the size of each cell must be greater than the length of a potential bird call, say, 5 seconds.
Recording mixing is the technique of adding two or more recordings together before passing them to a
model. In this case, the resulting recording contains all the birds from the initial recordings. Its weight
is also the sum of the initial weights. These augmentations make the training dataset more diverse.</p>
        <p>The model predictions may be post-processed using sliding window averaging. This approach
assumes that if there is a bird call in a certain time interval, the probability of the same bird call in the
neighboring intervals is also high. So, the final prediction for the current interval is a sum of predictions
for the current, previous, and next intervals with the weights of 1 −
coeficient  is an averaging parameter, which is often set to 0.25 in BirdCLEF competitions. The results
of using diferent  are presented in Table 4. It can be seen, that the value of  = 0.25 is indeed optimal,
2, 
, and  respectfully. The
however, the public score is maximized by  = 0.3.</p>
      </sec>
      <sec id="sec-2-6">
        <title>2.6. Inference time optimization</title>
        <p>The inference time on the CPU is one of the crucial factors in the competition. However, it was noticed
that the models became extremely slow after training. For example, the used model (seresnext26t_32x4d;
 = 768,  = 8,  = 2) processed the 240-seconds-long recording in 60-70 seconds after training,
while the untrained model did the same job in 3.5 seconds. The training procedure doesn’t change the
model architecture and only adjusts its weights.</p>
        <p>After some research, the problem was localized. In the model computational graph, some of the paths
are unnecessary, and the corresponding weights must be set to zero during training. However, using L2
regularization makes these weights very low but not exactly zero. As a result, not only do these paths
consume computational resources, but the computations with such low values are extremely slow. To
prevent this behavior, one has to retrain the model with an additional L1 regularization term, which
causes the small weights to be exactly zero. A simple solution for an already-trained model is weight
rounding. To do it, one has to convert the model precision from float32 to float16 and then back to float32.
Weight rounding can be performed with one line of code in PyTorch: model.half().float().</p>
        <p>Further acceleration is possible with the help of powerful frameworks, such as ONNX and OpenVino.
The model was converted to ONNX format and then exported to OpenVino. The inference time summary
is presented in Table 5. OpenVino seems faster than ONNX, but one may notice that it doesn’t use all
available CPU cores when running. To do so, hyperthreading (HT) must be switched on. The alternative
is to run multiple threads with multiprocessing. In Python, this can be done in several ways, for example,
by using the ThreadPoolExecutor class (TPE). As expected, the results are the same as with the use of
HT. It should be noted that many laptops have multiple cores and support HT now, so using it may
accelerate the model outside of the competition environment.</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>3. Results</title>
      <p>The main dificulty of the BirdCLEF 2024 competition is an unstable public score. The minor changes in
the model and its training procedure may significantly increase or decrease this score. This makes it
quite hard to test diferent approaches. For example, the results of training with exactly the same model
and data and diferent random seeds are presented in Table 6. On the one hand, the standard deviation
of the public score may be several times larger than the improvement of the score with the use of some
clever technique. On the other hand, this makes it hard to introduce a reliable CV. It can’t be consistent
with the public score because it is unstable, and with the private score because there is a huge domain
shift between the data. So, the reasonable way is to conduct many experiments, take the average of the
public score, and hope that this will not cause overfitting to the public score.</p>
      <p>The methods described in Section 2 were applied consequently, and the results are presented in
Table 7. The most significant improvement was caused by the use of geographical data.</p>
      <p>The inference time of the resulting models was 28 minutes, so an ensemble of four models with the
best public scores was made. The public score of this ensemble was 0.713 (13th place in the competition
public leaderboard). However, the selected models overfit, and the private score of the ensemble was
as low as 0.616. Despite the unlucky model selection, the presented methods seem good and may be
successfully used in the future competitions and applications.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>S.</given-names>
            <surname>Kahl</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Denton</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Klinck</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Ramesh</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Joshi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Srivathsa</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Anand</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Arvind</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>CP</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Sawant</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V. V.</given-names>
            <surname>Robin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Glotin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Goëau</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W.-P.</given-names>
            <surname>Vellinga</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Planqué</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Joly</surname>
          </string-name>
          , Overview of BirdCLEF 2024:
          <article-title>Acoustic identification of under-studied bird species in the western ghats</article-title>
          ,
          <source>Working Notes of CLEF 2024 - Conference and Labs of the Evaluation Forum</source>
          (
          <year>2024</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <article-title>Birdclef 2024 - birdcall species identification from audio</article-title>
          ,
          <year>2024</year>
          . URL: https://www.kaggle.com/ competitions/birdclef-2024.
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>A.</given-names>
            <surname>Joly</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Picek</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Kahl</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Goëau</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Espitalier</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Botella</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Deneu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Marcos</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Estopinan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Leblanc</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Larcher</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Šulc</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Hrúz</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Servajean</surname>
          </string-name>
          , et al.,
          <source>Overview of lifeclef</source>
          <year>2024</year>
          :
          <article-title>Challenges on species distribution prediction and identification</article-title>
          ,
          <source>in: International Conference of the CrossLanguage Evaluation Forum for European Languages</source>
          , Springer,
          <year>2024</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          <article-title>[4] Xeno-canto sharing wildlife sounds from around the world</article-title>
          ,
          <year>2024</year>
          . URL: https://xeno-canto.org.
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>M. V.</given-names>
            <surname>Shugaev</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Tanahashi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Dhingra</surname>
          </string-name>
          , U. Patel,
          <year>Birdclef 2021</year>
          :
          <article-title>building a birdcall segmentation model based on weak labels</article-title>
          ,
          <source>CEUR Workshop Proceedings</source>
          <volume>2936</volume>
          (
          <year>2021</year>
          ). URL: https://ceur-ws.
          <source>org/</source>
          Vol-
          <volume>2936</volume>
          /paper-141.pdf.
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>A.</given-names>
            <surname>Vaswani</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Shazeer</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Parmar</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Uszkoreit</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Jones</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A. N.</given-names>
            <surname>Gomez</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Kaiser</surname>
          </string-name>
          ,
          <string-name>
            <surname>I. Polosukhin</surname>
          </string-name>
          ,
          <article-title>Attention is all you need (</article-title>
          <year>2023</year>
          ). URL: https://arxiv.org/abs/1706.03762. arXiv:
          <volume>1706</volume>
          .
          <fpage>03762</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>K.</given-names>
            <surname>He</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Zhang</surname>
          </string-name>
          , S. Ren,
          <string-name>
            <given-names>J.</given-names>
            <surname>Sun</surname>
          </string-name>
          ,
          <article-title>Deep residual learning for image recognition</article-title>
          ,
          <source>CoRR abs/1512</source>
          .03385 (
          <year>2015</year>
          ). URL: http://arxiv.org/abs/1512.03385. arXiv:
          <volume>1512</volume>
          .
          <fpage>03385</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>J.</given-names>
            <surname>Hu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Shen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Albanie</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Sun</surname>
          </string-name>
          , E. Wu,
          <string-name>
            <surname>Squeeze-</surname>
          </string-name>
          and-excitation networks,
          <year>2019</year>
          . URL: http: //arxiv.org/abs/1709.01507. arXiv:
          <volume>1709</volume>
          .
          <fpage>01507</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          <article-title>[9] Vehicle type sound dataset</article-title>
          ,
          <year>2024</year>
          . URL: https://www.kaggle.com/datasets/brinkor/ vehicle
          <article-title>-type-sound-dataset.</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <article-title>Noise audio data dataset</article-title>
          ,
          <year>2024</year>
          . URL: https://www.kaggle.com/datasets/javohirtoshqorgonov/ noise-audio
          <article-title>-data.</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <article-title>Rainforest connection species audio detection data</article-title>
          ,
          <year>2021</year>
          . URL: https://www.kaggle.com/ competitions/rfcx
          <article-title>-species-audio-detection/data.</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <article-title>Hindi speech classification dataset</article-title>
          ,
          <year>2024</year>
          . URL: https://www.kaggle.com/datasets/vivmankar/ hindi-speech-classification.
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <article-title>Nocall manual classification dataset</article-title>
          ,
          <year>2024</year>
          . URL: https://www.kaggle.com/datasets/janmpia/ nocall-manual-classification.
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>