<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <journal-title-group>
        <journal-title>Conference and Labs of the Evaluation Forum, September</journal-title>
      </journal-title-group>
    </journal-meta>
    <article-meta>
      <title-group>
        <article-title>A Deep Learning Method for Visual Recognition of Snake Species</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Rail Chamidullin</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Milan Šulc</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Jiří Matas</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Lukáš Picek</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Department of Cybernetics, Faculty of Applied Sciences, University of West Bohemia</institution>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Department of Cybernetics, Faculty of Electrical Engineering, Czech Technical University in Prague</institution>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2021</year>
      </pub-date>
      <volume>2</volume>
      <fpage>1</fpage>
      <lpage>24</lpage>
      <abstract>
        <p>The paper presents a method for image-based snake species identification. The proposed method is based on deep residual neural networks - ResNeSt, ResNeXt and ResNet - fine-tuned from ImageNet pre-trained checkpoints. We achieve performance improvements by: discarding predictions of species that do not occur in the country of the query; combining predictions from an ensemble of classifiers; and applying mixed precision training, which allows training neural networks with larger batch size. We experimented with loss functions inspired by the considered metrics: soft F1 loss and weighted cross entropy loss. However, the standard cross entropy loss achieved superior results both in accuracy and in F1 measures. The proposed method scored third in the SnakeCLEF 2021 challenge, achieving 91.6% classification accuracy, Country F1 Score of 0.860, and F1 Score of 0.830.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;Snake Species Identification</kwd>
        <kwd>Fine-grained Classification</kwd>
        <kwd>Computer Vision</kwd>
        <kwd>Convolutional Neural Networks</kwd>
        <kwd>Deep Learning</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>
        The paper describes a method for automatic image-based snake species identification submitted
by the CMP team to the SnakeCLEF 2021 challenge [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ] – a part of LifeCLEF 2021 workshop [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ].
The problem of identifying snake species from images is dificult because the classification is
ifne-grained, some species look very similar, and up to hundreds of diferent snake species live
in one country.
      </p>
      <p>
        Taxonomic knowledge about snakes is crucial in diagnosis and medical response to snakebites.
Accurate identification of the snake species is important for the appropriate treatment of
snakebite victims since specific antivenoms are efective against specific venomous snakes.
Moreover, antivenoms should not be used to treat bites from non-venomous snakes because of
side efects such as allergic reactions [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]. Snakebites are a global health problem that kills or
disables half a million people a year in developing countries [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ].
      </p>
      <p>This paper is structured as follows: Section 2 describes related work focusing on snake species
identification. Section 3 introduces the input data and evaluation methodology of the SnakeCLEF
2021 challenge. Section 4 describes the adopted architecture of deep neural network and the
optimization procedure. Section 5 covers all experiments, ranging from preliminary experiments
to the final challenge submissions. Finally, the results are summarized in Section 6.</p>
    </sec>
    <sec id="sec-2">
      <title>2. Related Work</title>
      <p>
        Before the existence of large-scale image datasets for snake species classification, Abeysinghe
et al. [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ] proposed a one-shot learning approach for fine-tuning a Convolutional Neural
Network (CNN) for the task of snake species identification. The authors used a small dataset of 84
snake species, with most species having no more than 3 training images. The authors utilize
a Siamese network [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ] that ranks similarity between two inputs: The network is trained by
binary cross entropy minimization to estimate the probability of the query image belonging
to the same class as the reference image. At test time the query image is compared against all
annotated reference images of each class.
      </p>
      <p>
        In 2020, the first year of the SnakeCLEF challenge [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ], introduced a dataset with 287,632
images of 783 snake species taken in 145 countries. Only two teams presented their recognition
systems for identifying snake species.
      </p>
      <p>
        The best scoring team in SnakeCLEF 2020, gokuleloop [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ], fine-tuned ResNet-50-V2 [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ] from
ImageNet-1K and ImageNet-21K [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ] pre-trained checkpoints, the latter leading to better results.
The author applied the following training techniques:
• Gradient accumulation – a technique that accumulates gradients from small mini-batches
allowing larger efective mini-batch size.
• Mixup augmentation [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ] – an augmentation technique that combines random image
pairs from the training dataset.
• Group normalization [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ] – diferently from batch normalization, GN divides the channels
into groups and computes the mean and variance within each group.
      </p>
      <p>
        The second team in SnakeCLEF 2020, FHDO_BCSG [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ], first detected regions where snakes
occur using a Mask R-CNN [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ] object detector, and then classified the snake species in the
regions using EficientNet [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ]. The authors adjusted the output probabilities of EficientNet
based on the geographic location of the image: The softmax values for each image were
multiplied by the species a priori probability for a given geographic location. To clean the
training dataset from noisy samples, the authors utilized an ImageNet-1K pre-trained ResNet-50
network and discarded images not classified as snake and reptile classes.
      </p>
    </sec>
    <sec id="sec-3">
      <title>3. Challenge Description</title>
      <sec id="sec-3-1">
        <title>3.1. Dataset</title>
        <p>The training dataset provided by SnakeCLEF 2021 covers 772 snake species and contains
annotated images from three diferent sources: iNaturalist, HerpMapper and Flickr. Examples
of images are in Figure 1. The majority of images are from iNaturalist and HerpMapper, with
277,025 and 58,351 images, respectively. Their labels are confirmed by human annotators.
The Flickr dataset is the smallest, with 50,630 web-scraped images that contain noisy data.</p>
        <p>In total, 386,006 images with annotations were provided. Training with external data was not
allowed.</p>
        <p>The challenge organizers suggested a subset of 70,208 images, referenced as a mini-subset
in the rest of this paper, made of samples from INaturalist and HerpMapper. The experiments
described in Section 5 are based on the said subset.</p>
        <p>In addition to the images, the dataset contains metadata with information about the country
where the image was taken. In total, the training dataset includes images from 188 countries.
The dataset is fine-grained with a long tail class distribution. More than 22,000 images represent
the most frequent species, while the least frequent species have only 10 images. The least
represented species are often found in regions such as Middle and South America, South Africa
and Australia. Table 1 shows the distribution of images in geographical regions. For some
images, information about the geographical location is missing.</p>
        <p>Furthermore, the challenge organizers provided 28,418 images without annotations. Top one
species predictions for the test images were sent to the organizers to participate in the challenge.</p>
      </sec>
      <sec id="sec-3-2">
        <title>3.2. Data Preparation</title>
        <p>During the data exploration phase, we discovered that training and validation datasets contain
noisy data from Flickr. The noisy data are non-relevant images with various animal species
or objects. We estimate1 that the percentage of non-relevant images is 10.6 ± 0.1, with 95%
confidence interval. We decided to remove all Flickr images and proceeded with verified images
from iNaturalist and HerpMapper.</p>
        <p>1We used the Student’s t-distribution with  = 20 samples and  − 1 degrees of freedom where each sample
denotes the percentage of non-relevant images in a set of randomly selected 100 images.</p>
        <p>The challenge organizers suggested a data split with 90% training and 10% validation samples.
However, after removing Flickr images, it turned out that some species were not represented in
the proposed validation set. Table 2 displays the number of snake classes represented, i.e. classes
with at least one image, in the dataset sources. iNaturalist and HerpMapper combined have 768
classes which are all represented across the training set but only 733 classes in the validation
set. We thus created a new dataset split where all classes are represented in both training and
validation splits if more than one image of the species is available. If not, the image is placed in
the training set.</p>
        <p>Technically, the last 10% of images, ordered by metadata ID, for every species and country
combination were selected as the validation data. One validation image was selected for the cases
that had fewer than 10 images. We assume the ID ordering is random w.r.t. image content and
properties.
Mini-subset (introduced in Section 3.1)
iNaturalist + HerpMapper (new split)</p>
      </sec>
      <sec id="sec-3-3">
        <title>3.3. Evaluation Metrics</title>
        <p>The challenge used two metrics for the final evaluation. The primary metric is the macro
averaged F1 Score across countries ("Country F1 Score"), shown in equation 4. The secondary
metric is the macro averaged F1 Score ("F1 Score"), shown in equation 2.</p>
        <p>The F1 Score for each species  = 1, 2, ...,  is computed as a harmonic mean of precision 
and recall :</p>
        <p>The macro averaged F1 Score is the average of the 1 scores of all species:
where  is a  ×  matrix with elements  =
{︃1, country  is a habitat of species 
0, otherwise
Similarly, macro averaged Country F1 Score is obtained by averaging CF1 over all countries:
macro(1) = 1 ∑︁ 1.</p>
        <p>Country F1 Score 1 for each country  = 1, 2, ...,  is the macro averaged F1 Score
computed only for species living in country :
The macro averaged Country F1 Score thus increases the importance of species that appear in
more countries.</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>4. Methodology</title>
      <p>The proposed method is based on the state-of-the-art Convolutional Neural Networks (CNNs)
for image classification, described in Subsection 4.1. The following subsections describe the
optimization procedure, loss functions, the post-processing of the predictions, applying mixed
precision training and implementation details.</p>
      <sec id="sec-4-1">
        <title>4.1. Deep Residual Networks</title>
        <p>
          All experiments are based on deep residual neural networks, namely the original ResNet [
          <xref ref-type="bibr" rid="ref15">15</xref>
          ],
the ResNeXt [
          <xref ref-type="bibr" rid="ref16">16</xref>
          ], and the recent ResNeSt [
          <xref ref-type="bibr" rid="ref17">17</xref>
          ]. The ResNet architecture consists of a stack of
residual blocks – building modules with residual connections that combine input and output by
element-wise addition. The ResNeXt additionally includes a split-transform-merge strategy,
where each block performs a set of transformations with the same topology whose outputs are
aggregated by element-wise addition. For example, a single transformation can be a group of
convolutions. The ResNeSt incorporates a channel-wise attention strategy within each
splittransform-merge block: Each transformation consists of split groups over which the network
calculates the channel-wise split attention weights.
        </p>
        <p>
          All networks in our experiments were fine-tuned from ImageNet-1K [
          <xref ref-type="bibr" rid="ref18">18</xref>
          ] pre-trained
checkpoints. Residual networks typically [
          <xref ref-type="bibr" rid="ref15 ref16">15, 16</xref>
          ] use input size about 224 × 224, the pre-trained
ResNeSt-101 and ResNeSt-200 are available with a larger input sizes of 256 × 256 and 320 × 320,
respectively.
        </p>
      </sec>
      <sec id="sec-4-2">
        <title>4.2. Optimization Procedure</title>
        <p>
          We use two optimization algorithms for training CNN models: stochastic gradient descent with
momentum (SGD) and Adam [
          <xref ref-type="bibr" rid="ref19">19</xref>
          ]. Our preliminary experiments showed that Adam optimizer is
able to converge quickly, but the prediction score is inferior compared to SGD. The application of
the one cycle schedule policy [
          <xref ref-type="bibr" rid="ref20">20</xref>
          ] (one cycle) improved the results when applied with the Adam
optimizer while applying it with SGD did not work well in our preliminary experiments.
        </p>
        <p>The training hyper-parameters, such as learning rate, momentum and weight decay, are listed
in Table 3 and were set the same as in the network pre-training. Batch sizes were adjusted to fit
the network on the graphics processing unit (GPU). The input image size stays the same as in
the pre-trained networks.</p>
        <p>During the training, we select the best checkpoint based on the highest validation Country
F1 Score.</p>
      </sec>
      <sec id="sec-4-3">
        <title>4.3. Country-specific Removal of Predictions</title>
        <p>For each image, the dataset metadata include the country where the image was taken.
Additionally, the dataset comes with a list of countries and snake species that live there. We
utilize this information to adjust the model predictions to the country of the query as follows:
The classifier predictions are set to 0 for all species that do not live in the country of the query.
This adjustment is applied only at test time.</p>
      </sec>
      <sec id="sec-4-4">
        <title>4.4. Mixed Precision Training</title>
        <p>
          When training large CNN architectures, fitting the model into limited GPU memory is a
bottleneck. We considered the following workarounds: selecting a smaller batch size or applying
mixed precision training [
          <xref ref-type="bibr" rid="ref21">21</xref>
          ]. Both approaches have an accuracy trade-of.
        </p>
        <p>Mixed precision training is a technique that combines single-precision (32-bit floats, "FP32")
and half-precision (16-bit floats, "FP16") float numbers. In order to lower the memory
requirements, the forward and backward pass with the large batch size only use a half-precision version
of the model. Then, the gradient descent is applied to the single-precision version of the model.
In every training step following procedure is applied:
1. Apply the forward pass, compute the loss and apply backward pass on a model in FP16.
2. Convert the gradients from FP16 to FP32.
3. Apply the update on the primary model in FP32.</p>
        <p>4. Create a copy of the primary model in FP16.</p>
      </sec>
      <sec id="sec-4-5">
        <title>4.5. Loss Functions</title>
        <p>The baseline loss function for training the classifiers is the standard cross entropy loss:
ℓce = −</p>
        <p>∑︁ log , ,
=1
(5)
where  is the ground truth target and y are the classifier predictions for the -th example,
and , is the prediction for the ground truth class of the -th example.</p>
        <p>The following subsections describe the loss functions proposed to use the challenge metrics,
described in Section 3.3, as a loss measure.</p>
        <sec id="sec-4-5-1">
          <title>4.5.1. F1 Loss with Soft Assignments</title>
          <p>The F1 Score from Equation 2 is not diferentiable and thus cannot be utilized as a loss function
for back-propagation. We use an approximation of the F1 Score, referenced as soft F1 loss in
the rest of this paper, which uses soft assignments that make the function diferentiable:
• the true positives for species  are estimated using the softmax predictions y and one-hot

encoded target vector t as follows: T̂︁P = ∑︀ yt</p>
          <p>=1
• the false positives for species  are estimated using the softmax predictions y and one-hot

encoded target vector t as follows: F̂P︁ = ∑︀ y(1 − t)</p>
          <p>=1
• the false negatives for species  are estimated using the softmax predictions y and one-hot

encoded target vector t as follows: F̂N︁  = ∑︀ (1 − y)t
=1
Notice, that T̂︁P, F̂P︁, and F̂N︁ are now real valued. Soft F1 Score for species , ̂1︁, is obtained by
computing the harmonic mean of precision ̂︀ and recall :
̂︀
̂︀ = T̂︁PT̂+︁PF̂P︁ , ̂︀ = T̂︁PT̂+︁PF̂N︁  , ̂1︁ = ̂︀2̂+︀̂︀̂︀ .</p>
          <p>The macro averaged soft F1 Score is obtained by averaging ̂1︁ over all species:
macro(̂︁1) = 1 ∑︁ ̂1︁.
(9)</p>
          <p>The final loss function is ℓ̂︀1 = 1 − macro(̂︁1), so that it ranges from 0 (perfect) to 1 (worst).</p>
        </sec>
        <sec id="sec-4-5-2">
          <title>4.5.2. Weighted Cross Entropy</title>
          <p>Because the macro averaged Country F1 Score from Equation 4 increases the importance of
species appearing in more countries, we propose a weighted variant of the cross entropy loss
with species weights  based on the number of countries in which it appears:

ℓwce = − ∑︁  log , , (8)</p>
          <p>=1</p>
          <p>The Maximum Likelihood Estimation (MLE) of  would simply count the relative frequencies
 in the provided species-country incidence list. In order to avoid zero weights, we add Laplace
smoothing:
 =</p>
          <p>+ 1

∑︀ ( + 1)
=1
.</p>
        </sec>
      </sec>
      <sec id="sec-4-6">
        <title>4.6. Implementation Details</title>
        <p>
          The proposed method was developed using the PyTorch [
          <xref ref-type="bibr" rid="ref22">22</xref>
          ] machine learning framework and
the fastai framework [23] built on top of PyTorch. The code is available online2. All models
were fine-tuned from ImageNet-1K [
          <xref ref-type="bibr" rid="ref18">18</xref>
          ] pre-trained PyTorch Image Models [24] on one NVIDIA
Tesla V100 with 32GB graphic memory.
        </p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>5. Experiments</title>
      <sec id="sec-5-1">
        <title>5.1. Comparison of Residual Networks</title>
        <p>2https://github.com/chamidullinr/snake-species-identification</p>
      </sec>
      <sec id="sec-5-2">
        <title>5.2. Results of Mixed Precision Training</title>
        <p>As observed in the previous section, ResNeSt-101 with a higher input size achieves the highest
scores of the experimented residual networks. Since its deeper version, ResNeSt-200, does not fit
into our GPU memory with larger batch sizes, we experiment with the mixed precision training
from Section 4.4.</p>
        <p>Table 5 compares the training time and accuracy of ResNeSt-101 and ResNeSt-200 when
training with and without the mixed precision technique. Note that in our computational
environment, mixed precision runs slower than single precision. The prediction scores after
10 epochs show that mixed precision has little impact on prediction accuracy in setups with
the same architecture and batch size. Increasing the batch size from 32 to 64 has a much larger
impact on the accuracy. Thus the network trained on a larger batch size with mixed precision
achieves better scores than the single-precision network with a smaller batch size.</p>
      </sec>
      <sec id="sec-5-3">
        <title>5.3. Evaluation of Diferent Loss Functions</title>
        <p>The loss functions introduced in Section 4.5, namely the soft F1 loss and the weighted cross
entropy loss, resulted in inferior classification scores compared to cross entropy loss, see Table 6.
We, therefore, fine-tune the CNN classifiers with cross entropy loss, and then choose the best
training checkpoint based on the highest validation Country F1 Score.</p>
        <p>One possible explanation for the failure of the soft F1 loss is that the batch size of 64 is
significantly smaller than the total number of classes, 772. This leads to the classes not being
represented in every mini-batch, making the approximation of the F1 loss inaccurate. Figure 2
illustrates the inaccurate approximation of the F1 loss on an example, where the loss values are
mostly 0s or 1s.</p>
      </sec>
      <sec id="sec-5-4">
        <title>5.4. Evaluation of Country-specific Removal of Predictions</title>
        <p>We measure the prediction scores of ResNeSt-200 with and without the removal of species
predictions based on the country incidence information. Table 7 compares the prediction scores
on our validation set. The improvement is 0.150 in F1 Score and 0.193 in Country F1 Score.</p>
      </sec>
      <sec id="sec-5-5">
        <title>5.5. Challenge Submissions</title>
        <p>We submitted the following five runs to the SnakeCLEF 2021 challenge:
CMP_S1: ResNeSt-200 fine-tuned for 20 epochs on the full dataset with SGD.
CMP_S2: ResNeSt-200 from CMP_S1 fine-tuned for additional 10 epochs on the full dataset
with SGD.</p>
        <p>CMP_S3: ResNet-101 fine-tuned for 25 epochs on the full dataset with Adam and one cycle.
CMP_S4: ResNeXt-101 fine-tuned for 30 epochs on the mini-subset from Section 3.1 with Adam
and one cycle.</p>
        <p>CMP_S5: An ensemble of all four previous runs, combining the top one predictions by majority
voting strategy. In case of ties, predictions of CMP_S1 are preferred.</p>
        <p>Table 8 shows the final challenge scores on the test set. While diferent in accuracy, the CNN
architectures ResNeSt-200, ResNeXt-101 and ResNet-101 achieve similar results in the primary
challenge metric, the Country F1 Score. The highest scores are achieved by the ensemble.</p>
        <p>We recognize a shortcoming of the ensemble submission (CMP_S5), which inclines towards
the ResNeSt-200 submissions related to each other (CMP_S2 is fine-tuned from CMP_S1).
The remaining networks cannot outvote an agreement of CMP_S1 and CMP_S2.</p>
      </sec>
    </sec>
    <sec id="sec-6">
      <title>6. Conclusions</title>
      <p>The paper presents a deep learning method for image-based snake species identification, a
finegrained classification problem with a long tail class distribution. The method is based on
deep residual neural networks – ResNeSt, ResNeXt and ResNet – fine-tuned from ImageNet
pre-trained checkpoints. We achieve performance improvements by: discarding predictions of
species that do not occur in the country of the query; combining predictions from an ensemble
of classifiers; and applying mixed precision training, which allows training neural networks
with larger batch size.</p>
      <p>The experimented soft F1 loss and weighted cross entropy loss produced inferior results
compared to the standard cross entropy minimization. Thus, the competition submissions are
ifne-tuned with the standard cross entropy loss.</p>
      <p>The proposed method scored third in the SnakeCLEF 2021 challenge, achieving 91.6%
classification accuracy, Country F1 Score of 0.860, and F1 Score of 0.830.</p>
    </sec>
    <sec id="sec-7">
      <title>Acknowledgments</title>
      <p>This research was supported by the OP VVV funded project CZ.02.1.01/0.0/0.0/16_019/0000765.
LP was supported by the UWB grant, project No. SGS-2019-027.
[23] J. Howard, S. Gugger, Fastai: A Layered API for Deep Learning, Information 11 (2020) 108.</p>
      <p>URL: http://dx.doi.org/10.3390/info11020108. doi:10.3390/info11020108.
[24] R. Wightman, PyTorch Image Models, https://github.com/rwightman/
pytorch-image-models, 2019. doi:10.5281/zenodo.4414861, visited on
2021-0628.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>L.</given-names>
            <surname>Picek</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A. M.</given-names>
            <surname>Durso</surname>
          </string-name>
          , R. Ruiz De Castañeda,
          <string-name>
            <surname>I. Bolon</surname>
          </string-name>
          , Overview of SnakeCLEF 2021:
          <article-title>Automatic Snake Species Identification with Country-Level Focus</article-title>
          ,
          <source>in: Working Notes of CLEF 2021 - Conference and Labs of the Evaluation Forum</source>
          ,
          <year>2021</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>A.</given-names>
            <surname>Joly</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Goëau</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Kahl</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Picek</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Lorieul</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Cole</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Deneu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Servajean</surname>
          </string-name>
          , R. Ruiz De Castañeda,
          <string-name>
            <given-names>G. H.</given-names>
            <surname>Bolon</surname>
          </string-name>
          , Isabelle,
          <string-name>
            <given-names>R.</given-names>
            <surname>Planqué</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W.-P.</given-names>
            <surname>Vellinga</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Dorso</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Bonnet</surname>
          </string-name>
          , I. Eggel,
          <string-name>
            <given-names>H.</given-names>
            <surname>Müller</surname>
          </string-name>
          ,
          <source>Overview of LifeCLEF</source>
          <year>2021</year>
          :
          <article-title>a System-oriented Evaluation of Automated Species Identification and Species Distribution Prediction</article-title>
          ,
          <source>in: Proceedings of the Twelfth International Conference of the CLEF Association (CLEF</source>
          <year>2021</year>
          ),
          <year>2021</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>I.</given-names>
            <surname>Bolon</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A. M.</given-names>
            <surname>Durso</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S. Botero</given-names>
            <surname>Mesa</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Ray</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Alcoba</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Chappuis</surname>
          </string-name>
          , R. Ruiz de Castañeda,
          <article-title>Identifying the snake: First scoping review on practices of communities and healthcare providers confronted with snakebite across the world</article-title>
          ,
          <source>PLOS ONE 15</source>
          (
          <year>2020</year>
          ). URL: https: //doi.org/10.1371/journal.pone.0229989. doi:
          <volume>10</volume>
          .1371/journal.pone.
          <volume>0229989</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>C.</given-names>
            <surname>Abeysinghe</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Welivita</surname>
          </string-name>
          ,
          <string-name>
            <surname>I. Perera</surname>
          </string-name>
          ,
          <article-title>Snake Image Classification Using Siamese Networks</article-title>
          ,
          <source>in: Proceedings of the 2019 3rd International Conference on Graphics and Signal Processing</source>
          ,
          <year>2019</year>
          . URL: https://doi.org/10.1145/3338472.3338476.
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>G.</given-names>
            <surname>Koch</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Zemel</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Salakhutdinov</surname>
          </string-name>
          ,
          <article-title>Siamese Neural Networks for One-Shot Image Recognition</article-title>
          ,
          <source>in: ICML Deep Learning Workshop</source>
          ,
          <year>2015</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>L.</given-names>
            <surname>Picek</surname>
          </string-name>
          ,
          <string-name>
            <surname>I. Bolon</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A. M.</given-names>
            <surname>Durso</surname>
          </string-name>
          , R. Ruiz De Castañeda,
          <article-title>Overview of the snakeclef 2020: Automatic snake species identification challenge</article-title>
          ,
          <source>in: CLEF task overview</source>
          <year>2020</year>
          ,
          <source>CLEF: Conference and Labs of the Evaluation Forum</source>
          ,
          <year>2020</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>G. K.</given-names>
            <surname>Moorthy</surname>
          </string-name>
          ,
          <article-title>Impact of Pretrained Networks For Snake Species Classification</article-title>
          , in: CLEF working notes
          <year>2020</year>
          ,
          <source>CLEF: Conference and Labs of the Evaluation Forum</source>
          ,
          <year>2020</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>K.</given-names>
            <surname>He</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Zhang</surname>
          </string-name>
          , S. Ren,
          <string-name>
            <given-names>J.</given-names>
            <surname>Sun</surname>
          </string-name>
          ,
          <article-title>Identity Mappings in Deep Residual Networks</article-title>
          , in: Computer Vision - ECCV
          <year>2016</year>
          ,
          <year>2016</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>T.</given-names>
            <surname>Ridnik</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Ben-Baruch</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Noy</surname>
          </string-name>
          , L. Zelnik-Manor,
          <fpage>ImageNet</fpage>
          -21K
          <source>Pretraining for the Masses</source>
          ,
          <year>2021</year>
          . arXiv:
          <volume>2104</volume>
          .
          <fpage>10972</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>H.</given-names>
            <surname>Zhang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Cisse</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y. N.</given-names>
            <surname>Dauphin</surname>
          </string-name>
          ,
          <string-name>
            <surname>D.</surname>
          </string-name>
          <article-title>Lopez-Paz, mixup: Beyond Empirical Risk Minimization</article-title>
          , in: International Conference on Learning Representations,
          <year>2018</year>
          . URL: https://openreview.net/forum?id=
          <fpage>r1Ddp1</fpage>
          -
          <lpage>Rb</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>Y.</given-names>
            <surname>Wu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>He</surname>
          </string-name>
          , Group Normalization, in: Computer Vision - ECCV
          <year>2018</year>
          ,
          <year>2018</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>L.</given-names>
            <surname>Bloch</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Boketta</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Keibel</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Mense</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Michailutschenko</surname>
          </string-name>
          ,
          <string-name>
            <given-names>O.</given-names>
            <surname>Pelka</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Rückert</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Willemeit</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C. M.</given-names>
            <surname>Friedrich</surname>
          </string-name>
          ,
          <article-title>Combination of image and location information for snake species identification using object detection and EficientNets</article-title>
          , in: CLEF working notes
          <year>2020</year>
          ,
          <source>CLEF: Conference and Labs of the Evaluation Forum</source>
          ,
          <year>2020</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <given-names>K.</given-names>
            <surname>He</surname>
          </string-name>
          , G. Gkioxari,
          <string-name>
            <given-names>P.</given-names>
            <surname>Dollár</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Girshick</surname>
          </string-name>
          ,
          <string-name>
            <surname>Mask</surname>
            <given-names>R-CNN</given-names>
          </string-name>
          , in: 2017
          <source>IEEE International Conference on Computer Vision</source>
          (ICCV),
          <year>2017</year>
          . doi:
          <volume>10</volume>
          .1109/ICCV.
          <year>2017</year>
          .
          <volume>322</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <given-names>M.</given-names>
            <surname>Tan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Q.</given-names>
            <surname>Le</surname>
          </string-name>
          ,
          <article-title>EficientNet: Rethinking Model Scaling for Convolutional Neural Networks</article-title>
          ,
          <source>in: Proceedings of the 36th International Conference on Machine Learning</source>
          ,
          <year>2019</year>
          . URL: http://proceedings.mlr.press/v97/tan19a.html.
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <given-names>K.</given-names>
            <surname>He</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Zhang</surname>
          </string-name>
          , S. Ren,
          <string-name>
            <given-names>J.</given-names>
            <surname>Sun</surname>
          </string-name>
          ,
          <article-title>Deep Residual Learning for Image Recognition</article-title>
          ,
          <source>in: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR)</source>
          ,
          <year>2016</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [16]
          <string-name>
            <given-names>S.</given-names>
            <surname>Xie</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Girshick</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Dollar</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Tu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>He</surname>
          </string-name>
          ,
          <article-title>Aggregated Residual Transformations for Deep Neural Networks</article-title>
          ,
          <source>in: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR)</source>
          ,
          <year>2017</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [17]
          <string-name>
            <given-names>H.</given-names>
            <surname>Zhang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Wu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Zhang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Zhu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Lin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Zhang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Sun</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>He</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Mueller</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Manmatha</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Smola</surname>
          </string-name>
          , ResNeSt:
          <string-name>
            <surname>Split-Attention Networks</surname>
          </string-name>
          ,
          <year>2020</year>
          . arXiv:
          <year>2004</year>
          .08955.
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          [18]
          <string-name>
            <given-names>J.</given-names>
            <surname>Deng</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W.</given-names>
            <surname>Dong</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Socher</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.-J.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Fei-Fei</surname>
          </string-name>
          ,
          <article-title>ImageNet: A large-scale hierarchical image database</article-title>
          ,
          <source>in: 2009 IEEE Conference on Computer Vision and Pattern Recognition</source>
          ,
          <year>2009</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          [19]
          <string-name>
            <given-names>D. P.</given-names>
            <surname>Kingma</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Ba</surname>
          </string-name>
          ,
          <article-title>Adam: A Method for Stochastic Optimization</article-title>
          ,
          <source>CoRR abs/1412</source>
          .6980 (
          <year>2015</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          [20]
          <string-name>
            <given-names>L. N.</given-names>
            <surname>Smith</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Topin</surname>
          </string-name>
          , Super-Convergence:
          <article-title>Very Fast Training of Neural Networks Using Large Learning Rates</article-title>
          ,
          <year>2018</year>
          . arXiv:
          <volume>1708</volume>
          .
          <fpage>07120</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          [21]
          <string-name>
            <given-names>P.</given-names>
            <surname>Micikevicius</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Narang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Alben</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Diamos</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Elsen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Garcia</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Ginsburg</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Houston</surname>
          </string-name>
          ,
          <string-name>
            <given-names>O.</given-names>
            <surname>Kuchaiev</surname>
          </string-name>
          , G. Venkatesh, H. Wu, Mixed Precision Training, in: International Conference on Learning Representations,
          <year>2018</year>
          . URL: https://openreview.net/forum?id=r1gs9JgRZ.
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          [22]
          <string-name>
            <given-names>A.</given-names>
            <surname>Paszke</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Gross</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Massa</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Lerer</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Bradbury</surname>
          </string-name>
          , G. Chanan,
          <string-name>
            <given-names>T.</given-names>
            <surname>Killeen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Lin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Gimelshein</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Antiga</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Desmaison</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Kopf</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Yang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>DeVito</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Raison</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Tejani</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Chilamkurthy</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Steiner</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Fang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Bai</surname>
          </string-name>
          , S. Chintala,
          <article-title>PyTorch: An Imperative Style, High-Performance Deep Learning Library</article-title>
          , in: H.
          <string-name>
            <surname>Wallach</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          <string-name>
            <surname>Larochelle</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Beygelzimer</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          <article-title>d'Alché-</article-title>
          <string-name>
            <surname>Buc</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          <string-name>
            <surname>Fox</surname>
          </string-name>
          , R. Garnett (Eds.),
          <source>Advances in Neural Information Processing Systems</source>
          <volume>32</volume>
          ,
          <string-name>
            <surname>Curran</surname>
            <given-names>Associates</given-names>
          </string-name>
          , Inc.,
          <year>2019</year>
          , pp.
          <fpage>8024</fpage>
          -
          <lpage>8035</lpage>
          . URL: http://papers.neurips.cc/paper/ 9015-pytorch
          <article-title>-an-imperative-style-high-performance-deep-learning-library</article-title>
          .pdf.
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>