<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <journal-title-group>
        <journal-title>Bologna, Italy
$ picekl@kky.zcu.cz (L. Picek)</journal-title>
      </journal-title-group>
    </journal-meta>
    <article-meta>
      <title-group>
        <article-title>Overview of FungiCLEF 2022: Fungi Recognition as an Open Set Classification Problem</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Lukáš Picek</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Milan Šulc</string-name>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Jiří Matas</string-name>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Jacob Heilmann-Clausen</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Center for Macroecology, Evolution and Climate University of Copenhagen</institution>
          ,
          <country country="DK">Denmark</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Department of Cybernetics, Faculty of Applied Sciences, University of West Bohemia</institution>
          ,
          <country country="CZ">Czech Republic</country>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>Rossum.ai</institution>
          ,
          <country country="CZ">Czech Republic</country>
        </aff>
        <aff id="aff3">
          <label>3</label>
          <institution>The Center for Machine Perception Dept. of Cybernetics, FEE, Czech Technical University in Prague</institution>
          ,
          <country country="CZ">Czech Republic</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2022</year>
      </pub-date>
      <volume>000</volume>
      <fpage>0</fpage>
      <lpage>0002</lpage>
      <abstract>
        <p>The main goal of the new LifeCLEF challenge, FungiCLEF 2022: Fungi Recognition as an Open Set Classification Problem, was to provide an evaluation ground for end-to-end fungi species recognition in an open class set scenario. An AI-based fungi species recognition system deployed in the Atlas of Danish Fungi helps mycologists to collect valuable data and allows users to learn about fungi species identification. Advances in fungi recognition from images and metadata will allow continuous improvement of the system deployed in this citizen science project. The training set is based on the Danish Fungi 2020 dataset and contains 295,938 photographs of 1,604 species. For testing, we provided a collection of 59,420 expert-approved observations collected in 2021. The test set includes 1,165 species from the training set and 1,969 unknown species, leading to an open-set recognition problem. This paper provides (i) a description of the challenge task and datasets, (ii) a summary of the evaluation methodology, (iii) a review of the systems submitted by the participating teams, and (iv) a discussion of the challenge results.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;LifeCLEF</kwd>
        <kwd>FungiCLEF</kwd>
        <kwd>fine grained visual categorization</kwd>
        <kwd>metadata</kwd>
        <kwd>open-set recognition</kwd>
        <kwd>fungi</kwd>
        <kwd>species identification</kwd>
        <kwd>machine learning</kwd>
        <kwd>computer vision</kwd>
        <kwd>classification</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>
        Automatic recognition of fungi species assists mycologists, citizen scientists and nature
enthusiasts in species identification in the wild [
        <xref ref-type="bibr" rid="ref1 ref2">1, 2</xref>
        ]. Its availability supports the collection of valuable
biodiversity data. In practice, species identification typically does not depend solely on the visual
observation of the specimen but also on other information available to the observer — such as
habitat, substrate, location and time. The main goal for the new FungiCLEF competition was to
provide an evaluation ground for automatic methods for fungi recognition in an open class set
scenario, i.e, the submitted methods have to handle images of unknown species. Similarly to
previous LifeCLEF competitions, The competition was hosted on Kaggle primarily to attract
machine learning experts to participate and present their ideas. Thanks to rich metadata, precise
annotations, and baselines available to all competitors, the challenge provides a benchmark for
image recognition with the use of additional information.
      </p>
    </sec>
    <sec id="sec-2">
      <title>2. Challenge description</title>
      <p>
        The new FungiCLEF 2022 challenge: Fungi Recognition as an Open Set Classification Problem,
was organized in conjunction with the Conference and Labs of the Evaluation Forum (CLEF1)
and LifeCLEF2 research platform [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ], and FGVC9 Workshop3 — The Ninth Workshop on
FineGrained Visual Categorization organized within the CVPR conference.
      </p>
      <p>The main goal for this challenge was to return the species with the highest likelihood (or
"unknown") for each given test observation, consisting of a set of images and metadata —
the information about habitat, substrate, location, and more is provided for each observation.
Photographs of unknown fungi species had to be classified into an "unknown" class with label
id − 1. The baseline procedure to include metadata in the decision problem and baseline
pretrained image classifiers were provided as part of the task description to all participants. Sample
observations are visualized in Figure 1. Each row represents one observation.</p>
      <sec id="sec-2-1">
        <title>1 http://www.clef-initiative.eu/ 2 http://www.lifeclef.org/ 3 https://sites.google.com/view/fgvc9/home</title>
        <sec id="sec-2-1-1">
          <title>2.1. Dataset</title>
          <p>The FungiCLEF 2022 dataset is based on data collected through the Atlas of Danish Fungi Web4
and mobile (iOS5 and Android6) applications. The Atlas of Danish Fungi is a citizen science
platform with more than 4,000 actively contributing volunteers and with more than 1 million
content-checked observations of approximately 8,650 fungi species.</p>
          <p>
            Development set: For training, the competitors were provided with the DanishFungi 2020
(DF20) dataset [
            <xref ref-type="bibr" rid="ref4">4</xref>
            ]. DF20 contains 295,938 images — 266,344 for training and 29,594 for validation
— belonging to 1,604 species. All training samples passed an expert validation process,
guaranteeing high quality labels. Furthermore, rich observation metadata about habitat, substrate,
time, location, EXIF etc. are provided.
          </p>
          <p>Test set: The test dataset is constructed from all observations submitted in 2021, for which
expert-verified species labels are available. It includes observations collected across all substrate
and habitat types. The test set contains 59,420 observations with 118,676 images belonging
to 3,134 species: 1,165 known from the training set and 1,969 unknown species covering
approximately 30% of the test observations. The test set was further split into public (20%) and
private (80%) subsets — a common practice for Kaggle competitions to prevent participants
from overfitting to the leaderboard.</p>
        </sec>
        <sec id="sec-2-1-2">
          <title>2.2. Metadata</title>
          <p>The visual data is accompanied by metadata for approximately 99% of the image observations
and includes information about attributes related to the environment, place, time and taxonomy.
The provided metadata is acquired by citizen scientists and enables research directions on
combining visual data with metadata. We include 21 frequently filled-in attributes. The most
important attributes are listed and described below.</p>
          <p>Substrate: Substrates on which fungi live and fruit are an essential source of information that
helps diferentiate similarly-looking species. Each species or genus has its preferable substrate,
and it is rare to find it on other substrates. We provide one of 32 substrate types for more than
99% of images. We diferentiate wood of living trees, dead wood, soil, bark, stone, fruits and
others.</p>
        </sec>
      </sec>
      <sec id="sec-2-2">
        <title>4https://svampe.databasen.org/ 5https://apps.apple.com/us/app/atlas-of-danish-fungi/id1467728588 6https://play.google.com/store/apps/details?id=com.noque.svampeatlas</title>
        <p>Habitat: While substrate denotes the spots, the habitat indicates the overall environment
where fungi grow, which is vital for fungal recognition. We include the information about the
habitat for 99.5% of observations.</p>
        <p>Location: Fungi are highly location-dependent. We include multi-level location
information. Starting from GPS coordinates with included uncertainty, we further extracted information
about the country, region and district.</p>
        <p>Time-Stamp: Observation time is essential for fungi classification in the wild as
fruitbodies’ presence depends on seasonality or even the time in a day. Figure 2 shows the monthly
observation frequency for three genera.</p>
        <p>EXIF data: Since the camera device and its settings afect the resulting image, the image
classification models may be biased towards specific device attributes. To allow a deeper study
of such phenomena, we include the EXIF data for approximately 84% of images. We included
attributes such as White Balance, Color Space, Metering Mode, Aperture, Device, Exposure
Time and Shutter Speed.</p>
        <sec id="sec-2-2-1">
          <title>2.3. Timeline</title>
          <p>The competition and data were published in February 2022 through the LifeCLEF, Kaggle, and
FGVC challenge pages allowing anyone with research ambitions to register and participate in
the competition. The test data were provided jointly with the training data allowing continuous
evaluation. Each team could submit up to 2 submissions a day. The deadline for challenge
submissions was May 16, setting the competition for roughly three months. Participants submitted
CSV files containing the Top1 prediction for each fungi observation. Once the submission phase
was closed (mid-May), the participants were allowed to submit post-competition submissions
to evaluate any exciting findings.</p>
        </sec>
        <sec id="sec-2-2-2">
          <title>2.4. Evaluation Protocol</title>
          <p>The evaluation process consisted of two stages: (i) a public evaluation on the public subset
(20%) of the test set, which was available during the whole competition with a limit of two
submissions a day, and (ii) a final evaluation on the private test set (80%) after the challenge
deadline. The main evaluation metric for the competition was the F1, defined as the mean of
class-wise F1 scores:</p>
          <p>F1 = 1 ∑︁ 1 ,</p>
          <p>=1
where  represents the number of classes — in case of the Kaggle evaluation,  = 1, 165
(#classes in the test set) – and  is the species index. The F1 score for each class is calculated as
a harmonic mean of the class precision  and recall  :
1 = 2 ×  ×+  ,  =</p>
          <p>tp
tp + fp
,  =</p>
          <p>tp
tp + fn</p>
          <p>In single-label multi-class classification, the True Positives ( tp) of a species represents the
number of correct Top1 predictions of that species, False Positive (fp) denotes how many times
was diferent species predicted instead of the ( tp), and False Negatives (fn) indicates how many
images of species  have been wrongly classified.</p>
        </sec>
        <sec id="sec-2-2-3">
          <title>2.5. Working Notes</title>
          <p>All participants with valid submissions were asked to provide a Working Note paper — a technical
report with information needed to reproduce the results of all submissions. All
submitted Working Notes were reviewed by 2–3 reviewers. The review process was single-blind
and ofered up to two rebuttals. The acceptance rate was 75%.</p>
        </sec>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>3. Challenge Results</title>
      <p>The oficial challenge results, based on the F1 score, are displayed in Figure 3. The best
performing team — xiong — achieved F1 of 80.43% on the private test set and an accuracy of
65.69% on the complete test set. We note that the order would be diferent in terms of accuracy,
as shown in Figure 4, where the best accuracy of 67.08% on the full rest set was achieved by
team GG, primarily due to a high number of correctly identified out-of-scope observations. In
the case of the out-of-scope (OoS) identification performance, i.e. what proportion of
out-ofscope observations has been correctly classified as OoS, the best performing team with 44.55%
correctly categorized observations was one of the worst-performing teams in terms of F1. As
also displayed in Figure 4 participants identified less than 5% OoS observations and only four
teams achieved accuracy over 10% on out-of-scope observations. In Figure 5 we have evaluated
the species toxicity confusion on the full test set for all the participants, i.e., how often poisonous
species are confused for edible ones and vice versa. Interestingly, the more critical confusion
where poisonous fungi were misclassified as edible is relatively high even for the best scoring
models — 5.70% and 6.63% for team GG and team xiong, respectively.
(1)
(2)</p>
      <p>Full Test Set
Out-of-the Scope [Binary]
6
4
4 4
8
9 4 4
.
8
3
5 5</p>
      <p>Public leaderboard
6
.6 4</p>
    </sec>
    <sec id="sec-4">
      <title>4. Participants and Methods</title>
      <p>
        In total, 38 teams contributed with 701 valid submissions to the challenge evaluation on Kaggle.
The results on the public and private test sets (leaderboards) are displayed in Figure 3. Below
we summarize the approach of teams with published working notes. More details can be found
in the individual working notes of participants [
        <xref ref-type="bibr" rid="ref10 ref5 ref6 ref7 ref8 ref9">5, 6, 7, 8, 9, 10</xref>
        ] which passed the review process,
ensuring a suficient level of reproducibility and quality.
]
%
[40
n
o
i
s
u
fon30
C
s
e
i
c
e
p20
S
xiong [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ]: The winning submission by Xiong et al., achieving an impressive F1 score of 80.43%
on the private test set, used an ensemble of MetaFormer [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ] and ConvNext [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ] networks.
The provided metadata were utilized as inputs to the MetaFormer architecture. To battle the
long-tailed distribution of species, the authors used the Seasaw loss [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ]. Additional
improvements were achieved by test-time augmentation, adding a model trained with pseudo-labels to
the ensemble, and adding a thresholding post-processing to deal with out-of-scope observations.
USTC-IAT-United [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ]: The submission by Yu et al. used an ensemble of several CNN and
Transformer architectures: Metaformer [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ], SwinTransformer[
        <xref ref-type="bibr" rid="ref14">14</xref>
        ], EficientNet [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ], ViT (Vision
Transformer) [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ], BEiT [
        <xref ref-type="bibr" rid="ref17">17</xref>
        ]. The team scored 3rd with 79.06% of F1 score on the private test set.
In their working notes, the authors explore the impact of diferent data augmentation techniques,
model architectures, loss functions, and attention mechanisms on the classification performance.
GG [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ]: Shen et al. introduced a novel architecture CoLKANet based on VAN (Visual
Attention Network) [
        <xref ref-type="bibr" rid="ref18">18</xref>
        ] and CoAtNet [
        <xref ref-type="bibr" rid="ref19">19</xref>
        ]. It is a combination of large kernel attention and vision
transformer. The proposed CoLKANet outperforms Swin [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ] and VOLO [
        <xref ref-type="bibr" rid="ref20">20</xref>
        ] models in terms
of F1 by 2.3 and 1.9 percentage points, respectively. ConvNeXt [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ] performed similarly to
the proposed CoLKANet architecture. Furthermore, the team used techniques such as Label
Aware Smoothing [
        <xref ref-type="bibr" rid="ref21">21</xref>
        ], Pseudo labelling for tail classes and various augmentation techniques.
When TrivialAugment [
        <xref ref-type="bibr" rid="ref22">22</xref>
        ] was deployed during the middle stage of experimentation, the team
observed a rise in F1 of around 0.5%. Progressively, Random Erasing [
        <xref ref-type="bibr" rid="ref23">23</xref>
        ], CutMix [
        <xref ref-type="bibr" rid="ref24">24</xref>
        ] and
Mixup [
        <xref ref-type="bibr" rid="ref25">25</xref>
        ] were added, which helped with regularization. The final submission score was
achieved by an ensemble of five models: 2× ConvNeXt, VOLO, Swin, and CoLKANet. The novel
CoLKANet is an interesting contribution with potential outside this competition’s scope.
TeamSpirit [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ]: Fan et al., who scored sixth in the challenge with 77.58% F1 score, propose an
image classification method called Class-wise Weighted Prototype Classifier. CWPC decouples
closed-set training and open-set inference by constructing class centers from the training set
features and their prediction scores. A hard classes mining strategy and the LDAM loss [
        <xref ref-type="bibr" rid="ref26">26</xref>
        ]
were used to cope with the long-tailed distribution of species. This team encoded the metadata
using a multilingual BERT model [
        <xref ref-type="bibr" rid="ref27">27</xref>
        ] with RoBERTa [
        <xref ref-type="bibr" rid="ref28">28</xref>
        ].
      </p>
      <p>
        Stefan [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ]: Wolf and Beyerer refrained from using ensembles of multiple models, and —
for the sake of model simplicity — focused on developing a strong single-model submission.
The method is based on a Swin Transformer Large backbone [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ], a class-balancing training
scheme [
        <xref ref-type="bibr" rid="ref29">29</xref>
        ], heavy data augmentation [
        <xref ref-type="bibr" rid="ref30">30</xref>
        ] and thresholding the softmax scores to cope with
out-of-scope observations. The team scored 7th, in the challenge with 77.54% F1 score.
SSN [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ]:This team experimented with several ResNet [
        <xref ref-type="bibr" rid="ref31">31</xref>
        ], ResNeXt [
        <xref ref-type="bibr" rid="ref32">32</xref>
        ], and EficientNet [
        <xref ref-type="bibr" rid="ref33">33</xref>
        ]
architectures. For their best submission, feature vectors from two selected architectures,
EficientNetB4 and ResNeXt101, were concatenated with a categorical representation of metadata.
The resulted features were later used for training the XGBoost Ensemble Classifier [
        <xref ref-type="bibr" rid="ref34">34</xref>
        ]. An
interesting benefit of the XGBoost algorithm is that the relative importance of the ensembled
features is computed; thus, each feature might be observed and studied. With an absolute
F1 performance of 48.96%, the XGBoost algorithm with two CNN backbones poses a unique
approach for the classification, even though performing worst compared to other participants.
      </p>
    </sec>
    <sec id="sec-5">
      <title>5. Conclusions</title>
      <p>
        This paper presents an overview and results of the first edition of the FungiCLEF challenge
organized in conjunction with the Conference and Labs of the Evaluation Forum (CLEF7),
LifeCLEF8 research platform [
        <xref ref-type="bibr" rid="ref35">35</xref>
        ] and FGVC.
      </p>
      <p>
        All submissions with working notes were based on modern Convolutional Neural Network
(CNN) or transformer-inspired architectures, such as Metaformer [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ], Swin Transformer [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ],
and BEiT [
        <xref ref-type="bibr" rid="ref17">17</xref>
        ]. The best performing teams used ensembles of both CNNs and Transformers.
The winning team [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ] achieved 80.43% accuracy with a combination of ConvNext-large [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ] and
MetaFormer [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ]. The results were often improved by combining predictions belonging to the
same observation and by both training-time and test-time data augmentations.
      </p>
      <p>
        Participants experimented with a number of diferent training losses to battle the long tail
distribution and fine-grained classification with small inter-class diferences and large
intraclass diferences: besides the standard Cross Entropy loss function, we have seen successful
applications of the Seesaw loss [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ], Focal loss [
        <xref ref-type="bibr" rid="ref36">36</xref>
        ], Arcface loss [
        <xref ref-type="bibr" rid="ref37">37</xref>
        ], Sub-Center loss [
        <xref ref-type="bibr" rid="ref38">38</xref>
        ] and
Adaptive Margin [
        <xref ref-type="bibr" rid="ref39">39</xref>
        ].
      </p>
      <p>
        We were happy to see the participants experimented with diferent use of the provided
7 http://www.clef-initiative.eu/
8 http://www.lifeclef.org/
observation metadata, which often lead to improvements in the recognition scores. Besides the
probabilistic baseline published with the dataset [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ], we have seen hand-crafted encoding of the
metadata into feature vectors, as well as encoding of the metadata with a multilingual BERT
model [
        <xref ref-type="bibr" rid="ref27">27</xref>
        ] and RoBERTa [
        <xref ref-type="bibr" rid="ref28">28</xref>
        ]. The metadata were then combined with image features extracted
from a CNN or Transformer image classifier, or directly used as an input to Metaformer [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ].
      </p>
      <p>The results of participants’ comprehensive experiments with model architectures, loss
functions and usage of metadata in fine-grained image-classification will help to improve species
recognition services that aid researchers, citizen-science communities and nature enthusiasts.
As discussed in Section 3, there is still a great space for improvement in the recognition of
out-of-scope classes. Our evaluation of classification errors identified that confusion of
poisonous mushrooms for edible is much more common than confusion of edible mushrooms for
poisonous. This could be critical in applications that may afect the decision to consume a
mushroom, and presents an important aspect to address in the future work.</p>
    </sec>
    <sec id="sec-6">
      <title>Acknowledgments</title>
      <p>LP was supported by the UWB grant, project No. SGS-2022-017. LP was supported by the
Technology Agency of the Czech Republic, project No. SS05010008.
recognition, in: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern
Recognition, 2019, pp. 11947–11956.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>M.</given-names>
            <surname>Šulc</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Picek</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Matas</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Jeppesen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Heilmann-Clausen</surname>
          </string-name>
          ,
          <article-title>Fungi recognition: A practical use case</article-title>
          ,
          <source>in: Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision</source>
          ,
          <year>2020</year>
          , pp.
          <fpage>2316</fpage>
          -
          <lpage>2324</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>L.</given-names>
            <surname>Picek</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Šulc</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Matas</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Heilmann-Clausen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T. S.</given-names>
            <surname>Jeppesen</surname>
          </string-name>
          , E. Lind,
          <article-title>Automatic fungi recognition: Deep learning meets mycology</article-title>
          ,
          <source>Sensors</source>
          <volume>22</volume>
          (
          <year>2022</year>
          )
          <fpage>633</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>A.</given-names>
            <surname>Joly</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Goëau</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Kahl</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Picek</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Lorieul</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Cole</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Deneu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Servajean</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Durso</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Glotin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Planqué</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W.-P.</given-names>
            <surname>Vellinga</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Navine</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Klinck</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Denton</surname>
          </string-name>
          ,
          <string-name>
            <surname>I. Eggel</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Bonnet</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Šulc</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Hruz</surname>
          </string-name>
          , Overview of lifeclef
          <year>2022</year>
          :
          <article-title>an evaluation of machine-learning based species identification and species distribution prediction</article-title>
          ,
          <source>in: International Conference of the Cross-Language Evaluation Forum for European Languages</source>
          , Springer,
          <year>2022</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>L.</given-names>
            <surname>Picek</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Šulc</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Matas</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T. S.</given-names>
            <surname>Jeppesen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Heilmann-Clausen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Laessøe</surname>
          </string-name>
          , T. Frøslev,
          <article-title>Danish fungi 2020-not just another image recognition dataset</article-title>
          ,
          <source>in: Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision</source>
          ,
          <year>2022</year>
          , pp.
          <fpage>1525</fpage>
          -
          <lpage>1535</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>G.</given-names>
            <surname>Fan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Zining</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W.</given-names>
            <surname>Weiqiu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Yinan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Fei</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Zhicheng</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Hong</surname>
          </string-name>
          ,
          <article-title>Does closed-set training generalize to open-set recognition?</article-title>
          ,
          <source>in: Working Notes of CLEF 2022 - Conference and Labs of the Evaluation Forum</source>
          ,
          <year>2022</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>Z.</given-names>
            <surname>Xiong</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Ruan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Hu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Zhang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Zhu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Guo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W.</given-names>
            <surname>Zhu</surname>
          </string-name>
          ,
          <string-name>
            <surname>B. Han,</surname>
          </string-name>
          <article-title>An empirical study for ifne-grained fungi recognition with transformer and convnet</article-title>
          ,
          <source>in: Working Notes of CLEF 2022 - Conference and Labs of the Evaluation Forum</source>
          ,
          <year>2022</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>J.</given-names>
            <surname>Yu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Chang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Lu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Xie</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Zhang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Cai</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Du</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Wei</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Liu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Gao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Shuang</surname>
          </string-name>
          ,
          <article-title>Bag of tricks and a strong baseline for fgvc</article-title>
          ,
          <source>in: Working Notes of CLEF 2022 - Conference and Labs of the Evaluation Forum</source>
          ,
          <year>2022</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>Y.</given-names>
            <surname>Shen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Sun</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Zhu</surname>
          </string-name>
          ,
          <article-title>When large kernel meets vision transformer: A solution for snakeclef &amp; fungiclef</article-title>
          ,
          <source>in: Working Notes of CLEF 2022 - Conference and Labs of the Evaluation Forum</source>
          ,
          <year>2022</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>K.</given-names>
            <surname>Desingu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Bhaskar</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Palaniappan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E. A.</given-names>
            <surname>Chodisetty</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Bharathi</surname>
          </string-name>
          ,
          <article-title>Classification of fungi species: A deep learning based image feature extraction and gradient boosting ensemble approach</article-title>
          ,
          <source>in: Working Notes of CLEF 2022 - Conference and Labs of the Evaluation Forum</source>
          ,
          <year>2022</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>S.</given-names>
            <surname>Wolf</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Beyerer</surname>
          </string-name>
          ,
          <article-title>Transformer-based fine-grained fungi classification in an open-set scenario</article-title>
          ,
          <source>in: Working Notes of CLEF 2022 - Conference and Labs of the Evaluation Forum</source>
          ,
          <year>2022</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>Q.</given-names>
            <surname>Diao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Jiang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Wen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Sun</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Yuan</surname>
          </string-name>
          ,
          <article-title>Metaformer: A unified meta framework for ifne-grained recognition</article-title>
          ,
          <source>arXiv preprint arXiv:2203.02751</source>
          (
          <year>2022</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>Z.</given-names>
            <surname>Liu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Mao</surname>
          </string-name>
          ,
          <string-name>
            <surname>C.-Y. Wu</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          <string-name>
            <surname>Feichtenhofer</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          <string-name>
            <surname>Darrell</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          <string-name>
            <surname>Xie</surname>
          </string-name>
          ,
          <article-title>A convnet for the 2020s</article-title>
          ,
          <source>in: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition</source>
          ,
          <year>2022</year>
          , pp.
          <fpage>11976</fpage>
          -
          <lpage>11986</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <given-names>J.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W.</given-names>
            <surname>Zhang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Zang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Cao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Pang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Gong</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Chen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Liu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C. C.</given-names>
            <surname>Loy</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Lin</surname>
          </string-name>
          ,
          <article-title>Seesaw loss for long-tailed instance segmentation</article-title>
          ,
          <source>in: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition</source>
          ,
          <year>2021</year>
          , pp.
          <fpage>9695</fpage>
          -
          <lpage>9704</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <given-names>Z.</given-names>
            <surname>Liu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Lin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Cao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Hu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Wei</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Zhang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Lin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Guo</surname>
          </string-name>
          ,
          <article-title>Swin transformer: Hierarchical vision transformer using shifted windows</article-title>
          ,
          <source>in: Proceedings of the IEEE/CVF International Conference on Computer Vision</source>
          ,
          <year>2021</year>
          , pp.
          <fpage>10012</fpage>
          -
          <lpage>10022</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <given-names>M.</given-names>
            <surname>Tan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Q.</given-names>
            <surname>Le</surname>
          </string-name>
          , Eficientnet:
          <article-title>Rethinking model scaling for convolutional neural networks</article-title>
          ,
          <source>in: International conference on machine learning, PMLR</source>
          ,
          <year>2019</year>
          , pp.
          <fpage>6105</fpage>
          -
          <lpage>6114</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [16]
          <string-name>
            <given-names>A.</given-names>
            <surname>Dosovitskiy</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Beyer</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Kolesnikov</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Weissenborn</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Zhai</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Unterthiner</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Dehghani</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Minderer</surname>
          </string-name>
          , G. Heigold,
          <string-name>
            <given-names>S.</given-names>
            <surname>Gelly</surname>
          </string-name>
          , et al.,
          <article-title>An image is worth 16x16 words: Transformers for image recognition at scale</article-title>
          , arXiv preprint arXiv:
          <year>2010</year>
          .
          <volume>11929</volume>
          (
          <year>2020</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [17]
          <string-name>
            <given-names>H.</given-names>
            <surname>Bao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Dong</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Wei</surname>
          </string-name>
          , Beit:
          <article-title>Bert pre-training of image transformers</article-title>
          ,
          <source>arXiv preprint arXiv:2106.08254</source>
          (
          <year>2021</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          [18]
          <string-name>
            <surname>M.-H. Guo</surname>
          </string-name>
          ,
          <string-name>
            <surname>C.-Z. Lu</surname>
            ,
            <given-names>Z.-N.</given-names>
          </string-name>
          <string-name>
            <surname>Liu</surname>
          </string-name>
          , M.-M. Cheng, S.-M. Hu,
          <article-title>Visual attention network</article-title>
          ,
          <source>arXiv preprint arXiv:2202.09741</source>
          (
          <year>2022</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          [19]
          <string-name>
            <given-names>Z.</given-names>
            <surname>Dai</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Liu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Q. V.</given-names>
            <surname>Le</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Tan</surname>
          </string-name>
          ,
          <article-title>Coatnet: Marrying convolution and attention for all data sizes</article-title>
          ,
          <source>Advances in Neural Information Processing Systems</source>
          <volume>34</volume>
          (
          <year>2021</year>
          )
          <fpage>3965</fpage>
          -
          <lpage>3977</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          [20]
          <string-name>
            <given-names>L.</given-names>
            <surname>Yuan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Q.</given-names>
            <surname>Hou</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Jiang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Feng</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Yan</surname>
          </string-name>
          ,
          <article-title>Volo: Vision outlooker for visual recognition</article-title>
          ,
          <source>arXiv preprint arXiv:2106.13112</source>
          (
          <year>2021</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          [21]
          <string-name>
            <given-names>Z.</given-names>
            <surname>Zhong</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Cui</surname>
          </string-name>
          , S. Liu,
          <string-name>
            <given-names>J.</given-names>
            <surname>Jia</surname>
          </string-name>
          ,
          <article-title>Improving calibration for long-tailed recognition</article-title>
          ,
          <source>in: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition</source>
          ,
          <year>2021</year>
          , pp.
          <fpage>16489</fpage>
          -
          <lpage>16498</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          [22]
          <string-name>
            <given-names>S. G.</given-names>
            <surname>Müller</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Hutter</surname>
          </string-name>
          , Trivialaugment:
          <article-title>Tuning-free yet state-of-the-art data augmentation</article-title>
          ,
          <source>in: Proceedings of the IEEE/CVF International Conference on Computer Vision</source>
          ,
          <year>2021</year>
          , pp.
          <fpage>774</fpage>
          -
          <lpage>782</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          [23]
          <string-name>
            <given-names>Z.</given-names>
            <surname>Zhong</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Zheng</surname>
          </string-name>
          , G. Kang,
          <string-name>
            <given-names>S.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Yang</surname>
          </string-name>
          ,
          <article-title>Random erasing data augmentation</article-title>
          ,
          <source>in: Proceedings of the AAAI conference on artificial intelligence</source>
          , volume
          <volume>34</volume>
          ,
          <year>2020</year>
          , pp.
          <fpage>13001</fpage>
          -
          <lpage>13008</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          [24]
          <string-name>
            <given-names>S.</given-names>
            <surname>Yun</surname>
          </string-name>
          , D. Han,
          <string-name>
            <given-names>S. J.</given-names>
            <surname>Oh</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Chun</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Choe</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Yoo</surname>
          </string-name>
          , Cutmix:
          <article-title>Regularization strategy to train strong classifiers with localizable features</article-title>
          ,
          <source>in: Proceedings of the IEEE/CVF international conference on computer vision</source>
          ,
          <year>2019</year>
          , pp.
          <fpage>6023</fpage>
          -
          <lpage>6032</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref25">
        <mixed-citation>
          [25]
          <string-name>
            <given-names>H.</given-names>
            <surname>Zhang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Cisse</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y. N.</given-names>
            <surname>Dauphin</surname>
          </string-name>
          ,
          <string-name>
            <surname>D.</surname>
          </string-name>
          <article-title>Lopez-Paz, mixup: Beyond empirical risk minimization</article-title>
          ,
          <source>arXiv preprint arXiv:1710.09412</source>
          (
          <year>2017</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref26">
        <mixed-citation>
          [26]
          <string-name>
            <given-names>K.</given-names>
            <surname>Cao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Wei</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Gaidon</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Arechiga</surname>
          </string-name>
          , T. Ma,
          <article-title>Learning imbalanced datasets with labeldistribution-aware margin loss</article-title>
          ,
          <source>Advances in neural information processing systems</source>
          <volume>32</volume>
          (
          <year>2019</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref27">
        <mixed-citation>
          [27]
          <string-name>
            <given-names>J.</given-names>
            <surname>Devlin</surname>
          </string-name>
          , M.-
          <string-name>
            <given-names>W.</given-names>
            <surname>Chang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Lee</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Toutanova</surname>
          </string-name>
          , Bert:
          <article-title>Pre-training of deep bidirectional transformers for language understanding</article-title>
          , arXiv preprint arXiv:
          <year>1810</year>
          .
          <volume>04805</volume>
          (
          <year>2018</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref28">
        <mixed-citation>
          [28]
          <string-name>
            <given-names>Y.</given-names>
            <surname>Liu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Ott</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Goyal</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Du</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Joshi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Chen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>O.</given-names>
            <surname>Levy</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Lewis</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Zettlemoyer</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Stoyanov</surname>
          </string-name>
          ,
          <article-title>Roberta: A robustly optimized bert pretraining approach</article-title>
          , arXiv preprint arXiv:
          <year>1907</year>
          .
          <volume>11692</volume>
          (
          <year>2019</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref29">
        <mixed-citation>
          [29]
          <string-name>
            <given-names>A.</given-names>
            <surname>Gupta</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Dollar</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Girshick</surname>
          </string-name>
          ,
          <article-title>Lvis: A dataset for large vocabulary instance segmentation</article-title>
          ,
          <source>in: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition</source>
          ,
          <year>2019</year>
          , pp.
          <fpage>5356</fpage>
          -
          <lpage>5364</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref30">
        <mixed-citation>
          [30]
          <string-name>
            <given-names>E. D.</given-names>
            <surname>Cubuk</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Zoph</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Shlens</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Q. V.</given-names>
            <surname>Le</surname>
          </string-name>
          ,
          <article-title>Randaugment: Practical automated data augmentation with a reduced search space</article-title>
          ,
          <source>in: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition workshops</source>
          ,
          <year>2020</year>
          , pp.
          <fpage>702</fpage>
          -
          <lpage>703</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref31">
        <mixed-citation>
          [31]
          <string-name>
            <given-names>K.</given-names>
            <surname>He</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Zhang</surname>
          </string-name>
          , S. Ren,
          <string-name>
            <given-names>J.</given-names>
            <surname>Sun</surname>
          </string-name>
          ,
          <article-title>Deep Residual Learning for Image Recognition</article-title>
          ,
          <source>in: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR)</source>
          ,
          <year>2016</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref32">
        <mixed-citation>
          [32]
          <string-name>
            <given-names>S.</given-names>
            <surname>Xie</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Girshick</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Dollar</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Tu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>He</surname>
          </string-name>
          ,
          <article-title>Aggregated Residual Transformations for Deep Neural Networks</article-title>
          ,
          <source>in: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR)</source>
          ,
          <year>2017</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref33">
        <mixed-citation>
          [33]
          <string-name>
            <given-names>M.</given-names>
            <surname>Tan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Q.</given-names>
            <surname>Le</surname>
          </string-name>
          , Eficientnet:
          <article-title>Rethinking model scaling for convolutional neural networks</article-title>
          ,
          <source>in: International Conference on Machine Learning, PMLR</source>
          ,
          <year>2019</year>
          , pp.
          <fpage>6105</fpage>
          -
          <lpage>6114</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref34">
        <mixed-citation>
          [34]
          <string-name>
            <given-names>T.</given-names>
            <surname>Chen</surname>
          </string-name>
          , C. Guestrin,
          <article-title>XGBoost: A scalable tree boosting system</article-title>
          ,
          <source>in: Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, KDD '16</source>
          ,
          <string-name>
            <surname>ACM</surname>
          </string-name>
          , New York, NY, USA,
          <year>2016</year>
          , pp.
          <fpage>785</fpage>
          -
          <lpage>794</lpage>
          . URL: http://doi.acm.
          <source>org/10</source>
          .1145/ 2939672.2939785. doi:
          <volume>10</volume>
          .1145/2939672.2939785.
        </mixed-citation>
      </ref>
      <ref id="ref35">
        <mixed-citation>
          [35]
          <string-name>
            <given-names>A.</given-names>
            <surname>Joly</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Goëau</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Kahl</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Picek</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Lorieul</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Cole</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Deneu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Servajean</surname>
          </string-name>
          , R. Ruiz De Castañeda, I. Bolon,
          <string-name>
            <given-names>H.</given-names>
            <surname>Glotin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Planqué</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W.-P.</given-names>
            <surname>Vellinga</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Dorso</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Klinck</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Denton</surname>
          </string-name>
          ,
          <string-name>
            <surname>I. Eggel</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Bonnet</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Müller</surname>
          </string-name>
          , Overview of lifeclef
          <year>2021</year>
          :
          <article-title>a system-oriented evaluation of automated species identification and species distribution prediction</article-title>
          ,
          <source>in: Proceedings of the Twelfth International Conference of the CLEF Association (CLEF</source>
          <year>2021</year>
          ),
          <year>2021</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref36">
        <mixed-citation>
          [36]
          <string-name>
            <surname>T.-Y. Lin</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          <string-name>
            <surname>Goyal</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          <string-name>
            <surname>Girshick</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          <string-name>
            <surname>He</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          <string-name>
            <surname>Dollár</surname>
          </string-name>
          ,
          <article-title>Focal loss for dense object detection</article-title>
          ,
          <source>in: Proceedings of the IEEE international conference on computer vision</source>
          ,
          <year>2017</year>
          , pp.
          <fpage>2980</fpage>
          -
          <lpage>2988</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref37">
        <mixed-citation>
          [37]
          <string-name>
            <given-names>J.</given-names>
            <surname>Deng</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Guo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Xue</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Zafeiriou</surname>
          </string-name>
          , Arcface:
          <article-title>Additive angular margin loss for deep face recognition</article-title>
          ,
          <source>in: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition</source>
          ,
          <year>2019</year>
          , pp.
          <fpage>4690</fpage>
          -
          <lpage>4699</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref38">
        <mixed-citation>
          [38]
          <string-name>
            <given-names>J.</given-names>
            <surname>Deng</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Guo</surname>
          </string-name>
          , T. Liu,
          <string-name>
            <given-names>M.</given-names>
            <surname>Gong</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Zafeiriou</surname>
          </string-name>
          ,
          <article-title>Sub-center arcface: Boosting face recognition by large-scale noisy web faces</article-title>
          ,
          <source>in: European Conference on Computer Vision</source>
          , Springer,
          <year>2020</year>
          , pp.
          <fpage>741</fpage>
          -
          <lpage>757</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref39">
        <mixed-citation>
          [39]
          <string-name>
            <given-names>H.</given-names>
            <surname>Liu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Zhu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Lei</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S. Z.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <article-title>Adaptiveface: Adaptive margin and sampling for face</article-title>
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>