<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Overview of LifeCLEF Plant Identi cation task 2019: diving into data de cient tropical countries</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Herve Goeau</string-name>
          <email>herve.goeau@cirad.fr</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Pierre Bonnet</string-name>
          <email>pierre.bonnet@cirad.fr</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Alexis Joly</string-name>
          <email>alexis.joly@inria.fr</email>
          <xref ref-type="aff" rid="aff2">2</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>AMAP, Univ Montpellier, CIRAD, CNRS, INRA, IRD</institution>
          ,
          <addr-line>Montpellier</addr-line>
          ,
          <country country="FR">France</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>CIRAD, UMR AMAP</institution>
          ,
          <country country="FR">France</country>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>Inria ZENITH team</institution>
          ,
          <country country="FR">France</country>
        </aff>
        <aff id="aff3">
          <label>3</label>
          <institution>LIRMM</institution>
          ,
          <addr-line>Montpellier</addr-line>
          ,
          <country country="FR">France</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>Automated identi cation of plants has improved considerably thanks to the recent progress in deep learning and the availability of training data. However, this profusion of data only concerns a few tens of thousands of species, while the planet has nearly 369K. The LifeCLEF 2019 Plant Identi cation challenge (or "PlantCLEF 2019") was designed to evaluate automated identi cation on the ora of data de cient regions. It is based on a dataset of 10K species mainly focused on the Guiana shield and the Northern Amazon rainforest, an area known to have one of the greatest diversity of plants and animals in the world. As in the previous edition, a comparison of the performance of the systems evaluated with the best tropical ora experts was carried out. This paper presents the resources and assessments of the challenge, summarizes the approaches and systems employed by the participating research groups, and provides an analysis of the main outcomes.</p>
      </abstract>
      <kwd-group>
        <kwd>LifeCLEF</kwd>
        <kwd>PlantCLEF</kwd>
        <kwd>plant</kwd>
        <kwd>expert</kwd>
        <kwd>tropical ora</kwd>
        <kwd>Amazon rainforest</kwd>
        <kwd>Guiana Shield leaves</kwd>
        <kwd>species identi cation</kwd>
        <kwd>ne-grained classi cation</kwd>
        <kwd>evaluation</kwd>
        <kwd>benchmark</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>
        Automated identi cation of plants and animals has improved considerably in
the last few years. In the scope of LifeCLEF 2017 [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ] in particular, we measured
impressive identi cation performance achieved thanks to recent deep learning
models (e.g. up to 90 % classi cation accuracy over 10K species). Moreover, the
previous edition in 2018 showed that automated systems are not so far from the
human expertise [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ]. However, these 10K species are mostly living in Europe
and North America and only represent the tip of the iceberg. The vast majority
of the species in the world (about 369K species) actually lives in data de cient
countries in terms of collected observations and the performance of
state-of-theart machine learning algorithms on these species is unknown and presumably
much lower.
      </p>
      <p>
        The LifeCLEF 2019 Plant Identi cation challenge (or "PlantCLEF 2019")
presented in this paper was designed to evaluate automated identi cation on the
ora of such data de cient regions. The challenge was based on a new dataset
of 10K species mainly focused on the Guiana shield and the Northern Amazon
rainforest, an area known to have one of the greatest diversity of plants and
animals in the world. The average number of images per species in that new
dataset is signi cantly lower than the last dataset used in the previous edition
of PlantCLEF[
        <xref ref-type="bibr" rid="ref6">6</xref>
        ] (about 1 vs. 3), and many species contain very few images
or may even contain only one image. To make it worse, because of the lack of
illustrations of these species in the world, the data collected as a training set
su ers from several properties that do not facilitate the task: many images are
duplicated across di erent species leading to identi cation errors, some images
do not represent plants, and many images are drawings or digitalized herbarium
sheets that may be visually far from eld plant images. The test set, on the
other hand, does not present this type of noisy and biased content since it is
composed only of expert data identi ed in the eld with certainty. As these data
have never been published before, there is also no risk that they belong to the
training set.
      </p>
      <p>As in the 2018-th edition of PlantCLEF, a comparison of the performance of
the systems evaluated with the best tropical ora experts was carried out for
PlantCLEF 2019. In total, 26 deep-learning systems implemented by 6 di
erent research teams were evaluated with regard to the annotations of 5 experts
of the targeted tropical ora. This paper presents more precisely the resources
and assessments of the challenge, summarizes the approaches and systems
employed by the participating research groups, and provides an analysis of the main
outcomes.
2
2.1</p>
    </sec>
    <sec id="sec-2">
      <title>Dataset</title>
      <sec id="sec-2-1">
        <title>Training set</title>
        <p>
          We provided a new training data set of 10K species mainly focused on the Guiana
shield and the Amazon rainforest, known to be one of the largest collection of
living plants and animal species in the world (see gure1). As for the two
previous years, this training data was mainly aggregated by bringing together images
from complementary types of available sources, including expert data from the
international platform Encyclopedia of Life (EoL5) and images automatically
retrieved from the web using industrial search engines (Bing and Google) that
were queried with the binomial Latin name of the targeted species. Details
numbers of images and species per sub-dataset are provided in Table 1. A large part
of the images collected come from trusted websites, but they also contain a high
5 https://eol.org/
level of noise. It has been shown in previous editions of LifeCLEF, however, that
training deep learning models on such raw big data can be as e ective as training
models on cleaner but smaller expert data [
          <xref ref-type="bibr" rid="ref4">4</xref>
          ], [
          <xref ref-type="bibr" rid="ref5">5</xref>
          ]. The main objective of this
new study was to evaluate whether this inexpensive methodology is applicable
to the case of tropical oras that are much less observed and therefore much less
present on the web.
        </p>
        <p>One of the consequences of this change in target ora is that the average
number of images per species is much lower (about 1 vs. 3). Many species
contain only a few images and some of them even contain only 1 image. On the
other hand, few common species are still associated with several hundreds of
pictures mainly coming from EOL. In addition to this scarcity of data, the use
of web search engines to collect the images generates several types of noise:
Duplicate images and taxonomic noise: web search engines often return
the same image several times for di erent species. This typically happens when
an image is displayed in a web page that contains a list of several species. For
instance, in the Wikipedia web page of a genus, all child species are usually
listed but only a few of them are illustrated with a picture. As a consequence,
the available pictures are often retrieved for several species of the genus and not
only the correct one. We call this phenomenon taxonomic noise. The resulting
label errors are problematic but they can paradoxically have a certain usefulness
during the training. Indeed, species of the same genus often have a number of
common morphological traits, which can be visually similar. For the less
illustrated species, it is therefore often more cost-e ective to keep images of closely
related species rather than not having images at all. Technically, to help
managing this taxonomic noise, the duplicate images were replicated in the directories
of each species they belong to, but the same image name was used everywhere.
Non-photographic images of the plant (herbarium, drawings): because
of data scarcity, it often occurs that the only images available on the web for
a given species are not photographs but rather digitized herbarium sheets or
drawings from academic books. This typically happens for very rare or poorly
observed species. For the most extreme cases, these old testimonies, sometimes
more than a century old, are the only existing data. The usefulness of these
data for learning is again ambivalent. They can be visually very di erent from
a photograph of the plant but they still contain a rich information about the
appearance of the species.</p>
        <p>Atypical photographs of the plant: a number of images are related to the
target species but do not directly represent it. Typically, it can be a landscape
related to the habitat of the species, a photograph of the dissection of plant
organs, a handful of seeds on a blank sheet of paper, or a microscopic view.
Non-plant images: search engines sometimes return images that are not plants
and that have only a very indirect link with the target species: medicines,
animals, mushrooms, botanists, handcrafted-objects, logos, maps, ethnic (food,
craftsmanship), etc.
2.2</p>
      </sec>
      <sec id="sec-2-2">
        <title>Test set</title>
        <p>
          Unlike the training set, the labels of the images in the test set are of very high
quality to ensure the reliability of the evaluation. It is composed of 742 plant
observations, all of which have been identi ed in the eld by one of the best
experts on the ora in question. Marie-Franoise Prevost [
          <xref ref-type="bibr" rid="ref3">3</xref>
          ], the author of this
test set, is a french botanist, who has spent more than 40 years to study French
Guiana ora. She has an extensive eld work experience, and she has contributed
a lot to improve our current knowledge of this ora, by collecting numerous
herbarium specimens of great quality. Ten plant epithets have been dedicated
to her, which illustrates the acknowledgement of the taxonomists community to
her contribution to the tropical botany. For the re-annotation experiment by the
other 5 human experts, only a sub-set of 117 of these observations was used.
3
        </p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>Task Description</title>
      <p>The goal of the task was to identify the correct species of the 742 plants of the
test set. For every plant, the evaluated systems had to return a list of species,
ranked without ex-aequo. Each participating group was allowed to submit up to
10 run les built from di erent methods or systems (a run le is a formatted
text le containing the species predictions for all test items).</p>
      <p>The goal of the task was exactly the same for the 5 human experts, except
that we restricted the number of their responses to 3 species per test item to
reduce their e ort. The list of possible species was provided to them.</p>
      <p>The main evaluation metric for both the systems and the humans was the
Top1 accuracy, i.e. the percentage of test items for which the right species is
predicted in rst position. As complementary metrics, we also measured the
Top3 accuracy, Top5 accuracy and the Mean Reciprocal Rank (MRR), de ned
as the mean of the multiplicative inverse of the rank of the correct answer:
Q
MRR : 1 X</p>
      <p>Q</p>
      <p>1
q=1 rankq
4</p>
    </sec>
    <sec id="sec-4">
      <title>Participants and methods</title>
      <p>
        167 participants registered for the PlantCLEF challenge 2019 and downloaded
the data set, but only 6 research groups succeeded in submitting run les.
Details of the methods are developed in the individual working notes of most of
the participants (Holmes [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ], CMP [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ], MRIM-LIG [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]). We provide hereafter
a synthesis of the runs of the best performing teams:
CMP, Dept. of Cybernetics, Czech Technical University in Prague,
Czech Republic, 7 runs, [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ]: this team used an ensemble of 5 Convolutional
Neural Networks (CNNs) based on 2 state-of-the-art architectures
(InceptionResNet-v2 and Inception-v4). The CNNs were initialized with weights pre-trained
on the dataset used during ExpertCLEF2018 [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ] and then ne-tuned with di
erent hyper-parameters and with the use of data augmentation (random horizontal
ip, color distortions and random crops). Further performance improvements
were achieved by adjusting the CNN predictions according to the estimated
change of the classes distribution between the training set and the test set. The
running averages of the learned weights were used as the nal model and the
test was also processed with data augmentation (3 central crops at various scales
and their mirrored version). Regarding the data used for the training, the team
decided to remove all images estimated to be non oral data based on the
classication of a dedicated VGG net. This had the e ect of removing 300 species and
about 5,500 pictures. An important point is that additional training images were
downloaded from the GBIF platform6, in order to ll the missing species and
enrich the other species. This increased considerably the training dataset with
238,009 new pictures of good quality totalling then 666,711 pictures.
Unfortunately, the participants did not submit a run without the use of this new training
data, so that it not possible to measure accurately the impact of this addition.
The participant's focus was rather on evaluating di erent prior distribution of
the classes (Uniform, Maximum Likelihood Estimate, Maximum a Posteriori)
6 https://www.gbif.org/
to modify predictions in order to soften the impact of the highly unbalanced
distribution of the training set.
      </p>
      <p>
        Holmes, Neuon AI, Malaysia, 3 runs, [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]: This team used the same CNN
architectures than the CMP team (Inception-v4 and Inception-ResNet-v2). In
their case, however, the CNNs were initialized with weights pre-trained on
ImageNet rather than ExpertCLEF2018 [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ] and they did not use any additional
training data. An original feature of their system was the introduction of a
multitask classi cation layer, allowing to classify the images at the genus and family
levels in addition to the species. Complementary, this team spent some e orts
to clean the dataset. The 154,627 duplicate pictures were removed and they
automatically removed 15,196 additional near-duplicates based on a cosine
similarity in the feature space of the last layer of Inception-V4. Finally, they removed
13,341 non plant images automatically detected by using a plant vs. non plant
binary classi er (also based on Inception-V4). Overall, the whole training set
was decreased by nearly 42% of the images (whereas the CMP team increased
it by nearly 53%).
      </p>
      <p>
        Cross-validation experiments conducted by the authors [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ] show that removing
the duplicates and near-duplicates may allow to gain 4 points of accuracy. In
contrast, removing the non plant pictures does not provide any improvement.
The introduction of the multi-task classi er at the di erent taxonomic levels is
shown to provide one to two more points of accuracy.
      </p>
      <p>
        MRIM, LIG, France, 10 runs, [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]: this team based all the runs on DenseNet,
another state-of-art CNN architecture which has the advantage to have a
relatively low number of parameters compared to other popular CNNs. They
increased the initial model with a non-local block with the idea to model
interpixels correlations from dierent positions in the feature maps. They used a set
of data augmentation processes including random resize, random crop, random
ip and random brightness and contrast changes. To compensate the class
imbalance, they made use of oversampling and under-sampling strategies. Their
cross-validation experiments did show some signi cant improvements but these
bene ts were not con rmed on the nal test set, probably because of the
crossvalidation methodology (based on a subset of only 500 species among the most
populated ones).
      </p>
      <p>The three other remaining teams did not provide an extended description of their
system. According to the short description they provided, the datvo06 team from
Vietnam (1 run) used a similar approach to the MRIM team (DenseNet), the
Leowin team from India (2 runs) used Random Forest Boosted on the features
extracted from a ResNet, and the MLRG SSN team from India (3 runs), used a
ResNet 50 trained for 100 epochs with strati cation of batches.</p>
    </sec>
    <sec id="sec-5">
      <title>Results</title>
      <p>
        The detailed results of the evaluation are reported in Table 2. Figure 2 gives
a more graphical view of the comparison between the human experts and the
evaluated systems (on the dedicated test subset). Figure 3, on the other hand,
provides a comparison of the evaluated systems on the whole test set.
The main outcomes we can derive from that results are the following ones:
A very di cult task, even for experts: none of the botanist correctly
identi ed all observations. The top-1 accuracy of the experts is in the range
0:154 0:675 with a median value of 0:376. It illustrates the di culty of the
task, especially when reminding that the experts were authorized to use any
external resource to complete the task, with Flora books in particular. It shows
that a large part of the observations in the test may not contain enough
information to be identi ed with high con dence. The complete identi cation may
actually rely on other information such as the root shape, the smell of the plant,
type of habitat, the feeling of touch from certain parts, or the presence of
organs or feature that were not photographed of visible on the pictures. Only two
experts with an exceptional eld expertise were able to correctly identify more
than 60% of the observations. The other ones correctly identi ed less than 40%.
Tropical ora is much more di cult to identify. Results are signi cantly
lower than the previous edition of LifeCLEF con rming the assumption that
the tropical ora is inherently more di cult to identify than the more
generalist ora. The best accuracy obtained by an expert is 0.675 for the tropical
ora whereas it was 0.96 for the ora of temperate regions considered in 2018[
        <xref ref-type="bibr" rid="ref6">6</xref>
        ].
Comparison of medians (0.376 vs 0.8) and minimums (0.154 vs 0.613) over the
two years further highlights the gap. This can be explained by the fact that (i)
there is in general much more diversity in tropical regions compare to
temperate ones, for a same reference surface, (ii) tropical plants in high rainforests, are
much less accessible to humans who have much more di culties to improve their
knowledge on these ecosystems, (iii) the volume of available resources
(including herbarium specimens, books, web sites) is much less important on that oras.
Deep learning algorithms were defeated by far by the best experts. The
best automated system is half as good as the best expert with a gap of 0.365,
whereas last year the gap was only 0.12. Moreover, there is a strong
disparity in results between participants despite the use of popular and recent CNNs
(DensetNet, ResNet, Inception-ResNet-V2, Inception-V4), while during the last
four PlantCLEF editions the homogenization of high results forming a "skyline"
had often been observed. These di erences in accuracy can be explained in part
by the way participants managed the training set. Although previous work had
shown the e ectiveness of training from noisy data [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ],[
        <xref ref-type="bibr" rid="ref4">4</xref>
        ], most teams considered
that the training dataset was too noisy and too imbalanced. They made
consistent e orts for removing duplicates pictures (Holmes), for removing non plant
pictures (Holmes, CMP), for adding new pictures (CMP), or for reducing the
classes imbalance with smoothed re-sampling and other data sampling schemes
(MRIM). None of them attempted to simply run one of their models on the raw
data (as usually done in previous years). So that it is is not possible to conclude
on the bene t of such ltering methods this year.
6
      </p>
    </sec>
    <sec id="sec-6">
      <title>Complementary results</title>
      <p>Extending the training set with herbarium data may provide signi
cant improvements. As mentioned in section 4, the CMP team considerably
extended the training set by adding more than 238k images of the GBIF
platform, the vast majority of these images coming from the digitization of herbarium
collections. The performance of their system during the o cial evaluation was
not that good, but unfortunately, this was mainly due to a bug in the formatting
of their submissions. The corrected version of their submissions were evaluated
after the end of the o cial challenge and did achieve a top-1 accuracy of 41%,
10 points more than the best model of the evaluation and 3 points more than
the third human expert. It is likely that this high performance gain is mostly
due to the use of the additional training data. This opens up very interesting
perspectives for the massive use of herbarium collections, which are being
digitized at a very fast pace worldwide.</p>
      <p>Estimation of class purity in the training dataset. In order to evaluate
how the di erent types of noise (herbariums &amp; drawings, other plant pictures and
non plant pictures ) a ect the training set, we computed some statistics based on
a semi-automated classi cation of the training set. More precisely, we repeatedly
annotated some images by hand, trained dedicated classi ers and predicted the
missing labels. Figure 4 (left side) displays the average proportion of each noise
as a function of the number of training images per species. Complementary, on
the right side, we display the average proportion of duplicates still as a function
of the number of training images per species. If we look at the herbariums &amp;
drawings, it can be seen that their proportion signi cantly decreases up to 150
images per species and then increases again for the most populated species. This
evolution has to be correlated with the average proportion of web vs EoL data
in the training set. Indeed, the number of web images per species was limited to
150 so that the species with large amounts of training data are mainly illustrated
by EoL images. This data is highly trusted in terms of species labels but the
gure shows that it contains a high proportion of herbarium sheets. Concerning
the two other types of noise, the proportion of other plant pictures is globally
increasing likely because this type of picture is also more represented in the EoL
data. In contrast, the proportion of non plant pictures is decreasing above 150
images per species which means that this kind of noise is lower in EoL.
Concerning the duplicates, it can be seen that their proportion is strongly decreasing
with the number of images per species. This means that the absolute number
of duplicates is quite stable over all species but that its relative impact is much
more important for species having scarce training data.</p>
      <p>Average accuracy of the best evaluated systems as a function of
the number of training images per species To analyze the impact of the
di erent types of noise on the prediction performance, we computed the
specieswise performance of a fusion of the best run of each team (focusing on the three
teams who obtained the best results CMP, Holmes and MRIM). This was done
by rst averaging the scores returned by each system and then by computing
the mean rank of the correct answer over all the test images of a given species.
Figure 5 displays this mean rank for each species (using a color code, see legend)
as a function of the number of training images for this species and the estimated
proportions of the di erent noises. The following conclusions can be derive from
these graphs:
1. the more images, the better the performance: without surprise, all graphs
show that the mean rank of the correct species improves with the number
of images.
2. the presence of non plant pictures only a ects species with few training data
(as shown in the second sub-graph). Well populated species seem to be well
recognized even with a very high proportion of non plant pictures.
3. the presence of herbarium data and other plant pictures is not conclusive:
as discussed in section 2.1, these two types of contents are ambivalent. They
may bring some useful information but they may also disrupt the model.</p>
      <p>The graphs of 5 do not allow to conclude on this point.
4. a too high proportion of duplicates (above 20%) signi cantly degrades the
results, even for species having between 30 and 200 images.
7</p>
    </sec>
    <sec id="sec-7">
      <title>Conclusion</title>
      <p>This paper presented the overview and the results of the LifeCLEF 2019 plant
identi cation challenge following the eight previous editions conducted within
CLEF evaluation forum. The results reveal that the identi cation performance
on Amazonian plants is considerably lower than the one obtained on temperate
plants of Europe and North America. The performance of convolutional neural
networks fall due to the very low number of training images for most species
and the higher degree of noise that is occurring in such data. Human experts
themselves have much more di culty identifying the tropical specimens
evaluated this year compared to the more common species considered in previous
years. This shows that the small amount of data available for these species is
correlated with the lowest overall knowledge we have of them. An interesting
perspective for the future is to consider herbarium data as one solution to
overcome the lack of data. This material, collected by botanists for centuries, is the
most comprehensive knowledge that we have to date for a very large number
of species on earth. Thus, their massive ongoing digitization represents a great
opportunity. Nevertheless, they are very di erent from eld photographs, their
use will thus pose challenging domain adaptation problems.</p>
      <p>Acknowledgements We would like to thank very warmly Julien Engel, Remi
Girault, Jean-Francois Molino and the two other expert botanists who agreed to
participate in the task on plant identi cation. We also we would like to thank the
University of Montpellier and the Floris'Tic project (ANRU) who contributed
to the funding of the 2019-th edition of LifeCLEF.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Chulif</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Jing Heng</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Wei</surname>
            <given-names>Chan</given-names>
          </string-name>
          ,
          <string-name>
            <surname>T.</surname>
          </string-name>
          , Al Monnaf,
          <string-name>
            <given-names>M.A.</given-names>
            ,
            <surname>Chang</surname>
          </string-name>
          ,
          <string-name>
            <surname>Y.L.</surname>
          </string-name>
          :
          <article-title>Plant identi cation on amazonian and guiana shield ora: Neuon submission to lifeclef 2019 plant</article-title>
          . In: CLEF (Working Notes) (
          <year>2019</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <given-names>Dat</given-names>
            <surname>Nguyen Thanh</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.Q.</given-names>
            ,
            <surname>Goeuriot</surname>
          </string-name>
          ,
          <string-name>
            <surname>L.</surname>
          </string-name>
          :
          <article-title>Non-local densenet for plant clef 2019 contest</article-title>
          . In: CLEF (Working Notes) (
          <year>2019</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Delprete</surname>
            ,
            <given-names>P.G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Feuillet</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          :
          <article-title>Marie-francoise prevost fanchon(</article-title>
          <year>1941</year>
          {
          <year>2013</year>
          ).
          <source>Taxon</source>
          <volume>62</volume>
          (
          <issue>2</issue>
          ),
          <volume>419</volume>
          {
          <fpage>419</fpage>
          (
          <year>2013</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4. Goeau, H.,
          <string-name>
            <surname>Bonnet</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Joly</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          :
          <article-title>Plant identi cation based on noisy web data: the amazing performance of deep learning (lifeclef 2017)</article-title>
          .
          <source>In: Working Notes of CLEF</source>
          <year>2017</year>
          (
          <article-title>Cross Language Evaluation Forum) (</article-title>
          <year>2017</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5. Goeau, H.,
          <string-name>
            <surname>Bonnet</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Joly</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          :
          <article-title>Overview of expertlifeclef 2018: how far automated identi cation systems are from the best experts ?</article-title>
          <source>In: CLEF working notes 2018</source>
          (
          <year>2018</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6. Goeau, H.,
          <string-name>
            <surname>Bonnet</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Joly</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          :
          <article-title>Overview of expertlifeclef 2018: how far automated identi cation systems are from the best experts? lifeclef experts vs. machine plant identi cation task 2018</article-title>
          .
          <source>In: CLEF</source>
          <year>2018</year>
          (
          <year>2018</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>Joly</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          , Goeau, H.,
          <string-name>
            <surname>Botella</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Glotin</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bonnet</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Vellinga</surname>
            ,
            <given-names>W.P.</given-names>
          </string-name>
          , Muller, H.:
          <article-title>Overview of lifeclef 2018: a large-scale evaluation of species identi cation and recommendation algorithms in the era of ai</article-title>
          . In: Jones,
          <string-name>
            <given-names>G.J.</given-names>
            ,
            <surname>Lawless</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            ,
            <surname>Gonzalo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            ,
            <surname>Kelly</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            ,
            <surname>Goeuriot</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            ,
            <surname>Mandl</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            ,
            <surname>Cappellato</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            ,
            <surname>Ferro</surname>
          </string-name>
          , N. (eds.) CLEF:
          <article-title>CrossLanguage Evaluation Forum for European Languages</article-title>
          .
          <source>Experimental IR Meets Multilinguality, Multimodality, and Interaction</source>
          , vol.
          <source>LNCS</source>
          . Springer, Avigon, France (
          <year>Sep 2018</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>Joly</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          , Goeau, H.,
          <string-name>
            <surname>Glotin</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Spampinato</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bonnet</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Vellinga</surname>
            ,
            <given-names>W.P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lombardo</surname>
            ,
            <given-names>J.C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Planque</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Palazzo</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          , Muller, H.:
          <article-title>Lifeclef 2017 lab overview: multimedia species identi cation challenges</article-title>
          .
          <source>In: International Conference of the CrossLanguage Evaluation Forum for European Languages</source>
          . pp.
          <volume>255</volume>
          {
          <fpage>274</fpage>
          . Springer (
          <year>2017</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <surname>Krause</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sapp</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Howard</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zhou</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Toshev</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Duerig</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Philbin</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>FeiFei</surname>
          </string-name>
          , L.:
          <article-title>The unreasonable e ectiveness of noisy data for ne-grained recognition</article-title>
          .
          <source>In: European Conference on Computer Vision</source>
          . pp.
          <volume>301</volume>
          {
          <fpage>320</fpage>
          . Springer (
          <year>2016</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <surname>Picek</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sulc</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Matas</surname>
          </string-name>
          , J.:
          <article-title>Recognition of the amazonian ora by inception networks with test-time class prior estimation</article-title>
          .
          <source>In: CLEF (Working Notes)</source>
          (
          <year>2019</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <string-name>
            <surname>Sulc</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Picek</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Matas</surname>
          </string-name>
          , J.:
          <article-title>Plant recognition by inception networks with testtime class prior estimation</article-title>
          .
          <source>In: Working Notes of CLEF</source>
          <year>2018</year>
          (
          <article-title>Cross Language Evaluation Forum) (</article-title>
          <year>2018</year>
          )
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>