<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Say No to the Poisonous Fungi: An Efective Strategy for Reducing 0-1 Cost in FungiCLEF2024</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Bao-Feng Tan</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Yang-Yang Li</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Peng Wang</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Lin Zhao</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Xiu-Shen Wei</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>School of Computer Science and Engineering, Nanjing University of Science and Technology</institution>
          ,
          <addr-line>Nanjing</addr-line>
          ,
          <country country="CN">China</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>School of Computer Science and Engineering, and Key Laboratory of New Generation Artificial Intelligence Technology and Its Interdisciplinary Applications, Southeast University</institution>
          ,
          <addr-line>Nanjing</addr-line>
          ,
          <country country="CN">China</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>The FungiCLEF2024 competition endeavors to precisely identify fungi species leveraging both metadata and image analysis. Pivotal to the success of this competition are two crucial evaluation metrics: minimizing the error rate and the 0-1 cost loss resulting from misclassification. To reduce the identification error rate, we introduce a Dynamic MLP framework, drawing inspiration from [1]. This approach efectively integrates image and metadata embeddings through recursive blocks, utilizing matrix multiplication for deep fusion of information. To further address the issue of 0-1 cost, we devise a novel probability-based screening strategy, which initially consolidates poisonous fungi categories into a single class, then employs marginal expected loss and a threshold parameter  to optimize the recall rate for poisonous species. These approaches significantly reduce the error rate and 0-1 cost associated with misclassification and achieve a score of 0.5548 on the private leaderboard, securing the third-place ranking. The code is available at https://github.com/bftan1949/FungiCLEF2024.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;Fine-grained image recognition</kwd>
        <kwd>Open-Set</kwd>
        <kwd>0-1 cost loss</kwd>
        <kwd>Fungi Species Identification</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>
        Fine-grained visual classification, as a core challenge in the field of computer vision and pattern
recognition, plays a pivotal role in diverse practical applications [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]. The FungiCLEF2024 [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ] challenge,
serving as a crucial component of the LifeCLEF2024 [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ], aims to promote and incentivize in-depth
research on fungi identification algorithms, particularly in complex scenarios that integrate image
and metadata inputs. The achievement of this goal not only holds immense value for biodiversity
conservation, but also plays a crucial role in maintaining human health.
      </p>
      <p>
        Prior FungiCLEF challenges have achieved significant progress through deep learning models [
        <xref ref-type="bibr" rid="ref10 ref11 ref5 ref6 ref7 ref8 ref9">5, 6, 7,
8, 9, 10, 11</xref>
        ]. To further enhance the practical significance of the competition and efectively address
the challenges faced by developers, scientists, users, and the community, this year’s organizers have
introduced additional constraints. Therefore, the challenges faced by this year’s competition can be
summarized as follows:
• 0-1 Cost Loss: One of the core issues in FungiCLEF2024, which categorizes fungi into poisonous
and non-poisonous, is how to construct a model that minimizes the misclassification of poisonous
fungi as non-poisonous to ensure high reliability and safety in identification results.
• Hardware Constraints: All algorithms will be executed on the HuggingFace platform, subject
to strict limitations of 16GB of GPU memory and a two-hour runtime.
      </p>
      <p>
        The FungiCLEF2024 dataset is based on data collected through the Atlas of Danish Fungi mobile and
Web applications. All fungi specimen observation had to pass the expert validation process, therefore
guaranteeing high-quality labels. For the training dataset, it contains 295,938 images - belonging to
1,604 species. The validation dataset contains 60832 images belonging to 2,713 species, 1,084 known
from the training set and 1,629 unknown species. The dataset statistics are listed in Table 1.
For fine-grained image recognition and feature fusion, this paper employs Dynamic MLP [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ] to
fully fuse diverse feature information. Additionally, for open-set recognition, an entropy-based
approach is utilized, leveraging the model’s prediction confidence through entropy to identify open-set
images, surpassing previous methods. Furthermore, to minimize the critical 0-1 cost loss caused by
misclassifying poisonous fungi as non-poisonous, this paper proposes an easy but quite efective
way to mitigate 0-1 cost by utilizing a marginal expected loss function during training, which
significantly reduces the cost loss while maintaining accuracy. Details of methods will be discussed in Section 3.
The subsequent sections are organized as follows: In Section 2, we will provide a detailed explanation
and interpretation of the dataset and evaluation metrics used in the competition. Section 3 will outline
the methodology and core concepts adopted in this paper. Section 4 will focus on presenting our
experimental results and actual performance in the competition. Finally, in Section 5, we will provide a
comprehensive review and summary of the entire content.
      </p>
    </sec>
    <sec id="sec-2">
      <title>2. Related Work and Evaluation Metrics</title>
      <sec id="sec-2-1">
        <title>2.1. Related Work</title>
        <p>
          Fine-grained image classification: To enhance fine-grained image classification, several approaches
have been proposed. For instance, [
          <xref ref-type="bibr" rid="ref12 ref13 ref14 ref15">12, 13, 14, 15, 16</xref>
          ] detect the discriminative regions of an image to
exploit subtle details. SnapMix [17] utilizes the class activation map (CAM) [18] to mitigate label noise
in fine-grained data augmentation. Similarly, Attribute Mix [ 19] focuses on semantically meaningful
attribute features from two images to identify the same super-categories. FixRes [20] investigates data
augmentation and resolution strategies to boost classification performance. Other studies focus on
extracting more valuable features from multi-channel networks [21] or through contrastive learning [22].
Using additional information: Besides visual information, researchers have incorporated additional
information to enhance classification performance. Many existing works [ 23, 24, 25, 26] combine the
image features with additional multi-modal features directly through channel-wise concatenation.
Multi-modal features, including images, ages, and dates, were first introduced by Tang et al. [ 26], who
concatenated them from an MLP backbone network to make a joint prediction. Subsequently, Minetto
et al. [24] introduced metadata to the geo-spatial land classification task. Further, Salem et al. [ 25]
integrated dense overhead imagery with location and date into a general framework by concatenating
the outputs of the context network.
        </p>
        <p>Open-set recognition: Discriminative models are one of the most important ways for open-set
recognition [27]. Traditional methods, such as 1-vs-Set machines based on SVM [28], often sufer
from limitations stemming from the weak feature extraction ability of those traditional models. In
recent years, deep learning-based methods have garnered increasing attention due to their powerful
representation abilities. Bendale et al. [29] first proposed replacing the softmax layer in the network
with OpenMax, which calibrates the output probability using the Weibull distribution. A similar
work [30] replaced the softmax layer with one-vs-rest units. These methods have pioneered a new
direction for the research of open-set recognition.</p>
        <p>Previous FungiCLEF work: Most contributions to FungiCLEF2023 were centered on modern
Convolutional Neural Network (CNN) or transformer-inspired architectures, such as MetaFormer [31], Swin
Transformer [32], and Volo [33]. The winning team [34] achieved 79.28% accuracy using MetaFormer.
These results were often enhanced by combining predictions from the same observation and through
data augmentations applied during both training and testing. Techniques such as Seesaw loss [35],
Focal loss [36], Arcface loss [37], and Sub-Center loss [38] have achieved great success in addressing
the unbalanced class distribution. Additionally, metadata was combined with image features to classify
fungi categories, further improving the overall performance.</p>
      </sec>
      <sec id="sec-2-2">
        <title>2.2. Evaluation Metrics</title>
        <p>FungiCLEF2024 has set a total of 3 evaluation metrics, namely Track1, Track2, and Track3, which will
be introduced below.</p>
        <p>Track1: The rfist metric is the standard classification error, which is the average error in predicting
class labels. All categories not present in the training set should be correctly classified as the "unknown"
category (i.e., labeled as -1). The specific calculation formula is as follows:</p>
        <p>Track1(, ())) =
︂{ 0 if () = 
1 otherwise
Here,  represents the input image,  represents the true label of the input image, and  represents the
trained classification model.</p>
        <p>Track2: The second metric is the cost loss associated with confusing non-toxic and toxic species. Define
d(· ) as an indicator function, where if d() = 1, it indicates that category  is a toxic category, and
if d() = 0, it indicates that category  is a non-toxic category. The specific calculation formula for
Track2 is as follows:
⎧
⎪
Track2(, ()) = ⎨
0</p>
        <p>if d(()) = d()
  if d(()) = 0 &amp; d() = 1
⎪⎩ if d(()) = 1 &amp; d() = 0
In this competition,   = 100 and  = 1.</p>
        <p>Track3: The third metric is the sum of Track1 and Track2. The specific formula is as follows:</p>
        <p>Track3 = Track1 + Track2
The final ranking of the competition is based on the performance of Track3. Regardless of whether it is
Track1, Track2, or Track3, the lower the score, the higher the ranking.
(1)
(2)
(3)</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>3. Method</title>
      <p>In this section, we will introduce our method to handle the recognition problem with open-set and raise
an easy way to decrease the Track2 significantly.</p>
      <sec id="sec-3-1">
        <title>3.1. Fine-grained Image Recognition with Feature Fusion</title>
        <p>
          Feature fusion. To enhance the image representation and improve the result of image classification,
"Dynamic MLP" proposed in [
          <xref ref-type="bibr" rid="ref1">1</xref>
          ] is applied to fully release the potential of the meta information. Marking
the input image feature as  and the meta feature as  respectively, Dynamic MLP is designed to fuse
features through a matrix multiplication operation. Specifically, given the input image , we can obtain
the image feature through the backbone network1, and the meta (such as substrate and habitat) feature
through a well pre-trained clip text model [39], which can be described as following:
 = Backbone()
        </p>
        <p>), Clip(ℎ )))
 = MLP(Cat(Clip(
where Backbone(· ) denotes the model before the classification head, and  ∈ R, n is the output
dimension.  and ℎ  denote the substrate and habitat data, respectively. Clip denotes a well trained

clip text model. Cat(· ) denotes the channel-wise concatenation. MLP denotes a residual MLP network,
following the descriptions in PriorsNet [40]. After MLP, the meta feature is projected into the same
dimension as , i.e.,  ∈ R. Then, let the original  as 0, Dynamic MLP takes 0 and  as initial
inputs, and the enhanced image representation  is obtained after N recursive blocks. At last,  is
expanded to align the shape with 0 by a channel-increasing layer for classifying images. The process
of Dynamic MLP can be specifically summarized into three steps:
1. Taking into image and meta features and reshaping the meta feature from a 1-d vector to a 2-d
matrix, which can be formalized as following:</p>
        <p>= Reshape( ())
where Reshape(· ) denotes reshape operation, and f denotes a fully connected layer.
2. Obtaining the enhanced image feature  after N recursions are completed.</p>
        <p>+1 = ReLU(LN( ( @))),  = 0, 1, ...,</p>
        <p>ReLU(· ) and LN(· ) denote ReLU activation function and layer normalization, respectively. The
operator @ denotes the matrix multiplication.
3. Aligning the dimension of  and 0.</p>
        <p>ˆ = Layer( )
 = Head(ˆ )
(4)
(5)
(6)
(7)
(8)
(9)</p>
        <p>Layer(· ) denotes a channel-increasing layer.</p>
        <p>Fine-grained image classification. After the final enhanced image feature ˆ is obtained, we can
use it to classify the fine-grained Fungi images:
where Head is the last classification head which is used for recognizing images.
1In CNN backbones, a image feature are acquired after a pooling layer. In Vit based models, a image feature is the [CLS] token
in the last layer.</p>
      </sec>
      <sec id="sec-3-2">
        <title>3.2. Entropy Based Open-set Identifier</title>
        <p>In the Fungi competition, models are not only asked to correctly identify the species in the close-set, but
also the pictures in the open-set. As shown in Table 2, existing approaches such as [29, 41] show their
superiority on coarse-grained datasets, but are not suitable for large scale fine-grained datasets. So we
adopt an easy entropy based method to identify open-set images, which shows better result than [29, 41].
Entropy is defined following to measure the quality of the probability distribution:
entropy() = −

∑︁  log 
=1
where  = Softmax()
where Softmax(· ) denotes the softmax operation, C denotes the number of categories to be classified in
the close-set and  denotes the the probability of being identified as the category . In general, the
model is more confident for known categories, corresponding to a lower entropy. Whereas for unknown
categories the uncertainty is higher and hence the entropy will be higher. Thus, we can efectively
distinguish between known/unknown categories through a entropy threshold  . Once the threshold 
is determined, we can use the following formula to identify the open-set images:
 =
{︃</p>
        <p>− 1 if entropy() &gt; ,</p>
        <p>Argmax() if entropy() ≤ .
where Argmax(· ) denotes the argmax function. Since the choice of  determines the efect of a model,
we find the best threshold based on the validation set.</p>
      </sec>
      <sec id="sec-3-3">
        <title>3.3. Probability-guided Poisonous Recognizer</title>
        <p>The two tasks of the competition are improving the classification accuracy and reducing the cost of
classifying poisonous as non-poisonous, respectively. In fact, the latter task has a greater impact on the
ifnal score than the former. In this section, we will introduce an easy but quite efective way to reduce
the cost of the latter task.</p>
        <p>Put poisonous categories together. The Fungi dataset is a long-tail dataset with uneven distribution,
Figure 1 illustrates that the number of categories and the quantity of images for the poisonous are
significantly lower compared to edible ones. This imbalance poses a challenge for models to adequately
learn robust embeddings for poisonous species, thereby leading to misclassification where poisonous
species may be incorrectly labeled as edible. Such errors incur substantial costs. Therefore, we first
put all poisonous categories into a single class, efectively reducing the total number of categories
from 1604 to 1556 (comprising 1555 edible species and 1 aggregated poisonous class). This approach
facilitates the model’s focus on the general features of poisonous species, alleviating the need to discern
subtle distinctions. Next, two models will be trained separately to classify mixed categories (1555 edible
categories and a poisonous category) and only poisonous categories (49 poisonous categories).
(10)
(11)
(12)</p>
        <p>Counts of Poisonous vs Edible Fungi by Class ID</p>
        <p>Edible</p>
        <p>Poisonous
0
200
400
600</p>
        <p>Marginal expected loss. The most direct way to reduce cost is to optimize the cost function itself, but
since the calculation of cost is discrete, we use the marginal expected loss function here:
  =
|ℐ| ∈ℐ =1</p>
        <p>1 ∑︁ ∑︁  · gt()
where ℐ represents the training images and gt(· ) denotes the ground truth label of the image. As defined
in Section 2, d(· ) indicate poisonous species, where d() = 1 if the category  is poisonous, and d() = 0
if  is edible. According to the calculation formula of Track2, we can obtain the specific expression of</p>
        <p>gt(i), however, according to Track2, no matter what the ground truth of the picture is, as long as
the predicted label meets d(gt()) = d(), the same loss will be obtained, thus disrupting the correct
gradient descent direction to the ground truth. Hence, We amend the expression of 
gt(i) to introduce
a penalty specifically for instances where d(gt()) = d() while gt() ̸= , to address this issue. The
formula is as following:
gt() =
⎧ 0 if gt() = 
⎪
⎪⎪⎨ 5 if d(gt()) = d() &amp; gt() ̸= 
⎪ 10 if d(gt()) = 0 &amp; d() = 1
⎪
⎪⎩100 if d(gt()) = 1 &amp; d() = 0
MEloss will force the model to improve its ability to identify the poisonous and the edible while
improving the recognition accuracy.</p>
        <p>Increasing the recall rate of the poisonous. The primary aim of reducing Track2 lies in maximizing
the recall rate for the poisonous species. To achieve this, we set a probability-guided threshold  ,
where an image is considered poisonous as soon as the predicted probability of the poisonous category
exceeds  . As previously discussed, let define the model classifying mixed categories as ℎ and the model
classifying poisonous categories as , the ultimate decision method is outlined as Algorithm 1.</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>4. Experiments</title>
      <p>In this section, we will introduce the details and main results in detail.
(13)
(14)</p>
      <sec id="sec-4-1">
        <title>4.1. Implementation Details</title>
        <p>Basic settings. The thresholds of  and  are 1.5 and 0.01 respectively. The proposed method has been
developed utilizing the PyTorch framework [42]. Vit-large [43] and Eva02-large [44], implemented via
the timm library [45], serve as the model ℎ and , respectively. All the models have been pre-training
on the ImageNet dataset [46], and are conveniently accessible in HuggingFace. Fine-tuning of these
models was performed using 8 Nvidia RTX3090 GPUs. Input size of images is 336. The initial learning
rate was set to 2 × 10− 5, and the total number of training epochs was set to 15, with the first epoch
dedicated to warm-up by employing a learning rate of 2 × 10− 7. For optimal model training, we
employed the AdamW optimizer [47] in conjunction with a cosine learning rate scheduler [48], with
the weight decay set to 1 × 10− 2. Since the Fungi dataset with an unbalanced distribution, many
studies have suggested solutions [49, 50, 51, 52], considering the practical efect, we finally use seesaw
loss [35] and ME loss mentioned above to optimize the model.</p>
        <p>Data Augmentations. We employ a composed sequence of common augmentation techniques to
enhance results. During training, we first perform random cropping on the image, where the size of the
cropped region is randomly chosen between 50% and 100% of the original image size. Subsequently, the
slice is resized using the bicubic interpolation method and flipped horizontally and vertically with a
probability of 50%. Additionally, we incorporate hue-saturation and brightness-contrast augmentations
to randomly adjust the hue, saturation, value, brightness, and contrast of the input image. Finally,
standard normalization is applied to all input images. However, due to limitations in both running time
and GPU memory on HuggingFace, test-time augmentations are simplified by first resizing images to
336 and then normalizing them using the same mean and std as during training.</p>
      </sec>
      <sec id="sec-4-2">
        <title>4.2. Fungi Dataset Experiments</title>
        <p>The key experimental results are presented in Table 4. As evident from the table, Dynamic MLP
exhibits superior feature fusion ability, efectively reducing the error rate for Vit and Eva. Concurrently,
the strategy of consolidating all poisonous categories mitigates the model’s need to discern subtle
diferences between poisonous classes, enabling it to concentrate on macro diferences instead. Notably,
as observed in the last two lines of the table, Track2 and Track1 occupy opposing ends of the seesaw,
which is a natural consequence of the poisonous recognition strategy outlined in this paper. However,
from the standpoint of Track3, the advantages of optimizing Track2 outweigh those of Track1, thereby
reafirming the arguments put forth in section 3.</p>
        <p>Table 3 demonstrates the impact of diferent  on Track2. It is evident from the table that as the
threshold value decreases, Track2 steadily drops. This is because  directly afects the model’s recall
rate for poisonous classes; the smaller  is, the higher the recall rate of the model, thus reducing the
and all the methods mentioned in Section 3 are adopted. The results are reported on the validation set.

0.2
0.15
0.1
0.05
0.01</p>
        <sec id="sec-4-2-1">
          <title>Track2</title>
        </sec>
        <sec id="sec-4-2-2">
          <title>The Table shows results of diferent methods and settings, all of which are reported on the validation</title>
          <p>set. Track1 represents the error rate of classification (both on close-set and open-set). Track2 represents
the cost caused by misclassification and Track3 is equivalent to Track1 plus Track2. All these metrics are
as small as possible. DM represents Dynamic MLP, PPT represents for putting the poisonous categories
together, cat represents the concatenate operation in channel-wise with meta and image features.</p>
        </sec>
        <sec id="sec-4-2-3">
          <title>Feature fusion</title>
        </sec>
        <sec id="sec-4-2-4">
          <title>Open-set</title>
          <p>0-1 cost</p>
        </sec>
        <sec id="sec-4-2-5">
          <title>Backbone</title>
        </sec>
        <sec id="sec-4-2-6">
          <title>Vit-large</title>
        </sec>
        <sec id="sec-4-2-7">
          <title>Eva-large</title>
        </sec>
        <sec id="sec-4-2-8">
          <title>Vit-large</title>
        </sec>
        <sec id="sec-4-2-9">
          <title>Eva-large</title>
        </sec>
        <sec id="sec-4-2-10">
          <title>Vit-large &amp;</title>
        </sec>
        <sec id="sec-4-2-11">
          <title>Eva-large</title>
        </sec>
        <sec id="sec-4-2-12">
          <title>Vit-large &amp;</title>
        </sec>
        <sec id="sec-4-2-13">
          <title>Eva-large</title>
        </sec>
        <sec id="sec-4-2-14">
          <title>Vit-large &amp;</title>
        </sec>
        <sec id="sec-4-2-15">
          <title>Eva-large</title>
        </sec>
        <sec id="sec-4-2-16">
          <title>Vit-large &amp;</title>
        </sec>
        <sec id="sec-4-2-17">
          <title>Eva-large</title>
          <p>smaller the better.
cat
cat
DM
DM
DM
DM
DM
DM
Rank
1
2
3
4
5
6
7</p>
        </sec>
        <sec id="sec-4-2-18">
          <title>Entropy</title>
        </sec>
        <sec id="sec-4-2-19">
          <title>Entropy</title>
        </sec>
        <sec id="sec-4-2-20">
          <title>Entropy</title>
          <p>PPT
PPT</p>
        </sec>
        <sec id="sec-4-2-21">
          <title>ME-loss &amp; PPT</title>
        </sec>
        <sec id="sec-4-2-22">
          <title>ME-loss &amp; PPT &amp;</title>
          <p>threshold 
Track1
0.4623
0.4547
0.4581
0.4438
0.4313
0.3651
0.3809
0.3951
Track2
0.5095
0.5587
0.4854
0.4709
0.3234
0.4834
0.2868
Track3
0.9718
1.0134
0.9435
0.9147
0.7647
0.8485
0.6677
0.1226
0.5177</p>
        </sec>
        <sec id="sec-4-2-23">
          <title>Team</title>
          <p>IES
jack-etheredge
upupup(Our)
chirmy</p>
        </sec>
        <sec id="sec-4-2-24">
          <title>TingTing1999 glhr DS@GT Track1</title>
          <p>0.3107
0.2436
0.3898
0.2693
0.2749
0.4996
0.3907
Track2
0.0904
0.1629
0.0718
0.4149
0.4378
0.6511
1.604
Track3
0.3621
0.4075
0.513
0.6667
0.6934
1.1526
2.0443</p>
        </sec>
        <sec id="sec-4-2-25">
          <title>The table shows the final scores of diferent teams on the private leaderboard, Track1 represents the error rate of image recognition (including closed and open sets), Track2 represents the cost loss caused by wrong recognition, Track3 is equal to Track1 plus Track2, and the ranking is based on Track3, the</title>
          <p>probability of identifying poisonous classes as non-poisonous.
leaderboard, where we secured the 3rd position, surpassing the 4th place in Track2 and achieving a
performance even better than the 1st place. These outcomes collectively demonstrate the eficacy of
the method presented in Section 3.3. Nonetheless, due to the simplicity of the approach we adopted
for open-set recognition, we fell short in Track1 compared to other teams, highlighting an area that
requires enhancement in our future endeavors.</p>
        </sec>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>5. Conclusion</title>
      <p>The core challenges of FungiCLEF2024 are identifying fine-grained fungi images in an open-set
environment and minimizing the 0-1 cost of misclassification. To address the first challenge, we use
Dynamic MLP, a recursive structure utilizing matrix multiplication, for feature fusion to improve
accuracy. To mitigate the 0-1 cost, we propose an easy yet efective approach that first places poisonous
fungi categories into a single class and then employs ME loss and  to optimize the recall rate for
poisonous species. However, open-set fine-grained fungi recognition remains a significant challenge.
Our current approach relies solely on entropy for classifying open-set species, which has proven to be
overly simplistic and ineficient. Consequently, the open-set problem stands as an enduring challenge
that necessitates further investigation and innovation.
in: Computer Vision–ECCV 2014: 13th European Conference, Zurich, Switzerland, September
6-12, 2014, Proceedings, Part I 13, Springer, 2014, pp. 834–849.
[16] X.-S. Wei, C.-W. Xie, J. Wu, C. Shen, Mask-cnn: Localizing parts and selecting descriptors for
ifne-grained bird species categorization, Pattern Recognition 76 (2018) 704–714.
[17] S. Huang, X. Wang, D. Tao, Snapmix: Semantically proportional mixing for augmenting
finegrained data, in: Proceedings of the AAAI Conference on Artificial Intelligence, volume 35, 2021,
pp. 1628–1636.
[18] B. Zhou, A. Khosla, A. Lapedriza, A. Oliva, A. Torralba, Learning deep features for discriminative
localization, in: Proceedings of the IEEE conference on computer vision and pattern recognition,
2016, pp. 2921–2929.
[19] H. Li, X. Zhang, Q. Tian, H. Xiong, Attribute mix: Semantic data augmentation for fine grained
recognition, in: 2020 IEEE International Conference on Visual Communications and Image
Processing (VCIP), IEEE, 2020, pp. 243–246.
[20] H. Touvron, A. Vedaldi, M. Douze, H. Jégou, Fixing the train-test resolution discrepancy, Advances
in neural information processing systems 32 (2019).
[21] D. Chang, Y. Ding, J. Xie, A. K. Bhunia, X. Li, Z. Ma, M. Wu, J. Guo, Y.-Z. Song, The devil is in the
channels: Mutual-channel loss for fine-grained image classification, IEEE Transactions on Image
Processing 29 (2020) 4683–4695.
[22] Y. Gao, X. Han, X. Wang, W. Huang, M. Scott, Channel interaction networks for fine-grained
image categorization, in: Proceedings of the AAAI conference on artificial intelligence, volume 34,
2020, pp. 10818–10825.
[23] G. Mai, K. Janowicz, B. Yan, R. Zhu, L. Cai, N. Lao, Multi-scale representation learning for spatial
feature distributions using grid cells, arXiv preprint arXiv:2003.00824 (2020).
[24] R. Minetto, M. P. Segundo, S. Sarkar, Hydra: An ensemble of convolutional neural networks for
geospatial land classification, IEEE Transactions on Geoscience and Remote Sensing 57 (2019)
6530–6541.
[25] T. Salem, S. Workman, N. Jacobs, Learning a dynamic map of visual appearance, in: Proceedings
of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2020, pp. 12435–12444.
[26] K. Tang, M. Paluri, L. Fei-Fei, R. Fergus, L. Bourdev, Improving image classification with location
context, in: Proceedings of the IEEE international conference on computer vision, 2015, pp.
1008–1016.
[27] C. Geng, S.-j. Huang, S. Chen, Recent advances in open set recognition: A survey, IEEE transactions
on pattern analysis and machine intelligence 43 (2020) 3614–3631.
[28] W. J. Scheirer, A. de Rezende Rocha, A. Sapkota, T. E. Boult, Toward open set recognition, IEEE
transactions on pattern analysis and machine intelligence 35 (2012) 1757–1772.
[29] A. Bendale, T. E. Boult, Towards open set deep networks, in: Proceedings of the IEEE conference
on computer vision and pattern recognition, 2016, pp. 1563–1572.
[30] L. Shu, H. Xu, B. Liu, Doc: Deep open classification of text documents, arXiv preprint
arXiv:1709.08716 (2017).
[31] W. Yu, M. Luo, P. Zhou, C. Si, Y. Zhou, X. Wang, J. Feng, S. Yan, Metaformer is actually what
you need for vision, in: Proceedings of the IEEE/CVF conference on computer vision and pattern
recognition, 2022, pp. 10819–10829.
[32] Z. Liu, Y. Lin, Y. Cao, H. Hu, Y. Wei, Z. Zhang, S. Lin, B. Guo, Swin transformer: Hierarchical vision
transformer using shifted windows, in: Proceedings of the IEEE/CVF international conference on
computer vision, 2021, pp. 10012–10022.
[33] L. Yuan, Q. Hou, Z. Jiang, J. Feng, S. Yan, Volo: Vision outlooker for visual recognition, IEEE
transactions on pattern analysis and machine intelligence 45 (2022) 6575–6586.
[34] H. Ren, H. Jiang, W. Luo, M. Meng, T. Zhang, Entropy-guided open-set fine-grained fungi
recognition., in: CLEF (Working Notes), 2023, pp. 2122–2136.
[35] J. Wang, W. Zhang, Y. Zang, Y. Cao, J. Pang, T. Gong, K. Chen, Z. Liu, C. C. Loy, D. Lin, Seesaw
loss for long-tailed instance segmentation, 2021, pp. 9695–9704.
[36] T.-Y. Lin, P. Goyal, R. Girshick, K. He, P. Dollár, Focal loss for dense object detection, in: Proceedings
of the IEEE international conference on computer vision, 2017, pp. 2980–2988.
[37] J. Deng, J. Guo, N. Xue, S. Zafeiriou, Arcface: Additive angular margin loss for deep face recognition,
in: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2019, pp.
4690–4699.
[38] J. Deng, J. Guo, T. Liu, M. Gong, S. Zafeiriou, Sub-center arcface: Boosting face recognition
by large-scale noisy web faces, in: Computer Vision–ECCV 2020: 16th European Conference,
Glasgow, UK, August 23–28, 2020, Proceedings, Part XI 16, Springer, 2020, pp. 741–757.
[39] A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin,
J. Clark, et al., Learning transferable visual models from natural language supervision, in:
International conference on machine learning, PMLR, 2021, pp. 8748–8763.
[40] O. Mac Aodha, E. Cole, P. Perona, Presence-only geographical priors for fine-grained image
classification, in: Proceedings of the IEEE/CVF International Conference on Computer Vision,
2019, pp. 9596–9606.
[41] L. Neal, M. Olson, X. Fern, W.-K. Wong, F. Li, Open set learning with counterfactual images, in:</p>
      <p>Proceedings of the European conference on computer vision (ECCV), 2018, pp. 613–628.
[42] A. Paszke, S. Gross, F. Massa, A. Lerer, J. Bradbury, G. Chanan, T. Killeen, Z. Lin, N. Gimelshein,
L. Antiga, A. Desmaison, A. Kopf, E. Yang, Z. DeVito, M. Raison, A. Tejani, S. Chilamkurthy,
B. Steiner, L. Fang, J. Bai, S. Chintala, Pytorch: An imperative style, high-performance deep
learning library, in: Advances in Neural Information Processing Systems 32, Curran Associates,
Inc., 2019, pp. 8024–8035.
[43] A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani,
M. Minderer, G. Heigold, S. Gelly, et al., An image is worth 16x16 words: Transformers for image
recognition at scale, arXiv preprint arXiv:2010.11929 (2020).
[44] Y. Fang, Q. Sun, X. Wang, T. Huang, X. Wang, Y. Cao, Eva-02: A visual representation for neon
genesis, arXiv preprint arXiv:2303.11331 (2023).
[45] R. Wightman, Pytorch image models, https://github.com/rwightman/pytorch-image-models, 2019.
[46] J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, L. Fei-Fei, Imagenet: A large-scale hierarchical image
database, 2009, pp. 248–255.
[47] I. Loshchilov, F. Hutter, Fixing weight decay regularization in adam (2017).
[48] I. Loshchilov, F. Hutter, SGDR: Stochastic gradient descent with warm restarts, 2017.
[49] X.-S. Wei, S.-L. Xu, H. Chen, L. Xiao, Y. Peng, Prototype-based classifier learning for long-tailed
visual recognition, Science China Information Sciences 65 (2022) 160105.
[50] Y.-Y. He, J. Wu, X.-S. Wei, Distilling virtual examples for long-tailed recognition, in: Proceedings
of the IEEE/CVF International Conference on Computer Vision, 2021, pp. 235–244.
[51] Y. Zhang, X. Wei, B. Zhou, J. Wu, Bag of tricks for long-tailed visual recognition with deep
convolutional neural networks, 2021, pp. 3447–3455.
[52] B. Zhou, Q. Cui, X.-S. Wei, Z.-M. Chen, Bbn: Bilateral-branch network with cumulative learning
for long-tailed visual recognition, in: 2020 IEEE/CVF Conference on Computer Vision and Pattern
Recognition (CVPR), 2020, pp. 9716–9725. doi:10.1109/CVPR42600.2020.00974.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>L.</given-names>
            <surname>Yang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Song</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Zhao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Tao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Zhou</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Liang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Yang</surname>
          </string-name>
          ,
          <article-title>Dynamic mlp for fine-grained image classification by leveraging geographical and temporal information</article-title>
          ,
          <source>in: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition</source>
          ,
          <year>2022</year>
          , pp.
          <fpage>10945</fpage>
          -
          <lpage>10954</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>X.-S.</given-names>
            <surname>Wei</surname>
          </string-name>
          , Y.-
          <string-name>
            <given-names>Z.</given-names>
            <surname>Song</surname>
          </string-name>
          ,
          <string-name>
            <given-names>O. Mac</given-names>
            <surname>Aodha</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Wu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Peng</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Tang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Yang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Belongie</surname>
          </string-name>
          ,
          <article-title>Fine-grained image analysis with deep learning: A survey</article-title>
          ,
          <source>IEEE Transactions on Pattern Analysis and Machine Intelligence</source>
          <volume>44</volume>
          (
          <year>2021</year>
          )
          <fpage>8927</fpage>
          -
          <lpage>8948</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>L.</given-names>
            <surname>Picek</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Sulc</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Matas</surname>
          </string-name>
          , Overview of FungiCLEF 2024:
          <article-title>Revisiting fungi species recognition beyond 0-1 cost</article-title>
          ,
          <source>in: Working Notes of CLEF 2024 - Conference and Labs of the Evaluation Forum</source>
          ,
          <year>2024</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>A.</given-names>
            <surname>Joly</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Picek</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Kahl</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Goëau</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Espitalier</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Botella</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Deneu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Marcos</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Leblanc</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Larcher</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Šulc</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Hrúz</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Servajean</surname>
          </string-name>
          , et al.,
          <source>Overview of LifeCLEF</source>
          <year>2024</year>
          :
          <article-title>Challenges on species distribution prediction and identification</article-title>
          ,
          <source>in: International Conference of the CrossLanguage Evaluation Forum for European Languages</source>
          , Springer,
          <year>2024</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>L.</given-names>
            <surname>Picek</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Šulc</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Matas</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Heilmann-Clausen</surname>
          </string-name>
          ,
          <article-title>Overview of fungiclef 2022: Fungi recognition as an open set classification problem</article-title>
          , Working Notes of CLEF (
          <year>2022</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>J.</given-names>
            <surname>Yu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Chang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Lu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Xie</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Zhang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Cai</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Du</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Wei</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Liu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Gao</surname>
          </string-name>
          , et al.,
          <article-title>Bag of tricks and a strong baseline for FGVC</article-title>
          , Working Notes of CLEF (
          <year>2022</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>G.</given-names>
            <surname>Fan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Zining</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W.</given-names>
            <surname>Weiqiu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Yinan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Fei</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Zhicheng</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Hong</surname>
          </string-name>
          ,
          <article-title>Does closed-set training generalize to open-set recognition?</article-title>
          , Working Notes of CLEF (
          <year>2022</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>K.</given-names>
            <surname>Desingu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Bhaskar</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Palaniappan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E. A.</given-names>
            <surname>Chodisetty</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Bharathi</surname>
          </string-name>
          ,
          <article-title>Classification of fungi species: A deep learning based image feature extraction and gradient boosting ensemble approach</article-title>
          , Working Notes of CLEF (
          <year>2022</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>S.</given-names>
            <surname>Wolf</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Beyerer</surname>
          </string-name>
          ,
          <article-title>Transformer-based fine-grained fungi classification in an open-set scenario</article-title>
          , Working Notes of CLEF (
          <year>2022</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>F.</given-names>
            <surname>Hu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Duan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Zhu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.-S.</given-names>
            <surname>Wei</surname>
          </string-name>
          ,
          <article-title>A deep learning based solution to fungiclef2023</article-title>
          ,
          <string-name>
            <surname>Aliannejadi</surname>
          </string-name>
          et al.[
          <volume>1</volume>
          ] (
          <year>2023</year>
          )
          <fpage>2051</fpage>
          -
          <lpage>2059</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>F.</given-names>
            <surname>Hu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Duan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Zhu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Zhang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.-S.</given-names>
            <surname>Wei</surname>
          </string-name>
          ,
          <article-title>Watch out venomous snake species: A solution to snakeclef2023</article-title>
          ,
          <source>arXiv preprint arXiv:2307.09748</source>
          (
          <year>2023</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>A.</given-names>
            <surname>Behera</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Wharton</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P. R.</given-names>
            <surname>Hewage</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Bera</surname>
          </string-name>
          ,
          <article-title>Context-aware attentional pooling (cap) for finegrained visual classification</article-title>
          ,
          <source>in: Proceedings of the AAAI conference on artificial intelligence</source>
          , volume
          <volume>35</volume>
          ,
          <year>2021</year>
          , pp.
          <fpage>929</fpage>
          -
          <lpage>937</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <given-names>T.</given-names>
            <surname>Xiao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Xu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Yang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Zhang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Peng</surname>
          </string-name>
          ,
          <string-name>
            <surname>Z. Zhang,</surname>
          </string-name>
          <article-title>The application of two-level attention models in deep convolutional neural network for fine-grained image classification</article-title>
          ,
          <source>in: Proceedings of the IEEE conference on computer vision and pattern recognition</source>
          ,
          <year>2015</year>
          , pp.
          <fpage>842</fpage>
          -
          <lpage>850</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <given-names>Z.</given-names>
            <surname>Yang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Luo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Hu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Gao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <article-title>Learning to navigate for fine-grained classification</article-title>
          ,
          <source>in: Proceedings of the European conference on computer vision (ECCV)</source>
          ,
          <year>2018</year>
          , pp.
          <fpage>420</fpage>
          -
          <lpage>435</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <given-names>N.</given-names>
            <surname>Zhang</surname>
          </string-name>
          , J. Donahue,
          <string-name>
            <given-names>R.</given-names>
            <surname>Girshick</surname>
          </string-name>
          , T. Darrell,
          <article-title>Part-based r-cnns for fine-grained category detection,</article-title>
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>