<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Asymptotic Cross-Entropy Weighting and Guided-Loss in Supervised Hierarchical Setting using Deep Attention Networks</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Charles Kantor</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
          <xref ref-type="aff" rid="aff5">5</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Brice Rauby</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff5">5</xref>
          <xref ref-type="aff" rid="aff6">6</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Le´onard Boussioux</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
          <xref ref-type="aff" rid="aff5">5</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Emmanuel Jehanno</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff5">5</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Andre´-Philippe Drapeau Picard</string-name>
          <xref ref-type="aff" rid="aff4">4</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Maxim Larrive´e</string-name>
          <xref ref-type="aff" rid="aff4">4</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Hugues Talbot</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff5">5</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Ecole CentraleSupe ́lec Paris</institution>
          ,
          <country country="FR">France</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Inria Paris</institution>
          ,
          <country country="FR">France</country>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>Massachusetts Institute of Technology, Operations Research Center</institution>
          ,
          <addr-line>Cambridge, MA</addr-line>
          ,
          <country country="US">USA</country>
        </aff>
        <aff id="aff3">
          <label>3</label>
          <institution>Mila Artificial Intelligence Institute</institution>
          ,
          <addr-line>Montreal</addr-line>
          ,
          <country country="CA">Canada</country>
        </aff>
        <aff id="aff4">
          <label>4</label>
          <institution>Montreal Insectarium - Space for Life</institution>
          ,
          <country country="CA">Canada</country>
        </aff>
        <aff id="aff5">
          <label>5</label>
          <institution>Paris-Saclay University</institution>
          ,
          <country country="FR">France</country>
        </aff>
        <aff id="aff6">
          <label>6</label>
          <institution>Polytechnique Montreal</institution>
          ,
          <country country="CA">Canada</country>
        </aff>
      </contrib-group>
      <fpage>5157</fpage>
      <lpage>5166</lpage>
      <abstract>
        <p>This article reveals two main techniques for improving finegrained recognition and classification, defined as executing these tasks between items with similar general patterns, but that differ through small details. First, we build a preliminary automated segmentation algorithm to ignore the image's background for attention-guided classification. To do this, we wield segmentation-issued masks to make the classification network's training easier through an additional loss that penalizes attention given to features outside the mask using convolutional block. Furthermore, we proffer a hierarchical loss based on cross entropy penalizing parent-level classification to leverage the philology of each wildlife species. We applied our approaches in the particular context of butterfly recognition, which is of practical interest to entomologists.</p>
      </abstract>
      <kwd-group>
        <kwd>Hierarchical Classification</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>In this work, we deal with the issue of accurately
identifying large numbers of items in photographs, some of which
may differ only in minute details. This is a difficult problem
because both large and small differences must be taken into
account in order to recognize and classify.</p>
      <p>
        Among the images collected, a high percentage of species
remains unidentified and represents a time-consuming
labeling task for experts. Identifying an insect on the species level
is challenging and depends on the tiniest of details. Citizen
scientists can help collect a large amount of data such as
insect photographic documentation
        <xref ref-type="bibr" rid="ref3 ref7">(Horn et al. 2017;
Boussioux et al. 2019)</xref>
        , but accurate identification remains
restricting. Recent improvements in performance in a wide
range of classification tasks with deep learning methods
offer population monitoring opportunities and efficient and
large scale annotations. We worked with the eButterfly
        <xref ref-type="bibr" rid="ref7">(Prudic et al. 2017)</xref>
        citizen science program, which maintains a
fine-grained dataset of observations of all North American
butterflies species.
      </p>
      <p>We develop computer vision algorithms and propose
finegrained classification innovations using segmentation tools
to encourage the model to focus on areas of an image that
are salient for identification. We propose an additional loss
using the segmentation masks, penalizing attention given to
features outside the mask. We also design a specific loss
function that leverages the dataset’s hierarchical nature,
consequently improving the Family, Genus and Species level
accuracy.</p>
    </sec>
    <sec id="sec-2">
      <title>Related Work</title>
      <sec id="sec-2-1">
        <title>Preliminaries</title>
        <p>
          Fine-grained classification is a category of image
classification: the task is to distinguish between subtly rather
than grossly different items; for example between different
species of birds or dogs rather than giraffes vs. trucks. This
setting is more complex, requires better annotations, more
data and is not yet satisfactorily solved
          <xref ref-type="bibr" rid="ref12 ref4">(Xie et al. 2013;
Chai, Lempitsky, and Zisserman 2013)</xref>
          . A fundamental
difficulty is to induce the learning architecture to focus on small
but essential details without relying on overly complicated
annotations. A recent interesting approach has been to use
a deconstruction-reconstruction method to this end
          <xref ref-type="bibr" rid="ref5">(Chen
et al. 2019)</xref>
          and bipartite and bi-modal graphs
          <xref ref-type="bibr" rid="ref14">(Zhou and
Lin 2016; Song et al. 2020)</xref>
          .
        </p>
        <p>Segmentation is a fundamental task in computer vision.
Its objective is to find semantically consistent regions that
represent objects. Given enough data and annotations, deep
recurrent CNN architectures such as ResNet (He et al. 2016)
and recurrent auto-encoders like U-Net (Ronneberger,
Fischer, and Brox 2015) constitute the current state-of-the-art in
segmentation methods. In particular, U-Net and its variants
may learn a segmentation task from a few hundred labeled
inputs.</p>
        <p>
          The background of macro wildlife photos is typically full
of environmental details like grass or leaves that can
mislead the classification model and introduce bias. We noticed
experimentally via saliency maps that too much attention is
generally paid to the background rather than to the insect
itself. Consequently, we used automatic segmentation to help
focus on the foreground.
Hierarchical labels of a fine-grained dataset can be leveraged
to improve performance
          <xref ref-type="bibr" rid="ref13">(Zheng et al. 2020)</xref>
          .
        </p>
        <p>Tree-CNN is an architecture-based approach to this
setting in (Roy, Panda, and Roy 2019). Their study aims to
overcome the issue of catastrophic forgetting when
finetuning a model successively on each task (i.e., pre-training
on the task of classifying Families, then Genus and finally
Species). The architecture is built with a common trunk and
several finer branches corresponding to each task. Hence, the
model uses the hierarchy through the trunk while also
preserving the memory of each task. This Tree-CNN limits the
computation costs of retraining and can learn with a smaller
training effort than fine-tuning while retaining much of the
accuracy.</p>
        <p>
          <xref ref-type="bibr" rid="ref11">(Wu, Tygert, and LeCun 2019)</xref>
          present another approach
based on the loss instead of the architecture. They propose
a new loss that takes the hierarchy into account and is no
longer a flat loss compared to the cross-entropy (which
compares the same level classes). A convenient metric should
make all leaves equidistant from the root node.
        </p>
        <p>(Kosmopoulos et al. 2020) develops a different
methodology for tuning the loss. They compare different measures
presented in the literature and classify them into two
subgroups: pair-based and set-based measures. They consider
building a hierarchical loss as an optimization problem and
propose a pair-based metric, optimized with a max-flow
approach. The article also offers a set-based implementation
that approximates the sparsity-inducing `0 norm.
Considering the path from the root to the predicted node as a set of
nodes and identically for the ground truth node, they can
propose a measure that uses the intersection, union and
difference between the sets. This measure computes analogous
precision, recall and F1 scores. They implement a Lowest
Common Ancestor, which is a bridge between pair-based
and set-based measures. The corresponding new measure
performs very well and takes the advantages of both
approaches.</p>
      </sec>
      <sec id="sec-2-2">
        <title>Attention and Visualizing CNNs</title>
        <p>
          Attention mechanisms were introduced originally for
Neural Machine Translation in
          <xref ref-type="bibr" rid="ref1">(Bahdanau, Cho, and Bengio
2015)</xref>
          using recurrent neural networks. These mechanisms
utilize a divide-and-conquer approach to various AI tasks
by focusing features on relevant items. These tools have
been extensively used in Natural Language Processing tasks
(e.g., (Parikh et al. 2016) and
          <xref ref-type="bibr" rid="ref7">(Lin et al. 2017)</xref>
          ).
          <xref ref-type="bibr" rid="ref7">(Vaswani
et al. 2017)</xref>
          developed a model relying only on attention
and achieved state-of-the-art results in machine translation.
Later, the use of attention was extended to computer vision
tasks such as image classification and segmentation.
        </p>
        <p>
          The Convolutional Block Attention Module or CBAM
          <xref ref-type="bibr" rid="ref10 ref9">(Woo et al. 2018a)</xref>
          proposes a simple attention mechanism
for feed-forward Convolutional Neural Network (CNN)
architectures. Its lightweight structure and generality make it
suitable for many vision tasks that require large numbers of
parameters.
        </p>
        <p>
          Grad-CAM
          <xref ref-type="bibr" rid="ref7">(Selvaraju et al. 2017)</xref>
          is a popular technique
to make CNN models more explainable, showing on which
areas of the picture they focused on making the prediction,
using reverse gradient propagation descent. The
discriminative regions are localized through the areas of high gradient
flow within the network.
        </p>
      </sec>
      <sec id="sec-2-3">
        <title>Main source</title>
        <p>
          For our preliminary experiments, we use a data set of
pictures submitted across Canada, Mexico and the United
States, representing over seven hundred different species as
of June 2020. Among these observations, two-thirds have
been annotated by experts. The eButterfly program,
cofounded by the Montreal Insectarium, allows participants to
record sightings by uploading images with date and time
information
          <xref ref-type="bibr" rid="ref7">(Prudic et al. 2017)</xref>
          .
        </p>
      </sec>
      <sec id="sec-2-4">
        <title>Highly imbalanced classes</title>
        <p>Our data set is organized hierarchically. Each image has
three labels: a species belongs to one and only one genus,
belonging to one and only one family. This distribution of
labels enables us to have different complexity levels for
our classification task. Given more than two-thirds labeled
images, we anticipate being able to learn the family label
with the best precision, provide a slightly less accurate
estimate of the genus and a slightly worse again estimate of
the species. In the provided dataset, classes are highly
imbalanced, meaning we are facing a problem of fine-grained
classification with significantly under-represented classes.</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>Methods</title>
      <sec id="sec-3-1">
        <title>Guided Attention Mechanism</title>
        <p>As the shape of butterflies presents a limited variability, it
seems feasible to incorporate prior knowledge of the shape
of the object of interest. Due to the similarity between
butterflies’ overall shape, we posit that butterfly segmentation
is a simpler task than its fine-grained classification. We
propose to use masks obtained through an automatic
segmentation pipeline to improve the classification performance.
However, even though the segmentation is generally correct,
a few failure cases can deteriorate the classification’s
performance if used during test-time. For this reason, we
developed a method to leverage these masks during training
through an additional loss later called guided attention loss.</p>
      </sec>
      <sec id="sec-3-2">
        <title>Prior automated segmentation We used a pre-trained</title>
        <p>Mask R-CNN network to generate the segmentation masks
and fine-tuned it on a small subset of the dataset. This
approach is possible because the butterfly segmentation task is
sufficiently similar to the task of segmenting other objects
present in a common dataset, and therefore, pre-training is
very effective. We annotated a small subset (10%) used for
pre-training, and we qualitatively assessed the segmentation
performance. The segmentation results obtained were
satisfactory to be used in the guided attention.</p>
      </sec>
      <sec id="sec-3-3">
        <title>Foreword designed attention-based loss As our goal is</title>
        <p>
          to enforce the model’s attention on the butterfly, we use a
network that was explicitly implementing an attention
mechanism. For this reason, we used an attention model based on
the generic implementation of CBAM
          <xref ref-type="bibr" rid="ref10 ref9">(Woo et al. 2018b)</xref>
          .
This architecture is a good candidate for its correct
classification results on several benchmarks and because it
separates the spatial attention mask from the channel attention.
Therefore, we could penalize the high values of the spatial
attention mask located outside of the butterfly. Our loss can
be written as follows:
        </p>
        <p>L(M; S) =</p>
        <p>Pi;k;l Mik;l(1</p>
        <p>Sk;l)</p>
        <p>Pi;k;l Mik;l
with Mik;l the pixel intensity of index (k; l) along the
spatial dimension of the attention-issued mask (M ) for the
channel i and Sk;l the pixel’s value of index (k; l) along the
spatial dimension of the segmentation-issued mask (S).</p>
        <p>
          An attention-based loss is applied at each level to the
attention-issued masks (M ) computed spatially, at low
scale, with the segmentation-issued masks (S) max-pooled
to the correct spatial dimension (e.g., for M 2 [28
28] and S 2 [250 250], a max-pooling is applied to (S)).
Experimental set-up For our preliminary experiments,
we use a ResNet (He et al. 2016) pre-trained on Imagenet
(Deng et al. 2009). To prevent over-fitting, the weight
decay parameter was finetuned and a dropout layer
          <xref ref-type="bibr" rid="ref6">(Srivastava
et al. 2014)</xref>
          was added before the last fully-connected layer
with a keep-probability. Random rotation, flipping,
rescaling and cropping were added for data-augmentation during
training. The best weights on the validation set were saved
and the training was interrupted when no improvement was
noticed for more than 50 epochs. To obtain preliminary
results without changing the class-balancing parameters, we
restrained ourselves to a reduced dataset that was perfectly
balanced containing less than 100 species.
        </p>
        <p>We compare the proposed approach with the model
without attention mechanism (original ResNet) to the model with
attention mechanism trained without guided attention loss.
We witness the importance of the attention mechanism in the
classification task, highlighting the potential of the guided
attention approach. Our proposed approach yields almost as
good in top 1 and better in top 3 accuracies than ResNet with
CBAM.</p>
        <p>Analysis Our top 1 accuracy scores in training are near
perfect (better than 99% for the three models) which means
our loss is ineffective due to over-fitting. We will use more
training data to address the class imbalance in future work
as well as using an adaptive sampling strategy. Our strategy
with adaptive sampling is to use our uncertainty prediction
measure in our approach, called Over-CAM (Kantor et al.
2020): our measure rejects the predictions if the overlap
between binarized attention (or transformed saliency maps)
and object segmentation is not satisfactory. Indeed, this case
implies that the network likely based its prediction on at least
some regions outside the butterfly, i.e., both the background
and foreground. Then, we determine the overlap distribution
on the whole test set. Following that, several thresholds are
chosen regarding the distribution curve to determine from
which percentage we could ensure a corresponding certainty
of prediction.</p>
        <p>
          Indeed, with a correct prediction and a good overlap on
a given picture, it is reasonable to under-weigh this sample
in our training set. Furthermore, a good prediction with a
low overlap would mean the decision is based on irrelevant
features and therefore under-weighting the image can even
benefit the training. Indeed, we can imagine that the wrongly
used features would be forgotten later in training. Finally, we
can augment the weight of the images incorrectly predicted,
similarly to a hard-negative mining strategy (HNM)
(Felzenszwalb et al. 2009): it bases the sampling process on the
training results for each class. This is equivalent to providing
an uncertainty measure, which we can use to ameliorate the
class imbalance problem via an image adaptative sampling.
Regularization The most straightforward solution will be
to use more training data (another training set is already at
our disposal). Our future work will be to use more training
data: one can, for example, use all the training data (with
an adaptive sampling strategy in addressing the class
imbalance) or pre-train our model on other pre-existing
butterflies datasets. If unsuccessful in addressing the
overfitting issue, our approach would be to implement
stochastic depth as a regularization method, in addition to stronger
data-augmentation, such as methods of consistency training
applied in a semi-supervised configuration as in MixMatch
          <xref ref-type="bibr" rid="ref2">(Berthelot et al. 2019)</xref>
          .
        </p>
      </sec>
      <sec id="sec-3-4">
        <title>Hierarchical Classification</title>
        <p>Our data structure presents a hierarchical property since
each label is composed of different but related items. We
should make the best use of this knowledge to improve the
results. Indeed, classifying other families should be simpler
and more robust than classifying over species. Exploiting
this hierarchy can improve robustness. For example, if the
model is uncertain between two species A and B, which
respectively belong to families 1 and 2, while being certain it
belongs to family 1, it should predict species A.</p>
        <p>Learning underlying structure while preserving flat
classification Even if these hierarchies are common in the real
world, they are challenging to leverage to improve the
classification. On the one hand, using the parents-to-children
relation seems critical to extract relevant features and reduce
parent-level classification mistakes where the task should be
easier. On the other hand, over-penalizing parent-level
relationships can cause the classifier to under-perform on leaf
classes compared to flat classification. Therefore designing
a loss that enforces the learning of the underlying
hierarchy while preserving the flat classification performance is a
challenge we plan to address.</p>
        <p>We thus propose to evaluate the impact of the
weighting of the varying elements of a hierarchical loss on the
classification performance. In addition, we introduce a loss
WCE based on cross-entropy that improves the flat
classification performance while penalizing parent-level
classification mistakes. For a given sample, we use the following
loss:
with s; g; f being the weighted coefficients for species,
genus and family, cs; cg; cf being the species, genus and
family class labels and p a probability distribution function.</p>
      </sec>
      <sec id="sec-3-5">
        <title>The general hierarchical loss</title>
        <p>Problem definitions In supervised hierarchical
classification applied to images, we consider a data-set D
containing images whose labels belong to C, a set of classes
and we assume the existence and the knowledge of an
underlying tree-structures of height d &gt; 1. The leaves of
the structure represent all the classes of C. More precisely,
each leaf is a set containing only one element of C and
each element in C is contained in a different leaf. A parent
node is defined as the union of its children. We note as Ck
the collection of sets composed of the nodes at depth k
(each node being the set composed of the classes
descending from it). This way, we have C0, the root of the tree, a
collection of sets which union is equal to C. We assume that:
• 8c 2 C; 9! c0 2 Cd; c 2 c0
• 8c0 2 Cd; jc0j = 1
• 81
• 81
i
i
d; 8c 2 Ci; 8 c0 2 Ci 1; c \ c0 6= ; ) c
c0
d; 8c 2 Ci; 9! c0 2 Ci 1; c
c0</p>
        <p>Under these assumptions, it is possible to assess the
importance of a classification error. Given an image I 2 D
and its corresponding labels cI 2 C and considering the
prediction tI 2 C, we define the importance of the error :
k(I; t) = d maxf0 n d j 9c 2 Cn; tI 2 c ^ cI 2 cg.
We note that if tI = cI , we have k(I; t) = 0.</p>
        <p>With this setting, we are interested in learning that a
classifier reduces the number and the importance of the errors.
We will note p the predicted probability; it is defined over
the leaves of the structure and can be extended to every node
considering each parent node’s construction principle.
Indeed, at each level, every node has zero intersection with
its siblings and therefore, the predicted probability of a
parent node is equal to the sum of the predicted probability of
its children.</p>
        <p>Weighting-Impact method A natural and straightforward
approach to learn the hierarchical structure is through a
weighted classification loss. We consider the cross-entropy
loss for the nodes at each depth in the tree structure. For a
depth k 2 f1; ::; dg, we consider the cross entropy loss at
this depth defined as follow :</p>
        <p>CE(k; I) = 1cI 2c^ c2Ck (c) log(p(c))
with p the predicted probability of a node c.</p>
        <p>Given a tuple of weights = ( 1; : : : ; d) 2 R+d, we
compute the weighted cross-entropy:
d
W CE (I) = X j CE(j; I)</p>
        <p>j=1</p>
        <p>This weighted cross entropy loss is differentiable and
allows the optimization of the weights of a CNN through
gradient descent.</p>
        <p>Cross-entropy loss limitation It is critical to tune the
parameter properly, which requires a time-consuming
optimization or some expert knowledge. Moreover, the
crossentropy loss has inherent limitations that need to be
addressed for proper hierarchical learning. Indeed, as the
model weights converge during training, the predicted
probability of the target class converges to 1. Moreover, given
the labels’ underlying structure, the parent node’s predicted
probability is always greater than its children’s. Since the
cross-entropy loss is expressed as log p(ci), with ci the
label class, its gradient regarding p has a magnitude that
decreases as p augments. As a result, the cross-entropy loss
naturally under-weighs the optimization of parent-level
features with respect to the children and requires a weighting.
For this reason, we propose a loss in which gradient
magnitude is not decreasing while getting closer to 1. The
divergence in 0 implies a small impact of the weighting and the
convergence to 0 in 1 implies importance of the species
Loss properties We designed a loss function with the
shape shown in Figure 1. When both probabilities are close
to 1, it yields 0 and when probabilities are both close to 0,
it yields 1. The essence of that idea is that when the
gradient magnitude of the loss is close to 0, we have a gradient
magnitude higher in the direction of genus rather than
families. The reverse is observed when it is close to 1. Such
penalization is selected to hinder optimization on the genus
if family’s optimization is affected negatively.</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>Conclusion</title>
      <p>In this article, we propose a method for fine-grained
recognition and classification of wildlife images. In particular, we
propose to guide the convolutional neural networks by
leveraging attention masks along with segmentation as a means
to being less sensitive to the typical detail-rich environment.
This work shows improved results in top-3 accuracy in
comparison to the state of the art. Furthermore, we explore the
use of a hierarchical loss to leverage species philology. Our
approach is general enough to be adapted in broader
finegrained classification contexts. Our methodology can be of
great use for large-scale wildlife crowd-sourcing programs
that gather crucial census data to understand species
demographics and dynamics.</p>
      <p>Deng, J.; Dong, W.; Socher, R.; Li, L.-J.; Li, K.; and
FeiFei, L. 2009. ImageNet: A Large-Scale Hierarchical Image
Database. In CVPR09.</p>
      <p>Felzenszwalb, P. F.; Girshick, R. B.; McAllester, D.; and
Ramanan, D. 2009. Object detection with discriminatively
trained part-based models. IEEE transactions on pattern
analysis and machine intelligence 32(9): 1627–1645.
He, K.; Zhang, X.; Ren, S.; and Sun, J. 2016. Deep
residual learning for image recognition. In Proceedings of the
IEEE conference on computer vision and pattern
recognition, 770–778.</p>
      <p>Horn, G. V.; Aodha, O. M.; Song, Y.; Shepard, A.; Adam,
H.; Perona, P.; and Belongie, S. J. 2017. The iNaturalist
Challenge 2017 Dataset. CoRR abs/1707.06642. URL http:
//arxiv.org/abs/1707.06642.</p>
      <p>Kantor, C.; Rauby, B.; Boussioux, L.; Jehanno, E.; and
Talbot, H. 2020. Over-CAM : Gradient-Based Localization and
Spatial Attention for Confidence Measure in Fine-Grained
Recognition using Deep Neural Networks. doi:10.1109/
ICCV.2017.322. URL
https://hal.archives-ouvertes.fr/hal02974521. Working paper or preprint.</p>
      <p>Kosmopoulos, A.; Partalas, I.; Gaussier, E.; Paliouras, G.;
and Androutsopoulos, I. 2020. Evaluation Measures for
Hierarchical Classification: a Unified View and Novel
Approaches. URL http://www2.aueb.gr/users/ion/docs/dami
final manuscript.pdf.</p>
      <p>Lin, Z.; Feng, M.; Dos Santos, C.; Yu, M.; Xiang, B.; Zhou,
B.; and Bengio, Y. 2017. A Structured Self-attentive
Sentence Embedding .</p>
      <p>Parikh, A.; Ta¨ckstro¨m, O.; Das, D.; and Uszkoreit, J. 2016.
A Decomposable Attention Model for Natural Language
Inference. 2249–2255. doi:10.18653/v1/D16-1244.
Prudic, K. L.; McFarland, K. P.; Oliver, J. C.; Hutchinson,
R. A.; Long, E. C.; Kerr, J. T.; and Larrive´e, M. 2017.
eButterfly: Leveraging Massive Online Citizen Science for
Butterfly Conservation. Insects 8(2). ISSN 2075-4450.
doi:10.3390/insects8020053. URL https://www.mdpi.com/
2075-4450/8/2/53.</p>
      <p>Ronneberger, O.; Fischer, P.; and Brox, T. 2015. U-net:
Convolutional networks for biomedical image segmentation. In
International Conference on Medical image computing and
computer-assisted intervention, 234–241. Springer.
Roy, D.; Panda, P.; and Roy, K. 2019. Tree-CNN: A
Hierarchical Deep Convolutional Neural Network for Incremental
Learning. URL https://arxiv.org/pdf/1802.05800.pdf.
Selvaraju, R. R.; Cogswell, M.; Das, A.; Vedantam, R.;
Parikh, D.; and Batra, D. 2017. Grad-cam: Visual
explanations from deep networks via gradient-based localization. In
Proceedings of the IEEE international conference on
computer vision, 618–626.</p>
      <p>Song, K.; Wei, X.; Shu, X.; Song, R.; and Lu, J. 2020.
BiModal Progressive Mask Attention for Fine-Grained
Recog</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          <string-name>
            <surname>Bahdanau</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Cho</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ; and Bengio,
          <string-name>
            <surname>Y.</surname>
          </string-name>
          <year>2015</year>
          .
          <article-title>Neural Machine Translation by Jointly Learning to Align and Translate</article-title>
          .
          <source>CoRR abs/1409</source>
          .0473.
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          <string-name>
            <surname>Berthelot</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Carlini</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Goodfellow</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Oliver</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Papernot</surname>
          </string-name>
          , N.; and
          <string-name>
            <surname>Raffel</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          <year>2019</year>
          .
          <article-title>MixMatch: A Holistic Approach to Semi-Supervised Learning</article-title>
          . URL https://arxiv.org/ pdf/
          <year>1905</year>
          .02249.pdf.
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          <string-name>
            <surname>Boussioux</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Giro-Larraz</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Guille-Escuret</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Cherti</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ; and Ke´gl,
          <string-name>
            <surname>B.</surname>
          </string-name>
          <year>2019</year>
          .
          <article-title>InsectUp: Crowdsourcing Insect Observations to Assess Demographic Shifts and Improve Classification</article-title>
          . URL https://arxiv.org/pdf/
          <year>1906</year>
          .11898.pdf.
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          <string-name>
            <surname>Chai</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Lempitsky</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          ; and
          <string-name>
            <surname>Zisserman</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <year>2013</year>
          .
          <article-title>Symbiotic segmentation and part localization for fine-grained categorization</article-title>
          .
          <source>In Proceedings of the IEEE International Conference on Computer Vision</source>
          ,
          <fpage>321</fpage>
          -
          <lpage>328</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          <string-name>
            <surname>Chen</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Bai</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          ; Zhang, W.; and Mei,
          <string-name>
            <surname>T.</surname>
          </string-name>
          <year>2019</year>
          .
          <article-title>Destruction and construction learning for fine-grained image recognition</article-title>
          .
          <source>IEEE Transactions on Image Processing</source>
          <volume>29</volume>
          :
          <fpage>7006</fpage>
          -
          <lpage>7018</lpage>
          . doi:
          <volume>10</volume>
          .1109/TIP.
          <year>2020</year>
          .
          <volume>2996736</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          <string-name>
            <surname>Srivastava</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Hinton</surname>
          </string-name>
          , G.;
          <string-name>
            <surname>Krizhevsky</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Sutskever</surname>
            ,
            <given-names>I.;</given-names>
          </string-name>
          and Salakhutdinov,
          <string-name>
            <surname>R.</surname>
          </string-name>
          <year>2014</year>
          .
          <article-title>Dropout: A Simple Way to Prevent Neural Networks from Overfitting</article-title>
          .
          <source>Journal of Machine Learning Research</source>
          <volume>15</volume>
          (
          <issue>56</issue>
          ):
          <fpage>1929</fpage>
          -
          <lpage>1958</lpage>
          . URL http: //jmlr.org/papers/v15/srivastava14a.html.
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          <string-name>
            <surname>Vaswani</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ; et al.
          <year>2017</year>
          .
          <article-title>Attention is All you Need</article-title>
          . In Guyon, I.; Luxburg, U. V.;
          <string-name>
            <surname>Bengio</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ; Wallach,
          <string-name>
            <surname>H.</surname>
          </string-name>
          ; Fergus,
          <string-name>
            <surname>R.</surname>
          </string-name>
          ; Vishwanathan,
          <string-name>
            <surname>S.</surname>
          </string-name>
          ; and Garnett, R., eds.,
          <source>Advances in Neural Information Processing Systems</source>
          <volume>30</volume>
          ,
          <fpage>5998</fpage>
          -
          <lpage>6008</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          <string-name>
            <given-names>Curran</given-names>
            <surname>Associates</surname>
          </string-name>
          , Inc. URL http://papers.nips.cc/paper/ 7181-attention
          <article-title>-is-all-you-need</article-title>
          .pdf.
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          <string-name>
            <surname>Woo</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ; Park, J.;
          <string-name>
            <surname>Lee</surname>
            , J.-Y.; and
            <given-names>So</given-names>
          </string-name>
          <string-name>
            <surname>Kweon</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          <year>2018a</year>
          .
          <article-title>Cbam: Convolutional block attention module</article-title>
          .
          <source>In Proceedings of the European conference on computer vision (ECCV)</source>
          ,
          <fpage>3</fpage>
          -
          <lpage>19</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          <string-name>
            <surname>Woo</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ; Park, J.;
          <string-name>
            <surname>Lee</surname>
            , J.-Y.; and
            <given-names>So</given-names>
          </string-name>
          <string-name>
            <surname>Kweon</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          <year>2018b</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          <string-name>
            <surname>Wu</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Tygert</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ; and LeCun,
          <string-name>
            <surname>Y.</surname>
          </string-name>
          <year>2019</year>
          .
          <article-title>A hierarchical loss and its problems when classifying non-hierarchically</article-title>
          . URL https://arxiv.org/pdf/1709.01062.pdf.
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          <string-name>
            <surname>Xie</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Tian</surname>
            ,
            <given-names>Q.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Hong</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ; Yan,
          <string-name>
            <surname>S.</surname>
          </string-name>
          ; and Zhang,
          <string-name>
            <surname>B.</surname>
          </string-name>
          <year>2013</year>
          .
          <article-title>Hierarchical part matching for fine-grained visual categorization</article-title>
          .
          <source>In Proceedings of the IEEE International Conference on Computer Vision</source>
          ,
          <fpage>1641</fpage>
          -
          <lpage>1648</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          <string-name>
            <surname>Zheng</surname>
          </string-name>
          , H.;
          <string-name>
            <surname>Fu</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Zha</surname>
            ,
            <given-names>Z.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Luo</surname>
            , J.; and Mei,
            <given-names>T.</given-names>
          </string-name>
          <year>2020</year>
          .
          <article-title>Learning Rich Part Hierarchies With Progressive Attention Networks for Fine-Grained Image Recognition</article-title>
          .
          <source>IEEE Transactions on Image Processing</source>
          <volume>29</volume>
          :
          <fpage>476</fpage>
          -
          <lpage>488</lpage>
          . doi:
          <volume>10</volume>
          .1109/TIP.
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          <string-name>
            <surname>Zhou</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ; and
          <string-name>
            <surname>Lin</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          <year>2016</year>
          .
          <article-title>Fine-grained image classification by exploring bipartite-graph labels</article-title>
          .
          <source>In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition</source>
          ,
          <fpage>1124</fpage>
          -
          <lpage>1133</lpage>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>