<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Distribution-restrained Softmax Loss for the Model Robustness</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Hao Wang</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Chen Li</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Jinzhe Jiang</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Xin Zhang</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Yaqian Zhao</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Weifeng Gong</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff4">4</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Inspur Electronic Information Industry Co., Ltd</institution>
          ,
          <addr-line>1036 Langchao Rd., Jinan</addr-line>
          ,
          <country country="CN">China</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>School of Computer and Artificial Intelligence, Zhengzhou University</institution>
          ,
          <addr-line>100 Science Rd., Zhengzhou</addr-line>
          ,
          <country country="CN">China</country>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>Shandong Hailiang Information Technology Institutes</institution>
          ,
          <addr-line>1768 Xinli St., Jinan</addr-line>
          ,
          <country country="CN">China</country>
        </aff>
        <aff id="aff3">
          <label>3</label>
          <institution>State Key Laboratory of High-End Server &amp; Storage Technology</institution>
          ,
          <addr-line>1036 Langchao Rd., Jinan</addr-line>
          ,
          <country country="CN">China</country>
        </aff>
        <aff id="aff4">
          <label>4</label>
          <institution>Zhengzhou Yunhai Information Technology Co., Ltd</institution>
          ,
          <addr-line>278 Xinyi Rd., Zhengzhou</addr-line>
          ,
          <country country="CN">China</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>Recently, the issue of robustness in deep learning models has garnered considerable attention, prompting the development of diverse methods aimed at enhancing model robustness. These approaches encompass adversarial training, architectural modifications, the design of novel loss functions, as well as certified defenses, among others. Despite these efforts, a comprehensive understanding of the underlying principles governing robustness against attacks remains elusive, and related research in this area is still not sufficient. Here, we have identified a significant factor that affects the robustness of models: the distribution characteristics of softmax values for non-real label samples. We found that the results after an attack are highly correlated with the distribution characteristics. Leveraging this observation, we introduce a novel loss function that effectively mitigates the diversity in softmax distribution. Extensive experimental evaluations demonstrate that our proposed method significantly enhances model robustness without significant time consumption. These findings are poised to contribute valuable insights to the realm of AI safety.</p>
      </abstract>
      <kwd-group>
        <kwd>1 Loss function</kwd>
        <kwd>Adversarial attack</kwd>
        <kwd>model robustness</kwd>
        <kwd>AI safety</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>
        While deep neural networks (DNNs) have
demonstrated remarkable performance in a
variety of applications including computer vision
[
        <xref ref-type="bibr" rid="ref1">1</xref>
        ], speech recognition [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ], and natural language
processing [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ], they are vulnerable to adversarial
attacks, which involve the addition of small
perturbations to input examples resulting in
incorrect results with high confidence [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ]-[
        <xref ref-type="bibr" rid="ref8">8</xref>
        ]. As
DNNs are being widely used in various domains,
ensuring their security by improving their
robustness against adversarial attacks has become
a critical research area.
      </p>
      <p>
        Lots of defense techniques have been
developed to enhance the adversarial robustness
of DNNs [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ]-[
        <xref ref-type="bibr" rid="ref15">15</xref>
        ]. A recent trend in adversarial
defense is the use of certified defenses [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ][
        <xref ref-type="bibr" rid="ref17">17</xref>
        ],
which attempt to provide a guarantee that the
model will not be fooled within a norm ball radius
around original images. Nevertheless, this type of
certified defense have some limitations, such as
high computational cost, lower robustness and
higher requirements for data distribution.
Adversarial training has been identified as the
most effective approach [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ][
        <xref ref-type="bibr" rid="ref10">10</xref>
        ]. Nonetheless, it is
important to acknowledge that the adversarial
training techniques is time-consuming and entails
significant computational resources, often
resulting in an adverse impact on standard
accuracy metrics [
        <xref ref-type="bibr" rid="ref18">18</xref>
        ]. Numerous investigations
have observed the existence of a trade-off
relationship between standard accuracy and
robust accuracy for adversarial training
[
        <xref ref-type="bibr" rid="ref10">10</xref>
        ],[
        <xref ref-type="bibr" rid="ref18">18</xref>
        ][
        <xref ref-type="bibr" rid="ref20">20</xref>
        ].
      </p>
      <p>
        Conversely, a number of studies show the
evidence indicating that the insights derived
solely from adversarial training are not
universally reliable or infallible [
        <xref ref-type="bibr" rid="ref21">21</xref>
        ][
        <xref ref-type="bibr" rid="ref22">22</xref>
        ]. Gillmer
et al. [
        <xref ref-type="bibr" rid="ref21">21</xref>
        ] pointed out that under the given setting,
even small standard errors imply that most points
can be proven to have misclassified points in their
neighbouring region. In this setting, achieving
perfect standard accuracy, which can be easily
implemented with a simple classifier, is sufficient
to achieve perfect adversarial robustness. Starting
from Gaussian mixture model, Hu et al. [
        <xref ref-type="bibr" rid="ref22">22</xref>
        ]
revealed two distinct effects: the first effect is a
direct consequence of the constraint of adversarial
robustness, which results in a degradation of the
standard accuracy due to the optimizing direction
change. The second effect is related to the class
imbalance ratio between the two classes being
considered, which leads to an increase in the
difference of accuracy compared to standard
training due to a reduction of “norm”.
      </p>
      <p>
        Hence, to overcome the standard accuracy
problems, other methods have been proposed
[
        <xref ref-type="bibr" rid="ref4">4</xref>
        ][
        <xref ref-type="bibr" rid="ref11">11</xref>
        ]-[
        <xref ref-type="bibr" rid="ref15">15</xref>
        ]. Goodfellow et al. [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ] demonstrated
that radial basis activation functions are more
resistant to perturbations, but their deployment
requires significant modifications to existing
architectures. Papernot et al. [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ] proposed a
method to enhance the DNN perturbation
robustness using distilled models. However, this
method exhibits certain limitations: (1) it requires
dual training, which is costly, and (2) theoretically,
the second model cannot be more accurate than
the first model, which means that it will inevitably
lead to some destroy of accuracy (although the test
in the paper shows that the accuracy actually
improved after distillation on CIFAR10, possibly
because the baseline accuracy was not high, only
81.39%).
      </p>
      <p>
        Wang et al. [
        <xref ref-type="bibr" rid="ref23">23</xref>
        ] proposed to use dropout
during inference to introduce stochasticity and
defend against adversarial attacks. On the other
hand, Gu et al. [
        <xref ref-type="bibr" rid="ref24">24</xref>
        ] argued that the key issue of
adversarial defense is to propose a suitable
training process and objective function that can
effectively enable the network to learn invariant
regions around the training data. To this end, they
proposed a deep contractive network to explicitly
learn invariant features at each layer and restrict
the change of dy with respect to dx by adding a
term ||dy/dx||2 to the loss function, ensuring that
the perturbations on x have little impact on y.
They showed some promising preliminary results
[
        <xref ref-type="bibr" rid="ref24">24</xref>
        ]. However, this penalty limits the ability of the
deep contractive network compared to traditional
DNNs [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ].
      </p>
      <p>
        It is worth noting that Rice et al. [
        <xref ref-type="bibr" rid="ref25">25</xref>
        ] argued
that currently no method in isolation improves
distinctly than early stopping. Additionally, Wu et
al. [
        <xref ref-type="bibr" rid="ref26">26</xref>
        ] pointed out that early stopping can lead to
a flatter weight loss landscape, which can result in
a smaller robust generalization gap. But only if the
training process is sufficiently, it can be beneficial
to the test robustness.
      </p>
      <p>
        In this paper, we present a novel approach for
enhancing adversarial robustness through the
introduction of a distribution-restrained softmax
loss. Through extensive experiments involving
adversarial attack tests, we find that pre- and
postattack softmax are highly correlated. Leveraging
this observed relationship, we propose a loss
function designed to mitigate the diversity in
softmax distribution. Leveraging this observed
relationship, we propose a loss function designed
to mitigate the diversity in softmax distribution.
Although our research shares certain similarities
with previous studies [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ][
        <xref ref-type="bibr" rid="ref26">26</xref>
        ], we believe that our
loss function serves as a valuable complement to
address distinct scenarios. Further comparative
discussion will be conducted later.
      </p>
    </sec>
    <sec id="sec-2">
      <title>2. Method</title>
      <p>In this section, we introduce the method of
iterative optimization to generate a transformation
robust visualization images. In general, the input
image is transformed by a certain operation, and
the optimization is performed on this transformed
image. Following, the transformation invariance
is tested on the concerned model. The process is
carried out iteratively until the convergence
condition is achieved.</p>
      <p>Previous versions of optimization with the
zero image produced less recognizable images,
while our method can give more informational
visualization results.
2.1.</p>
    </sec>
    <sec id="sec-3">
      <title>Adversarial Attack</title>
      <p>
        The target for adversary is to find an
adversarial example  ′ that can fool DNNs to
make incorrect predictions. To make it
unconspicuous for human,  ′ should not be far
away from   , by ‖ ′ −   ‖ ≤  . There are
many types of attacks have been proposed
[
        <xref ref-type="bibr" rid="ref27">27</xref>
        ][
        <xref ref-type="bibr" rid="ref30">30</xref>
        ]. To verify our method, Iterative Fast Gradient
Sign Method (I-FGSM) and Projected Gradient
Descent (PGD) method are used. Only the
IFGSM is shown in the main text for brevity.
      </p>
      <p>Fast Gradient Sign Method (FGSM). FGSM
ϵ along the gradient direction:
perturbs the natural example   by the step size of
 ′ =   +  ∙ 
(∇   
( (  ),   ))
(1)
where f is the function of DNN model.</p>
      <p>
        I-FGSM. Also, FGSM can be extended to an
iterative version [
        <xref ref-type="bibr" rid="ref30">30</xref>
        ], which has been proposed by
Kurakin et al.:
  +1 = 

 , {
      </p>
      <p>+ ∙ 
(∇  
( (  ),   ))} (2)
where N is the N-th iteration, 
 , indicates
the attacked image is clipped within the ϵ-ball of
the last step, α is the value of perturbation.</p>
      <p>Projected Gradient Descent (PGD). PGD is an
iterative method that perturbs the natural example
  by the certain value of η，and after each step
of perturbation, it projects the adversarial example
back to the adjacent of   :
 ′( +1) = ∏ (</p>
      <p>′( )
+ ∙ 
(∇ ′ ( ( 
′( )) ,   ))) (3)
where L is the loss function, ∏ (∙) is the
the adversarial attack.
projection operation, and  
′( ) is the k-th step of
2.2.</p>
    </sec>
    <sec id="sec-4">
      <title>Distribution-restrained</title>
    </sec>
    <sec id="sec-5">
      <title>Softmax Loss Function (DRSL)</title>
      <p>It is common to use cross entropy as a loss</p>
      <p>( ( ;  )) (4)
function for DNN:
 ( ( ;  ),  ) = −1 log(
where θ is the set of parameters of the classifier,
f(∙) is the function defined by DNN, and 1
denotes the one-hot encoding of y.</p>
      <p>
        However, there are some studies indicate that
the softmax cross entropy doesn’t guarantee a
good robustness [
        <xref ref-type="bibr" rid="ref26">26</xref>
        ],[
        <xref ref-type="bibr" rid="ref31">31</xref>
        ],[
        <xref ref-type="bibr" rid="ref37">37</xref>
        ]-[
        <xref ref-type="bibr" rid="ref40">40</xref>
        ]. Following
these works, we investigate the pre- and
postattack softmax probabilities on the
dataset [
        <xref ref-type="bibr" rid="ref35">35</xref>
        ]. Figure
1
shows the
      </p>
      <p>MNIST
softmax
probabilities on the MNIST dataset. Here, we
compares two cases: (1) the probabilities of
softmax outputs after attack, (2) the probabilities
of second largest softmax outputs before attack. It
is obvious to identify the pattern that they are
roughly similar (except for the true label 1), which
implies after attacks, the second largest softmax
probabilities trends to be the largest one. This
observation lead us to hypothesize that as the
second largest softmax probabilities decrease, it
becomes more difficult for an adversarial attack to
manipulate them into becoming the largest one.</p>
      <p>Taking this hypothesis a step further, we posit that
the distribution of softmax probabilities may
affect the robustness of the model.
before attack on the MNIST dataset. (0)-(9) Corresponds to true labels of 0-9. The probability after
attack is shown in green bar, and the second maximum softmax probabilities before attack is shown
in red bar.</p>
      <p>
        As we discussed before, trade-off between
standard and robust accuracy is not a universal
rule, especially for non- adversarial training
method [
        <xref ref-type="bibr" rid="ref21">21</xref>
        ]-[
        <xref ref-type="bibr" rid="ref24">24</xref>
        ]. Gu et al. [
        <xref ref-type="bibr" rid="ref24">24</xref>
        ] argued that the
sensitivity of neural networks to adversarial
examples is more related to inherent flaws in the
training process and objective function, rather
than the model architecture.
      </p>
      <p>To further consolidate our findings, a standard
accuracy-softmax distribution experiment is
performed. Here, we adopt distance as a metric for
stochastic of softmax distribution.</p>
      <p>
        Euclidean Distance. Euclidean Distance is a
well-used distance measure in the
multidimensions of space [
        <xref ref-type="bibr" rid="ref41">41</xref>
        ]. It is used to compare the
absolute distances between two points in the
dimensions of space:
      </p>
      <p>= √∑ =1(  −   )2 (5)
where a and b are both n dimensional vectors.</p>
      <p>In the metric for stochasticity of softmax
distribution case, the vector a and b are the real
distribution of the output softmax of models and
the ideal distribution, respectively. Here, we
define the ideal case as a totally average
distribution (For example, if there is a
fourcategory classification, the ideal average
distribution of the output softmax should be [0.25,
0.25, 0.25, 0.25]).</p>
      <p>
        Two architectures of DNN are tested:
regularized Convolutional neural network VGG
[
        <xref ref-type="bibr" rid="ref32">32</xref>
        ] and multi-head attention based ViT [
        <xref ref-type="bibr" rid="ref33">33</xref>
        ]. All
the models are established in a comparable size,
which will be discussed in the Section 3. Figure 2
gives the distance metric for stochasticity of
softmax distribution. It shows our method rise the
stochasticity of softmax distribution exactly, both
on VGG and ViT models. Also, the result reveals
that there are no significant positive correlations
between standard accuracy and softmax
distribution, that is, with the model accuracy
increase, the softmax distribution doesn’t become
more sharpness. For cases of VGG, it shows even
a negative correlation.
      </p>
      <p>Inspired by the preliminary result above, we
assume the softmax distribution is a key factor of
adversarial robustness, which is perhaps
orthogonal to the standard optimizing direction
and will not harm the accuracy significantly. Here,
we propose the distribution-restrained softmax
loss function towards robust DNN models:
 ( ( ;  ),  ) = −1 log( ( ( ;  ))</p>
      <p>+ ∙  ( ( ( ;  )),  ) (3)
where d(∙,∙) is the distance function, τ is the weight
of the distance, avg is the average distribution.</p>
    </sec>
    <sec id="sec-6">
      <title>3. Experiments</title>
      <p>In this section, we show the analysis of
softmax distribution. Then cases of the robustness
results are shown. Different loss functions are
compared by adversarial robustness in Section 3.1,
respect to different models and datasets. Also, the
random noise robustness is tested in Section 3.2.</p>
      <p>
        All networks used ReLUs in the hidden layers
and softmax layers at the output. All reported
experiments were repeated five times with
random initialization of neural network
parameters. We compared the proposed functions
with Cross Entropy loss (CE), Generalized Cross
Entropy loss (GCE) [
        <xref ref-type="bibr" rid="ref31">31</xref>
        ] and ours.
      </p>
      <p>
        All experiments were conducted with
identical optimization procedures and
architectures, changing only the loss functions.
All the parameters size are restricted around 1.6
M. And the accuracy deviation of models are
restricted to 0.5% (most cases are less than 0.1%).
We conducted experiments using VGG [
        <xref ref-type="bibr" rid="ref32">32</xref>
        ] and
ViT [
        <xref ref-type="bibr" rid="ref33">33</xref>
        ] models optimized with the default
setting of Adam [
        <xref ref-type="bibr" rid="ref34">34</xref>
        ] on the MNIST [
        <xref ref-type="bibr" rid="ref35">35</xref>
        ] and
CIFAR-10 [
        <xref ref-type="bibr" rid="ref36">36</xref>
        ] datasets.
      </p>
    </sec>
    <sec id="sec-7">
      <title>3.1. Adversarial Robustness</title>
    </sec>
    <sec id="sec-8">
      <title>Different Loss Functions with</title>
      <p>Experimental results of adversarial robustness
with different loss functions are shown in Figure
3. It reveals that the Distribution-restrained
Softmax Loss outperform other loss functions in
the VGG model. Besides MNIST dataset,
CIFAR10 are also used to train the models. In comparison,
we also tested the loss functions for the ViT model.
It can be observed that the advantage holds in the
well-used ViT model.</p>
      <p>
        Although our idea is based on analysis of
adversarial perturbation, we are also curious about
the effect of the noise. That’s because the
perturbation can be regard as a kind of noise,
many works have pay close attention to their
relation [
        <xref ref-type="bibr" rid="ref42">42</xref>
        ]-[
        <xref ref-type="bibr" rid="ref44">44</xref>
        ]. In this part, we will show a case
study of label noise robustness with different loss
functions [
        <xref ref-type="bibr" rid="ref45">45</xref>
        ]. To visually view the effect of loss
function, output dimensions of DNNs are reduced
from 10 to 2. Here, several dimensionality
reduction method are used [
        <xref ref-type="bibr" rid="ref46">46</xref>
        ]. Figure 4 shows
the result in MNIST dataset by t-SNE method.
      </p>
      <p>As shown in Figure 4(c), with the noise
intensity increase, the test accuracy of all the
models decrease. While the DRSL maintains the
accuracy better than others. To make a further
understanding of this phenomenon,
dimensionality reduction methods are applied to
visualize the softmax distribution. As shown in
Figure 4(d) we can find that after a dimensionality
reduction, DRSL gives a distinctive structure on
the softmax distribution compared to others.</p>
      <p>Difference to clusters of CE and GCE that are
“caterpillar-like”, DRSL gives “nematode-like”
clusters. After attacks, adversarial examples of
other two models are distributed in the reduced
space, while DRSL’s are concentrated at the tip of
clusters (Shown in Figure 4(d), second row.
Adversarial examples are shown as black dots).
That makes our method promising to discriminate
adversarial examples, which need a further study
in future. And under a label noise disturbed,
models with other loss functions show a scene of
chaos, while DRSL maintain a curve segment
shape in the reduced space, which is relatively
easy to distinguish of different categories.
Therefore, this analysis shows that the DRSL
function is not only helpful to adversarial
robustness, but can also capture the label noise
robustness.</p>
    </sec>
    <sec id="sec-9">
      <title>4. Conclusions</title>
      <p>Robustness is a vital property for DNN models.
In this paper, we have identified a significant
factor that affects the robustness of models: the
distribution characteristics of softmax values for
non-real label samples. We found that the results
after an attack are highly correlated with the
distribution characteristics. And after the
distribution diversity of softmax is suppressed in
loss function, we find a significant improvement
of model robustness. Although there are already
some techniques to address model robustness, we
believe that our loss function can serve as a
valuable complement. For instance, after
dimensionality reduction, DRSL show its
potential to recognize the contaminated data. Also,
DRSL can be applied not only to classification
models but also to other softmax-inclusive models,
such as generative models, which inspires us to
further investigate and explore of the method.</p>
    </sec>
    <sec id="sec-10">
      <title>5. References</title>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>Kaiming</given-names>
            <surname>He</surname>
          </string-name>
          , Xiangyu Zhang, Shaoqing Ren, and
          <string-name>
            <given-names>Jian</given-names>
            <surname>Sun</surname>
          </string-name>
          .
          <article-title>Deep residual learning for image recognition</article-title>
          .
          <source>In CVPR</source>
          ,
          <year>2016</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>Yisen</given-names>
            <surname>Wang</surname>
          </string-name>
          , Xuejiao Deng, Songbai Pu, and
          <string-name>
            <given-names>Zhiheng</given-names>
            <surname>Huang</surname>
          </string-name>
          .
          <article-title>Residual convolutional ctc networks for automatic speech recognition</article-title>
          .
          <source>arXiv preprint arXiv:1702.07793</source>
          ,
          <year>2017</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>Ashish</given-names>
            <surname>Vaswani</surname>
          </string-name>
          , Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones,
          <string-name>
            <given-names>Aidan N.</given-names>
            <surname>Gomez</surname>
          </string-name>
          , Lukasz Kaiser,
          <string-name>
            <given-names>Illia</given-names>
            <surname>Polosukhin</surname>
          </string-name>
          .
          <article-title>Attention is all you need</article-title>
          .
          <source>arXiv preprint arXiv:1706.03762</source>
          ,
          <year>2017</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <surname>Ian</surname>
            <given-names>J Goodfellow</given-names>
          </string-name>
          , Jonathon Shlens, and
          <string-name>
            <given-names>Christian</given-names>
            <surname>Szegedy</surname>
          </string-name>
          .
          <article-title>Explaining and harnessing adversarial examples</article-title>
          .
          <source>In ICLR</source>
          ,
          <year>2015</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>A.</given-names>
            <surname>Kurakin</surname>
          </string-name>
          , I. Goodfellow,
          <string-name>
            <given-names>S.</given-names>
            <surname>Bengio</surname>
          </string-name>
          ,
          <article-title>Adversarial examples in the physical world</article-title>
          ,
          <source>arXiv preprint arXiv:1607.02533</source>
          ,
          <year>2016</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>S.</given-names>
            <surname>Komkov</surname>
          </string-name>
          and
          <string-name>
            <given-names>A.</given-names>
            <surname>Petiushko</surname>
          </string-name>
          ,
          <source>AdvHat: RealWorld Adversarial Attack on ArcFace Face ID System, 2020 25th International Conference on Pattern Recognition (ICPR)</source>
          ,
          <year>2021</year>
          , pp.
          <fpage>819</fpage>
          -
          <lpage>826</lpage>
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>I.</given-names>
            <surname>Evtimov</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Eykholt</surname>
          </string-name>
          , E. Fernandes,
          <string-name>
            <given-names>T.</given-names>
            <surname>Kohno</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Prakash</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Rahmati</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Song</surname>
          </string-name>
          ,
          <source>Robust Physical-World Attacks on Deep Learning Models, arXiv preprint arXiv:1707.08945</source>
          ,
          <year>2017</year>
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>N.</given-names>
            <surname>Morgulis</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Kreines</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Mendelowitz</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Weisglass</surname>
          </string-name>
          ,
          <article-title>Fooling a Real Car with Adversarial Traffic Signs</article-title>
          , arXiv preprint arXiv:
          <year>1907</year>
          .00374,
          <year>2019</year>
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>Aleksander</given-names>
            <surname>Madry</surname>
          </string-name>
          , Aleksandar Makelov, Ludwig Schmidt, Dimitris Tsipras, and
          <string-name>
            <given-names>Adrian</given-names>
            <surname>Vladu</surname>
          </string-name>
          .
          <article-title>Towards deep learning models resistant to adversarial attacks</article-title>
          .
          <source>In International Conference on Learning Representations (ICLR)</source>
          ,
          <year>2018</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <surname>Hongyang</surname>
            <given-names>Zhang</given-names>
          </string-name>
          , Yaodong Yu, Jiantao Jiao, Eric P Xing, Laurent El Ghaoui, and
          <string-name>
            <given-names>Michael I</given-names>
            <surname>Jordan</surname>
          </string-name>
          .
          <article-title>Theoretically principled trade-off between robustness and accuracy</article-title>
          .
          <source>In International Conference on Machine Learning (ICML)</source>
          ,
          <year>2019b</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <surname>Shixiang</surname>
            <given-names>Gu</given-names>
          </string-name>
          ,
          <string-name>
            <given-names>Luca</given-names>
            <surname>Rigazio</surname>
          </string-name>
          .
          <article-title>Toward deep neural network architectures robust to adversarial examples</article-title>
          .
          <source>In International Conference on Learning Representations (ICLR)</source>
          ,
          <year>2015</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <surname>Nicolas</surname>
            <given-names>Papernot</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Patrick</surname>
            <given-names>McDaniel</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Xi Wu</surname>
            , Somesh Jha, and
            <given-names>Ananthram</given-names>
          </string-name>
          <string-name>
            <surname>Swami</surname>
          </string-name>
          .
          <article-title>Distillation as a Defense to Adversarial Perturbations against Deep Neural Networks</article-title>
          .
          <source>arXiv preprint arXiv:1511.04508</source>
          ,
          <year>2015</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <surname>Siyue</surname>
            <given-names>Wang</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Xiao</surname>
            <given-names>Wang</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Pu Zhao</surname>
            ,
            <given-names>Wujie</given-names>
          </string-name>
          <string-name>
            <surname>Wen</surname>
            , David Kaeli,
            <given-names>Peter</given-names>
          </string-name>
          <string-name>
            <surname>Chin</surname>
            ,
            <given-names>Xue</given-names>
          </string-name>
          <string-name>
            <surname>Lin</surname>
          </string-name>
          .
          <article-title>Defensive Dropout for Hardening Deep Neural Networks under Adversarial Attacks</article-title>
          . arXiv preprint arXiv:
          <year>1809</year>
          .05165,
          <year>2018</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <given-names>A.</given-names>
            <surname>Raghunathan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Steinhardt</surname>
          </string-name>
          , and
          <string-name>
            <given-names>P.</given-names>
            <surname>Liang</surname>
          </string-name>
          .
          <article-title>Certified defenses against adversarial examples</article-title>
          .
          <source>ICLR</source>
          ,
          <year>2018</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <given-names>E.</given-names>
            <surname>Wong</surname>
          </string-name>
          and
          <string-name>
            <given-names>Z.</given-names>
            <surname>Kolter</surname>
          </string-name>
          .
          <article-title>Provable defenses against adversarial examples via the convex outer adversarial polytope</article-title>
          .
          <source>in International Conference on Machine Learning. PMLR</source>
          ,
          <year>2018</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [16]
          <string-name>
            <given-names>A.</given-names>
            <surname>Raghunathan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Steinhardt</surname>
          </string-name>
          , and
          <string-name>
            <given-names>P.</given-names>
            <surname>Liang</surname>
          </string-name>
          .
          <article-title>Certified defenses against adversarial examples</article-title>
          .
          <source>ICLR</source>
          ,
          <year>2018</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [17]
          <string-name>
            <given-names>E.</given-names>
            <surname>Wong</surname>
          </string-name>
          and
          <string-name>
            <given-names>Z.</given-names>
            <surname>Kolter</surname>
          </string-name>
          .
          <article-title>Provable defenses against adversarial examples via the convex outer adversarial polytope</article-title>
          .
          <source>In International Conference on Machine Learning. PMLR</source>
          ,
          <year>2018</year>
          , pp.
          <fpage>5286</fpage>
          -
          <lpage>5295</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          [18]
          <string-name>
            <surname>Dimitris</surname>
            <given-names>Tsipras</given-names>
          </string-name>
          , Shibani Santurkar, Logan Engstrom, Alexander Turner,
          <string-name>
            <given-names>Aleksander</given-names>
            <surname>Madry</surname>
          </string-name>
          . Robustness May Be at Odds with Accuracy. arXiv preprint arXiv:
          <year>1805</year>
          .12152
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          [19]
          <string-name>
            <surname>Aditi</surname>
            <given-names>Raghunathan</given-names>
          </string-name>
          , Sang Michael Xie,
          <string-name>
            <given-names>Fanny</given-names>
            <surname>Yang</surname>
          </string-name>
          , John C. Duchi,
          <string-name>
            <given-names>Percy</given-names>
            <surname>Liang</surname>
          </string-name>
          .
          <article-title>Adversarial Training Can Hurt Generalization</article-title>
          . arXiv preprint arXiv:
          <year>1906</year>
          .06032,
          <year>2019</year>
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          [20]
          <string-name>
            <surname>Dong</surname>
            <given-names>Su</given-names>
          </string-name>
          , Huan Zhang, Hongge Chen, Jinfeng Yi,
          <string-name>
            <surname>Pin-Yu Chen</surname>
            , and
            <given-names>Yupeng</given-names>
          </string-name>
          <string-name>
            <surname>Gao</surname>
          </string-name>
          .
          <article-title>Is Robustness the Cost of Accuracy? A Comprehensive Study on the Robustness of 18 Deep Image Classification Models</article-title>
          . arXiv preprint arXiv:
          <year>1808</year>
          .01688,
          <year>2018</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          [21]
          <string-name>
            <surname>Justin</surname>
            <given-names>Gilmer</given-names>
          </string-name>
          , Luke Metz, Fartash Faghri, Samuel S. Schoenholz, Maithra Raghu,
          <string-name>
            <given-names>Martin</given-names>
            <surname>Wattenberg</surname>
          </string-name>
          , &amp;
          <article-title>Ian Goodfellow. The Relationship Between High-Dimensional Geometry and Adversarial Examples</article-title>
          . arXiv preprint arXiv:
          <year>1801</year>
          .02774,
          <year>2018</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          [22]
          <string-name>
            <surname>Yuzheng</surname>
            <given-names>Hu</given-names>
          </string-name>
          , FanWu, Hongyang Zhang, Han Zhao.
          <source>Understanding the Impact of Adversarial Robustness on Accuracy Disparity. arXiv preprint arXiv:2211.15762</source>
          ,
          <year>2022</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          [23]
          <string-name>
            <surname>Siyue</surname>
            <given-names>Wang</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Xiao</surname>
            <given-names>Wang</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Pu Zhao</surname>
            ,
            <given-names>Wujie</given-names>
          </string-name>
          <string-name>
            <surname>Wen</surname>
            , David Kaeli,
            <given-names>Peter</given-names>
          </string-name>
          <string-name>
            <surname>Chin</surname>
            ,
            <given-names>Xue</given-names>
          </string-name>
          <string-name>
            <surname>Lin</surname>
          </string-name>
          .
          <article-title>Defensive Dropout for Hardening Deep Neural Networks under Adversarial Attacks</article-title>
          . arXiv preprint arXiv:
          <year>1809</year>
          .05165,
          <year>2018</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          [24]
          <string-name>
            <surname>Shixiang</surname>
            <given-names>Gu</given-names>
          </string-name>
          ,
          <string-name>
            <given-names>Luca</given-names>
            <surname>Rigazio</surname>
          </string-name>
          .
          <article-title>Towards Deep Neural Networks Architectures Robust to Adversarial Examples</article-title>
          .
          <source>In Proceedings of the 2015 International Conference on Learning Representations. Computational and Biological Learning Society</source>
          ,
          <year>2015</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref25">
        <mixed-citation>
          [25]
          <string-name>
            <surname>Leslie</surname>
            <given-names>Rice</given-names>
          </string-name>
          ,
          <string-name>
            <given-names>Eric</given-names>
            <surname>Wong</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. Zico</given-names>
            <surname>Kolter</surname>
          </string-name>
          .
          <article-title>Overfitting in adversarially robust deep learning</article-title>
          .
          <source>arXiv preprint arXiv:2002.11569</source>
          ,
          <year>2020</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref26">
        <mixed-citation>
          [26]
          <string-name>
            <surname>Dongxian</surname>
            <given-names>Wu</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Shu-Tao</surname>
            <given-names>Xia</given-names>
          </string-name>
          ,
          <string-name>
            <given-names>Yisen</given-names>
            <surname>Wang</surname>
          </string-name>
          .
          <article-title>Adversarial Weight Perturbation Helps Robust Generalization</article-title>
          .
          <source>In 34th Conference on Neural Information Processing Systems</source>
          ,
          <year>2020</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref27">
        <mixed-citation>
          [27]
          <string-name>
            <surname>Ian J Goodfellow</surname>
            , Jonathon Shlens, and
            <given-names>Christian</given-names>
          </string-name>
          <string-name>
            <surname>Szegedy</surname>
          </string-name>
          .
          <article-title>Explaining and harnessing adversarial examples</article-title>
          .
          <source>In ICLR</source>
          ,
          <year>2015</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref28">
        <mixed-citation>
          [28]
          <string-name>
            <surname>Aleksander</surname>
            <given-names>Madry</given-names>
          </string-name>
          , Aleksandar Makelov, Ludwig Schmidt, Dimitris Tsipras, and
          <string-name>
            <given-names>Adrian</given-names>
            <surname>Vladu</surname>
          </string-name>
          .
          <article-title>Towards deep learning models resistant to adversarial attacks</article-title>
          .
          <source>In ICML</source>
          ,
          <year>2018</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref29">
        <mixed-citation>
          [29]
          <string-name>
            <given-names>Nicholas</given-names>
            <surname>Carlini</surname>
          </string-name>
          and
          <string-name>
            <given-names>David</given-names>
            <surname>Wagner</surname>
          </string-name>
          .
          <article-title>Towards evaluating the robustness of neural networks</article-title>
          . In S&amp;P,
          <year>2017</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref30">
        <mixed-citation>
          [30]
          <string-name>
            <given-names>A.</given-names>
            <surname>Kurakin</surname>
          </string-name>
          ,
          <string-name>
            <surname>I. Goodfellow</surname>
          </string-name>
          , and
          <string-name>
            <surname>S. Bengio.</surname>
          </string-name>
          <article-title>Adversarial machine learning at scale</article-title>
          .
          <source>In International Conference on Learning Representations</source>
          ,
          <year>2017</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref31">
        <mixed-citation>
          [31]
          <string-name>
            <given-names>Zhilu</given-names>
            <surname>Zhang Mert R. Sabuncu</surname>
          </string-name>
          .
          <article-title>Generalized Cross Entropy Loss for Training Deep Neural Networks with Noisy Labels</article-title>
          .
          <source>In 32nd Conference on Neural Information Processing Systems</source>
          ,
          <year>2018</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref32">
        <mixed-citation>
          [32]
          <string-name>
            <surname>Simonyan</surname>
            <given-names>K</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zisserman</surname>
            <given-names>A</given-names>
          </string-name>
          .
          <article-title>Very Deep Convolutional Networks for Large-Scale Image Recognition</article-title>
          .
          <source>arXiv preprint arXiv:1409.1556</source>
          ,
          <year>2014</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref33">
        <mixed-citation>
          [33]
          <string-name>
            <surname>Alexey</surname>
            <given-names>Dosovitskiy</given-names>
          </string-name>
          , Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit,
          <string-name>
            <given-names>Neil</given-names>
            <surname>Houlsby</surname>
          </string-name>
          .
          <article-title>An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale</article-title>
          . arXiv preprint arXiv:
          <year>2010</year>
          .11929,
          <year>2020</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref34">
        <mixed-citation>
          [34]
          <string-name>
            <surname>Diederik</surname>
            <given-names>P</given-names>
          </string-name>
          <string-name>
            <surname>Kingma and Jimmy Ba</surname>
          </string-name>
          .
          <article-title>Adam: A method for stochastic optimization</article-title>
          .
          <source>arXiv preprint arXiv:1412.6980</source>
          ,
          <year>2014</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref35">
        <mixed-citation>
          [35]
          <article-title>LeCun, Yann and Cortes, Corinna and Burges, CJ. MNIST handwritten digit database</article-title>
          .
          <year>2010</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref36">
        <mixed-citation>
          [36]
          <string-name>
            <given-names>Alex</given-names>
            <surname>Krizhevsky</surname>
          </string-name>
          and
          <string-name>
            <given-names>Geoffrey</given-names>
            <surname>Hinton</surname>
          </string-name>
          .
          <article-title>Learning multiple layers of features from tiny images</article-title>
          .
          <year>2009</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref37">
        <mixed-citation>
          [37]
          <string-name>
            <surname>Tianyu</surname>
            <given-names>Pang</given-names>
          </string-name>
          , Kun Xu, Yinpeng Dong, Chao Du, Ning Chen,
          <string-name>
            <given-names>Jun</given-names>
            <surname>Zhu</surname>
          </string-name>
          .
          <article-title>Rethinking softmax cross-entropy loss for adversarial robustness</article-title>
          .
          <source>ICLR</source>
          <year>2020</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref38">
        <mixed-citation>
          [38]
          <string-name>
            <surname>Weiyang</surname>
            <given-names>Liu</given-names>
          </string-name>
          , Yandong Wen,
          <string-name>
            <given-names>Zhiding</given-names>
            <surname>Yu</surname>
          </string-name>
          , and
          <string-name>
            <given-names>Meng</given-names>
            <surname>Yang</surname>
          </string-name>
          .
          <article-title>Large-margin softmax loss for convolutional neural networks</article-title>
          .
          <source>In International Conference on Machine Learning (ICML)</source>
          ,
          <year>2016</year>
        </mixed-citation>
      </ref>
      <ref id="ref39">
        <mixed-citation>
          [39]
          <string-name>
            <surname>Hao</surname>
            <given-names>Wang</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Yitong</surname>
            <given-names>Wang</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zheng Zhou</surname>
            , Xing Ji, Dihong Gong, Jingchao Zhou,
            <given-names>Zhifeng</given-names>
          </string-name>
          <string-name>
            <surname>Li</surname>
          </string-name>
          , and Wei Liu. Cosface:
          <article-title>Large margin cosine loss for deep face recognition</article-title>
          .
          <source>In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR)</source>
          , pp.
          <fpage>5265</fpage>
          -
          <lpage>5274</lpage>
          ,
          <year>2018</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref40">
        <mixed-citation>
          [40]
          <string-name>
            <surname>Jiankang</surname>
            <given-names>Deng</given-names>
          </string-name>
          , Jia Guo, Niannan Xue, and
          <string-name>
            <given-names>Stefanos</given-names>
            <surname>Zafeiriou</surname>
          </string-name>
          . Arcface:
          <article-title>Additive angular margin loss for deep face recognition</article-title>
          .
          <source>In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR)</source>
          , pp.
          <fpage>4690</fpage>
          -
          <lpage>4699</lpage>
          ,
          <year>2019</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref41">
        <mixed-citation>
          [41]
          <string-name>
            <surname>Christopher Bishop M. Pattern Recognition</surname>
          </string-name>
          and
          <source>Machine Learning</source>
          . Springer,
          <year>2006</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref42">
        <mixed-citation>
          [42]
          <string-name>
            <surname>Akshay</surname>
            <given-names>Agarwal</given-names>
          </string-name>
          , Mayank Vatsa,
          <string-name>
            <given-names>Richa</given-names>
            <surname>Singh</surname>
          </string-name>
          , and
          <string-name>
            <surname>Nalini</surname>
            <given-names>K.</given-names>
          </string-name>
          <string-name>
            <surname>Ratha</surname>
          </string-name>
          .
          <article-title>Noise is Inside Me! Generating Adversarial Perturbations with Noise Derived from Natural Filters</article-title>
          .
          <source>In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition Workshop (CVPR)</source>
          ,
          <year>2020</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref43">
        <mixed-citation>
          [43]
          <string-name>
            <surname>Chengjun</surname>
            <given-names>Tang</given-names>
          </string-name>
          , Kun Zhang, Chunfang Xing, Yong Ding, Zengmin Xu.
          <source>Perlin Noise Improve Adversarial Robustness. arXiv preprint arXiv:2112.13408</source>
          ,
          <year>2021</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref44">
        <mixed-citation>
          [44]
          <string-name>
            <surname>Fei</surname>
            <given-names>Wu</given-names>
          </string-name>
          , Wenxue Yang, Limin Xiao and
          <string-name>
            <given-names>Jinbin</given-names>
            <surname>Zhu</surname>
          </string-name>
          .
          <article-title>Adaptive Wiener Filter and Natural Noise to Eliminate Adversarial Perturbation</article-title>
          .
          <source>Electronics</source>
          <volume>9</volume>
          ,
          <issue>1634</issue>
          ,
          <year>2020</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref45">
        <mixed-citation>
          [45]
          <string-name>
            <surname>Aritra</surname>
            <given-names>Ghosh</given-names>
          </string-name>
          , Himanshu Kumar,
          <string-name>
            <given-names>P. S.</given-names>
            <surname>Sastry</surname>
          </string-name>
          .
          <article-title>Robust Loss Functions under Label Noise for Deep Neural Networks</article-title>
          .
          <source>arXiv preprint arXiv:1712.09482</source>
          ,
          <year>2017</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref46">
        <mixed-citation>
          [46]
          <string-name>
            <surname>Maaten</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          v.d. and
          <string-name>
            <surname>Hinton</surname>
          </string-name>
          ,
          <source>G. Journal of Machine Learning Research</source>
          ,
          <volume>9</volume>
          ,
          <fpage>2579</fpage>
          -
          <lpage>2605</lpage>
          ,
          <year>2008</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>