<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Detection of Adversarial Supports in Few-Shot Classifiers U sing S elf-Similarity a nd Filtering</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Yi Xiang Marcus Tan</string-name>
          <xref ref-type="aff" rid="aff2">2</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Penny Chong</string-name>
          <xref ref-type="aff" rid="aff2">2</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Jiamei Sun</string-name>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Ngai-Man Cheung</string-name>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Yuval Elovici</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Alexander Binder</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Department of Informatics, University of Oslo</institution>
          ,
          <country country="NO">Norway</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Department of Software and Information Systems Engineering, Ben-Gurion University of the Negev</institution>
          ,
          <addr-line>Be'er Sheva</addr-line>
          ,
          <country country="IL">Israel</country>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>Information Systems Technology and Design pillar, Singapore University of Technology and Design</institution>
          ,
          <country country="SG">Singapore</country>
        </aff>
        <aff id="aff3">
          <label>3</label>
          <institution>ST Engineering-SUTD Cyber Security Laboratory</institution>
        </aff>
      </contrib-group>
      <abstract>
        <p>Few-shot classifiers excel under limited training samples, making them useful in applications with sparsely user-provided labels. Their unique relative prediction setup offers opportunities for novel attacks, such as targeting support sets required to categorise unseen test samples, which are not available in other machine learning setups. In this work, we propose a detection strategy to identify adversarial support sets, aimed at destroying the understanding of a few-shot classifier for a certain class. We achieve this by introducing the concept of self-similarity of a support set and by employing filtering of supports. Our method is attack-agnostic, and we are the first to explore adversarial detection for support sets of few-shot classifiers to the best of our knowledge. Our evaluation of the miniImagenet (MI) and CUB datasets exhibits good attack detection performance despite conceptual simplicity, showing high AUROC scores. We show that self-similarity and filtering for adversarial detection can be paired with other filtering functions, constituting a generalisable concept.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;adversarial machine learning</kwd>
        <kwd>adversarial defence</kwd>
        <kwd>detection</kwd>
        <kwd>few-shot</kwd>
        <kwd>self-similarity</kwd>
        <kwd>filtering</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>plored, albeit gaining traction [7, 8]. This is compared
to models under the standard classification setting, where
An open topic in machine learning is the transferability such phenomenon had been widely explored [9, 10]. The
of a trained model to a new set of prediction categories relative nature of predictions in few-shot setups allows
without retraining efforts, in particular when some classes going beyond crafting adversarial test samples.
have very few samples. Few-shot learning algorithms The attacker could craft adversarial perturbations for
have been proposed to address this, where prediction and all -shot support samples of the attacked class and insert
training are based on the concept of an episode. Each them into the deployment phase of the model. The goal is
episode (task) comprises several labelled training samples to misclassify test samples of the attacked class regardless
per class (i.e. 1 or 5), denoted as the support set, and of the samples drawn in the other classes. In this work,
query samples for episodic testing, called the query set. we consider the impact on the few-shot accuracy of the
Unlike other setups, the prediction in few-shot models is attacked class, in the presence of adversarial
perturbarelative to the support set classes of an episode [1, 2, 3, 4]. tions, even when different samples were drawn for the
The label categories vary in each episode and training is non-attacked classes. This is a highly realistic scenario
performed by drawing randomised sets of classes, thus it- as the victim could unknowingly draw such adversarial
erating over varying prediction tasks when learning model support sets during the evaluation phase once they were
parameters. Effectively, this learns a class-agnostic simi- inserted by the attacker. The use of adversarial samples to
larity metric which generalises to novel categories [5, 6]. attack other settings than the one trained for are known as</p>
      <p>Unfortunately, the adversarial susceptibility of mod- transferability attacks.
els under the few-shot paradigm remains relatively unex- Prior methods proposed to mitigate such adverse effects
through the lenses of detection [11, 12] and model
robustInternational Workshop on Safety &amp; Security of Deep Learning, ness [13, 14]. Though these methods work well for neural
A"ugmuastrc2u0s2ta1nyx16@gmail.com (Y. X. M. Tan); networks under the conventional classification setting,
pennychong94@gmail.com (P. Chong); sunjiamei.hit@gmail.com they will fail on few-shot classifiers due to limited data.
(J. Sun); ngaiman&lt;underscore&gt;cheung@sutd.edu.sg (N. Cheung); Furthermore, these defences were not trained to transfer
elovici@bgu.ac.il (Y. Elovici); alexabin@ifi.uio.no (A. Binder) its pre-existing knowledge towards a novel distribution
~ https://sites.google.com/site/mancheung0407/ (N. Cheung); of class samples, contrary to few-shot classifiers. With
(hAtt.psB:/i/nsdcehro)lar.google.de/citations?user&lt;equal&gt;5B8CTlEAAAAJ the aforementioned drawbacks in mind, we propose a
conCPWrEooUrckReshdoinpgs IhStpN:/c1e6u1r3-w-0s.o7r3g ©CCo2Em0m2U1onRCsoLpWiycreingoshertkAfotsrtrhtihbouistpipoanpP4e.rr0obIyncteietrsneaadtuiotihnnoarlgs(.sCUC(sBCeYpEe4r.Um0)i.tRted-WundSer.Coreragti)ve cdeeptetuctailolyn soifmapdlveemrseatrhiaoldsfuoprppoerrtfosarmmipnlgesatitnacthki-sagsnetotsintigc.
We exploit the concept of support and query sets of few- activations from an autoencoder. [13] proposed using
shot classifiers to measure the similarity of samples within an autoencoder to reconstruct input samples such that
a support set after filtering. We perform this by randomly only the necessary signals remain for classification. Their
splitting the original support set randomly into auxiliary method requires fine-tuning the decoder based on the
classupport and query sets, followed by filtering the auxiliary sification loss of the input with respect to the ground
support and predicting on the query. If the samples are truth. However, under the few-shot setting, such
finenot self-similar, we will flag the support set as adversarial. tuning based on the classification loss should be avoided
To this end, we make the following contributions in our as we would require large enough samples from each
work: class for this step. [14] attempts to stabilise sensitive
1. We propose a novel attack-agnostic detection neurons which might be more prone to the effects of
admechanism against adversarial support sets in the versarial perturbations, by enforcing similar behaviours of
domain of few-shot classification. This is based these neurons between clean and adversarial inputs. Their
on self-similarity under randomised splitting of method requires adversarial samples during the training
the support set and filtering, and is the first, to process which potentially makes defending against novel
the best of our knowledge, for the detection of attacks challenging. Hence, we proposed a detection
apadversarial support sets in few-shot classifiers. proach that does not make use of any adversarial samples.
2. We investigate the effects of a unique white-box Though we employed the concept of feature preserving as
adversary against few-shot frameworks, through one of our various filtering functions, our approach is
difthe lens of transferability attacks. Rather than ferent from [14] as it does not suffer from this limitation.
crafting adversarial query samples similar to stan- Hence, in our work, we adopted an approach that does not
dard machine learning setups, we optimise adver- require any labelled data.
sarial supports sets, in a setting where all
nontarget classes are varying.
3. We provide further analysis on the detection per- 3. Background
formance of our algorithm when using
differing filtering functions and also different formula- 3.1. Few-shot classifiers Used
tion variants of the aforementioned self-similarity
quantity.</p>
      <sec id="sec-1-1">
        <title>A majority of the few-shot classifiers are trained with</title>
        <p>episodes sampled from the training set. Each episode
con* with  labelled</p>
        <p>The remaining of our paper is structured as follows: sists of a support set  = {, }=1
Section 2 discusses prior literature and Section 3 provides samples per  classes, and a query set  = {}=1
readers with the background to this study. We dive into with  unlabelled samples from the same  classes
our method in Section 4 and describe our experimental to be classified, denoted as a -way  -shot task. The
settings and evaluation results in Section 5. We provide metric-based classifiers learn a distance metric that
comfurther in-depth discussion in Section 6 and we conclude pares the features of support samples  and query sample
in Section 7 with summary and future work.  and generates similarity scores for classification.
During inference, the episodes are sampled from the test set
that has no overlapping categories with the training set.
2. Related Works In this work, we explored two known metric-based
fewshot classifiers, namely the RelationNet (RN) [ 5] and a
Poisoning of Support Sets: There is limited literature state-of-the-art model, the cross-attention network (CAN)
examining the poisoning of support sets in meta-learning. [6]. As illustrated in Figure 1, the support and query
sam[7] proposed an attack routine, Meta-Attack, extending ples are first encoded by a backbone CNN to get the image
from the highly explored Projected Gradient Descent
features { |  = 1, . . . , } and , respectively. The
(taPcGkDer)iasttuancakb[le10to]. oTbhtaeiynafseseudmbaecdkafsrcoemnathrieocwlahsesrieficatthieonat- feature vectors  and  ∈ R ,ℎ , , where  , ℎ ,
and  are the channel dimension, height, and width of
of the query set. Hence, the authors used the empirical the image features. If  &gt; 1,  will be the averaged
loss on the support set to generate adversarial support feature of the support samples from class . To measure
samples to induce misclassification behaviours to unseen the similarity between  and , the RN model
concatequeries. nates  and  along the channel dimension pairwise
and uses a relation module to calculate the similarities.
2.1. Autoencoder-based and Feature The CAN model adopts a cross-attention module that
Preserving-based Defences generates attention weights for every {, } pair. The
attended image features are further classified with cosine
similarity in the spirit of dense classification [15].
[12] performs detection of such attacks using
Nonparametric Scan Statistics (NPSS), based on hidden node
0 =  +   (−, ),</p>
        <p>(1)
 = ,{−1 +  (∇(ℎ(−1), ℎ))},</p>
        <p>(2)
where ℎ(.) is a prediction logits for classifier ℎ of some
input sample, ℎ is the ground truth label,  is the
loss used during training (i.e. cross entropy with softmax),
 is the step size and  is the adversarial strength which
limits the adversarial candidate  within an -bounded
ℓ∞ ball.</p>
        <p>The Carlini-Wagner (CW) attack [16] finds the smallest
 that successfully fools a target model using the Adam
optimiser. Their attack solves the following objective
function:</p>
        <p>min ||||2 +  · ( + , ),
.. (′, ) = max(−, max(ℎ(′)̸=) − ℎ(′)).</p>
        <p />
        <p>− ∼   ( ∖ {}),

−, − ∼   (| ∈ −),
 ∼   (| = ),</p>
        <p>= (1 , . . . , ℎ ),
ℎ() = ℎ(1 , . . . , ℎ , −, , −),
where  is the set of all classes, and − the random set of
classes used in the episode together with class . The last
line in (4) indicates that the few-shot classifier ℎ takes in a
support set made up of  and − and a query set made up
of  and −, which is a simplification to the expression,
to relate to (2) and (3). The adversarial perturbation  and
the underlying gradients are computed only for each of
the support samples  of the target class.</p>
        <p>(4)</p>
      </sec>
    </sec>
    <sec id="sec-2">
      <title>4. Defence Methodology</title>
      <p>The first term penalises  from being too large while the
second term ensures misclassification. The value 
is a weighting factor that controls the trade-off between
ifnding a low  and having a successful misclassification.
ℎ(·) refers to the logits of prediction index  and  refers
to the target prediction.  is the confidence value that
influences the logits score differences between the target
prediction  and the next best prediction .</p>
      <sec id="sec-2-1">
        <title>3.3. Threat Model</title>
        <p>We assume that the attacker wants to destroy the few-shot
classifier’s notion of a targeted class, , unlike
conventional machine learning frameworks where one is
optimising single test samples to be misclassified. The attacker
wants to find an adversarially perturbed set of support
images, such that misclassification of most query
samples from class  occurs, regardless of the class labels of
The defence is based on three components: sampling of
auxiliary query and support sets, filtering the auxiliary
support sets, and measuring the accuracy on the unfiltered
auxiliary query set. We denote a statistic either averaged
over all possible splits or for a randomly drawn split of a
support set into auxiliary sets with filtering of the
auxiliary supports as self-similarity. We elaborate further on
auxiliary sets below.</p>
      </sec>
      <sec id="sec-2-2">
        <title>4.1. Auxiliary Sets</title>
        <p>Few-shot classifiers’ support and query sets can be freely
chosen, implying that any sample can be used as either
a support or query. Given a support set for class , we
randomly split it into auxiliary sets, where  might be
clean or adversarial:
 ∪  =    ∩  = ∅,

.. || = ℎ − 1  || = 1.</p>
        <p>(5)</p>
        <sec id="sec-2-2-1">
          <title>The few-shot learner is now faced with a randomly drawn</title>
          <p>( − 1)-shot problem, evaluating on one query sample per
way, with the option to average the  possible splits.</p>
        </sec>
      </sec>
      <sec id="sec-2-3">
        <title>4.2. Detection of Adversarial Support</title>
      </sec>
      <sec id="sec-2-4">
        <title>Sets</title>
        <p>Our detection mechanism flags a support set as adversarial
when support samples within a class are highly different
from each other, as shown in Figure 2. Given a support set
of class , , we split it randomly into two auxiliary sets
 and . We filter  using a function (·) and
use the resultant samples as the new auxiliary support set
to evaluate . Following which, we obtain the logits
of  both before and after the filtering of the
auxiliary support (i.e. using  and () respectively)
and compute the ℓ1 norm difference between them. The
adversarial score  is given in Eq. (6) where ℎ is the
few-shot classifier
 =‖ℎ((), ) − ℎ(, )‖1, (6)

and  is any filtering function which maps a support set
onto its own space. The filter  is chosen such that it
causes smaller impact to clean samples, while inducing
larger dissimilarity between the auxiliary support and
query sets under adversariality. We observe that we obtain
already very high AUROC detection scores when
computing  without averaging over ℎ draws, which
we elaborate further later. We flag a support set  as
adversarial if the adversarial score goes above a certain
threshold1 (i.e.  &gt;  ). Different statistics can be
used to compute , with Eq. (6) being one of many.
Our main contribution lies rather in the proposal of using
self-similarity of a support set for such detection.</p>
        <sec id="sec-2-4-1">
          <title>1 can be chosen by examining the  on clean support</title>
          <p>samples, according to a desired threshold based on False Positive
Rates (e.g. @5% FPR).
We explored using an autoencoder (AE) as a filtering
function (·), for the detection of adversarial samples in
the support set, motivated by [13]. Initially we trained
a standard autoencoder to reconstruct the clean samples
in the image space using the MSE loss. However, the
standard autoencoder performed poorly in detecting
adversarial supports since it did not learn to preserve the
feature space representation of image samples. Therefore,
we switched to a feature-space preserving autoencoder
which additionally reconstructs the images in the feature
space of the few-shot classifier, contrary to prior work
where they fine-tuned their AE on the classification loss.
We argue that using classification loss for fine-tuning is
inapplicable in few shot learning due to having very few
labelled samples. We minimise the following objective
function for the feature-space preserving autoencoder:
′</p>
          <p>‖ − ˆ‖22
ℒℱ = 1 ′ ∑=︁1 0.01 · ()1/2
+ ‖ − ˆ‖22
()1/2
(7)
where  and ˆ are the original and reconstructed
image samples, respectively, and,  and ˆ are the feature
representation of the original and reconstructed image
obtained from the few-shot model before any metric module
(i.e. features from CNN backbone). The second loss term
ensures that the reconstructed image features are
similar to those of original image in the feature space of the
few-shot models. We train the feature-space preserving
autoencoder by fine-tuning the weights from the standard
autoencoder.</p>
        </sec>
      </sec>
      <sec id="sec-2-5">
        <title>4.4. Median Filtering (FeatS)</title>
        <p>In our work, we also explored an alternative filtering
function. We adopted a feature squeezing (FeatS) filter from
[11] where it was used in a conventional classifier. It
essentially performs local spatial smoothing of images by
having the centre pixel taking the median value among its
neighbours within a 2x2 sliding window. As their
detection performance was reasonably high using this filter, we
decided to use it as an alternative to FPA as an explorative
step. However, their approach performs filtering on each
individual test sample whereas we use it on the auxiliary
support set. Since we would like to demonstrate our
detection principle and the FPA performs already very well,
we leave further filtering functions to future research.</p>
      </sec>
      <sec id="sec-2-6">
        <title>5.2. Baseline Accuracy of Few-shot</title>
      </sec>
      <sec id="sec-2-7">
        <title>Classifiers</title>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>5. Experiments and Results</title>
      <sec id="sec-3-1">
        <title>5.1. Experimental Settings</title>
        <sec id="sec-3-1-1">
          <title>We evaluated our classifiers by taking the average and standard deviation accuracy over 2000 episodes across all models and datasets, reported in Table 1, to show that we were attacking reasonably performing few-shot classifiers.</title>
          <p>Datasets: MiniImagenet (MI) [17] and CUB [18] datasets
were used in our experiments. We prepared them
following prior benchmark splits [19, 20], with 64/16/20
categories for the train/val/test sets of MI and 100/50/50
categories for the train/val/test sets of CUB. In our attack Table 1
and detection evaluation, we chose an exemplary set of Baseline classification accuracy of the chosen models on
10 and 25 classes from the test set for MI and CUB re- the two datasets, under a 5-way 5-shot setting, computed
spectively, and we report the average metrics across them. across 2000 randomly sampled episodes. We report the
This is purely for computational efcfiiency. For the RN mean with 95% confidence intervals for the accuracy.
model, we used image sizes of 224 while using image RN - 5 shot CAN - 5 shot
sizes of 96 for the CAN model across both datasets. We MI 0.727 ± 0.0037 0.787 ± 0.0033
shrank the image size for the CAN model due to memory CUB 0.842 ± 0.0032 0.890 ± 0.0026
usage issues.</p>
          <p>Attacks: In our work, we used two different attack
routines, one being PGD while the other being a slight variant
of the CW attack. This variant uses a normal Stochastic 5.3. Attack Evaluation Metrics
Gradient Descent optimiser instead of Adam as we did not We evaluated the success of our attacks via computing
yield good performing adversarial samples with the latter. the Attack Success Rate (ASR), measuring the
proporWe still used the objective function defined in Eq. (3) to tion of samples that had adversarial candidates generated
optimise our CW adversarial samples, while using Eq. (2) from attacks that successfully cause misclassification. We
to perform a perturbation step less the clipping and sign only considered samples from the targeted class when
functions. We name this attack CW-SGD. For our PGD at- measuring ASR:
tack, we limit the ℓ∞ norm of the perturbation to 12/255
and a step size of  = 0.05 (see Eq. (2)). For our CW-  = E,∼{ (argmax (ℎ ( + , ))
SGD attack, clipping was not used due to the optimisation
over ||||2 while  = 0.1 and  = 50. We would like to ̸= )}.
stress that optimising for the best set of hyperparameters (8)
for generating attacks is not the main focus of our work
as we are more interested in obtaining viable adversarial The remaining ( − 1) classes were sampled randomly.
samples. In both settings, we generate 50 sets of adver- In the evaluation of the detection performances, we used
sarial perturbations for each of the 10 and 25 exemplary the Area Under the Receiver Operating Characteristic
classes for MI and CUB respectively. We also attack all  (AUROC) metric, since detection problems are binary
support samples for the targeted class . (whether an adversarial sample is present or not), and true</p>
          <p>Autoencoder: We used a ResNet-50 [21] architecture and false positives can be collected at various predefined
for the autoencoders2. For the MI dataset, we trained the threshold values ( ).
standard autoencoder from scratch with a learning rate of
1e-4. For CUB, we trained the standard encoder initialised 5.4. Transferability Attack Results
from ImageNet with a learning rate of 1e-4, and the
standard decoder from scratch with a learning rate of 1e-3. For We conducted transferability experiments to evaluate how
ifne-tuning of the feature-space preserving autoencoder, well the attacker generalised their generated adversarial
we used a learning rate of 1e-4. We employed a decaying perturbation under two unique scenarios: i) transfer with
learning rate with a step size of 10 epochs and  = 0.1. ifxed supports and ii) transfer with new supports. Setting
We used the Adam [22] optimiser with a weight decay of (i) assumes that we have the same adversarial support set
1e-4. In both settings, we used the train split for training for class  and we evaluated the ASR over newly drawn
and the validation split for selecting our best performing query sets. Setting (ii) relaxes this assumption, and we
inset of autoencoder weights out of 150 epochs of training. stead applied the generated adversarial perturbation, that
It is implemented in PyTorch [23]. was stored during the attack phase, on newly drawn
support sets for class , similarly evaluating over newly drawn
query sets. Contrary to transferability attacks in
conventional setups where a sample is generated on one model
2Autoencoder architecture adapted from GitHub repository and evaluated on another, we performed transferability to
https://github.com/Alvinhech/resnet-autoencoder. new tasks, by drawing randomly sets of non-target classes
0.8
0.6</p>
        </sec>
      </sec>
      <sec id="sec-3-2">
        <title>5.5. Detection of Adversarial Supports</title>
        <p>together with their support sets, and new query sets for
the few-shot paradigm. (b) CAN (5-shot) model.</p>
        <p>As illustrated in Figure 3, the PGD generated adversar- Figure 4: AUROC scores for the various filter functions
ial samples showed higher transferability than the CW- (normal distributed noise, median filtering from Feature
SGD attack, across both models and under both scenarios. Squeezing (FeatS), FPA) across our experiment settings,
The exceptionally high transfer ASR we observed under for RN and CAN models. Higher is better.
scenario (i) implies that once the attacker had obtained
an adversarial support set targeting a specific class,
successful attacks can be carried out on new tasks for which
the target class is present. This further reinforces the
motivation to investigate defence methods for few-shot
classifiers. Under scenario (ii), where the support set of
the target class is also randomised, we see lower transfer
ASR across the chosen classes. We would like to remind
readers that the adversarial samples were optimised
explicitly using setting (i) and not for (ii). Even though the
ASR in scenario (ii) is lower than in (i), there still exist
classes, where unpleasantly high ASR occurs.
tion against adversarial samples. Though "FeatS"
exhibits already a good detection performance, our FPA
approach consistently outperforms it across all settings.</p>
        <p>The "Noise" approach, however, does not detect well. We
see highly varied detection performances across the
different settings, which makes this approach highly unreliable3.</p>
        <p>This result is hardly surprising since such methods require
substantial manual fine-tuning of its noise parameters.</p>
        <p>This is not ideal as newer attacks can be introduced in
the future and also, being in a few-shot framework, the
optimal noise parameters between different task instances
might not be consistent as the data might be different.</p>
        <p>However, our FPA filter approach exhibits such
robustness even in such scenarios as it still achieved favourable
AUROC scores. For clean samples, our FPA managed to
reconstruct  such that the logits of  before and
after filtering remained consistent, even when the FPA did
not encounter classes from the novel split during training.</p>
        <p>We compared our explored approaches against a simple
ifltering function for (·), since prior detection methods
for adversarial samples in few-shot classifiers do not
exist. We experimented with using normal distributed noise
as a filter, in which we computed the channel-wise
variance for drawing normal distributed noise to be added to
the images. Being in the context of detection, we report
the AUROC scores to evaluate the effectiveness of our
detection algorithm.</p>
        <p>Our results in Figure 4 shows that FPA exhibits good de- 3Cases with AUROC score less than 0.5 indicates that more
tection performance. This indicates that the self-similarity favourable detection effectiveness can be achieved by flipping the
of clean samples under filtering of the auxiliary support detection threshold (i.e.  &gt;  to  &lt;  ). However, it
set is preserved to a degree, which allows discrimina- will not be experimentally consistent. This is also a clear indication
of the lack of reliability of using "Noise" as a filtering function.</p>
        <p>FPA
FeatS
Noise
FPA
FeatS</p>
        <p>Noise
of FPA
pronounced also in cases when the prediction label does
not switch.</p>
      </sec>
      <sec id="sec-3-3">
        <title>6.2. Varying Degrees of Regularisation</title>
        <sec id="sec-3-3-1">
          <title>We observe lower AUROC scores for the RN model than</title>
          <p>the CAN model in Figure 4. As such, we question if
this difference can be attributed to the FPA’s ability to
reconstruct clean samples effectively. Recalling from
Eq. (7), we define an additional regularisation term to
enforce stricter reconstruction requirements to also include
class distribution reconstruction. More specifically, we
minimise the following objective function:
(9)
ℒℱ′ =
′
=1
 ′</p>
          <p>‖ − ˆ‖22
1 ∑︁ 0.01 · ()1/2
+ ‖ − ˆ‖22
()1/2</p>
        </sec>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>6. Discussion</title>
      <sec id="sec-4-1">
        <title>6.1. Study of Self-Similarity</title>
      </sec>
      <sec id="sec-4-2">
        <title>Computation Methods</title>
        <p>In Section 4.2, we described one of the possible detection
mechanisms based on logits differences. An alternative
would be to use hard label predictions. Thus, we
investigate the effect of a differing scheme as a justification for
our choice . For the case of hard label predictions,
we perform the following: we compute the average
accuracy of , across the different permutated partitions
of , illustrated in Figure 5. This results in the statistic
′:
′ =
1
ℎ
∑︁
ℎ =1
1[(ℎ((,), ,)) ̸= ],</p>
        <p>where ℎ is the few-shot classifier,  is the filtering function,
and 1 is the indicator function. Similarly, we flag the
support set as adversarial when ′ &gt;  , such that it
goes beyond a certain threshold.</p>
        <p>…




iliary sets  and  is performed. Best viewed in
Higher is better.</p>
        <p>AUROC scores for the two detection mechanisms (
and ′) using our FPA across our experiment settings.</p>
        <p>Model</p>
        <p>RN
(5-shot)</p>
        <p>CAN
(5-shot)</p>
        <p>Dataset</p>
        <p>MI
CUB
MI
CUB</p>
        <p>PGD</p>
        <p>CW-SGD

as  consistently outperforms ′, with the differ- ing to few-shot classifiers. To this end, we propose a novel
ence being bigger for RN. Differences in logits can be
adversarial attack detection algorithm on support sets in
+ ‖ − ˆ‖22 ,
()1/2</p>
        <p>(10)
where  and ˆ are the original and reconstructed image
samples, respectively, and  and ˆ are the feature
representation of the original and reconstructed image obtained
from the few-shot model before any metric module, and
 and ˆ are the logits of the original and reconstructed
image. We refer to this variant as   ′. Similarly, we
train   ′ by fine-tuning the weights from the standard
autoencoder.
AUROC results comparing    and   ′ for the RN.</p>
        <p>We computed the results across the 10 and 25
exemplary classes for MI and CUB respectively, and 50 sets of
adversarial perturbations.</p>
        <p>Dataset</p>
        <p>MI
CUB</p>
        <p>PGD</p>
        <p>CW-SGD</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>7. Conclusion</title>
      <sec id="sec-5-1">
        <title>Adversarial attacks against the support sets can be damag</title>
        <p>the few-shot framework, which has not been explored [10] A. Madry, A. Makelov, L. Schmidt, D. Tsipras,
prior, to the best of our knowledge. Our algorithm works A. Vladu, Towards deep learning models resistant
by using the concept of self-similarity among samples in to adversarial attacks, in: International Conference
the support set and filtering. We obtained high detection on Learning Representations, 2018.
AUROC scores in the CAN and RN models, across MI [11] W. Xu, D. Evans, Y. Qi, Feature squeezing:
Detectand CUB datasets, with FPA and FeatS filtering functions, ing adversarial examples in deep neural networks,
though FPA is superior. We have also found that using arXiv preprint arXiv:1704.01155 (2017).
differences of the logits scores yield better detection per- [12] C. Cintas, S. Speakman, V. Akinwande, W. Ogallo,
formances and a higher degree of regularisation of FPA K. Weldemariam, S. Sridharan, E. McFowland,
Dedoes not guarantee better detection results. Future work tecting adversarial attacks via subset scanning of
can explore the efficacy of our detection for black-box at- autoencoder activations and reconstruction error, in:
tack settings and the detection performances with different IJCAI, 2020.
ifltering functions. [13] J. Folz, S. Palacio, J. Hees, A. Dengel, Adversarial
defense based on structure-to-signal autoencoders,
in: 2020 IEEE Winter Conference on Applications
Acknowledgments of Computer Vision (WACV), IEEE, 2020, pp. 3568–
3577.</p>
        <p>This research is supported by both ST Engineering Elec- [14] C. Zhang, A. Liu, X. Liu, Y. Xu, H. Yu, Y. Ma, T. Li,
tronics and National Research Foundation, Singapore, Interpreting and improving adversarial robustness of
under its Corporate Laboratory @ University Scheme deep neural networks with neuron sensitivity, IEEE
(Programme Title: STEE Infosec-SUTD Corporate Labo- Transactions on Image Processing 30 (2020) 1291–
ratory). 1304.
[15] Y. Lifchitz, Y. Avrithis, S. Picard, A. Bursuc, Dense
References classification and implanting for few-shot learning,
in: Proceedings of the IEEE Conference on
Com[1] C. Finn, P. Abbeel, S. Levine, Model-agnostic meta- puter Vision and Pattern Recognition, 2019, pp.
learning for fast adaptation of deep networks, in: 9258–9267.</p>
        <p>Proceedings of the 34th ICML Volume 70, JMLR. [16] N. Carlini, D. Wagner, Towards evaluating the
roorg, 2017, pp. 1126–1135. bustness of neural networks, in: 2017 IEEE
Sym[2] N. Mishra, M. Rohaninejad, X. Chen, P. Abbeel, posium on Security and Privacy, IEEE, 2017, pp.</p>
        <p>A simple neural attentive meta-learner, in: ICLR, 39–57.</p>
        <p>2018. [17] O. Vinyals, C. Blundell, T. Lillicrap, D. Wierstra,
[3] A. Nichol, J. Schulman, Reptile: a scal- et al., Matching networks for one shot learning, in:
able metalearning algorithm, arXiv preprint NIPS, 2016, pp. 3630–3638.</p>
        <p>arXiv:1803.02999 2 (2018) 4. [18] C. Wah, S. Branson, P. Welinder, P. Perona, S.
Be[4] Q. Sun, Y. Liu, T.-S. Chua, B. Schiele, Meta-transfer longie, The caltech-ucsd birds-200-2011 dataset
learning for few-shot learning, in: Proceedings of (2011).</p>
        <p>the IEEE CVPR, 2019, pp. 403–412. [19] S. Ravi, H. Larochelle, Optimization as a model for
[5] F. Sung, Y. Yang, L. Zhang, T. Xiang, P. H. Torr, few-shot learning, in: ICLR, 2017.</p>
        <p>T. M. Hospedales, Learning to compare: Relation [20] H.-Y. Tseng, H.-Y. Lee, J.-B. Huang, M.-H. Yang,
network for few-shot learning, in: Proceedings of Cross-domain few-shot classification via learned
the IEEE CVPR, 2018, pp. 1199–1208. feature-wise transformation, in: ICLR, 2020.
[6] R. Hou, H. Chang, M. Bingpeng, S. Shan, X. Chen, [21] K. He, X. Zhang, S. Ren, J. Sun, Deep residual
Cross attention network for few-shot classification, learning for image recognition, in: Proceedings of
in: NIPS, 2019, pp. 4005–4016. the IEEE conference on computer vision and pattern
[7] H. Xu, Y. Li, X. Liu, H. Liu, J. Tang, Yet meta recognition, 2016, pp. 770–778.
learning can adapt fast, it can also break easily, 2020. [22] D. Kingma, J. Ba, Adam: A method for
stochasarXiv:2009.01672. tic optimization, arXiv preprint arXiv:1412.6980
[8] M. Goldblum, L. Fowl, T. Goldstein, Robust (2014).</p>
        <p>few-shot learning with adversarially queried meta- [23] A. Paszke, S. Gross, S. Chintala, G. Chanan,
learners, arXiv preprint arXiv:1910.00982 (2019). E. Yang, Z. DeVito, Z. Lin, A. Desmaison,
[9] C. Szegedy, W. Zaremba, I. Sutskever, J. Bruna, L. Antiga, A. Lerer, Automatic differentiation in
D. Erhan, I. Goodfellow, R. Fergus, Intriguing pytorch (2017).
properties of neural networks, arXiv preprint
arXiv:1312.6199 (2013).</p>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list />
  </back>
</article>