<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>DAoB: A Transferable Adversarial Attack via Boundary Information for Speaker Recognition Systems</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Junjian Zhang</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Binyue Deng</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Hao Tan</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Le Wang</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Zhaoquan Gu</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Cyberspace Institute of Advanced Technology, Guangzhou University</institution>
          ,
          <addr-line>Guangzhou</addr-line>
          ,
          <country country="CN">China</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Department of New Networks, Peng Cheng Laboratory</institution>
          ,
          <addr-line>Shenzhen</addr-line>
          ,
          <country country="CN">China</country>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>School of Computer Science and Technology, Harbin Institute of Technology (Shenzhen)</institution>
          ,
          <addr-line>Shenzhen</addr-line>
          ,
          <country country="CN">China</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2023</year>
      </pub-date>
      <fpage>12</fpage>
      <lpage>17</lpage>
      <abstract>
        <p>Audio deepfakes pose significant security threats to speaker recognition systems (SRSs), particularly with the growing threat of adversarial attacks. Existing black-box attack methods mostly rely on ensembling multiple datasets and models to search for adversarial examples (AEs) with good transferability, but they ignore the limitations of such search algorithms. In this paper, we comprehensively analyze diferent iterative-based adversarial attack methods and explain diferent transferability from the perspective of optimizing the search space. Furthermore, we propose a difusion-based attack method located on the boundary (DAoB for short), which takes boundary information into consideration to achieve better transferability. Specifically, DAoB starts searching for an appropriate AE from the boundary of the search space instead of the original example, then it guides the search process by difused audio and the gradients of multiple white-box models to obtain better gradient directions. To validate the efectiveness, we conducted experiments on seven state-of-the-art SRSs and DAoB outperforms others. Remarkably, even in the black-box scenario, the attack success rate of DAoB attains an impressive 97.2%, in close proximity to the rate achieved in the white-box scenario.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;Speaker recognition</kwd>
        <kwd>adversarial attack</kwd>
        <kwd>boundary information</kwd>
        <kwd>difusion</kwd>
        <kwd>transferability</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>incorrect output [7]. This attack can create examples that
deceive deep learning models without being easily
deWith the rapid development of information technolo- tected by humans. This reveals the common and serious
gies, identity recognition have become more intelligent vulnerabilities of deep learning models and promotes the
and convenient. Typical methods such as facial recogni- further development of various intelligent technologies.
tion, fingerprint recognition, iris recognition, and speaker Among them, adversarial attacks against SRSs are more
recognition can quickly perform identity verification dificult than those against images and have started later
with high success ratio[1, 2]. Especially with the emer- [2].
gence of deep learning technologies, identity recogni- Commonly speaking, there are three attack scenarios:
tion has made significant progress. However, Deepfake white-box, gray-box, and black-box. In the white-box
technology verifies the challenge of identity recognition scenario, attackers have access to all the details of the
systems’ reliability [3, 4]. Compared to image-based deep- SRSs, so they often employ gradient-based attack
methfake technology, speech-based deepfake technology is ods to directly generate adversarial examples. Recent
a relatively new field. For example, recently emerging studies [8, 9, 10] show that adversarial attacks can fully
technologies [5, 6] such as speech recording and replay, overcome almost all white-box SRSs. In the gray-box
sceText-to-Speech (TTS), and Voice Conversion (VC) can nario, attackers need to continuously query the victim
deceive SRSs to a certain extent. system and obtain corresponding score vectors to
opti</p>
      <p>
        Various deep-learning models are shown to be vul- mize adversarial examples. For example, the FAKEBOB
nerable to adversarial attacks which add small amounts method proposed in [
        <xref ref-type="bibr" rid="ref9">11</xref>
        ] can efectively attack most SRSs
of noise to the benign example and mislead high- in the gray-box situation. However, this strategy requires
performance deep neural networks (DNNs) to produce frequent access to the victim system, which may expose
the attacking intent and weaken the attack’s
concealIJCAI 2023 Workshop on Deepfake Audio Detection and Analysis ment. In real scenarios, the internal information of the
(DADA 2023), August 19, 2023, Macao, S.A.R model is often unknown, and a normal SRS only outputs
* Corresponding author. a speaker identity label, not a score vector. Therefore,
2$11221110261001600@69e@.gzeh.guz.heduu.e.dcnu.(cBn. (DJ.eZnhga);ntga)n; hh198@gmail.com the black-box attack scenario has attracted more
atten(H. Tan); wangle@gzhu.edu.cn (L. Wang); guzhaoquan@hit.edu.cn tion from researchers. Currently, efective attacks against
(Z. Gu) black-box scenarios are mostly based on the
transferabil0009-0002-1180-0786 (J. Zhang) ity of adversarial examples, by generating adversarial
© 2023 Copyright for this paper by its authors. Use permitted under Creative Commons License examples on existing white-box models or retraining
CPWrEooUrckReshdoinpgs IhStpN:/c1e6u1r3-w-0s.o7r3g ACttEribUutRion W4.0oInrtekrnsahtioonpal (PCCroBYce4.0e).dings (CEUR-WS.org)
substitute models to generate the adversarial examples
against the target black-box model. Therefore, improving
the transferability of adversarial examples has become a
unanimous goal of researchers.
spaces, and explain the diferences in
transferability of diferent attack methods through the
diferences in search spaces;
• We propose DAoB, a powerful technique for
generating adversarial examples, which can generate
adversarial examples that pose a strong threat to
black-box SRSs;
• We demonstrate the efectiveness of our method
by attacking seven state-of-the-art SRSs. Our
method can achieve attack efects that are close
to the white-box scenario without accessing the
target model in a fully black-box setting.
      </p>
      <sec id="sec-1-1">
        <title>The rest of the paper is organized as follows. The</title>
        <p>next section highlights the related work in the field of
adversarial attacks against speaker recognition systems.
Then, we demonstrate DAoB and its theoretical analysis
in Section 3. We describe the setup of the experiment
in Section 4 and show the experiment results and
discussions in Section 5. Finally, we conclude the paper in
Section 6.</p>
      </sec>
    </sec>
    <sec id="sec-2">
      <title>2. Related work</title>
      <p>
        Many works on enhancing transferability [12, 13, 14]
have achieved positive results, but there are still some
unresolved or undiscovered issues. First, the development 2.1. FakeBob
of transferability for adversarial examples is still insuf- FakeBob [
        <xref ref-type="bibr" rid="ref9">11</xref>
        ] first proposed a black-box adversarial attack
ifcient, and the efectiveness of attacks is significantly on SRSs. By assuming that the attacker can obtain the
diferent from that of white-box attacks. Second, with the scoring of the model, and designing a loss function based
limitation of the algorithmic constraints, searching and on the score vector, the attack success rate can be close
generating adversarial examples always in a smaller real to that of a white-box attack. However, this method still
search space than expected. Third, current works have requires model scores to attack in actual attacks, and it
limited research on model ensembles, mostly integrat- is not a complete black box. At the same time, a large
ing only 2-3 models, and lack research on multi-model number of query target models are required during the
ensembles. attack process, which is dificult to achieve in practical
      </p>
      <p>In this paper, we propose a difusion attack located on applications.
the boundary (called DAoB for short), which adopts the
boundary information to address the above-mentioned
problems. The workflow of DAoB is described in Fig. 1. 2.2. Transferable Adversarial Attack
Specifically, we place the search region for adversarial Zhang et al. [13] proposed an integrated attack at the
examples directly on the attack boundary, which is more logits level can be used to achieve a black-box attack
likely to produce high transferability at the beginning of with a higher success rate, but they did not give a
reasoneach iteration. Secondly, when calculating gradient infor- able solution to the problem of inconsistent logits ranges
mation, we obtain multiple gradient information that has between diferent models in the paper. The method in
been difused by taking an information difusion space [12] uses spatial momentum to calculate the gradient for
near the example and taking the average of these gradi- integrated attack, that is, the gradient information in the
ents as guidance for this iteration. Finally, considering previous iteration process will be used in this iteration,
that the logits of diferent models have large diferences, and the use of momentum can efectively improve the
we adopt gradient-level model integration for the attack. aggressiveness and transferability of adversarial
exam</p>
      <p>We summarize the paper’s main contributions as fol- ples. The STA-MDCT [14] method introduces discrete
lows: cosine transform into the adversarial attack of speaker
• We conduct theoretical analysis and experimental recognition, it firstly converts the audio to the frequency
comparisons of existing adversarial attack meth- domain and adds random adversarial noise, and then
ods, propose definitions of ideal and actual search converts it back to the time domain for model
integration attack. This method improves the interpretability
and transferability of the attack in the black-box envi- boundary epsilon to prevent excessive noise. For
multironment, but the additional transformation increases a step attacks (using the most commonly used parameter
lot of computing time, and the current models are all selection, step size  = / ), as the gradient direction
end-to-end input, they would not pay more attention to changes, the attacker is unable to search near the set
the frequency domain part of the audio. boundaries.</p>
      <p>For transfer-based black-box attacks, many
experiments have shown that the single-step attack method
3. METHODOLOGY has certain advantages in transferability. It can also be
inferred from experience that the further an adversarial
3.1. DAoB example is from the original example, i.e., the closer it
As shown in the Fig. 1, we divide each iteration of DAoB is to the boundary, the greater the diference in speaker
into two phases: a module for computing the gradient of characteristics between the adversarial example and the
the current iteration, and another for using the gradient original example. Therefore, we believe that the
transto guide the search for the adversarial example in this it- ferability of adversarial examples in the search space
eration. Specifically, in the gradient computation module, near the boundary is stronger than that of adversarial
we input the difused example into multiple white-box examples in the search space near the original example.
models to obtain their corresponding embeddings and In summary, we propose an attack method that starts
calculate the gradient accordingly. In the example update from the boundary. Based on the initial gradient
informamodule, we add the gradient obtained in this iteration tion, we directly add a large amount of noise to the
origito the momentum gradient. Then, we extract the direc- nal example to reach the boundary. Then, we search for
tion of the momentum gradient to guide the update of adversarial examples from the boundary with a smaller
the current example. Specifically, we update the original step size.
example starting from the search boundary with a small
step size for multiple iterations, eventually obtaining an 3.1.2. Difusion attack
adversarial example with strong transferability, which
can cause a black-box victim model to output incorrect
results. Below, we will detail the theoretical basis and
specific operations of DAoB.</p>
      <sec id="sec-2-1">
        <title>The difusion model has played a huge role in the fields</title>
        <p>of image and speech generation. The process of a single
difusion step is to add noise to the example  to obtain
+1, while the reverse difusion is the process of using
 and +1 as guidance to learn the generation method.</p>
        <p>Algorithm 1 DAoB Inspired by the difusion model, we optimize the process
Require: of generating adversarial examples iteratively obtained
Set of several white-box models  = { | = from , by difuse it to , and then using  as guidance
1, · · · , }, clean input  with target labels , pa- to iterate to obtain +1. Considering that the process of
rameters  = {,  , , ,  } generating adversarial examples lacks support from data
Ensure: volume, we perform multiple difusions to better guide</p>
        <p>Adversarial example ; the direction of iteration.
1: 1 ← , 0 ← 0,  0 ← ;</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>2: for iteration time  ← 1 to  do 3.1.3. DAoB with model ensembles</title>
      <p>3: for number of sub-example  ← 1 to  do
4:  = (,  )
5: for number of model  ← 1 to  do</p>
      <p>Calculate gradient  of (,  )
6: end for
7: end for
8: Update  by Eq. (3-4)
9: Update  by Eq. (5-7)
10: end for
11: return 
To improve transferability, we obtain gradients from
multiple white-box models. Unlike ensemble methods based
on logits level, we set the same loss function for each
white-box model. We believe that the size of the gradient
in any dimension reflects the influence that the gradient
direction can have in that dimension. Therefore, we do
not set weights for the gradients of diferent white-box
models in the ensemble gradient. We simply add the
gradients obtained from each white-box model and then take
the direction of the ensemble gradient as the direction
for the current iteration.</p>
      <p>We introduced the details of DAoB, which is described
in Algorithm 1. We formalize the specific attack process
as follows:
3.1.1. Take the boundary as the starting point</p>
      <sec id="sec-3-1">
        <title>Previous attack methods generate adversarial examples by updating from the original example. Attackers set a</title>
        <p>=
=1,· ,</p>
        <p>∑︁
=1,· ,
∇
︁(
,  ,
︁)
where  (︀ , ︀) is the derivative of the loss function

with respect to the input example  .</p>
      </sec>
      <sec id="sec-3-2">
        <title>Then we obtain the gradient for this iteration as:</title>
        <p>=  · −1
+</p>
        <p>‖‖1
.</p>
      </sec>
      <sec id="sec-3-3">
        <title>We use the first obtained gradient 1 to guide the orig</title>
        <p>inal example to reach the boundary of the search space:
1 =  +  · sign (1) ,
where  is the constraint of pertubation.</p>
        <p>Afterwards, we use the momentum gradient to
generate smaller perturbation, which is then used to create
the adversarial audio. This process can be formalized as
follows:</p>
        <p>=  ·  −1 ,
 = Clip {︁−1 +   · sign (+1)}︁ ,
where   is the attack step size, and  is the attenuation
factor.
(4)
(5)
(6)</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>4. Experiment set up</title>
      <sec id="sec-4-1">
        <title>4.1. Datasets and Models</title>
        <sec id="sec-4-1-1">
          <title>The datasets generated according to libriSpeech [15] are</title>
          <p>consistent with the Spk10-enroll, Spk10-test, and
Spk10imposter datasets published in [16].To validate the
effectiveness of our method, we selected seven strong
victim models: Res34-L [17], Res34-V [17], TDy_HR [18],
TDy_QR [18], TDy_VGG [18], XV-plda [19], ECAPA [20].
iteration, the audio 
difused into a :</p>
          <p>The input consists of a set of several white-box models
 = { | = 1, · · ·</p>
          <p>, },clean input  with target
labels , and a series of required parameters. In each
 from the previous iteration is
 = (
,  ) = 
 + (
,  ),
(1)
 (0,  2).
where (</p>
          <p>,  ) is Gaussian noise with the same
shape as  and follows a normal distribution</p>
          <p>The gradients computed using  and  in each
iteration can be formalized as:</p>
        </sec>
      </sec>
      <sec id="sec-4-2">
        <title>4.2. Evaluation Metrics</title>
        <sec id="sec-4-2-1">
          <title>We use the attack success rate (ASR) of adversarial audio on black-box victim models as the metric to evaluate the transferability of adversarial audio.</title>
        </sec>
        <sec id="sec-4-2-2">
          <title>In addition to focusing on the transferability of adver</title>
          <p>sarial audio, we also pay attention to the imperceptibility
of adversarial audio. Therefore, we refer to [12] and use</p>
        </sec>
        <sec id="sec-4-2-3">
          <title>Signal-to-Noise Ratio (SNR), Perceptual Evaluation of</title>
          <p>of the imperceptibility of adversarial audio.</p>
          <p>Speech Quality (PESQ), and 2 −  as the indicators</p>
        </sec>
      </sec>
      <sec id="sec-4-3">
        <title>4.3. Parameter Details</title>
        <p>During the process of adversarial attacks, all attacks were
(2) targeted attacks in the open-set recognition scenario, and
the attack targets were fixed simple targets, with the
specific settings being the same as in [ 21]. The number
of iterations  for all iterative attack methods was 10. We
conducted experiments with diferent perturbation limits,
but due to space constraints, the perturbation limits 
shown in the experimental results below are all set to
(3) 0.002. The fixed step size  is set to / . In particular,
there is some randomness in difusion of DAoB, so all
results reported for our methods are the average of 10
repeated experiments.</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>5. Experiment Results and</title>
    </sec>
    <sec id="sec-6">
      <title>Analysis</title>
      <sec id="sec-6-1">
        <title>5.1. Results of Diferent Attack Methods</title>
        <p>We systematically studied the transferability of
adversarial examples generated by DAoB and reported the
experimental results in Tab. 1, where the DAoB’s
parameter setting is { = 3,</p>
        <p>= 0.8}. Specifically, all attack
methods adopted the strategy of the gradient-level model
ensemble. The gray part in the table is the attack success
rate of adversarial examples on the white-box model. We
can see that for any victim model, the attack success rate
of DAoB is higher than that of MI-FGSM, with the highest
increase of 16.9 percentage points. Compared with the</p>
        <sec id="sec-6-1-1">
          <title>PGD and MI-FGSM, the average attack success rate from</title>
          <p>six group black-box attack experiments has improved by
43.65% and 23.04%, respectively. At the same time, the
audio quality of DAoB is between that of MI-FGSM and
FGSM, which is consistent with our description of the
search space in the previous text: DAoB can search the
area near the boundary, while MI-FGSM cannot.
Nevertheless, the adversarial audio obtained by DAoB still
achieves a high signal-to-noise ratio and PESQ. In
addition, we did not include xv-plda as a white-box model in
the table because we found during the experiment that
xv-plda would have a negative impact on the attack. We
have not found a specific reason for this phenomenon
and we will try to explain it in future work.</p>
        </sec>
      </sec>
      <sec id="sec-6-2">
        <title>5.2. Ablation Study</title>
        <sec id="sec-6-2-1">
          <title>We conducted ablation studies by combining MI-FGSM with difusion attack and boundary attack, and Fig. 2 re</title>
        </sec>
      </sec>
    </sec>
    <sec id="sec-7">
      <title>6. Conclusion</title>
      <sec id="sec-7-1">
        <title>We study efective adversarial attacks for black-box SRSs.</title>
        <p>By adopting the boundary information, we propose
DAoB, a new adversarial example generation strategy
that efectively improves the transferability of
adversarial examples. Experimental results show that DAoB can
achieve a high success rate under the black-box scenario
that is close to that of white-box attacks. In future work,
we will try to reduce the size of adversarial perturbations
and limit the area of adversarial perturbations to enhance
the stealth of attacks without afecting the success rate
of attacks.</p>
      </sec>
    </sec>
    <sec id="sec-8">
      <title>Acknowledgement</title>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          <year>2020</year>
          , pp.
          <fpage>1738</fpage>
          -
          <lpage>1742</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          <article-title>This work is supported in part by the Major Key Project</article-title>
          [11]
          <string-name>
            <given-names>G.</given-names>
            <surname>Chen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Chenb</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Fan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Du</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Zhao</surname>
          </string-name>
          ,
          <string-name>
            <surname>F.</surname>
          </string-name>
          <article-title>Song, of PCL (Grant No. PCL2022A03) and the Guangdong Y</article-title>
          . Liu,
          <article-title>Who is real bob? adversarial attacks on Provincial Key Laboratory of Novel Security Intelligence speaker recognition systems, in: 2021 IEEE SympoTechnologies (2022B1212010005). sium on Security and Privacy (SP)</article-title>
          , IEEE,
          <year>2021</year>
          , pp.
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          <string-name>
            <surname>References</surname>
            [12]
            <given-names>H.</given-names>
          </string-name>
          <string-name>
            <surname>Tan</surname>
            ,
            <given-names>Z.</given-names>
          </string-name>
          <string-name>
            <surname>Gu</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          <string-name>
            <surname>Wang</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          <string-name>
            <surname>Zhang</surname>
            ,
            <given-names>B. B.</given-names>
          </string-name>
          <string-name>
            <surname>Gupta</surname>
          </string-name>
          ,
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          <string-name>
            <given-names>Z.</given-names>
            <surname>Tian</surname>
          </string-name>
          , Improving adversarial transferability by [1]
          <string-name>
            <given-names>J.</given-names>
            <surname>Priesnitz</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Rathgeb</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Buchmann</surname>
          </string-name>
          ,
          <string-name>
            <surname>C.</surname>
          </string-name>
          <article-title>Busch, temporal and spatial momentum in urban speaker</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          <string-name>
            <given-names>M.</given-names>
            <surname>Margraf</surname>
          </string-name>
          ,
          <article-title>An overview of touchless 2d finger- recognition systems</article-title>
          , Computers and Electrical En-
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          <article-title>print recognition</article-title>
          ,
          <source>EURASIP Journal on Image and gineering 104</source>
          (
          <year>2022</year>
          )
          <fpage>108446</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          <source>Video Processing</source>
          <year>2021</year>
          (
          <year>2021</year>
          )
          <fpage>1</fpage>
          -
          <lpage>28</lpage>
          . [13]
          <string-name>
            <given-names>Y.</given-names>
            <surname>Zhang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Jiang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Villalba</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Dehak</surname>
          </string-name>
          , Black-box [2]
          <string-name>
            <given-names>H.</given-names>
            <surname>Tan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Zhang</surname>
          </string-name>
          , J. Zhang, M.
          <article-title>Shafiq, attacks on spoofing countermeasures using trans-</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          <article-title>speaker recognition systems: A survey</article-title>
          ,
          <year>Electronics 2020</year>
          , pp.
          <fpage>4238</fpage>
          -
          <lpage>4242</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          <volume>11</volume>
          (
          <year>2022</year>
          )
          <fpage>2183</fpage>
          . [14]
          <string-name>
            <given-names>J.</given-names>
            <surname>Yao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Luo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.-L.</given-names>
            <surname>Zhang</surname>
          </string-name>
          , Interpretable spec[3]
          <string-name>
            <given-names>Z.</given-names>
            <surname>Gu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W.</given-names>
            <surname>Hu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Zhang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Lu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Yin</surname>
          </string-name>
          , L. Wang,
          <article-title>trum transformation attacks to speaker recognition,</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          <article-title>Gradient shielding: towards understanding vulner-</article-title>
          arXiv
          <source>preprint arXiv:2302.10686</source>
          (
          <year>2023</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          <article-title>ability of deep neural networks</article-title>
          , IEEE transactions [15]
          <string-name>
            <given-names>V.</given-names>
            <surname>Panayotov</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Chen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Povey</surname>
          </string-name>
          , S. Khudanpur,
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          <source>on network science and engineering 8</source>
          (
          <year>2020</year>
          )
          <fpage>921</fpage>
          -
          <lpage>Librispeech</lpage>
          :
          <article-title>an asr corpus based on public domain</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          932. audio books, in: 2015 IEEE international confer[4]
          <string-name>
            <given-names>Z.</given-names>
            <surname>Gu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Khan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Deng</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Du</surname>
          </string-name>
          ,
          <string-name>
            <surname>M.</surname>
          </string-name>
          <article-title>Guizani, ence on acoustics, speech and signal processing</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          <string-name>
            <given-names>Z.</given-names>
            <surname>Tian</surname>
          </string-name>
          ,
          <article-title>Iepsbp: A cost-eficient image encryp- (ICASSP)</article-title>
          , IEEE,
          <year>2015</year>
          , pp.
          <fpage>5206</fpage>
          -
          <lpage>5210</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          <article-title>tion algorithm based on parallel chaotic system for</article-title>
          [16]
          <string-name>
            <given-names>G.</given-names>
            <surname>Chen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Zhao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Song</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Chen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Fan</surname>
          </string-name>
          , Y. Liu,
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          <article-title>green iot, IEEE Transactions on Green Communi- Sec4sr: a security analysis platform for speaker</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          <source>cations and Networking</source>
          <volume>6</volume>
          (
          <year>2021</year>
          )
          <fpage>89</fpage>
          -
          <lpage>106</lpage>
          . recognition,
          <source>arXiv preprint arXiv:2109.01766</source>
          (
          <year>2021</year>
          ). [5]
          <string-name>
            <given-names>L.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Chen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T. F.</given-names>
            <surname>Zheng</surname>
          </string-name>
          , A study on re- [17]
          <string-name>
            <given-names>J. S.</given-names>
            <surname>Chung</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Huh</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Mun</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Lee</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H. S.</given-names>
            <surname>Heo</surname>
          </string-name>
          ,
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          <string-name>
            <surname>verification</surname>
          </string-name>
          ,
          <source>arXiv preprint arXiv:1706.02101</source>
          (
          <year>2017</year>
          ).
          <article-title>fence of metric learning for speaker recognition</article-title>
          , [6]
          <string-name>
            <given-names>Y.</given-names>
            <surname>Zhang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Jiang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Duan</surname>
          </string-name>
          ,
          <article-title>One-class learning arXiv preprint</article-title>
          arXiv:
          <year>2003</year>
          .
          <volume>11982</volume>
          (
          <year>2020</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          <article-title>towards synthetic voice spoofing detection</article-title>
          , IEEE [18]
          <string-name>
            <given-names>S.-H.</given-names>
            <surname>Kim</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Nam</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.-H.</given-names>
            <surname>Park</surname>
          </string-name>
          , Temporal dy-
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          <source>Signal Processing Letters</source>
          <volume>28</volume>
          (
          <year>2021</year>
          )
          <fpage>937</fpage>
          -
          <lpage>941</lpage>
          .
          <article-title>namic convolutional neural network for text[7</article-title>
          ]
          <string-name>
            <given-names>Y.</given-names>
            <surname>Dong</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Liao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Pang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Su</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Zhu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Hu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <article-title>independent speaker verification and phonemic</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          <article-title>Boosting adversarial attacks with momentum, in: analysis</article-title>
          ,
          <source>in: ICASSP</source>
          <year>2022</year>
          -2022 IEEE International
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          <source>vision and pattern recognition</source>
          ,
          <year>2018</year>
          , pp.
          <fpage>9185</fpage>
          -
          <lpage>9193</lpage>
          . cessing (ICASSP), IEEE,
          <year>2022</year>
          , pp.
          <fpage>6742</fpage>
          -
          <lpage>6746</lpage>
          . [8]
          <string-name>
            <given-names>Z.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Wu</surname>
          </string-name>
          , J. Liu,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Chen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Yuan</surname>
          </string-name>
          , Advpulse: [19]
          <string-name>
            <given-names>D.</given-names>
            <surname>Snyder</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Garcia-Romero</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Sell</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Povey</surname>
          </string-name>
          ,
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          <article-title>adversarial attacks via subsecond perturbations, in: for speaker recognition</article-title>
          , in: 2018 IEEE international
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          <source>Proceedings of the 2020 ACM SIGSAC Conference conference on acoustics, speech and signal process-</source>
        </mixed-citation>
      </ref>
      <ref id="ref25">
        <mixed-citation>
          <source>on Computer and Communications Security</source>
          ,
          <year>2020</year>
          , ing (ICASSP), IEEE,
          <year>2018</year>
          , pp.
          <fpage>5329</fpage>
          -
          <lpage>5333</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref26">
        <mixed-citation>
          pp.
          <fpage>1121</fpage>
          -
          <lpage>1134</lpage>
          . [20]
          <string-name>
            <given-names>B.</given-names>
            <surname>Desplanques</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Thienpondt</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Demuynck</surname>
          </string-name>
          , [9]
          <string-name>
            <given-names>W.</given-names>
            <surname>Zhang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Zhao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Liu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Li</surname>
          </string-name>
          , X. Cheng, T. F.
          <article-title>Ecapa-tdnn: Emphasized channel attention</article-title>
          , propa-
        </mixed-citation>
      </ref>
      <ref id="ref27">
        <mixed-citation>
          <article-title>cation system using universal adversarial pertur- ifcation</article-title>
          , arXiv preprint arXiv:
          <year>2005</year>
          .
          <volume>07143</volume>
          (
          <year>2020</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref28">
        <mixed-citation>
          bations, in: ICASSP 2021
          <article-title>-</article-title>
          2021 IEEE International [21]
          <string-name>
            <given-names>J.</given-names>
            <surname>Zhang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Tan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Deng</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Hu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Zhu</surname>
          </string-name>
          , L. Huang,
        </mixed-citation>
      </ref>
      <ref id="ref29">
        <mixed-citation>
          <article-title>cessing (ICASSP)</article-title>
          , IEEE,
          <year>2021</year>
          , pp.
          <fpage>2575</fpage>
          -
          <lpage>2579</lpage>
          .
          <article-title>method for generating adversarial examples for</article-title>
          [10]
          <string-name>
            <given-names>Y.</given-names>
            <surname>Xie</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Shi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Liu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Chen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Yuan</surname>
          </string-name>
          ,
          <article-title>speaker recognition</article-title>
          ,
          <source>in: 2022 7th IEEE</source>
          Interna-
        </mixed-citation>
      </ref>
      <ref id="ref30">
        <mixed-citation>
          <article-title>against speaker recognition systems, in: ICASSP (DSC)</article-title>
          , IEEE,
          <year>2022</year>
          , pp.
          <fpage>167</fpage>
          -
          <lpage>174</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref31">
        <mixed-citation>2020-2020 IEEE international conference on acous-</mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>