<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Description of a Multi-Stage Audio Spoofing System in ADD Challenge 2023</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Hua Hua</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Jingze Lu</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Peiyang Shi</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Zengqiang Shang</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Yuxiang Zhang</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Xuyuan Li</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Pengyuan Zhang</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Institude of Acoustics, Chinese Academy of Sciences</institution>
          ,
          <addr-line>19 N.4th Ring West Rd., Haidian Dist., Beijing</addr-line>
          ,
          <country country="CN">China</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2023</year>
      </pub-date>
      <fpage>49</fpage>
      <lpage>57</lpage>
      <abstract>
        <p>This paper is a detailed description of a multi-stage audio synthesis system that participated in Track1.1 (spoofing) in ADD Challenge 2023 in which we play the role of an attacker. These stages include a fully end-to-end text-to-speech model, a fully end-to-end any-to-many voice conversion model, and an adversarial attacking model. We believe that such a design can reduce the possibility of artifact exposure at multiple levels and dimensions so that it is closer to real speech and is becoming hard to extinguish. Besides, we adopt post-processing methods to further improve our spoofing capability against detection methods for non-speech parts. Our system won 3rd place in the total score of Track1.1 in this challenge, especially, the performance in attacking black-box systems ranked 1st.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;Speech synthesis</kwd>
        <kwd>Anti-spoofing</kwd>
        <kwd>Deception success rate</kwd>
        <kwd>Adversarial attacking</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>is judged to be more authentic.</p>
      <p>The system designed in our paper adopts a multi-stage
Audio deep synthesis techniques have been able to gener- speech generation method as shown in Figure 1. The first
ate high-quality speech whose authenticity is dificult for stage is a state-of-art text-to-speech (TTS) module based
humans to recognize. Meanwhile, as the quality of deeply on hierarchical feature modeling, the second stage is an
synthesized speech approaches human’s natural voice, end-to-end any-to-many voice conversion (VC) module
anti-spoofing systems have emerged and been continu- leveraging intermediate features from automatic speech
ously upgraded for security purposes. So audio deepfake recognition (ASR), and the third stage is a gradient-based
and anti-spoofing are in a game of attacking and defend- adversarial attacking (AA) module. In addition, we also
ing. So far, there have been several authoritative chal- come up with some countermeasures against non-speech
lenges for speech deepfake and anti-spoofing tasks such part detection, which we call the post-processing (PP)
as ASVSpoof-2017/2019/2021[1, 2, 3] and ADD-2022[4]. module.</p>
      <p>
        T
        <xref ref-type="bibr" rid="ref6">his ADD-2023</xref>
        Challenge also continues this theme[5]. Why are we building such a heavy cascading pipeline
      </p>
      <p>Track 1.1 of this challenge focuses on speech spoofing with so many stages? The motivation along with
reattacks against anti-spoofing systems. Each participating lated works will be discussed in Section 2. The system
team is required to design a speech generation system, in- framework and detailed methods adopted in diferent
put the given text and speaker information, and generate components will be described in Section 3. Training
setthe corresponding fake speech. Synthesized audio will tings will be described in Section 4. Experiments and our
be sent to a group of black-box anti-spoofing systems rankings in the ADD Challenge will be demonstrated in
(from submissions of other tracks) and a white-box anti- Section 5.
spoofing system (the oficial baseline) for detection, and
the deception success rate (DSR) will be used to measure
the performance of a synthesizing system. A higher DSR 2. Related works and motivation
means stronger deception ability, in other words, it can
be an indication that the audio generated by the system
TTS and VC are techniques that convert certain types
of inputs to human speech waveforms – TTS accepts
raw text or phoneme sequences, while VC takes in the
voice of a source speaker. With the development of deep
learning, research on TTS and VC based on deep neural
networks has been carried out in large numbers.</p>
      <p>TTS technology has made remarkable progress in its
goal of generating natural speech which is close to
human speaking[6, 7]. Representative contributions can be</p>
      <p>Sentence</p>
      <p>Word
Phone</p>
      <p>Hierarchical
latent representation
Text</p>
      <p>Hier-TTS
Text-to-speech (TTS)</p>
      <p>Voice conversion (VC)</p>
      <p>E2EVC</p>
      <p>Target
speaker
PPG intermediate
features</p>
      <p>Attacking
attempt</p>
      <p>White-box
Anti-spoofing</p>
      <p>System</p>
      <p>Return</p>
      <p>Gradient
characteristics
Project Gradient</p>
      <p>Descent
Adversarial attacking (AA)</p>
      <p>Post-processing (PP)
Silence Replace</p>
      <p>Global Additive
Real silence
sample</p>
      <p>Parabolic
crossfading</p>
      <p>Final Waveform
based series[12, 13], fully end-to-end VITS[14], difusion- through the hierarchical VAE structure, so that the
genbased models[15, 16], etc. erated speech has good long-term features at multiple</p>
      <p>Considering VC, models based on generative adversar- levels such as sentence level, phrase level, word level,
ial networks (GAN) along with their variants, such as and frame level.</p>
      <p>StarGAN-VC[17] and CycleGAN-VC[18] use a generator Additionally, machine learning models may misclassify
or a conditional generator that transforms the source examples that are only slightly diferent from correctly
speaker’s voice features to the target speaker’s directly. classified examples drawn from the data distribution[ 33].
Auto-encoder-based VC systems such as AutoVC[19] and In many cases, a wide variety of models with diferent
VQVC[20] learn to reproduce their input as the output, architectures trained on diferent subsets of the training
exhibiting efectiveness in disentangling speaker iden- data misclassify the same adversarial example. This
sugtity information from linguistic content. Benefiting from gests that adversarial examples expose fundamental blind
automatic speech recognition (ASR) systems pre-trained spots, so we can take advantage of this characteristic to
with large corpora, models based on phonetic posteri- create specific samples to attack the classification system,
orgrams (PPGs) are considered to have an advantage in which is called an adversarial attack (AA). The
repreextracting speaker-independent acoustic features from sentative methods include Fast Gradient Sign Method
source speech[21, 22]. (FGSM)[33], and the Projected Gradient Descent (PGD)</p>
      <p>
        From the perspective of spoofing, we understand that algorithm[34] involved in this paper is exactly a variant
most anti-spoofing systems do not target specific arti- of FGSM. Such an AA stage may further compensate
facts, but use more general features[23, 24, 25, 26]. There- for the shortcomings of the aforementioned generative
fore, when synthesizing speech, we will try our best to models, hiding possible exposed artifacts under
adversarmove in a direction that is more authentic to human per- ial perturbations especially confronting white-box
antiception. But in the last year or two, the interpretability of spoofing systems.
antispoofing has become stronger and stronger, and some This is not enough. Some anti-spoofing systems even
representative works have focused on the detection of pay attention to non-speech parts like breath sounds
artifacts generated by vocoders[27, 28]. Therefore, in this and silent segments because generative models usually
article, we consider making the VC stage a complete end- lack attention to these non-speech parts[35, 36]. Our
to-end form[29], and adopt certain strategies to alleviate VC system can reduce the false creation of unreal
harartifacts such as chessboard efects. In addition, some monic structures in breath sounds because its front end
counterfeiting studies have begun to notice the inade- is an ASR-based phoneme classifier. Nevertheless,
requacy of the acoustic model, for example, detecting some garding silence, it is no
        <xref ref-type="bibr" rid="ref17">t that easy. In the ASVspoof 2019</xref>
        unnatural aspects of synthesized speech i
        <xref ref-type="bibr" rid="ref20">n long-term fea- and 2021</xref>
        databases, the silence diferences are
importures that traditional anti-spoofing methods do not con- tant artifacts that influence countermeasures[ 2, 3]. We
sider, such as prosody, style, and emotion[30, 31]. There- know that one real speech recording usually has natural
fore, our TTS stage adopts the latest Hier-TTS model silent segments at the beginning and the end, as well
of our team[32], which is mainly trained step by step as between syllables, words, or phrases, which is
determined by semantic or prosodic boundaries. On the other a fine-to-coarse manner, and then the HAD reconstructs
hand, synthesized speech especially through VC, tends the speech waveform from coarse-to-fine leveraging
exto exhibit some anomalies diferent from real speech: no tracted hierarchical latent variables. To introduce text
silence, completely zero amplitude, or unexpected noises. information, HCE obtains linguistic and phonological
inAlthough this will not cause severe problems in common formation at diferent scales from phoneme and character
application scenarios as human listeners will probably sequences and then injects them into each hierarchy of
not find this problem, there is a high chance that these ar- HAE and HAD. For the modeling of duration, we inject
tifacts will be detected by an anti-spoofing system. There- phoneme-scale durations at the phoneme level encoder
fore, we adopt two non-neural post-processing (PP) algo- and reconstruct the durations using the phoneme decoder.
rithms that aim at implementing near-realistic silences Thus, the duration and waveform reconstruction share
and pauses against silence detection[29]. part of the hierarchical hidden variables, which facilitates
learning more consistent prosody.
      </p>
    </sec>
    <sec id="sec-2">
      <title>3. System framework and methods</title>
      <p>The system designed is a multi-stage waveform
generation pipeline. The first stage is a state-of-art
text-tospeech TTS module based on hierarchical feature
modeling. This module analyzes the input text to predict
prosody and style features at the word level, phrase level,
and sentence level, and finally generates speech
waveforms with high naturalness in long-term features. The
second stage is an end-to-end any-to-many VC module
leveraging PPG features extracted from a pre-trained
ASR. We introduce the desired voiceprint of the target
speaker at this stage and transfer the aforementioned
waveforms to the target speaker’s timbre. The third AA
stage is optional: It is dificult for us to impose this attack
on the unknown black-box antispoofing systems, so we
skip the stage; but confronting the white-box baseline
system whose model is known, we use the adversarial
attack method to further improve our spoofing ability.
Finally, the generated speech waveform will also be
processed in the PP stage to enhance the resistance to the
detection of non-speech parts.</p>
      <sec id="sec-2-1">
        <title>3.1. TTS with hierarchical variational autoencoders</title>
        <p>We adopt Hier-TTS [32] as the very first stage of the
whole model that analyzes and processes text inputs and
returns raw speech waveform. Because of its powerful
ability to model the unified space of text and audio,
HierTTS can reconstruct factors such as prosody and style
better than common TTS models, therefore it may achieve
better performance against some recent anti-spoofing
systems which take inner-speaker long-time consistency
of prosody or style into consideration.</p>
        <p>Figure 2 is the overall structure which contains a
Hierarchical Audio Encoder (HAE), a Hierarchical Context
Encoder (HCE), and a Hierarchical Audio Decoder (HAD).
Hier-TTS introduce five latent variables at diferent
temporal resolution, including sentence, word, subword,
phoneme, and frame level. First, HAE extracts the
hierarchical latent variables from the linear spectrogram in</p>
      </sec>
      <sec id="sec-2-2">
        <title>3.2. Any-to-many VC based on phonetic posteriorgrams</title>
        <p>The VC model features the structure proposed in [29],
which can be divided into five components as shown
in Figure 3. In short, it incorporates a conformer
encoder from an auto-speech-recognition (ASR) model and
a series of transformer blocks on one stream. On
another stream, a posterior encoder upon linear
spectrograms is constructed. Outputs from those two modules
are constrained to be subject to the same distribution,
hence reducing the gap between the generated latent
variables and the distribution of the real features. Then
a re-parameterization module followed by a GAN-based
decoder is leveraged to convert hidden features to the
target waveform. This waveform is then post-processed
to become the final generated speech. For the ASR, we
utilize an encoder a hybrid CTC/Attention model[37] to
extract the features of phonetic posteriorgrams(PPGs)
from given speech.</p>
        <p>We use the non-causal WaveNet[38] residual blocks
as the posterior encoder. Such residual block consists of
several dilated convolutions with skip connection and
gated activation. For multi-speaker VC, we additionally
feed speaker embedding into residual blocks through the
method of global conditioning.</p>
        <p>The conversion decoder is almost essentially the
HiFiGAN[39] generator. The generator is a stack of
convolution blocks, which include transpose convolution
layers and a multi-receptive field fusion module (MRF).
The MRF is composed of a series of residual blocks that
have receptive fields of diferent sizes. To avoid
possible checkerboard artifacts [40] caused by the transpose
convolution process, we rebuild the upsampling layer
using temporal nearest interpolation followed by a 1D
convolution. As with GAN-based vocoders, we also add
a discriminator D that attempts to distinguish audio
generated by the generator G from the ground truth. Similar
to the design in Hifi-GAN[ 39] and MelGAN[41], the
discriminators include a multi-period discriminator(MPD)
and a multi-scale discriminator (MSD). MPD is a mixture
of window-based sub-discriminators, where diferent
pe</p>
        <sec id="sec-2-2-1">
          <title>Hierarchical Audio Encoder (HAE)</title>
        </sec>
        <sec id="sec-2-2-2">
          <title>Word2Sentence Encoder</title>
        </sec>
        <sec id="sec-2-2-3">
          <title>Subword2Word Encoder</title>
        </sec>
        <sec id="sec-2-2-4">
          <title>Phone2Subword Encoder</title>
          <p>target</p>
        </sec>
        <sec id="sec-2-2-5">
          <title>Frame2Phone Encoder</title>
        </sec>
        <sec id="sec-2-2-6">
          <title>Frame Encoder</title>
        </sec>
        <sec id="sec-2-2-7">
          <title>Linear Spectrogram</title>
        </sec>
        <sec id="sec-2-2-8">
          <title>Hierarchical Audio Decoder (HAD) Hierarchical Context Encoder (HCE)</title>
          <p>word
sub−
−
−
−
sample
sample
sample</p>
        </sec>
        <sec id="sec-2-2-9">
          <title>Sentence2Word Decoder</title>
        </sec>
        <sec id="sec-2-2-10">
          <title>Word2Subword Decoder</title>
        </sec>
        <sec id="sec-2-2-11">
          <title>Subword2Phone Decoder sample</title>
          <p>predict</p>
        </sec>
        <sec id="sec-2-2-12">
          <title>Phone2Frame Decoder sample</title>
        </sec>
        <sec id="sec-2-2-13">
          <title>Frame Decoder</title>
        </sec>
        <sec id="sec-2-2-14">
          <title>Waveform</title>
        </sec>
        <sec id="sec-2-2-15">
          <title>Word2Sentence Encoder</title>
        </sec>
        <sec id="sec-2-2-16">
          <title>Subword2Word Encoder</title>
        </sec>
        <sec id="sec-2-2-17">
          <title>Character Embedding</title>
        </sec>
        <sec id="sec-2-2-18">
          <title>Phone2Subword Encoder</title>
        </sec>
        <sec id="sec-2-2-19">
          <title>Phoneme Embedding</title>
          <p>Feature combination
Calculating KL Divergence
For training
For inference</p>
          <p>+1 = ∏︁ ( +  (∇(, ,  )))</p>
          <p>(1)
+
Where  represents the original sample,  denotes
sample label/class,  is the number of iterations,  is the
maximum perturbation range,  is the perturbation size
riodic patterns of waveform are operated. MSD directly for each iteration, and L represents the loss function of
functions on diferent scales, which consecutively evalu- the target model. The PGD attack generates the
adverates audio samples at diferent levels that help to capture sarial perturbation that maximizes the sample loss value
consecutive patterns and long-term dependencies. through multiple iterations. Adversarial training can be
understood as solving an optimization problem with an
3.3. Adversarial attacking external minimum and an internal maximum. It solves
decision boundaries that are robust to maximum
adverProjected Gradient Descent (PGD) [34] is considered to sarial perturbations. For speech, adversarial attacks can
be the strongest attack method based on gradient infor- target both the original audio and the acoustic features
mation. When the model is robust to PGD attacks, it is used by the model, which need to be considered at the
robust to most gradient-based attack methods. The PGD same time during adversarial training. The optimization
attack uses the gradient of the model to the input to find function for adversarial training is shown below:
the disturbance that maximizes the loss value. A typical
PGD attack on network  is mathematically expressed as
the following recursive formula:
 , ℎ ( ) = (,)[ ∈ (,  + ,  )]</p>
          <p>(2)
Among them,  ( ) is the external minimum optimization
function, which optimizes the network parameters for
the most adversarial example of a clean sample, so that
the network can be classified correctly.
Source
1
12 (2 − )</p>
          <p>&gt; 
0 ≤  ≤ 
 2() =</p>
          <p>︂{
for  from 0 to :
for  from 0 to ℎ:
() ←</p>
          <p>1(, )()
ℎ() ←</p>
          <p>2()ℎ()
Connect(,,ℎ)=
⎧
⎪
⎪
⎪
⎪
⎪
⎪
⎪
⎪
⎪
⎪
⎪
⎪
⎨ () + ℎ( +  −</p>
          <p>)
()
0 ≤  &lt;  −  ,
 −</p>
          <p>≤  &lt; ,
ℎ( +  −</p>
          <p>)
⎩  ≤</p>
          <p>≤  + ℎ − 
where  and ℎ denote the raw silent segments
to be concatenated.</p>
          <p>and ℎ represent the
lengths of these two segments.  and  in the
formulas are the subscripts of the time dimension iteration. 
stands for the overlapping duration. The final result is
calculated by Connect(·).</p>
          <p>Further more, we assume that a decline in speech
quality may cause the performance of some anti-spoofing
systems to degrade correspondingly and thus increase
the probability of a successful attack. We reduce the</p>
          <p>Mel
Spectrogram</p>
          <p>Conformer
Encoder
×N</p>
          <p>Transformer</p>
          <p>Blocks
×N</p>
          <p>Latent
Representation
Re-parameterize</p>
          <p>Coding
Vector</p>
          <p>Decoder
(Generator)
recordings of the target speaker. We crop those real seg- for  from 0 to +ℎ- :</p>
        </sec>
      </sec>
      <sec id="sec-2-3">
        <title>3.4. Post-processing</title>
        <p>Similar to what is proposed in our previous work[29],
we introduce some postprocessing countermeasures
against silence detection and check-out employed by
antispoofing. We experiment with two methods of appending
silence: In the first approach, we leverage a non-neural
voice-activity-detecting(VAD) system that quickly finds
silent segments in fake speech and gets the timestamps.</p>
        <p>Then we randomly select real silent segments from the
ments to the same length as calculated from the
timestamps given by VAD and then simply replace them. The
other is global superposition without relying on VAD.</p>
        <p>We also randomly select multiple real silent segments
and normalize their amplitudes to an average level. After
that, we connect them end to end through the algorithm
of parabolic cross-fading till the total length surpasses
the length of generated audio. Finally, we directly paste
this stitched silent sample onto our synthetic speech like
additive noise. The parabolic cross-fading algorithm is
formulated as follows:
 1(, ) =
⎧
⎪
⎪
⎨
⎪
⎩
⎪ 12 ( −</p>
        <p>1
 &lt;  −  ,
 −  ≤  ≤ 
2 − )( − )
(3)
speech quality of speech by adding white Gaussian noise anchored in:
and using the signal-to-noise ratio (SNR) to characterize 1. A complete end-to-end training method has been
our speech quality. We notice that when the SNR is too adopted–the boundary between the acoustic model and
low, the intelligibility of speech drops so rapidly that it the vocoder is ambiguous so the part that acts as a
becomes nearly unintelligible and cannot pass ASR, so vocoder in the synthesis pipeline is somewhat diferent
there is probably a trade-of and we manage to maintain from a traditional autoregressive/GAN/DPM model. This
the SNR at a relatively high level after perturbation. may have led to anti-spoofing models submitted by many
teams that have not seen the patterns of our generated
speech.
4. Experiments and results 2. We have considered the existence of some long-term
artifacts based on human hearing perception, and the
We tested the spoofing ability of our synthesized audio on HierTTS and E2EVC models adopted have been designed
several diferent mainstream anti-spoofing models along to suppress these artifacts to a certain extent.
with various training features. For each anti-spoofing 3. AA and PP methods have increased the complexity
model, we sent 1000 pieces of the generated waveform of generating speech, masking some artifacts and causing
to perform our test, and the equal error rate (EER) was possible bias in the feature extraction of the anti-spoofing
considered to be the performance indicator. Results are model. Methods that use non-speech parts as the main
shown below in Table 1. Judging from the results, AA basis for detection are also targeted.
and PP methods have achieved some efects in the face For the given AsvSpoof baseline released by the
chalof a bunch of anti-spoofing systems. lenge organizers, the adversarial attack method we adopt</p>
        <p>In the original design that was abandoned, we also has achieved some efect. ∼ 1e-5 and 0.99 represent the
tried to incorporate a vocoder of Difusion Probabilis- average detection score before and after adding the
adtic Models (DPM) after the VC module, because in dif- versarial attack respectively. We are informed that the
fusion models, a segment of generated signal is iter- judging threshold is 0.1, so the performance against
baseated from white noise. Unlike most vocoders that use a line is supposed to be DSR 100%. However, according to
deconvolution-like upsampling method, difusion models the Challenge results, our spoofing performance against
will be less likely to produce artifacts in the frequency the white-box baseline remains only DSR 23%. The
redomain. However, the actual efect is not obvious or may sults are inconsistent. Our follow-up tests suggest that
reduce performance, according to Table 2, which demon- the PGD noise used in the AA method may cause some
strates the results of this part of the ablation experiment. negative efects after being enhanced by PP approaches,</p>
        <p>Our team’s system achieved the 3rd place in the overall but we are not confident enough.
score in this spoofing track. It is worth noting that our We understand that in any case, there is still room
synthesized speech had the highest DSR% against all for improvement in the attack capabilities of white-box
black-box anti-spoofing systems in two rounds of testing. and black-box systems, such as further dismantling the
Table 3 gives the specific performance data of all teams. GAN-like vocoder structure of the complete end-to-end
Our team is A03. VC part, or solving the problem that PGD may play a</p>
        <p>We believe that the reasons why our multi-stage sys- negative role.
tem is more powerful against black boxes are mainly</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>5. Conclusions</title>
      <p>
        This paper is a detailed description of a multi-stage audio
synthesis system that participated in Track1.1 (spoofing)
in ADD C
        <xref ref-type="bibr" rid="ref6">hallenge 2023</xref>
        . The system we build takes into
account the ability to hide artifacts at diferent feature
levels. Moreover, we use adversarial sample attacks and
post-processing methods to further improve our spoofing
capabilities. Experiments and challenge results show that
our ability to attack black-box anti-spoofing systems is
relatively good, but some modules may have a negative
efect on the white-box baseline. In the follow-up work,
we will continue to solve these problems to strengthen
the ofensive and defensive capabilities in the field of
speech deepfake and anti-spoofing.
      </p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          <string-name>
            <surname>tion</surname>
          </string-name>
          ,
          <year>2021</year>
          , pp.
          <fpage>47</fpage>
          -
          <lpage>54</lpage>
          . doi:
          <volume>10</volume>
          .21437/ASVSPOOF.
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2021-
          <fpage>8</fpage>
          . [4]
          <string-name>
            <given-names>J.</given-names>
            <surname>Yi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Fu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Tao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Nie</surname>
          </string-name>
          , H. Ma, C. Wang,
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          <string-name>
            <given-names>S.</given-names>
            <surname>Zhang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Yan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Xu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Wen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <year>Add 2022</year>
          :
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          <year>2022</year>
          , pp.
          <fpage>9216</fpage>
          -
          <lpage>9220</lpage>
          . doi:
          <volume>10</volume>
          .1109/ICASSP43922.
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          <year>2022</year>
          .
          <volume>9746939</volume>
          . [5]
          <string-name>
            <given-names>J.</given-names>
            <surname>Yi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Tao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Fu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Yan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Wang</surname>
          </string-name>
          , C. Y.
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          <string-name>
            <given-names>H.</given-names>
            <surname>Gu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Wen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Liang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Lian</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Li</surname>
          </string-name>
          , Add2023: the
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          <string-name>
            <surname>cepted by</surname>
            <given-names>IJCAI</given-names>
          </string-name>
          2023 Workshop on Deepfake Audio
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          <string-name>
            <surname>Detection</surname>
            and
            <given-names>Analysis (DADA</given-names>
          </string-name>
          <year>2023</year>
          ),
          <year>2023</year>
          . [6]
          <string-name>
            <given-names>X.</given-names>
            <surname>Tan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Qin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Soong</surname>
          </string-name>
          , T.-Y. Liu, A survey on neu-
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          <source>ral speech synthesis</source>
          ,
          <year>2021</year>
          . arXiv:
          <volume>2106</volume>
          .
          <fpage>15561</fpage>
          . [7]
          <string-name>
            <given-names>B.</given-names>
            <surname>Sisman</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Yamagishi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>King</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <article-title>An overview</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          (
          <year>2020</year>
          ). [8]
          <string-name>
            <given-names>Y.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Skerry-Ryan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Stanton</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Wu</surname>
          </string-name>
          , R. J.
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          <year>2017</year>
          . arXiv:
          <volume>1703</volume>
          .
          <fpage>10135</fpage>
          . [9]
          <string-name>
            <given-names>J.</given-names>
            <surname>Shen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Pang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R. J.</given-names>
            <surname>Weiss</surname>
          </string-name>
          , M. Schuster,
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          <source>ing wavenet on mel spectrogram predictions</source>
          ,
          <year>2018</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          <source>arXiv:1712</source>
          .
          <fpage>05884</fpage>
          . [1]
          <string-name>
            <given-names>T.</given-names>
            <surname>Kinnunen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Sahidullah</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Delgado</surname>
          </string-name>
          , M. Todisco,
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          <article-title>2017 challenge: Assessing the limits of replay</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          <source>spoofing attack detection</source>
          ,
          <year>2017</year>
          . doi:
          <volume>10</volume>
          .21437/
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          <string-name>
            <surname>Interspeech.</surname>
          </string-name>
          <year>2017</year>
          -
          <volume>1111</volume>
          . [2]
          <string-name>
            <given-names>M.</given-names>
            <surname>Todisco</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Vestman</surname>
          </string-name>
          , M. Sahidullah,
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          <string-name>
            <given-names>T.</given-names>
            <surname>Kinnunen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K. A.</given-names>
            <surname>Lee</surname>
          </string-name>
          ,
          <year>Asvspoof 2019</year>
          : Fu-
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          <string-name>
            <surname>tection</surname>
          </string-name>
          ,
          <year>2019</year>
          , pp.
          <fpage>1008</fpage>
          -
          <lpage>1012</lpage>
          . doi:
          <volume>10</volume>
          .21437/
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          <string-name>
            <surname>Interspeech.</surname>
          </string-name>
          <year>2019</year>
          -
          <volume>2249</volume>
          . [3]
          <string-name>
            <given-names>J.</given-names>
            <surname>Yamagishi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Todisco</surname>
          </string-name>
          , M. Sahidullah,
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          <string-name>
            <given-names>N.</given-names>
            <surname>Evans</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Delgado</surname>
          </string-name>
          ,
          <year>Asvspoof 2021</year>
          :
          <fpage>accelerat</fpage>
          -
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          <article-title>ing progress in spoofed and deepfake speech detec-</article-title>
          [10]
          <string-name>
            <given-names>Y.</given-names>
            <surname>Ren</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Ruan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Tan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Qin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Zhao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Zhao</surname>
          </string-name>
          , T.-
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          <article-title>to speech, in: Proceedings of the 33rd International Aasist: Audio anti-spoofing using integrated</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          <source>Conference on Neural Information Processing Sys- spectro-temporal graph attention networks</source>
          ,
          <year>2021</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          <string-name>
            <surname>tems</surname>
          </string-name>
          ,
          <year>2019</year>
          , pp.
          <fpage>3171</fpage>
          -
          <lpage>3180</lpage>
          . arXiv:
          <volume>2110</volume>
          .
          <fpage>01200</fpage>
          . [11]
          <string-name>
            <given-names>Y.</given-names>
            <surname>Ren</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Hu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Tan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Qin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Zhao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Zhao</surname>
          </string-name>
          , T.-Y. [24]
          <string-name>
            <given-names>H.</given-names>
            <surname>Tak</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Todisco</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Wang</surname>
          </string-name>
          , J. weon Jung, J. Yamag-
        </mixed-citation>
      </ref>
      <ref id="ref25">
        <mixed-citation>
          <string-name>
            <surname>Liu</surname>
          </string-name>
          ,
          <article-title>Fastspeech 2: Fast and high-quality end-to-end ishi, N. Evans, Automatic speaker verification spoof-</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref26">
        <mixed-citation>
          text to speech,
          <year>2022</year>
          . arXiv:
          <year>2006</year>
          .
          <article-title>04558. ing and deepfake detection using wav2vec 2.0</article-title>
          and [12]
          <string-name>
            <given-names>C.</given-names>
            <surname>Miao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Liang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Chen</surname>
          </string-name>
          , J. Ma,
          <string-name>
            <given-names>S.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Xiao</surname>
          </string-name>
          , data augmentation,
          <year>2022</year>
          . arXiv:
          <volume>2202</volume>
          .
          <fpage>12233</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref27">
        <mixed-citation>
          <article-title>Flow-tts: A non-autoregressive network for text to</article-title>
          [25]
          <string-name>
            <given-names>H.</given-names>
            <surname>Tak</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Patino</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Todisco</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Nautsch</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Evans</surname>
          </string-name>
          ,
        </mixed-citation>
      </ref>
      <ref id="ref28">
        <mixed-citation>
          <article-title>speech based on flow</article-title>
          , in: ICASSP 2020
          <article-title>- 2020 IEEE A. Larcher, End-to-end anti-spoofing with rawnet2,</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref29">
        <mixed-citation>
          <source>International Conference on Acoustics, Speech and 2021</source>
          . arXiv:
          <year>2011</year>
          .01108.
        </mixed-citation>
      </ref>
      <ref id="ref30">
        <mixed-citation>
          <source>Signal Processing (ICASSP)</source>
          ,
          <year>2020</year>
          , pp.
          <fpage>7209</fpage>
          -
          <lpage>7213</lpage>
          . [26]
          <string-name>
            <given-names>H.</given-names>
            <surname>Tak</surname>
          </string-name>
          , J. weon Jung,
          <string-name>
            <given-names>J.</given-names>
            <surname>Patino</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Todisco</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Evans</surname>
          </string-name>
          ,
        </mixed-citation>
      </ref>
      <ref id="ref31">
        <mixed-citation>
          <source>doi:10.1109/ICASSP40776</source>
          .
          <year>2020</year>
          .
          <volume>9054484</volume>
          .
          <article-title>Graph attention networks for anti-</article-title>
          spoofing,
          <year>2021</year>
          . [13]
          <string-name>
            <given-names>J.</given-names>
            <surname>Kim</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Kim</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Kong</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Yoon</surname>
          </string-name>
          ,
          <article-title>Glow-tts: A gener-</article-title>
          arXiv:
          <fpage>2104</fpage>
          .
          <fpage>03654</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref32">
        <mixed-citation>
          <article-title>ative flow for text-to-speech via monotonic align-</article-title>
          [27]
          <string-name>
            <given-names>X.</given-names>
            <surname>Yan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Yi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Tao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Wang</surname>
          </string-name>
          , H. Ma, T. Wang,
        </mixed-citation>
      </ref>
      <ref id="ref33">
        <mixed-citation>
          <source>ment search</source>
          ,
          <year>2020</year>
          . arXiv:
          <year>2005</year>
          .11129.
          <string-name>
            <given-names>S.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Fu</surname>
          </string-name>
          , An initial investigation for de[14]
          <string-name>
            <given-names>J.</given-names>
            <surname>Kim</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Kong</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Son</surname>
          </string-name>
          ,
          <article-title>Conditional variational au- tecting vocoder fingerprints of fake audio</article-title>
          , in:
        </mixed-citation>
      </ref>
      <ref id="ref34">
        <mixed-citation>
          <article-title>toencoder with adversarial learning for end-to-end</article-title>
          <source>Proceedings of the 1st International Workshop on</source>
        </mixed-citation>
      </ref>
      <ref id="ref35">
        <mixed-citation>
          <string-name>
            <surname>Machine</surname>
            <given-names>Learning</given-names>
          </string-name>
          , PMLR,
          <year>2021</year>
          , pp.
          <fpage>5530</fpage>
          -
          <lpage>5540</lpage>
          . '22,
          <string-name>
            <surname>Association</surname>
            for Computing Machinery, New [15]
            <given-names>M.</given-names>
          </string-name>
          <string-name>
            <surname>Jeong</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          <string-name>
            <surname>Kim</surname>
            ,
            <given-names>S. J.</given-names>
          </string-name>
          <string-name>
            <surname>Cheon</surname>
            ,
            <given-names>B. J.</given-names>
          </string-name>
          <string-name>
            <surname>Choi</surname>
            ,
            <given-names>N. S.</given-names>
          </string-name>
          <string-name>
            <surname>Kim</surname>
          </string-name>
          , York, NY, USA,
          <year>2022</year>
          , p.
          <fpage>61</fpage>
          -
          <lpage>68</lpage>
          . URL: https://doi.org/
        </mixed-citation>
      </ref>
      <ref id="ref36">
        <mixed-citation>
          <article-title>Dif-tts: A denoising difusion model for text</article-title>
          -to-
          <volume>10</volume>
          .1145/3552466.3556525. doi:
          <volume>10</volume>
          .1145/3552466.
        </mixed-citation>
      </ref>
      <ref id="ref37">
        <mixed-citation>
          <string-name>
            <surname>speech</surname>
          </string-name>
          ,
          <year>2021</year>
          . arXiv:
          <volume>2104</volume>
          .
          <fpage>01409</fpage>
          . 3556525. [16]
          <string-name>
            <given-names>V.</given-names>
            <surname>Popov</surname>
          </string-name>
          ,
          <string-name>
            <given-names>I.</given-names>
            <surname>Vovk</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Gogoryan</surname>
          </string-name>
          , T. Sadekova, [28]
          <string-name>
            <given-names>C.</given-names>
            <surname>Sun</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Jia</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Hou</surname>
          </string-name>
          , E. AlBadawy, S. Lyu, Exposing
        </mixed-citation>
      </ref>
      <ref id="ref38">
        <mixed-citation>
          <article-title>abilistic model for text-to-</article-title>
          <string-name>
            <surname>speech</surname>
          </string-name>
          ,
          <year>2021</year>
          . artifacts,
          <year>2023</year>
          . arXiv:
          <volume>2302</volume>
          .
          <fpage>09198</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref39">
        <mixed-citation>
          <source>arXiv:2105</source>
          .
          <fpage>06337</fpage>
          . [29]
          <string-name>
            <given-names>H.</given-names>
            <surname>Hua</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Chen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Zhang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Zhang</surname>
          </string-name>
          , Improv[17]
          <string-name>
            <given-names>H.</given-names>
            <surname>Kameoka</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Kaneko</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Tanaka</surname>
          </string-name>
          ,
          <string-name>
            <surname>N.</surname>
          </string-name>
          <article-title>Hojo, ing spoofing capability for end-to-end any-to-many</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref40">
        <mixed-citation>
          <article-title>Stargan-vc: Non-parallel many-to-many voice con- voice conversion</article-title>
          ,
          <source>in: Proceedings of the 1st Interna-</source>
        </mixed-citation>
      </ref>
      <ref id="ref41">
        <mixed-citation>
          <source>in: 2018 IEEE Spoken Language Technology Work- Multimedia</source>
          , Association for Computing Machinery,
        </mixed-citation>
      </ref>
      <ref id="ref42">
        <mixed-citation>
          <article-title>shop (SLT)</article-title>
          , IEEE,
          <year>2018</year>
          , pp.
          <fpage>266</fpage>
          -
          <lpage>273</lpage>
          .
          <year>2022</year>
          . [18]
          <string-name>
            <given-names>T.</given-names>
            <surname>Kaneko</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Kameoka</surname>
          </string-name>
          , Cyclegan-vc: Non-parallel [30]
          <string-name>
            <given-names>C.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Yi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Tao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Zhang</surname>
          </string-name>
          , S. Zhang,
        </mixed-citation>
      </ref>
      <ref id="ref43">
        <mixed-citation>
          networks,
          <source>in: 2018 26th European Signal Processing based on prosodic and pronunciation features</source>
          ,
          <year>2023</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref44">
        <mixed-citation>
          <string-name>
            <surname>Conference (EUSIPCO),</surname>
            <given-names>IEEE</given-names>
          </string-name>
          ,
          <year>2018</year>
          , pp.
          <fpage>2100</fpage>
          -
          <lpage>2104</lpage>
          . arXiv:
          <volume>2305</volume>
          .
          <fpage>13700</fpage>
          . [19]
          <string-name>
            <given-names>K.</given-names>
            <surname>Qian</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Zhang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Chang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Yang</surname>
          </string-name>
          , M. Hasegawa- [31]
          <string-name>
            <given-names>E.</given-names>
            <surname>Conti</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Salvi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Borrelli</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Hosler</surname>
          </string-name>
          , P. Bestagini,
        </mixed-citation>
      </ref>
      <ref id="ref45">
        <mixed-citation>
          <source>Conference on Machine Learning, PMLR</source>
          ,
          <year>2019</year>
          , pp. nition
          <article-title>: A semantic approach</article-title>
          , in: ICASSP
        </mixed-citation>
      </ref>
      <ref id="ref46">
        <mixed-citation>
          5210-
          <fpage>5219</fpage>
          .
          <fpage>2022</fpage>
          - 2022 IEEE International Conference on [20]
          <string-name>
            <surname>D.-Y. Wu</surname>
            ,
            <given-names>Y.-H.</given-names>
          </string-name>
          <string-name>
            <surname>Chen</surname>
            , H.-
            <given-names>Y.</given-names>
          </string-name>
          <string-name>
            <surname>Lee</surname>
          </string-name>
          , Vqvc+:
          <article-title>One-shot Acoustics, Speech and Signal Processing (ICASSP),</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref47">
        <mixed-citation>
          <article-title>voice conversion by vector quantization</article-title>
          and u-net
          <year>2022</year>
          , pp.
          <fpage>8962</fpage>
          -
          <lpage>8966</lpage>
          . doi:
          <volume>10</volume>
          .1109/ICASSP43922.
        </mixed-citation>
      </ref>
      <ref id="ref48">
        <mixed-citation>
          <string-name>
            <surname>architecture</surname>
          </string-name>
          ,
          <year>2020</year>
          . arXiv:
          <year>2006</year>
          .04154.
          <year>2022</year>
          .
          <volume>9747186</volume>
          . [21]
          <string-name>
            <given-names>H.</given-names>
            <surname>Zheng</surname>
          </string-name>
          , W. Cai,
          <string-name>
            <given-names>T.</given-names>
            <surname>Zhou</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Zhang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Li</surname>
          </string-name>
          , Text- [32]
          <string-name>
            <given-names>Z.</given-names>
            <surname>Shang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Shi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Zhang</surname>
          </string-name>
          , L. Wang,
          <string-name>
            <surname>G</surname>
          </string-name>
          . Zhao, Hiertts:
        </mixed-citation>
      </ref>
      <ref id="ref49">
        <mixed-citation>
          <article-title>network based phonetic level features, in: 2016 23rd multi-scale hierarchical variational auto-encoder,</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref50">
        <mixed-citation>
          <source>International Conference on Pattern Recognition Applied Sciences 13.2</source>
          (
          <year>2023</year>
          ):
          <volume>868</volume>
          (
          <year>2023</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref51">
        <mixed-citation>
          <source>(ICPR)</source>
          ,
          <year>2016</year>
          , pp.
          <fpage>2872</fpage>
          -
          <lpage>2877</lpage>
          . doi:
          <volume>10</volume>
          .1109/ICPR. [33]
          <string-name>
            <given-names>I. J.</given-names>
            <surname>Goodfellow</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Shlens</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Szegedy</surname>
          </string-name>
          , Explain-
        </mixed-citation>
      </ref>
      <ref id="ref52">
        <mixed-citation>
          <year>2016</year>
          .7900072. ing and harnessing adversarial examples,
          <year>2015</year>
          . [22]
          <string-name>
            <given-names>L.</given-names>
            <surname>Sun</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Kang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Meng</surname>
          </string-name>
          , Phonetic arXiv:
          <volume>1412</volume>
          .
          <fpage>6572</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref53">
        <mixed-citation>
          <article-title>posteriorgrams for many-to-one voice conversion [34]</article-title>
          <string-name>
            <given-names>A.</given-names>
            <surname>Madry</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Makelov</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Schmidt</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Tsipras</surname>
          </string-name>
          ,
        </mixed-citation>
      </ref>
      <ref id="ref54">
        <mixed-citation>
          <source>without parallel data training</source>
          ,
          <year>2016</year>
          , pp.
          <fpage>1</fpage>
          -
          <lpage>6</lpage>
          . doi:10.
          <string-name>
            <given-names>A.</given-names>
            <surname>Vladu</surname>
          </string-name>
          ,
          <article-title>Towards deep learning models resistant</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref55">
        <mixed-citation>
          1109/ICME.
          <year>2016</year>
          .
          <volume>7552917</volume>
          . to adversarial attacks,
          <year>2019</year>
          . arXiv:
          <volume>1706</volume>
          .
          <fpage>06083</fpage>
          . [23]
          <string-name>
            <surname>J. weon Jung</surname>
            ,
            <given-names>H.</given-names>
            -S. Heo, H.
          </string-name>
          <string-name>
            <surname>Tak</surname>
            , H. jin Shim, [35]
            <given-names>N. M.</given-names>
          </string-name>
          <string-name>
            <surname>Müller</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          <string-name>
            <surname>Dieckmann</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          <string-name>
            <surname>Czempin</surname>
          </string-name>
          , R. Canals,
        </mixed-citation>
      </ref>
      <ref id="ref56">
        <mixed-citation>
          <source>learn?</source>
          ,
          <source>arXiv preprint arXiv:2106.12914</source>
          (
          <year>2021</year>
          ). [36]
          <string-name>
            <given-names>Y.</given-names>
            <surname>Zhang</surname>
          </string-name>
          , W. Wang,
          <string-name>
            <given-names>P.</given-names>
            <surname>Zhang</surname>
          </string-name>
          , The efect of
        </mixed-citation>
      </ref>
      <ref id="ref57">
        <mixed-citation>
          <string-name>
            <surname>system</surname>
          </string-name>
          ,
          <year>2021</year>
          , pp.
          <fpage>4279</fpage>
          -
          <lpage>4283</lpage>
          . doi:
          <volume>10</volume>
          .21437/
        </mixed-citation>
      </ref>
      <ref id="ref58">
        <mixed-citation>
          <string-name>
            <surname>Interspeech.</surname>
          </string-name>
          <year>2021</year>
          -
          <volume>1281</volume>
          . [37]
          <string-name>
            <given-names>H.</given-names>
            <surname>Miao</surname>
          </string-name>
          , G. Cheng, C. Gao,
          <string-name>
            <given-names>P.</given-names>
            <surname>Zhang</surname>
          </string-name>
          , Y. Yan,
        </mixed-citation>
      </ref>
      <ref id="ref59">
        <mixed-citation>
          <article-title>to-end speech recognition architecture</article-title>
          ,
          <year>2020</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref60">
        <mixed-citation>
          arXiv:
          <year>2001</year>
          .
          <volume>08290</volume>
          . [38]
          <string-name>
            <surname>A.</surname>
          </string-name>
          v. d. Oord,
          <string-name>
            <given-names>S.</given-names>
            <surname>Dieleman</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Zen</surname>
          </string-name>
          , K. Simonyan,
        </mixed-citation>
      </ref>
      <ref id="ref61">
        <mixed-citation>
          <article-title>raw audio</article-title>
          ,
          <source>arXiv preprint arXiv:1609.03499</source>
          (
          <year>2016</year>
          ). [39]
          <string-name>
            <given-names>J.</given-names>
            <surname>Kong</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Kim</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Bae</surname>
          </string-name>
          , Hifi-gan: Generative ad-
        </mixed-citation>
      </ref>
      <ref id="ref62">
        <mixed-citation>
          <source>Processing Systems</source>
          <volume>33</volume>
          (
          <year>2020</year>
          ). [40]
          <string-name>
            <given-names>J.</given-names>
            <surname>Pons</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Pascual</surname>
          </string-name>
          , G. Cengarle,
          <string-name>
            <given-names>J.</given-names>
            <surname>Serrà</surname>
          </string-name>
          , Upsampling
        </mixed-citation>
      </ref>
      <ref id="ref63">
        <mixed-citation>
          <article-title>artifacts in neural audio synthesis</article-title>
          ,
          <source>in: ICASSP</source>
          <year>2021</year>
          -
        </mixed-citation>
      </ref>
      <ref id="ref64">
        <mixed-citation>2021 IEEE International Conference on Acoustics,</mixed-citation>
      </ref>
      <ref id="ref65">
        <mixed-citation>
          <string-name>
            <surname>Speech and Signal Processing</surname>
          </string-name>
          (ICASSP), IEEE,
          <year>2021</year>
          ,
        </mixed-citation>
      </ref>
      <ref id="ref66">
        <mixed-citation>
          pp.
          <fpage>3005</fpage>
          -
          <lpage>3009</lpage>
          . [41]
          <string-name>
            <given-names>K.</given-names>
            <surname>Kumar</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Kumar</surname>
          </string-name>
          , T. de Boissiere, L. Gestin, W. Z.
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>