<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>The USTC-NERCSLIP System for the Track 1.2 of Audio Deepfake Detection (ADD 2023) Challenge⋆</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Haochen Wu</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Zhuhai Li</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Luzhen Xu</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Zhentao Zhang</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Wenting Zhao</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Bin Gu</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Yang Ai</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Yexin Lu</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Jie Zhang</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Zhenhua Ling</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Wu Guo</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>China Merchants Bank</institution>
          ,
          <addr-line>Shenzhen, 518057</addr-line>
          ,
          <country country="CN">China</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>University of Science and Technology of China</institution>
          ,
          <addr-line>Hefei, 230027</addr-line>
          ,
          <country country="CN">China</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2023</year>
      </pub-date>
      <fpage>119</fpage>
      <lpage>124</lpage>
      <abstract>
        <p>This paper describes the system of USTC-NERCSLIP submitted to the track 1.2 of the second Audio Deepfake Detection Challenge (ADD 2023). Our system consists of a wav2vec2.0-based front-end feature extractor and an AASIST-based back-end classifier. To further solve the problem of the gap in the noise and synthesis algorithms between the training and evaluation sets, we propose a multi-level data augmentation method. Specifically, we add a variety of noises to the training set to simulate the noise environment of the evaluation set. Besides, we use several vocoders to synthesize fake audio based on the genuine audio in the training set to enrich the synthesis algorithms. Results show that the proposed method achieves a WEER of 12.45% on the two-round evaluation with a single system, which ranks the top among all submissions.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;ADD 2023</kwd>
        <kwd>data augmentation</kwd>
        <kwd>wav2vec 2</kwd>
        <kwd>0</kwd>
        <kwd>speech synthesis</kwd>
        <kwd>vocoder</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>conversion (VC) [9], speech replay and impersonation.</p>
      <p>To further address diversified and challenging attack
sitOver the last decades, the development of artificial intel- uations in realistic applications, the first Audio Deepfake
ligence (AI) has in turn contributed to rapid advances in Detection Challenge (ADD 2022) [10] extends the attack
speech synthesis and voice conversion applications. Deep situations of fake audio detection. Diferent from ADD
learning models can generate realistic and human-like 2022, the second ADD Challenge (ADD 2023) [11]
fospeech, which has a wide range of applications in human- cuses on surpassing the constraints of binary real/fake
computer interaction, smart home, entertainment, educa- classification, localizing the manipulated intervals in a
tion, etc [1]. Nevertheless, they also bring a potential to partially fake speech as well as pinpointing the source
pose a serious threat to the society if someone misuses algorithm used to generate the fake audio. The ADD
it, e.g., using fake audio to commit fraud or mislead pub- 2023 Challenge contains three tracks, among which the
lic opinions. Therefore, in order to improve the speech Track 1 is an audio fake game (FG) consisting of an audio
security, detecting fake audio is essential to reduce the generation task (Track 1.1) and a fake audio detection
threat posed by the disinformation embedded in speech. task (Track 1.2). For Track 1.1, participants aim to
genAlso, deepfake audio detection can help to address the erate fake audio that can spoof the detection systems
serious vulnerability of automated speaker verification of Track 1.2. For Track 1.2, participants aim to detect
systems against various malicious spoofing attacks [2]. fake utterances, especially the fake samples generated</p>
      <p>
        Guarding against such abuse and misuse, deepfake au- from Track 1.1. The two tracks represent a more realistic
dio detection is therefore an interesting emerging topic in situation of what anti-spoofing researchers need to deal
the AI community. The ASVspoof Challenge [3, 4, 5, 6] is with day to day.
held every two years dedicated to spoofing speech detec- Recently, self-supervised learning (SSL) has achieved
tion, including text to speech synthesis (TTS) [7, 8], voice significant advances in the fields of natural language
proIJCAI 2023 Workshop on Deepfake Audio Detection and Analysis cessing (NLP) [12], automatic speech recognition [13] as
        <xref ref-type="bibr" rid="ref25">(DADA 2023)</xref>
        , August 19, 2023, Macao, S.A.R well as speaker verification [ 14]. It was shown that
build⋆ This work was supported by the USTC-CMB Joint Laboratory of ing a general pre-trained model based on the exploitation
Artificial Intelligence, the National Natural Science Foundation of a large mount of unlabeled data can be quite essential
of China (62101523), Hefei Municipal Natural Science Foundation to boost the performance of many downstream tasks,
(2022012) and USTC Research Funds of the Double First-Class reduce data labeling eforts and lower entry barriers for
* CInoirtiraetsipvoen(YdiDn2g1a0u0t0h0o2r0.08). individual tasks. To our knowledge, only a few works
† These authors contributed equally. have used self-supervised pre-trained models for fake
$ whc1414858026@mail.ustc.edu.cn (H. Wu); audio detection. In this work, we thus make eforts to
snowsea@mail.ustc.edu.cn (Z. Li); jzhang6@ustc.edu.cn (J. Zhang) use an open-sourced self-supervised pre-trained model as
© 2023 Copyright for this paper by its authors. Use permitted under Creative Commons License the feature extractor to help build a robust ADD system.
CPWrEooUrckReshdoinpgs IhStpN:/c1e6u1r3-w-0s.o7r3g ACttEribUutRion W4.0oInrtekrnsahtioonpal (PCCroBYce4.0e).dings (CEUR-WS.org)
      </p>
      <p>This paper presents our submitted system to the Track from 10 dB to 30 dB.
1.2 of the ADD 2023 Challenge. Due to the diversified For the reverberation, we add distortion to the clean
noise and rich synthesis algorithms in the evaluation set, training set using the public room impulse response (RIR)
it is dificult to obtain the desired performance using the [19]. It should be pointed out that the reverberation
training set directly. We thus propose a multi-level data and noise are added to the audio sequentially. First, we
augmentation (DA) method to address this problem. The add reverberation at a certain probability , and then
ifrst-level DA aims to develop a robust system to noise, noise is added at a certain probability . Besides, unlike
reverberation and channel variation. We use the pub- traditional DA, which enlarge the training set, we add
lic noise, reverberation datasets and RawBoost DA tool nuisance variability to the existing training data online.
[15] to conduct online DA. On the other hand, we adopt For the channel variation, we apply a variety of codec
the speed perturbation [16] and compression coding [17] algorithms [20], including MP3, OGG, AAC OPUS, a-law
methods for ofline DA. We find that these techniques and µ-law. Besides, to mock the telephony transmission
have a little contribution to fix the gap in the number of loss, audio samples are first downsampled to 8kHz and
synthesis algorithms between the training and evalua- then upsampled back to 16kHz. To add more spices, we
tion sets. Therefore, the second level of our DA method also consider speed perturbation to help improve the
aims to increase the variety of the synthesis algorithms performance. Considering the high computational costs
in the training set. Specifically, we use several vocoders of compression coding and speed perturbation, we use
to synthesize fake audio based on the genuine audio in them in an ofline manner. The amount of the resulting
the training set. Experimental results show that the pro- training set is increased by three times compared to the
posed method outperforms all other submissions with a original training data.
weighted equal error rate (WEER) of 12.45% in Track 1.2.</p>
      <p>The remainder of this paper is organized as follows. 2.2. DA for Synthesis Algorithms
Section 2 and 3 describe the proposed multi-level DA
method and model architecture in detail, respectively.</p>
      <p>Section 4 introduces the experimental setup, followed
by experimental results in Section 5. Finally, Section 6
concludes this work.</p>
    </sec>
    <sec id="sec-2">
      <title>2. Data Augmentation</title>
      <sec id="sec-2-1">
        <title>In this section, we will show a detailed description of the proposed multi-level DA method from two aspects, i.e., for diverse noise and synthesized speech, respectively.</title>
        <sec id="sec-2-1-1">
          <title>2.1. DA for Diversified Noise</title>
        </sec>
      </sec>
      <sec id="sec-2-2">
        <title>To reduce over-fitting and bias caused by diversified noise</title>
        <p>in real scenes, we apply augmentation methods from
three aspects: noise, reverberation and channel variation.</p>
        <p>For the noise, on one hand, we add recorded noises
and distortion from MUSAN [18] dataset to the clean
audio, which is commonly used in other speech-related
ifelds. On the other hand, by utilizing the RawBoost DA
technique, we add diferent nuisance noises dependent
on the corresponding raw waveform inputs. Based on a
variety of convolutional and additive noises, RawBoost
models nuisance variability stemming such as encoding,
transmission, microphones and amplifiers, as well as
linear and nonlinear distortions. RawBoost consists of three
independent noises [15]: 1) linear and non-linear
convolutive noise; 2) impulsive signal-dependent additive
noise; 3) stationary signal-independent additive noise.</p>
        <p>The details are available in [15]. Finally, the noises are
mixed at a random signal-to-noise ratio (SNR) ranging
Most DA methods focus on improving the
generalization of the system in real scenes, such as the methods
in Section 2.1, which, however, cannot cover the speech
synthesis algorithms in the training set. To tackle this
issue, we use several vocoders to synthesize fake audio,
including traditional and neural vocoders. The traditional
vocoder can directly synthesize fake audio after
determining the parameters of the algorithm, while the neural
vocoder can only synthesize audio with the generator
trained on the training set. In this work, we choose three
traditional vocoders (TV) and one neural vocoder (NV).</p>
        <p>The details are given as follows.</p>
        <p>• Grifin-Lim [21]: This is a traditional vocoder,
which synthesizes audio using the true
melspectrum. It uses a phase reconstruction method
based on the redundancy of the short-time
Fourier transform and promotes the consistency
of a spectrogram. The mel-spectrum is first used
to estimate the amplitude spectrum, which is then
used to estimate the audio waveform.
• WORLD [22]: World shows a superiority in not
only the sound quality, but also the complexity
(thus being appropriate for real-time cases). It
includes three parameter estimation modules (for
estimating F0, spectral envelop, aperiodic
parameter extraction, respectively), followed by a
synthesis module to generate the speech-like signals.
• RAAR [23]: RAAR is extensively used in optics,
and its efectiveness has been validated in many
applications. In [23], Tomoki et al. applied a
phase reconstruction algorithm to acoustic
scenarios, where the evaluated acoustical metrics</p>
        <p>show that RAAR is robust against noise and per- and XLS-R follows the same model architecture, the
feaforms well for both small and large number of tures extracted from XLS-R can be more general due to
iterations. the more fruitful data resource, leading to a stronger
fea• HiFi-GAN [8]: As one of the state-of-the-art neu- ture reliability and domain robustness. Therefore, we use
ral vocoders, HiFi-GAN generates audio based on XLS-R as the feature extractor for the ADD task.
generative adversarial networks (GANs) using The XLS-R model mainly includes three stages. Firstly,
the true mel-spectrum. It includes one generator the raw waveform is sent into a feature encoder
comand two discriminators: multi-scale and multi- posed of several convolutional layers (CNN). The feature
period discriminators, which can achieve eficient encoder extracts vector representations of size 1024
evand high-fidelity speech synthesis. The generator ery 20ms and the receptive field is 25ms. Secondly, these
and discriminators are trained adversarially by encoder embeddings are fed into the context encoder,
incorporating two additional losses to improve which contains 24 transformer block layers and is used
the training stability and model performance. to explore the contextual information contained in the
input speech. At the third stage, the feature encoder
representation is processed by a quantization module to
obtain a quantized representation. Then, the model is
trained in a self-supervised manner with a contrastive
loss by using the contextual representations to predict
the masked counterparts at certain positions.</p>
        <sec id="sec-2-2-1">
          <title>3.2. AASIST Back-End</title>
        </sec>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>3. Methodology</title>
      <sec id="sec-3-1">
        <title>The diagram of our system for the ADD 2023 Challenge</title>
        <p>is illustrated in Figure 1, which follows a fully automated
end-to-end pipeline consisting of two modules: feature
extractor and classifier. The wav2vec2 based feature
extractor aims to extract a high-level representation of the
speech. The classification module is a modified version of
AASIST [24] with the removal of the sinc convolutional
layer based front-end.</p>
      </sec>
      <sec id="sec-3-2">
        <title>AASIST is an end-to-end Audio Anti-Spoofing system</title>
        <p>using Integrated Spectro-Temporal graph attention
networks, which won the top rank in ASVspoof 2019 logical
access (LA). It is an extension of RawGAT-ST [27] with
three modifications: 1) a novel heterogeneous stacking
3.1. Wav2vec2 Front-End graph attention layer, which models artefacts spanning
Wav2vec2 [25] is a self-supervised pre-trained model, heterogeneous temporal and spectral domains with a
which can extract speech representations or embeddings heterogeneous attention mechanism and a stack node,
from raw waveform. It has shown an impressive per- 2) a max graph operation that involves a competitive
formance on many downstream tasks, particularly on selection of artefacts, and 3) a modified readout scheme.
automatic speaker recognition. As a variant, XLS-R [26] AASIST uses a sinc convolutional layer based front-end,
is a new self-supervised cross-lingual speech representa- and thus can extract representations directly from raw
tion model based on wav2vec2, which scales the number waveform inputs.
of languages, the amount of training data as well as the In [28] Tak et al. tried to improve the
generalizamodel size. XLS-R is pre-trained on 436K hours of unan- tion and domain robustness using a pre-trained,
selfnotated speech in 128 languages. Although wav2vec2 supervised model with fine-tuning. Specifically, the sinc
convolution layer of AASIST is replaced by the
aforevocoder HiFi-GAN, we only use all of the 5319 genuine
audios for training, without using any other external
data. We directly use the trained model to synthesize
fake audio based on the genuine audio used for training.</p>
        <p>Note that HiFi-GAN takes the mel-spectrum as input.</p>
        <p>Each vocoder can synthesize 5319 fake audios from all
the genuine audios. The resulting 21276 fake audios are
incorporated altogether in the training set.</p>
        <sec id="sec-3-2-1">
          <title>4.3. Implementation Details</title>
          <p>mentioned wav2vec2 model. Besides, a fully connected
layer after the pre-trained model is used to reduce the
representation dimension from 1024 to 128. In this work,
we adopt the same model architecture as in [28].</p>
        </sec>
      </sec>
      <sec id="sec-3-3">
        <title>In experiments, the audio streams are truncated or re</title>
        <p>peated to a duration of 6 seconds during the train stage.</p>
        <p>The probabilities  and  in Section 2.1 are set to 13
and 15 , respectively. During fine-tuning, the pre-trained
4. Experimental Setup wav2vec2 model is optimized jointly with the AASIST
via the back-propagation. We use the standard Adam
This section introduces the dataset used in experiments, optimizer [29], which adopts a mini-batch size of 16 and
parameter configurations for our multi-level DA method a learning rate of 10−5 with a weight decay of 10−4 to
and implementation details of our system. avoid over-fitting. Since the result after two epochs of
training is always worse than that obtained from only
one epoch, all models are fine-tuned for only one epoch
4.1. Dataset on two RTX 3090 GPUs. Considering the imbalance
beThe Track 1.2 of ADD 2023 aims to distinguish fake au- tween the genuine and fake audios in the training set, we
dios from genuine ones, which are fully fake utterances use the weighted cross entropy to minimize the training
generated by text-to-speech or voice-conversion algo- loss. The weights associated with the genuine and fake
rithms. The training set consists of 3012 genuine utter- categories are set to 0.9 and 0.1, respectively.
ances and 24072 fake utterances. The development set
consists of 2307 genuine utterances and 26017 fake utter- 5. Results
ances. Besides, there are 111976 and 118477 utterances in
the first and second round evaluation sets, respectively, In this section, we show the performance of the proposed
where the second contains noise and fake audio gener- multi-level DA method based ADD system. The results
ated by the teams participated in the Track 1.1 using depend on the equal error rate (EER) and the final score
unknown synthesis algorithms. is weighted equal error rate (WEER), which is defined as</p>
        <p>We find that the audio in the training and
development sets are quite clean and have identical data distri- WEER =  * EERR1 +  * EERR2 (1)
bution, making the trained model performs well on the
development set. To be more general, we thus choose to
re-partition the training and development sets. First, we
combine the training set and development set as a larger
dataset. Then, we randomly select 50% of the fake audio
and all the genuine audio from the dataset and combine
them into the new training set. Furthermore, we apply
our DA method on the new training set to enable the
ifnal system with a high robustness and performance.
where  = 0.4,  = 0.6, EERR1 and EERR2 denotes
the EERs obtained in the two rounds of Track 1.2.</p>
        <p>Model Comparison: As our system adopts the
wav2vec 2.0 front-end and AASIST back-end, in order to
show the respective function we conduct several
comparisons. First, we compare the AASIST with the Light
Convolutional Neural Networks (LCNN), which adopts
the same architecture as [30] in the ADD 2022 Track 3.2.</p>
        <p>Besides, the LCNN takes the STFT as input features
in4.2. Vocoders stead of the raw waveform. The results for the AASIST
with the sinc-layer front-end or the wav2vec 2.0
frontThe three traditional vocoders synthesize fake audio end are presented in Table 2. For the training set, DA
from diferent input features. The mel-spectrum and for diversified noise and the three traditional vocoders
Fourier amplitude spectrum are used for the Grifin-Lim are used. Besides, we only show the results on the first
and RAAR, respectively, while fundamental frequency, round of the evaluation set.
spectral envelope and aperiodic parameter estimated by Since the LCNN and AASIST with the sinc-layer
frontWORLD are used to synthesize audio. More details about end do not use the pre-trained model, we train them
the vocoders are summarized in Table 1. For the neural diferently. Specifically, the Adam optimizer is adopted
with  1 = 0.9,  2 = 0.98 and weight decay 10−4 . The
batch size is set to 32. The learning rate is initially set to
0.0003 with 50% decay for every 10 epochs. We train the
network for 100 epochs and select the model with the
lowest EER on the development set as the final model for
evaluation.</p>
        <p>From Table 2, we can see that AASIST with the
sinclayer front-end outperforms the LCNN. However, the
EER is still unacceptably high with a poor robustness.</p>
        <p>Replacing the sinc-layer with the wav2vec 2.0 front-end
can reduce the EER by 7.26%, which validates the benefit
of using pre-trained model for the ADD task.</p>
        <p>Comparison of Vocoders: Then, we conduct ablation
studies on the proposed vocoder-based DA method.
According to the types of vocoders, we combine the audios
generated by the three traditional vocoders (TV) into the
TV set, while the audios generated by the neural vocoder
(NV) are regarded as the NV set. We train our system
on diferent datasets by combining the training set with
the TV set or NV set. Besides, the noise-based DA is
used, but we only show the results on the first round of
the evaluation set in Table 3. It is clear that the EER on
the evaluation set is very high (e.g., 40.53%) even using
the DA for diversified noise and the wav2vec 2.0
frontend. Applying the traditional vocoders, the EER drops to
25.45%, which can be further reduced to 11.56% in case
of training on both the TV and NV sets. This reveals
that the vocoders in combination with DA techniques
are rather helpful to improve the ADD performance.</p>
        <p>Comparison of MUSAN and RawBoost: Apart from
synthesis algorithms, the diferences in the background
noise, reverberation and channel variety between the
training and evaluation sets also play an important role.</p>
        <p>However, due to the tight challenge schedule, we ignore
the efect of the RIR, speed perturbation and compression
coding. Here, we compare the influence of the MUSAN
and RawBoost used for the online DA in Table 4, where
note that vocoders are incorporated. We can see that
using online noises leads to a significant EER decrease
from 19.15% to 11.56%. It also shows that MUSAN is
more beneficial than RawBoost to increase the diversity
of noise in the training set.</p>
        <p>Finally, it should be noticed that the performance
slightly drops on the Round 2 evaluation (from 11.56% to
13.05%). This is probability due to the fact that the fake
audio synthesized by participants in Track 1.1 are added
into the evaluation set, which further increases the data
diversity. More importantly, our system still shows its
superiority and ranks the top in Round 2. In Figure 2, we
summarize the overall WEER of the top 10 participants,
where the teams are anonymized accordingly.</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>6. Conclusions</title>
      <sec id="sec-4-1">
        <title>This paper presents the detailed system description of</title>
        <p>USTC-NERCSLIP submitted to the ADD Challenge 2023,
which involves the wav2vec2-based feature extractor and
the AASIST-based classifier. In addition, we proposed a
multi-level DA method for the diversified noise and
synthesis algorithms in the evaluation set, which was shown
to largely improve the performance and robustness. Due
to the new data source in the second-round evaluation,
the performance slightly drops, but our system still ranks
the first place in Track 1.2.</p>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [16]
          <string-name>
            <given-names>T.</given-names>
            <surname>Ko</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Peddinti</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Povey</surname>
          </string-name>
          , et al.,
          <string-name>
            <surname>Audio</surname>
          </string-name>
          augmenta-
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          <article-title>tion for speech recognition</article-title>
          ,
          <source>in: Proc. INTERSPEECH</source>
          [1]
          <string-name>
            <given-names>N. M.</given-names>
            <surname>Müller</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Pizzi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Williams</surname>
          </string-name>
          , Human per- 2015,
          <year>September 2015</year>
          , p.
          <fpage>3586</fpage>
          -
          <lpage>3589</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          <article-title>ception of audio deepfakes</article-title>
          ,
          <source>in: Proc. DDAM</source>
          <year>2022</year>
          , [17]
          <string-name>
            <surname>T.-L. Vu</surname>
            ,
            <given-names>Z.</given-names>
          </string-name>
          <string-name>
            <surname>Zeng</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          <string-name>
            <surname>Xu</surname>
          </string-name>
          , et al.,
          <source>Audio codec simula-</source>
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          <source>October</source>
          <year>2022</year>
          , pp.
          <fpage>85</fpage>
          -
          <lpage>91</lpage>
          .
          <article-title>tion based data augmentation for telephony speech</article-title>
          [2]
          <string-name>
            <given-names>C.</given-names>
            <surname>Zhang</surname>
          </string-name>
          , S. Ranjan,
          <string-name>
            <given-names>J. H.</given-names>
            <surname>Hansen</surname>
          </string-name>
          ,
          <article-title>An analysis recognition</article-title>
          ,
          <source>in: Proc. APSIPA ASC</source>
          ,
          <year>September 2019</year>
          ,
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          <article-title>of transfer learning for domain mismatched text</article-title>
          - pp.
          <fpage>198</fpage>
          -
          <lpage>203</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          <article-title>independent speaker verification</article-title>
          ,
          <source>in: Proc. Odyssey</source>
          [18]
          <string-name>
            <given-names>D.</given-names>
            <surname>Snyder</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Chen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Povey</surname>
          </string-name>
          , MUSAN: A mu-
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          <string-name>
            <surname>Workshop</surname>
          </string-name>
          , June 2018, pp.
          <fpage>182</fpage>
          -
          <lpage>186</lpage>
          . sic, speech, and
          <article-title>noise corpus</article-title>
          , in: arXiv preprint [3]
          <string-name>
            <given-names>Z.</given-names>
            <surname>Wu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Kinnunen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Evans</surname>
          </string-name>
          , et al.,
          <source>ASVspoof arXiv:1510.08484</source>
          ,
          <year>2015</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          <year>2015</year>
          :
          <article-title>the first automatic speaker verification spoof-</article-title>
          [19]
          <string-name>
            <given-names>T.</given-names>
            <surname>Ko</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Peddinti</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Povey</surname>
          </string-name>
          , et al.,
          <source>A study on</source>
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          <source>TERSPEECH</source>
          <year>2015</year>
          ,
          <year>September 2015</year>
          , pp.
          <fpage>2037</fpage>
          -
          <lpage>2041</lpage>
          .
          <article-title>speech recognition</article-title>
          ,
          <source>in: Proc. ICASSP</source>
          ,
          <year>March 2017</year>
          , [4]
          <string-name>
            <given-names>T.</given-names>
            <surname>Kinnunen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Sahidullah</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Delgado</surname>
          </string-name>
          , et al., The pp.
          <fpage>5220</fpage>
          -
          <lpage>5224</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          <article-title>ASVspoof 2017 Challenge: assessing the limits</article-title>
          of [20]
          <string-name>
            <surname>R. K. Das</surname>
          </string-name>
          ,
          <article-title>Known-unknown data augmentation</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          <source>SPEECH</source>
          <year>2017</year>
          ,
          <year>August 2017</year>
          , pp.
          <fpage>2</fpage>
          -
          <lpage>6</lpage>
          . access and speech deepfake attacks:
          <source>Asvspoof</source>
          <year>2021</year>
          , [5]
          <string-name>
            <given-names>M.</given-names>
            <surname>Todisco</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Vestman</surname>
          </string-name>
          , et al.,
          <source>ASVspoof Proc. ASVspoof2021 Workshop</source>
          (
          <year>2021</year>
          )
          <fpage>29</fpage>
          -
          <lpage>36</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          2019:
          <article-title>Future horizons in spoofed</article-title>
          and fake audio [21]
          <string-name>
            <given-names>N.</given-names>
            <surname>Perraudin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Balazs</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P. L.</given-names>
            <surname>Søndergaard</surname>
          </string-name>
          , A fast
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          detection, in
          <source>: Proc. INTERSPEECH</source>
          <year>2019</year>
          ,
          <article-title>September grifin-lim algorithm</article-title>
          ,
          <source>in: Proc. WASPAA</source>
          , October
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          <year>2019</year>
          , pp.
          <fpage>1008</fpage>
          -
          <lpage>1012</lpage>
          .
          <year>2013</year>
          , pp.
          <fpage>1</fpage>
          -
          <lpage>4</lpage>
          . [6]
          <string-name>
            <given-names>J.</given-names>
            <surname>Yamagishi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Todisco</surname>
          </string-name>
          , et al.,
          <source>ASVspoof</source>
          [22]
          <string-name>
            <given-names>M.</given-names>
            <surname>Morise</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Yokomori</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Ozawa</surname>
          </string-name>
          , World: a
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          2021:
          <article-title>Accelerating progress in spoofed and deep- vocoder-based high-quality speech synthesis sys-</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          <article-title>fake speech detection</article-title>
          ,
          <source>in: Proc. ASVspoof2021</source>
          Work
          <article-title>- tem for real-time applications</article-title>
          , IEICE Transactions
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          <string-name>
            <surname>shop</surname>
          </string-name>
          ,
          <source>2021. on Information and Systems</source>
          <volume>99</volume>
          (
          <year>2016</year>
          )
          <fpage>1877</fpage>
          -
          <lpage>1884</lpage>
          . [7]
          <string-name>
            <surname>A.</surname>
          </string-name>
          v. d. Oord,
          <string-name>
            <given-names>S.</given-names>
            <surname>Dieleman</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Zen</surname>
          </string-name>
          , et al.,
          <source>Wavenet</source>
          <volume>:</volume>
          [23]
          <string-name>
            <given-names>T.</given-names>
            <surname>Kobayashi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Tanaka</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Yatabe</surname>
          </string-name>
          , et al.,
          <source>Acoustic</source>
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          2016. optics,
          <source>in: Proc. ICASSP</source>
          , May
          <year>2022</year>
          , pp.
          <fpage>6212</fpage>
          -
          <lpage>6216</lpage>
          . [8]
          <string-name>
            <given-names>J.</given-names>
            <surname>Kong</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Kim</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Bae</surname>
          </string-name>
          , Hifi-gan: Generative ad- [24]
          <string-name>
            <surname>J</surname>
            .-w. Jung,
            <given-names>H.</given-names>
            -S. Heo, H.
          </string-name>
          <string-name>
            <surname>Tak</surname>
            , et al.,
            <given-names>AASIST</given-names>
          </string-name>
          : Au-
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          <source>Processing Systems</source>
          <volume>33</volume>
          (
          <year>2020</year>
          )
          <fpage>17022</fpage>
          -
          <lpage>17033</lpage>
          .
          <year>2022</year>
          , pp.
          <fpage>6367</fpage>
          -
          <lpage>6371</lpage>
          . [9]
          <string-name>
            <given-names>T.</given-names>
            <surname>Kaneko</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Kameoka</surname>
          </string-name>
          , Cyclegan-vc:
          <article-title>Non-parallel [25]</article-title>
          <string-name>
            <given-names>A.</given-names>
            <surname>Baevski</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Zhou</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Mohamed</surname>
          </string-name>
          , et al.,
          <source>Wav2vec</source>
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          <article-title>voice conversion using cycle-consistent adversarial 2.0: A framework for self-supervised learning of</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          networks,
          <source>in: Proc. EUSIPCO</source>
          ,
          <year>2018</year>
          , pp.
          <fpage>2100</fpage>
          -
          <lpage>2104</lpage>
          . speech representations,
          <source>Advances in neural infor</source>
          [10]
          <string-name>
            <given-names>J.</given-names>
            <surname>Yi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Fu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Tao</surname>
          </string-name>
          , et al.,
          <source>ADD</source>
          <year>2022</year>
          :
          <article-title>the first audio mation processing systems 33 (</article-title>
          <year>2020</year>
          )
          <fpage>12449</fpage>
          -
          <lpage>12460</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          <article-title>deep synthesis detection challenge</article-title>
          ,
          <source>in: Proc. ICASSP</source>
          , [26]
          <string-name>
            <given-names>A.</given-names>
            <surname>Babu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Tjandra</surname>
          </string-name>
          , et al.,
          <string-name>
            <surname>XLS-R:</surname>
          </string-name>
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          <source>May</source>
          <year>2022</year>
          , pp.
          <fpage>9216</fpage>
          -
          <lpage>9220</lpage>
          .
          <article-title>Self-supervised cross-lingual speech representation</article-title>
          [11]
          <string-name>
            <given-names>J.</given-names>
            <surname>Yi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Tao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Fu</surname>
          </string-name>
          , et al.,
          <source>ADD</source>
          <year>2023</year>
          :
          <article-title>the second learning at scale</article-title>
          ,
          <source>in: Proc. INTERSPEECH</source>
          <year>2022</year>
          ,
          <year>2022</year>
          ,
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          <article-title>audio deepfake detection challenge</article-title>
          ,
          <source>in: Proc. IJCAI</source>
          pp.
          <fpage>2278</fpage>
          -
          <lpage>2282</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref25">
        <mixed-citation>
          <source>Workshop on DADA</source>
          <year>2023</year>
          ,
          <year>August 2023</year>
          . [27]
          <string-name>
            <given-names>H.</given-names>
            <surname>Tak</surname>
          </string-name>
          , J.-w. Jung,
          <string-name>
            <given-names>J.</given-names>
            <surname>Patino</surname>
          </string-name>
          , et al.,
          <string-name>
            <surname>End-</surname>
            to-end [12]
            <given-names>T.</given-names>
          </string-name>
          <string-name>
            <surname>Brown</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          <string-name>
            <surname>Mann</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          <string-name>
            <surname>Ryder</surname>
          </string-name>
          , et al.,
          <article-title>Language mod- spectro-temporal graph attention networks for</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref26">
        <mixed-citation>
          <source>mation processing systems</source>
          <volume>33</volume>
          (
          <year>2020</year>
          )
          <fpage>1877</fpage>
          -
          <lpage>1901</lpage>
          .
          <article-title>fake detection</article-title>
          ,
          <source>in: Proc. ASVspoof Workshop</source>
          ,
          <year>2021</year>
          . [13]
          <string-name>
            <given-names>C.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Wu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Qian</surname>
          </string-name>
          , et al.,
          <source>Unispeech: Unified</source>
          [28]
          <string-name>
            <given-names>H.</given-names>
            <surname>Tak</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Todisco</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Wang</surname>
          </string-name>
          , et al.,
          <source>Automatic</source>
        </mixed-citation>
      </ref>
      <ref id="ref27">
        <mixed-citation>
          <article-title>unlabeled data</article-title>
          ,
          <source>in: Proc. ICML</source>
          ,
          <year>2021</year>
          , pp.
          <fpage>10937</fpage>
          -
          <article-title>tion using wav2vec 2.0 and data augmentation</article-title>
          , in:
        </mixed-citation>
      </ref>
      <ref id="ref28">
        <mixed-citation>
          10947.
          <string-name>
            <given-names>The</given-names>
            <surname>Speaker</surname>
          </string-name>
          and Language Recognition Workshop, [14]
          <string-name>
            <given-names>Z.</given-names>
            <surname>Fan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Zhou</surname>
          </string-name>
          , et al.,
          <source>Exploring wav2vec 2.0 June</source>
          <year>2022</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref29">
        <mixed-citation>
          <article-title>on speaker verification and language identification</article-title>
          , [29]
          <string-name>
            <given-names>D. P.</given-names>
            <surname>Kingma</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. L.</given-names>
            <surname>Ba</surname>
          </string-name>
          ,
          <article-title>Adam: A method for stochas-</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref30">
        <mixed-citation>
          <source>in: Proc. INTERSPEECH</source>
          ,
          <year>2021</year>
          , pp.
          <fpage>1509</fpage>
          -
          <lpage>1513</lpage>
          . tic optimization,
          <source>in: Proc. ICLR</source>
          ,
          <year>2015</year>
          , pp.
          <fpage>1</fpage>
          -
          <lpage>15</lpage>
          . [15]
          <string-name>
            <given-names>H.</given-names>
            <surname>Tak</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Kamble</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Patino</surname>
          </string-name>
          , et al.,
          <string-name>
            <surname>Rawboost</surname>
            : A raw [30]
            <given-names>H.</given-names>
          </string-name>
          <string-name>
            <surname>Tak</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          <string-name>
            <surname>Todisco</surname>
            ,
            <given-names>X.</given-names>
          </string-name>
          <string-name>
            <surname>Wang</surname>
          </string-name>
          , othets, Deepfake
        </mixed-citation>
      </ref>
      <ref id="ref31">
        <mixed-citation>
          <article-title>data boosting and augmentation method applied detection system for the add challenge track 3.2</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref32">
        <mixed-citation>
          <article-title>to automatic speaker verification anti-spoofing, in: based on score fusion</article-title>
          ,
          <source>in: Proc. DDAM</source>
          <year>2022</year>
          , October
        </mixed-citation>
      </ref>
      <ref id="ref33">
        <mixed-citation>
          <string-name>
            <surname>Proc. ICASSP</surname>
          </string-name>
          , May
          <year>2022</year>
          , pp.
          <fpage>6382</fpage>
          -
          <lpage>6386</lpage>
          .
          <year>2022</year>
          , pp.
          <fpage>43</fpage>
          -
          <lpage>52</lpage>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>