<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>The NeteaseGames System for fake audio generation task of 2023 Audio Deepfake Detection Challenge</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Haoyue Zhan</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Yang Zhang</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Xinyuan Yu</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>NetEase Games AI Lab</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Guangzhou</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>China</string-name>
        </contrib>
      </contrib-group>
      <pub-date>
        <year>2023</year>
      </pub-date>
      <abstract>
        <p>This paper presents the description of our speech synthesis system for the Audio Deep Synthesis Detection Challenge (ADD 2023). We utilizes a FastPitch-based model augmented with a BERT-based prosody feature and an utterance embedding predictor to model the generated speech. We incorporate these components to improve one-to-many generation modeling. Evaluation results indicate a significant advantage over some false speech detection models, earning a second-place ranking in the competition overall.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;speech synthesis</kwd>
        <kwd>fake audio</kwd>
        <kwd>post-processing</kwd>
        <kwd>ADD challenge</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>
        The proliferation of deepfake technology has under- Speaker Embedding Decoder
scored the importance of reliable methods to detect and
fparkeeveDnetttehcetimonalCichioaullsenugsee o(AfDdeDep2f0a2k3e)s.isTahedeAeupdlieoaDrneienpg- [
        <xref ref-type="bibr" rid="ref2 ref3">3, 2</xref>
        ] Repeat
tceocmtipoentiatniodnanthaalytsaisimosf dtoeepprfoakmeoatuedrieos[e1a]r.cThhienptrhime adrey- PDruerdaitcitoonr PrPediticchtor PErendeircgtyor
objective of the competition is to accelerate the
development of more robust and reliable deepfake detection tools
for audio and encourage the exploration of cutting-edge UttBerearntcEemEbmedbdeidndging Encoder Variance Adaptors
techniques in this field. Phoneme Num
      </p>
      <p>
        The competition comprises four tracks, each with a
unique focus and set of challenges. In this paper, we LengPthhoRneemguelator
describe the speech synthesis system we employed in the [
        <xref ref-type="bibr" rid="ref2 ref3">2, 3</xref>
        ]
fake audio generation task (Track 1.1). This Track cen- Phoneme Phoneme Length
ters on "Adversarial Attacks," which involves generating Embedding Sequence
speech to deceive a model with false detection capabili- Figure 1: Acoustic Model
ties. Participants are required to train their models using
the provided training dataset and evaluate their
performance on a test dataset. The success rate of generating
speech that can fool the detection model determines the the performance of their models against those of other
results of this task. participants in a fair and transparent setting.
      </p>
      <p>
        The primary challenge of Track 1.1 is the develop- Speech synthesis technology, based on deep learning,
ment of efective adversarial attacks that can generate has the capability to generate counterfeit speech that
speech that is dificult for detection models to distinguish mimics a target speaker’s voice from text and the target
from genuine audio. Deepfake audio generated using ad- speaker’s voice data[
        <xref ref-type="bibr" rid="ref2 ref3 ref4">2, 3, 4</xref>
        ]. Currently, speech synthesis
vanced machine learning techniques can be challenging is primarily achieved through two methods: multi-stage
to detect, even for humans. The competition ofers an synthesis and end-to-end synthesis. Multi-stage
syntheopportunity for participants to benchmark and compare sis can be categorized further into autoregressive models
based on Tacotron[
        <xref ref-type="bibr" rid="ref5">5</xref>
        ] and non-autoregressive models
IJCAI 2023 Workshop on Deepfake Audio Detection and Analysis based on FastSpeech[
        <xref ref-type="bibr" rid="ref6">6</xref>
        ]. For our competition system, we
(DADA 2023), August 19, 2023, Macao, S.A.R have employed the latter, FastPitch-based model as the
* Corresponding author. acoustic model framework[
        <xref ref-type="bibr" rid="ref7">7</xref>
        ].
$ zhanhaoyue@corp.netease.com (H. Zhan); This paper is organized as follows: Section 2 outlines
zyhuaxningyyuaanng0029@@ccoorrpp..nneetteeaassee..ccoomm ((XY.. ZYuh)ang); our data preprocessing process, Section 3 describes our
© 2023 Copyright for this paper by its authors. Use permitted under Creative Commons License model structure, Section 4 provides detail of our
compeCPWrEooUrckReshdoinpgs IhStpN:/c1e6u1r3-w-0s.o7r3g ACttEribUutRion W4.0oInrtekrnsahtioonpal (PCCroBYce4.0e).dings (CEUR-WS.org) tition results, and finally, the conclusion is presented.
      </p>
      <sec id="sec-1-1">
        <title>Encoder</title>
      </sec>
      <sec id="sec-1-2">
        <title>Output</title>
      </sec>
      <sec id="sec-1-3">
        <title>Conformer</title>
      </sec>
      <sec id="sec-1-4">
        <title>Block</title>
        <p>×2</p>
      </sec>
      <sec id="sec-1-5">
        <title>Conformer</title>
      </sec>
      <sec id="sec-1-6">
        <title>Block</title>
      </sec>
      <sec id="sec-1-7">
        <title>Conformer</title>
      </sec>
      <sec id="sec-1-8">
        <title>Block</title>
      </sec>
      <sec id="sec-1-9">
        <title>Phoneme</title>
      </sec>
      <sec id="sec-1-10">
        <title>Input</title>
      </sec>
      <sec id="sec-1-11">
        <title>Repeat</title>
      </sec>
      <sec id="sec-1-12">
        <title>Stop Gradient Add F C</title>
        <p>F
C</p>
      </sec>
      <sec id="sec-1-13">
        <title>Utterance</title>
      </sec>
      <sec id="sec-1-14">
        <title>Embedding</title>
      </sec>
      <sec id="sec-1-15">
        <title>Utterance Embedding</title>
      </sec>
      <sec id="sec-1-16">
        <title>Predictor</title>
      </sec>
      <sec id="sec-1-17">
        <title>Bert Embedding</title>
      </sec>
      <sec id="sec-1-18">
        <title>Phoneme Num</title>
        <p>(a) Overview</p>
      </sec>
    </sec>
    <sec id="sec-2">
      <title>2. Data preprocessing</title>
      <p>
        long pauses to align punctuation with these two types
of pauses. Redundant punctuation was tagged with no
In accordance with the requirements of Track 1.1 (gener- duration. To account for the silence mark that may
ation task), the AISHELL-3[
        <xref ref-type="bibr" rid="ref8">8</xref>
        ] dataset has been utilized. exist at the beginning of a sentence, we introduced a
This extensive Chinese speech corpus comprises over placeholder in the text and phoneme sequence to ensure
88 thousand utterances, which amount to 85 hours of length matching. Additionally, since a Chinese
characspeech. ter in the MFA phoneme set may correspond to an
indefinite length of phoneme representation, we recorded
2.1. Encoder Input Word corresponding Phoneme Num to ensure that the
length after upsampling the BERT embedding matches
Rhythm modeling in speech synthesis plays a crucial role the phoneme sequence exactly. Although AISHELL3 is a
in enhancing the naturalness and fluency of generated purely Chinese dataset, we adapted the language
depenspeech, particularly with respect to rhythm pauses in dent phoneme(LDP) in MFA to IPA and added a phoneme
unpunctuated long sentences. To address this, we have length regulator to the preprocessing process to
accomintroduced a pre-trained embedding modeling method modate potential cross-lingual TTS scenarios[
        <xref ref-type="bibr" rid="ref7">7</xref>
        ].
based on BERT 1. However, the length of the BERT em- Speech synthesis is a clear one-to-many generation
bedding is commensurate with the number of words, task, wherein diferences in prosodic pauses may occur
while the base model of FastSpeech-based aligns longer in addition to variations in speech rate, pitch, and energy,
rhythm pauses as silence in the speech-text alignment even for the same text and speaker. To better capture
data. This misaligns the input BERT embedding sequence the variances of generated speech, we have incorporated
with the phoneme sequence length, necessitating some an utterance predictor modeling method, in addition to
adjustments to the data preprocessing process. To this BERT embedding. However, unlike the delightful TTS
end, we utilized the MFA[
        <xref ref-type="bibr" rid="ref9">9</xref>
        ] tool to force-align speech model[
        <xref ref-type="bibr" rid="ref10">10</xref>
        ], we have utilized the intermediate vector
reptext. Silences below 250 ms were marked as rhythm resentation of a pre-trained emotion classification model 2
pauses and merged with the previous phoneme, while as the supervisory target for our utterance predictor. The
pauses ranging from 250 to 500 ms were marked as regu- positions of these Encoder inputs in our model are
illuslar pauses, and those exceeding 500 ms were marked as trated in Figures 1 and 2.
      </p>
      <sec id="sec-2-1">
        <title>1https://github.com/Executedone/Chinese-FastSpeech2</title>
      </sec>
      <sec id="sec-2-2">
        <title>2https://github.com/audeering/w2v2-how-to</title>
        <p>2.2. Pitch and Energy</p>
      </sec>
      <sec id="sec-2-3">
        <title>In the first step, we extracted energy features from the</title>
        <p>
          linear spectrogram. Subsequently, we utilized the aligned
LDP duration to average the frame-level energy sequence
and generate the LDP-level energy sequence. The
resulting sequence was quantized to aid in subsequent
processing. In the second step, we utilized the WORLD[
          <xref ref-type="bibr" rid="ref11">11</xref>
          ] tool
to extract the pitch sequence for each speaker’s speech.
Following this, we normalized the sequence using the
speaker’s average and variance. Subsequently, we
obtained the LDP-level normalized real-valued pitch
sequence based on the aligned LDP duration.
        </p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>3. Acoustic Modeling</title>
      <p>Wav
Output
HiFiGAN
Vocoder</p>
      <p>
        Post
Processing
3.1. Encoder
3.2. Explicit modelings
To overcome any potential pronunciation interference
that may arise due to direct superposition of the BERT 3.3. Decoder
embedding input and the aggregated phoneme input, we The Decoder architecture, akin to the Encoder, is
comemploy a Conformer block to fuse the information. This prised of four Conformer blocks. In this approach, we
is accomplished by leveraging a linear layer to adjust the fuse the Encoder output, pitch embedding, and energy
dimension of the BERT embedding input. Analogously, embedding, and pass the resultant through two
Conthe utterance embedding is similarly inserted and pro- former blocks. Subsequently, we incorporate the speaker
cessed through a Conformer block. Ultimately, we obtain embedding obtained from the lookup table into the
arthe Encoder output by passing the input through two chitecture. Finally, we pass this through two Conformer
Conformer blocks. The Conformer block utilized in our blocks and a linear layer to obtain the mel spectrum.
approach is akin to the one described in the delightful
TTS. However, we have made certain modifications, such 3.4. Vocoder and Post-processing
as removing the Depthwise convolution and
substituting the self-attention with Relative Position Multi-head
Attention[
        <xref ref-type="bibr" rid="ref12">12</xref>
        ] to circumvent any potential instability in
pronunciation that may arise towards the close of long
sentences due to absolute position information. The
detail of the Encoder inputs in the model are illustrated in
Figure ??.
      </p>
      <sec id="sec-3-1">
        <title>The log-mel spectrogram generated by the proposed ap</title>
        <p>
          proach are converted into speech signals using a
universal and fine-tuned HiFi-GAN vocoder[
          <xref ref-type="bibr" rid="ref18">18</xref>
          ]. This vocoder
has been pre-trained on the AISHELL3 dataset. In our
experimental setup, we make a post-processing for the
input log-mel spectrum before feeding it to the vocoder.
        </p>
        <p>
          This trick has been demonstrated to alleviate some
highfrequency noise. The processing steps are outlined below:
During the training stage, we adopt the approach
proposed in [
          <xref ref-type="bibr" rid="ref13">13</xref>
          ] and employ a 1-D convolution layer to  = ( +  − 1) * ,    = 1.2 (1)
transform the pitch value into pitch embedding.
Similarly, we use a lookup table to convert quantized energy while  denotes the value of log-mel spectrogram
values into energy embedding. During inference stage, and  is a factor used to adjust . This adjustment
we leverage variance adaptors, including the duration increases the diference between values around 0, which
predictor, in line with the methodology described in [
          <xref ref-type="bibr" rid="ref14">14</xref>
          ]. may explain why high-frequency noise will be reduced.
To ensure that the training of these variance adaptors However, if the coeficient is multiplied by excessively
does not negatively impact the training of the encoder, large value, it can cause an increase in low-frequency
we implement a stop-gradient operation on the input of energy, which changes the speech quality and reduces
43.2 41.72 40.12 39.68 38.14 37.91 37.8 36.63 34.83 33.16
30.71
()
%
R
S
D
W
perceived quality. We used default values based on our
experience. Additionally, we added an ofset due to the
logarithmic distribution not being symmetric about the
zero-point.
domain.
4.2. Metrics
        </p>
      </sec>
      <sec id="sec-3-2">
        <title>Track 1.1 of the ADD challenge is an adversarial game</title>
        <p>that requires participants to generate adversarial
sam4. Results ples and enhance the anti-attack capabilities of the audio
deepfake detection model from two opposing sides. The
4.1. Experimental setup generation and detection tasks of Track 1 are evaluated
separately. For the generation task (Track 1.1), the
decepIn our experimental setup, we downsample the AISHELL- tion success rate (DSR) is chosen as the metric. The DSR
3 audio from 44.1 kHz to 24 kHz for the TTS model and metric quantifies the extent to which the audio deepfake
to 16 kHz for all evaluations. We represent the audio detection model is deceived by the generated utterances
features as a sequence of 80-dimensional log-mel spec- and is defined as follows:
trogram frames. These frames are computed from 40 ms
wthiendCoownsfotrhmaterarbeloschkifsteind obuyr1p0rmopso.sTehdemhoidddeelnisssizeet toof  = ·  (2)
128. Each feed-forward layer of the Conformer blocks where W denotes the count of wrong detection
samcomprises of two 1-D convolution layers, each with a ker- ples by all the detection models on the condition of
nel size of 3 and 1024 intermediate channels. Regarding achieving their respective equal error rate (EER)
perthe acoustic model, we utilize an l1 loss function for the formance, A is the number of evaluation samples, and N
mel spectrum, while the mean-squared error (MSE) loss represents the number of detection models. In the first
function is applied to other components. The duration round, the DSR against the Track 1.2 submissions
constiand energy predictor compute the loss in the logarithmic tutes the overall generation performance metric. In the
second round, weighted consideration is also given to the
DSR against the detection model that we release.
Specifically, in the second round, the generation performance
metric is defined as:
  =   +</p>
        <p>(3)
where  =0.4 and  =0.6, and they denote the respective
weights for EER and DSR in our consideration.
4.3. Evaluations</p>
      </sec>
      <sec id="sec-3-3">
        <title>In Track 1.1, participants are tasked with generating at</title>
        <p>tack samples while adhering to the specified text and
speaker identities. Our team id is A02. As depicted in
Figures 4(a) and 4(b), our deception success rate (DSR)
ranks 7th and 8th in the first and second rounds,
respectively. Notwithstanding, in Figure 4(c), which displays
the DSR scores of the baseline models provided by the
organizers in the second round, we secured the first
position and outperformed the second-ranking system by a
significant margin. Ultimately, our speech synthesis
system obtained the second position in the overall scoring
shown in Figure 4(d).</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>5. Conclusion</title>
      <sec id="sec-4-1">
        <title>This paper presents our current speech synthesis system,</title>
        <p>which has yielded promising results in the adversarial
process with the fake audio detection model. Specifically,
our system has demonstrated significant advantages over
baseline models, thereby validating its eficacy in the task
of fake audio generation.</p>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>Y.</given-names>
            <surname>Jiangyan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Jianhua</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Ruibo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Xinrui</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W.</given-names>
            <surname>Chenglong</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W.</given-names>
            <surname>Tao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Chuyuan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Xiaohui</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Yan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Yong</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Le</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Junzuo</surname>
          </string-name>
          , G. Hao,
          <string-name>
            <given-names>W.</given-names>
            <surname>Zhengqi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Shan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Zheng</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Shuai</surname>
          </string-name>
          , L. Haizhou,
          <year>Add 2023</year>
          :
          <article-title>the second audio deepfake detection challenge</article-title>
          ,
          <source>in: IJCAI 2023 Workshop on Deepfake Audio Detection and Analysis (DADA</source>
          <year>2023</year>
          ),
          <year>2023</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>J.</given-names>
            <surname>Kim</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Kong</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Son</surname>
          </string-name>
          ,
          <article-title>Conditional variational autoencoder with adversarial learning for end-to-end text-to-speech</article-title>
          ,
          <source>in: International Conference on Machine Learning</source>
          , ICML,
          <year>2021</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>M.</given-names>
            <surname>Jeong</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Kim</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S. J.</given-names>
            <surname>Cheon</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B. J.</given-names>
            <surname>Choi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N. S.</given-names>
            <surname>Kim</surname>
          </string-name>
          ,
          <article-title>Dif-tts: A denoising difusion model for text-tospeech</article-title>
          ,
          <source>arXiv preprint arXiv:2104.01409</source>
          (
          <year>2021</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>X.</given-names>
            <surname>Tan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Chen</surname>
          </string-name>
          , H. Liu,
          <string-name>
            <given-names>J.</given-names>
            <surname>Cong</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Zhang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Liu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Leng</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Yi</surname>
          </string-name>
          , et al.,
          <article-title>Naturalspeech: Endto-end text-to-speech synthesis with human-level quality</article-title>
          ,
          <source>arXiv preprint arXiv:2205.04421</source>
          (
          <year>2022</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>Y.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Skerry-Ryan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Stanton</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Wu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R. J.</given-names>
            <surname>Weiss</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Jaitly</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Yang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Xiao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Chen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Bengio</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Q.</given-names>
            <surname>Le</surname>
          </string-name>
          , et al.,
          <article-title>Tacotron: Towards end-to-end speech synthesis</article-title>
          ,
          <source>in: Proc. Interspeech</source>
          ,
          <year>2017</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>Y.</given-names>
            <surname>Ren</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Ruan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Tan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Qin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Zhao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Zhao</surname>
          </string-name>
          , T.-Y. Liu,
          <article-title>Fastspeech: Fast, robust and controllable text-to-speech</article-title>
          ,
          <source>in: Advances in Neural Information Processing Systems(NeurIPS)</source>
          ,
          <year>2019</year>
          , pp.
          <fpage>3171</fpage>
          -
          <lpage>3180</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>H.</given-names>
            <surname>Zhan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Yu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Zhang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Zhang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Lin</surname>
          </string-name>
          ,
          <article-title>Exploring Timbre Disentanglement in NonAutoregressive Cross-Lingual Text-to-</article-title>
          <string-name>
            <surname>Speech</surname>
          </string-name>
          ,
          <source>in: Proc. Interspeech</source>
          <year>2022</year>
          ,
          <year>2022</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>Y.</given-names>
            <surname>Shi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Bu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Xu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Zhang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <article-title>Aishell-3: A multi-speaker mandarin tts corpus and the baselines</article-title>
          , arXiv preprint arXiv:
          <year>2010</year>
          .
          <volume>11567</volume>
          (
          <year>2020</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>M.</given-names>
            <surname>McAulife</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Socolof</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Mihuc</surname>
          </string-name>
          , et al.,
          <article-title>Montreal forced aligner: Trainable text-speech alignment using kaldi</article-title>
          ., in: Interspeech,
          <year>2017</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>Y.</given-names>
            <surname>Liu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Xu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Chen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Tan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>He</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Zhao</surname>
          </string-name>
          ,
          <string-name>
            <surname>Delightfultts:</surname>
          </string-name>
          <article-title>The microsoft speech synthesis system for blizzard challenge 2021</article-title>
          , arXiv preprint arXiv:
          <volume>2110</volume>
          .12612 (
          <year>2021</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>M.</given-names>
            <surname>Morise</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Yokomori</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Ozawa</surname>
          </string-name>
          , World:
          <article-title>a vocoder-based high-quality speech synthesis system for real-time applications</article-title>
          ,
          <source>IEICE Transactions on Information and Systems</source>
          <volume>99</volume>
          (
          <year>2016</year>
          )
          <fpage>1877</fpage>
          -
          <lpage>1884</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <surname>C.-Z. A. Huang</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Vaswani</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          <string-name>
            <surname>Uszkoreit</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          <string-name>
            <surname>Shazeer</surname>
            , I. Simon,
            <given-names>C.</given-names>
          </string-name>
          <string-name>
            <surname>Hawthorne</surname>
            ,
            <given-names>A. M.</given-names>
          </string-name>
          <string-name>
            <surname>Dai</surname>
            ,
            <given-names>M. D.</given-names>
          </string-name>
          <string-name>
            <surname>Hofman</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          <string-name>
            <surname>Dinculescu</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          <string-name>
            <surname>Eck</surname>
          </string-name>
          , Music transformer, arXiv preprint arXiv:
          <year>1809</year>
          .
          <volume>04281</volume>
          (
          <year>2018</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <given-names>A.</given-names>
            <surname>Łańcucki</surname>
          </string-name>
          ,
          <article-title>Fastpitch: Parallel text-to-speech with pitch prediction</article-title>
          ,
          <source>in: 2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)</source>
          ,
          <year>2021</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <given-names>Y.</given-names>
            <surname>Ren</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Hu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Tan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Qin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Zhao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Zhao</surname>
          </string-name>
          , T.-Y. Liu,
          <article-title>Fastspeech 2: Fast and high-quality end-to-end text-to-speech</article-title>
          ,
          <source>in: International Conference on Learning Representations(ICLR)</source>
          ,
          <year>2021</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <given-names>J.</given-names>
            <surname>Kim</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Kim</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Kong</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Yoon</surname>
          </string-name>
          ,
          <article-title>Glow-tts: A generative flow for text-to-speech via monotonic alignment search</article-title>
          ,
          <source>in: Advances in Neural Information Processing Systems</source>
          ,
          <year>2020</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [16]
          <string-name>
            <given-names>T.</given-names>
            <surname>Raitio</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Rasipuram</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Castellani</surname>
          </string-name>
          ,
          <string-name>
            <surname>Controllable Neural</surname>
          </string-name>
          Text-to-
          <article-title>Speech Synthesis Using Intuitive Prosodic Features</article-title>
          ,
          <source>in: Proc. Interspeech</source>
          ,
          <year>2020</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [17]
          <string-name>
            <given-names>C.</given-names>
            <surname>Gong</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Ling</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Guo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Zhang</surname>
          </string-name>
          , J. Dang,
          <article-title>Improving naturalness and controllability of sequence-to-sequence speech synthesis by learning local prosody representations</article-title>
          ,
          <source>in: 2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)</source>
          ,
          <year>2021</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          [18]
          <string-name>
            <given-names>J.</given-names>
            <surname>Kong</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Kim</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Bae</surname>
          </string-name>
          , Hifi-gan:
          <article-title>Generative adversarial networks for eficient and high fidelity speech synthesis</article-title>
          ,
          <source>Advances in Neural Information Processing Systems</source>
          (
          <year>2020</year>
          ).
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>