<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <journal-title-group>
        <journal-title>ORCID:</journal-title>
      </journal-title-group>
    </journal-meta>
    <article-meta>
      <title-group>
        <article-title>Antispoofing Countermeasures in Modern Voice Authentication Systems</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Michael V. Evsyukov</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Michael M. Putyato</string-name>
          <email>putyato.m@gmail.com</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Alexander S. Makaryan</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Kuban State Technological University</institution>
          ,
          <addr-line>2, Moskovskaya Str., Krasnodar, 350072</addr-line>
          ,
          <country country="RU">Russia</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>1801</year>
      </pub-date>
      <volume>000</volume>
      <fpage>0</fpage>
      <lpage>0001</lpage>
      <abstract>
        <p>The article provides an overview of modern voice authentication systems. The relevance and importance of voice authentication systems are substantiated. This article focuses on spoofing methods and antispoofing countermeasures because the main challenge in developing voice authentication systems is protection against spoofing. The most commonly used approaches to spoofing, their efficiency, and most important features are described. The existing types of countermeasures against spoofing are presented. The experience of their use against various types of spoofing is described, and links to specific implementations are provided. Improved classification of antispoofing countermeasures is proposed. Promising areas of research in the field of voice authentication are highlighted. According to the authors, the promising areas of research are countermeasures with high generalizing ability, countermeasures specialized against replay attacks, text-dependent countermeasures, development and improvement of joint approaches to assessing the efficiency of an automatic speaker verification system with integrated countermeasures. biometrics, authentication, voice, spoofing, information security, artificial intelligence, GMM, Proceedings of VI International Scientific and Practical Conference Distance Learning Technologies (DLT-2021), September 20-22,</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>SVM.</p>
    </sec>
    <sec id="sec-2">
      <title>1. Introduction</title>
      <p>
        According to Google data, 500 million people use Google Assistant monthly [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. Apple claims the
voice assistant Siri handles 25 billion requests every month. [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]. Ease of use and time savings are the
main reasons why voice assistants are gaining popularity. In addition, the emergence of a wide range
of smart devices and the rapid development of the Internet of Things (IoT) make the voice interface
even more in demand, as it can provide the most comfortable user experience. Voice control is
implemented, for example, in the Yandex.Station smart speaker, Tesla cars, as well as in various smart
home systems.
      </p>
      <p>Thus, voice assistants have entered the daily life of numerous users, and the next natural step is their
introduction to payment systems and banking. The main driver for the development of voice solutions
is personalization because voice interaction can provide valuable information about the needs and
behavior of customers. It allows banks and FinTech companies to offer services that precisely meet the
expectations of a particular user.</p>
      <p>
        In a 2017 survey by Business Insider Intelligence in the United States, 8% of respondents said that
they used voice commands to buy goods, pay bills, and perform P2P transactions. According to their
forecast, by 2022 the number of users of voice interfaces is expected to reach 31% of the US adult
population [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ].
      </p>
      <p>The graph of growth in the millions of US users of voice interfaces is presented in Figure 1.</p>
      <p>2021 Copyright for this paper by its authors.
61,1</p>
      <p>77,9
23,3</p>
      <p>33,3
18,4
2017
2018
2019
2020
2021
2022</p>
      <p>The high consumer value of voice payments is driving banks and payment providers such as PayPal,
Amazon, Apple, and Google to develop artificial intelligence technologies dedicated to voice
processing.</p>
      <p>
        However, information security problems are the main obstacle that prevents voice payments from
gaining the full confidence of banks and users. For them to become as natural as interaction with a
merchant or a bank employee, it is necessary to improve the existing methods of protection and
authentication [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ].
      </p>
      <p>Automatic speaker verification algorithms are well studied, easy to use, and applicable for both
continuous and one-time authentication. However, due to the widespread availability of inexpensive
audio recording and reproducing devices, they are susceptible to spoofing, i.e. vulnerable to the actions
of intruders aimed at impersonating another person. In this regard, the study of antispoofing methods is
the main direction in the development of voice authentication systems.</p>
      <p>The purpose of this article is to review the current state of research in the field of antispoofing
countermeasures for voice authentication systems.</p>
    </sec>
    <sec id="sec-3">
      <title>2. Types of Spoofing</title>
      <p>Spoofing refers to the actions of an attacker aimed at successfully authenticating in the system under
the guise of another person. Due to the widespread availability of high-quality sound recording and
reproducing equipment, voice authentication systems are highly susceptible to spoofing.</p>
      <p>
        The main types of spoofing are [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ]:
1. Impersonation
      </p>
      <p>This type of spoofing involves one person mimicking the vocal characteristics of another.
Impersonation differs from other types of spoofing in that the attacker does not need to use auxiliary
technical means and methods to implement it. In this regard, counteraction to this type of spoofing does
not require additional countermeasures and is performed by an automatic speaker verification system
itself.</p>
      <p>2. Speech recording (replay attack)</p>
      <p>Speech recording is a simple and efficient form of spoofing. According to a large number of
researchers, it poses the most serious threat to automatic speaker verification systems. Its
implementation consists in recording a fragment of a person’s speech and then replaying it to an
authentication system during verification.</p>
      <p>3. Voice conversion</p>
      <p>Voice conversion involves the use of specialized software that modifies a person’s voice in such a
way that it becomes similar to another person’s voice.</p>
      <p>
        Estimating resistance of various countermeasures against this type of spoofing was the subject of
the ASVSpoof 2015 competition [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ].
      </p>
      <p>
        The following speech conversion algorithms were used during ASVSpoof 2015:
 exemplar-based unit selection for voice conversion utilizing temporal information [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ];
 adjusting the first Mel-cepstral coefficient to make the voice specter resemble that of the target
(one of the simplest algorithms) [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ];
 voice conversion based on Gaussian Mixture Model (the most commonly used) [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ];
 voice conversion based on the tensor representation of speaker space [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ];
 voice conversion, using dynamic kernel partial least squares regression [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ].
4. Speech synthesis
      </p>
      <p>This method involves the generation of artificial speech based on an arbitrary text that resembles the
voice of a certain person.</p>
      <p>
        Estimating the ability of various countermeasures to resist this type of spoofing was the subject of
the ASVSpoof 2015. The following speech synthesis algorithms were used during the competition:
 unit-selection concatenative speech synthesis (this type of spoofing was found to be the most
efficient one) [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ];
 statistical speech synthesis based on the Hidden Markov Model [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ].
      </p>
      <p>
        Deepfake technology, which implies the use of generative-adversarial neural networks, is a
promising method for speech synthesis and voice conversion. The assessment of the ability of
countermeasures to resist deepfake-based spoofing will be carried out during the ASVSpoof 2021
competition [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ].
      </p>
      <p>
        5. Disguised attacks on speech processing systems that exploit the human perception of sound
The study [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ] considers 4 ways of transforming a voice recording in such a way that it becomes
unintelligible to a person, but so that its essential acoustic vocal features remain unchanged, and the
recording can pass a speaker verification system or be processed by a speech recognition system.
      </p>
    </sec>
    <sec id="sec-4">
      <title>3. Antispoofing Countermeasures</title>
      <p>To resist various spoofing technologies, the voice authentication system must include an additional
antispoofing mechanism called countermeasure. The purpose of the countermeasure is to notice the fact
that the system is undergoing a spoofing attack.</p>
      <p>
        A general classification of countermeasures against spoofing is presented in [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ]. We offer an updated
version of it.
      </p>
      <p>1. Challenge-response-based countermeasures</p>
      <p>Challenge-response-based countermeasures involve explicit user interaction with the system during
authentication. As a rule, when implementing them, the system generates random text that needs to be
read by the user. To authenticate the user, a text-independent verification algorithm is used, and the
correctness of reading the text is checked by a speech recognition algorithm.</p>
      <p>This type of countermeasure is highly effective against the most dangerous type of spoofing – replay
attacks. This is because the attacker, as a rule, cannot pre-record the user’s speech in such a way that it
would be possible to quickly compose a random utterance from its fragments.</p>
      <p>
        2. Acoustic-features-based countermeasures for detection of synthesized and converted speech
This type of countermeasure aims at extracting imperfections from voice recordings that indicate
that a piece of speech has been obtained using speech synthesis or voice conversion techniques. Such
countermeasures were the subject of research during the ASVSpoof 2015 competition [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ].
      </p>
      <p>Countermeasures of this type, like speaker verification methods, are mainly based on the use of
short-term spectral characteristics. The most widely used classifiers for these countermeasures are the
Gaussian Mixture Model, Support Vector Machines, and Artificial Neural Networks.
3. Voice liveness detection methods based on features of the human vocal tract</p>
      <p>Since spoofing involves using loudspeakers, the task of voice authentication can be represented as
a combination of the following two tasks:
 authentic a person by his voice characteristics (verification);
 confirm that the source of the voice is a live person (countermeasure).</p>
      <p>This type of countermeasure relies on the characteristics of the human vocal tract that cause acoustic
effects that are difficult to record and reproduce using artificial means.</p>
      <p>
        The examples of such countermeasures are:
 voice liveness detection algorithms based on pop noise caused by the human breath [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ];
 phoneme localization-based liveness detection for voice authentication on smartphones [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ].
      </p>
      <p>
        The scheme of phoneme localization-based liveness detection for voice authentication on
smartphones is presented in Figure 2 [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ].
 articulatory gesture-based liveness detection [
        <xref ref-type="bibr" rid="ref17">17</xref>
        ].
4. Voice liveness detection methods based on features of loudspeakers
      </p>
      <p>
        At the moment, a wide range of users have access to sound recording and playback devices capable
of copying a human voice with very high quality, which continues to evolve. When using the previously
described voice characteristics (for example, MFCC), it is extremely difficult to distinguish an artificial
voice from a live one, as evidenced by the results of the ASVSpoof 2017 competition [
        <xref ref-type="bibr" rid="ref18">18</xref>
        ].
      </p>
      <p>
        In this regard, it is proposed to use discriminative features of another nature that make it possible to
understand if the source of a voice during an authentication attempt is a loudspeaker. For example, in
[
        <xref ref-type="bibr" rid="ref19">19</xref>
        ] it is proposed to use a magnetic field to distinguish a loudspeaker from a live person.
5. Multi-modal biometric-based methods
      </p>
      <p>This group of methods implies increasing the efficiency of the authentication system and resistance
to spoofing by using two or more unrelated biometric characteristics.</p>
      <p>
        For example, work [
        <xref ref-type="bibr" rid="ref20">20</xref>
        ] proposes a bimodal identity verification system using Gaussian Mixture
Model with Universal Background Model for voice authentication and a face verification system using
Gabor’s features and linear discriminant analysis. Also, in [
        <xref ref-type="bibr" rid="ref21">21</xref>
        ] the possibility of using keyboard
handwriting in conjunction with other types of authentication is considered.
      </p>
    </sec>
    <sec id="sec-5">
      <title>4. Conclusion</title>
      <p>Currently, there is a large number of ways to perform spoofing against a voice authentication system.
On the other hand, there is also a variety of anti-spoofing approaches. Different types of antispoofing
are differently efficient against certain types of spoofing. For example, the challenge-response approach
works well against voice replay attacks, but it does not apply to other types of spoofing. In turn,
countermeasures that rely on the search for acoustic imperfections of the voice used for spoofing are
effective against speech synthesis and voice conversion, but much less effective against replay attacks.
In addition, different systems perform differently in countering different spoofing algorithms of the
same type.</p>
      <p>Therefore, a promising area of research is countermeasures with a high generalizing ability and
countermeasures specialized against replay attacks.</p>
      <p>In addition, during previous ASVSpoof competitions, only text-independent countermeasures were
considered, however, the development of text-dependent countermeasures can have certain advantages
as well.</p>
      <p>Another promising area of research is the development and improvement of joint approaches to
assessing the efficiency of automatic speaker verification systems with integrated countermeasures.</p>
    </sec>
    <sec id="sec-6">
      <title>5. References</title>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>L.</given-names>
            <surname>Eadicicco</surname>
          </string-name>
          ,
          <article-title>Google just revealed that half a billion people around the world are using the Google Assistant as it battles with Amazon to conquer the smart home</article-title>
          ,
          <year>2020</year>
          . URL: https://www.businessinsider.com/google-assistant-500
          <string-name>
            <surname>-</surname>
          </string-name>
          million
          <article-title>-users-challenges-amazon-alexa2020-1</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>B.</given-names>
            <surname>Kinsella</surname>
          </string-name>
          , Apple Still in Holding Pattern on Voice,
          <source>Siri Used 25 Billion Times Per Month But New Features Limited</source>
          ,
          <year>2020</year>
          : URL: https://voicebot.ai/
          <year>2020</year>
          /06/22/apple-still
          <article-title>-in-holding-patternon-voice-siri-used-25-billion-times-per-month-but-new-features-limited/</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>D.V.</given-names>
            <surname>Dyke</surname>
          </string-name>
          ,
          <article-title>Soon nearly a third of US consumers will regularly make payments with their voice</article-title>
          ,
          <year>2017</year>
          . URL: https://www.businessinsider.
          <article-title>com/the-voice-payments-</article-title>
          <string-name>
            <surname>report-</surname>
          </string-name>
          2017-6?r=US&amp;IR=T
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>B.</given-names>
            <surname>Hao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Hei</surname>
          </string-name>
          ,
          <article-title>Voice Liveness Detection for Medical Devices</article-title>
          , in: D.R. Kisku (Ed.), P. Gupta (Ed.),
          <string-name>
            <given-names>J.K.</given-names>
            <surname>Sing</surname>
          </string-name>
          (Ed.),
          <source>Design and Implementation of Healthcare Biometric Systems</source>
          ,
          <year>2019</year>
          , pp.
          <fpage>109</fpage>
          -
          <lpage>136</lpage>
          . DOI:
          <volume>10</volume>
          .4018/978-1-
          <fpage>5225</fpage>
          -7525-2.
          <fpage>ch005</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>Z.</given-names>
            <surname>Wu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Yamagishi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Kinnunen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Hanilc</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Sahidullah</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Sizov</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Evans</surname>
          </string-name>
          ,
          <string-name>
            <surname>M.</surname>
          </string-name>
          <article-title>Todisco, ASVspoof: the Automatic Speaker Verification Spoofing and Countermeasures Challenge</article-title>
          ,
          <source>IEEE Journal of Selected Topics in Signal Processing</source>
          ,
          <volume>6</volume>
          (
          <issue>1</issue>
          ) (
          <year>2016</year>
          )
          <fpage>588</fpage>
          -
          <lpage>604</lpage>
          . DOI:
          <volume>10</volume>
          .1109/JSTSP.
          <year>2017</year>
          .2671435
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>Z.</given-names>
            <surname>Wu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Virtanen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Kinnunen</surname>
          </string-name>
          , E. Chng,
          <string-name>
            <given-names>H.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <article-title>Exemplar-based unit selection for voice conversion utilizing temporal information</article-title>
          ,
          <source>in: Proceedings of 14th Annual Conference of the International Speech Communication Association</source>
          ,
          <year>Interspeech 2013</year>
          ,
          <year>2013</year>
          , pp.
          <fpage>3057</fpage>
          -
          <lpage>3061</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>T.</given-names>
            <surname>Fukada</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Tokuda</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Kobayashi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Imai</surname>
          </string-name>
          ,
          <article-title>An adaptive algorithm formel-cepstral analysis of speech</article-title>
          ,
          <source>in: Proceedings of International Conference on Acoustics, Speech, and Signal Processing, ICASSP-92</source>
          ,
          <year>1992</year>
          , pp.
          <fpage>137</fpage>
          -
          <lpage>140</lpage>
          . DOI:
          <volume>10</volume>
          .1109/ICASSP.
          <year>1992</year>
          .225953
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>T.</given-names>
            <surname>Toda</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.W.</given-names>
            <surname>Black</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Tokuda</surname>
          </string-name>
          ,
          <article-title>Voice conversion based on maximum-likelihood estimation of spectral parameter trajectory</article-title>
          ,
          <source>IEEE Transactions on Audio, Speech, and Language Processing</source>
          <volume>15</volume>
          (
          <issue>8</issue>
          ) (
          <year>2007</year>
          )
          <fpage>2222</fpage>
          -
          <lpage>2235</lpage>
          . DOI:
          <volume>10</volume>
          .1109/TASL.
          <year>2007</year>
          .907344
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>D.</given-names>
            <surname>Saito</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Yamamoto</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Minematsu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Hirose</surname>
          </string-name>
          ,
          <article-title>One-to-many voice conversion based on the tensor representation of speaker space</article-title>
          ,
          <source>in: 12th Annual Conference of the International Speech Communication Association, INTERSPEECH</source>
          <year>2011</year>
          , Florence, Italy,
          <year>2011</year>
          , pp.
          <fpage>653</fpage>
          -
          <lpage>656</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>E.</given-names>
            <surname>Helander</surname>
          </string-name>
          , H. Sil´en, T. Virtanen,
          <string-name>
            <given-names>M.</given-names>
            <surname>Gabbouj</surname>
          </string-name>
          ,
          <article-title>Voice conversion using dynamic kernel partial least squares regression</article-title>
          ,
          <source>IEEE Transactions on Audio Speech and Language Processing</source>
          <volume>20</volume>
          (
          <issue>3</issue>
          ) (
          <year>2012</year>
          )
          <fpage>806</fpage>
          -
          <lpage>817</lpage>
          . DOI:
          <volume>10</volume>
          .1109/TASL.
          <year>2011</year>
          .2165944
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>A.J.</given-names>
            <surname>Hunt</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.W.</given-names>
            <surname>Black</surname>
          </string-name>
          ,
          <article-title>Unit selection in a concatenative speech synthesis system using a large speech database</article-title>
          , in: IEEE International Conference on Acoustics, Speech, and Signal Processing Conference Proceedings,
          <year>1996</year>
          . DOI:
          <volume>10</volume>
          .1109/ICASSP.
          <year>1996</year>
          .541110
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>J.</given-names>
            <surname>Yamagishi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Kobayashi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Nakano</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Ogata</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Isogai</surname>
          </string-name>
          ,
          <article-title>Analysis of speaker adaptation algorithms for HMM-based speech synthesis and a constrained smaplr adaptation algorithm</article-title>
          ,
          <source>IEEE Trans. Audio, Speech and Language Processing</source>
          ,
          <volume>17</volume>
          (
          <issue>1</issue>
          ) 2009
          <fpage>66</fpage>
          -
          <lpage>83</lpage>
          . DOI:
          <volume>10</volume>
          .1109/TASL.
          <year>2008</year>
          .2006647
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <given-names>H.</given-names>
            <surname>Delgado</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Evans</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Kinnunen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.A.</given-names>
            <surname>Lee</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Liu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Nautsch</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Patino</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Sahidullah</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Todisco</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <surname>J</surname>
          </string-name>
          . Yamagishi,
          <string-name>
            <surname>ASVspoof</surname>
          </string-name>
          <year>2021</year>
          :
          <article-title>Automatic Speaker Verification Spoofing and Countermeasures Challenge Evaluation Plan</article-title>
          , ASVspoof consortium,
          <year>2021</year>
          . URL: https://www.asvspoof.org/asvspoof2021/asvspoof2021_evaluation_plan.pdf
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <given-names>H.</given-names>
            <surname>Abdullah</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W.</given-names>
            <surname>Garcia</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Peeters</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Traynor</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Butler</surname>
          </string-name>
          , J. Wilson,
          <article-title>Practical Hidden Voice Attacks against Speech and Speaker Recognition Systems, The Network and Distributed System Security Symposium</article-title>
          (NDSS),
          <year>2019</year>
          . URL: https://www.ndss-symposium.org/wpcontent/uploads/2019/02/ndss2019_
          <fpage>08</fpage>
          -
          <lpage>1</lpage>
          _Abdullah_paper.pdf
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <given-names>S.</given-names>
            <surname>Shiota</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Villavicencio</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Yamagishi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Ono</surname>
          </string-name>
          , I. Echizen, T. Matsui,
          <article-title>Voice liveness detection algorithms based on pop noise caused by human breath for automatic speaker verification</article-title>
          ,
          <source>in 16th Annual Conference of the International Speech Communication Association</source>
          ,
          <year>Interspeech 2015</year>
          ,
          <year>2015</year>
          . pp.
          <fpage>239</fpage>
          -
          <lpage>243</lpage>
          . DOI:
          <volume>10</volume>
          .21437/Interspeech.2015-
          <fpage>92</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [16]
          <string-name>
            <given-names>L.</given-names>
            <surname>Zhang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Tan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Yang</surname>
          </string-name>
          , et al.,
          <article-title>VoiceLive: A phoneme localization-based liveness detection for voice authentication on smartphones</article-title>
          ,
          <source>in: Proceedings of the 2016 ACM SIGSAC Conference on Computer and Communications Security, CCS'16</source>
          ,
          <year>2016</year>
          , pp.
          <fpage>1080</fpage>
          -
          <lpage>1091</lpage>
          . DOI:
          <volume>10</volume>
          .1145/2976749.2978296
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [17]
          <string-name>
            <given-names>L.</given-names>
            <surname>Zhang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Tan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Yang</surname>
          </string-name>
          ,
          <article-title>Hearing your voice is not enough: An articulatory gesture-based liveness detection for voice authentication</article-title>
          ,
          <source>in: Proceedings of the 2016 ACM SIGSAC Conference on Computer and Communications Security, CCS'17</source>
          ,
          <year>2017</year>
          , pp.
          <fpage>55</fpage>
          -
          <lpage>71</lpage>
          . DOI:
          <volume>10</volume>
          .1145/3133956.3133962
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          [18]
          <string-name>
            <given-names>T.</given-names>
            <surname>Kinnunen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Sahidullah</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Delgado</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Todisco</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Evans</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Yamagishi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.A.</given-names>
            <surname>Lee</surname>
          </string-name>
          , 19th Annual Conference of the International Speech Communication Association,
          <year>Interspeech 2018</year>
          , Stockholm, Sweden,
          <year>2018</year>
          . DOI:
          <volume>10</volume>
          .21437/Interspeech.2017-
          <fpage>1111</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          [19]
          <string-name>
            <given-names>L.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Chen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Wang</surname>
          </string-name>
          , et al.,
          <article-title>A study on replay attack and anti-spoofing for automatic speaker verification</article-title>
          ,
          <source>in: Proceedings of 18th Annual Conference of the International Speech Communication Association</source>
          ,
          <year>Interspeech 2017</year>
          , Stockholm, Sweden,
          <year>2017</year>
          , pp.
          <fpage>92</fpage>
          -
          <lpage>96</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          [20]
          <string-name>
            <given-names>A.</given-names>
            <surname>Usoltsev</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Petrovska-Delacrétaz</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Houssemeddine</surname>
          </string-name>
          ,
          <article-title>Full Video Processing for Mobile Audio-Visual Identity Verification</article-title>
          ,
          <source>in: Proceedings of the 5th International Conference on Pattern Recognition Applications and Methods</source>
          ,
          <string-name>
            <surname>ICPRAM</surname>
          </string-name>
          <year>2016</year>
          ,
          <year>2016</year>
          , pp.
          <fpage>552</fpage>
          -
          <lpage>557</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          [21]
          <string-name>
            <surname>M.M. Putyato</surname>
            ,
            <given-names>A.S.</given-names>
          </string-name>
          <string-name>
            <surname>Makaryan</surname>
            ,
            <given-names>S.M.</given-names>
          </string-name>
          <string-name>
            <surname>Chich</surname>
            ,
            <given-names>V.K.</given-names>
          </string-name>
          <string-name>
            <surname>Markova</surname>
          </string-name>
          ,
          <article-title>System development for identification and confirmation of access legitimacy based on biometric authentication dynamic methods</article-title>
          ,
          <source>Caspian Journal: Control and High Technologies</source>
          ,
          <volume>3</volume>
          (
          <issue>51</issue>
          ) (
          <year>2020</year>
          )
          <fpage>83</fpage>
          -
          <lpage>93</lpage>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>