<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Leonardo Badino Center for Translational Neurophysiology of Speech and Communication Istituto Italiano di Tecnologia - Italy</article-title>
      </title-group>
      <pub-date>
        <year>2016</year>
      </pub-date>
      <abstract>
        <p />
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>English. Despite the impressive results
achieved by ASR technology in the last few
years, state-of-the-art ASR systems can still
perform poorly when training and testing
conditions are different (e.g., different acoustic
environments). This is usually referred to as
the mismatch problem. In the ArtiPhon task at
Evalita 2016 we wanted to evaluate phone
recognition systems in mismatched speaking
styles. While training data consisted of read
speech, most of testing data consisted of
single-speaker hypo- and hyper-articulated
speech. A second goal of the task was to
investigate whether the use of speech
production knowledge, in the form of measured
articulatory movements, could help in building
ASR systems that are more robust to the
effects of the mismatch problem. Here I report
the result of the only entry of the task and of
baseline systems.</p>
      <p>Italiano. Nonostante i notevoli risultati
ottenuti recentemente nel riconoscimento
automatico del parlato (ASR) le prestazioni dei
sistemi ASR peggiorano significativamente in
quando le condizioni di testing sono differenti
da quelle di training (per esempio quando il
tipo di rumore acustico è differente). Un
primo gol della ArtiPhon task ad Evalita 2016 è
quello di valutare il comportamento di sistemi
di riconoscimento fonetico in presenza di un
mismatch in termini di registro del parlato.
Mentre il parlato di training consiste di frasi
lette ad un velocita; di eloquio “standard”, il
parlato di testing consiste di frasi sia
iperche ipo-articolate. Un secondo gol della task
è quello di analizzare se e come l’utilizzo di
informazione concernente la produzione del
parlato migliora l’accuratezza dell’ASR e in
particolare nel caso di mismatch a livello di
registri del parlato. Qui riporto risultati
dell’unico sistema che è stato sottomesso e di
una baseline.
1</p>
    </sec>
    <sec id="sec-2">
      <title>Introduction</title>
      <p>
        In the last five years ASR technology has
achieved remarkable results, thanks to
increased training data, computational resources,
and the use of deep neural networks (DNNs,
(LeCun et al., 2015)). However, the
performance of connectionist ASR degrades when
testing conditions are different from training
conditions (e.g., acoustic environments are
different)
        <xref ref-type="bibr" rid="ref8">(Huang et al., 2014)</xref>
        . This is usually
referred to as the training-testing mismatch
problem. This problem is partly masked by
multi-condition training
        <xref ref-type="bibr" rid="ref14">(Seltzer et al., 2013)</xref>
        that consists in using very large training
datasets (up to thousands of hours) of transcribed
speech to cover as many sources of variability
as possible (e.g., speaker’s gender, age and
accent, different acoustic environments).
      </p>
      <p>One of the two main goals of the ArtiPhon
task at Evalita 2016 was to evaluate phone
recognition systems in mismatched speaking
styles. Between training and testing data the
speaking style was the condition that differed.
More specifically, while the training dataset
consists of read speech where the speaker was
required to keep a constant speech rate, testing
data range from slow and hyper-articulated
speech to fast and hypo-articulated speech.
Training and testing data are from the same
speaker.</p>
      <p>The second goal of the ArtiPhon task was to
investigate whether the use of speech
production knowledge, in the form of measured
articulatory movements, could help in building
ASR systems that are more robust to the
effects of the mismatch problem.</p>
      <p>
        The use of speech production knowledge,
i.e., knowledge about how the vocal tract
behaves when it produces speech sounds, is
motivated by the fact that complex phenomena
observed in speech, for which a simple purely
acoustic description has still to be found, can
be easily and compactly described in speech
production-based representations. For
example, in Articulatory Phonology
        <xref ref-type="bibr" rid="ref4">(Browman and
Goldstein, 1992)</xref>
        or in the distinctive features
framework
        <xref ref-type="bibr" rid="ref9">(Jakobson et al., 1952)</xref>
        coarticulation effects can be compactly modeled as
temporal overlaps of few vocal tract gestures. The
vocal tract gestures are regarded as invariant,
i.e., context- and speaker-independent,
production targets that contribute to the realization of
a phonetic segment. Obviously the invariance
of a vocal tract gesture partly depends on the
degree of abstraction of the representation but
speech production representations offer
compact descriptions of complex phenomena and
of phonetic targets that purely acoustic
representations are not able to provide yet
        <xref ref-type="bibr" rid="ref10">(Maddieson, 1997)</xref>
        .
      </p>
      <p>
        Recently, my colleagues and I have
proposed DNN-based “articulatory” ASR where
the DNN that computes phone probabilities is
forced, during training, to learn/use motor
features. We have proposed strategies that allow
motor information to produce an inductive bias
on learning. The bias resulted in improvements
over strong DNN-based purely auditory
baselines, in both speaker-dependent
        <xref ref-type="bibr" rid="ref2 ref3">(Badino et al.,
2016)</xref>
        and speaker-independent settings
(Badino, 2016)
      </p>
      <p>
        Regarding the Artiphon task, unfortunately
only one out of the 6 research groups that
expressed an interest in the task actually
participated (Piero Cosi from ISTC at CNR,
henceforth I will refer to this participant as ISTC)
        <xref ref-type="bibr" rid="ref7">(Cosi, 2016)</xref>
        . The ISTC system did not use
articulatory data.
      </p>
      <p>In this report I will present results of the
ISTC phone recognition systems and of
baseline systems that also used articulatory data.
2</p>
    </sec>
    <sec id="sec-3">
      <title>Data</title>
      <p>
        The training and testing datasets used for the
ArtiPhon task were selected from voice cnz of
the Italian MSPKA corpus
(http://www.mspkacorpus.it/)
        <xref ref-type="bibr" rid="ref5">(Canevari et al.,
2015)</xref>
        , which was collected in 2015 at the
Istituto Italiano di Tecnologia (IIT).
      </p>
      <p>
        The training dataset corresponds to the
666utterance session 1 of MPSKA, where the
speaker was required to keep a constant speech
rate. The testing dataset was a 40-utterance
subset selected from session 2 of MPSKA.
Session 2 of MPSKA contains a continuum of
ten descending articulation degrees, from
hyper-articulated to hypo-articulated speech.
Details on the procedure used to elicit this
continuum are provided in
        <xref ref-type="bibr" rid="ref5">(Canevari et al., 2015)</xref>
        .
      </p>
      <p>Articulatory data consist of trajectories of 7
vocal tract articulators and recorded with the
NDI (Northern Digital Instruments, Canada)
wave speech electromagnetic articulography
system at 400 Hz.</p>
      <p>Seven 5-Degree-of-freedom (DOF) sensor
coils were attached to upper and lower lips
(UL and LL), upper and lower incisors (UI and
LI), tongue tip (TT), tongue blade (TB) and
tongue dorsum (TD). For head movement
correction a 6-DOF sensor coil was fixed on the
bridge of a pair of glasses worn by the
speakers.</p>
      <p>The NDI system tracks sensor coils in 3D
space providing 7 measurements per each coil:
3 positions (i.e. x; y; z) and 4 rotations (i.e.
Q0;Q1;Q2;Q3) in quaternion format with Q0 =
0 for 5-DOF sensor coils.</p>
      <p>Contrarily to other articulographic systems
(e.g. Carstens 2D AG200, AG100) speakers
head is free to move. That increases comfort
and the naturalness of speech.</p>
      <p>During recordings speakers were asked to
read aloud each sentence that is prompted on a
computer screen. In order to minimize
disfluencies speakers had time to silently read each
sentence before reading out.</p>
      <p>The audio files of the MSPKA corpus are
partly saturated.</p>
      <p>
        The phone set consists of 60 phonemes,
although the participants could collapsed them to
48 phonemes as proposed in
        <xref ref-type="bibr" rid="ref5">(Canevari et al.,
2015)</xref>
        .
3
      </p>
    </sec>
    <sec id="sec-4">
      <title>Sub-tasks</title>
      <p>In Artiphon sub-tasks are phone recognition
tasks. The participants were asked to:
 train phone recognition systems on the
training dataset and then run them on the
test dataset;
 (optional) use articulatory data to build
“articulatory” phone recognition systems.</p>
      <p>Articulatory data were also provided in the
test dataset thus three different scenarios were
possible:
 Scenario 1. Articulatory data not available
 Scenario 2. Articulatory data available at
both training and testing.
 Scenario 3. Articulatory data available
only at training.</p>
      <p>Note that only scenarios 1 and 3 are realistic
ASR scenarios as during testing articulatory
data are very difficult to access.</p>
      <p>Participants could build purely acoustic and
articulatory phone recognition systems starting
from the Matlab toolbox developed at IIT,
available at
https://github.com/robotology/natural-speech.
4</p>
    </sec>
    <sec id="sec-5">
      <title>Phone recognition systems</title>
      <p>Baseline systems are hybrid DNN-HMM
systems while ISTC systems are either
GMMHMM or DNN-HMM systems with
DNNHMM.</p>
      <p>The ISTC systems were trained using the
KALDI ASR engine. ISTC systems used either
the full phone set (with 60 phone labels) or a
reduced phone set (with 29 phones). In the
reduced phone set all phones that are not actual
phonemes in current Italian were correctly
removed. However, important phonemes were
also arbitrarily removed, most importantly,
geminates and corresponding non-geminate
phones were collapsed into a single phone
(e.g., /pp/ and /p/ were both represented by
label /p/).</p>
      <p>ISCT systems used either monophones or
triphones.</p>
      <p>
        ISCT systems were built using KALDI
        <xref ref-type="bibr" rid="ref11 ref12">(Povey et al., 2011)</xref>
        with TIMIT recipes
adapted to the APASCI dataset
        <xref ref-type="bibr" rid="ref1">(Angelini &amp;
al., 1994)</xref>
        . Two training datasets where used:
 the single-speaker dataset provided within
the ArtiPhon task;
 the APASCI dataset.
      </p>
      <p>In all cases only acoustic data were used
(scenario 1), so the recognition systems were
purely acoustic recognition systems.
Henceforth I will refer to ISTC systems trained on
the ArtiPhon single-speaker training dataset as
speaker-dependent ISTC systems (as the
speaker in training and testing data is the same)
and to ISTC systems trained on the APASCI
dataset as speaker-independent ISTC systems.
Baseline systems were built using the
aforementioned Matlab toolbox and only trained on
the ArtiPhon training dataset (so they are all
speaker-dependent systems).</p>
      <p>
        Baseline systems used a 48 phone set and
three-state monophones
        <xref ref-type="bibr" rid="ref5">(Canevari et al., 2015)</xref>
        .
Baseline systems were trained and tested
according to all three aforementioned three
scenarios. The articulatory data considered only
refer to x-y positions of 6 coils (the coil
attached to the upper teeth was exluded).
5
      </p>
    </sec>
    <sec id="sec-6">
      <title>Results</title>
      <p>Here I report some of the most relevant results
regarding ISTC and baseline systems.</p>
      <p>Baseline systems and ISTC systems are not
directly comparable as very different
assumptions were made, most importantly they use
different phone sets.</p>
      <p>Additionally, ISCT systems were mainly
concerned with exploring the best performing
systems (created using well-known KALDI
recipes for ASR) and comparing them in the
speaker-dependent and in the
speakerindependent case.</p>
      <p>Baselines systems were created to
investigate the utility of articulatory features in
mismatched speaking styles.
5.1</p>
    </sec>
    <sec id="sec-7">
      <title>ISCT systems</title>
      <p>Here I show results on ISCT systems trained
and tested on the 29 phone set. Table 1 shows
results of the speaker-dependent systems while
Table 2 shows results in the
speakerindependent case.</p>
      <p>
        The results shown in the two tables refer to
the various training and decoding experiments,
see
        <xref ref-type="bibr" rid="ref13">(Rath et al., 2013)</xref>
        for all acronyms
references:
 MonoPhone (mono);
 Deltas + Delta-Deltas (tri1);
 LDA + MLLT (tri2);
 LDA + MLLT + SAT (tri3);
 SGMM2 (sgmm2_4);
 MMI + SGMM2
(sgmm2_4_mmi_b0.14);
 Dan’s Hybrid DNN (tri4-nnet),
 system combination, that is Dan’s DNN +
      </p>
      <p>SGMM (combine_2_1-4);
 Karel’s Hybrid DNN
(dnn4_pretraindbn_dnn);
 system combination that is Karel’s DNN
+ sMBR (dnn4_pretrain-dbn_dnn_1-6).</p>
      <p>
        The most interesting result is that while
DNN-HMM systems outperform GMM-HMM
systems in the speaker-independent case (as
expected), GMM-HMM and more specifically
sub-space GMM-HMM
        <xref ref-type="bibr" rid="ref11 ref12">(Povey et al., 2011)</xref>
        ,
outperform the DNN-based systems in the
speaker dependent case.
      </p>
      <p>
        Another interesting result is that sequence
based training strategies
        <xref ref-type="bibr" rid="ref15">(Vesely et al., 2013)</xref>
        did not produce any improvement over
framebased training strategies.
5.2 Baseline systems – acoustic vs.
articulatory results
The baseline systems addressed the two main
questions that motivated the design of the
ArtiPhon task: (i) how does articulatory
information contribute to phone recognition?; (ii)
how does the phone recognition system
performance vary along the continuum from
hyper-articulated speech to hypo-articulated
speech?
      </p>
      <p>Figure 1 shows phone recognition results of
3 different systems over the 10 degrees of
articulation from hyper-articulated to
hypoarticulated speech.</p>
      <p>The three systems, reflecting the three
aforementioned training-testing scenarios, are:
 phone recognition system that only uses
acoustic feature, specifically mel-filtered
spectra coefficients (MFSCs, scenario 1)
 articulatory phone recognition system
where actual measured articulatory/vocal
tract features (VTFs) are appended to the
input acoustic vector during testing
(scenario 2)
 articulatory phone recognition system
where reconstructed VTFs are appended to
the input acoustic vector during testing
(scenario 3)</p>
      <p>
        The last system reconstructs the articulatory
features using an acoustic-to-articulatory
mapping learned during training (see, e.g.,
        <xref ref-type="bibr" rid="ref6">(Canevari et al., 2013)</xref>
        for details).
      </p>
      <p>All systems used a 48-phone set as in cc.</p>
      <p>One first relevant result is that all systems
performed better at high levels of
hyperarticulation than at “middle” levels (i.e., levels
5-6) which mostly corresponds to the training
condition (Canevari et al., forthcoming). In all
systems performance degraded from hyper- to
hypo-articulated speech.</p>
      <p>Reconstructed VTFs always decrease the
phone error rate. Appending recovered VTFs
to the acoustic feature vector produces a
relative PER reduction that ranges from 4.6% in
hyper-articulated speech, to 5.7% and 5.2% in
middle- and hypo-articulated speech
respectively.</p>
      <p>Actual VTFs provide a relative PER
reduction up to 23.5% in hyper-articulated speech,
whereas, unexpectedly, no improvements are
observed when actual VTFs are used in
middle- and hypo-articulated speech. That might
due to the fact that sessions 1 and 2 of the
MSPKA corpus took place in different days so
EMA coils could be in slightly different
positions.</p>
    </sec>
    <sec id="sec-8">
      <title>Conclusions</title>
      <p>This paper described the ArtiPhon task at
Evalita 2016 and showed and discussed results
of baseline phone recognition systems and of
the submitted systems.
tures and Their Correlates. Cambridge,MA:
MIT Press.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          <string-name>
            <surname>Angelini</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Brugnara</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Falavigna</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Giuliani</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gretter</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Omologo</surname>
            ,
            <given-names>M..</given-names>
          </string-name>
          <year>1994</year>
          .
          <article-title>Speaker Independent Continuous Speech Recognition Using an Acoustic-Phonetic Italian Corpus</article-title>
          .
          <source>In Proceedings of ICSLP. Yokohama</source>
          , Japan.
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          <string-name>
            <surname>Badino</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          <year>2016</year>
          .
          <article-title>Phonetic Context Embeddings for DNN-HMM Phone Recognition</article-title>
          .
          <source>In Proceedings of Interspeech</source>
          . San Francisco, CA.
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          <string-name>
            <surname>Badino</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Canevari</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Fadiga</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Metta</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          <year>2016</year>
          .
          <article-title>Integrating Articulatory Data in Deep Neural Network-based Acoustic Modeling</article-title>
          .
          <source>Computer Speech and Language</source>
          ,
          <volume>36</volume>
          ,
          <fpage>173</fpage>
          -
          <lpage>195</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          <string-name>
            <surname>Browman</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Goldstein</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          <year>1992</year>
          .
          <article-title>Articulatory phonology: an overview</article-title>
          .
          <source>Phonetica</source>
          <volume>49</volume>
          (
          <issue>3-4</issue>
          ),
          <fpage>155</fpage>
          -
          <lpage>180</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          <string-name>
            <surname>Canevari</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Badino</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Fadiga</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          <year>2015</year>
          .
          <article-title>A new Italian dataset of parallel acoustic and articulatory data</article-title>
          .
          <source>Proceedings of Interspeech. Dresden.</source>
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          <string-name>
            <surname>Canevari</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Badino</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Fadiga</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Metta</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          <year>2013</year>
          .
          <article-title>Cross-corpus and cross-linguistic evaluation of a speaker-dependent DNN-HMM ASR sytem using EMA data</article-title>
          .
          <source>In Proceedings of Workshop on Speech Production for Automatic Speech Recognition. Lyon</source>
          , France.
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          <string-name>
            <surname>Cosi</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          <year>2016</year>
          .
          <article-title>Phone Recognition Experiments on ArtiPhon with KALDI</article-title>
          .
          <source>In Proceedings of Third Italian Conference on Computational Linguistics</source>
          (CLiC-it
          <year>2016</year>
          ) &amp;
          <article-title>Fifth Evaluation Campaign of Natural Language Processing and Speech Tools for Italian</article-title>
          .
          <source>Final Workshop (EVALITA</source>
          <year>2016</year>
          ).
          <article-title>Associazione Italiana di Linguistica Computazionale (AILC).</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          <string-name>
            <surname>Huang</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Yu</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Liu</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Gong</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          <year>2014</year>
          .
          <article-title>A Comparative Analytic Study on the Gaussian Mixture and Context Dependent Deep Neural Network Hidden Markov Models</article-title>
          .
          <source>Proceedings of Interspeech. Singapore.</source>
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          <string-name>
            <surname>Jakobson</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Fant</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Halle</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          <year>1952</year>
          .
          <article-title>Preliminaries to Speech Analysis: The Distinctive FeaLeCun</article-title>
          ,
          <string-name>
            <given-names>Y.</given-names>
            ,
            <surname>Bengio</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            , and
            <surname>Hinton</surname>
          </string-name>
          ,
          <string-name>
            <surname>G. E.</surname>
          </string-name>
          <year>2015</year>
          .
          <article-title>Deep Learning</article-title>
          .
          <source>Nature</source>
          ,
          <volume>521</volume>
          ,
          <fpage>436</fpage>
          -
          <lpage>444</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          <string-name>
            <surname>Maddieson</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          <year>1997</year>
          .
          <article-title>Phonetic universals</article-title>
          . In W. Hardcastle, and
          <string-name>
            <given-names>J.</given-names>
            <surname>Laver</surname>
          </string-name>
          ,
          <source>The Handbook of Phonetic Sciences</source>
          . pp.
          <fpage>619</fpage>
          --
          <lpage>639</lpage>
          . Oxford: Blackwell Publishers.
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          <string-name>
            <surname>Povey</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Burget</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Agarwal</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Akyazi</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kai</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ghoshal</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Glembek</surname>
            ,
            <given-names>O.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Goel</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Karafiát</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rastrow</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rosei</surname>
            ,
            <given-names>R. C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Schwarzb</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Thomash</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          <year>2011</year>
          .
          <article-title>The Subspace Gaussian Mixture Model - a Structured Model for Speech Recognition</article-title>
          .
          <source>ComputerSpeech and Language</source>
          .
          <volume>25</volume>
          (
          <issue>2</issue>
          ),
          <fpage>404</fpage>
          -
          <lpage>439</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          <string-name>
            <surname>Povey</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ghoshal</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          , &amp; al.
          <year>2011</year>
          .
          <article-title>The KALDI Speech Recognition Toolkit</article-title>
          .
          <source>Proceedings of ASRU2011.</source>
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          <string-name>
            <surname>Rath</surname>
            ,
            <given-names>S. P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Povey</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Vesely</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Cernocky</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          <year>2013</year>
          .
          <article-title>Improved feature processing for Deep Neural Networks</article-title>
          .
          <source>Proceedings of Interspeech</source>
          , pp.
          <fpage>109</fpage>
          --
          <lpage>113</lpage>
          . Lyon, France.
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          <string-name>
            <surname>Seltzer</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Yu</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Wang</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          <year>2013</year>
          .
          <article-title>An Investigation of Deep Neural Networks for Noise Robust Speech Recognition</article-title>
          .
          <source>Proceedings of ICASSP</source>
          . Vancouver, Canada.
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          <string-name>
            <surname>Vesely</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ghoshal</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Burget</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Povey</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          <year>2013</year>
          .
          <article-title>Sequence-discriminative training of deep neural networks</article-title>
          .
          <source>Proceeding of Interspeech. Lyon</source>
          , France.
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>