<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Development of Visual and Audio Speech Recognition Systems Using Deep Neural Networks</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Denis Ivanko</string-name>
          <email>EMAILdenis.ivanko11@gmail.com</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Dmitry Ryumin</string-name>
          <email>ryumin.d@iias.spb.su</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>St. Petersburg Federal Research Center of the Russian Academy of Sciences (SPC RAS)</institution>
          ,
          <addr-line>14th lin. Vasilievsky Island, 39, St. Petersburg, 199178</addr-line>
          ,
          <country country="RU">Russia</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>In this paper we design end-to-end neural network for the low-resource lip-reading task and audio speech recognition task using 3D CNNs, pre-trained CNN weights of several state-ofthe-art models (e.g. VGG19, InceptionV3, MobileNetV2, etc.) and LSTMs. We present two phrase-level speech recognition pipelines: for lip-reading and acoustic speech recognition. We evaluate different combinations of front-end and back-end modules on the RUSAVIC dataset. We compare our results with traditional 2D CNN approach and demonstrate the increase in recognition accuracy up to 14%. Moreover, we carefully studied existing state-of-the-art models to be use for augmentation. Based on the conducted analysis we have chosen 5 most promising model's architectures and evaluated them on own data. We have tested our systems on a real-word data of two different scenarios: recorded in idling vehicle and during actual driving. Our independently trained systems demonstrated acoustic speech accuracy up to 90% and lip-reading accuracy up to 61%. Future work will focus on the fusion of visual and audio speech modalities and on speaker adaptation. We expect that fused multi-modal information will help to further improve recognition performance compared to a single modality. Another possible direction could be the research of different NN-based architectures to better tackle end-to-end lip-reading task.</p>
      </abstract>
      <kwd-group>
        <kwd>1 Computer vision</kwd>
        <kwd>automated lip-reading</kwd>
        <kwd>speech recognition</kwd>
        <kwd>end-to-end</kwd>
        <kwd>CNN</kwd>
        <kwd>LSTM</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>The ability to use a natural-to-human way of communication greatly improve the interaction quality
of modern computer vision-based assistive systems. Speech is the usual way for humans to
communicate. At the same time, the accuracy and robustness of automatic speech recognition (ASR)
systems is not satisfactory in many practical conditions of use (e.g. in acoustically noisy conditions,
while driving a car or being in a crowded place, etc.). In these cases, the advantage of using visual
information about speech (lip-movements) in addition to audio is undeniable and is used in a number
of state-of-the-art systems.</p>
      <p>In current research, we tried to approach the problem of automatic audio-visual speech recognition
from a computer vision and machine learning perspective. We developed and research two independent
integral (end-to-end) systems for automatic recognition of Russian speech with limited vocabulary
using CNN-based deep neural networks architectures. Moreover, we tried to consider the problem of
acoustic speech recognition as a purely computer vision task by using images of speech spectrograms
in order to train the networks.</p>
      <p>There is no doubt that in recent years the active development of machine learning field has pushed
the results in many other areas with automated lip-reading is no exception. However, despite all the
achieved progress, the development of end-to-end speech recognition systems based on audio and visual
information is still a new direction. Practically no research has been carried out in this field for the
Russian language. There is no out-of-the-box solution accepted by researchers to the development of
such systems. There are no representative open-access datasets for training NN models that have the
required parameters, such as a sufficient number of speakers, phone-viseme labelling, vocabulary size
adequate for the task, etc. (there are almost no public datasets available for languages other than
English). The combination of these factors allows us to state a significant gap in the field of research.
One of the main goals of this study is to bring the recognition efficiency of automatic systems closer to
the level of human speech perception in noisy conditions, which is an extremely important task.</p>
      <p>
        In this paper, we present lip-reading pipeline and acoustic speech recognition pipeline with the use
of deep 3D CNNs. We trained and evaluated our models using RUSAVIC [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ] dataset on a limited
vocabulary of 50 phrases. To handle the over-fitting problem due to the increased number of parameters
from the 3D kernels, we applied the idea from [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ] to inflate the pre-trained weights of the several
stateof-the-art models, such as MobileNetV2 [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ], DenseNet121 [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ], NASNetMobile [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ], etc.
      </p>
    </sec>
    <sec id="sec-2">
      <title>2. Related works</title>
      <p>
        The classical approach towards audio-visual speech recognition involves a two-stage pipeline,
including informative features extraction and classification using the sequence model. Usually the
processes in two stages are independent. The most popular feature extraction approaches were based
on dimension reduction and compression, such as Discrete Cosine Transform [
        <xref ref-type="bibr" rid="ref6 ref7">6, 7</xref>
        ]. Followed in the
second stage by a sequence model (e.g. Hidden Markov Model) to tackle the temporal dependency from
the extracted features for classification [
        <xref ref-type="bibr" rid="ref10 ref8 ref9">8-10</xref>
        ].
      </p>
      <p>
        Although this two-stage pipeline methods have made significant progress over the decades, all such
methods directly separate the feature extraction process from the classifier’s training process, resulting
in the extracted features might not be the optimal for classification. In the recent years, deep learning
approaches have been proposed and achieved the state-of-the-art performance [
        <xref ref-type="bibr" rid="ref11 ref12 ref13 ref14 ref15 ref16">11-16</xref>
        ]. To date, the
classical approach to AV speech recognition have been gradually replaced by the end-to-end trainable
neural networks. In a raw approximation they behave somewhat similar to the traditional methods: a
sequence of the mouth images is fed into the convolutional network to extract the features [
        <xref ref-type="bibr" rid="ref17 ref18">17,18</xref>
        ],
which a further passed to a back-end model (RNN, LSTM, GRU or other) to account for the temporal
dependency for classification [
        <xref ref-type="bibr" rid="ref19 ref20 ref21">19 - 21</xref>
        ]. Since the calculated gradient can be send back from back-end
model to the front-end, the entire network is end-to-end trainable. Recent advances demonstrated that
the learned features are more suitable for speech recognition and lip-reading than the standalone features
calculated by traditional methods [
        <xref ref-type="bibr" rid="ref22">22</xref>
        ].
      </p>
      <p>
        The major advantage of modern approach is that entire system consists of an end-to-end trainable
front-end and back-end neural network (so two-stage process no longer exists). Thus, the learned
features are more connected to the task that the network is trained on. The first work which proposed
to use the CNNs to replace the independent features extractor was [
        <xref ref-type="bibr" rid="ref23">23</xref>
        ]. In turn, the first work that
proposed to use the LSTM for classification and achieved a significant improvement was [
        <xref ref-type="bibr" rid="ref24">24</xref>
        ]. Other
researchers in [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ] proposed to take advantages of large-scale lip-reading dataset to train a front-end
followed by the LSTM module at the end for classification. The researchers in [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ] proposed a neural
network to extract the audio features and tried to fuse them with video information.
      </p>
      <p>
        It is generally accepted that visual features extracted from images by 2D CNNs are suitable for some
computer vision tasks (e.g. image classification, lip-reading, gesture recognition etc.) [
        <xref ref-type="bibr" rid="ref25 ref26">25,26</xref>
        ]. However,
it is more natural to learn spatio-temporal features by using a 3D CNNs as the front-end for feature
extraction. Nonetheless, according to our knowledge only a few works have researched the use of 3D
CNN for lip-reading [
        <xref ref-type="bibr" rid="ref15 ref16 ref27 ref28">15, 16, 27, 28</xref>
        ]. In addition, these works usually apply only a shallow version of
3D CNNs with no more than 3 convolutional layers. Obviously, this approach contradicts to the
common rule that a deep network is expected to do better than a shallow one. Thus, the question of how
to train a deep neural network without over-fitting on standard audio-visual datasets is still open and
practically not studied.
      </p>
      <p>
        Convolutional neural network architectures have been developed for image and video processing for
a long time. In the work [
        <xref ref-type="bibr" rid="ref29">29</xref>
        ] an approach to extract features independently from each frame using a 2D
CNN have been proposed to re-use the pre-trained weights of ImageNet [
        <xref ref-type="bibr" rid="ref30">30</xref>
        ] model. In the work [
        <xref ref-type="bibr" rid="ref31">31</xref>
        ]
the 3D CNNs for video action recognition have been introduced as a natural extension of the 2D
convolution. The researchers in the work [
        <xref ref-type="bibr" rid="ref32">32</xref>
        ] went beyond using the shallow CNN version and explored
the deep 3D CNN version with replacing all 2D operations with their 3D counterparts.
      </p>
    </sec>
    <sec id="sec-3">
      <title>3. Data &amp; Preprocessing</title>
      <p>
        Almost no publicly-accessible audio-visual Russian speech datasets are available and suitable for
NN training. The most recent one was introduced in the work [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ] and was specifically designed for the
task of robust speech recognition in acoustically-noisy car environment.
      </p>
      <p>
        The multi-speaker audio-visual corpus RUSAVIC (RUSsian Audio-Visual speech In Cars) includes
a continuous Russian speech with multi-angle video and audio data. It contains recordings of 20 native
Russian speakers. The database stores audio and video recordings of Russian speech, as well as labelling
information. The recording and labelling of audio-visual data was carried out using the created software
package [
        <xref ref-type="bibr" rid="ref33">33</xref>
        ] designed to capture, synchronize and combine audio and video data from two or more
smartphones located in the vehicle cabin. Recording of the speech corpus was carried out both in traffic
conditions and in the idling of a vehicle, i.e. in full-scale and semi-natural conditions, as close as
possible to the real conditions of functioning.
      </p>
      <p>Each speaker performed 10 recording sessions and was captured by three smartphones from three
different angles with FullHD 1920 × 1080 video resolution and 60 fps recording rate. During each
recording session speaker uttered 50 phrases, which are the most frequent driver requests for
smartphones (according to open source data of several state-of-the-art speech recognition engines, such
as AlexaAuto, YandexDrive, GoogleDrive, etc.). The basic structure of the corpus is depicted in the
Figure 1, and some snapshots of the speakers during the recordings are shown in Figure 2.
3.1.</p>
    </sec>
    <sec id="sec-4">
      <title>Visual data</title>
      <p>
        Detecting a region-of-interest (ROI) that contains the mouth motion is the first and very important
step in building a reliable automated lip-reading system. Thus, our first target is to crop this ROI (mouth
region) from each frame of the video. To this end we applied the state-of-the-art solution of MediaPipe
Face Mesh [
        <xref ref-type="bibr" rid="ref34">34</xref>
        ] that is able to estimate 468 3D face landmarks.
      </p>
      <p>Face Mesh employs machine learning to infer the 3D surface geometry and provides real-time
performance critical for real-life speech recognition scenarios. The general pipeline consists of two
deep neural network models working together: (1) A detector that computes face locations (operates on
the full image) and (2) a 3D face landmark model (operates on the detected locations) that predicts the
approximate surface geometry via regression. The basic structure of this pipeline depicted in the Figure
3.</p>
      <p>
        In addition, the mouth region crops can also be generated based on the face landmarks identified in
the previous frame and only when the Model could no longer detect face presence the face detector is
invoked to relocalize the face region. The pipeline is implemented as a graph that uses face landmark
subgraph from the Face detection Module (Fig. 3, left) and renders using face renderer subgraph. We
used the same BlazeFace detector as in original work [
        <xref ref-type="bibr" rid="ref35">35</xref>
        ]. The 3D face landmark model employed
transfer learning and was trained with several objectives: it simultaneously predicts 3D landmark
coordinates on synthetic rendered data and 2D semantic contours on annotated real-word data.
      </p>
      <p>The 3D landmark network (Fig. 3, stage 2) receives as input a cropped frame and outputs the
positions of the 3D points, as well as the probability of a face being present and aligned in the input.
The Face landmark module performs a face landmark detection in the screen coordinate space, where
the X- and Y- coordinates are normalized screen coordinates. Example of the detected 468 face
landmarks is shown in the Figure 3, right.</p>
    </sec>
    <sec id="sec-5">
      <title>Acoustic data</title>
      <p>
        One of the first works, that treated raw acoustic signal as an image was [
        <xref ref-type="bibr" rid="ref36">36</xref>
        ]. The authors proved
that between the first two convolutional layers, the CNN learns (in parts) and models the phone-specific
spectral envelope information of 2-4 ms speech. They demonstrated advantages of using the
CNNbased approach to yield ASR performance.
      </p>
      <p>
        In current research we handle acoustic speech processing by obtaining spectrograms of the uttered
phrases from the raw audio data with its further processing by the integral CNN-LSTM network. We
implement this using librosa library [
        <xref ref-type="bibr" rid="ref37">37</xref>
        ]. It is a python package for music and audio analysis. It is
structured as collection of submodules. A spectrogram is calculated by computing the fast fourier
transform (FFT) over a series of overlapping windows extracted from the raw audio signal. The process
of dividing the signal in short term sequences of fixed size and applying FFT on those independently is
called Short-time Fourier transform (STFT). The spectrogram is then calculated as the squared complex
magnitude of the STFT. The general process of calculating spectrogram from the raw acoustic signal is
depicted in the Figure 4.
      </p>
    </sec>
    <sec id="sec-6">
      <title>4. Proposed methodology</title>
      <p>End-to-end approach to automatic speech recognition assumes the training of only one neural
network that combines all the stages of the traditional approach. At the same time, this presupposes the
presence of certain structural blocks of the network, which we divide into four sequential processing
stages:
1. Inputs, which is a sequence of cropped mouth images in case of lip-reading or a spectrogram
images in case of acoustic speech recognition.
2. Front-end-module, to extract features from the inputs. We used a 3-4 3D CNN layers for the
visual features extraction in the lip-reading system and a number of pre-trained CNNs for acoustic
speech recognition.
3. Back-end module, to model the temporal dependency and summarize the features into a single
vector that represents the score for each phrase.
4. Classification module, to compute the probabilities of each phrase. In both systems represented
by a softmax layer.</p>
      <p>The most of the existing end-to-end systems fall into this structure. In this paper we focused on the
inputs (visual and audio data preprocessing) described in Section 3 and front-end modules which will
be introduced in the following sections.</p>
    </sec>
    <sec id="sec-7">
      <title>Building automated lip-reading system</title>
      <p>
        General network architecture of our 3D CNN-based automated lip-reading system is presented in
the figure 5. We compare our lipreading results with the recent work [
        <xref ref-type="bibr" rid="ref33">33</xref>
        ] that used deep 2D CNNs,
which were originally proposed to solve image-base tasks. The general approach with applying 2D
CNNs on lip-reading data is to concatenate the features independently extracted from each frame. On
the other hand, 3D convolution can process the dynamics (at least short-term dynamics) and is proven
to be useful in many other computer vision-related tasks followed by the recurrent network at the
backend. However, due to the difficulty of training a vast number of parameters introduced by the
threedimensional kernel in current research we explore only 3 to 4 layers network with 3D convolution.
      </p>
      <p>Cropped mouth frames sequences are first normalized to the size of 224×224 and then split into
batches of 30 frames with 50% overlap (15 frames) before fed into the network. On all 3D CNN layers,
we use three-dimensional kernel, followed by the batch normalization, Rectified linear units and 3D
max-pooling. Specifically, in case of 3-layer network the number of kernels were 32, 64 and 128
respectively for each layer. In case of 4-year network the number of kernels were 32, 32, 64 and 128
respectively for each layer. The front-end visual features extraction part of our model ends with one
densely connected layer with 512 neurons in it.</p>
      <p>The back-end of the model consists of 2 (Long-short term memory) LSTM layers. LSTM is a type
of recurrent neural networks, which are well-known for the ability to model temporal dependency and
are typical back-end modules used in many computer vision and speech recognition tasks. Among
RNNs, LSTM is proven to be useful when dealing with the exploding and vanishing gradient problem
[38]. Specifically, we use a two-layer LSTM with a hidden state dimension of 512 for each cell in the
first layer and 256 in the second, followed by 50 phrases classificatory represented by densely connected
linear layer.</p>
      <p>We trained and evaluated the proposed network on the phrase-level lip-reading dataset RUSAVIC.
The number of target phrases is 50. We took 8 repetition of each phrase for the training and 2 for the
testing for each speaker. Hence, the network has learned to discriminate between 50 target phrases
based purely on the lip movements information.</p>
    </sec>
    <sec id="sec-8">
      <title>Building acoustic speech recognition system</title>
      <p>Basic architecture of the end-to-end 2D CNN spectrogram-based acoustic speech recognition system
is depicted in the Figure 6. We preprocess the raw acoustic data and obtain phrase-level spectrograms
in accordance with the pipeline presented in the section 3.2. This step is followed by spectrogram
normalization (we tested 2 types of input dimensions 224×224 and 299×299 depending on the
pretrained model).</p>
      <p>
        Pre-trained weights are proven to be useful in many image-based tasks. Therefore, we tried to get
the best use of modern transfer learning approaches and applied five different pre-trained deep CNN
architectures, namely VGG19 [39], InceptionV3 [40], MobileNetV2 [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ], DenseNet121 [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ] and
NASNetMobile [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ].
      </p>
      <p>VGG19 is 26-layer deep convolutional neural network with &gt;143 million of trainable parameters.
The default input size for this model is 224×224. The model is trained for the large-scale image
recognition scenarios.</p>
      <p>InceptionV3 is 159-layer deep network with &gt;23 million of trainable parameters, developed
specifically for mobile vision scenarios and big-data scenarios.</p>
      <p>MobileNetV2 is 88-layer CNN with &gt;3.5 million of trainable parameters. It uses inverted residual
blocks with bottlenecking features and has a drastically lower parameter count than the original
MobileNet. MobileNets support any input size greater than 32×32, with larger image sizes offering
better performance.</p>
      <p>DenseNet121 is 121-layer network with &gt;8 million of trainable parameters. DenseNets have several
compelling advantages: they alleviate the vanishing-gradient problem, strengthen feature propagation,
encourage feature reuse, and substantially reduce the number of parameters.</p>
      <p>NASNetMobile has &gt;5 million of trainable parameters. It is a scalable architecture for image
classification and consist of two repeated building blocks termed Normal Cell and Reduction Cell. In
current research we applied the latest 769 layers architecture.</p>
      <p>The layers with pre-trained weights are followed by a 50 neuron softmax classification layers, that
provides final recognition result.</p>
    </sec>
    <sec id="sec-9">
      <title>5. Evaluation experiments</title>
      <p>In this section we evaluate and compare the proposed architectures on different speakers of the
RUSAVIC dataset. Maximum number of epochs was 50 and training was interrupted if the accuracy
does not increase for 5 epochs. For each speaker the train and test data were splitted into 80 : 20 percent
ratio. In total we trained a three speaker-dependent lip-reading system and three acoustic speech
recognition systems, based on the available amount of data.</p>
      <p>We summarize the lip-reading recognition results in the Table 1 and acoustic (spectrogram-based)
speech recognition results in the Table 2. The first two systems (ID1 and ID2) were trained on the data,
recorded in the idle vehicle, parked on the busy crossroads. The system #3 was trained with the actual
driving data.</p>
      <p>
        According to the table 1, the 3D CNN-based architecture clearly outperforms the traditional CNN
for lip-reading both: in vehicle idling conditions 61% vs (47 to 55 %) and 57% vs (46 to 53 %) accuracy
on the phrase-level recognition, and driving conditions 59% vs (51 to 54 %) accuracy. The 2D CNN
results were from the recent paper [
        <xref ref-type="bibr" rid="ref33">33</xref>
        ] evaluating speaker-dependent recognition systems on
RUSAVIC dataset. Thus, in all systems 3D CNNs demonstrated significant improvements in the terms
of recognition accuracy. Interestingly, despite the fact that driving requires rather active head turns we
did not found much difference in recognition accuracy between the models trained on driving data
versus models trained on data recorded in a parked vehicle.
      </p>
      <p>Another interesting finding was that using a slightly deeper 3D CNN (from 3 to 4
spatioconvolutional layers) results in increasing recognition accuracy up to 3% absolute. However, due to the
limited amount of Russian lip-reading data available, further increase in network’s depth does not lead
to further improvements of recognition accuracy.</p>
      <p>It can be seen from the Table 2, that spectrogram-based acoustic speech recognition generally
performed better than the lip-reading. These results are within expectation range since acoustic
information usually convey much more speech-related information that the lips movements. We
achieved the maximum result of 90% recognition accuracy on the speaker #2 using pre-trained weights
of the VGG19 model. In turn, the lowest recognition results were demonstrated by the model with
NASNetMobile pre-trained weights (59%), which was trained on the driving data.</p>
      <p>In addition to that, we perform experimental study and assess several state-of-the-art model
architectures in order to research which of them provides better pre-trained weights for the task of
automated speech recognition with using spectrograms as the network input. According to the obtained
results, the most suitable for this task was VGG19 model, that achieved from 79 to 90% recognition
accuracy. On the other hand, the lowest recognition results demonstrated NASNetMobile architecture,
with only 59 to 61% speech recognition accuracy on all three systems. These results are almost the
same that the one achieved by the 3D CNN lip-reading system with four spatio-temporal layers.</p>
      <p>However, the advantage of VGG19 model is easily explained by fact that it has more than 143
million of trainable parameters, when the NASNetMobile only provides slightly more than 5 million
trainable parameters. Thus, it is natural that VGG19 generalized better on the provided lip-reading data,
since it was initially trained on much bigger amount of visual data. However, the disadvantage of using
this architecture might be its resource-costly to the computational power of the device. E.g. it is not
optimal to use it on smartphones or similar resource-dependent devices.
ID</p>
      <sec id="sec-9-1">
        <title>Architecture</title>
      </sec>
      <sec id="sec-9-2">
        <title>Number of 3DCNN layers</title>
      </sec>
      <sec id="sec-9-3">
        <title>Recognized classes Accuracy, % (epoch)</title>
        <p>1) 512 neurons with L2
regularization = 0,001
2) 256 neurons with L2
regularization = 0,001
1) 512 neurons with L2
regularization = 0,001
2) 256 neurons with L2
regularization = 0,001
50
58 (21)
61 (17)
47-55
55 (20)
57 (15)
46-53
56 (39)
59 (34)</p>
      </sec>
    </sec>
    <sec id="sec-10">
      <title>6. Conclusions</title>
      <p>In this paper we have successfully demonstrated the capability and feasibility of designing an
endto-end neural network for the low-resource lip-reading and audio speech recognition task using 3D
CNNs, pre-trained CNN weights of several state-of-the-art models (e.g. VGG19, InceptionV3,
MobileNetV2, etc.) and LSTMs. We were able to achieve a state-of-the-art accuracy of 90 % for
acoustic speech and 61% for lip-reading with 50 recognizable classes. To the best of our knowledge
current research is one of the first attempts to work with Russian audio-visual speech.</p>
      <p>We presented two phrase-level speech recognition pipelines: for lip-reading and acoustic speech
recognition. We evaluated different combinations of front-end and back-end modules on the RUSAVIC
dataset. We compared our results with traditional 2D CNN approach and demonstrated that even
shallow 3 to 4 spatio-convolution layer network can outperform traditional approach up to 14 %
recognition accuracy. Moreover, we carefully studied existing state-of-the-art models to be used for
augmentation and transfer learning in the field of image processing and computer vision. Based on the
conducted analysis we have chosen 5 most promising model’s architectures and provided recognition
results for each.</p>
      <p>In the current research, we have studied Russian audio-visual speech from a computer vision
perspective. We have tested our systems on real-word data of two different scenarios: idling vehicle
and actual driving. Our independently trained systems demonstrated acoustic speech accuracy up to
90% and lip-reading accuracy up to 61%. Future work will focus on the fusion of visual and audio
speech recognition systems and on speaker adaptation. We expect that fused multi-modal information
will help to further improve recognition performance compared to a single modality.</p>
    </sec>
    <sec id="sec-11">
      <title>7. Acknowledgements</title>
      <p>This research is financially supported by the Russian Science Foundation (project No. 21-71-00132).</p>
    </sec>
    <sec id="sec-12">
      <title>8. References</title>
      <p>[38] X. Weng, K. Kitani: Learning spatio-temporal features with two-stream deep 3D CNNs for
lipreading. arXiv preprint (2019). arXiv:1905.02540.
[39] K. Simonyan, A. Zisserman: Very deep convolutional networks for large-scale image recognition.</p>
      <p>In arXiv preprint (2019). arXiv:1409.1556.
[40] C. Szegedy, V. Vanhoucke, S. Ioffe, J. Shlens, Z. Wojna: Rethinking the inception architecture for
computer vision. In Proceedings of the IEEE conference on computer vision and pattern
recognition (2016) 2818-2826. doi: 10.1109/CVPR.2016.308.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>A.</given-names>
            <surname>Kashevnik</surname>
          </string-name>
          et al.:
          <article-title>Multimodal Corpus Design for Audio-Visual Speech Recognition in Vehicle Cabin</article-title>
          .
          <source>In IEEE Access</source>
          , vol.
          <volume>9</volume>
          (
          <year>2021</year>
          )
          <fpage>34986</fpage>
          -
          <lpage>35003</lpage>
          . doi:
          <volume>10</volume>
          .1109/ACCESS.
          <year>2021</year>
          .
          <volume>3062752</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>J.</given-names>
            <surname>Carreira</surname>
          </string-name>
          ,
          <string-name>
            <surname>A.</surname>
          </string-name>
          <article-title>Zisserman: Quo Vadis, Action Recognition? A New Model and the Kinetics Dataset</article-title>
          .
          <source>In proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR)</source>
          (
          <year>2017</year>
          )
          <fpage>6299</fpage>
          -
          <lpage>6308</lpage>
          . doi:
          <volume>10</volume>
          .1109/CVPR.
          <year>2017</year>
          .
          <volume>502</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>A.</given-names>
            ,
            <surname>Howard</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Zhu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Chen</surname>
          </string-name>
          , et al.:
          <article-title>MobileNets: Efficient Convolutional Neural Networks for Mobile Vision Applications</article-title>
          . In arXiv:
          <volume>1704</volume>
          .04861, pp.
          <fpage>1</fpage>
          -
          <lpage>9</lpage>
          (
          <year>2017</year>
          ). arXiv:
          <volume>1704</volume>
          .
          <fpage>04861</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>G.</given-names>
            <surname>Huang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Liu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L. Van Der</given-names>
            <surname>Maaten</surname>
          </string-name>
          ,
          <string-name>
            <surname>K.</surname>
          </string-name>
          <article-title>Weinberger: Densely Connected Convolutional Networks</article-title>
          .
          <source>In IEEE Conference on Computer Vision and Pattern Recognition (CVPR)</source>
          (
          <year>2018</year>
          )
          <fpage>2261</fpage>
          -
          <lpage>2269</lpage>
          . doi:
          <volume>10</volume>
          .1109/CVPR.
          <year>2017</year>
          .
          <volume>243</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>Z.</given-names>
            <surname>Barret</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Vijay</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Jonathon</surname>
          </string-name>
          ,
          <string-name>
            <surname>L.</surname>
          </string-name>
          <article-title>Quoc: Learning Transferable Architectures for Scalable Image Recognition</article-title>
          .
          <source>In Conference on Computer Vision and Pattern Recognition (CVPR)</source>
          (
          <year>2018</year>
          )
          <fpage>8697</fpage>
          -
          <lpage>8710</lpage>
          . doi:
          <volume>10</volume>
          .1109/CVPR.
          <year>2018</year>
          .
          <volume>00907</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>A.</given-names>
            <surname>Pass</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Zhang</surname>
          </string-name>
          ,
          <string-name>
            <surname>D.</surname>
          </string-name>
          <article-title>Stewart: An Investigation into Features for Multi-View Lipreading</article-title>
          .
          <source>In 2010 IEEE International Conference on Image Processing</source>
          (
          <year>2010</year>
          )
          <fpage>2417</fpage>
          -
          <lpage>2420</lpage>
          . doi:
          <volume>10</volume>
          .1109/ICIP.
          <year>2010</year>
          .
          <volume>5650963</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>X.</given-names>
            <surname>Hong</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Yao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Wan</surname>
          </string-name>
          ,
          <string-name>
            <surname>R.</surname>
          </string-name>
          <article-title>Chen: A PCA Based Visual DCT Feature Extraction Method for Lipreading</article-title>
          .
          <source>International Conference on Intelligent Information Hiding and Multimedia Signal Processing</source>
          , (
          <year>2006</year>
          )
          <fpage>321</fpage>
          -
          <lpage>326</lpage>
          . doi:
          <volume>10</volume>
          .1109/IIH-MSP.
          <year>2006</year>
          .
          <volume>265008</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>V.</given-names>
            <surname>Estellers</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Gurban</surname>
          </string-name>
          , J. Thiran:
          <article-title>On Dynamic Stream Weighting for Audio-Visual Speech Recognition</article-title>
          .
          <source>IEEE Transactions on Audio, Speech, and Language Processing</source>
          ,
          <volume>20</volume>
          ,
          <issue>4</issue>
          (
          <year>2012</year>
          )
          <fpage>1145</fpage>
          -
          <lpage>1157</lpage>
          . doi:
          <volume>10</volume>
          .1109/TASL.
          <year>2011</year>
          .
          <volume>2172427</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>D.</given-names>
            <surname>Ivanko</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Karpov</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Ryumin</surname>
          </string-name>
          , et al.:
          <article-title>Using a high-speed video Camera for robust audio-visual speech recognition in acoustically noisy conditions</article-title>
          .
          <source>In International Conference on Speech and Computer</source>
          , (
          <year>2017</year>
          )
          <fpage>757</fpage>
          -
          <lpage>766</lpage>
          . doi:
          <volume>10</volume>
          .1007/978-3-
          <fpage>319</fpage>
          -66429-3_
          <fpage>76</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>D.</given-names>
            <surname>Stewart</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Seymour</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Pass</surname>
          </string-name>
          ,
          <string-name>
            <surname>J.</surname>
          </string-name>
          <article-title>Ming: Robust Audio-Visual Speech Recognition under Noisy Audio-Video Conditions</article-title>
          .
          <source>IEEE transactions on cybernetics</source>
          ,
          <volume>44</volume>
          ,
          <issue>2</issue>
          (
          <year>2014</year>
          )
          <fpage>175</fpage>
          -
          <lpage>184</lpage>
          . doi:
          <volume>10</volume>
          .1109/TCYB.
          <year>2013</year>
          .
          <volume>2250954</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>A.</given-names>
            <surname>Zisserman</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Chung</surname>
          </string-name>
          .
          <article-title>Lip Reading in the Wild</article-title>
          .
          <source>In Asian conference on computer vision</source>
          , (
          <year>2016</year>
          )
          <fpage>87</fpage>
          -
          <lpage>103</lpage>
          . doi:
          <volume>10</volume>
          .1007/978-3-
          <fpage>319</fpage>
          -54184-
          <issue>6</issue>
          _
          <fpage>6</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>H.</given-names>
            <surname>Ninomiya</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Kitaoka</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Tamura</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Iribe</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Takeda</surname>
          </string-name>
          .
          <article-title>Integration of Deep Bottleneck Features for Audio-Visual Speech Recognition. In 16th annual conference of the international speech communication association</article-title>
          , (
          <year>2015</year>
          )
          <fpage>575</fpage>
          -
          <lpage>582</lpage>
          . doi:
          <volume>10</volume>
          .1109/APSIPA.
          <year>2015</year>
          .
          <volume>7415335</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <given-names>D.</given-names>
            <surname>Ivanko</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Karpov</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Fedotov</surname>
          </string-name>
          et al.:
          <article-title>Multimodal speech recognition: increasing accuracy using high speed video data</article-title>
          .
          <source>Journal of Multimodal User Interfaces</source>
          , vol.
          <volume>12</volume>
          , (
          <year>2018</year>
          )
          <fpage>319</fpage>
          -
          <lpage>328</lpage>
          . doi:
          <volume>10</volume>
          .1007/s12193-018-0267-1.
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <given-names>D.</given-names>
            <surname>Ivanko</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Ryumin</surname>
          </string-name>
          ,
          <string-name>
            <surname>A.</surname>
          </string-name>
          <article-title>Karpov: An Experimental Analysis of Different Approaches to AudioVisual Speech Recognition</article-title>
          and Lip-Reading.
          <source>In Proceedings of 15th International Conference on Electromechanics and Robotics</source>
          , Singapore, (
          <year>2021</year>
          )
          <fpage>197</fpage>
          -
          <lpage>209</lpage>
          . doi:
          <volume>10</volume>
          .1007/
          <fpage>978</fpage>
          -981-15-5580-0_
          <fpage>16</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <given-names>S.</given-names>
            <surname>Petridis</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Stafylakis</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Ma</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Cai</surname>
          </string-name>
          , G. Tzimiropoulos M.
          <article-title>Pantic: End-to-End Audiovisual Speech Recognition</article-title>
          .
          <source>In IEEE international conference on acoustics, speech and signal processing (ICASSP)</source>
          (
          <year>2018</year>
          )
          <fpage>6548</fpage>
          -
          <lpage>6552</lpage>
          . doi:
          <volume>10</volume>
          .1109/ICASSP.
          <year>2018</year>
          .
          <volume>8461326</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [16]
          <string-name>
            <given-names>G.</given-names>
            <surname>Tzimiropoulos</surname>
          </string-name>
          ,
          <string-name>
            <surname>T.</surname>
          </string-name>
          <article-title>Stafylakis: Combining Residual Networks with LSTMs for Lipreading</article-title>
          .
          <source>In arXiv preprint arXiv:1703</source>
          .
          <fpage>04105</fpage>
          . (
          <year>2017</year>
          ). doi:
          <volume>10</volume>
          .21437/INTERSPEECH.2017-
          <volume>85</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [17]
          <string-name>
            <given-names>B.</given-names>
            <surname>Xu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Lu</surname>
          </string-name>
          ,
          <string-name>
            <surname>Y.</surname>
          </string-name>
          <article-title>Guo: Watch to Listen Clearly: Visual Speech Enhancement Driven Multi-modality Speech Recognition</article-title>
          .
          <source>In Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision</source>
          , (
          <year>2020</year>
          )
          <fpage>1637</fpage>
          -
          <lpage>1646</lpage>
          . doi:
          <volume>10</volume>
          .1109/WACV45572.
          <year>2020</year>
          .
          <volume>9093314</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          [18]
          <string-name>
            <given-names>E.</given-names>
            <surname>Ryumina</surname>
          </string-name>
          et al.:
          <article-title>A Novel Method for Protective Face Mask Detection Using Convolutional Neural Networks and Image Histograms</article-title>
          .
          <source>In ISPRS-International Archives of the Photogrammetry, Remote Sensing and Spatial Information Sciences</source>
          ,
          <volume>4421</volume>
          (
          <year>2021</year>
          ) pp.
          <fpage>177</fpage>
          -
          <lpage>182</lpage>
          . doi:
          <volume>10</volume>
          .5194/isprsarchives-XLIV-2
          <string-name>
            <surname>-W1-</surname>
          </string-name>
          2021-177-
          <year>2021</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          [19]
          <string-name>
            <given-names>X.</given-names>
            <surname>Chen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Du</surname>
          </string-name>
          .,
          <string-name>
            <surname>H.</surname>
          </string-name>
          <article-title>Zhang: Lipreading with DenseNet and resBi-LSTM, Signal</article-title>
          ,
          <source>Image and Video Processing</source>
          ,
          <volume>14</volume>
          (
          <issue>5</issue>
          ) (
          <year>2020</year>
          )
          <fpage>981</fpage>
          -
          <lpage>989</lpage>
          . doi:
          <volume>10</volume>
          .1007/s11760-019-01630-1.
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          [20]
          <string-name>
            <given-names>E.</given-names>
            <surname>Ryumina</surname>
          </string-name>
          ,
          <string-name>
            <surname>A</surname>
          </string-name>
          . Karpov:
          <article-title>Facial expression recognition using distance importance scores between facial landmarks</article-title>
          .
          <source>CEUR Workshop Proceedings</source>
          , vol.
          <volume>2744</volume>
          . (
          <year>2020</year>
          )
          <fpage>1</fpage>
          -
          <lpage>10</lpage>
          . doi:
          <volume>10</volume>
          .51130/graphicon-2020-2-3-32.
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          [21]
          <string-name>
            <given-names>D.</given-names>
            <surname>Ivanko</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Ryumin</surname>
          </string-name>
          ,
          <string-name>
            <surname>I. Kipyatkova</surname>
          </string-name>
          et al.:
          <article-title>Lip-Reading Using pixel-based and geometry-based features for multimodal human-robot interfaces</article-title>
          .
          <source>In Proceedings of 14th International Conference on Electromechanics and Robotics</source>
          , Singapore, (
          <year>2020</year>
          )
          <fpage>477</fpage>
          -
          <lpage>486</lpage>
          . doi:
          <volume>10</volume>
          .1007/
          <fpage>978</fpage>
          -981-13-9267- 2_
          <fpage>39</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          [22]
          <string-name>
            <given-names>I.</given-names>
            <surname>Fung</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Mak</surname>
          </string-name>
          :
          <article-title>End-to-end low-resource lip-reading with maxout CNN and LSTM</article-title>
          .
          <source>In IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)</source>
          (
          <year>2018</year>
          )
          <fpage>2511</fpage>
          -
          <lpage>2515</lpage>
          . doi:
          <volume>10</volume>
          .1109/ICASSP.
          <year>2018</year>
          .
          <volume>8462280</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          [23]
          <string-name>
            <given-names>K.</given-names>
            <surname>Noda</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Yamaguchi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Nakadai</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Okuno</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Ogata</surname>
          </string-name>
          .
          <article-title>Lipreading Using Convolutional Neural Network</article-title>
          .
          <source>In proceedings of the Annual Conference of the International Speech Communication Association</source>
          , (
          <year>2014</year>
          )
          <fpage>1149</fpage>
          -
          <lpage>1153</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          [24]
          <string-name>
            <given-names>S.</given-names>
            <surname>Hochreiter</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Schmidhuber</surname>
          </string-name>
          . Long
          <string-name>
            <surname>Short-Term Memory</surname>
          </string-name>
          .
          <source>Neural computation</source>
          ,
          <volume>9</volume>
          (
          <issue>8</issue>
          ) (
          <year>1997</year>
          )
          <fpage>1735</fpage>
          -
          <lpage>1780</lpage>
          . doi:
          <volume>10</volume>
          .1162/neco.
          <year>1997</year>
          .
          <volume>9</volume>
          .8.1735.
        </mixed-citation>
      </ref>
      <ref id="ref25">
        <mixed-citation>
          [25]
          <string-name>
            <given-names>D.</given-names>
            <surname>Ivanko</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Ryumin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Axyonov</surname>
          </string-name>
          ,
          <string-name>
            <surname>M.</surname>
          </string-name>
          <article-title>Železný: Designing Advanced Geometric Features for Automatic Russian Visual Speech Recognition</article-title>
          .
          <source>In International Conference on Speech and Computer</source>
          , pp.
          <fpage>245</fpage>
          -
          <lpage>254</lpage>
          (
          <year>2018</year>
          ). doi:
          <volume>10</volume>
          .1007/978-3-
          <fpage>319</fpage>
          -99579-3_
          <fpage>26</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref26">
        <mixed-citation>
          [26]
          <string-name>
            <given-names>D.</given-names>
            <surname>Ivanko</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Ryumin</surname>
          </string-name>
          :
          <string-name>
            <given-names>A Novel</given-names>
            <surname>Task-Oriented Approach Toward Automated</surname>
          </string-name>
          Lip-Reading System Implementation.
          <source>International Archives of the Photogrammetry, Remote Sensing and Spatial Information Sciences</source>
          ,
          <volume>44</volume>
          (
          <issue>2</issue>
          ) (
          <year>2021</year>
          )
          <fpage>85</fpage>
          -
          <lpage>89</lpage>
          . doi:
          <volume>10</volume>
          .5194/isprs-archives-XLIV-2
          <string-name>
            <surname>-W1-</surname>
          </string-name>
          2021-85-
          <year>2021</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref27">
        <mixed-citation>
          [27]
          <string-name>
            <given-names>T.</given-names>
            <surname>Afouras</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Chung</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Senior</surname>
          </string-name>
          ,
          <string-name>
            <given-names>O.</given-names>
            <surname>Vinyals</surname>
          </string-name>
          ,
          <string-name>
            <surname>A.</surname>
          </string-name>
          <article-title>Zisserman: Deep Audio-Visual Speech Recognition</article-title>
          .
          <source>IEEE transactions on pattern analysis and machine intelligence</source>
          (
          <year>2018</year>
          ). doi:
          <volume>10</volume>
          .1109/TPAMI.
          <year>2018</year>
          .
          <volume>2889052</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref28">
        <mixed-citation>
          [28]
          <string-name>
            <given-names>Y. M.</given-names>
            <surname>Assael</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Shillingford</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Whiteson</surname>
          </string-name>
          , N. De Freitas.
          <article-title>LipNet: End-to-End Sentence-Level Lipreading</article-title>
          .
          <source>arXiv preprint arXiv:1611.01599</source>
          ,
          <issue>2</issue>
          (
          <issue>4</issue>
          ) (
          <year>2017</year>
          ). arXiv:
          <volume>1611</volume>
          .
          <fpage>01599</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref29">
        <mixed-citation>
          [29]
          <string-name>
            <given-names>A.</given-names>
            <surname>Karpathy</surname>
          </string-name>
          , G. Toderici,
          <string-name>
            <given-names>S.</given-names>
            <surname>Shetty</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Leung</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Sukthankar</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Li</surname>
          </string-name>
          :
          <article-title>Large-Scale Video Classification with Convolutional Neural Networks</article-title>
          .
          <source>In Proceedings of the IEEE conference on Computer Vision</source>
          and Pattern
          <string-name>
            <surname>Recognition</surname>
          </string-name>
          (
          <year>2014</year>
          )
          <fpage>1725</fpage>
          -
          <lpage>1732</lpage>
          . doi:
          <volume>10</volume>
          .1109/CVPR.
          <year>2014</year>
          .
          <volume>223</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref30">
        <mixed-citation>
          [30]
          <string-name>
            <given-names>A.</given-names>
            <surname>Krizhevsky</surname>
          </string-name>
          , I. Sutskever,
          <string-name>
            <given-names>G.E.</given-names>
            <surname>Hinton</surname>
          </string-name>
          .:
          <article-title>Imagenet Classification with Deep Convolutional Neural Networks</article-title>
          .
          <source>Advances in neural information processing systems</source>
          ,
          <volume>25</volume>
          , (
          <year>2012</year>
          )
          <fpage>1097</fpage>
          -
          <lpage>1105</lpage>
          . doi:
          <volume>10</volume>
          .1145/3065386.
        </mixed-citation>
      </ref>
      <ref id="ref31">
        <mixed-citation>
          [31]
          <string-name>
            <given-names>D.</given-names>
            <surname>Tran</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Bourdev</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Fergus</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Torresani</surname>
          </string-name>
          ,
          <string-name>
            <surname>M.</surname>
          </string-name>
          <article-title>Paluri: Learning Spatiotemporal Features with 3D Convolutional Networks</article-title>
          .
          <source>In Proceedings of the IEEE international conference on computer vision</source>
          , (
          <year>2015</year>
          )
          <fpage>4489</fpage>
          -
          <lpage>4497</lpage>
          . doi:
          <volume>10</volume>
          .1109/ICCV.
          <year>2015</year>
          .
          <volume>510</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref32">
        <mixed-citation>
          [32]
          <string-name>
            <given-names>K.</given-names>
            <surname>Liu</surname>
          </string-name>
          , W. Liu,
          <string-name>
            <given-names>C.</given-names>
            <surname>Gan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Tan</surname>
          </string-name>
          , H. Ma: T-C3D:
          <article-title>Temporal Convolutional 3D Network for Realtime Action Recognition</article-title>
          .
          <source>In Proceedings of the AAAI conference on artificial intelligence</source>
          ,
          <volume>32</volume>
          (
          <issue>1</issue>
          ), (
          <year>2018</year>
          )
          <fpage>7138</fpage>
          -
          <lpage>7145</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref33">
        <mixed-citation>
          [33]
          <string-name>
            <given-names>D.</given-names>
            <surname>Ivanko</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Ryumin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Axyonov</surname>
          </string-name>
          ,
          <string-name>
            <surname>A.</surname>
          </string-name>
          <article-title>Kashevnik: Speaker-Dependent Visual Command Recognition in Vehicle Cabin: Methodology and Evaluation</article-title>
          . In International Conference on Speech and
          <source>Computer</source>
          (
          <year>2021</year>
          ). To appear.
        </mixed-citation>
      </ref>
      <ref id="ref34">
        <mixed-citation>
          [34]
          <string-name>
            <given-names>Y.</given-names>
            <surname>Kartynnik</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Ablavatski</surname>
          </string-name>
          , I. Grishchenko,
          <string-name>
            <surname>M.</surname>
          </string-name>
          <article-title>Grundmann: Real-time Facial Surface Geometry from Monocular Video on Mobile GPUs</article-title>
          . In: CVPR Workshop on Computer Vision for Augmented and Virtual Reality, IEEE (
          <year>2019</year>
          )
          <fpage>1</fpage>
          -
          <lpage>4</lpage>
          . arXiv:
          <year>1907</year>
          .06724.
        </mixed-citation>
      </ref>
      <ref id="ref35">
        <mixed-citation>
          [35]
          <string-name>
            <given-names>V.</given-names>
            <surname>Bazarevsky</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Kartynnik</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Vakunov</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Raveendran</surname>
          </string-name>
          ,
          <string-name>
            <surname>M.</surname>
          </string-name>
          <article-title>Grundmann: Blazeface: Submillisecond neural face detection on mobile gpus</article-title>
          .
          <source>arXiv preprint</source>
          (
          <year>2019</year>
          ). arXiv:
          <year>1907</year>
          .05047
        </mixed-citation>
      </ref>
      <ref id="ref36">
        <mixed-citation>
          [36]
          <string-name>
            <given-names>D.</given-names>
            <surname>Palaz</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Magimai-Doss</surname>
          </string-name>
          ,
          <string-name>
            <surname>R.</surname>
          </string-name>
          <article-title>Collobert: Analysis of cnn-based speech recognition system using raw speech as input</article-title>
          .
          <source>In 16th annual conference of the international speech communication association</source>
          , (
          <year>2015</year>
          ), pp.
          <fpage>686</fpage>
          -
          <lpage>692</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref37">
        <mixed-citation>
          [37]
          <string-name>
            <given-names>B.</given-names>
            <surname>McFee</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Colin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Dawen</surname>
          </string-name>
          , E. Daniel,
          <string-name>
            <given-names>M.</given-names>
            <surname>McVicar</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Battenberg</surname>
          </string-name>
          ,
          <string-name>
            <surname>O.</surname>
          </string-name>
          <article-title>Nieto: librosa: Audio and music signal analysis in python</article-title>
          .
          <source>In Proceedings of the 14th python in science conference</source>
          (
          <year>2015</year>
          )
          <fpage>18</fpage>
          -
          <lpage>25</lpage>
          . doi:
          <volume>10</volume>
          .25080/Majora-7b98e3ed-003.
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>