<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <journal-title-group>
        <journal-title>GraphiCon</journal-title>
      </journal-title-group>
    </journal-meta>
    <article-meta>
      <title-group>
        <article-title>Automatic Detection of Certain Unwanted Driver Behavior</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Boris Faizov</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Vlad Shakhuro</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Anton Konushin</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Lomonosov Moscow State University</institution>
          ,
          <addr-line>Leninskiye Gory, 1, Moscow, 119991</addr-line>
          ,
          <country country="RU">Russia</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>NRU Higher School of Economics</institution>
          ,
          <addr-line>Pokrovsky Bulvar, 11, Moscow, 109028</addr-line>
          ,
          <country country="RU">Russia</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2021</year>
      </pub-date>
      <volume>31</volume>
      <fpage>27</fpage>
      <lpage>30</lpage>
      <abstract>
        <p>This work is devoted to the automatic detection of unwanted driver behavior such as smoking, using a mobile phone, and eating. The various existing datasets are practically unsuitable for this task. We did not find suitable training data with RGB video sequences shot from the position of the inner mirror. So we investigated the possibility of training the algorithms for this task on an out-of-domain set of people faces images. We also filmed our own test video sequence in a car to test the algorithms. We investigated diferent existing algorithms working both with one frame and with video sequences and conducted an experimental comparison of them. The availability of temporal information improved quality. Another important aspect is metrics for assessing the quality of the resulting system. We showed that experimental evaluation in this task should be performed on the entire video sequences. We proposed an algorithm for detecting undesirable driver actions and showed its efectiveness.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;Computer Vision</kwd>
        <kwd>Action detection</kwd>
        <kwd>Action classification</kwd>
        <kwd>Driver distraction</kwd>
        <kwd>Domain adaptation</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>A large number of road accidents occur because the driver is distracted while driving. Diferent
countries have fines for such violations, but they are very dificult to track. Modern computer
vision methods can solve this problem. Nowadays state cameras automatically fine only for
easy-to-track violations. Improving drivers’ actions controlling is the next step in building such
systems. Also, it can be used in monitoring car-sharing and taxi drivers as they often employ
a large number of unqualified persons who often break the rules. Examples of distraction
actions which we focused on in our work are: talking on the phone, using a smartphone, eating,
smoking.</p>
      <p>
        For such a system to work, a video camera must be installed in the cab of the car, which
will record the actions of the driver. The most convenient location seems to us near the inner
mirror, which will shoot the driver from the front-top. However, we did not find any publically
available datasets with RGB images from the specified camera position. Moreover, we did not
ifnd datasets in which all the actions that we focused on in this work would be presented. In our
experiments we tried these sets: State-Farm’s competition [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ], Drive&amp;Act [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ], out-of-domain set
of people faces images provided by "Tevian" [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ] company. To test the algorithm, we filmed in a
car our own video sequence. In it, we alternated performing distractions with normal driving.
      </p>
      <p>
        Also, such a system must be efective to work in real-time. We tried diferent methods. To
classify by one frame, we used ResNet-50 [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ] and MobileNetV2 [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ]. After frame-by-frame
classification to improve accuracy we proposed an algorithm for the aggregation of consecutive
frames. To process video fragments, we tried: a complex Inflated 3D Model [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ] architecture
and a simple model in which the frame-by-frame outputs of the convolutional backbone are
aggregated by LSTM or a fully connected layer. In our experiments, we showed that 5 frames
per second is enough speed for such system.
      </p>
      <p>Another important aspect that we investigated is metrics for assessing the quality of the
system. In our task we need to determine what action is currently taking place, if necessary,
taking into account past frames. So we decided that our task is similar to online action detection.
In practice, quality is usually assessed either on separate frames from the test sample or on
already cut video fragments with known boundaries of actions. We concluded that experimental
evaluation of such methods should be performed on the entire video sequence. We don’t have
to detect distraction on each separate frame and it is enough to find distraction in at least one
frame. It is also important to monitor the false positive rate and per-class recall.</p>
      <p>To conclude our paper we proposed a method for car driver’s action classification. This
method evaluates separate frames and then aggregates temporal information over the last
frames. We showed that data from other domains could be successfully used for training such
an algorithm.</p>
    </sec>
    <sec id="sec-2">
      <title>2. Related work</title>
      <p>The approaches for solving this problem could be divided into frame-by-frame classification
and video action detection.</p>
      <p>
        Most of the papers and techniques for this task work with a single frame. Conventional
convolution neural network architectures such as VGG[
        <xref ref-type="bibr" rid="ref7">7</xref>
        ], ResNet[
        <xref ref-type="bibr" rid="ref4">4</xref>
        ], MobileNet [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ] etc. can be
used for the classification of driver’s images. Simple specialized CNN architecture was proposed
in [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ]. They also collected two sets of data to account for the poor illuminations and diferent
road conditions. The evaluation was performed both on their new set and on (SEU) dataset [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ].
Another fast CNN model used Principal Component Analysis (PCA) technology to whiten the
driver’s image in [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ]. To reduce computational complexity and memory requirements for
VGG [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ] network researchers introduce modifications in its architecture in [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ] specifically for
the task of driver monitoring. Several papers [12, 13] used a genetically weighted ensemble of
convolutional neural networks with diferent CNN’s classifying diferent cropped body parts
      </p>
      <p>
        Another approach is to use video sequence classification methods such as [
        <xref ref-type="bibr" rid="ref6">14, 15, 6</xref>
        ]. However,
these approaches contain 3D-convolutions and have too high computational complexity for
practical use in real-time systems. But it is possible to aggregate information extracted by
conventional CNNs from the last few frames. For example, a model can submit the extracted
features to the LSTM [16, 17]. Also, SoTA-models [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ] often analyze the optical flow in parallel
with RGB-frames. Optical flow represents the motion direction of each pixel between two image
frames. It is a powerful idea and it has been used to improve accuracy when classifying videos. It
helps algorithms to prioritize motion as a key characteristic of the scene. But in a real embedded
system calculating the optical flow can be too slow. In [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ] the authors applied the video models
to their new dataset and got the best quality with the Inflated 3D Model (I3D) [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ]. Authors
of I3D architecture suggested converting successful 2D image classification models into 3D
ConvNets. They used the Inception-V1 [18] model and inflated all the filters and pooling kernels
– endowing them with an additional temporal dimension. Parameters also were bootstrapped
from the pre-trained ImageNet model.
      </p>
    </sec>
    <sec id="sec-3">
      <title>3. Datasets</title>
      <p>Collecting video sequences for this task is challenging. Firstly, it is unsafe to collect completely
real data in which a person is really driving a car and is asked to perform distractions. This
approach can result in an emergency situation. Secondly, the collected data should be
representative and contain a large number of people. The models shouldn’t learn to work only on a few
specific persons.</p>
      <p>
        There are several publically available datasets for driver action recognition:
1. State-Farm Kaggle competition dataset [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]: 26 participants, 10 classes, 22424 training
images. For safety reasons the truck was dragging the car around on the streets — so
these drivers weren’t really driving. Unfortunately, the usage of StateFarm’s dataset is
limited to the purposes of the competition. Frame examples are shown in figure 4(a).
2. Distracted Driver Dataset [12]: 44 participants, 10 classes, 14478 frames. It was inspired
by State-Farm’s dataset and has similar classes.
3. Drive&amp;Act [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ] dataset: 15 participants, hierarchical annotation scheme with 83
finegrained categories. Unlike previous datasets, it consists of video sequences which is much
useful in a real scenario. Every person was asked to perform twice a pre-defined sequence
of actions. Set has RGB and depth images from a side view, 3D body poses information
from front-top view, and infrared images from six diferent views. Frame examples are
shown in figure 1.
      </p>
      <p>There are many diferent kinds of distractions. Some actions of interest may be missing. By
mixing diferent datasets, the model can learn to identify diferent sets of classes by environment
and driver, rather than by their actions. Therefore, with a new class, it might be necessary to
collect the entire dataset again.</p>
      <p>
        In our experiments, we used proprietary data provided by "Tevian" [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ] company. It has RGB
images of neutral, smoking, and talking on the phone people. In the experiments, we divided
the sample into training (75%) and validation (25%) so that the proportions of the classes were
preserved. We needed these data to train our model find people who smoke while driving
because there is no such class in other datasets. Examples of frames are in figure 2. The big
disadvantage of this set in this task is that we have to cut out the faces of people in the car. This
could lead to the loss of contextual information in the frame and deterioration of the possible
quality.
      </p>
      <p>We also filmed our own test video sequence. For safety, the car wasn’t moving during the
shooting. It presents classes for safe driving, eating (food, drinks), telephone (texting, calls),
smoking. The length of the sequence is 14 minutes. Examples of frames are shown in figure 3.</p>
      <p>
        Drivers in [
        <xref ref-type="bibr" rid="ref1 ref9">1, 12, 9</xref>
        ] and RGB-images from Drive&amp;Act [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ] are photographed from the side
view. It seems to us that this choice is not very successful, since in a real system (especially in a
taxi car) the camera with such an arrangement can be obstructed by a neighboring passenger
or luggage.
      </p>
    </sec>
    <sec id="sec-4">
      <title>4. Proposed method</title>
      <p>Low-performance models such as Inflated 3D Model can be used if training takes place on
labeled video sequences. We decided to try a more eficient approach. We propose to first
train a frame-by-frame classifier of driver actions based on MobileNetV2. We took the network
pre-trained on the ImageNet and fine-tuned it on the training set images. Further, sequential
frames may be independently classified and then aggregated with an additional neural network.
We considered such frame aggregational neural networks:
1. A single-layer perceptron that accepts concatenated frame attributes as input and outputs
an action class.
2. Recurrent LSTM network with 128 features.
3. Transformer [19] network with 4 layers, 2 heads and 128 features.</p>
      <p>When training on a frame-by-frame dataset, there is no way to train a video model. But to
improve the quality of the method in the real-time system, we still need to aggregate the results
of the last few frames. For this, we have proposed our method:
1. Consider that the system has a speed equal to 5 frames per second.
2. The system classifies each frame independently of the others.
3. The last 10 frames are considered. If at least 8 of them were classified as some class, which
is not safe driving, then we consider that we have detected a distraction at the moment.</p>
      <p>Otherwise, we consider the current moment of time to be safe driving.</p>
    </sec>
    <sec id="sec-5">
      <title>5. Metrics</title>
      <p>We used several metrics to measure the performance of diferent approaches. For per-frame
classification we used:
• Accuracy — the ratio of correctly classified images to the total number of images.
• False positives — the ratio of safe driving images, classified as driver distractions to the
total number of safe driving images.</p>
      <p>• Per-class recall.</p>
      <p>For actions in video sequences, we used another set of metrics. By action, we mean a set of
frames with one separate video event as a whole. The system does not need to be triggered at
every point in time during a distraction event. It is enough to find at least one moment with
each action so that the driver could be fined. We need to take into consideration that each frame
has previous frames and we can use this data. So metrics should be calculated after aggregation
of the last few frames. Also, we need to have metrics that combine information from an entire
video.</p>
      <p>• Accuracy on frames after aggregation with previous frames.
• False positive on frames after aggregation with previous frames.
• Per-class recall on frames after aggregation with previous frames.
• False positive on actions. This means that for each ground truth action we collect a set of
all triggered classes during this action. Then if this action is safe driving, but we have
distracted class in a set of triggered classes, then this action is a false positive.
• Per-class recall on actions. This is similar to the previous metric, but we check if our
model triggered a real class of each action.
• Per-class mAp. First, the frames are ranked according to their confidence from high to
  ()
low.  @ =   ()+  () , where   () and   () are the number of true
positives and false positives accordingly at the first  frames. The average precision of a
class is then defined as
 = ∑︀   @*1[   ] . The mean of the  over all classes is final mAP.</p>
      <p>When testing, we took into account that it is possible to trigger an action in the range of
+ − 40 frames from it. This type of misclassification was still considered correct.</p>
    </sec>
    <sec id="sec-6">
      <title>6. Results</title>
      <sec id="sec-6-1">
        <title>6.1. Experiments with State-Farm dataset</title>
        <p>
          We tried neural network models on State-Farm competition [
          <xref ref-type="bibr" rid="ref1">1</xref>
          ] dataset. Among the conventional
models we have chosen ResNet-50 [
          <xref ref-type="bibr" rid="ref4">4</xref>
          ] and MobileNetV2 [
          <xref ref-type="bibr" rid="ref5">5</xref>
          ]. Having studied the existing
solutions, we decided to try not only RGB pictures, but also the segmentation of people in the
frames. The segmentation model was similar to the RGB model, segmentation masks were
passed to the input instead of color images. We also tried to concatenate outputs of RGB and
segmentation models and add one fully connected layer on top of these features to obtain
classifications. The segmentation was obtained using the human parsing model [ 20]. Examples
of frames and segmentation are shown in figure 4(a,b). In addition to the usual augmentations
(turns, crops, noise) we used augmentation where two diferent pictures of the same class are
mixed, an example in figure 4(c). This additional augmentation improved the quality. It seems
to us that this happened because such transformation prevented the model from remembering
specific people and there were not enough diferent persons in the frames of the original training
set. However, we noticed that marked and test sets have subsets of the same people in images.
Therefore, we divided the labeled sample into training and validation sets according to people
identifiers: 19 persons in the training set, 26 in the validation set. The results are shown in
table 1. The operating time here and further was measured on the Intel Core I7-3770K CPU
processor.
        </p>
      </sec>
      <sec id="sec-6-2">
        <title>6.2. Basic experiments on Drive&amp;Act</title>
        <sec id="sec-6-2-1">
          <title>6.2.1. Inflated 3D Model.</title>
          <p>
            To begin with, we tried to reproduce the results of paper [
            <xref ref-type="bibr" rid="ref2">2</xref>
            ] for the classification of fine-grained
activities of 34 classes by Inflated 3D Model with frame resolution 224 × 224. The model was
trained separately on RGB images from side view and optical flow. When training an RGB
stream, we took 36 frames with a stride of 2 frames, and when training an optical flow stream,
we take 18 frames with a stride of 2 frames. During testing, all frames of a video fragment are
taken. The results are shown in table 2. Our average per-class accuracy on validation 68.97 has
almost reached the value from the original paper 69.57. But in a low-performance real-time
system, it is unlikely that it will be possible to calculate the optical flow.
          </p>
        </sec>
        <sec id="sec-6-2-2">
          <title>6.2.2. Three selected classes.</title>
          <p>In our experiments, we decided to focus on three important activities: using a smartphone
(talking on the phone, interacting with phone), food (eating, drinking), safe driving (all other
classes). In addition to the Inflated 3D Model model, we decided to train the frame-by-frame
model with MobileNetV2 architecture. We also tried to train MobileNetV2 only on cropped
heads. Heads were cropped from segmentations obtained with human parsing model [20]. The
results are shown in the table 3.</p>
        </sec>
      </sec>
      <sec id="sec-6-3">
        <title>6.3. Experiments on Drive&amp;Act</title>
        <p>In this section, we focused on results on three selected classes of the Drive&amp;Act dataset (safe
driving, eating, phone). The best results when training and testing on the Drive&amp;Act set were
obtained with the following training setup:
1. Convert original markup that was discontinuous, merging consecutive clips of similar
classes.
2. Training a frame-by-frame classifier for all 34 classes on full (not cropped)[522] frames.
3. Extraction of frame features from the outputs of the penultimate layer of the classifier.
4. Training LSTM network / One-layer linear perceptron / small transformer on 3 selected
classes of interest with zero class weight equal to 4.0. The network will classify by the
last 10 frames with FPS=5.</p>
        <p>All consecutive frames with the preceding ones in the video are viewed independently,
therefore, there are as many triggers in the confusion matrix as there are frames in the video.</p>
        <p>Results are shown in table 4. Best of all, in our opinion, is the LSTM network, because it has
a low false-positive rate and good recall/accuracy. The problem with the transformer was that
it was heavily overfitted due to a lack of data. We tried to reduce the number of transformer
layers or the dimension of the features, but we were unable to improve the quality. Maybe
transformers don’t fit well for this particular dataset. We also investigated frames, where the
LSTM model had false positives and concluded that they were normal in most cases. They
occurred mostly in boundaries between actions in the dataset due to the not accurate markup.
Aggregating frames without neural networks reduce the false-positive rate, but it worsens recall
on distraction classes.</p>
        <p>Learning only on people’s heads worked much worse. It seems to us that this is because the
context is not visible.</p>
      </sec>
      <sec id="sec-6-4">
        <title>6.4. Experiments on faces images and our test images</title>
        <p>
          Our test video consists of RGB images from the front-top view. But in Drive&amp;Act videos from
the front-top view are infrared. Therefore, this data is not suitable for testing our method.
Data with face images provided by "Tevian"[
          <xref ref-type="bibr" rid="ref3">3</xref>
          ] company is more acceptable — these are RGB
pictures with people’s faces and three classes: nothing, phone, smoking. Therefore, we trained
Method
        </p>
        <p>LSTM
Linear perceptron</p>
        <p>Transformer
Aggregating frames
by last 10 frames
without NN
the MobileNetV2 classifier on people faces. Then we applied it to our test video. In this case,
we need to cut out the head of a person. We again did this using the human parsing model [20].
The results obtained during frame-by-frame testing are in table 5. And the results on our test
video are in table 6. Our result on the test video still contained 10 false-positive frames after
aggregation of the last 10 frames, but they all were just before the person in the video started
simulating smoking. Per-frame recall values may seem low, but metrics on action level are
better. Even if we didn’t trigger on a class at every time moment of action, we still found 2 of 5
smoking acts and 4 of 11 phone using acts.</p>
      </sec>
    </sec>
    <sec id="sec-7">
      <title>7. Conclusion</title>
      <p>In this paper, we considered the problem of recognizing distracted drivers. We proposed several
methods for driver’s action classification using frames and video sequences. In the absence
of video sequences suitable for the training set, we proposed a method for aggregating the
last frames of a video sequence and experimentally demonstrated its efectiveness. We have
shown that the driver distraction classification problem can be solved using out-of-domain
training data. The proposed model was trained on photographs of people’s faces and showed
high quality on our test video sequence.
work, in: Proceedings of the IEEE conference on computer vision and pattern recognition
workshops, 2018, pp. 1032–1038.
[12] Y. Abouelnaga, H. M. Eraqi, M. N. Moustafa, Real-time distracted driver posture
classification, arXiv preprint arXiv:1706.09498 (2017).
[13] H. M. Eraqi, Y. Abouelnaga, M. H. Saad, M. N. Moustafa, Driver distraction identification
with an ensemble of convolutional neural networks, Journal of Advanced Transportation
2019 (2019).
[14] D. Tran, L. Bourdev, R. Fergus, L. Torresani, M. Paluri, Learning spatiotemporal features
with 3d convolutional networks, in: Proceedings of the IEEE international conference on
computer vision, 2015, pp. 4489–4497.
[15] Z. Qiu, T. Yao, T. Mei, Learning spatio-temporal representation with pseudo-3d residual
networks, in: proceedings of the IEEE International Conference on Computer Vision, 2017,
pp. 5533–5541.
[16] S. Yeung, O. Russakovsky, N. Jin, M. Andriluka, G. Mori, L. Fei-Fei, Every moment counts:
Dense detailed labeling of actions in complex videos, International Journal of Computer
Vision 126 (2018) 375–389.
[17] N. Srivastava, E. Mansimov, R. Salakhudinov, Unsupervised learning of video
representations using lstms, in: International conference on machine learning, PMLR, 2015, pp.
843–852.
[18] C. Szegedy, W. Liu, Y. Jia, P. Sermanet, S. Reed, D. Anguelov, D. Erhan, V. Vanhoucke,
A. Rabinovich, Going deeper with convolutions, in: Proceedings of the IEEE conference
on computer vision and pattern recognition, 2015, pp. 1–9.
[19] A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. Kaiser, I.
Polosukhin, Attention is all you need, arXiv preprint arXiv:1706.03762 (2017).
[20] P. Li, Y. Xu, Y. Wei, Y. Yang, Self-correction for human parsing, IEEE Transactions on
Pattern Analysis and Machine Intelligence (2020).</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>State</given-names>
            <surname>Farm Distracted Driver Detection</surname>
          </string-name>
          ,
          <source>State farm distracted driver detection</source>
          ,
          <year>2016</year>
          . URL: https://www.kaggle.com/c/state-farm
          <article-title>-distracted-driver-detection/overview.</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>M.</given-names>
            <surname>Martin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Roitberg</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Haurilet</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Horne</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Reiß</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Voit</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Stiefelhagen</surname>
          </string-name>
          ,
          <article-title>Drive&amp;act: A multi-modal dataset for fine-grained driver behavior recognition in autonomous vehicles</article-title>
          ,
          <source>in: Proceedings of the IEEE/CVF International Conference on Computer Vision</source>
          ,
          <year>2019</year>
          , pp.
          <fpage>2801</fpage>
          -
          <lpage>2810</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <surname>Tevian</surname>
          </string-name>
          , Tevian,
          <year>2021</year>
          . URL: https://tevian.ai/.
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>K.</given-names>
            <surname>He</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Zhang</surname>
          </string-name>
          , S. Ren,
          <string-name>
            <given-names>J.</given-names>
            <surname>Sun</surname>
          </string-name>
          ,
          <article-title>Deep residual learning for image recognition</article-title>
          ,
          <source>in: Proceedings of the IEEE conference on computer vision and pattern recognition</source>
          ,
          <year>2016</year>
          , pp.
          <fpage>770</fpage>
          -
          <lpage>778</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>M.</given-names>
            <surname>Sandler</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Howard</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Zhu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Zhmoginov</surname>
          </string-name>
          , L.-C.
          <article-title>Chen, Mobilenetv2: Inverted residuals and linear bottlenecks</article-title>
          ,
          <source>in: Proceedings of the IEEE conference on computer vision and pattern recognition</source>
          ,
          <year>2018</year>
          , pp.
          <fpage>4510</fpage>
          -
          <lpage>4520</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>J.</given-names>
            <surname>Carreira</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Zisserman</surname>
          </string-name>
          ,
          <article-title>Quo vadis, action recognition? a new model and the kinetics dataset</article-title>
          ,
          <source>in: proceedings of the IEEE Conference on Computer Vision and Pattern Recognition</source>
          ,
          <year>2017</year>
          , pp.
          <fpage>6299</fpage>
          -
          <lpage>6308</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>K.</given-names>
            <surname>Simonyan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Zisserman</surname>
          </string-name>
          ,
          <article-title>Very deep convolutional networks for large-scale image recognition</article-title>
          ,
          <source>arXiv preprint arXiv:1409.1556</source>
          (
          <year>2014</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>C.</given-names>
            <surname>Yan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Coenen</surname>
          </string-name>
          ,
          <string-name>
            <surname>B. Zhang,</surname>
          </string-name>
          <article-title>Driving posture recognition by convolutional neural networks</article-title>
          ,
          <source>IET Computer Vision</source>
          <volume>10</volume>
          (
          <year>2016</year>
          )
          <fpage>103</fpage>
          -
          <lpage>114</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>C.</given-names>
            <surname>Zhao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Zhang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>He</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Lian</surname>
          </string-name>
          ,
          <article-title>Recognition of driving postures by contourlet transform and random forests</article-title>
          ,
          <source>IET Intelligent Transport Systems</source>
          <volume>6</volume>
          (
          <year>2012</year>
          )
          <fpage>161</fpage>
          -
          <lpage>168</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>X.</given-names>
            <surname>Rao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Lin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Chen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Zhao</surname>
          </string-name>
          ,
          <article-title>Distracted driving recognition method based on deep convolutional neural network</article-title>
          ,
          <source>Journal of Ambient Intelligence and Humanized Computing</source>
          (
          <year>2019</year>
          )
          <fpage>1</fpage>
          -
          <lpage>8</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>B.</given-names>
            <surname>Baheti</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Gajre</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Talbar</surname>
          </string-name>
          ,
          <article-title>Detection of distracted driver using convolutional neural net-</article-title>
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>