<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Spatio-Temporal CNN Baseline Method for the Sports Video Task of MediaEval 2021 Benchmark Pierre-Etienne Martin CCP Department, Max Planck Institute for Evolutionary Anthropology, D-04103 Leipzig, Germany pierre_etienne_martin@eva.mpg.de</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>FC SoftMax</string-name>
        </contrib>
      </contrib-group>
      <pub-date>
        <year>2021</year>
      </pub-date>
      <fpage>13</fpage>
      <lpage>15</lpage>
      <abstract>
        <p>This paper presents the baseline method proposed for the Sports Video task part of the MediaEval 2021 benchmark. This task proposes a stroke detection and a stroke classification subtasks. This baseline addresses both subtasks. The spatio-temporal CNN architecture and the training process of the model are tailored according to the addressed subtask. The method has the purpose of helping the participants to solve the task and is not meant to reach stateof-the-art performance. Still, for the detection task, the baseline is performing better than the other participants, which stresses the dificulty of such a task.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>49
30
60
30
Pool
24
Conv
(3x3x3)
ReLU
80
(3Cxo3nxv3)Pool
ReLU
1</p>
      <p>FC
ReLU
500</p>
    </sec>
    <sec id="sec-2">
      <title>INTRODUCTION</title>
      <p>Most recent action detection and classification methods developed
in the literature have been using deep learning approaches and
high-dimensional spaces. In the domain of image classification,
a specific kind of Neural Network has become very popular: the
Convolutional Neural Networks (CNNs). Since the breakthrough
at the 2012 ImageNet Challenge, CNNs have demonstrated a great
improvement for image classification.</p>
      <p>
        For video applications in general and action recognition in
particular, the first models proposed were a direct extension of image
classification methods [
        <xref ref-type="bibr" rid="ref1 ref20">1, 20</xref>
        ] using 2D convolutions. However, to
better capture the temporal information proper to video content,
the use of 3D convolutions has emerged [
        <xref ref-type="bibr" rid="ref7 ref8">7, 8</xref>
        ]. One can also consider
temporal information using the motion extracted from successive
frames, such as the optical flow. The latest can be used i) as a single
modality or in parallel with the RGB information [
        <xref ref-type="bibr" rid="ref2 ref20 ref21 ref3">2, 3, 20, 21</xref>
        ]; or ii)
to train a network for extracting motion features to perform
classiifcation at a later stage [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ]. These methods also raise the question of
how to fuse the diferent modalities [
        <xref ref-type="bibr" rid="ref10 ref5">5, 10</xref>
        ]. In [
        <xref ref-type="bibr" rid="ref19 ref9">9, 19</xref>
        ], the estimated
pose is used jointly with these two modalities to perform action
classification. In [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ] all the three modalities are used and fused in
order to perform stroke classification.
      </p>
      <p>
        As part of the task organization, for the first time since the
beginning of the Sports Video task (in 2019 [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ]), we decided to provide
a baseline to alleviate minor aspects of the task, such as video and
xml processing; and help the participants in their submission. The
baseline method uses a 3D CNN inspired from [
        <xref ref-type="bibr" rid="ref13 ref14">13, 14</xref>
        ]. We adjusted
the method to answer both proposed subtasks of this year’s
edition [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ]: stroke detection and stroke classification from videos
of the TTStroke-21 corpus. The implementation of the method is
available publicly on Github1.
2
      </p>
    </sec>
    <sec id="sec-3">
      <title>METHOD</title>
      <p>
        In order to perform classification and detection, we consider the
model architecture presented in Fig. 1. For each subtask, a distinct
model has been trained on the train set. We train both using a
stochastic gradient approach with a Nesterov momentum of 0.5 [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ],
a weight decay of 0.005 [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ] and a constant learning rate of 0.0001.
Both models are trained over 500 epochs. The objective function
is the cross-entropy loss of the output processed by the softmax
function (eq. 1) summing over the batch:
 (′ )
L (, ) = −( Í  ( ) )
(1)
      </p>
      <p>At each epoch, the model is validated on the validation set. The
model performing the best on this set is saved and then evaluated on
the test set. The model is fed with the video frames resized to 120 ×
120 and staked successively in cuboids of length 98, representing
approximately 0.82 seconds.</p>
      <p>For the detection task, we inferred Non-stroke segments from the
annotated Stroke segments. We considered only segments between
two consecutive strokes greater than 200 frames. Such a segment
is divided in successive blocks of 200 frames, non overlapping, and
added has a negative sample for training the model. The split using
200 frames allows a correct number of negative samples: from the
783 train and 234 validation segments, we inferred respectively for
each set 1196 and 260 negative segments. No negative segments
have been inferred from the test set. Stroke detection is tackled
1https://github.com/ccp-eva/SportTaskME21
as a classification task by considering two classes: Stroke on
Nonstroke. From the test set, which has no temporal boundaries, we
created window proposals of length 150 every 150 frames for all the
videos. This size was chosen empirically and meant to be revised to
achieve good performance. For the classification task, all the classes
were not represented in the dataset but we still consider all the 20
possible stroke classes.</p>
      <p>To train the model, we inputted the RGB cuboids composed of the
successive frames from the starting frame of the considered segment.
The desired output is the class vector summing to one and binary
at training time. Its length is the number of considered classes: 2
for detection and 20 for classification. Each element represents the
probability of belonging to a class. During inference, we follow
a similar procedure, and the class decision is the argmax of the
output vector.
3</p>
    </sec>
    <sec id="sec-4">
      <title>RESULTS</title>
      <p>
        This section presents the results per subtask according to the
metrics presented in [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ].
3.1
      </p>
    </sec>
    <sec id="sec-5">
      <title>Subtask 1 - Stroke Detection</title>
      <p>The detection subtask was tackled as a classification task,
considering the strokes and non-strokes samples. After 500 epochs, the
model reached 98.3% and 75.7% of accuracy, respectively, on the
train and validation sets. On the test set, the model is evaluated
using the mAP metric. This metric takes into account the number
of actions detected and their overlapping with the ground truth.
The baseline achieves an mAP of 0.0173, which the two participants
of this subtask did not outperform.</p>
      <p>Runs are also evaluated using a global IoU that considers only
the frame-wise overlap of the detected strokes with the ground
truth annotations. The number of strokes detected is no longer
taken into account in the evaluation. The baseline achieves a Global
IoU of 0.144, which was outperformed by one participant.</p>
      <p>The method’s performance is quite low due to the method being
relatively simple. It also relies on a straightforward and non-eficient
window proposal to segment the strokes without fusing the output
decision. Indeed, two consecutive windows, part of the same stroke
and classified as strokes, will be classified as two diferent strokes
and not a single one, and will therefore have an impact on the mAP
metric. The method can easily be improved by considering better
proposals and fusing the output decisions.
3.2</p>
    </sec>
    <sec id="sec-6">
      <title>Subtask 2 - Stroke Classification</title>
      <p>The results for the stroke classification subtask on the test set are
reported in the table 1. This table is divided into diferent sections
for considering diferent refined classifications. After training, the
model reached only 25.2% and 28.9% of accuracy, respectively, on
the train and validation sets.</p>
      <p>• “Global” consider all the 20 classes
• “Type” consider only the type of the stroke: Defensive,</p>
      <p>Ofensive or Service
• “Hand-Side” consider only Forehand and Backhand
superclasses
• “Type and Hand-Sided” consider the intersection of the
two last clusters leading to 6 classes.</p>
      <p>The confusion matrix of the “Type” and “Hand-Side” are also
depicted in Fig. 2 for further analysis.</p>
      <p>Pierre-Etienne Martin</p>
      <p>
        From table 1, we can state that the performance of the baseline,
considering all classes, is limited. This may be improved by
further analysis of the corpus and further training. Indeed, only 18
classes over the 20 possible were present in the corpus this year,
which simplifies the complexity of the task and could have been
taken into account in the model’s design. Fig. 2.a reveals that the
services have not been learned at all, which is undoubtedly due to
the input processing during training which considers only the 100
ifrst frames and is therefore unable to capture features from these
longer strokes. Finally, Fig. 2.b underlines the main weakness of
the model: being unable of distinguishing Forehand and Backhand
strokes. The pipeline’s method could consider higher level
categories, following a cascade method, to improve the performance.
Two of the three participants have outperformed by far the baseline
performance [
        <xref ref-type="bibr" rid="ref17 ref18">17, 18</xref>
        ].
4
      </p>
    </sec>
    <sec id="sec-7">
      <title>CONCLUSION</title>
      <p>This baseline intends to help the participants solving the Sports
Video Task. The baseline performance remains limited, but its
publicly available implementation allows the participants to not start
from scratch. Many aspects of the method may be improved, such as
the data processing: a spatial and temporal ROI may increase the
performance. Similarly with the architecture of the model, which was
kept very simple, or the training method that could have merged
the train and validation sets before inferring on the test set.</p>
      <p>The detection subtask seems to be challenging. No participants
were able to beat the baseline performance with regard to the mAP
metric, which is the ranking metric. This subtask is new in the
Sports Video Task, which also explains the low results obtained.
However we believe much improvement can be obtained since our
method has tackled it as a classification task. The window proposal
is also very crude and can easily be improved.</p>
      <p>The classification subtask has gathered more participants with,
overall, more successful performance. This may be explained by
the task’s non-novelty in the history of the MediaEval benchmark
and the more active investigation in this field.</p>
      <p>Next year we plan to gather ideas from this year’s submissions
to improve the baseline and give a more substantial base to the new
participants joining the Sports Video Task.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>Hakan</given-names>
            <surname>Bilen</surname>
          </string-name>
          , Basura Fernando, Efstratios Gavves, and
          <string-name>
            <given-names>Andrea</given-names>
            <surname>Vedaldi</surname>
          </string-name>
          .
          <year>2018</year>
          .
          <article-title>Action Recognition with Dynamic Image Networks</article-title>
          .
          <source>IEEE Trans. Pattern Anal. Mach. Intell</source>
          .
          <volume>40</volume>
          ,
          <issue>12</issue>
          (
          <year>2018</year>
          ),
          <fpage>2799</fpage>
          -
          <lpage>2813</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>Jordan</given-names>
            <surname>Calandre</surname>
          </string-name>
          , Renaud Péteri, and
          <string-name>
            <given-names>Laurent</given-names>
            <surname>Mascarilla</surname>
          </string-name>
          .
          <year>2019</year>
          .
          <article-title>Optical Flow Singularities for Sports Video Annotation: Detection of Strokes in Table Tennis</article-title>
          . In
          <source>MediaEval (CEUR Workshop Proceedings)</source>
          , Vol.
          <volume>2670</volume>
          . CEUR-WS.org.
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>João</given-names>
            <surname>Carreira</surname>
          </string-name>
          and
          <string-name>
            <given-names>Andrew</given-names>
            <surname>Zisserman</surname>
          </string-name>
          .
          <year>2017</year>
          .
          <article-title>Quo Vadis, Action Recognition? A New Model and the Kinetics Dataset</article-title>
          .
          <source>In CVPR. IEEE Computer Society</source>
          ,
          <fpage>4724</fpage>
          -
          <lpage>4733</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>Nieves</given-names>
            <surname>Crasto</surname>
          </string-name>
          , Philippe Weinzaepfel, Karteek Alahari, and
          <string-name>
            <given-names>Cordelia</given-names>
            <surname>Schmid</surname>
          </string-name>
          .
          <year>2019</year>
          .
          <article-title>MARS: Motion-Augmented RGB Stream for Action Recognition</article-title>
          .
          <source>In CVPR. IEEE Computer Society</source>
          ,
          <fpage>7882</fpage>
          -
          <lpage>7891</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>Christoph</given-names>
            <surname>Feichtenhofer</surname>
          </string-name>
          , Axel Pinz, and
          <string-name>
            <given-names>Andrew</given-names>
            <surname>Zisserman</surname>
          </string-name>
          .
          <year>2016</year>
          .
          <article-title>Convolutional Two-Stream Network Fusion for Video Action Recognition</article-title>
          .
          <source>In CVPR. IEEE Computer Society</source>
          , 1933-
          <fpage>1941</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>Stephen</given-names>
            <surname>Jose</surname>
          </string-name>
          Hanson and Lorien Y. Pratt.
          <year>1988</year>
          .
          <article-title>Comparing Biases for Minimal Network Construction with Back-Propagation</article-title>
          .
          <source>In NIPS</source>
          .
          <volume>177</volume>
          -
          <fpage>185</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>Ho</given-names>
            <surname>Joon</surname>
          </string-name>
          <string-name>
            <given-names>Kim</given-names>
            ,
            <surname>Joseph</surname>
          </string-name>
          <string-name>
            <given-names>S.</given-names>
            <surname>Lee</surname>
          </string-name>
          , and Hyun Seung Yang.
          <year>2007</year>
          .
          <article-title>Human Action Recognition Using a Modified Convolutional Neural Network</article-title>
          .
          <source>In ISNN (2) (Lecture Notes in Computer Science)</source>
          , Vol.
          <volume>4492</volume>
          . Springer,
          <fpage>715</fpage>
          -
          <lpage>723</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>Tiago</given-names>
            <surname>Lima</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Bruno J. T.</given-names>
            <surname>Fernandes</surname>
          </string-name>
          , and
          <string-name>
            <surname>Pablo</surname>
            <given-names>V. A.</given-names>
          </string-name>
          <string-name>
            <surname>Barros</surname>
          </string-name>
          .
          <year>2017</year>
          .
          <article-title>Human action recognition with 3D convolutional neural network</article-title>
          .
          <source>In LA-CCI. IEEE</source>
          , 1-
          <fpage>6</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <surname>Diogo</surname>
            <given-names>C.</given-names>
          </string-name>
          <string-name>
            <surname>Luvizon</surname>
            , David Picard,
            <given-names>and Hedi</given-names>
          </string-name>
          <string-name>
            <surname>Tabia</surname>
          </string-name>
          .
          <year>2018</year>
          .
          <article-title>2D/3D Pose Estimation and Action Recognition Using Multitask Deep Learning</article-title>
          .
          <source>In CVPR. IEEE Computer Society</source>
          ,
          <fpage>5137</fpage>
          -
          <lpage>5146</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <surname>Pierre-Etienne Martin</surname>
          </string-name>
          .
          <year>2020</year>
          .
          <article-title>Fine-Grained Action Detection and Classification from Videos with Spatio-Temporal Convolutional Neural Networks</article-title>
          .
          <article-title>Application to Table Tennis. (Détection et classification fines d'actions à partir de vidéos par réseaux de neurones à convolutions spatio-temporelles</article-title>
          .
          <source>Application au tennis de table)</source>
          .
          <source>Ph.D. Dissertation</source>
          . University of La Rochelle, France. https://tel.archives-ouvertes.fr/ tel-03128769
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <surname>Pierre-Etienne</surname>
            <given-names>Martin</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Jenny</surname>
          </string-name>
          Benois-Pineau, Boris Mansencal, Renaud Péteri, Laurent Mascarilla,
          <string-name>
            <surname>Jordan Calandre</surname>
            , and
            <given-names>Julien</given-names>
          </string-name>
          <string-name>
            <surname>Morlier</surname>
          </string-name>
          .
          <year>2019</year>
          .
          <article-title>Sports Video Annotation: Detection of Strokes in Table Tennis Task for MediaEval 2019</article-title>
          . In
          <source>MediaEval (CEUR Workshop Proceedings)</source>
          , Vol.
          <volume>2670</volume>
          . CEUR-WS.org.
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <surname>Pierre-Etienne</surname>
            <given-names>Martin</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Jenny</surname>
          </string-name>
          Benois-Pineau, Boris Mansencal, Renaud Péteri, Laurent Mascarilla,
          <string-name>
            <surname>Jordan Calandre</surname>
            , and
            <given-names>Julien</given-names>
          </string-name>
          <string-name>
            <surname>Morlier</surname>
          </string-name>
          .
          <year>2021</year>
          .
          <article-title>Sports Video: FineGrained Action Detection and Classification of Table Tennis Strokes from videos for MediaEval 2021</article-title>
          . In MediaEval (CEUR Workshop Proceedings).
          <source>CEUR-WS.org.</source>
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <surname>Pierre-Etienne</surname>
            <given-names>Martin</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Jenny</surname>
            Benois-Pineau, Boris Mansencal, Renaud Péteri, and
            <given-names>Julien</given-names>
          </string-name>
          <string-name>
            <surname>Morlier</surname>
          </string-name>
          .
          <year>2019</year>
          .
          <article-title>Siamese Spatio-Temporal Convolutional Neural Network for Stroke Classification in Table Tennis Games</article-title>
          . In
          <source>MediaEval (CEUR Workshop Proceedings)</source>
          , Vol.
          <volume>2670</volume>
          . CEUR-WS.org.
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <surname>Pierre-Etienne</surname>
            <given-names>Martin</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Jenny</surname>
            Benois-Pineau,
            <given-names>Renaud</given-names>
          </string-name>
          <string-name>
            <surname>Péteri</surname>
            , and
            <given-names>Julien</given-names>
          </string-name>
          <string-name>
            <surname>Morlier</surname>
          </string-name>
          .
          <year>2018</year>
          .
          <article-title>Sport Action Recognition with Siamese Spatio-Temporal CNNs: Application to Table Tennis</article-title>
          .
          <source>In CBMI. IEEE</source>
          , 1-
          <fpage>6</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <surname>Pierre-Etienne</surname>
            <given-names>Martin</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Jenny</surname>
            Benois-Pineau,
            <given-names>Renaud</given-names>
          </string-name>
          <string-name>
            <surname>Péteri</surname>
            , and
            <given-names>Julien</given-names>
          </string-name>
          <string-name>
            <surname>Morlier</surname>
          </string-name>
          .
          <year>2021</year>
          .
          <article-title>Three-Stream 3D/1D CNN for Fine-Grained Action Classification and Segmentation in Table Tennis</article-title>
          .
          <source>CoRR abs/2109</source>
          .14306 (
          <year>2021</year>
          ). arXiv:
          <volume>2109</volume>
          .14306 https://arxiv.org/abs/2109.14306
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [16]
          <string-name>
            <surname>Yurii</surname>
            <given-names>E.</given-names>
          </string-name>
          <string-name>
            <surname>Nesterov</surname>
          </string-name>
          .
          <year>2004</year>
          .
          <article-title>Introductory Lectures on Convex Optimization - A Basic Course</article-title>
          .
          <source>Applied Optimization</source>
          , Vol.
          <volume>87</volume>
          . Springer.
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [17]
          <string-name>
            <surname>Trong-Tung</surname>
            <given-names>Nguyen</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Thanh-Son</surname>
            <given-names>Nguyen</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gia-Bao Dinh</surname>
            <given-names>Ho</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hai-Dang Nguyen</surname>
          </string-name>
          , and
          <string-name>
            <surname>Minh-Triet Tran</surname>
          </string-name>
          .
          <year>2021</year>
          . HCMUS at MediaEval 2021:
          <article-title>Ensembles of Action Recognition Networks with Prior Knowledge for Table Tennis Strokes Classification Task</article-title>
          . In
          <source>MediaEval (CEUR Workshop Proceedings)</source>
          .
          <source>CEUR-WS.org.</source>
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          [18]
          <string-name>
            <surname>Yijun</surname>
            <given-names>Qian</given-names>
          </string-name>
          , Lijun Yu, Wenhe Liu, and
          <string-name>
            <surname>Alexander</surname>
            <given-names>G.</given-names>
          </string-name>
          <string-name>
            <surname>Hauptmann</surname>
          </string-name>
          .
          <year>2021</year>
          .
          <article-title>Learning Unbiased Transformer for Long-Tail Sports Action Classification</article-title>
          . In MediaEval (CEUR Workshop Proceedings).
          <source>CEUR-WS.org.</source>
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          [19]
          <string-name>
            <surname>Grégory</surname>
            <given-names>Rogez</given-names>
          </string-name>
          , Philippe Weinzaepfel, and
          <string-name>
            <given-names>Cordelia</given-names>
            <surname>Schmid</surname>
          </string-name>
          .
          <year>2020</year>
          .
          <article-title>LCR-Net++: Multi-Person 2D and 3D Pose Detection in Natural Images</article-title>
          .
          <source>IEEE Trans. Pattern Anal. Mach. Intell</source>
          .
          <volume>42</volume>
          ,
          <issue>5</issue>
          (
          <year>2020</year>
          ),
          <fpage>1146</fpage>
          -
          <lpage>1161</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          [20]
          <string-name>
            <given-names>Karen</given-names>
            <surname>Simonyan</surname>
          </string-name>
          and
          <string-name>
            <given-names>Andrew</given-names>
            <surname>Zisserman</surname>
          </string-name>
          .
          <year>2014</year>
          .
          <article-title>Two-Stream Convolutional Networks for Action Recognition in Videos</article-title>
          . In NIPS.
          <volume>568</volume>
          -
          <fpage>576</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          [21]
          <string-name>
            <surname>Xuanhan</surname>
            <given-names>Wang</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lianli Gao</surname>
          </string-name>
          , Peng Wang,
          <string-name>
            <surname>Xiaoshuai Sun</surname>
          </string-name>
          , and Xianglong Liu.
          <year>2018</year>
          .
          <article-title>Two-Stream 3-D convNet Fusion for Action Recognition in Videos With Arbitrary Size and Length</article-title>
          .
          <source>IEEE Trans. Multimedia</source>
          <volume>20</volume>
          ,
          <issue>3</issue>
          (
          <year>2018</year>
          ),
          <fpage>634</fpage>
          -
          <lpage>644</lpage>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>