<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Two Stream Network for Stroke Detection in Table Tennis</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Anam Zahra</string-name>
          <email>anam_zahra@eva.mpg.de</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Pierre-Etienne Martin</string-name>
          <email>pierre_etienne_martin@eva.mpg.de</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>CCP Department, Max Planck Institute for Evolutionary Anthropology</institution>
          ,
          <addr-line>D-04103 Leipzig</addr-line>
          ,
          <country country="DE">Germany</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2021</year>
      </pub-date>
      <fpage>13</fpage>
      <lpage>15</lpage>
      <abstract>
        <p>This paper presents a table tennis stroke detection method from videos. The method relies on a two-stream Convolutional Neural Network processing in parallel the RGB Stream and its computed optical flow. The method has been developed as part of the MediaEval 2021 benchmark for the Sport task. Our contribution did not outperform the provided baseline on the test set but has performed the best among the other participants with regard to the mAP metric.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>INTRODUCTION</title>
      <p>
        With the advent of Convolutional Neural Networks (CNNs),
especially after the success of AlexNet [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ], object detection,
localization, and classification from images and videos have greatly
progressed [
        <xref ref-type="bibr" rid="ref22 ref3 ref5 ref9">3, 5, 9, 22</xref>
        ].The development of computer vision
methods has motivated broader applications in the academic world. Our
team is currently working on egocentric recordings from children
in kindergarten and at home. The analysis of these recordings shall
give us an automatic overview of their interactions on a daily basis.
We hope to link these interactions with their cognitive development
and, thereby, better understand early child development. With our
participation in the Sports Video Task [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ], in the stroke detection
subtask, we hope to perfect our knowledge in event detection and
transpose it to our project.
      </p>
      <p>
        The diversity of applications and visual data in sport, makes
sports video analysis attractive for researchers. Automated sport
event detection and action classification, especially from low-resolution
videos, are helpful for monitoring and training purposes. For
example in [
        <xref ref-type="bibr" rid="ref10 ref6">6, 10</xref>
        ], the authors automate the performance analysis for
the training optimization of players. Similarly, Sports Video task at
MediaEval 2021 benchmark aims at improving athlete performance
and training experience through the first steps of stroke detection
and classification from videos.
      </p>
      <p>
        Event detection in videos is the first step to many other
hottopics such as video summarizing [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ], automated semantic
segmentation [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ] and action recognition [
        <xref ref-type="bibr" rid="ref16 ref4">4, 16</xref>
        ]. These methods may be
used to build summary, selecting highlights, and assisting players
in training sessions. One way to approach the problem of event
detection in sports with balls, can be through ball detection and
tracking. Several researchers have tried to get the 2D, and 3D ball
trajectories in order to achieve so [
        <xref ref-type="bibr" rid="ref17 ref18 ref20">17, 18, 20</xref>
        ].
      </p>
      <p>
        Inspired from [
        <xref ref-type="bibr" rid="ref14 ref15 ref21 ref8">8, 14, 15, 21</xref>
        ], this method combines the optical
lfow and features learned from the RGB stream in order to detect a
stroke in table tennis and assess its duration. This implementation
is an extension of the baseline code provided by the Sport Task
organizers [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ].
2
      </p>
    </sec>
    <sec id="sec-2">
      <title>APPROACH</title>
      <p>
        Initially, we sought to use ball detection and tracking to perform
stroke detection. The first implementation used the pretrained
model TTNet [
        <xref ref-type="bibr" rid="ref21">21</xref>
        ]. However, the model failed to adapt to the
acquisition conditions from TTStroke-21 [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ], on which the task is built
upon, and no fine-tuning was possible since no ball coordinates
are available in the provided annotations. Therefore we decided to
train a model from scratch.
      </p>
      <p>In this section, we first present the preparation of the videos
and then the model presenting the processed data. Both processes
are depicted in Fig. 1. Post processing is performed to form a final
decision.
2.1</p>
    </sec>
    <sec id="sec-3">
      <title>Data Preparation</title>
      <p>
        In video content analysis, the motion of objects of interest between
frames can be of significant interest in order to understand their
evolution in space. As such, we decided to use optical flow as a
modality to perform stroke detection. Inspired by [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ], we decided
to use DeepFlow method [
        <xref ref-type="bibr" rid="ref23">23</xref>
        ] to compute the optical flow from
consecutive frames. The optical flow is computed from frames
resized to 320 × 128. This size was initially chosen to keep the
ball at least two pixels big, as it has previously been done in [
        <xref ref-type="bibr" rid="ref21">21</xref>
        ].
Both the RGB and optical flow frames are consecutively stacked
in a tensor of length 75. As in [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ], stroke detection is tackled as a
classification problem with two classes: “Stroke” and “Non-stroke”.
2.2
      </p>
    </sec>
    <sec id="sec-4">
      <title>Model</title>
      <p>
        As shown in figure 1, our Two-Stream model is composed of two
branches of the same length. Each branch is a succession of four
blocks and each block is composed of a convolutional layer with
3 × 3 × 3 filters, followed by a ReLU activation function, and a
2 × 2 × 2 pooling layer. The output of each branch is then flattened
and fed to a fully connected layer that outputs a feature vector of
length 500. Both feature vectors are then concatenated and fed into
a final fully connected layer of length two to predict the “Stroke”
and “Non-storke” classes. One branch takes RGB frames of the
video and the other computed optical flow. The model is trained
using a stochastic gradient descent method over 250 epochs with a
learning rate of 0.001, a batch size of 10, a weight decay of 0.005, and
a Nesterov momentum [
        <xref ref-type="bibr" rid="ref19">19</xref>
        ] of 0.5. The negative samples creation
and input processing is the same as the baseline [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ].
2.3
      </p>
    </sec>
    <sec id="sec-5">
      <title>Post Processing</title>
      <p>Our model classifies 75 consecutive frames. In order to create stroke
segments over the whole video, we classify every 75 frames of the
videos, which leads to applying a sliding window without overlap.
If two consecutive segments are classified as stroke, the segments
are fused to create only one stroke.
3</p>
    </sec>
    <sec id="sec-6">
      <title>RESULTS AND ANALYSIS</title>
      <p>
        The metrics for evaluating the detection performance are described
in [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ]. Our approach reached a mean Average Precision (mAP)
of 0.00124 and a Global Intersection over Union (G-IoU) of 0.0700.
It falls behind the baseline which reaches respectively 0.0173 and
0.144. Our other attempts using early concatenation of the RGB
and Optical Flow modalities - meaning an input of size 5 × 320 × 128
in one branch model - or training method without shufling of the
data, reached even lesser performance.
      </p>
      <p>Nevertheless, from a classification point of view, and according
to the Fig. 2, our model learned the stroke features and can perform
reasonable results when stroke boundaries are known: 86.4% of
accuracy on the validation set after only 60 epochs. Which may
indicates that the main failure is coming from the post processing
method.</p>
      <p>Indeed, by looking at the stroke distribution across the diferent
sets, see table 1, we may notice how little the inferred stroke ratio is
on the test set: 0.57 strokes for 1000 frames, whereas the stroke rate
is 1.85 and 2.28 for 1000 frames in the training and validation sets.
Furthermore, our post processing was not limited in term of stroke
duration, leading to everlasting strokes: 4500 frames - meaning the
fusions of 60 consecutive video segments. These points indicate
that our post-processing method can be improved.</p>
      <p>
        A better separation of the stroke may be reached by defining
the event using ball tracking and the ball motion [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]. This was our
initial attempt, inspired by [
        <xref ref-type="bibr" rid="ref21">21</xref>
        ], but the available pretrained model
considers a diferent point of view and was unable to adapt to the
TTStroke-21 videos point of view.
4
      </p>
    </sec>
    <sec id="sec-7">
      <title>CONCLUSION</title>
      <p>The Sports Video Task, and more specifically the stroke detection
subtask, has proven to be challenging. Even if our implementation
has learned to classify strokes, we were not able to outperform
the baseline performance. We have underlined the importance of
the post processing step through a stroke concentration and
duration analysis. Furthermore, our failure to adapt a pretrained model
on similar dataset, but with a diferent acquisition point of view,
stresses the dificulty of the deep trained models to adapt to a
change of scene, which is inherent to the fine-grained aspect of
the classification subtask. As first time participants, we thought
to tackle only one task to ease our submission. However, we now
believe that a method tackling both the detection and classification
may be the best for solving the Sport Video subtasks.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>Lamberto</given-names>
            <surname>Ballan</surname>
          </string-name>
          , Marco Bertini,
          <source>Alberto Del Bimbo</source>
          ,
          <string-name>
            <given-names>Lorenzo</given-names>
            <surname>Seidenari</surname>
          </string-name>
          , and
          <string-name>
            <given-names>Giuseppe</given-names>
            <surname>Serra</surname>
          </string-name>
          .
          <year>2011</year>
          .
          <article-title>Event detection and recognition for semantic annotation of video</article-title>
          .
          <source>Multimedia tools and applications 51</source>
          ,
          <issue>1</issue>
          (
          <year>2011</year>
          ),
          <fpage>279</fpage>
          -
          <lpage>302</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>Jordan</given-names>
            <surname>Calandre</surname>
          </string-name>
          , Renaud Péteri, Laurent Mascarilla, and
          <string-name>
            <given-names>Benoit</given-names>
            <surname>Tremblais</surname>
          </string-name>
          .
          <year>2021</year>
          .
          <article-title>Table Tennis ball kinematic parameters estimation from non-intrusive single-view videos</article-title>
          .
          <source>In 2021 International Conference on Content-Based Multimedia Indexing (CBMI)</source>
          .
          <source>IEEE</source>
          , 1-
          <fpage>6</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>Joao</given-names>
            <surname>Carreira</surname>
          </string-name>
          and
          <string-name>
            <given-names>Andrew</given-names>
            <surname>Zisserman</surname>
          </string-name>
          .
          <year>2017</year>
          .
          <article-title>Quo vadis, action recognition? a new model and the kinetics dataset</article-title>
          .
          <source>In proceedings of the IEEE Conference on Computer Vision and Pattern Recognition</source>
          .
          <fpage>6299</fpage>
          -
          <lpage>6308</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <surname>Chandni</surname>
            <given-names>J</given-names>
          </string-name>
          <string-name>
            <surname>Dhamsania and Tushar V Ratanpara</surname>
          </string-name>
          .
          <year>2016</year>
          .
          <article-title>A survey on human action recognition from videos</article-title>
          .
          <article-title>In 2016 online international conference on green engineering and technologies (IC-GET)</article-title>
          .
          <source>IEEE</source>
          , 1-
          <fpage>5</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>Ross</given-names>
            <surname>Girshick</surname>
          </string-name>
          .
          <year>2015</year>
          .
          <article-title>Fast r-cnn</article-title>
          .
          <source>In Proceedings of the IEEE international conference on computer vision</source>
          . 1440-
          <fpage>1448</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <surname>Mike</surname>
            <given-names>D</given-names>
          </string-name>
          <string-name>
            <surname>Hughes and Roger M Bartlett</surname>
          </string-name>
          .
          <year>2002</year>
          .
          <article-title>The use of performance indicators in performance analysis</article-title>
          .
          <source>Journal of sports sciences 20</source>
          ,
          <issue>10</issue>
          (
          <year>2002</year>
          ),
          <fpage>739</fpage>
          -
          <lpage>754</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <surname>Yasmin</surname>
            <given-names>S Khan</given-names>
          </string-name>
          and
          <string-name>
            <given-names>Soudamini</given-names>
            <surname>Pawar</surname>
          </string-name>
          .
          <year>2015</year>
          .
          <article-title>Video summarization: survey on event detection and summarization in soccer videos</article-title>
          .
          <source>International Journal of Advanced Computer Science and Applications</source>
          <volume>6</volume>
          ,
          <issue>11</issue>
          (
          <year>2015</year>
          ),
          <fpage>256</fpage>
          -
          <lpage>259</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>Gregory</given-names>
            <surname>Koch</surname>
          </string-name>
          , Richard Zemel, Ruslan Salakhutdinov, and others.
          <source>2015</source>
          .
          <article-title>Siamese neural networks for one-shot image recognition</article-title>
          .
          <source>In ICML deep learning workshop</source>
          , Vol.
          <volume>2</volume>
          . Lille.
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>Alex</given-names>
            <surname>Krizhevsky</surname>
          </string-name>
          , Ilya Sutskever, and
          <string-name>
            <given-names>Geofrey E</given-names>
            <surname>Hinton</surname>
          </string-name>
          .
          <year>2012</year>
          .
          <article-title>Imagenet classification with deep convolutional neural networks</article-title>
          .
          <source>Advances in neural information processing systems</source>
          <volume>25</volume>
          (
          <year>2012</year>
          ),
          <fpage>1097</fpage>
          -
          <lpage>1105</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>Adrian</given-names>
            <surname>Lees</surname>
          </string-name>
          .
          <year>2003</year>
          .
          <article-title>Science and the major racket sports: a review</article-title>
          .
          <source>Journal of sports sciences 21</source>
          ,
          <issue>9</issue>
          (
          <year>2003</year>
          ),
          <fpage>707</fpage>
          -
          <lpage>732</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <surname>Pierre-Etienne Martin</surname>
          </string-name>
          .
          <year>2021</year>
          .
          <article-title>Spatio-Temporal CNN baseline method for the Sports Video Task of MediaEval 2021 benchmark</article-title>
          . In
          <source>MediaEval (CEUR Workshop Proceedings)</source>
          .
          <source>CEUR-WS.org.</source>
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <surname>Pierre-Etienne</surname>
            <given-names>Martin</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Jenny</surname>
          </string-name>
          Benois-Pineau, Boris Mansencal, Renaud Péteri, Laurent Mascarilla,
          <string-name>
            <surname>Jordan Calandre</surname>
            , and
            <given-names>Julien</given-names>
          </string-name>
          <string-name>
            <surname>Morlier</surname>
          </string-name>
          .
          <year>2021</year>
          .
          <article-title>Sports Video: Fine-Grained Action Detection and Classification of Table Tennis Strokes from videos for MediaEval 2021</article-title>
          . In MediaEval (CEUR Workshop Proceedings).
          <source>CEUR-WS.org.</source>
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <surname>Pierre-Etienne</surname>
            <given-names>Martin</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Jenny</surname>
            Benois-Pineau,
            <given-names>Renaud</given-names>
          </string-name>
          <string-name>
            <surname>Péteri</surname>
            , and
            <given-names>Julien</given-names>
          </string-name>
          <string-name>
            <surname>Morlier</surname>
          </string-name>
          .
          <year>2018</year>
          .
          <article-title>Sport Action Recognition with Siamese Spatio-Temporal CNNs: Application to Table Tennis</article-title>
          .
          <source>In CBMI. IEEE</source>
          , 1-
          <fpage>6</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <surname>Pierre-Etienne</surname>
            <given-names>Martin</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Jenny</surname>
            Benois-Pineau,
            <given-names>Renaud</given-names>
          </string-name>
          <string-name>
            <surname>Péteri</surname>
            , and
            <given-names>Julien</given-names>
          </string-name>
          <string-name>
            <surname>Morlier</surname>
          </string-name>
          .
          <year>2019</year>
          .
          <article-title>Optimal Choice of Motion Estimation Methods for Fine-Grained Action Classification with 3D Convolutional Networks</article-title>
          .
          <source>In 2019 IEEE International Conference on Image Processing, ICIP</source>
          <year>2019</year>
          , Taipei, Taiwan,
          <source>September 22-25</source>
          ,
          <year>2019</year>
          . IEEE,
          <fpage>554</fpage>
          -
          <lpage>558</lpage>
          . https://doi.org/ 10.1109/ICIP.
          <year>2019</year>
          .8803780
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <surname>Pierre-Etienne</surname>
            <given-names>Martin</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Jenny</surname>
            Benois-Pineau,
            <given-names>Renaud</given-names>
          </string-name>
          <string-name>
            <surname>Péteri</surname>
            , and
            <given-names>Julien</given-names>
          </string-name>
          <string-name>
            <surname>Morlier</surname>
          </string-name>
          .
          <year>2020</year>
          .
          <article-title>3D attention mechanisms in Twin Spatio-Temporal Convolutional Neural Networks</article-title>
          .
          <article-title>Application to action classification in videos of table tennis games.</article-title>
          .
          <source>In 25th International Conference on Pattern Recognition (ICPR2020</source>
          <string-name>
            <surname>) - MiCo Milano Congress Center</surname>
          </string-name>
          , Italy,
          <fpage>10</fpage>
          -
          <lpage>15</lpage>
          January
          <year>2021</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [16]
          <string-name>
            <surname>Pierre-Etienne</surname>
            <given-names>Martin</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Jenny</surname>
            Benois-Pineau,
            <given-names>Renaud</given-names>
          </string-name>
          <string-name>
            <surname>Péteri</surname>
            , and
            <given-names>Julien</given-names>
          </string-name>
          <string-name>
            <surname>Morlier</surname>
          </string-name>
          .
          <year>2020</year>
          .
          <article-title>Fine grained sport action recognition with twin spatiotemporal convolutional neural networks</article-title>
          .
          <source>Multimedia Tools and Applications</source>
          <volume>79</volume>
          ,
          <issue>27</issue>
          (
          <year>2020</year>
          ),
          <fpage>20429</fpage>
          -
          <lpage>20447</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [17]
          <string-name>
            <surname>Hnin</surname>
            <given-names>Myint</given-names>
          </string-name>
          , Patrick Wong, Laurence Dooley, and
          <string-name>
            <given-names>Adrian</given-names>
            <surname>Hopgood</surname>
          </string-name>
          .
          <year>2015</year>
          .
          <article-title>Tracking a table tennis ball for umpiring purposes</article-title>
          .
          <source>In 2015 14th IAPR International Conference on Machine Vision Applications</source>
          (MVA). IEEE,
          <fpage>170</fpage>
          -
          <lpage>173</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          [18]
          <string-name>
            <surname>Hnin</surname>
            <given-names>Myint</given-names>
          </string-name>
          , Patrick Wong, Laurence Dooley, and
          <string-name>
            <given-names>Adrian</given-names>
            <surname>Hopgood</surname>
          </string-name>
          .
          <year>2016</year>
          .
          <article-title>Tracking a table tennis ball for umpiring purposes using a multi-agent system</article-title>
          . (
          <year>2016</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          [19]
          <string-name>
            <surname>Ilya</surname>
            <given-names>Sutskever</given-names>
          </string-name>
          , James Martens, George Dahl, and
          <string-name>
            <given-names>Geofrey</given-names>
            <surname>Hinton</surname>
          </string-name>
          .
          <year>2013</year>
          .
          <article-title>On the importance of initialization and momentum in deep learning</article-title>
          .
          <source>In International conference on machine learning. PMLR</source>
          ,
          <fpage>1139</fpage>
          -
          <lpage>1147</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          [20]
          <string-name>
            <given-names>Sho</given-names>
            <surname>Tamaki</surname>
          </string-name>
          and
          <string-name>
            <given-names>Hideo</given-names>
            <surname>Saito</surname>
          </string-name>
          .
          <year>2013</year>
          .
          <article-title>Reconstruction of 3d trajectories for performance analysis in table tennis</article-title>
          .
          <source>In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition Workshops</source>
          .
          <fpage>1019</fpage>
          -
          <lpage>1026</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          [21]
          <string-name>
            <surname>Roman</surname>
            <given-names>Voeikov</given-names>
          </string-name>
          , Nikolay Falaleev, and
          <string-name>
            <given-names>Ruslan</given-names>
            <surname>Baikulov</surname>
          </string-name>
          .
          <year>2020</year>
          .
          <article-title>TTNet: Real-time temporal and spatial video analysis of table tennis</article-title>
          .
          <source>In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops</source>
          .
          <fpage>884</fpage>
          -
          <lpage>885</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          [22]
          <string-name>
            <given-names>Heng</given-names>
            <surname>Wang</surname>
          </string-name>
          and
          <string-name>
            <given-names>Cordelia</given-names>
            <surname>Schmid</surname>
          </string-name>
          .
          <year>2013</year>
          .
          <article-title>Action recognition with improved trajectories</article-title>
          .
          <source>In Proceedings of the IEEE international conference on computer vision</source>
          . 3551-
          <fpage>3558</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          [23]
          <string-name>
            <surname>Philippe</surname>
            <given-names>Weinzaepfel</given-names>
          </string-name>
          , Jerome Revaud, Zaid Harchaoui, and
          <string-name>
            <given-names>Cordelia</given-names>
            <surname>Schmid</surname>
          </string-name>
          .
          <year>2013</year>
          .
          <article-title>DeepFlow: Large Displacement Optical Flow with Deep Matching</article-title>
          .
          <source>In 2013 IEEE International Conference on Computer Vision</source>
          . 1385-
          <fpage>1392</fpage>
          . https://doi.org/10.1109/ICCV.
          <year>2013</year>
          .175
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>