<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Sports Video: Fine-Grained Action Detection and Classification of Table Tennis Strokes from Videos for MediaEval 2021</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Pierre-Etienne Martin</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Jordan Calandre</string-name>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Boris Mansencal</string-name>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Jenny Benois-Pineau</string-name>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Renaud Péteri</string-name>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Laurent Mascarilla</string-name>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Julien Morlier</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>CCP Department, Max Planck Institute for Evolutionary Anthropology</institution>
          ,
          <addr-line>D-04103 Leipzig</addr-line>
          ,
          <country country="DE">Germany</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>IMS, University of Bordeaux</institution>
          ,
          <addr-line>Talence</addr-line>
          ,
          <country country="FR">France</country>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>MIA, La Rochelle University</institution>
          ,
          <addr-line>La Rochelle</addr-line>
          ,
          <country country="FR">France</country>
        </aff>
        <aff id="aff3">
          <label>3</label>
          <institution>Univ. Bordeaux</institution>
          ,
          <addr-line>CNRS, Bordeaux INP, LaBRI, Talence</addr-line>
          ,
          <country country="FR">France</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2021</year>
      </pub-date>
      <fpage>13</fpage>
      <lpage>15</lpage>
      <abstract>
        <p>Sports video analysis is a prevalent research topic due to the variety of application areas, ranging from multimedia intelligent devices with user-tailored digests up to analysis of athletes' performance. The Sports Video task is part of the MediaEval 2021 benchmark. This task tackles fine-grained action detection and classification from videos. The focus is on recordings of table tennis games. Running since 2019, the task has ofered a classification challenge from untrimmed video recorded in natural conditions with known temporal boundaries for each stroke. This year, the dataset is extended and ofers, in addition, a detection challenge from untrimmed videos without annotations. This work aims at creating tools for sports coaches and players in order to analyze sports performance. Movement analysis and player profiling may be built upon such technology to enrich the training experience of athletes and improve their performance.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>INTRODUCTION</title>
      <p>
        Action detection and classification are one of the main challenges
in computer vision [
        <xref ref-type="bibr" rid="ref20">20</xref>
        ]. Over the last few years, the number of
datasets and their complexity dedicated to action classification has
drastically increased [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ]. Sports video analysis is one branch of
computer vision and applications in this area range from
multimedia intelligent devices with user-tailored digests, up to analysis of
athletes’ performance [
        <xref ref-type="bibr" rid="ref21 ref28 ref4">4, 21, 28</xref>
        ]. A large amount of work is devoted
to the analysis of sports gestures using motion capture systems.
However, body-worn sensors and markers could disturb the natural
behavior of sports players. This issue motivates the development
of methods for game analysis using non-invasive equipment such
as video recordings from cameras.
      </p>
      <p>The Sports Video Classification project was initiated by the
Sports Faculty (STAPS) and the computer science laboratory LaBRI
of the University of Bordeaux, and the MIA laboratory of La
Rochelle University1. This project aims to develop artificial
intelligence and multimedia indexing methods for the recognition of
table tennis activities. The ultimate goal is to evaluate the
performance of athletes, with a particular focus on students, to develop
1This work was supported by the New Aquitania Region through CRISP project
ComputeR vIsion for Sports Performance and the MIRES federation.
optimal training strategies. To that aim, the video corpus named
TTStroke-21 was recorded with volunteer players.</p>
      <p>
        Datasets such as UCF-101 [
        <xref ref-type="bibr" rid="ref26">26</xref>
        ], HMDB [
        <xref ref-type="bibr" rid="ref6 ref8">6, 8</xref>
        ], AVA [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ] and
Kinetics [
        <xref ref-type="bibr" rid="ref1 ref2 ref25 ref7 ref9">1, 2, 7, 9, 25</xref>
        ] are being use in the scope of action recognition
with, year after year, an increasing number of considered videos
and classes. Few datasets focus on fine-grained classification in
sports such as FineGym [
        <xref ref-type="bibr" rid="ref24">24</xref>
        ] and TTStroke21 [
        <xref ref-type="bibr" rid="ref19">19</xref>
        ].
      </p>
      <p>
        To tackle the increasing complexity of the datasets, we have on
one hand methods getting the most of the temporal information: for
example, in [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ], where spatio-temporal dependencies are learned
from the video using only RGB data. And on the other hand,
methods combining other modalities extracted from videos, such as the
optical flow [
        <xref ref-type="bibr" rid="ref18 ref27 ref3">3, 18, 27</xref>
        ]. The inter-similarity of actions - strokes - in
TTStroke-21 makes the classification task challenging, and both
cited aspects shall be used to improve performance.
      </p>
      <p>The following sections present the Sport task this year and its
specific terms of use. Complementary information on the task may
be found on the dedicated page from the MediaEval website2.
2</p>
    </sec>
    <sec id="sec-2">
      <title>TASK DESCRIPTION</title>
      <p>
        This task uses the TTStroke-21 database [
        <xref ref-type="bibr" rid="ref19">19</xref>
        ]. This dataset is
constituted of recordings of table tennis players performing in natural
conditions. This task ofers researchers an opportunity to test their
ifne-grained classification methods for detecting and classifying
strokes in table tennis videos. Compared to the Sports Video 2020
edition, this year, we extend the task with detection, and enrich the
data set with new and more diverse stroke samples. The task now
ofers two subtasks. Each subtask has its own split of the dataset,
leading to diferent train, validation, and test sets.
      </p>
      <p>
        Participants can choose to participate in only one or both
subtasks and submit up to five runs for each. The participants must
provide one XML file per video file present in the test set for each
run. The content of the XML file varies according to the subtask.
Runs may be submitted as an archive (zip file), with each run in a
diferent directory for each subtask. Participants should also submit
a working notes paper, which describes their method and indicates
if any external data, such as other datasets or pretrained networks,
was used to compute their runs. The task is considered fully
automatic: once the videos are provided to the system, results should be
produced without human intervention. Participants are encouraged
to release their code publicly with their submission. This year, a
baseline for both subtasks was shared publicly [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ].
2https://multimediaeval.github.io/editions/2021/tasks/sportsvideo/
2.1
      </p>
    </sec>
    <sec id="sec-3">
      <title>Subtask 1 - Stroke Detection</title>
      <p>Participants must build a system that detects whether a stroke
has been performed, whatever its class, and extract its temporal
boundaries. The aim is to distinguish between moments of interest
in a game (players performing strokes) from irrelevant moments
(time between strokes, picking up the ball, having a break. . . ). This
subtask can be a preliminary step for later recognizing a stroke that
has been performed.</p>
      <p>Participants have to segment regions where a stroke is performed
in the provided videos. Provided XML files contain the stroke
temporal boundaries (frame index of the videos) related to the train and
validation sets. We invite the participants to fill an XML file for each
test video in which each stroke should be temporally segmented
frame-wise following the same structure.</p>
      <p>
        For this subtask, the videos are not shared across train, validation,
and test sets; however, a same player may appear in the diferent
sets. The Intersection over Union (IoU) and Average Precision (AP)
metrics will be used for evaluation. Both are usually used for image
segmentation but are adapted for this task:
• Global IoU: the frame-wise overlap between the ground
truth and the predicted strokes across all the videos.
• Instance AP: each stroke represents an instance to be
detected. Detection is considered True when the IoU between
prediction and ground truth is above an IoU threshold. 20
thresholds from 0.5 to 0.95 with a step of 0.05 are
considered, similarly to the COCO challenge [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ]. This metric
will be used for the final ranking of participants.
2.2
      </p>
    </sec>
    <sec id="sec-4">
      <title>Subtask 2 - Stroke Classification</title>
      <p>
        This subtask is similar to the main task of the previous edition [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ].
This year the dataset is extended, and a validation set is provided.
      </p>
      <p>Participants are required to build a classification system that
automatically labels video segments according to a performed stroke.
There are 20 possible stroke classes. The temporal boundaries of
each stroke are supplied in the XML files accompanying each video
in each set. The XML files dedicated to the train and validation sets
contain the stroke class as a label, while in the test set, the label
is set to “Unknown”. Hence for each XML file in the test set, the
participants are invited to replace the default label “Unknown” by
the stroke class that the participant’s system has assigned according
to the given taxonomy.</p>
      <p>For this subtask, the videos are shared across the sets following
a random distribution of all the strokes with the proportions of
60%, 20% and 20% respectively for the train, validation and test sets.
All submissions will be evaluated in terms of global accuracy for
ranking and detailed with per-class accuracy.</p>
      <p>
        Last year, the best global accuracy (31.4%) was obtained by [
        <xref ref-type="bibr" rid="ref22">22</xref>
        ]
using Channel-Separated CNN. [
        <xref ref-type="bibr" rid="ref17">17</xref>
        ] is second (26.6%) using 3D
attention mechanism and [
        <xref ref-type="bibr" rid="ref23">23</xref>
        ] third (16.7%) using pose information
and cascade labelling method. Improvement has been observed
compared to the previous edition [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ] with a best accuracy of 22.9% [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ].
This improvement seems to be correlated by various factors such
as: i) multi-modal methods, ii) deeper and more complex CNN
capturing simultaneously spatial and temporal features, and iii) class
decision following a cascade method.
      </p>
    </sec>
    <sec id="sec-5">
      <title>DATASET DESCRIPTION</title>
      <p>The dataset has been recorded at the STAPS using lightweight
equipment. It is constituted of player-centered videos recorded in
natural conditions without markers or sensors, see Fig 1.
Professional table tennis teachers designed a dedicated taxonomy. The
dataset comprises 20 table tennis stroke classes: height services, six
ofensive strokes, and six defensive strokes. The strokes may be
divided in two super-classes: Forehand and Backhand.</p>
      <p>All videos are recorded in MPEG-4 format. We blurred fhe faces
of the players for each original video frame using OpenCV deep
learning face detector, based on the Single Shot Detector (SSD)
framework with a ResNet base network. A tracking method has
been implemented to decrease the false positive rate. The detected
faces are blurred, and the video is re-encoded in MPEG-4.</p>
      <p>Compared with Sports Video 2020 edition, this year, the data set
is enriched with new and more diverse video samples. A total of
100 minutes of table tennis games across 28 videos recorded at 120
frames per second is considered. It represents more than 718 000
frames in HD (1920 × 1080). An additional validation set is also
provided for better comparison across participants. This set may be
used for training when submitting the test set’s results. Twenty-two
videos are used for the Stroke Classification subtask, representing
1017 strokes randomly distributed in the diferent sets following the
previously given proportions. The same videos are used in the train
and validation sets of the Segmentation subtask, and six additional
videos, without annotations, are dedicated to its test set.
4</p>
    </sec>
    <sec id="sec-6">
      <title>SPECIFIC TERMS OF USE</title>
      <p>Although faces are automatically blurred to preserve anonymity,
some faces are misdetected, and thus some players remain
identifiable. In order to respect the personal data of the players, this dataset
is subject to a usage agreement, referred to as Special Conditions.
The complete acceptance of these Special Conditions is a mandatory
prerequisite for the provision of the Images as part of the MediaEval
2021 evaluation campaign. A complete reading of these conditions
is necessary and requires the user, for example, to obscure the faces
(blurring, black banner) in the video before use in any publication
and to destroy the data by October 1st, 2022.
5</p>
    </sec>
    <sec id="sec-7">
      <title>DISCUSSIONS</title>
      <p>This year the Sports Video task of MediaEval proposes two subtasks:
i) Detection and ii) Classification of strokes from videos. Even if
the players’ faces are blurred, the provided videos still fall under
particular usage conditions that the participants need to accept.
Participants are encouraged to share their dificulties and their
results even if they seem not suficiently good. All the investigations,
even when not successful, may inspire future methods.</p>
    </sec>
    <sec id="sec-8">
      <title>ACKNOWLEDGMENTS</title>
      <p>Many thanks to the players, coaches, and annotators who
contributed to TTStroke-21.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>João</given-names>
            <surname>Carreira</surname>
          </string-name>
          , Eric Noland, Andras Banki-Horvath,
          <string-name>
            <given-names>Chloe</given-names>
            <surname>Hillier</surname>
          </string-name>
          , and
          <string-name>
            <given-names>Andrew</given-names>
            <surname>Zisserman</surname>
          </string-name>
          .
          <year>2018</year>
          . A Short Note about Kinetics-
          <volume>600</volume>
          . CoRR abs/
          <year>1808</year>
          .01340 (
          <year>2018</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>João</given-names>
            <surname>Carreira</surname>
          </string-name>
          , Eric Noland, Chloe Hillier, and
          <string-name>
            <given-names>Andrew</given-names>
            <surname>Zisserman</surname>
          </string-name>
          .
          <year>2019</year>
          .
          <article-title>A Short Note on the Kinetics-700 Human Action Dataset</article-title>
          . CoRR abs/
          <year>1907</year>
          .06987 (
          <year>2019</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>João</given-names>
            <surname>Carreira</surname>
          </string-name>
          and
          <string-name>
            <given-names>Andrew</given-names>
            <surname>Zisserman</surname>
          </string-name>
          .
          <year>2017</year>
          .
          <article-title>Quo Vadis, Action Recognition? A New Model and the Kinetics Dataset</article-title>
          .
          <source>In CVPR. IEEE Computer Society</source>
          ,
          <fpage>4724</fpage>
          -
          <lpage>4733</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>Moritz</given-names>
            <surname>Einfalt</surname>
          </string-name>
          , Dan Zecha, and
          <string-name>
            <given-names>Rainer</given-names>
            <surname>Lienhart</surname>
          </string-name>
          .
          <year>2018</year>
          .
          <article-title>Activity-Conditioned Continuous Human Pose Estimation for Performance Analysis of Athletes Using the Example of Swimming</article-title>
          . In WACV.
          <volume>446</volume>
          -
          <fpage>455</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>Chunhui</given-names>
            <surname>Gu</surname>
          </string-name>
          , Chen Sun,
          <string-name>
            <given-names>David A.</given-names>
            <surname>Ross</surname>
          </string-name>
          , Carl Vondrick, Caroline Pantofaru,
          <string-name>
            <given-names>Yeqing</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Sudheendra</given-names>
            <surname>Vijayanarasimhan</surname>
          </string-name>
          , George Toderici, Susanna Ricco, Rahul Sukthankar, Cordelia Schmid, and
          <string-name>
            <given-names>Jitendra</given-names>
            <surname>Malik</surname>
          </string-name>
          .
          <year>2018</year>
          .
          <article-title>AVA: A Video Dataset of Spatio-Temporally Localized Atomic Visual Actions</article-title>
          . (
          <year>2018</year>
          ),
          <fpage>6047</fpage>
          -
          <lpage>6056</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>Hueihan</given-names>
            <surname>Jhuang</surname>
          </string-name>
          , Juergen Gall, Silvia Zufi, Cordelia Schmid, and
          <string-name>
            <given-names>Michael J.</given-names>
            <surname>Black</surname>
          </string-name>
          .
          <year>2013</year>
          .
          <article-title>Towards Understanding Action Recognition</article-title>
          .
          <source>In ICCV. IEEE Computer Society</source>
          ,
          <fpage>3192</fpage>
          -
          <lpage>3199</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>Will</given-names>
            <surname>Kay</surname>
          </string-name>
          , João Carreira, Karen Simonyan, Brian Zhang, Chloe Hillier, Sudheendra Vijayanarasimhan, Fabio Viola, Tim Green, Trevor Back, Paul Natsev, Mustafa Suleyman, and
          <string-name>
            <given-names>Andrew</given-names>
            <surname>Zisserman</surname>
          </string-name>
          .
          <year>2017</year>
          .
          <article-title>The Kinetics Human Action Video Dataset</article-title>
          .
          <source>CoRR abs/1705</source>
          .06950 (
          <year>2017</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>Hildegard</given-names>
            <surname>Kuehne</surname>
          </string-name>
          , Hueihan Jhuang, Estíbaliz Garrote, Tomaso A.
          <string-name>
            <surname>Poggio</surname>
            , and
            <given-names>Thomas</given-names>
          </string-name>
          <string-name>
            <surname>Serre</surname>
          </string-name>
          .
          <year>2011</year>
          .
          <article-title>HMDB: A large video database for human motion recognition</article-title>
          .
          <source>In ICCV. IEEE Computer Society</source>
          ,
          <fpage>2556</fpage>
          -
          <lpage>2563</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>Ang</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Meghana</given-names>
            <surname>Thotakuri</surname>
          </string-name>
          ,
          <string-name>
            <given-names>David A.</given-names>
            <surname>Ross</surname>
          </string-name>
          , João Carreira, Alexander Vostrikov, and
          <string-name>
            <given-names>Andrew</given-names>
            <surname>Zisserman</surname>
          </string-name>
          .
          <year>2020</year>
          .
          <article-title>The AVA-Kinetics Localized Human Actions Video Dataset</article-title>
          . CoRR abs/
          <year>2005</year>
          .00214 (
          <year>2020</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <surname>Tsung-Yi Lin</surname>
            ,
            <given-names>Michael</given-names>
          </string-name>
          <string-name>
            <surname>Maire</surname>
            ,
            <given-names>Serge J.</given-names>
          </string-name>
          <string-name>
            <surname>Belongie</surname>
            , James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and
            <given-names>C. Lawrence</given-names>
          </string-name>
          <string-name>
            <surname>Zitnick</surname>
          </string-name>
          .
          <year>2014</year>
          .
          <string-name>
            <surname>Microsoft</surname>
            <given-names>COCO</given-names>
          </string-name>
          :
          <article-title>Common Objects in Context</article-title>
          . In Computer Vision - ECCV 2014 - 13th European Conference, Zurich, Switzerland, September 6-
          <issue>12</issue>
          ,
          <year>2014</year>
          , Proceedings,
          <string-name>
            <surname>Part V</surname>
          </string-name>
          (Lecture Notes in Computer Science),
          <string-name>
            <given-names>David J.</given-names>
            <surname>Fleet</surname>
          </string-name>
          , Tomás Pajdla,
          <source>Bernt Schiele, and Tinne Tuytelaars (Eds.)</source>
          , Vol.
          <volume>8693</volume>
          . Springer,
          <fpage>740</fpage>
          -
          <lpage>755</lpage>
          . https://doi.org/10.1007/ 978-3-
          <fpage>319</fpage>
          -10602-1_
          <fpage>48</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>Zheng</given-names>
            <surname>Liu</surname>
          </string-name>
          and
          <string-name>
            <given-names>Haifeng</given-names>
            <surname>Hu</surname>
          </string-name>
          .
          <year>2019</year>
          .
          <article-title>Spatiotemporal Relation Networks for Video Action Recognition</article-title>
          .
          <source>IEEE Access</source>
          <volume>7</volume>
          (
          <year>2019</year>
          ),
          <fpage>14969</fpage>
          -
          <lpage>14976</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <surname>Pierre-Etienne Martin</surname>
          </string-name>
          .
          <year>2020</year>
          .
          <article-title>Fine-Grained Action Detection and Classification from Videos with Spatio-Temporal Convolutional Neural Networks</article-title>
          .
          <article-title>Application to Table Tennis. (Détection et classification fines d'actions à partir de vidéos par réseaux de neurones à convolutions spatio-temporelles</article-title>
          .
          <source>Application au tennis de table)</source>
          .
          <source>Ph.D. Dissertation</source>
          . University of La Rochelle, France. https://tel.archives-ouvertes.fr/ tel-03128769
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <surname>Pierre-Etienne Martin</surname>
          </string-name>
          .
          <year>2021</year>
          .
          <article-title>Spatio-Temporal CNN baseline method for the Sports Video Task of MediaEval 2021 benchmark</article-title>
          . In
          <source>MediaEval (CEUR Workshop Proceedings)</source>
          .
          <source>CEUR-WS.org.</source>
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <surname>Pierre-Etienne</surname>
            <given-names>Martin</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Jenny</surname>
          </string-name>
          Benois-Pineau, Boris Mansencal, Renaud Péteri, Laurent Mascarilla,
          <string-name>
            <surname>Jordan Calandre</surname>
            , and
            <given-names>Julien</given-names>
          </string-name>
          <string-name>
            <surname>Morlier</surname>
          </string-name>
          .
          <year>2019</year>
          .
          <article-title>Sports Video Annotation: Detection of Strokes in Table Tennis Task for MediaEval 2019</article-title>
          . In
          <source>MediaEval (CEUR Workshop Proceedings)</source>
          , Vol.
          <volume>2670</volume>
          . CEUR-WS.org.
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <surname>Pierre-Etienne</surname>
            <given-names>Martin</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Jenny</surname>
          </string-name>
          Benois-Pineau, Boris Mansencal, Renaud Péteri, Laurent Mascarilla,
          <string-name>
            <surname>Jordan Calandre</surname>
            , and
            <given-names>Julien</given-names>
          </string-name>
          <string-name>
            <surname>Morlier</surname>
          </string-name>
          .
          <year>2020</year>
          .
          <article-title>Sports Video Classification: Classification of Strokes in Table Tennis for MediaEval 2020</article-title>
          . In
          <source>MediaEval (CEUR Workshop Proceedings)</source>
          , Vol.
          <volume>2882</volume>
          . CEUR-WS.org.
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [16]
          <string-name>
            <surname>Pierre-Etienne</surname>
            <given-names>Martin</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Jenny</surname>
            Benois-Pineau, Boris Mansencal, Renaud Péteri, and
            <given-names>Julien</given-names>
          </string-name>
          <string-name>
            <surname>Morlier</surname>
          </string-name>
          .
          <year>2019</year>
          .
          <article-title>Siamese Spatio-Temporal Convolutional Neural Network for Stroke Classification in Table Tennis Games</article-title>
          . In
          <source>MediaEval (CEUR Workshop Proceedings)</source>
          , Vol.
          <volume>2670</volume>
          . CEUR-WS.org.
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [17]
          <string-name>
            <surname>Pierre-Etienne</surname>
            <given-names>Martin</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Jenny</surname>
            Benois-Pineau, Boris Mansencal, Renaud Péteri, and
            <given-names>Julien</given-names>
          </string-name>
          <string-name>
            <surname>Morlier</surname>
          </string-name>
          .
          <year>2020</year>
          .
          <article-title>Classification of Strokes in Table Tennis with a Three Stream Spatio-Temporal CNN for MediaEval 2020</article-title>
          . In
          <source>MediaEval (CEUR Workshop Proceedings)</source>
          , Vol.
          <volume>2882</volume>
          . CEUR-WS.org.
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          [18]
          <string-name>
            <surname>Pierre-Etienne</surname>
            <given-names>Martin</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Jenny</surname>
            Benois-Pineau,
            <given-names>Renaud</given-names>
          </string-name>
          <string-name>
            <surname>Péteri</surname>
            , and
            <given-names>Julien</given-names>
          </string-name>
          <string-name>
            <surname>Morlier</surname>
          </string-name>
          .
          <year>2019</year>
          .
          <article-title>Optimal Choice of Motion Estimation Methods for Fine-Grained Action Classification with 3D Convolutional Networks</article-title>
          .
          <source>In ICIP. IEEE</source>
          ,
          <fpage>554</fpage>
          -
          <lpage>558</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          [19]
          <string-name>
            <surname>Pierre-Etienne</surname>
            <given-names>Martin</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Jenny</surname>
            Benois-Pineau,
            <given-names>Renaud</given-names>
          </string-name>
          <string-name>
            <surname>Péteri</surname>
            , and
            <given-names>Julien</given-names>
          </string-name>
          <string-name>
            <surname>Morlier</surname>
          </string-name>
          .
          <year>2020</year>
          .
          <article-title>Fine grained sport action recognition with Twin spatio-temporal convolutional neural networks</article-title>
          .
          <source>Multim. Tools Appl</source>
          .
          <volume>79</volume>
          ,
          <fpage>27</fpage>
          -
          <lpage>28</lpage>
          (
          <year>2020</year>
          ),
          <fpage>20429</fpage>
          -
          <lpage>20447</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          [20]
          <string-name>
            <surname>Pierre-Etienne</surname>
            <given-names>Martin</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Jenny</surname>
            Benois-Pineau, Renaud Péteri, Akka Zemmari, and
            <given-names>Julien</given-names>
          </string-name>
          <string-name>
            <surname>Morlier</surname>
          </string-name>
          .
          <year>2021</year>
          .
          <article-title>3D Convolutional Networks for Action Recognition: Application to Sport Gesture Recognition</article-title>
          . Springer International Publishing.
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          [21]
          <string-name>
            <surname>Marion</surname>
            <given-names>Morel</given-names>
          </string-name>
          , Catherine Achard, Richard Kulpa, and
          <string-name>
            <given-names>Séverine</given-names>
            <surname>Dubuisson</surname>
          </string-name>
          .
          <year>2017</year>
          .
          <article-title>Automatic evaluation of sports motion: A generic computation of spatial and temporal errors</article-title>
          .
          <source>Image Vis. Comput</source>
          .
          <volume>64</volume>
          (
          <year>2017</year>
          ),
          <fpage>67</fpage>
          -
          <lpage>78</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          [22]
          <string-name>
            <given-names>Hai</given-names>
            <surname>Nguyen-Truong</surname>
          </string-name>
          , San Cao,
          <string-name>
            <given-names>N. A.</given-names>
            <surname>Khoa</surname>
          </string-name>
          <string-name>
            <given-names>Nguyen</given-names>
            ,
            <surname>Bang-Dang</surname>
          </string-name>
          <string-name>
            <given-names>Pham</given-names>
            , Hieu Dao,
            <surname>Minh-Quan</surname>
          </string-name>
          <string-name>
            <given-names>Le</given-names>
            ,
            <surname>Hoang-Phuc</surname>
          </string-name>
          Nguyen-Dinh,
          <article-title>Hai-Dang Nguyen, and</article-title>
          <string-name>
            <given-names>Minh-Triet</given-names>
            <surname>Tran</surname>
          </string-name>
          .
          <year>2020</year>
          . HCMUS at MediaEval 2020:
          <article-title>Ensembles of Temporal Deep Neural Networks for Table Tennis Strokes Classification Task</article-title>
          . In
          <source>MediaEval (CEUR Workshop Proceedings)</source>
          , Vol.
          <volume>2882</volume>
          . CEUR-WS.org.
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          [23]
          <string-name>
            <given-names>Soichiro</given-names>
            <surname>Sato</surname>
          </string-name>
          and
          <string-name>
            <given-names>Masaki</given-names>
            <surname>Aono</surname>
          </string-name>
          .
          <year>2020</year>
          .
          <article-title>Leveraging Human Pose Estimation Model for Stroke Classification in Table Tennis</article-title>
          . In
          <source>MediaEval (CEUR Workshop Proceedings)</source>
          , Vol.
          <volume>2882</volume>
          . CEUR-WS.org.
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          [24]
          <string-name>
            <surname>Dian</surname>
            <given-names>Shao</given-names>
          </string-name>
          ,
          <string-name>
            <given-names>Yue</given-names>
            <surname>Zhao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Bo</given-names>
            <surname>Dai</surname>
          </string-name>
          , and
          <string-name>
            <given-names>Dahua</given-names>
            <surname>Lin</surname>
          </string-name>
          .
          <year>2020</year>
          .
          <article-title>FineGym: A Hierarchical Video Dataset for Fine-Grained Action Understanding</article-title>
          .
          <source>In CVPR. IEEE</source>
          ,
          <fpage>2613</fpage>
          -
          <lpage>2622</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref25">
        <mixed-citation>
          [25]
          <string-name>
            <surname>Lucas</surname>
            <given-names>Smaira</given-names>
          </string-name>
          , João Carreira, Eric Noland, Ellen Clancy,
          <string-name>
            <surname>Amy Wu</surname>
            , and
            <given-names>Andrew</given-names>
          </string-name>
          <string-name>
            <surname>Zisserman</surname>
          </string-name>
          .
          <year>2020</year>
          . A Short Note on the Kinetics-700
          <string-name>
            <surname>-2020 Human Action</surname>
          </string-name>
          <article-title>Dataset</article-title>
          . CoRR abs/
          <year>2010</year>
          .10864 (
          <year>2020</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref26">
        <mixed-citation>
          [26]
          <string-name>
            <surname>Khurram</surname>
            <given-names>Soomro</given-names>
          </string-name>
          , Amir Roshan Zamir, and
          <string-name>
            <given-names>Mubarak</given-names>
            <surname>Shah</surname>
          </string-name>
          .
          <year>2012</year>
          .
          <article-title>UCF101: A Dataset of 101 Human Actions Classes From Videos in The Wild</article-title>
          .
          <source>CoRR abs/1212</source>
          .0402 (
          <year>2012</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref27">
        <mixed-citation>
          [27]
          <string-name>
            <surname>Gül</surname>
            <given-names>Varol</given-names>
          </string-name>
          , Ivan Laptev, and
          <string-name>
            <given-names>Cordelia</given-names>
            <surname>Schmid</surname>
          </string-name>
          .
          <year>2018</year>
          .
          <article-title>Long-Term Temporal Convolutions for Action Recognition</article-title>
          .
          <source>IEEE Trans. Pattern Anal. Mach. Intell</source>
          .
          <volume>40</volume>
          ,
          <issue>6</issue>
          (
          <year>2018</year>
          ),
          <fpage>1510</fpage>
          -
          <lpage>1517</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref28">
        <mixed-citation>
          [28]
          <string-name>
            <surname>Roman</surname>
            <given-names>Voeikov</given-names>
          </string-name>
          , Nikolay Falaleev, and
          <string-name>
            <given-names>Ruslan</given-names>
            <surname>Baikulov</surname>
          </string-name>
          .
          <year>2020</year>
          .
          <article-title>TTNet: Real-time temporal and spatial video analysis of table tennis</article-title>
          . (
          <year>2020</year>
          ),
          <fpage>3866</fpage>
          -
          <lpage>3874</lpage>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>