<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Human activity recognition using deep learning approaches</article-title>
      </title-group>
      <contrib-group>
        <aff id="aff0">
          <label>0</label>
          <institution>School of Mathematics</institution>
          ,
          <addr-line>Statistics and Computer Science</addr-line>
          ,
          <institution>University of KwaZulu-Natal</institution>
          ,
          <addr-line>Westville Campus, Private Bag X54001, Durban 4000</addr-line>
          ,
          <country country="ZA">South Africa</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>Human activity recognition using video data has been an active research area in computer vision for many years. This research presents an architecture that employs deep learning techniques to effectively solve human activity recognition of single individuals using visual information from videos. The architecture adopts an Octave Convolutional neural network as a feature extractor and a weighting strategy that regresses the importance of each video segment. This architecture was trained and evaluated on the KTH human activity dataset. The results are promising and validates the approach taken.</p>
      </abstract>
      <kwd-group>
        <kwd>human activity recognition (HAR)</kwd>
        <kwd>octave convolution neural network</kwd>
        <kwd>temporal segment network</kwd>
        <kwd>deep adaptive temporal pooling</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Introduction</title>
      <p>Human activity recognition is the process of identifying the actions and goals of
individuals from a series of observations of activities performed within a given
environment. This recognition process can be applied to many areas that have the goal of
monitoring human actions such as surveillance systems for detecting illegal activities
or sporting arenas for detecting foul play. This recognition task is solvable using
either video information, sensor data, or both.</p>
      <p>
        The application of deep learning has shown tremendous performance
improvements over traditional techniques. The use of deep learning has an advantage through
its automatic feature extraction abilities [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. In this paper, a deep learning architecture
that can successfully classify human activities is presented. The performance of the
architecture on classifying activities of the KTH human activity dataset is also given.
The proposed architecture utilizes a temporal segment network (TSN) as the base
architecture. A video V is divided into N segments of equal lengths. A frame from the
middle of each video segment serves as input for the spatial stream. The motion
representations of each segment were also extracted using the TVL1 optical flow
algorithm on a stack of 10 consecutive frames. An OctResNet50 model pre-trained on the
CIFAR-100 dataset was used as the frame-level feature extractor. The OctResNet50
model is a ResNet50 model that utilizes the octave convolution [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ]. The temporal
features are stacked together and parsed to a deep adaptive temporal pooling (DATP)
module to generate the importance of each temporal segment [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]. The two-component
Gaussian Mixture Model (GMM) was chosen as the weights-generator employed by
the DATP module. The generated weights were assigned to the temporal segments
and a late fusion strategy was adopted for fusing both the spatial and temporal
streams. The fused vector was then parsed to a feedforward neural network that
outputs the predicted class label for the entire video. All trainable parameters in this
architecture was trained through backpropagation.
3
      </p>
    </sec>
    <sec id="sec-2">
      <title>Results</title>
      <p>Most of the boxing, handclapping and handwaving videos were classified correctly.
The jogging and walking actions were misclassified as running. This misclassification
may have occurred because the motion representations of the 3 actions were similar.
The results obtained are not as good as the state-of-the-art results, but they are
promising and validates the approach taken.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Khurana</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          and
          <string-name>
            <surname>Kushwaha</surname>
            ,
            <given-names>A.K.S.</given-names>
          </string-name>
          ,
          <year>2019</year>
          .
          <article-title>Delving Deeper with Dual-Stream CNN for Activity Recognition</article-title>
          . In Recent Trends in Communication, Computing, and
          <string-name>
            <surname>Electronics</surname>
          </string-name>
          (pp.
          <fpage>333</fpage>
          -
          <lpage>342</lpage>
          ). Springer, Singapore.
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Song</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Cheung</surname>
            ,
            <given-names>N.M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chandrasekhar</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          and
          <string-name>
            <surname>Mandal</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <year>2018</year>
          .
          <article-title>Deep Adaptive Temporal Pooling for Activity Recognition</article-title>
          . arXiv preprint arXiv:
          <year>1808</year>
          .07272.
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Arunnehru</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chamundeeswari</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          and
          <string-name>
            <surname>Bharathi</surname>
            ,
            <given-names>S.P.</given-names>
          </string-name>
          ,
          <year>2018</year>
          .
          <article-title>Human action recognition using 3D convolutional neural networks with 3D motion cuboids in surveillance videos</article-title>
          .
          <source>Procedia computer science</source>
          ,
          <volume>133</volume>
          , pp.
          <fpage>471</fpage>
          -
          <lpage>477</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Chen</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Fang</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Xu</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Yan</surname>
            ,
            <given-names>Z.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kalantidis</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rohrbach</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Yan</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          and
          <string-name>
            <surname>Feng</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <year>2019</year>
          .
          <article-title>Drop an Octave: Reducing Spatial Redundancy in Convolutional Neural Networks with Octave Convolution</article-title>
          . arXiv preprint arXiv:
          <year>1904</year>
          .05049.
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5. Nada.kth.se. (
          <year>2019</year>
          ).
          <article-title>Recognition of human actions</article-title>
          . [online] Available at: http://www.nada.kth.se/cvap/actions/ [Accessed 14 Aug.
          <year>2019</year>
          ]
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>