<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Cepstral Polynomial Regression For Sequential Detection Of Impulsive Waveform In Video Sound-Track</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Cyril Hory</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>William J. Christmas</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Anil Kokaram</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Cyril Hory is with Laboratoire Traitement et Communication de l'Information, CNRS-GET/Te ́le ́com-Paris</institution>
          ,
          <addr-line>37-39 rue Dareau</addr-line>
          ,
          <institution>75014 Paris. William J. Christmas is with University of Surrey, Centre for Vision Speech and Signal Processing</institution>
          ,
          <addr-line>Guilford GU2 7XH</addr-line>
          ,
          <country>UK Anil</country>
          <institution>Kokaram is with University of Dublin, Trinity College, EEE Department, College Green</institution>
          ,
          <addr-line>Dublin 2 Ireland. fellowship</addr-line>
        </aff>
      </contrib-group>
      <abstract>
        <p>- A new set of features is introduced for characterization of impulsive events in video clips from the audio signal. The discriminative power of these features to detect and isolate racket hits in tennis video clip is discussed.</p>
      </abstract>
      <kwd-group>
        <kwd>cepstral features</kwd>
        <kwd>sequential classification</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>I. INTRODUCTION</title>
      <p>
        In digital video analysis it is now apparent that audio cues
extracted from a video, along with the visual cues, can provide
relevant information for semantic understanding of the content.
In [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ] for example, audio and visual features are combined
within an HMM framework for parsing tennis video.
Audiovisual cooperation is ensured through multi-modal conditional
density estimation in [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. Mel-cepstral coefficients are used
in [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ] to identify specific sounds in a baseball game
videoclip in order to detect commercials, speech or music using the
maximum entropy method.
      </p>
      <p>In many application domains it is possible to identify some
critically informative elementary short-terms event. However,
even though visual attributes can be as informative as audio
attributes for characterizing short-terms event, audio data are
more convenient to handle in terms of computation load.
Among short term events impulsive waveforms can be
peculiarly informative. For instance accurate percussive sound
detection can help with beat analysis. Racket hits are
particularly informative events for the understanding a tennis game.
From the detection and characterisation of racket hits, it is
possible to extract information such as the score, player fitness
and skills, or the strategy.</p>
      <p>
        We have proposed in [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ] a semi-supervised sequential scheme
for detecting events from the audio stream of a video sequence
using the Generalized CUSUM procedure. Following this
detection step, a system for event identification can be triggered
when an event is detected.
      </p>
      <p>II. CEPSTRAL FEATURES EXTRACTION</p>
      <p>
        We focus here on the identification of impulsive waveforms
from the audio content after detection. We propose a new
(1)
(3)
set of features based on cepstral analysis of the recorded
events. Denote by c = [c1, c2, . . . , cN ]T the vector of cepstral
coefficients of the event e [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]:
      </p>
      <p>c = |FT−1{log(|FT{e}|)}| ,
where FT{.} is the discrete Fourier transform. Assume there
exists a vector a(p) = [a0, a1, . . . , ap]T such that
c = Q(p)a(p) + ν(p) ,
(2)
where Q(p) is a N ×(p+1) matrix with element gj−1(qi−1) on
the ith row and jth column, where gj can be any conveniently
chosen polynomial of order j and qi the ith quefrency index,
and ν(p) is an N × 1 vector of random perturbations. The
mean-square error estimate aˆ(p) of the vector of regression
coefficients a(p) is:</p>
      <p>aˆ(p) = R(−p1)Q(Tp)c ,
with R(p) = Q(Tp)Q(p). The Cepstral Regression Coefficients
(CRC) aˆ(p) are descriptors of the cepstrum content of the
detected events.</p>
      <p>In unsupervised sequential classification and learning the
amount of available data is often limited and small. Low
dimensional feature spaces must be considered in order to
cope with the curse of dimensionality. In such a situation the
CRC’s allow for the encoding of the whole information carried
by the cepstrum in a small number of coefficients. Moreover
the inversion of the (p + 1) × (p + 1) matrix R(p) in (3) can be
performed recursively by using a block matrix decomposition
and the Schur complement of R(p−1). The dimensionality of
the feature space can thus be adaptively updated to match the
size of the dataset. The recursive computation of the regression
coefficients makes these features particularly appealing for the
implementation of a sequential classification system.</p>
      <p>III. EXPERIMENTAL VALIDATION</p>
      <p>
        A classification experiment was carried on excerpts of tennis
video clip to evaluate the capability of the proposed features to
discriminate impulsive waveforms. The impulsive waveform
(target class of the classification experiment) are the racket
hits. A Biased Discriminant Analysis [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ] was performed within
a supervised learning framework. The experiment has been
carried on the second and third game of the Australian Open
final tennis game of 2003. The training set contains 203 events
that were detected by the CUSUM test including 20 actual
racket hits. The test data contains 479 events that were detected
by the CUSUM test including 61 racket hits.
      </p>
      <p>On Fig. 1 are displayed the ROC curves of the classifier based
1
0.9
0.8
itilyb00..76
a
b
o
r
p0.5
n
o
i
tc0.4
e
t
e
D0.3
0.2
0.1
00
1
0.9
0.8
y0.7
ti
ilb0.6
a
b
o
rp0.5
n
o
ti
tec0.4
e
D
0.3
0.2
0.1
00
1
0.9
0.8
tliiyb00..76
a
b
o
r
p0.5
n
o
i
tc0.4
e
t
e
D0.3
0.2
0.1
00
0.1 0.2 0.3 False alarm probability 0.7 0.8 0.9</p>
      <p>0.4 0.5 0.6
(a) First Cepstral Coefficients (liftering)
0.1 0.2 0.3 Fa0ls.4e alarm probability 0.7 0.8 0.9</p>
      <p>0.5 0.6
(b) Cepstral Regression Coefficients
0.1 0.2 0.3 False alarm probability 0.7 0.8 0.9</p>
      <p>0.4 0.5 0.6
(c) Spectral Regression Coefficients
1
1
1
on the First Cepstral Coefficients (FCC), CRC’s, and Spectral
Regression Coefficients (SRC) of the extracted events. The
3-dimension CRC vector outperforms the FCC’s whether the
training set or the test set is classified. The 3 FCC’s fail to
encode enough information about the event.</p>
      <p>The test set classification performances of the CRC’s and
SRC’s are equivalent although the CRC’s outperforms the
SRC’s when applied on the training set.</p>
      <p>
        The 7-dimension FCC’s performs better than the 3-dimension
FCC’s when applied to the training set although the
performances on the test set are similar whatever the feature space
dimension. When applied to the training set, the 7-dimension
FCC vector behaves as the CRC’s and outperforms the SRC’s.
However the performances of the FCC’s dramatically
deteriorate when applied to the test set. This shows that a classifier
based on the FCC’s is spoiled by an over-fitting phenomenon.
If modelling the waveform as the convolution of a source
waveform and a filter impulse response, FCC’s encode
information about the filter while high quefrency coefficients are
characteristic of the source [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]. In the experiment the impulse
response of the filter depends on the acoustic characteristics of
the hall and on the electronic recording device. It is a common
to all the extracted events. Thus the filter characteristics can
not provide relevant information for discriminating the various
events.
      </p>
      <p>The CRC of order 6 exhibits a better ability to characterise and
discriminate the racket hits than the CRC of order 2 except
at high false alarm probability. This shows that the CRC of
order 6 tends to perform an over-fitting of the training set. As
a consequence, outliers racket hits are more often taken into
account in the model. In this case, the outliers are two lifted
shots. The high probability of detection obtained, even though
the characteristic of the class has evolved from the second to
the third game, shows that the CRC features are relevant to
encode the cepstrum content of a racket hit in a non-stationary
context.</p>
    </sec>
    <sec id="sec-2">
      <title>IV. CONCLUSION</title>
      <p>The CRC’s perform better than the standard FCC’s in
a low-dimension feature space but the discrepancy between
performances of the two feature vectors decreases when the
dimension increases. One can conclude that the CRC’s seem
more appropriate for classification of impulsive waveform
when dealing with a small data set. In a higher dimension
feature space performances are equivalent but feature vector
computed from the polynomial regression (CRC’s and SRC’s)
tends to provide less over-fitting than the standard FCC’s.
The high discriminating power in a small dimensional feature
space and the recursive computation of the features allows for
their integration in a sequential and adaptive learning system.
Work is currently being carried to show how the proposed
features could improve the retrieval results obtained here in a
static supervised learning context.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>R.</given-names>
            <surname>Dahyot</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Kokaram</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Rea</surname>
          </string-name>
          , and
          <string-name>
            <given-names>H.</given-names>
            <surname>Denman</surname>
          </string-name>
          , “
          <article-title>Joint audio visual retrieval for tennis broadcasts</article-title>
          ,”
          <source>in Proceedings of IEEE ICASSP'03</source>
          ,
          <year>2003</year>
          , pp.
          <fpage>561</fpage>
          -
          <lpage>564</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>J. R.</given-names>
            <surname>Deller</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. G.</given-names>
            <surname>Proakis</surname>
          </string-name>
          , and
          <string-name>
            <given-names>J. H. L.</given-names>
            <surname>Hansen</surname>
          </string-name>
          ,
          <source>Discrete-Time Processing of Speech Signals. MacMillan</source>
          ,
          <year>1993</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>C.</given-names>
            <surname>Hory</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Kokaram</surname>
          </string-name>
          , and
          <string-name>
            <given-names>W. J.</given-names>
            <surname>Christmas</surname>
          </string-name>
          , “
          <article-title>Threshold learning from samples drawn from the null hypothesis for the GLR CUSUM test</article-title>
          ,”
          <source>in Proc. IEEE MLSP</source>
          ,
          <year>2005</year>
          , pp.
          <fpage>111</fpage>
          -
          <lpage>116</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>W.</given-names>
            <surname>Hua</surname>
          </string-name>
          , M. Han, and
          <string-name>
            <given-names>Y.</given-names>
            <surname>Gong</surname>
          </string-name>
          , “
          <article-title>Baseball scene classification using multimedia features,”</article-title>
          <source>in Proceedings of IEEE ICME'02</source>
          ,
          <year>2002</year>
          , pp.
          <fpage>821</fpage>
          -
          <lpage>824</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>E.</given-names>
            <surname>Kijak</surname>
          </string-name>
          , G. Gravier,
          <string-name>
            <given-names>P.</given-names>
            <surname>Gros</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Oisel</surname>
          </string-name>
          , and
          <string-name>
            <given-names>F.</given-names>
            <surname>Bimbot</surname>
          </string-name>
          , “
          <article-title>HMM based structuring of tennis videos using visual and audio cues,”</article-title>
          <source>in Proc. IEEE ICME'03</source>
          ,
          <year>2003</year>
          , pp.
          <fpage>309</fpage>
          -
          <lpage>312</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>X. S.</given-names>
            <surname>Zhou</surname>
          </string-name>
          and
          <string-name>
            <given-names>T. S.</given-names>
            <surname>Huang</surname>
          </string-name>
          , “
          <article-title>Small Sample Learning during Multimedia Retrieval using BiasMap,”</article-title>
          <source>in Proceedings of IEEE CVPR'01, December</source>
          <year>2001</year>
          , pp.
          <fpage>11</fpage>
          -
          <lpage>17</lpage>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>