<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>NTU System at MediaEval 2015: Zero Resource Query by Example Spoken Term Detection Using Deep and Recurrent Neural Networks</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Cheng-Tao Chung</string-name>
          <email>b97901182@gmail.com</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Yang-de Chen</string-name>
          <email>yongde0108@gmail.com</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Graduate Institute of Communication, Engineering, National Taiwan University</institution>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Graduate Institute of Electrical Engineering, National Taiwan University</institution>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2015</year>
      </pub-date>
      <fpage>14</fpage>
      <lpage>15</lpage>
      <abstract>
        <p>This note serves as a documentation describing the methods the authors of this paper implemented for the Query by Example Search on Speech Task (QUESST) as a part of MediaEval 2015. In this work, we combined DTW, DNN and RNN in one framework to perform query by example spoken term detection in a zero resource setting.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. INTRODUCTION</title>
      <p>
        Participants of the task were asked to implement a query
by example spoken term detection system on a corpus
provided by the organizers. The queries were divided into the
development set and the evaluation set, and the list of
correct documents are given for the development queries. A
soft score and a hard decision for every query-document
pair in the evaluation set has to be provided. Note that
in this task, only whether or not the query appears in the
document is considered, when the query appears in the
document is not important. For more information please refer
to the overview paper [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ].
      </p>
      <p>
        In this work, we approached the task under a zero resource
setting [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ] using neural networks. This means we did not
use any other information than the corpus itself. For the
task considered here, we need to formulate an objective that
compares two feature sequences of a query and a document
both of varying length and return a score. The Deep Neural
Network [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ] (DNN) is a state-of-the-art architecture that
has been widely applied in speech recognition. However, it
is limited to framewise objectives where the length of the
input feature has to be xed. Hence we need to focus on
two issues in the work: dealing with the varying sequence
length and the formulation of a sequence objective.
      </p>
      <p>
        Dynamic Time Warping [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ] (DTW) is one of the earliest
techniques applied in the eld and can nd the alignment
of two sequences, hence transforming both sequences into
a feature representation of the same length. DTW solves
the problem of varying sequence length. On the other hand,
we use Recurrent Neural Networks [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ] (RNN) to generate
a sequence objective for the query-document pair. In this
work, we combine DTW, DNN and RNN in one framework.
      </p>
    </sec>
    <sec id="sec-2">
      <title>OBJECTIVE AND APPROACH</title>
      <p>
        Let the feature sequence of an utterance be denoted as
X 2 RS F , where S denotes the length of the sequence, and
F denotes the number of dimensions of the feature. We
extract 39 dimensional MFCCs with energy, delta and double
deltas with HTK[
        <xref ref-type="bibr" rid="ref6">6</xref>
        ] for our features in this work. By
performing sub-sequence Dynamic Time Warping on the two
feature sequences of a document Xd and that of a query
Xq, we can nd the aligned warping sequences Wd 2 RT F
in the document and Wq 2 RT F in the query, where T
denotes the length of the warping sequence.
      </p>
      <p>Wd; Wq = DT W (Xd; Xq):</p>
      <p>We forward the feature frames of both Xd and Xq through
the same deep neural network. The number of neurons in the
DNN is 39; 100; 100; 39 on each layer from input to output,
We use tanh as the activation function of the network.</p>
      <p>Hd
Hq
=
=</p>
      <p>DN N (Wd);</p>
      <p>DN N (Wq):
On the nal layer of the DNN, the output features are then
concatenated, then forwarded to a recurrent neural network.
The number of neurons in the hidden layer of the RNN is 50,
and the sigmoid function is used as the activation function.
The output of the RNN at the time frame t is a single score
st(q; d), s(q; d) is the average of the score along the entire
sequence, and T is the length of the sequence:</p>
      <p>RN N ([Hd; Hq]) = [s0(d; q); s1(d; q); :::; sT 1(d; q)] (4)
s(d; q) =</p>
      <p>t
1 X st(d; q):
T
(1)
(2)
(3)
(5)
(6)
(7)</p>
      <p>For every query q, the score of a positive document dp
containing q should be high; the score of a negative document
dn containing q should be low. Therefore, the following is
the objective which we wish to minimize:</p>
      <p>The nal objective we wish to minimize is the sum of the
objective of all the queries:</p>
      <p>Lq =</p>
      <p>X s(dn; q)
p;n</p>
      <p>s(dp; q):
L = X X s(dn; q)
q p;n
s(dp; q):
We train the entire network including both the DNN and
RNN using back-propagation algorithm. We take s(d; q)
as the score for the document-query pair (d; q), and query
would be considered to be in a document if s(d; q) &gt; 0:5.
set
dev
eval
method
dtw
rnn
dtw
rnn</p>
    </sec>
    <sec id="sec-3">
      <title>EXPERIMENTS AND RESULTS</title>
      <p>
        The entire corpus was trained only using the corpus of
QUESST 2015. We derived two sets of scores from the
method above. The rst set of scores is the DTW scores
generated when we initially align the features between query
and document. The second set of scores is the score
generated from our RNN in equation 4. We did not perform any
pretraining on the network and used random initialization
for all the weights. The neural network was implemented
using the Theano library [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ]. Positive examples for
querydocument pairs were selected from all query types(T1, T2,
T3) in the development set, negative examples were
randomly generated query-document pairs. The results of our
experiments are shown in Table 1. We only show the actual
normalized cross entropy (Cnxe). From the results, we see
that the RNN did not perform better than the DTW, and
neither system seemed to have performed well. This could
be due to error in the implementation or insu cient number
of epochs during the training for the RNN. Since the results
of the RNN were based on the results of the DTW, it is
unclear of what caused the problem.
      </p>
    </sec>
    <sec id="sec-4">
      <title>CONCLUSION</title>
      <p>The authors of the paper attempted a framework to
combine DNN, RNN and DTW under a single zero resource
neural network framework for query by example spoken term
detection.</p>
    </sec>
    <sec id="sec-5">
      <title>5. SUPPLEMENTARY MATERIAL</title>
      <p>Since we were encouraged to discuss other systems that
we've tried, we've included several versions of our system
through di erent development iterations. All of these
systems except for the correspondence auto-encoder have been
implemented, yet not all have been evaluated on the corpus.
The philosophy behind all these designs was to map
querydocument pairs to a single trainable objective so the error
back-propagates through the entire network.</p>
      <p>In the rst attempt, we learned feature transformation on
the acoustic features (MFCC) using a DNN. The objective
of the DNN was a the warping distance after DTW, and the
error of the DTW was back-propagated into the network.
This design was abandoned due to unreasonably slow
calculation: DNN is GPU intensive while DTW is CPU intensive,
making this hybrid system hard to implement using Theano.</p>
      <p>
        In the second attempt, we replaced DTW with
Convolutional Neural Networks (CNN) [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ]. CNNs are GPU friendly
architectures which xes the problem of the previous hybrid
systems. The 2D feature representation of every utterance
was treated as an image. The acoustic features from
different queries/documents were padded with zero vectors to
be the same length. This feature transformation CNN only
has convolutional layers. The end result the rst CNN was
another 2D feature representation with one axis being time.
We treated the features at the output as if they were
acoustic features and plot the warping matrix (pairwise consine
similarity). However, instead of applying DTW, we treated
the warping matrix itself as another image and forward it
through another CNN. The target of the CNN was whether
or not the document contains the query. This design didn't
work because the error on the testing set didn't converge,
maybe due to serious over- tting.
      </p>
      <p>In the third attempt, we removed the second CNN to
reduce the number of parameters. The rst CNN has a fully
connected layer in this design, and we took the inner
product of the fully connected layer from the document and the
query to be the error. Although the number of parameters
have been reduced, the error on the testing set still didn't
converge. Maybe the number of correct training pairs just
wasn't enough to train such systems.</p>
      <p>
        Finally, we decided to consider using correspondence
autoencoders [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ] for this task. The plan was to use
correspondence autoencoders as an zero-resource feature extractor.
We planned to perform another DTW on these extracted
features. Although the system contains both DTWs and
DNNs, they are not jointly trained so the performance
problem of the rst design doesn't occur. However, DTWs are
extremely time consuming operations opposed to RNNs. We
ran out of time to perform the second DTW so decided to
replace it with a RNN which is the system that we submitted
in the end.
      </p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>Igor</given-names>
            <surname>Szoke</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Luis J</given-names>
            .
            <surname>Rodriguez-Fuentes</surname>
          </string-name>
          , Andi Buzo, Xavier Anguera, Florian Metze, Jorge Proenca,
          <string-name>
            <given-names>Martin</given-names>
            <surname>Lojka</surname>
          </string-name>
          , and
          <string-name>
            <given-names>Xiao</given-names>
            <surname>Xiong</surname>
          </string-name>
          .
          <article-title>Query by example search on speech at mediaeval 2015</article-title>
          .
          <source>In Working Notes Proceedings of the Mediaeval 2015 Workshop</source>
          , Wurzen, Germany,
          <year>2015</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>Aren</given-names>
            <surname>Jansen</surname>
          </string-name>
          , Emmanuel Dupoux, Sharon Goldwater, Mark Johnson, Sanjeev Khudanpur,
          <string-name>
            <given-names>Kenneth</given-names>
            <surname>Church</surname>
          </string-name>
          , Naomi Feldman, Hynek Hermansky, Florian Metze, Richard Rose, et al.
          <article-title>A summary of the 2012 JHU CLSP workshop on zero resource speech technologies and models of early language acquisition</article-title>
          .
          <year>2013</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>Li</given-names>
            <surname>Deng</surname>
          </string-name>
          , Geo rey Hinton, and
          <string-name>
            <given-names>Brian</given-names>
            <surname>Kingsbury</surname>
          </string-name>
          .
          <article-title>New types of deep neural network learning for speech recognition and related applications: An overview</article-title>
          .
          <source>In Acoustics, Speech and Signal Processing (ICASSP)</source>
          ,
          <year>2013</year>
          IEEE International Conference on, pages
          <volume>8599</volume>
          {
          <fpage>8603</fpage>
          . IEEE,
          <year>2013</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>Meinard</given-names>
            <surname>Mu</surname>
          </string-name>
          <article-title>ller. Dynamic time warping. Information retrieval for music and motion</article-title>
          , pages
          <volume>69</volume>
          {
          <fpage>84</fpage>
          ,
          <year>2007</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>Tony</given-names>
            <surname>Robinson</surname>
          </string-name>
          , Mike Hochberg, and
          <string-name>
            <given-names>Steve</given-names>
            <surname>Renals</surname>
          </string-name>
          .
          <article-title>The use of recurrent neural networks in continuous speech recognition</article-title>
          .
          <source>In Automatic speech and speaker recognition</source>
          , pages
          <volume>233</volume>
          {
          <fpage>258</fpage>
          . Springer,
          <year>1996</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>Steve</given-names>
            <surname>Young</surname>
          </string-name>
          , Gunnar Evermann, Mark Gales, Thomas Hain, Dan Kershaw, Xunying Liu, Gareth Moore, Julian Odell, Dave Ollason,
          <string-name>
            <given-names>Dan</given-names>
            <surname>Povey</surname>
          </string-name>
          , et al.
          <source>The HTK book</source>
          , volume
          <volume>2</volume>
          . Entropic Cambridge Research Laboratory Cambridge,
          <year>1997</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>James</given-names>
            <surname>Bergstra</surname>
          </string-name>
          , Olivier Breuleux, Frederic Bastien, Pascal Lamblin, Razvan Pascanu, Guillaume Desjardins, Joseph Turian, David Warde-Farley, and
          <string-name>
            <given-names>Yoshua</given-names>
            <surname>Bengio</surname>
          </string-name>
          .
          <article-title>Theano: a CPU and GPU math expression compiler</article-title>
          .
          <source>In Proceedings of the Python for scienti c computing conference (SciPy)</source>
          , volume
          <volume>4</volume>
          , page 3. Austin, TX,
          <year>2010</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>Alex</given-names>
            <surname>Krizhevsky</surname>
          </string-name>
          , Ilya Sutskever, and
          <string-name>
            <surname>Geo</surname>
            rey
            <given-names>E</given-names>
          </string-name>
          <string-name>
            <surname>Hinton</surname>
          </string-name>
          .
          <article-title>Imagenet classi cation with deep convolutional neural networks</article-title>
          .
          <source>In Advances in neural information processing systems</source>
          , pages
          <volume>1097</volume>
          {
          <fpage>1105</fpage>
          ,
          <year>2012</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>Herman</given-names>
            <surname>Kamper</surname>
          </string-name>
          , Micha Elsner, Aren Jansen, and
          <string-name>
            <given-names>Sharon</given-names>
            <surname>Goldwater</surname>
          </string-name>
          .
          <article-title>Unsupervised neural network based feature extraction using weak top-down constraints</article-title>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>