<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>CUHK System for QUESST Task of MediaEval 2014</article-title>
      </title-group>
      <contrib-group>
        <aff id="aff0">
          <label>0</label>
          <institution>Haipeng Wang, Tan Lee DSP-STL, Dept. of EE The Chinese University of Hong Kong</institution>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2014</year>
      </pub-date>
      <fpage>16</fpage>
      <lpage>17</lpage>
      <abstract>
        <p>This paper describes a spoken keyword search system developed at the Chinese University of Hong Kong (CUHK) for the query by example search on speech (QUESST) task of MediaEval 2014. This system utilizes posterior features and dynamic time warping (DTW) for keyword matching. Multiple types of posterior features are generated with different tokenizers, and then fused by a linear combination on the DTW distance matrices. The main contribution of this year's system is a multiview segment clustering (MSC) approach for unsupervised ASM tokenizer construction. The Cnxe and ATWV of our submitted results on the Evaluation set are 0.682 and 0.412, respectively.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. INTRODUCTION</title>
      <p>
        The query by example search on speech (QUESST) task
aims at detecting the keyword occurrences in a unlabeled
speech collection using spoken queries in a language
independent fashion. In this year's QUESST dataset, the speech
collection involves about 23 hours of speech data from 6
languages, and the query set includes 560 development queries
and 555 evaluation queries. The average duration of queries
is about 0.9 second after voice activity detection (VAD).
More details about the task description can be found in [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ].
      </p>
      <p>
        Our system was designed only for the type 1 query
matching. It followed the posteriorgram-based template matching
framework [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ], in which speech tokenizers were used to
generate posteriorgrams, and DTW was applied for keyword
detection. The tokenizers were either built from the
searching speech collection given in the task, or developed from
some resource-rich languages. In order to exploit the
complementary information of multiple tokenizers, the DTW
matrix combination method [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ] was used. Raw DTW
detection scores were then normalized to zero mean and unit
variance. On the evaluation set, the Cnxe and ATWV of
our submission are 0.682 and 0.412. If only considering the
type 1 query matching, the Cnxe and ATWV are 0.611 and
0.526.
2.1
      </p>
    </sec>
    <sec id="sec-2">
      <title>SYSTEM DESCRIPTION</title>
    </sec>
    <sec id="sec-3">
      <title>System Overview</title>
      <p>
        In this year's evaluation, our system employs a similar
framework as our previous system for spoken web search
task in 2012 [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ]. The system involves seven tokenizers,
including a GMM tokenizer, ve phoneme recognizers, and an
ASM tokenizer [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ]. Using these tokenizers, the query
examples and test utterances are converted into frame-level
posteriorgrams. Di erent tokenizers may use di erent
algorithms to generate posteriorgrams. Let Qi denote the query
posteriorgram generated by the ith tokenizer, and let Ti
denote the corresponding test posteriorgram. The distance
matrix Di was computed as the inner-product [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ],
Di =
log(QiT
      </p>
      <p>Ti) i = 1; 2; :::; 7:
(1)</p>
      <p>To exploit the complementary information from di erent
tokenizers, the distance matrices were combined linearly to
give a new distance matrix D,</p>
      <p>7
D = X wiDi;</p>
      <p>i=1
where wi denotes the weighting coe cients for Di and was
simply set to 17 .</p>
      <p>Subsequently, DTW detection was applied to the
combined distance matrix D to locate the top matching regions.
DTW detection was performed with a sliding window with
a window shift of 5 frames. The adjustment window
constraint was imposed on the DTW alignment path. Let dq;t
denote the normalized DTW alignment distance between the
qth query on the tth hit region. The raw detection score was
computed by</p>
      <p>sq;t = exp( dq;t= );
where the scaling factor was set to 0.6. To calibrate the
score distribution of di erent queries, a 0/1 normalization
was used,
s^q;t = (sq;t
q)= q;
where s^q;t is the calibrated score, and q and q2 are the
mean and variance of the raw scores of the qth query.
2.2</p>
    </sec>
    <sec id="sec-4">
      <title>GMM Tokenizer</title>
      <p>The GMM tokenizer was trained from the given searching
speech collection. It contained 1024 Gaussian components.
The input of the GMM tokenizer was 39-dimensional MFCC
feature vector. The MFCC features were processed with
VAD and utterance-based mean and variance normalization
(MVN). Vocal tract length normalization (VTLN) was then
applied to the MFCC features o alleviate the in uence of
speaker variation.</p>
      <p>
        The warping factors of VTLN were estimated iteratively
as proposed in [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ]. The iteration started with training a
(2)
(3)
(4)
GMM from the unwarped MFCC features. Then the
warping factors were estimated with a maximum-likelihood grid
search using the GMM. A new GMM was trained using
the warped features, and new warping factors were then
re-estimated. This process was iterated four times in our
implementation. The usefulness of VTLN for this task was
experimentally demonstrated in our previous paper [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ].
2.3
      </p>
    </sec>
    <sec id="sec-5">
      <title>Phoneme Recognizers</title>
      <p>
        Our system involved ve phoneme recognizers, namely
Czech, Hungarian, Russian, English and Mandarin phoneme
recognizers. All these phoneme recognizers used the split
temporal context network structure [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ]. The Czech,
Hungarian, Russian phoneme recognizers were developed at Brno
University of Technology (BUT) and released in [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. The
English phoneme recognizer was trained on about 15-hour
speech data from the Fisher corpus and Swichboard
Cellular corpus. The Mandarin phoneme recognizer was trained
on about 15-hour speech data from the CallHome corpus
and the CallFriend corpus. These phoneme recognizers were
used to generate mono-phone state-level posteriorgrams
without any language model constraint.
2.4
      </p>
    </sec>
    <sec id="sec-6">
      <title>ASM Tokenizer</title>
      <p>Acoustic segment modeling (ASM) is a way to build an
HMM-based speech tokenizer from unlabeled speech data. It
consists of three steps, namely initial segmentation, segment
labeling, and iterative training and decoding. Initial
segmentation searches for the acoustic discontinuities and
partitions speech utterances into short-time speech segments.
In our implementation, we simply used the one-best
recognition results of the Hungarian phoneme recognizer to get
the hypothesised segment boundaries.</p>
      <p>
        Segment labeling is to assign a label to each short-time
speech segment. We used a multiview segment clustering
(MSC) approach for segment labeling. The MSC approach
took in multiple segment-level posterior features, computed
the similarity matrix and Laplacian matrix of the speech
segments for each type of posterior feature, and made a linear
combination on the Laplacian matrices. With the combined
Laplacian matrix, eigen-decomposition was performed to
derive the spectral embedding representations, and k-means
was applied to nd 100 clusters. Details of the MSC
approach are described in [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ].
      </p>
      <p>The cluster labels were used as initializations for iterative
training and decoding, in which HMM training and decoding
were performed iteratively until converge.</p>
    </sec>
    <sec id="sec-7">
      <title>RESULTS</title>
      <p>Table 1 shows the results obtained by our system on
evaluation queries. Based on our previous experience on TWV
values, we only submitted a small portion of the scores which
were higher than a threshold. This gives us the results of
System 1. However, if all the scores of all the trials are
considered, we obtain the results of System 2, which gives
obvious reductions on the Cnxe values. Similar observations
can also be made when only considering the type 1 query
matching. Corresponding results are shown in Table 2. The
di erence between Cnxe and TWV metrics needs to be
carefully examined in the future.</p>
      <p>To run the experiments, we used a computer with Intel
i73770K CPU (3.50GHz, 4 cores), 32GB RAM and 1T hard
drive. In the online searching process, all the posteriorgrams
were stored in the memory. This caused very high memory
cost (&gt;10GB). The computation cost in the searching
process was mainly caused by DTW detection. The searching
speed factor of our system was about 0.021. The slow
searching speed is one main drawback of our system and needs to
be improved.
4.</p>
    </sec>
    <sec id="sec-8">
      <title>CONCLUSION</title>
      <p>We have described an overview of the CUHK system
submitted to the MediaEval 2014 QUESST task along with the
evaluation results. Our system involves seven tokenizers
and uses DTW matrix combination for fusion. Only type
1 query matching is considered in the system development.
The main highlight of our system lies in the MSC approach
in the ASM tokenizer construction. In general we think the
performances for type 1 query matching are acceptable, but
the slow searching speed and high memory cost need to be
substantially improved.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1] http://speech. t.vutbr.cz/software/phonemerecognizer
          <article-title>-based-long-temporal-context.</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>X.</given-names>
            <surname>Anguera</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Rodriguez-Fuentes</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A. B. I.</given-names>
            <surname>Szoke</surname>
          </string-name>
          , and
          <string-name>
            <given-names>F.</given-names>
            <surname>Metze</surname>
          </string-name>
          .
          <article-title>Query by example search on speech at mediaeval 2014</article-title>
          .
          <source>In Working Notes Proceedings of the Mediaeval 2014 Workshop</source>
          , Barcelona, Spain,
          <source>October 16-17</source>
          ,
          <year>2014</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>T.</given-names>
            <surname>Hazen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W.</given-names>
            <surname>Shen</surname>
          </string-name>
          , and
          <string-name>
            <given-names>C.</given-names>
            <surname>White</surname>
          </string-name>
          .
          <article-title>Query-by-example spoken term detection using phonetic posteriorgram templates</article-title>
          .
          <source>In ASRU</source>
          , pages
          <volume>421</volume>
          {
          <fpage>426</fpage>
          ,
          <year>2009</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>P.</given-names>
            <surname>Schwarz</surname>
          </string-name>
          .
          <article-title>Phoneme recognition based on long temporal context</article-title>
          ,
          <source>PhD thesis</source>
          .
          <year>2009</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>H.</given-names>
            <surname>Wang</surname>
          </string-name>
          and
          <string-name>
            <given-names>T.</given-names>
            <surname>Lee</surname>
          </string-name>
          .
          <article-title>CUHK system for the spoken web search task at mediaeval 2012</article-title>
          .
          <source>In Working Notes Proceedings of the Mediaeval 2012 Workshop</source>
          , Pisa, Italy, October 4-
          <issue>5</issue>
          ,
          <year>2012</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>H.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Lee</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.-C.</given-names>
            <surname>Leung</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Ma</surname>
          </string-name>
          , and
          <string-name>
            <given-names>H.</given-names>
            <surname>Li</surname>
          </string-name>
          .
          <article-title>Acoustic segment modeling with spectral clustering methods</article-title>
          . in submission to IEEE/ASM TASLP.
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>H.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Lee</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.-C.</given-names>
            <surname>Leung</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Ma</surname>
          </string-name>
          , and
          <string-name>
            <given-names>H.</given-names>
            <surname>Li</surname>
          </string-name>
          .
          <article-title>Using parallel tokenizers with DTW matrix combination for low-resource spoken term detection</article-title>
          .
          <source>In ICASSP</source>
          ,
          <year>2013</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>H.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <surname>C.-C. Leung</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          <string-name>
            <surname>Lee</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          <string-name>
            <surname>Ma</surname>
            , and
            <given-names>H.</given-names>
          </string-name>
          <string-name>
            <surname>Li</surname>
          </string-name>
          .
          <article-title>An acoustic segment modeling approach to query-by-example spoken term detection</article-title>
          .
          <source>In ICASSP</source>
          ,
          <year>2012</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>S.</given-names>
            <surname>Wegmann</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>McAllaster</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Orlo</surname>
          </string-name>
          , and
          <string-name>
            <given-names>B.</given-names>
            <surname>Peskin</surname>
          </string-name>
          .
          <article-title>Speaker normalization on conversational telephone speech</article-title>
          .
          <source>In ICASSP</source>
          ,
          <year>1996</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>