<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>GTTS Systems for the SWS Task at MediaEval 2013</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Luis J. Rodriguez-Fuentes</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Amparo Varona</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Mikel Penagarikano</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Germán Bordel</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Mireia Diez</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Software Technologies Working Group (</institution>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2013</year>
      </pub-date>
      <fpage>18</fpage>
      <lpage>19</lpage>
      <abstract>
        <p>This paper brie y describes the systems presented by the Software Technologies Working Group (http://gtts.ehu.es, GTTS) of the University of the Basque Country (UPV/EHU) to the Spoken Web Search (SWS) task at MediaEval 2013. GTTS systems consist of four main modules: (1) feature extraction; (2) speech activity detection; (3) DTW-based query matching; and (4) score calibration and fusion. The most remarkable contributions are the use of phone loglikelihood ratio features, the normalization of the DTW distance matrix and the calibration/fusion approach (which is imported from language/speaker veri cation).</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. INTRODUCTION</title>
      <p>The MediaEval 2013 Spoken Web Search (SWS) task
consists of searching for a spoken query within a set of audio
documents [4]. The locations and durations of all the
occurrences of spoken queries in the audio documents must
be obtained. System performance is primarily measured in
terms of the Average Term-Weighted Value (ATWV) [5],
but also in terms of a normalized cross-entropy metric and
the processing resources (real-time factor and peak
memory usage) required by the submitted systems [6]. For more
details on the SWS task at MediaEval 2013, see [2].
2.1</p>
    </sec>
    <sec id="sec-2">
      <title>SYSTEM OVERVIEW</title>
    </sec>
    <sec id="sec-3">
      <title>Feature extraction</title>
      <p>The Brno University of Technology (BUT) phone decoders
for Czech, Hungarian and Russian [7] are applied to
decode both the spoken queries and the audio documents.
BUT decoders are trained on 8 kHz SpeechDat(E) databases
recorded over xed telephone networks, containing 12, 10
and 18 hours of speech and featuring 45, 61 and 52 units for
Czech, Hungarian and Russian, respectively (three of them
being non-phonetic units that stand for short pauses and
noises).</p>
      <p>
        Given an input signal of length T , the decoder outputs
the posterior probability of each state s (1 s S) of each
unit i (1 i M ) at each frame t (1 t T ), pi;s(t),
where M is the number of units and S the number of states
per unit. The posterior probability of each unit i at each
frame t are computed by adding the posteriors of its states:
pi(t) = X pi;s(t)
8s
(
        <xref ref-type="bibr" rid="ref3">1</xref>
        )
2.2
      </p>
    </sec>
    <sec id="sec-4">
      <title>Speech Activity Detection</title>
      <p>Given an audio signal, Speech Activity Detection (SAD) is
performed by discarding those phone posterior feature
vectors for which the non-speech posterior is the highest. The
remaining vectors, along with their corresponding time o
sets, are stored for further use, but the component
corresponding to the non-speech unit is deleted. If the number of
speech vectors is too low (in this evaluation, that threshold
was arbitrarily set to 10, that is, 0.1 seconds), the whole
signal is discarded, to save time and to avoid false alarms.
2.3</p>
    </sec>
    <sec id="sec-5">
      <title>DTW-based query matching</title>
      <p>Given two SAD- ltered sequences of feature vectors
corresponding to a spoken query q and a spoken document x, the
cosine distance is computed between each pair of vectors,
q[i] and x[j] as follows:
d(q[i]; x[j]) =
log
with dmin(j) = min d(q[i]; x[j]) and dmax(j) = max d(q[i]; x[j]).</p>
      <p>i i
In this way, matrix values are all comprised between 0 and
1, so that a perfect match would produce a quasi-diagonal
sequence of zeroes.</p>
      <p>The best match of a query q of length m in a spoken
document x of length n is de ned as that minimizing the
average distance in a crossing path of the matrix dnorm. A
crossing path starts at any given frame of x, k1 2 [1; n],
then traverses a region of x which is optimally aligned to
q (involving L vector alignments), and ends at frame k2 2
[k1; n]. The average distance in this crossing path is:</p>
      <p>
        L
davg(q; x) = 1 X dnorm(q[il]; x[jl]) (
        <xref ref-type="bibr" rid="ref2 ref6">4</xref>
        )
      </p>
      <p>
        L l=1
where il and jl are the indices of the vectors of q and x
in the alignment l, for l = 1; 2; : : : ; L. Note that i1 = 1,
iL = m, j1 = k1 and jL = k2. The minimization operation
(
        <xref ref-type="bibr" rid="ref4">2</xref>
        )
(
        <xref ref-type="bibr" rid="ref1 ref5">3</xref>
        )
c2-late
PMUi
0.023
is accomplished by means of a dynamic programming
procedure, which is (n m d) in time (d: size of feature vectors)
and (n m) in space. The detection score is computed as
1 davg(q; x). The starting time and the duration of each
detection are obtained by retrieving the time o sets
corresponding to frames k1 and k2 in the SAD- ltered spoken
document.
      </p>
      <p>
        This procedure is iteratively applied to nd not only the
best match but also less likely matches in the same
document. To that end, a queue of search intervals is de ned
and initialized with (1; n). Let us consider an interval (a; b),
and assume that the best match is found at (a0; b0), then
the intervals (a; a0) and (b0; b) are added to the queue (for
further processing) if: (
        <xref ref-type="bibr" rid="ref3">1</xref>
        ) the score of the current match is
greater than a given threshold (in this evaluation, 0:85); (
        <xref ref-type="bibr" rid="ref4">2</xref>
        )
the interval is long enough (in this evaluation, half the query
length); and (
        <xref ref-type="bibr" rid="ref1 ref5">3</xref>
        ) the number of matches (already computed
+ pendant) is less than a given maximum (in this evaluation,
7). Finally, the list of matches for each query is truncated to
the N with the highest scores (in this evaluation, N = 1000).
      </p>
      <p>Under the extended (multiple examples) condition, only
the examples passing SAD ltering (i.e. with enough speech
samples) are considered for each query. The longest example
is taken as reference and DTW-aligned to the other
available examples. Finally, the vectors aligned at each frame
are averaged and a single average example is obtained and
processed as in the required (single example) condition.</p>
    </sec>
    <sec id="sec-6">
      <title>2.4 Score calibration and fusion</title>
      <p>System scores are transformed according to [1], which is an
adaptation of the discriminative calibration/fusion approach
commonly applied in speaker and language recognition.</p>
      <p>First, the so-called q-norm (query normalization) is
applied, so that zero-mean and unit-variance scores are
obtained per query. Then, if n di erent systems are fused,
detections are aligned so that only those supported by n=2
or more systems are retained for further processing (this is
known as majority voting validation). Let us consider one
of such validated detections, corresponding to a query q; if a
system A does not provide a score for it, we use instead the
minimum score that A has output for q. The same value
is assigned to missed detections and non-target trials. In
this way, a complete set of scores is prepared, which besides
the ground truth (target/non-target labels) can be used to
discriminatively estimate a linear transformation that
produces well-calibrated scores that can be linearly combined
to get fused scores. Under this approach, the Bayes
optimal threshold |given by the e ective prior (0:0148 for this
evaluation)| is applied. The BOSARIS toolkit [3] is used
to estimate and apply the calibration/fusion models.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          <article-title>3. RESULTS Tables 1 and 2 show the results (performance and processing resources) for GTTS systems in the required and extended conditions, respectively. All the experiments have been carried out on a 2 Xeon E5-2450 ( 8 core, 2 HT) @2</article-title>
          .10GHz, 64GB,
          <source>under Linux Fedora 3.3.4-5.fc17.x86 64</source>
          .
          <article-title>The indexing phase involves just applying BUT decoders to extract phone posterior features. ISF, SSF and PMU values have been computed as if all the computation had been performed sequentially in a single processor (see [6]). Calibration and fusion costs have been neglected</article-title>
          .
          <source>The contrastive systems 2</source>
          ,
          <article-title>3 and 4 (c2, c3 and c4) use the BUT decoders for Czech, Hungarian and Russian, respectively. The contrastive system 1 (c1) uses the concatenation of phone posteriors from the three decoders as features (and the average of non-speech posteriors for SAD)</article-title>
          .
          <article-title>The primary system (p), which is the fusion of the four contrastive systems, increases MTWV in 5 absolute points (15% relative) with regard to the best contrastive (c1). In all cases, calibration and fusion parameters have been estimated on the development set. Late submissions xed a bug in the fusion script (which did not count missed detections), thus leading to better calibrated systems. Note also that a 15% relative MTWV increase (nearly 4 absolute points) is obtained by using multiple examples under the approach described above (system c2-late</article-title>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>4. REFERENCES</mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>A.</given-names>
            <surname>Abad</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L. J. Rodriguez</given-names>
            <surname>Fuentes</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Penagarikano</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Varona</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Diez</surname>
          </string-name>
          , and
          <string-name>
            <given-names>G.</given-names>
            <surname>Bordel</surname>
          </string-name>
          .
          <article-title>On the calibration and fusion of heterogeneous spoken term detection systems</article-title>
          .
          <source>In Interspeech</source>
          <year>2013</year>
          , Lyon, France,
          <year>August</year>
          25-29
          <year>2013</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>X.</given-names>
            <surname>Anguera</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Metze</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Buzo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>I. Szoke</given-names>
            , and L.
            <surname>-J.</surname>
          </string-name>
          Rodriguez-Fuentes.
          <article-title>The Spoken Web Search Task</article-title>
          . In MediaEval 2013 Workshop, Barcelona, Spain, October
          <volume>18</volume>
          -19
          <year>2013</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>N.</given-names>
            <surname>Bru</surname>
          </string-name>
          <article-title>mmer</article-title>
          and E. de Villiers.
          <article-title>The BOSARIS Toolkit User Guide: Theory, Algorithms and Code for Binary Classi er Score Processing</article-title>
          .
          <source>Technical report</source>
          ,
          <year>2011</year>
          . https://sites.google.com/site/bosaristoolkit/.
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>MediaEval</given-names>
            <surname>Benchmarking</surname>
          </string-name>
          <article-title>Initiative for Multimedia Evaluation</article-title>
          .
          <article-title>The 2013 Spoken Web Search Task</article-title>
          ,
          <year>June 2013</year>
          . http://www.multimediaeval.org/mediaeval2013/sws2013/.
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [5]
          <string-name>
            <surname>NIST.</surname>
          </string-name>
          <article-title>The Spoken Term Detection (STD) 2006 Evaluation Plan</article-title>
          ,
          <year>September 2006</year>
          . http://www.itl.nist.gov/iad/mig/tests/std/2006/.
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>L.-J.</given-names>
            <surname>Rodriguez-Fuentes</surname>
          </string-name>
          and
          <string-name>
            <given-names>M.</given-names>
            <surname>Penagarikano. MediaEval 2013 Spoken Web</surname>
          </string-name>
          <article-title>Search Task: System Performance Measures</article-title>
          .
          <source>Technical report</source>
          , GTTS, UPV/EHU, May
          <year>2013</year>
          . http://gtts.ehu.es/gtts/NT/fulltext/rodriguezmediaeval13.pdf.
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>P.</given-names>
            <surname>Schwarz</surname>
          </string-name>
          .
          <article-title>Phoneme recognition based on long temporal context</article-title>
          .
          <source>PhD thesis</source>
          , FIT,
          <string-name>
            <surname>BUT</surname>
          </string-name>
          , Brno, Czech Republic,
          <year>2008</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>