<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>On Localizing Keywords in Continuous Speech using Mismatched Crowd</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>P. Radadia</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>T. Madhesia</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>K. Kalra</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>A. Sriraman</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>M. Patwardhan</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>S. Karande TCS Research</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>-B Hadpasar Industrial Estate</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>India shirish.karande@tcs.com</string-name>
        </contrib>
      </contrib-group>
      <abstract>
        <p>We report results on use of a mismatched crowd in spotting</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Introduction</title>
      <p>
        The language demography on crowdsourcing platforms is
significantly different from the actual world population
        <xref ref-type="bibr" rid="ref3">(Pavlick et al. 2014)</xref>
        . Therefore, recent papers have
explored the use of mismatched crowds for speech annotation.
        <xref ref-type="bibr" rid="ref1 ref2">(Jyothi and Hasegawa-Johnson 2015b)</xref>
        have shown that, for
isolated word recognition, even though the accuracies of
individual mismatched workers may be poor, the annotation
can be aggregated to provide significantly improved
accuracy. Nevertheless,
        <xref ref-type="bibr" rid="ref1 ref2">(Jyothi and Hasegawa-Johnson 2015a)</xref>
        observed that accurate annotations in continuous speech is
significantly harder . In this paper, we consider a task whose
perceptual difficulty lies between the two.
      </p>
      <p>
        We utilize the mismatched crowd for keyword
localization (e.g. see
        <xref ref-type="bibr" rid="ref5">(Sanders, Neville, and Woldorff 2002)</xref>
        ) in
continuous speech. The task required a worker to listen to
a speech utterance and then label the data as follows: (1)
Choose a keyword from the list of 5 options. Among them,
4 were words and the fifth option allowed the worker to
choose ‘None of above’.(2) If the worker choses one of the
words then she is required to mark boundaries on the
waveform of the utterance. Thus a worker’s label consists of
keyword identity and its markings (Figure 1). Our setup, while
not exactly mapped to a specific usecase, is motivated by
low resource scenarios where a keyword spotting engine is
desired for a limited dictionary of words. An (in-loop) ASR
is not assumed to be available for this work-in-progress.
      </p>
      <p>
        Crowd consensus has been an active research area
        <xref ref-type="bibr" rid="ref4">(Raykar et al. 2010; Zhou et al. 2014)</xref>
        . We build upon
(Welinder and Perona 2010). Our key contributions are:
(1) Demonstration of a non-trivial ability amongst a
mismatched crowd to spot and segment word utterances . (2) A
Copyright c 2018 for this paper by its authors. Copying permitted
for private and academic purposes.
framework that explicitly models spammer bias, leading to
improved accuracies in label aggregation.
      </p>
    </sec>
    <sec id="sec-2">
      <title>Task Creation and Data Collection</title>
      <p>To create tasks, we considered utterances from four
languages viz, Arabic, German, Hindi and Russian. The
utterances are extracted from online news videos which
provided subtitles along with them. The Arabic utterances are
extracted from TED talks. The subtitle text of an utterance
is used as a ground truth transcription for that utterance.
We extracted 50 utterances, for each language, of on
average duration of 4.2 second. To generate a task where a
keyword is present in the utterance, we picked a random word
from the utterance transcription and remaining 3 words are
sampled randomly from word corpus specific to the
underlying language. We ensured that only one keyword from
the options will be present in the utterance. Note that the
ground truth markings of the keywords have been generated
by native speakers hired from Upwork. In case of ‘None of
above’ tasks, we presented the words that have not been
spoken in the utterance. We generated an equal number
of tasks with a keyword present or absent. Since the
mismatched worker may not be familiar with any of the four
languages, we transliterated non-roman scripted
transcriptions (Arabic, Hindi and Russian in our case) to Roman
(English) using Google’s read phonetic utility.</p>
      <p>We got responses from 100 CrowdFlower workers. Each
worker was asked to label 25 tasks. Each task was
attempted by 10 different workers under a random
allocation.The workers were asked to identify their native
language before attempting the task, and responses where the
workers were not mismatched have been discarded from
the study. The workers were spread across 21 languages
e.g. English, Spanish, Czech, Lithuanian, Serbian, Tagalog,
Ukrainian etc., however the majority of them identified
English (34%) or Spanish (35%) as their native language.</p>
      <sec id="sec-2-1">
        <title>Bias in Crowd Annotations</title>
        <p>The overall accuracy of the crowd (in %) for each language
is shown in Table 1. We consider the worker’s annotation
correct only when her choice matches the ground truth and
the marking overlaps at least 50% with the ground truth
markings. We report accuracy for cases when (1) (KW)
keyword is present in utterance, (2) (NKW) none of the word
options are present and (3) overall performance. The
results from Table 1 indicate a strong bias toward choosing
the ‘None of above’ option. Consequently, in this work, we
explicitly model this bias, and demonstrate the utility of it
to obtain improved accuracies in aggregation.</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>Modeling and Label Aggregation</title>
      <sec id="sec-3-1">
        <title>Model for crowd worker</title>
        <p>Let the workers be indexed by j 2 W = f1; 2; :::M g, the
tasks be indexed by i 2 T = f1; 2; :::; N g and lij be the
annotation provided by jth worker for the ith task. Moreover,
each task i has an underlying target label (to be estimated),
zi 2 Si. The set Si (Z+ Z+ Oi) [ N K where
Oi represents the set of keyword identities in the option list</p>
      </sec>
      <sec id="sec-3-2">
        <title>Arabic German Hindi Russian</title>
        <p>KW 24:82 29:45 27:45 24:81
NKW 71:01 72:51 74:16 73:91
Overall 57:88 48:15 35:81 47:40
Table 1: Language wise performance of crowd
shown to the worker for task i, N K represents the label
where keyword is not present in the utterance and the
integer tuple represents the boundaries of a marked keyword.
We assume that lij belong to the same set as zi.</p>
        <p>We model the parameters of worker j as j = f j ; j g,
where j represents the probability that worker behaves
like a spammer. When behaving like a spammer, j
represents the probability of her bias towards choosing ‘None
of above’ option. The spammer chooses all the remaining
keyword option with equal probability. Furthermore, when
a spammer chooses a keyword option, he is expected to
mark the boundaries. We assume that the spammer
randomly chooses the keyword boundaries. On the contrary,
when a working is not spamming, we assume it behaves
like a Hammer, i.e. if a word is not present in an utterance
it always correctly chooses the ”None of the above” option,
and if a word is present it always chooses the correct option
from the list. However, we assume that while identifying
the boundaries a slight error may occur. This error is
modeled by a bi-variate Gaussian distribution which is common
for the entire crowd population. One can refer to Figure 2 to
understand the above described notation and the likelihood
expression presented in the remainder of the section.</p>
        <p>Consider the case when correct keyword is absent but
worker chooses the wrong keyword x from the list and
provides markings of it on waveform as a1 and a2 respectively.
The likelihood is given by:
p (lij = (a1; a2; x)jzi = N K; j ) =</p>
        <p>1 1
j ) 2 D
(1)
where represents the number of samples in the utterance.
D represents number of words in the presented option list.
The above equation states that worker is spamming but
is not inclined towards choosing ‘None of above’ option.
However, since the worker is spamming, we can assume
that it choose the keyword options with equal with uniform
probability 1=D and her spam boundary markings are also
governed with uniform probability 1= 2. Meanwhile, when
a worker rejects all the words when actually there is no
keyword in the utterance, the likelihood probability is given by:
p (lij = N Kjzi = N K; j ) = (1
j ) + j j
(2)
Above equation states that either the worker did not spam
the label (hence probability 1 ) or might have spammed
but been biased to choose the ‘None of above’ option
(probability j j ). Similarly when worker spams the label given
that a keyword is present in the utterance, its likelihood
probability is computed as:
p (lij = N Kjzi = (t1; t2; y); j ) =
j j
Further, when the worker does provide a correct label for
the keyword this may occur because he is honest, or a
spammer has accidentally made the right choice:</p>
        <p>p (lij = (a1; a2; y)jzi = (t1; t2; y); j ) =
(1
j )N ((a1; a2); (t1; t2); ) + j (1
1 1
j ) 2 D
We assume that boundary markings by an honest worker are
normally distributed about the ground truth.
(3)
(4)</p>
        <p>Finally, we account for the possibility that provided word
choice does not match the groundtruth label. This event can
occur only when the worker is spamming:
p (lij = (a1; a2; x 6= y)jzi = (t1; t2; y); j ) =
j (1</p>
        <p>1 1
j ) 2 D
(5)</p>
      </sec>
      <sec id="sec-3-3">
        <title>EM algorithm for label aggregation</title>
        <p>Using the likelihood equations we can setup an EM to
estimate the latent true label, zi, and worker’s parameters. We
utilize the following notation:Each worker j provides labels
Lj = flij gi2Tj for the subset Tj T . Similarly, each task
i has labels Li = flij gj2Wi provided by subset of workers
Wi W. The set of all workers’ labels is denoted as L.</p>
        <p>E-step: Assuming current estimate of model parameters
^, we approximate the posterior on each target value zi:
p^(zi) / p(zij ) Y</p>
        <p>p (lij jzi; j )
j2Wi
where the label prior has uniform distribution. The
probability of target values can be decomposed as follows:
p^(zi = (t1; t2; y)) =</p>
        <p>0
p^(zi = N K) = (1
j2Wi
1
Y p (lij jzi = (t1; t2; y); j )A
) Y p (lij jzi = N K; j )</p>
        <p>j2Wi
where is the prior probability of the cases when keyword
is present. We estimate the target label as follows:
^zi = arg max (p^(zi = (t1; t2; y)); p^(zi = (N K)))
zi
To avoid slow sampling, we approximate the posterior on
zi with delta function (Welinder and Perona 2010),
p^(zi) = (^zi)
(6)
(7)
(8)
(9)</p>
        <p>M-step: The parameters for worker j are estimated by
maximizing the expectation of logarithm of posterior on j
with respect to p^(zi):
We compare the proposed model against: (1) Majority
Voting (MV), (2) EM for multiclass as described in
(Welinder and Perona 2010) and (3) Model without considering
bias (i.e. j = 0). Note that baseline methods (1) and (2),
rather unfairly, only consider the keyword identity and not
the markings while evaluating the accuracies.</p>
        <p>Table 2 shows that the proposed model provides a gain
of 30%, 35.46% and 20.91% over MV, Multiclass and the
unbiased method in terms of overall accuracy. Furthermore,
Table 3 shows that there is a significant gain for all
languages. The improvement can be attributed to: (1) The
location annotation has higher dimension making it resilient
to random spam (2) The workforce is biased towards
selecting the ’None of above’ option. The introduction of the bias
parameter is responsible for providing a gain of 5%.</p>
      </sec>
      <sec id="sec-3-4">
        <title>Model Parameter Estimation</title>
        <p>Figure 3 compares the worker parameters estimated by the
EM against those obtained from the ground truth. It can be
observed that overall reliability cannot be the only
parameter to characterize the workers. In Figure 3, red points (blue
points) represent spam workers (honest annotators) as
identified by the proposed EM setup when we use 0.8 as a
filtering threshold on j ; j values. It was observed that the EM
was able to correctly identify 36 out of 46 spammers.</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>Future work</title>
      <p>The list of options determines the difficult of the task. In
practice, there can be great efficiency in employing an ASR
in the loop, where, the word options are generated by the
ASR engine. We anticipate these words to have greater
phonetic proximity compared to our task. We intend to extend
our experiments and systems to study the above scenario.
Welinder, P., and Perona, P. 2010. Online
crowdsourcing: rating annotators and obtaining cost-effective labels. In
IEEE computer society conference on computer vision and
pattern recognition workshops (CVPRW), 25–32. IEEE.
Zhou, D.; Liu, Q.; Platt, J. C.; and Meek, C. 2014.
Aggregating ordinal labels from crowds by minimax conditional
entropy. In ICML, 262–270.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          <string-name>
            <surname>Jyothi</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Hasegawa-Johnson</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          <year>2015a</year>
          .
          <article-title>Acquiring speech transcriptions using mismatched crowdsourcing</article-title>
          .
          <source>In AAAI</source>
          ,
          <fpage>1263</fpage>
          -
          <lpage>1269</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          <string-name>
            <surname>Jyothi</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Hasegawa-Johnson</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          <year>2015b</year>
          .
          <article-title>Transcribing continuous speech using mismatched crowdsourcing</article-title>
          .
          <source>In Sixteenth Annual Conference of the International Speech Communication Association.</source>
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          <string-name>
            <surname>Pavlick</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Post</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Irvine</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Kachaev</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ; and CallisonBurch,
          <string-name>
            <surname>C.</surname>
          </string-name>
          <year>2014</year>
          .
          <article-title>The language demographics of amazon mechanical turk</article-title>
          .
          <source>Transactions of the Association for Computational Linguistics</source>
          <volume>2</volume>
          :
          <fpage>79</fpage>
          -
          <lpage>92</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          <string-name>
            <surname>Raykar</surname>
            ,
            <given-names>V. C.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Yu</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Zhao</surname>
            ,
            <given-names>L. H.</given-names>
          </string-name>
          ; Valadez,
          <string-name>
            <surname>G. H.</surname>
          </string-name>
          ; Florin,
          <string-name>
            <given-names>C.</given-names>
            ;
            <surname>Bogoni</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            ; and
            <surname>Moy</surname>
          </string-name>
          ,
          <string-name>
            <surname>L.</surname>
          </string-name>
          <year>2010</year>
          .
          <article-title>Learning from crowds</article-title>
          .
          <source>Journal of Machine Learning Research</source>
          <volume>11</volume>
          (Apr):
          <fpage>1297</fpage>
          -
          <lpage>1322</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          <string-name>
            <surname>Sanders</surname>
            ,
            <given-names>L. D.</given-names>
          </string-name>
          ; Neville,
          <string-name>
            <surname>H. J.;</surname>
          </string-name>
          and Woldorff,
          <string-name>
            <surname>M. G.</surname>
          </string-name>
          <year>2002</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          <article-title>Speech segmentation by native and non-native speakersthe use of lexical, syntactic, and stress-pattern cues</article-title>
          .
          <source>Journal of Speech, Language, and Hearing Research</source>
          <volume>45</volume>
          (
          <issue>3</issue>
          ):
          <fpage>519</fpage>
          -
          <lpage>530</lpage>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>