<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Overview of MediaEval 2011 Rich Speech Retrieval Task and Genre Tagging Task</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Christoph Kofler c.kofler@tudelft.nl</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
          <xref ref-type="aff" rid="aff4">4</xref>
          <xref ref-type="aff" rid="aff5">5</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Delft University of Technology</institution>
          ,
          <country country="NL">Netherlands</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Gareth J.</institution>
          <addr-line>F. Jones</addr-line>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>Maria Eskevich</institution>
        </aff>
        <aff id="aff3">
          <label>3</label>
          <institution>Martha Larson</institution>
        </aff>
        <aff id="aff4">
          <label>4</label>
          <institution>Roeland Ordelman</institution>
        </aff>
        <aff id="aff5">
          <label>5</label>
          <institution>Sebastian Schmiedeke</institution>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2011</year>
      </pub-date>
      <fpage>1</fpage>
      <lpage>2</lpage>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. INTRODUCTION</title>
      <p>
        The MediaEval 2011 Rich Speech Retrieval Tasks and
Genre Tagging Tasks are two new tasks o ered in MediaEval
2011 that are designed to explore the development of
techniques for semi-professional user generated content (SPUG).
They both use the same data set: the MediaEval 2010 Wild
Wild Web Tagging Task (ME10WWW). The ME10WWW
data set contains Creative Commons licensed video collected
from blip.tv in 2009. It was created by the PetaMedia
Network of Excellence (http://www.petamedia.eu) in order to
test retrieval algorithms for video content as it occurs `in the
wild' on the Internet and, in particular, for user contributed
multimedia that is embedded within a social network. The
data set was initially described in [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]. In this overview
paper, we repeat the essential characteristics of the data set,
describe the tasks and specify how they are evaluated.
      </p>
    </sec>
    <sec id="sec-2">
      <title>THE ME10WWW BLIP.TV DATA SET</title>
      <p>The ME10WWW data set consists of video episodes from
blip.tv and a social network comprised of Twitter users
tweeting about them. Videos were collected for shows for which
the link to one of their episodes had been tweeted. Topsy
(http://topsy.com) was used to collect blip.tv links from
tweets. Their licenses were checked to con rm that they
were Creative Commons and then the videos were
downloaded from blip.tv. Topsy was then searched again to gather
all users mentioning any one of the videos. We crawled the
tweets and social network of these users for two steps (i.e.,
collected their interlocutors and their interlocutors's
interlocutors). The data set contains 1974 episodes (247
development and 1727 test) comprising a total of ca. 350 hours of
data. The development set is small with respect to the test
set and is not intended for training, but rather for
parameter tuning. The episodes were chosen from 460 di erent
shows|shows with less than four episodes were not
considered for inclusion in the data set. Each video is associated
with metadata record including uploader assigned
information (title, description, license, tags, uploaded ID/series ID).</p>
      <p>
        The data set is accompanied by automatic speech
recognition (ASR) transcripts [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ], which were generously
provided by LIMSI (http://www.limsi.fr/) and Vocapia
Research (http://www.vocapia.com/) to MediaEval. In order
to be included in the ME10WWW set, a video needed to
have been transcribed by the ASR-system with an average
word-level con dence score of &gt; 0.7. The set is
predominantly English with approximate 6 hours of non-English
content divided over French, Spanish and Dutch. Speci
cally for 2011, LIMSI/Vocapia also provided a second set of
\confusion networks", meaning that for a given time-point
(time code), the transcripts may contain more than one
hypotheses of the ASR-system.
      </p>
      <p>
        The data set is also accompanied by the output of a shot
detection system developed at the Technische Universitat
Berlin [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. Both shot boundaries and extracted keyframes
(one per shot) are included.
3.
      </p>
    </sec>
    <sec id="sec-3">
      <title>RICH SPEECH RETRIEVAL TASK</title>
      <p>The Rich Speech Retrieval task is a known-item retrieval
task and requires participants to return a ranked list of
results in response to a query. The features used can be
derived from speech, audio, visual content or metadata. The
task can be considered to be a `new generation' spoken
content retrieval task in two respects. First, instead of requiring
the identi cation of spoken documents or speech segments,
the task requires the return of jump-in points, time points
in the video at which users must start watching to view
material relevant to the query. Second, the queries used
express the user information need along three dimensions: the
topical content, the speech act and the visual content.</p>
      <p>
        The speech act dimension is, to our knowledge,
investigated for the rst time with this data set. When speakers
speak, they are, on the one hand, pronouncing words, but
on the other hand they are also actually `doing' something.
Treating spoken content in terms of `illocutionary speech
acts' (http://en.wikipedia.org/wiki/Speech acts) emphasizes
what speakers are accomplishing by speaking. The ve speech
acts used are `apology', `de nition', `opinion', `promise' and
`warning'. These categories are similar to the `speech-act
like' units that have ben used in dialogue act modeling [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ].
We are motivated to investigate this dimension because we
conjecture that a connection exists between the reason for
which a speaker makes an utterance and the reason for which
a user would later search for the utterance. We assume that
this connection could be used to improve spoken content
retrieval systems.
      </p>
      <p>The data set includes 30 queries associated with the
development set and 50 queries associated with the test set.
The form of queries includes a long form (&lt;title&gt;), a short
form (&lt;short_title&gt;) and a label indicating the speech act
type (act). An example of an apology is the following. Long
form: `How does Peter Busch, sta member of the Morning
Swim Show, save face after the faux pas he made during his
interview with Terry Denison?'. Short form: `Peter Busch
president chairman Denison morning Swim Show'. Although
some queries in the data set make reference to visual
characteristics (i.e., the speaker's clothing), most resemble this
example and do not clearly need a visual contribution.
Participants are required to make one submission for all test
topics that uses only the provided ASR transcript (2010),
and a second one using any combination of techniques which
they believe will give the best overall performance. In the
required run the participants must use the full queries.</p>
      <p>The query set was created by having human annotators
locate portions of the videos that they would be interested
in sharing online and then formulating a comment on what
the video portion was about (long form) and a query that
would allow them to re- nd jump-in points corresponding
to those portions (short form). Online-sharing was
chosen as the scenario, since it seemed to be the most natural,
widely-understood reason for which users would be
attempting to re- nd previously seen video segments (i.e., known
items). Access to a su cient number of human annotators
was secured by using a crowdsourcing platform, Amazon's
Mechanical Turk (http://www.mturk.com). The task was
carefully designed, the context of online video-sharing was
clearly described and the language kept simple, i.e., \When
you come across something interesting you might want to
share on Facebook, Twitter or your favorite social network."
Examples were given of quotes falling into each of the speech
act categories. We note that `opinions' were much more
frequently located in the corpus than other categories such as
`warning' or `promise'. The pairs of query+jump-in-point
returned by the crowdsourcing workers were subjected to
a standard quality control procedure (which eliminated ca.
40% of the results) and then by a screening for suitability
and completeness (which eliminated a further ca. 40%).</p>
      <p>
        The o cial evaluation metric of the task is mGAP, cf. [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ],
which generalizes the relevance of hypothesized jump-in points
in relation to ground truth points by imposing a symmetric
step-wise linearly decaying penalty function within a window
of tolerance (10s, 30s, 60s windows are used). Since RSR is a
known-item task, the metric is e ectively a `mGRR' (mean
Generalize Reciprocal Rank).
      </p>
    </sec>
    <sec id="sec-4">
      <title>GENRE TAGGING TASK</title>
      <p>Genre information in the form of genre tags can
provide valuable support for users searching and browsing the
Internet for video. However, much video|and especially
SPUG video|is not accurately or adequately tagged. This
task attempts to automatically generate genre labels such
as they are used to organize videos on video platforms such
as blip.tv. Genre tagging is related to the genre classi
cation task set out by Google as an ACM Multimedia Grand
Challenge task in 2009 and 2010.</p>
      <p>The Genre Tagging task requires participants to
automatically assign genre tags to videos using features derived from
speech, audio, visual content or associated textual or social
information (Twitter social network). The tag set users are
required to predict contains 26 genre tags.1 Each video is
associated with only one genre. Genre tags were collected
for each episode using the API provided by blip.tv. A genre
tag is represented by the eld categoryName in the JSON
output provided by the API. Whitespace in the genre tags
was replaced by the underscore character (` '), ampersands
(`&amp;') by the word `and'.</p>
      <p>Participants submit up to ve runs in total representing
ve di erent approaches to the task (i.e., experimental
conditions). They are required to complete one run using ASR
transcripts only and one run including metadata. Within the
metadata, we assume that user-assigned tags might already
contain explicit genre information. Since we are interested in
algorithms that work under the (likely) event that no tags
are available, one run is also required with metadata, but
no tags. Upon request, groups with a particular focus were
excused from speci c required runs. The o cial evaluation
metric is MAP (Mean Average Precision).
5.</p>
    </sec>
    <sec id="sec-5">
      <title>ACKNOWLEDGMENTS</title>
      <p>Thank you to Isabelle Ferrane (IRIT) who served as
organizer of the Genre Tagging Task on the behalf of Quaero
(http://www.quaero.org) and gave the task overview at the
MediaEval 2011 workshop. Thank you to Thijs Verschoor
(Twente) for his contribution to the RSR Task. This
MediaEval e ort received funding from EC FP7 NoE PetaMedia
and Science Foundation Ireland RFP 2008.</p>
    </sec>
    <sec id="sec-6">
      <title>6. REFERENCES</title>
      <p>1art, autos and vehicles, business, citizen journalism,
comedy, conferences and other events, default category,
documentary, educational food and drink,
gaming, health, literature, movies and television,
music and entertainment, personal or auto-biographical,
politics, religion, school and education, sports,
technology, the environment, the mainstream media, travel,
videoblogging, web development and sites</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>P.</given-names>
            <surname>Kelm</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Schmiedeke</surname>
          </string-name>
          , and
          <string-name>
            <given-names>T.</given-names>
            <surname>Sikora</surname>
          </string-name>
          .
          <article-title>Feature-based video key frame extraction for low quality video sequences</article-title>
          .
          <source>In WIAMIS '09</source>
          , pages
          <fpage>25</fpage>
          {
          <fpage>28</fpage>
          ,
          <year>2009</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>L.</given-names>
            <surname>Lamel</surname>
          </string-name>
          and
          <string-name>
            <given-names>J.-L.</given-names>
            <surname>Gauvain</surname>
          </string-name>
          .
          <article-title>Speech processing for audio indexing</article-title>
          .
          <source>In Advances in Natural Language Processing, LNCS 5221</source>
          , pages
          <fpage>4</fpage>
          <lpage>{</lpage>
          15. Springer,
          <year>2008</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>M.</given-names>
            <surname>Larson</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Soleymani</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Serdyukov</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Rudinac</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Wartena</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Murdock</surname>
          </string-name>
          , G. Friedland,
          <string-name>
            <given-names>R.</given-names>
            <surname>Ordelman</surname>
          </string-name>
          , and
          <string-name>
            <given-names>G. J. F.</given-names>
            <surname>Jones</surname>
          </string-name>
          .
          <article-title>Automatic tagging and geotagging in video collections and communities</article-title>
          .
          <source>In ACM ICMR '11</source>
          , pages
          <issue>51:1</issue>
          {
          <issue>51</issue>
          :
          <fpage>8</fpage>
          ,
          <year>2011</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>B.</given-names>
            <surname>Liu</surname>
          </string-name>
          and
          <string-name>
            <given-names>D. W.</given-names>
            <surname>Oard</surname>
          </string-name>
          .
          <article-title>One-sided measures for evaluating ranked retrieval e ectiveness with spontaneous conversational speech</article-title>
          .
          <source>In SIGIR '06</source>
          , pages
          <fpage>673</fpage>
          {
          <fpage>674</fpage>
          ,
          <year>2006</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>A.</given-names>
            <surname>Stolcke</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Coccaro</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Bates</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Taylor</surname>
          </string-name>
          , C. Van
          <string-name>
            <surname>Ess-Dykema</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          <string-name>
            <surname>Ries</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          <string-name>
            <surname>Shriberg</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          <string-name>
            <surname>Jurafsky</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          <string-name>
            <surname>Martin</surname>
            ,
            <given-names>and M.</given-names>
          </string-name>
          <string-name>
            <surname>Meteer</surname>
          </string-name>
          .
          <article-title>Dialogue act modeling for automatic tagging and recognition of conversational speech</article-title>
          .
          <source>Comput. Linguist.</source>
          ,
          <volume>26</volume>
          :
          <fpage>339</fpage>
          {
          <fpage>373</fpage>
          ,
          <year>September 2000</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>