<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Evaluation of a Video Annotation Tool Based on the LSCOM Ontology</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Emilie Garnaud</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Alan F. Smeaton</string-name>
          <email>Alan.Smeaton@dcu.ie</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Markus Koskela</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>E. Garnaud is with Institut EURECOM</institution>
          ,
          <addr-line>2229, Route des Creteˆs, BP 193 - 06904 Sophia Antipolis Cedex</addr-line>
          ,
          <country>France and</country>
          <institution>A. Smeaton and M. Koskela are with the Centre for Digital Video Processing, Dublin City University</institution>
        </aff>
      </contrib-group>
      <abstract>
        <p>- In this paper we present a video annotation tool to index video, we need to support different ways for the user based on the LSCOM ontology [1] which contains more than 800 to navigate it in order to complete the annotation process. semantic concepts. The tool provides four different ways for the In our annotation tool there are four distinct ways to suesaerrcthoblyoctahteemaep, ptrreoeprtriaatveercsoanlcaenpdtsotnoe uwshe,icnhaumseelsypbrea-scicomsepaurtcehd, annotate content, described as follows. concept similarities to recommend concepts for the annotator to rueslea.tiAve seeftfeoctfivuesneerssexopfetrhiemdeniftfseriesntreappoprrtoeadchdeesm.onstrating the A. Basic search</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>Index Terms— Video annotation, ontology, LSCOM, semantic
concept distances.</p>
    </sec>
    <sec id="sec-2">
      <title>I. INTRODUCTION</title>
      <p>In visual media processing, a lot of progress has been made
in automatically analysing low level visual features in order
to obtain a description of the content. However, annotations
by humans are still often needed to extract accurate deep
semantic information from within. Indeed manual tagging of
visual content has become widespread on the internet through
what is known as “folksonomy” in which human annotators
provide descriptive content tags.</p>
      <p>One of the challenges in the area of human annotation is
generating consistency across annotations in terms of both the
vocabulary used and the way it is used. The common approach
here is to provide users with an ontology, or an organisation
of allowable semantic tags or concepts. This is popular in
enterprises such as photo and video stock archives where only
a small number of people actually perform the annotation and
thus they are familiar with the ontology and the way it is
used. In more open-ended applications such as social tagging
or tagging by untrained users then ontologies are regarded
as too restrictive and too hard to learn in a short period of
time and so such applications favour free form tagging at the
expense of the consistency the use of an ontology brings.</p>
      <p>
        Here, we address the issue of how an untrained user could
use a pre-defined ontology to index video in the domain of
broadcast TV news. Specifically, we use the LSCOM ontology
[
        <xref ref-type="bibr" rid="ref1">1</xref>
        ], of about 850 concepts to help index media by semantics.
      </p>
    </sec>
    <sec id="sec-3">
      <title>II. VIDEO ANNOTATION TOOL</title>
      <p>Traditional annotation tools based on a lexicon or ontology
usually provide a full list of concepts with no, or very poor
ways to navigate it. This works quite well for a small lexicon
or for users who are trained to use it, but this is not scalable
to a larger ontology or the case where the users are untrained.
Thus in order to use the LSCOM or any other large ontology
An alphabetically-ordered list of the ontology and a search
box to find matching concepts is provided which is simple but
effective when users have a good knowledge of the ontology.</p>
      <sec id="sec-3-1">
        <title>B. Search by themes</title>
        <p>More than 700 concepts of the ontology have been arranged
into 19 different themes such as Arts &amp; Entertainment,
Business &amp; Commerce, News, Politics, Wars &amp; Conflicts . . . so an
annotator can search for a concept by first selecting a theme
that seems to fit with the shot.</p>
      </sec>
      <sec id="sec-3-2">
        <title>C. Recommended concepts</title>
        <p>
          In previous work introduced in [
          <xref ref-type="bibr" rid="ref2">2</xref>
          ] we computed similarity
among all pairs of concepts in the LSCOM ontology using
a combination of usage co-occurrence as the ontology was
used to index a corpus of 80 hours of video, combined with
visual shot-shot (and by implication, annotation-annotation)
similarities. We used these concept-concept co-occurrences to
generate “recommended concepts” at any point after
annotation by at least 1 concept. This worked by determining the 15
concepts most similar to the set of concepts already used to
annotate a shot, and this top-15 was refreshed every time an
additional concept was used in annotating a shot.
        </p>
      </sec>
      <sec id="sec-3-3">
        <title>D. Tree organization</title>
        <p>An hierachical version of the ontology has recently been
completed so we introduced some of its elements in our tool
by creating an area where a user can navigate among different
trees of the ontology.</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>III. EXPERIMENTS AND ANALYSIS</title>
      <p>We performed preliminary experiments involving 10 native
English-speaking users who each annotated 40 shots using
different functionalities of the tool, either in a restricted
timeframe or with unlimited time to complete. To replicate
the scenario of an untrained user annotating material on the
internet, our users did not receive any special training in
using the annotation tool. Shots to be annotated were selected
randomly and people used functionalities in a Latin squares
protocol so as not to bias the results. We analyzed four
different aspects of the annotation process namely the overall
time spent on annotating, the number of annotations per shot,
the shot annotation rate, and the number of annotations during
the first minute. Results are shown below. The best annotation
the same performance but after the first minute people lost time
searching the ontology for additional concepts as they did not
have enough knowledge to know when to stop as searching the
ontology does not provide any kind of closure to the process.
Average time per
shot
# annotations per
shot (Avg)
Annotation rate
Avg annotations
in 1st minute</p>
      <p>Search
Only
1m 53
6.9
6.1
6.3</p>
      <p>Search +
Themes
performance is obtained using the “recommended concepts”
feature because the time spent in free annotation is the same
as the “search only” version (representing the traditional
approach) but the number of annotations is greater when
recommendations are used. Using the“themes” feature seems
to slow down the annotation process without increasing the
number of annotations, probably due to a lack of knowledge
of the ontology and the way concepts had been organised
into different themes. Also, some shots are really good for
annotation by themes but others are not, which is why they
are a good complement to searching for concepts to annotate.</p>
      <p>We also found an unexpected result from the “entire tool”
experiment which surprisingly doesn’t seem to be the most
effective ! Once more, this seems to be due to a lack of
knowledge of the tool by users. Our whole point of
using untrained users is to replicate the common situation of
untrained users annotating resources on the internet. If we
examine the number of annotations done during the first
minute then “recommended concepts” and “entire tool” have</p>
    </sec>
    <sec id="sec-5">
      <title>IV. CONCLUSIONS AND FUTURE WORK</title>
      <p>The approach of using recommended concepts as a way
of annotating seems to be promising though the size of our
experiment is small. The “recommended concepts” could be
improved by collecting more data to link associated concepts.
Indeed, some associated concepts are really good (like ”store”,
”landlines”, ”bank”, ”office” and ”female person” for
”administrative assistant”) but some others are not, such as
(”harbors”, ”boat ship”, ”business people”, ”canal” and ”lakes” for
”house of worship”).</p>
      <p>The tool seems to be powerful for various user profiles. For
beginners, it helps them to learn the ontology and for experts
it provides a way to annotate concepts that they are not used
to annotating which improve their knowledge of the ontology.</p>
    </sec>
    <sec id="sec-6">
      <title>ACKNOWLEDGMENT This work was supported by Science Foundation Ireland under grant 03/IN.3/I361, by the EC under contract FP6-027026 (KSpace) and by IRCSET.</title>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>M.</given-names>
            <surname>Naphade</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.R.</given-names>
            <surname>Smith</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Tesic</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.-F.</given-names>
            <surname>Chang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W.</given-names>
            <surname>Hsu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Kennedy</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Hauptmann</surname>
          </string-name>
          and
          <string-name>
            <given-names>J.</given-names>
            <surname>Curtis</surname>
          </string-name>
          .
          <article-title>Large-Scale Concept Ontology for Multimedia</article-title>
          , IEEE Multimedia,
          <volume>13</volume>
          (
          <issue>3</issue>
          )
          <string-name>
            <surname>July-Sept</surname>
          </string-name>
          ,
          <year>2006</year>
          , pp.
          <fpage>86</fpage>
          -
          <lpage>91</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>M.</given-names>
            <surname>Koskela</surname>
          </string-name>
          and
          <string-name>
            <given-names>A.F.</given-names>
            <surname>Smeaton</surname>
          </string-name>
          .
          <source>Clustering-Based Analysis of Semantic Concept Models for Video Shots In Proc. IEEE International Conference on Multimedia &amp; Expo (ICME</source>
          <year>2006</year>
          ). Toronto, Canada.
          <source>July</source>
          <year>2006</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>