<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Video Stream Structuring and Annotation Using Electronic Program Guides</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Jean-Philippe Poli</string-name>
          <email>jppoli@ina.fr</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Jean Carrive</string-name>
          <email>jcarrive@ina.fr</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Institut National de l'Audiovisuel 4</institution>
          ,
          <addr-line>avenue de l'Europe 94366 Bry-sur-Marne Cedex</addr-line>
          ,
          <country country="FR">France</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Laboratoire des Sciences de l'Information et des Systèmes LSIS (UMR CNRS 6168) Campus scientifique de Saint Jérôme Avenue Escadrille Normandie Niemen 13397</institution>
          <addr-line>Marseille Cedex 20</addr-line>
          ,
          <country country="FR">France</country>
        </aff>
      </contrib-group>
      <fpage>137</fpage>
      <lpage>141</lpage>
      <abstract>
        <p>The French National Audiovisual Institute (INA) is in charge of archiving continuously the video stream of every French television channel. In order to provide an access to particular programs in these streams, each program must be described. Presently, a stream is manually structured and annotated. Our work focuses on the use of electronic program guides to structure and annotate a video stream. We present in this article a way for our system to index a video stream using such program guides.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1 Introduction</title>
      <p>The French National Audiovisual Institute1 is dealing with huge video databases since
it is in charge of archiving each program broadcasted on French TV 24/7. In order to
provide an efficient way to consult its archives, the institute is used to describing
manually both the structure and the content of each document. For many years, the
videos are digitally acquired and that makes possible to automate many kinds of
treatments like subtitles extraction, transcription or restoration.</p>
      <p>Our work focuses on video streams (for instance a whole broadcasted week)
structuring. Our system finds both the boundaries of the various programs and
commercials, and then annotates them using different knowledge sources. This
document structuring can lead to practical applications in archives consulting and is a
first necessary phase to programs description or automatic video indexing since we
will not look for the same semantic features in broadcast news and in soap operas.</p>
      <p>
        The video indexing community, to our knowledge, seldom looked for solutions to
this problem and rather focuses on semantic features extraction [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ] whereas the
stream’s structure can guide this extraction. Moreover, these methods are often suited
1 http://www.ina.fr
for small videos and cannot be effective in our case since it will lead to great times of
computations [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ].
      </p>
      <p>
        We propose an original approach which uses a maximum of knowledge on TV
programs in order to minimize calculations: the most important a priori information is
a program guide related to the stream [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ]. But unlikely, it cannot be used in its rough
state for the structuring because of its being incomplete and imprecise. Firstly, in
association with the past schedules, the program guide is used to predict telecasts’
boundaries in the stream thanks to the constancy of the programs schedules from one
year to another. [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ] shows most of television channels have entered in a mode of
competition to attract the most audience as possible, which explains they need to
develop the loyalty of their public. In the same way, a channel determines the
advertisements fares depending on the audience that must be as stable as possible [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ].
That implies them to be regular in their programs schedules and that’s why our
assumption is totally realistic. We will see in the next section how we can use this
constancy to improve program guides. Secondly, program guides can provide
semantic information about a program and we will discuss about this kind of
information in the last section.
      </p>
      <p>We propose in this article a way to take advantage of such program guide on an
automatic video structuring and annotation task. The objective is to determine
temporal windows within which the system can limit its search for program’s
boundaries. Once programs are delimited and classified, the system can then extract
semantic features by using automatic indexing tools. We will start by presenting the
contribution and the difficulties of automatic structuring comparing to manual
segmentation, and then we will see what kind of information we can get from forecast
program schedules.</p>
    </sec>
    <sec id="sec-2">
      <title>2 From Manual to Automatic Video Stream Structuring</title>
      <p>As we saw in the previous section, to index a stream we need to structure it and divide
it into telecasts and commercials.</p>
      <p>Manual video structuring poses few problems: for one broadcast day and for every
channel, it is necessary to monitor the stream in order to find its structure. A manual
structuring is still in use and is made from the forecast schedules; that implies such
structures are not very accurate. Automatic video structuring would permit to have a
better precision and especially be more rapid. The difficulty of this automation is
explained by the replacement or the cancellation of telecasts; during our study of
schedules, we saw that during the night programs could be shrunken. All these
particularities may affect the results of an automatic video structuring.</p>
      <p>
        Another difficulty is the INA’s hierarchical taxonomy of the different programs,
which is really too precise for an automatic recognition: for instance, we cannot
distinguish by their audio and video features a movie from a TV film, or more
specifically an action film from a detective film, even if some works are carried out
[
        <xref ref-type="bibr" rid="ref2">2</xref>
        ][
        <xref ref-type="bibr" rid="ref9">9</xref>
        ]. Since our work has to be compatible with the present taxonomy, we have
chosen to aggregate some specific terms that cannot be easily computationally
distinguished.
      </p>
      <p>Predicted
beginning
9:01:15 am
9:23:31 am
9:28:08 am
10:53:31 am
11:01:13 am
11:36:36 am
12:08:15 am
12:12:29 am
12:53:02 am
12:55:05 am
12:58:19 am
1:45:36 pm
1:49:35 pm
1:55:20 pm
2:56:08 pm
4:00:46 pm</p>
      <p>Real
beginning
9:00:11 am
9:22:59 am
9:27:45 am
10:51:49 am
10:58:41 am
11:34:23 am
12:07:04 am
12:12:35 am
12:51:22 am
12:55:15 am
12:58:36 am
1:47:22 pm
1:51:05 pm
1:57:26 pm
2:58:56 pm
4:03:09 pm</p>
      <sec id="sec-2-1">
        <title>Kind of</title>
      </sec>
      <sec id="sec-2-2">
        <title>Program</title>
        <p>SOAP OPERA
MAGAZINE
MAGAZINE
NEWS
GAME
GAME
MAGAZINE
GAME
MAGAZINE
WEATHER
NEWS
WEATHER
SERVICE
TV SHOW
TV SHOW
TV SHOW</p>
        <p>
          Furthermore, we need to be sure that our work can be applied at the institute and
this constrained us to find a way to structure a broadcasted week in an efficient time
and with the guaranty of exhaustiveness, and that works on a set of documents as
biggest and vastest as possible. For example, some programs are not separated at all
(two episodes of the same TV show); hence black frames and silence detections like
in [
          <xref ref-type="bibr" rid="ref4">4</xref>
          ] are not enough to locate telecasts boundaries on every channels at every hours.
        </p>
        <p>About time efficiency, we can barely imagine a shot or scene detection followed by
an extraction of semantic features on the whole document, mobilizing frame analysis
– for instance: face detection, tracking, or even recognition, motion extraction – and
audio analysis – for example: speaker recognition, applause and laugh detection, or
transcription – before the results are integrated in order to give a label to each part of
the document. Our will is to reduce the number of detections that occurred during the
structuring process. We use knowledge to produce hypotheses about the document
structure and then we check them by local detections. By the way, program guides,
which are delivered to the institute and TV magazines at least one week before the
broadcast, already give a first idea of the global structure since it presents the main
programs’ hours. These program guides are obviously imprecise but the real difficulty
is their incompleteness: short programs like magazines, weather forecast, services or
lotteries don’t appear in these guides, advertisements are occulted, and sometimes
planed programs are canceled or replaced by another one.</p>
        <p>
          In our work, we consider the telecasts scheduling as a markovian process [
          <xref ref-type="bibr" rid="ref5">5</xref>
          ].
Classical Markov models didn’t fit the real broadcasting because of its being
timeindependent: programs succession doesn’t only depend on the kind of the last
program but on the hour and the day it was broadcasted. For instance, the early news
is followed by a soap opera whereas the news at the prime-time is followed by short
magazines. Further details on our model can be found in [
          <xref ref-type="bibr" rid="ref7">7</xref>
          ].
        </p>
        <p>
          The learning phase is initialized with INA’s program schedules of the past years.
After the learning phase, the model can improve a forecast program guide by adding
all the telecasts which don’t appear in: it’s a statistical completion. Table 1 presents
the predicted schedule of a broadcast day. Improved schedules give a temporal
window within witch the system can find a program boundary and can improve the
method used in [
          <xref ref-type="bibr" rid="ref8">8</xref>
          ] by eliminating false alarms and undetected boundaries. To be sure
a telecast of the stream fit a predicted telecast, we can use detections like theme, face
or logo recognition.
        </p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>3 Towards Semantic Extraction from Program Guides</title>
      <p>The second phase of the manual video stream indexing at INA consists in describing
telecasts of interest like news or magazines. Automatic indexing can provide
information on telecasts which were not manually indexed, and can help to describe
the program during the manual indexing of the others.</p>
      <p>We saw program guides can be used to structure the video stream; but they can
also provide information about telecasts. We can distinguish two classes of metadata:
objective ones are very useful since it concerns production details like year,
nationality, title and actors, and subjective ones like the summary which is less useful
for INA but constitute a good departure for the program description.</p>
      <p>There are many ways to get automatically these metadata: we can either get them
on a TV magazine website or with metadata broadcast with the stream or online
accessible like TV-ANYTIME or XML-TV (fig. 1) born with the numerical TV
growth.</p>
      <p>&lt;tv&gt;
&lt;programme channel="fr2" start="20010829095500 BST"&gt;
&lt;title&gt;King of the Hill&lt;/title&gt;
&lt;sub-title&gt;Meet the Propaniacs&lt;/sub-title&gt;
&lt;desc&gt; Bobby tours with a comedy troupe […] &lt;/desc&gt;
&lt;credits&gt;</p>
      <p>&lt;actor&gt;Mike Judge&lt;/actor&gt;
&lt;/credits&gt;
&lt;/programme&gt;
&lt;/tv&gt;
We are still implementing the system and we have just finished the learning module
and improving the one for the statistical completion. We have to finish and
experiment the part of the system that gets online metadata with the XML-TV
grabber.</p>
      <p>The next task will be to implement the boundary detector in order to experiment
the efficiency of the use of semantic features from electronic program guides for the
boundaries detection.
5</p>
    </sec>
    <sec id="sec-4">
      <title>Conclusion</title>
      <p>In order to create a system that can separate, label and describe the different programs
in a stream representing a broadcast week, it is useful to know the different structures
it can have. Detecting advertisements in this stream allows isolating the majority of
the programs but not all, and the predicted schedules can give the system a temporal
window within which it can find the programs boundaries.</p>
      <p>We presented in this article a way to improve the veracity of forecast schedules
which reflects the structure of the document with a Markov model which is used to
learn the broadcast habits of a channel. Our experiments show, in spite of the
combinatory explosion, we can make a forecast schedule exhaustive even if it stays
temporally imprecise.</p>
      <p>Finally, we presented also a way to get from online or electronic program guides
metadata about the various telecasts and show how it can help to structure the stream.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <given-names>R.</given-names>
            <surname>Chaniac</surname>
          </string-name>
          , and
          <string-name>
            <given-names>J.P.</given-names>
            <surname>Jezequel</surname>
          </string-name>
          <article-title>: La television</article-title>
          .
          <source>La découverte</source>
          (
          <year>2005</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <given-names>S.</given-names>
            <surname>Fischer</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Lienhart</surname>
          </string-name>
          , and W. Effelsberg:
          <article-title>Automatic recognition of film genres</article-title>
          .
          <source>Proc. ACM Multimedia</source>
          (
          <year>1995</year>
          )
          <fpage>295</fpage>
          -
          <lpage>304</lpage>
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3. L. Fonnet:
          <article-title>La programmation d'une chaîne de television</article-title>
          .
          <source>Dixit</source>
          (
          <year>2003</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <given-names>X.</given-names>
            <surname>Naturel</surname>
          </string-name>
          , and P. Gros: Etiquetage automatique de programmes de television.
          <source>CORESA</source>
          (
          <year>2005</year>
          ) [to appear]
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <given-names>J.R.</given-names>
            <surname>Norris</surname>
          </string-name>
          :
          <article-title>Markov chains</article-title>
          . Cambridge Series in Statitical and Probabilistic
          <string-name>
            <surname>Mathematics</surname>
          </string-name>
          (
          <year>1997</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <given-names>J.P.</given-names>
            <surname>Poli</surname>
          </string-name>
          and
          <string-name>
            <given-names>J.</given-names>
            <surname>Carrive</surname>
          </string-name>
          :
          <article-title>Proposition d'une architecture pour un système de structuration de flux audiovisuals</article-title>
          .
          <source>CORESA</source>
          (
          <year>2005</year>
          ) [to appear]
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <given-names>J.P.</given-names>
            <surname>Poli</surname>
          </string-name>
          <article-title>: Predicting program guides for video structuring</article-title>
          .
          <source>ICTAI</source>
          (
          <year>2005</year>
          ) [to appear]
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>C. G. M. Snoek</surname>
            , and
            <given-names>M.</given-names>
          </string-name>
          <article-title>Worring: Multimodal video indexing : a review of the state-of-theart</article-title>
          .
          <source>Multimedia Tools ans Applications</source>
          vol.
          <volume>25</volume>
          (
          <year>2005</year>
          )
          <fpage>5</fpage>
          -
          <lpage>35</lpage>
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <given-names>Y.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Liu</surname>
          </string-name>
          and
          <string-name>
            <given-names>J.</given-names>
            <surname>Huan</surname>
          </string-name>
          :
          <article-title>Mutimedia content analysis using both audio and visual clues</article-title>
          .
          <source>IEEE Signal processing magazine</source>
          , vol.
          <volume>17</volume>
          (
          <year>2000</year>
          )
          <fpage>12</fpage>
          -
          <lpage>36</lpage>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>