<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>A Unified, Modular and Multimodal Approach to Search and Hyperlinking Video</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Jonathon Hare jsh</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>@ecs.soton.ac.uk</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Jamie Davies Neha Jain jagd</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>@ecs.soton.ac.uk nj</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>@ecs.soton.ac.uk</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>David Dupplaw</institution>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Electronics and Computer Science, University of Southampton</institution>
          ,
          <country country="UK">United Kingdom</country>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>John Preston</institution>
        </aff>
        <aff id="aff3">
          <label>3</label>
          <institution>Sina Samangooei</institution>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2013</year>
      </pub-date>
      <fpage>18</fpage>
      <lpage>19</lpage>
      <abstract>
        <p>This paper describes a modular architecture for searching and hyperlinking clips of TV programmes. The architecture aimed to unify the combination of features from di erent modalities through a common representation based on a set of probability density functions over the timeline of a programme. The core component of the system consisted of analysis of sections of transcripts based on a textual query. Results show that search is made worse by the addition of other components, whereas in hyperlinking precision is increased by the addition of visual features.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. INTRODUCTION</title>
      <p>
        The 2013 MediaEval search an hyperlinking task [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]
tackles two problems; search across and within video collections,
and hyperlinking of short video segments relevant to a given
anchor segment. This paper describes the system we built
to address the two tasks. The motivation behind the design
of this system was to provide a uniform way of combining
features across di erent modalities of data.
      </p>
    </sec>
    <sec id="sec-2">
      <title>OVERALL APPROACH</title>
      <p>Our overall idea was to represent each programme by a
probability density function (PDF) over the timeline of the
programme. The area under the PDF between two time
points essentially represents the probability of that portion
of the programme to being relevant to the query. By
constructing PDFs for each programme with a given query, we
can then locate the high-probability segments of the PDFs,
which in turn tell us the beginning and end times of hits
that can be returned to the user. With respect to the search
and hyperlinking tasks, the primary di erence is the form of
the query.</p>
      <p>The architecture backing the approach was modular, and
we constructed various modules to incorporate data from
di erent data sources. We developed two forms of
module; the rst was capable of using the query to generate
new data points for the PDFs for a set of programmes, and
the second was capable of weighting the entire PDF for a
given programme (i.e. to increase or decrease the global
relevance). The modules responsible for adding to the PDF
worked by placing Gaussian functions at the point of
interest on the timeline, with variance proportional to the length
of the segment of interest. At the end, the overall PDF for a
programme can be computed from summation of the
Gaussians at every time point; this entire process can be viewed
as a variable bandwidth kernel density estimation.</p>
      <p>Most of the work is done by modules working on the
textual data (transcripts, synopsis, program titles). This
information was indexed using Lucene1 with separate elds for
each source. Each module is described brie y below.
2.1</p>
    </sec>
    <sec id="sec-3">
      <title>Generating Modules</title>
      <p>The transcript search module was the most important
component of the system. The module searched for
keywords (taken from the query string) across all transcripts of
a certain kind (LIMSI/Vocapia, LIUM, or subtitles) using
the Lucene index. Keyword matches were extracted from
each transcript in turn and grouped using hierarchical
agglomerative clustering over the time di erences between hits
within a programme. When a cluster's separation
(calculated as the temporal distance between the average values
of the cluster's left and right children) fell below a
specied threshold, a cluster was formed, and used to build a
Gaussian whose amplitude was calculated from</p>
      <p>jQj w2WQ
= jWQj X boost(w) idf(w)
where Q was the set of all possible keywords from the query,
WQ was the subset of query keywords that appeared in the
transcript, idf : W ! R was a function mapping each
keyword on to its inverse document frequency, and boost : W !
R was a function mapping each keyword on to its Lucene
query boost. Additionally, the true amplitude was scaled
by the normalised score returned by Lucene when searching
for transcript documents matching the query. The
amplitude of the Gaussian captures the relevance of all keywords
in the cluster with respect to the document, as well as how
completely the cluster covers the set of all possible query
terms. The Gaussians were centred on the midpoint of the
range covered by the cluster, and the standard deviation of
the Gaussian was chosen as one third of the temporal size
of the cluster plus 60 seconds.</p>
      <p>The concept module analysed the query text and visual
cues for known concepts that could be added to timelines.
The amplitude for the concept module's Gaussians was
determined from the normalised con dence for each concept
1http://lucene.apache.org
Run code
S M Mod
U M Mod
I M Mod
S MV ModCon
U MV ModCon
I MV ModCon
S MV ModConLSH
U MV ModConLSH
I MV ModConLSH
detection, and the standard deviation was a constant 5
seconds.</p>
      <p>
        The visual information module worked by nding shots
that were visually similar to other shots with high con
dence. For each programme, the most stable key-frame of
each shot was extracted and SIFT features were calculated.
Each SIFT feature was hashed using locality-sensitive
hashing (LSH) and a graph was constructed where the vertices
were keyframes and edges were created if pairs of key-frames
contained colliding features [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]. The module found sections
of timelines corresponding to shots whose integrals exceeded
a threshold (i.e. shots already deemed relevant by the
preceding modules), and added Gaussians centred on the shots
whose keyframes were directly connected to this keyframe
on the LSH graph. The base amplitude of the Gaussians
was determined as the fraction of functions under which the
two keyframes collided to the largest number of collisions.
A constant width of 60 seconds was used. This module was
implemented using OpenIMAJ 2 [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ].
2.2
      </p>
    </sec>
    <sec id="sec-4">
      <title>Weighting modules</title>
      <p>The synopsis and title modules increased the weight of
timelines belonging to programmes whose keywords matched
keywords in the synopsis and title elds of the index. The
channel lter module performed nave NLP on the query:
if a channel was mentioned in the query, then any
timelines corresponding to programmes on other channels were
removed from the timeline set (i.e. setting the weight to
zero).</p>
    </sec>
    <sec id="sec-5">
      <title>3. SEARCHING AND HYPERLINKING</title>
      <p>The architecture described in the previous section was
used to facilitate both the search and hyperlinking tasks.
For the search task, the system was con gured to take the
query and pass it directly to each module. Concepts were
inferred from the query text and the visual information
module was used for query expansion, through the detection of
visually similar segments to high-con dence detections from
the other modules. For hyperlinking, the transcript of the
anchor segment was used as the query text (together with
the synopsis, title and channel of the programme from which
the anchor was drawn). Concepts detected in the anchor
were used as input to the concept module. The visual
information module was used to nd segments that were visually
similar to the anchor.</p>
    </sec>
    <sec id="sec-6">
      <title>RESULTS</title>
      <p>The results from the search task are summarised in
Table 1. Runs using the subtitles gave the best performance
in each category, which is understandable due to the more
accurate nature of subtitles compared with speech-to-text
transcripts. It is interesting to observe that the performance
of the system decreased as additional features were brought
in, which may indicate that these additional modules were
not scaled properly, or otherwise very noisy in the context of
the query. This is surprising for concept detection, as in the
search task concepts were directly picked from the query.</p>
      <p>Table 2 shows the results for the hyperlinking task. It can
be seen that the baseline results with just the subtitle
information used give the best results. Adding the concept
detections causes a drop in performance, indicating that similar
visual concepts do not necessarily indicate relevant content.
The addition of the SIFT-LSH features looks like it might
slightly improve performance; this needs further veri cation
without the concepts.
5.</p>
    </sec>
    <sec id="sec-7">
      <title>CONCLUSION</title>
      <p>The system performed better at the hyperlinking sub-task
than the search sub-task; this was slightly unexpected as the
search performance on the development data was higher.
This may in part be due to fundamental limits within the
transcript module: textual queries have a low bandwidth
and describe many features that are not discernible from a
programme's transcript, and thus a more complex approach
might be required to improve performance. Additional NLP
on queries, along with person detection (i.e. face
recognition/veri cation), could also improve performance in the
search domain.</p>
      <p>For both tasks the addition of visual information tended
to harm overall performance. Again, this was slightly
unexpected, as we saw improvements in the development search
queries when using this information. One possible reason for
this is that the visual information is only useful for certain
types of query (or anchor). It would be interesting to explore
this further; a starting point for this would be to analyse the
results on a per-item basis (rather than just looking at the
overall averages).</p>
    </sec>
    <sec id="sec-8">
      <title>ACKNOWLEDGMENTS 6. 7.</title>
    </sec>
    <sec id="sec-9">
      <title>ADDITIONAL AUTHORS</title>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>M.</given-names>
            <surname>Eskevich</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G. J.</given-names>
            <surname>Jones</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Chen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Aly</surname>
          </string-name>
          , and
          <string-name>
            <given-names>R.</given-names>
            <surname>Ordelman</surname>
          </string-name>
          .
          <article-title>The Search and Hyperlinking Task at MediaEval 2013</article-title>
          . In MediaEval 2013 Workshop, Barcelona, Spain, October
          <volume>18</volume>
          -19
          <year>2013</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>J. S.</given-names>
            <surname>Hare</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Samangooei</surname>
          </string-name>
          , and
          <string-name>
            <given-names>D. P.</given-names>
            <surname>Dupplaw</surname>
          </string-name>
          .
          <article-title>OpenIMAJ and ImageTerrier: Java libraries and tools for scalable multimedia analysis and indexing of images</article-title>
          .
          <source>In ACM MM'11</source>
          , pages
          <fpage>691</fpage>
          {
          <fpage>694</fpage>
          . ACM,
          <year>2011</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>J. S.</given-names>
            <surname>Hare</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Samangooei</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D. P.</given-names>
            <surname>Dupplaw</surname>
          </string-name>
          , and
          <string-name>
            <given-names>P. H.</given-names>
            <surname>Lewis</surname>
          </string-name>
          .
          <article-title>Twitter's visual pulse</article-title>
          .
          <source>In ICMR'13</source>
          , pages
          <fpage>297</fpage>
          {
          <fpage>298</fpage>
          , New York, NY, USA,
          <year>2013</year>
          . ACM.
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>