<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>DCU Linking Runs at MediaEval 2014: Search and Hyperlinking Task</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Shu Chen</string-name>
          <email>shu.chen4@mail.dcu.ie</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Gareth J. F. Jones</string-name>
          <email>gjones@computing.dcu.ie</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Noel E. O'Connor</string-name>
          <email>Noel.OConnor@dcu.ie</email>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>CNGL, School of Computing, Dublin City University</institution>
          ,
          <addr-line>Dublin 9</addr-line>
          ,
          <country country="IE">Ireland</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>INSIGHT Research Center /, CNGL, Dublin City University</institution>
          ,
          <addr-line>Dublin 9</addr-line>
          ,
          <country country="IE">Ireland</country>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>INSIGHT Research Center, Dublin City University</institution>
          ,
          <addr-line>Dublin 9</addr-line>
          ,
          <country country="IE">Ireland</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2014</year>
      </pub-date>
      <fpage>16</fpage>
      <lpage>17</lpage>
      <abstract>
        <p>We describe Dublin City University (DCU)'s participation in the Hyperlinking sub-task of the Search and Hyperlinking task MediaEval 2014. The investigation focuses on how to e ciently identify target segments in a large BBC TV dataset. In our submission, Linear Discriminant Analysis is used to estimate fusion weights for multimodal features with the objective of improving the reliability of creating the potential links.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. INTRODUCTION</title>
      <p>For our participation in the Hyperlinking sub-task of Search
and Hyperlinking of Television Content in MediaEval 2014,
we developed a new strategy of target segment
determination based on speaker identi cation. Linear Discriminant
Analysis was used to estimate fusion weights. The paper is
organized as follows: Section 2 describes a new strategy to
determine target segments, Section 3 describes fusion weight
determination based on Linear Discriminant Analysis,
Section 4 gives our experimental results, and Section 5
concludes the paper.</p>
    </sec>
    <sec id="sec-2">
      <title>SPEAKER-BASED TARGET SEGMENT</title>
    </sec>
    <sec id="sec-3">
      <title>DETERMINATION</title>
      <p>
        This is the third year of the MediaEval Search and
Hyperlinking task [
        <xref ref-type="bibr" rid="ref7 ref8">8, 7</xref>
        ]. The previous editions showed that
state-of-the-art IR techniques can be applied to multimodal
hyperlinking. However, identifying e ective target segments
is still an open issue. According to [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ], a target segment
should be moderate in length 10 to 120 seconds. In this
paper, a simple and e cient target segment determination
algorithm was developed by using the speaker identi cation
from the LIMSI transcript [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ].
      </p>
      <p>
        Similar to searching model, the retrieval list of target
segments are determined by the relevance to the query anchor.
Existing researches [
        <xref ref-type="bibr" rid="ref3 ref5">5, 3</xref>
        ] use a sliding window to cover a
variety of identifying target segments. We annotate that a
target segment is a collection of multimodal features with time
stamps, and those containing relevant information should
be allocated a higher rank. As a result, we identify target
segments as following. 1) Separate a video into a number
of clips. The separating benchmark is de ned according to
the multimodal feature distribution. Each separated clip is
de ned as a seed segment. 2) Each seed segment increases
the size of itself by merging adjacent segments. It stops
when the length of newly merged segment reaches the
target segment standard de ned in [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ]. 3) Repeat step 2 until
all the seed segments have been expanded. 4) Identify each
expanded segment as the target segment.
      </p>
      <p>
        The kernel of the algorithm is how to determine the seed
segments in a video and how to de ne the benchmark to
merge adjacent segments. Existing researches purpose that
the speech transcript often plays an important role in
effective hyperlinking [
        <xref ref-type="bibr" rid="ref3 ref5">5, 3</xref>
        ]. Our investigation is based on
LIMSI transcripts. Identifying a seed segment involves the
transcript and speaker analysis. We assume that a group
of spoken words, if presented by the same speaker, is
content related. LIMSI transcript provides speaker information
with each sentence. De ne each transcript sentence as seni
and its speaker ID as spki. Each video Vj is presented as
a vector Vj = [sen1; sen2; :::seni]. If k continuous sentences
[senn; senn+1; :::senn+k] satisfy spkn = spkn+1 = ::spkn+k,
they are merged to create a new seed segment sdl, whose
duration jsdlj is determined by the start time of senn and end
time of senn+k. Our algorithm traverses all sentences to
convert Vj into a new vector of seed segments [sd1; sd2; :::sdl].
We assume that two transcripts presented by di erent
speakers, if close enough, are possible to relate to the same topic.
As a result, our algorithm to create target segments is
dened as:
      </p>
      <p>Step 1: For each sdi in Vj , compare the two nearby
speech transcripts. Merge the one with lower time
interval to create a new target segment mgi. In this
paper, the time interval threshold is set to be 10 seconds,
following the minimum length of a target segment.
Step 2: Continue step 1 until the length of merged
segment mgi is longer than 2 minutes.</p>
      <p>Step 3: Save the merged segment mgi and process the
sdi+1 in Vj .</p>
      <p>Step 4: Finally all merged segments [mg1; mg2; :::mgk]
are regarded as potential target segments.
3.</p>
    </sec>
    <sec id="sec-4">
      <title>ESTIMATE FUSION WEIGHTS USING</title>
    </sec>
    <sec id="sec-5">
      <title>LINEAR DISCRIMINANT ANALYSIS</title>
      <p>
        Linear Discriminant Analysis (LDA) algorithm [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ]. was
used to estimate linear fusion weights. The linear
combination in LDA wT x can determine the fusing weights for
the multimodal features. We de ne a training group as jXj
and its corresponding sample groups as jY j, where yi 2 jY j
could be 0 or 1, meaning that the video segment to a speci c
query anchor can be either related or non-related.
Therefore, the training data was separated into two classi cations:
jX1j where x 2 X1 satis es p(xijyi = 0), and jX2j, where
x 2 X2 satis es p(xijyi = 1). Assuming both dataset
follows the normal distribution, we have the means 1, 2, the
covariance 1 and 2. The ratio between them was de ned
as the class variance Sb, the within class variance Sw, and
w as the vector of fusing weights. Sw and Sb can be de ned
as:
      </p>
      <p>Sb = (w
1
w</p>
      <p>2)2
Sw = (wT
1 w + wT
2 w)</p>
      <p>A linear combination was calculated by maximizing the
criterion of between class variance and within class variance.
De ne the criterion c as following:
c =
wT ( 1 + 2)
2)2</p>
      <p>
        According to [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ], a maximum separation can be achieved
by maximizing c by making the weight vector to follow :
w / ( 1 + 2) 1 ( 1
2)
      </p>
      <p>MAP evaluation on DCU Hyperlinking</p>
    </sec>
    <sec id="sec-6">
      <title>EXPERIMENT</title>
      <p>
        The experiment data is introduced in MediaEval 2014
overview paper [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ]. Our experiment is constructed based
on only LIMSI transcripts. Two runs are investigated to
evaluate the performance of speaker-based target segment
determination and fusion weight estimation using LDA. All
2 runs creates the target segments using speaker-based
determination algorithm. RUN 1 involves the late fusion model
on metadata and LIMSI transcripts. RUN 2 involves only
LIMSI transcripts. The fusion equation is de ned as
following:
      </p>
      <p>FusionScore = wm</p>
      <p>MetadataScore + wt LIMSIScore (5)</p>
      <p>
        In RUN 1, the linear fusion weights wm and wm are
determined by LDA algorithm. The hyperlinking retrieval on
metadata and LIMSI transcripts are processed using
TDIDF model on text features. Lucene 4.9.0 is applied to
implement indexing and searching. The details are provided in
[
        <xref ref-type="bibr" rid="ref4">4</xref>
        ]. The retrieval model in RUN 2 uses Lucene
implementation as well.
      </p>
      <p>
        Table 1 shows the Mean Average Precision (MAP) value of
each runs. The evaluation metrics is de ned in [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. The
experiment reveals that LDA estimation on linear fusion work
could improve the hyperlinking quality, as presented in
Table 1 that RUN 1 has an overall increase on MAP, MAP bin
and MAP tol.
(1)
(2)
(3)
(4)
6.
7.
      </p>
    </sec>
    <sec id="sec-7">
      <title>CONCLUSIONS</title>
      <p>
        This paper presented details of DCU's participation in the
TV Data Hyperlinking task of MediaEval 2014. The
evaluation shows that LDA algorithm can improve the
hyperlinking performance by estimating multimodal fusion weights.
Our future work will be placed on comparing the
speakerbased target segment determination with other strategies
described in previous MediaEval research [
        <xref ref-type="bibr" rid="ref7 ref8">8, 7</xref>
        ].
      </p>
    </sec>
    <sec id="sec-8">
      <title>ACKNOWLEDGEMENT</title>
      <p>This work is funded by the European Commission's
Seventh Framework Programme (FP7) as part of the AXES
project (ICT-269980).</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>R.</given-names>
            <surname>Aly</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Eskevich</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Ordelman</surname>
          </string-name>
          , and
          <string-name>
            <given-names>G. J. F.</given-names>
            <surname>Jones</surname>
          </string-name>
          .
          <article-title>Adapting Binary Information Retrieval Evaluation Metrics for Segment-based Retrieval tasks</article-title>
          .
          <year>2013</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>G.</given-names>
            <surname>Balakrishnama. Linear Discriminant Analysis - A Brief Tutorial</surname>
          </string-name>
          ,
          <year>1998</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>S.</given-names>
            <surname>Chen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Eskevich</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G. J. F.</given-names>
            <surname>Jones</surname>
          </string-name>
          , and
          <string-name>
            <given-names>N. E. O'</given-names>
            <surname>Connor</surname>
          </string-name>
          .
          <article-title>An Investigation into Feature E ectiveness for Multimedia Hyperlinking</article-title>
          .
          <source>In MultiMedia Modeling</source>
          , pages
          <volume>251</volume>
          {
          <fpage>262</fpage>
          . Dublin, Ireland,
          <year>2014</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>S.</given-names>
            <surname>Chen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G. J. F.</given-names>
            <surname>Jones</surname>
          </string-name>
          , and
          <string-name>
            <given-names>N. E. O'</given-names>
            <surname>Connor. DCU Linking</surname>
          </string-name>
          <article-title>Runs at Mediaeval 2013: Search and Hyperlinking Task</article-title>
          . In MediaEval 2013 Workshop, Barcelona, Spain,
          <year>2013</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>S.</given-names>
            <surname>Chen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>McGuinness</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Aly</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N. O</given-names>
            <surname>'Connor</surname>
          </string-name>
          , and F. de Jong.
          <article-title>The AXES-lite video search engine</article-title>
          .
          <source>In Proceedings of WIAMIS 2012</source>
          , pages
          <issue>1</issue>
          {
          <fpage>4</fpage>
          , Dublin, Ireland,
          <year>2012</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>M.</given-names>
            <surname>Eskevich</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Aly</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Ordelman</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D. N.</given-names>
            <surname>Racca</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Chen</surname>
          </string-name>
          , and
          <string-name>
            <given-names>G. J. F.</given-names>
            <surname>Jones</surname>
          </string-name>
          .
          <source>The Search and Hyperlinking Task at Mediaeval 2014. In In Proceedings of the MediaEval 2014 Multimedia Benchmark Workshop</source>
          , Barcelona, Spain,
          <year>2014</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>M.</given-names>
            <surname>Eskevich</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G. J. F.</given-names>
            <surname>Jones</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Chen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Aly</surname>
          </string-name>
          , and
          <string-name>
            <given-names>R.</given-names>
            <surname>Ordelman</surname>
          </string-name>
          .
          <article-title>Search and Hyperlinking Task at MediaEval 2013</article-title>
          . In MediaEval 2013 Workshop, Barcelona, Spain,
          <year>2013</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>M.</given-names>
            <surname>Eskevich</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G. J. F.</given-names>
            <surname>Jones</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Chen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Aly</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Ordelman</surname>
          </string-name>
          , and
          <string-name>
            <given-names>M.</given-names>
            <surname>Larson</surname>
          </string-name>
          .
          <article-title>Search and Hyperlinking Task at MediaEval 2012</article-title>
          . In MediaEval 2012 Workshop, Pisa, Italy,
          <year>2012</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <surname>R. A. FISHER.</surname>
          </string-name>
          <article-title>The Use of Multimple Measurements in Taxonomic Problems</article-title>
          . pages
          <fpage>179</fpage>
          {
          <fpage>188</fpage>
          ,
          <year>1936</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>L.</given-names>
            <surname>Lamel</surname>
          </string-name>
          and
          <string-name>
            <given-names>J.-L.</given-names>
            <surname>Gauvain</surname>
          </string-name>
          .
          <article-title>Speech processing for audio indexing</article-title>
          .
          <source>In Advances in Natural Language Processing (LNCS 5221)</source>
          , pages
          <fpage>4</fpage>
          <lpage>{</lpage>
          15. Springer-Verlag,
          <year>2008</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>