<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Intelligent Search in a Collection of Video Lectures</article-title>
      </title-group>
      <contrib-group>
        <aff id="aff0">
          <label>0</label>
          <institution>University of Trento, Dept. of Information and Communication Tech.</institution>
          ,
          <addr-line>Via Sommarive 10, 38050 Trento</addr-line>
          ,
          <country country="IT">Italy</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>In recent years, the use of streamed digital video as a teaching and learning resource has become an increasingly attractive option for many educators as an innovation which expands the range of learning resources available to students by moving away from static text-and-graphic resources towards a video-rich learning environment. Streamed video is already widely used in some universities and it is mostly being used for transmitting unenhanced recordings of live lectures. What we are proposing is a way of enriching this video streaming scenario in the eLearning context. We want to extract information from the video and the correlated materials and make them searchable. Thus, the aim of this thesis is to create a semantically searchable collection of video lectures. In the literature, surprisingly little information can be found about speech and document retrieval in combination with lecture recording. There are interesting examples of e-lecture creation and delivery e.g. [5], audio retrieval of lecture recording [3] that explore automatic processing of speech, or systems such as the eLecture portal [1] which indexes the audio and also the text of the lecture slides. But to the best of our knowledge there is no system which combines and synchronizes the different modalities in a searchable collection. What we propose is enabling the search and navigation through the different media types presented in a frontal lecture with the addition of the video recording. In video indexing domain Snoek and Worring [4] have proposed to define multimodality as ”the capacity of an author of the video document to express a predefined semantic idea, by combining a layout with a specific content, using at least two information channels”. The channels or modalities of a video document described in [4] are the visual, auditory and textual modality. We believe that using more than one modality - as explained in [2] - could increase productivity also in the context of e-learning, where it is really frequent to scan for information. Our main focus is to enable search on two modalities; in particular we will index the auditory modality of video lecture content based on transcription obtained with automatic speech recognition tools and on textual modality using text indexing on the related materials. Furthermore, we do not just want to present an enhanced version of the current state of the art in</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>e-lecture retrieval but we also envision to add semantic capabilities to the search
functionalities in order to provide a superior learning experience to the student
(personalized search, personalized learning path, relevant contextualization,
automatic video content profiling...).</p>
      <p>The application we are proposing is a way to provide a tool for students
to enable more flexibility in e-Lectures consumption. The student could seek
inside a collection of learning lectures and related materials (desktop
activities recording, PowerPoint presentations, interactive whiteboard tracks...). The
search could also be personalized to meet the student demands. For each hit the
system would display the lectures video-recording and the temporally
synchronized learning materials. Another benefit we want to archive using Semantic Web
techniques is to present a profile information of the content of the video lecture,
this could lead to an improvement of the state of the art in the video-indexing
field allowing automatic profile annotation of the content of the video.</p>
      <p>The research work specifically addressed by this thesis will investigate the
following challenges:
– Finding an innovative way for mastering the gap between information
extraction and knowledge representation in our context. For each video and
related learning resource an RDF representation would be extracted. The
created graph would be navigated during the search task to find the
requested information and suggest related topics using ontology linkage. We
will use ontologies for high level lecture description and query understanding.
– Automatic content description of the presented learning material. A textual
description of the content of a video result, lecture or course would be
presented at the user. This could be realized presenting a profile information of
the knowledge extracted from the video and the related material.
– Evaluation of the tool value for improving the student performance and for
shortening the learning time we will conduct before and after the semantic
enhancement.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>Christoph</given-names>
            <surname>Hermann</surname>
          </string-name>
          , Wolfgang Hu¨rst, and Martina Welte.
          <article-title>The electure portal: An advanced archive for lecture recordings</article-title>
          .
          <source>In Informatics Education Europe Conference</source>
          ,
          <year>2006</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>D.</given-names>
            <surname>Jones</surname>
          </string-name>
          ,
          <string-name>
            <surname>Shen A.</surname>
          </string-name>
          , and
          <string-name>
            <given-names>D W.</given-names>
            ,
            <surname>Reynolds</surname>
          </string-name>
          .
          <article-title>Two experiments comparing reading with listening for human processing of conversational telephone</article-title>
          .
          <source>In Interspeech 2005 Eurospeech Conference</source>
          ,
          <year>2005</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>A.</given-names>
            <surname>Park</surname>
          </string-name>
          , T. Hazen, and
          <string-name>
            <given-names>J.</given-names>
            <surname>Glass</surname>
          </string-name>
          .
          <article-title>Automatic processing of audio lectures for information retrieval: Vocabulary selection and language modeling</article-title>
          . In ICASSP, March
          <year>2005</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>C.G.M.</given-names>
            <surname>Snoek</surname>
          </string-name>
          and
          <string-name>
            <given-names>M.</given-names>
            <surname>Worring</surname>
          </string-name>
          .
          <article-title>Multimodal video indexing: A review of the stateof-the-art</article-title>
          .
          <source>In Multimedia Tools and Applications</source>
          , number
          <volume>25</volume>
          , pages
          <fpage>5</fpage>
          -
          <lpage>35</lpage>
          ,
          <year>2005</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>Z.</given-names>
            <surname>Zhu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>McKittrick</surname>
          </string-name>
          , and
          <string-name>
            <given-names>W.</given-names>
            <surname>Li</surname>
          </string-name>
          .
          <article-title>Virtualized classroom automated production, media integration and user-customized presentation</article-title>
          .
          <source>In Multimedia Data and Document Engineering</source>
          ,
          <year>July 2004</year>
          . Semantic Web.
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>