<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Multimedia Summarization in Law Courts: An Environment for Browsing and Consulting</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>E. Fersini</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>G. Arosio</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>E. Messina</string-name>
          <email>messinag@disco.unimib.it</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>F. Archetti</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>D. Toscani</string-name>
          <email>toscanig@milanoricerche.it</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Consorzio Milano Ricerche</institution>
          ,
          <addr-line>Via Cicognara 7 - 20129 Milano</addr-line>
          ,
          <country country="IT">Italy</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>DISCo, Universita degli Studi di Milano-Bicocca</institution>
          ,
          <addr-line>Viale Sarca, 336 - 20126 Milano</addr-line>
          ,
          <country country="IT">Italy</country>
        </aff>
      </contrib-group>
      <fpage>73</fpage>
      <lpage>80</lpage>
      <abstract>
        <p>Digital videos represent a fundamental informative source of those events that occur during a penal proceedings, which thanks to the technologies available nowadays, can be stored, organized and retrieved in short time and with low cost. Considering the dimension that a video source can assume with respect to a courtroom recording, various necessities have been highlighted by the main judicial actors: fast navigation of the stream, e cient access to data inside and e ective representation of relevant contents. One of the possible solutions to these requirements is represented by multimedia summarization aimed at deriving a synthetic representation of audio/video contents, characterized by a limited loss of meaningful information. In this paper a multimedia summarization environment is proposed for de ning a storyboard for proceedings celebrated into courtrooms.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>Multimedia summarization techniques analyze several informative sources
comprises into a multimedia document, with the aim at extracting a semantic
abstract. Multimedia summarization techniques available in literature can be
divided in three main categories: (1) internal techniques, which exploit low level
features of audio, video and text; (2) external techniques, which refer to the
information typically associated with a viewing activity and interaction with the
user; (3) hybrid techniques, which combine internal and external information.
These techniques are focused on di erent types of features: (a) domain speci c,
i.e. typical characteristics of a given domain known a priori and (b) non-domain
speci c, i.e. non-generic features associated with a particular context.</p>
      <p>
        With respect to internal techniques the main goal is to analyze low-level
features derived from text, images and audio content within a multimedia
document. In [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ] the performance related to the identi cation of special events are
increased by combining scene recognition techniques with OCR-based approaches
for subtitles recognition in baseball video documents. In [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ] entire scenes
containing text are recognized using OCR techniques for a subsequent identi cation
Post-proceedings of the 2nd International Conference
on ICT Solutions for Justice (ICT4Justice 2009)
of key events through audio and video features in football matches. In [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ] the
semantics of objects and events occurring within news video are extracted from
subtitles and used to specialize / improve the systems of automatic speech
recognition.
      </p>
      <p>
        In order to reduce the semantic gap between low level features and semantic
concept for producing a meaningful summary, research is moving towards the
inclusion of external information that usually include knowledge of the context
in which evolves a multimedia document and user-based information. The
techniques able to generate a video summary on the basis of external information
are limited to three case studies [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ] [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ] [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ] focused on using domain speci c
features. In [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ] a summarization technique is proposed in order to gather context
information from the acquisition/registration phase and collected by monitoring
the movement of citizens around their houses. Cameras at a speci c position and
pressure sensors are used to track users. Since users are not required to provide
any kind of information, the summary produced by analyzing data concerning
the movement (such as the distance between steps and direction changes). In
[
        <xref ref-type="bibr" rid="ref5">5</xref>
        ] and [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ] semantic annotations collected during the production phase of the
video and described by the standard MPEG-7 are analyzed. In particular in
[
        <xref ref-type="bibr" rid="ref6">6</xref>
        ] a sequence of audio-video segments is produced on the basis of annotations
from video sports (baseball matches), such as players' names or speci c events
occurring during the match. In [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ] a video, characterized by a set of MPEG-7
macro-semantic annotations collected during the acquisition phase, is further
annotated by users in order to indicate their level of interest in each video
segment. The associations between preferences and the macro-annotations are then
modelled by using supervised learning approaches to enable the generation of
automatic summary of new multimedia documents.
      </p>
      <p>
        An attempt that tries to combine the peculiarities of the previous techniques
is represented by Hybrid Techniques. Hybrid summarisation techniques combine
the advantages provided by internal and external approaches by analyzing a
combination of internal and external information. As overviewed for the
previous techniques, the hybrid ones can be distinguished in domain speci c and
non-domain speci c. Examples of domain-speci c hybrid techniques are related
to music videos [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ], broadcast news [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ] and movies [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ]. In non-domain speci c
approaches we can nd:
- in [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ] the summarization approach could be described by two stages: (1) frames
are grouped by a clustering approach, using colour image features; (2) in editing
phase, manual annotations of representative frame for each cluster, with
subsequent spread to frames of the same cluster, are required. Summary for a new
video is then generated by using representative elements of each cluster that
generates a matching with user query.
- in [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ] an annotation tool is used during the editing phase in order to propagate
semantic descriptors to non-labelled contents. During the summary generation
phase, the user pro le is considered in order to create a customized synthetic
representation.
      </p>
      <p>By analyzing the state of the art related to multimedia summarization, no
evidences about summary over courtroom proceedings are given. In this paper
we are mainly addressing the problem of deriving a storyboard of a multimedia
document coming from penal proceedings recordings, by proposing an external
summarization technique based on the unsupervised clustering algorithm named
Induced Bisecting K-Means. The main outline of this paper is the following.
In section 2 the proposed multimedia summarization environment is presented.
In section 2.1 the work ow for deriving a storyboard for the judicial actors is
described. In section 3 details about the exploited clustering algorithm are given.
Finally, in section 4 conclusions are derived.
2</p>
    </sec>
    <sec id="sec-2">
      <title>Multimedia Summarization Environment</title>
      <p>In order to address the problem of de ning a short and meaningful representation
of a debate that is celebrated within a law courtroom, we propose a
multimedia summarization environment based on unsupervised learning. The main goal
of this environment is to create a storyboard of either a hearing or an entire
proceedings, by taking into account the semantic information embedded into a
courtroom recording.</p>
      <p>In particular, the main information sources used for producing a multimedia
summary are represented by:
{ automatic speech transcriptions that correspond to what is uttered by the
actors involved into hearings/proceedings
{ automatic audio annotations coming from emotional states recognition (for
example fear, neutral, anger)
{ automatic video annotations that correspond to what happen during a
debate (for instance change of witness posture, new witness, behavior of a given
actor)</p>
      <p>The Multimedia Summarization Environment includes two di erent modules:
the acquisition module and the summarization module.</p>
      <p>{ The acquisition module, given a query speci ed by the end user, retrieves
multimedia information from the Multimedia Database in terms of
audiovideo track(s), speech transcription and semantic annotations.
{ The summarization module is aimed at producing on-demand storyboard by
exploiting the information retrieved by the acquisition module. The summary
is created by focusing on maximally query-relevant passages and reducing
cross-document redundancy.</p>
      <p>A simple overview of the modules involved into the multimedia summarization
environment is depicted in gure 1 (a).
2.1</p>
      <p>Multimedia Summarization Work ow
In order to summarize a multimedia document according to the user needs, a
query statement is de ned for acquiring requirements in terms of query and to
start the entire work ow (see gure 1 (b)).
ProIDc e s  
HeaIDr i ng </p>
      <p>Preprocessing 
 
Multime dia 
DB 
Data 
Representatio</p>
      <p>n 
Multimedia 
Summarizatio
n 
Algorithm 
Data 
structure 
Storyboard 
Construction 
(a) Overview of the multimedia (b) Overview of the multimedia
summarization macro-modules summarization work ow</p>
      <p>The user query is speci ed at the graphic interface level, where a list of
trials are available, in terms of keywords in which we are interested (whatever
is uttered by the involved speaker, the emotional state of actors, etc...).
Once the query has been speci ed, it is submitted to the pre-processing module.
The aim of this module is to optimize the user query by eliminating noise and
by reducing the size of vocabulary, i.e. stop words removal and stemming are
performed to enhance retrieval performance on transcription and annotations.</p>
      <p>After the preprocessing activity, the query will be ready to be submitted to
the retrieval module, which is aimed at accessing to the multimedia database
in order to retrieve all the information that match the user query in terms of
transcription of the debate, audio and videos annotations (audio and video tag).
At this level, two possibilities are given to the end user: to retrieve and summarize
an entire trial that matches the query or to summarize only those sub-parts of
the proceedings that match the query. In the rst case the user query is used
to retrieve the multimedia documents related to a trial by executing a
highlevel skimming of the overall database. After this initial step all the clips of the
retrieved hearing are considered for producing the summary. In the second case
the query is used to scan the database in a more exhaustive way so that, within
a given trial, only the audio, video and textual clips that completely match the
user query are retrieved.</p>
      <p>In both cases we refer to a (audio, video and textual) clip as a consecutive
portion of a debate in which there is one speaker whom is active, i.e. there exist
a sequence of words uttered by the same speaker without breaking due to other
speakers. In this way we have one clip of audio/video tag and transcription for
each speaker period.</p>
      <p>The next step in the multimedia summarization work ow relates to data
representation module. The aim of this module is to combine information coming
from di erent sources in order to create a uni ed representation. This activity
is performed through a feature vector representation, where all the information
able to characterize the audio, video and textual clip of interest are managed as
features and weights. Examples of features exploited by this representation are
given by the textual transcription, the audio and video tag, the start and end
time of the relevant sub-parts of the debate.</p>
      <p>
        Given the data structure that has been created, the multimedia
summarization module may start the summary generation. The core component is based
on a clustering algorithm named Induced Bisecting K-means [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ]. The
algorithm creates a hierarchical organization of (audio, video and textual) clips, by
grouping in several clusters hearings or sub-parts of them according to a given
similarity metric. This algorithm is able to build a dichotomic tree in which
coherent concepts are grouped together, i.e. each cluster created by the algorithm
contain a set of audio, video and textual clips containing similar concepts that
are also coherent with the query submitted by user.
      </p>
      <p>The last step related to the storyboard construction, where the nal storyboard is
derived from the dichotomic tree structure coming from the previous multimedia
summarization activity. Given the dichotomic tree, a pruning step is performed
in order to choose only those clusters that respect a given intra-cluster
similarity threshold. Suppose that the pruning activity after the Induced Bisecting
K-means returns a set of clusters as reported in gure 2 where C1, C2 and C3
are the resulting clusters and the clips named 1; : : : ; 9 represent the sub-parts of
the debate. The storyboard construction activity considers the representative
elements of each cluster (centroids) as the relevant audio, video and textual clips
for the summary. The storyboard is generated starting from the centroids by
presenting to the end user the rst frame, together with the references of the
given audio, video and textual clip, references of the trial/hearing, start and end
time of the segments and so on. By referencing gure 2, only the rst frames
related to segments 2, 5 and 7 (representative of the obtained 3 clusters) are
presented to the end user as pictures that could be clicked to start the
corresponding audio-video portion.</p>
      <p>In the following subsection details about the core component of the
multimedia summarization environment, i.e. the Induced Bisecting K-Means clustering
algorithm, are given.
3</p>
    </sec>
    <sec id="sec-3">
      <title>The hierarchical clustering algorithm</title>
      <p>
        The approaches proposed in the literature for hierarchical clustering were mostly
statistical with a high computational complexity . A novel approach, Bisecting
k-Means was proposed in [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ], has a linear complexity and is relatively e cient
and scalable. It starts with a single cluster of multimedia clips and works in the
following way:
Algorithm 1 Bisecting K-Means
1: Pick a cluster S of clips ml to split
2: Select two random seeds which are the initial representative clips (centroids)
3: Find 2 sub-clusters S1 and S2 using the basic k-Means algorithm
4: Repeat step 2 and 3 for ITER times and take the split that produces the clustering
with the highest Intra Cluster Similarity (ICS)
5: ICS(Sk) = 1 P cosine(mi; mj)
      </p>
      <p>jSkj2 mi;mj2Sk
6: Repeat steps 1, 2 and 3 until the desired number of clusters is reached.</p>
      <p>
        The major disadvantage of this algorithm is that it requires the a priori
speci cation of the number of clusters K and the parameter ITER for creating
several splits of the same group in order to choose the best one. An incorrect
estimation of K and ITER may lead to poor clustering accuracy. Moreover, the
algorithm is sensitive to the noise which may a ect the computation of cluster
centroids. For any given cluster let N be the number of clips belonging to that
cluster and R the set of their indices. In fact, the jth element of a cluster centroid
1
used by the k-Means algorithm during step 3 is computed as cj = P mij (r)
N r2R
where N represents the number of clips belonging to the cluster and mrj (r) is
the vectorial representation of j features of clip i belonging to cluster r. The
centroid c may contain also the contribution of noisy features contained into the
clip representation which the pre-processing phase have not been able to remove.
To overcome these two problems we exploit an extended version of the Standard
Bisecting k-Means, named Induced Bisecting k-Means [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ], whose main steps
are described as follows:
Algorithm 2 Induced Bisecting K-Means
1: Set the Intra Cluster Similarity (ICS) threshold parameter
2: Build a distance matrix A whose elements are given by the Euclidean distance
between clips aij = rP(mik mjk)2 where i and j are clips
      </p>
      <p>k
3: Select, as centroids, the two clips i and j s.t. aij = max Alm
l;m
4: Find 2 sub-clusters S1 and S2 using the basic k-Means algorithm
5: Check the ICS of S1 and S2 as
6: If the ICS value of a cluster is smaller than , then reapply the divisive process
to this set, starting form step 2
7: If the ICS value of a cluster is over a given threshold, then stop. 6. The entire
process will nish when there are no sub-clusters to divide.</p>
      <p>
        The main di erences of this algorithm with respect to the Standard Bisecting
k- Means consist in: (1) how the initial centroids are chosen: as centroids of the
two child clusters we select the clips of the parent cluster having the greatest
distance between them; (2) the cluster splitting rule: a cluster is split in two if
its Intra Cluster Similarity is smaller than a threshold parameter . Therefore,
the optimal number of cluster K is controlled by the parameter and therefore
no input parameters K and ITER must be speci ed by the user. Our algorithm
outputs a binary tree of clips, where each node represents a clip collection which
elements are similar. This structure is processed according to [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ], in order to
obtain a at representation of clusters.
4
      </p>
    </sec>
    <sec id="sec-4">
      <title>Conclusion and Future Work</title>
      <p>In this paper a multimedia summarization environment has been presented in
order to allow judicial actors to browse and navigate multimedia documents
related to penal hearings/proceedings. The main component of this environment
is represented by the summarization module, which create a storyboard for the
end user by exploiting several semantic information embedded into a courtroom
recording. In particular, automatic speech transcriptions joint with automatic
audio and video annotations have been used for deriving a compressed and
meaningful representation of what happens into a law courtroom. Our work is now
focused on creating a testing environment for a quality assessment of the
storyboard produced by our environment.</p>
    </sec>
    <sec id="sec-5">
      <title>Acknowledgment</title>
      <p>This work has been supported by the European Community FP-7 under the
JUMAS Project (ref.: 214306).</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <given-names>J.</given-names>
            <surname>Kim</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Chang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Kang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Kim</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Kim</surname>
          </string-name>
          ,
          <article-title>Summarization of news video and its description for content-based access</article-title>
          ,
          <source>International Journal of Imaging Systems and Technology</source>
          ,
          <volume>267</volume>
          -
          <fpage>274</fpage>
          ,
          <year>2004</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <given-names>C.</given-names>
            <surname>Liang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Kuo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W.</given-names>
            <surname>Chu</surname>
          </string-name>
          , J. Wu,
          <article-title>Semantic units detection and summarization of baseball videos</article-title>
          ,
          <source>in: Proc. of the 47th Midwest Symposium on Circuits and Systems</source>
          , pp.
          <fpage>297</fpage>
          -
          <lpage>300</lpage>
          ,
          <year>2004</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <given-names>D.W.</given-names>
            <surname>Tjondronegoro</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Chen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Pham</surname>
          </string-name>
          ,
          <article-title>Classi cation of selfconsumable highlights for soccer video summaries</article-title>
          ,
          <source>in Proc. of the IEEE International Conference on Multimedia and Expo</source>
          , pp.
          <fpage>579</fpage>
          -
          <lpage>582</lpage>
          ,
          <year>2003</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4. G. de Silva, T. Yamasaki,
          <string-name>
            <given-names>K.</given-names>
            <surname>Aizawa</surname>
          </string-name>
          ,
          <article-title>Evaluation of video summarization for a large number of cameras in ubiquitous home</article-title>
          ,
          <source>in: Proc. of the 13th Annual ACM International Conference on Multimedia</source>
          , pp.
          <fpage>820</fpage>
          -
          <lpage>828</lpage>
          ,
          <year>2005</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <given-names>A.</given-names>
            <surname>Jaimes</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Echigo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Teraguchi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Satoh</surname>
          </string-name>
          ,
          <article-title>Learning personalized video highlights from detailed MPEG-7 metadata</article-title>
          , in
          <source>: Proc. of the IEEE International Conference on Image Processing</source>
          , pp.
          <fpage>133</fpage>
          -
          <lpage>136</lpage>
          ,
          <year>2002</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <given-names>Y.</given-names>
            <surname>Takahashi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Nitta</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Babaguchi</surname>
          </string-name>
          ,
          <article-title>Video Summarization for Large Sports Video Archives</article-title>
          ,
          <source>in Proc. of the IEEE International Conference on Multimedia and Expo</source>
          , pp.
          <fpage>1170</fpage>
          -
          <lpage>1173</lpage>
          ,
          <year>2005</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <given-names>L.</given-names>
            <surname>Agnihotri</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Dimitrova</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.R.</given-names>
            <surname>Kender</surname>
          </string-name>
          ,
          <article-title>Design and evaluation of a music video summarization system</article-title>
          ,
          <source>in Proc. of the IEEE International Conference on Multimedia and Expo</source>
          , pp.
          <fpage>1943</fpage>
          -
          <lpage>1946</lpage>
          ,
          <year>2004</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <given-names>H.</given-names>
            <surname>Yang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Chaisorn</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Zhao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Neo</surname>
          </string-name>
          , T. Chua,
          <article-title>VideoQA: question answering on news video</article-title>
          ,
          <source>in: Proc. of the 11th Annual ACM International Conference on Multimedia</source>
          , pp.
          <fpage>632</fpage>
          -
          <lpage>641</lpage>
          ,
          <year>2003</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <given-names>T.</given-names>
            <surname>Moriyama</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Sakauchi</surname>
          </string-name>
          ,
          <article-title>Video summarization based on the psychological unfolding of drama, Systems and</article-title>
          Computers in Japan, pp
          <fpage>1122</fpage>
          -
          <lpage>1131</lpage>
          ,
          <year>2002</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <given-names>Y.</given-names>
            <surname>Rui</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.X.</given-names>
            <surname>Zhou</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.S.</given-names>
            <surname>Huang</surname>
          </string-name>
          ,
          <article-title>E cient access to video content in a uni ed framework</article-title>
          ,
          <source>in: Proc. of the IEEE International Conference on Multimedia Computing and Systems</source>
          , pp.
          <fpage>735</fpage>
          -
          <lpage>740</lpage>
          ,
          <year>1999</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <string-name>
            <given-names>B.L.</given-names>
            <surname>Tseng</surname>
          </string-name>
          ,
          <string-name>
            <surname>C.-Y.L.J.R. Smith</surname>
          </string-name>
          ,
          <string-name>
            <surname>Using</surname>
            <given-names>MPEG</given-names>
          </string-name>
          <article-title>-7 and MPEG-21 for personalizing video</article-title>
          ,
          <source>IEEE Transactions on Multimedia, pp42-52</source>
          ,
          <year>2004</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          12.
          <string-name>
            <surname>M. Steinbach</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          <string-name>
            <surname>Karypis</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          <string-name>
            <surname>Kumar</surname>
          </string-name>
          ,
          <article-title>A comparison of Document Clustering Techniques</article-title>
          ,
          <source>In KDD Workshop on Text Mining</source>
          ,
          <year>2000</year>
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          13.
          <string-name>
            <given-names>F.</given-names>
            <surname>Archetti</surname>
          </string-name>
          , E. Fersini,
          <string-name>
            <given-names>P.</given-names>
            <surname>Campanelli</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Messina</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A Hierarchical</given-names>
            <surname>Document</surname>
          </string-name>
          <article-title>Clustering Environment Based on the Induced Bisecting k-Means</article-title>
          ,
          <source>in Proc. of Flexible Query Answering Systems</source>
          ,
          <volume>4027</volume>
          /2006
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>