<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Content Based Video Retrieval System for Distorted Video Queries</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Boris Tseytlin</string-name>
          <email>batseytlin@edu.hse.ru</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Ilya Makarov</string-name>
          <email>iamakarov@hse.ru</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>National Research University, Higher School of Economics</institution>
          ,
          <addr-line>Moscow</addr-line>
          ,
          <country country="RU">Russia</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>We consider the task of content-based video retrieval (CBVR) given a query video, which is expected to match if it is a distorted short subsequence of a reference video from a database. In this paper, we present a CBVR system architecture that is both robust and scalable. We use a modified rHash frame fingerprint generation method. It is both, extremely robust to distortions and fast to compute. We utilize the Faiss library, developed by Facebook Research, to index fingerprint binary vectors. The VCDB dataset is used for benchmarking.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;information retrieval</kwd>
        <kwd>multimedia database</kwd>
        <kwd>content-based video retrieval</kwd>
        <kwd>rHash</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>
        In this paper we consider the task of searching a database of reference videos given a query video. A
query video is expected to match if it is a distorted short subsequence of a reference video (also referred
to as partial copy). In literature the task is known as content-based video retrieval (CBVR) [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ],
contentbased video search [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ], near-duplicate video matching [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ] or content based video copy detection [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ].
In all of these formulations the task involves searching for videos that contain a subsequence similar
to the query video. To avoid confusion, we will refer to the task as CBVR further on.
      </p>
      <p>
        The CBVR community benefited a lot from the advancements in content-based image retrieval
(CBIR) [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ], which is focused on searching a database of images for partial copies. It’s common to use
frame features as the basis for CBVR systems. A CBVR problem can be approached as a CBIR problem,
however using temporal information present in video sequences have proven to provide better results
as seen in [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ].
      </p>
      <p>The task of content-based video retrieval is challenging because of various distortions and content
variations that query videos might be subject to. These include noise, compression artifacts, rotation,
framerate alternation, logo attacks, frame removal, scale and lightning changes. Scalability is also
a major concern, because videos contain thousands of frames, which leads to long query times and
large memory requirements. Most approaches to CBVR require calculating the distance from the
query video to all subsequences of all reference videos. Clearly, this approach becomes infeasible for
real-world applications as database volume grows.</p>
      <p>Robust hashing, also known as video fingerprinting, is a popular approach for partial copy
detection. It involves generating a content-based signature assigned to video frames, and matching query
video signatures to reference video signatures. In robust hashing, the hash function should be stable
in regards to visual distortions.</p>
      <p>
        In this paper we present a CBVR system architecture that is both robust and scalable. We use a
modified rHash [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ] frame fingerprint generation method. It’s both extremely robust to distortions
and fast to compute. We utilize the Faiss library [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ], developed by Facebook Research, to index
fingerprint binary vectors. This is done to speed up database queries. The VCDB dataset is used for
benchmarking.
      </p>
      <p>The rest of the paper is organized as follows. A brief overview of the related studies is presented in
Section 2. In Section 3, the system architecture is presented. The video file preprocessing steps, the
ifngerprint generation approach, database indexing and searching strategies are described. In Section
4 we present our evaluation results on the VCDB dataset. Finally, Section 5 presents our conclusions.</p>
    </sec>
    <sec id="sec-2">
      <title>2. Related work</title>
      <p>Feature extraction is a central part of any CBVR system. Pairwise comparison of frames using
pixelby-pixel measures is ineficient. It’s necessary to find a compact representation of video sequences
in lower dimensionsionality. This is known as feature extraction or feature generation. A feature
generation approach aims to strike a balance between matching accuracy, fast searching and compact
database size.</p>
      <p>Video feature generation approaches can be classified into categories:
• color-based features
• temporal features
• spatial features
• deep learning based features
• fusion of diferent types of features</p>
      <p>
        An overview of image feature extraction techniques is provided in [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ]. Color-based features are
usually derived from histograms of pixel intensities or colors. Color histograms are sensitive to
specific color spaces, which makes them sensitive to video formats. They are also highly afected by
noise, saturation changes. However, a local region color histogram feature can be highly resistant to
distortions. In that approach, the image is divided into non-overlapping regions, a color histogram is
calculated for each region, and then histogram bin counts are concatenated together to form a feature
vector.
      </p>
      <p>
        Temporal features are based on time and frame diferences. They are extracted over time from a
video sequence. In [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ] a temporal based sequence matching method was proposed. In this method
each frame is divided into a grid of average pixel intensity values, and grids are stored in a ranking
sequence. This provided a global and local description of temporal variation.
      </p>
      <p>
        Spatial features are generated from individual frames. In video fingerprinting, the feature vector of
an image is a binary hash. Average hash (aHash), proposed in [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ] is one of the most commonly used
ifngerprinting methods. This method is simple and fast to compute. It uses the mean intensity value
of a grayscale image to binarize it by thresholding. The output is a short binary sequence representing
an image. [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ] proposed an extension to this method called rHash. The extension is simple yet efective
- each frame is divided into non-overlapping blocks, block mean intensity values are calculated, and
each block is binarized by thresholding on the median of block mean values. The result is a short
binary sequence of predefined size. This hash has proven to be more robust than aHash. It’s also easy
to compute. [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ] proposed DCT hash, which is efective, but requires complex transformations.
      </p>
      <p>
        Image descriptors utilizing SIFT, SURF and other key point extraction algorithms are often used
to generate spatial features. Roth et al. [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ] proposed a spatial feature based on SURF point counts.
They divided each frame into non-overlapping blocks and counted the amount of key points in each
frame. The resulting key point counts were used as frame descriptions. SURF points are known to
be extremely robust to image manipulations, including mirroring and rotation. In [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ] the idea was
expanded by adding a temporal component. SURF point counts of each frame were used to generate
a block ranking matrix. Block ranks produced a vector that described the whole video sequence.
      </p>
      <p>
        Usually generating features from each video frame is redudant. Commonly, key frames and frame
rate downsampling are used for speeding up feature extraction. In 2010 [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ] proposed the concept
of temporally informative representative images (TIRI). This method aims to aggregate sequences of
frames together without loosing much information. It has been used in many studies ever since [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ]
[
        <xref ref-type="bibr" rid="ref12">12</xref>
        ] [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ].
      </p>
      <p>
        There have been multiple attempts at using neural networks to extract features for video retrieval
tasks. A deep learning approach achieved state-of-the-art results on the VCDB dataset [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ]. [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ] used
a VGG-16 architecture network, together with dimensionality reduction via PCA, to generate features
for video copy detection. [
        <xref ref-type="bibr" rid="ref17">17</xref>
        ] describes a few cases of "Siamese twin" networks being used for feature
generation. An approach using rHash [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ] managed to obtain results comparable to deep-learning
generated features, whilist keeping the computation costs low.
      </p>
      <p>
        The scalability issue of CBVR was studied thoroughly. A CBVR system should be able to retrieve
similar video sequences quickly while operating on a very large database of signatures. Inverted
ifle indexes are commonly used in practice [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ] [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ] [
        <xref ref-type="bibr" rid="ref17">17</xref>
        ]. Faiss [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ] by Facebook Research and Annoy
[
        <xref ref-type="bibr" rid="ref18">18</xref>
        ] by Spotify are tools that provide in-memory inverted file indexes. In Locality Sensitive Hashing
(LSH) higher dimensional data is projected into a lower dimensionality representation using random
projections [
        <xref ref-type="bibr" rid="ref19">19</xref>
        ]. [20] proposed an extension to LSH for cases of searching for set queries, which is
more eficient when dealing with feature sets, such as SURF descriptors.
      </p>
      <p>
        The overview of CBVR datasets conducted by [
        <xref ref-type="bibr" rid="ref17">17</xref>
        ] shows that most works on topic benchmark
their work on either TRECVID [21] or VCDB [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ] datasets. VCDB is relatively new, however it’s the
only dataset containing real partial copies. Other datasets include only simulated copies. This made
VCDB the standard dataset for benchmarking CBVR systems.
      </p>
    </sec>
    <sec id="sec-3">
      <title>3. System overview</title>
      <p>The basic architecture of the proposed CBVR system is presented on Figure 1. The diagram shows
two scenarios - the index generation flow indicated by red arrows, when the database is populated
with reference videos, and the querying flow indicated by blue arrows.</p>
      <p>First, videos are transcoded to a predefined size and framerate, and converted to grayscale, to make
the system robust to framerate alternations. Then TIRI frames are generated. A modified rHash
called Quadrant rHash is obtained from each TIRI frame. In case of index generation, the obtained
hash vectors are stored in a binary vector index using Faiss. In case of querying, the obtained hash
vectors are matched against the database. In the final step, fingerprint matching results are used to
make a decision on whether the query video matches any reference videos.</p>
      <sec id="sec-3-1">
        <title>3.1. TIRI extraction</title>
        <p>
          Temporally informative representative images (TIRI) [
          <xref ref-type="bibr" rid="ref13">13</xref>
          ] are obtained as weighted average of
sequential frames. First, a sequence of frames  1,  2, … ,   extracted from a video is divided into
nonoverlapping blocks of length  . For each block a TIRI frame  ′ is computed as a weighted average
of pixel intensities of frames in block. There are many ways to pick weights   . We have chosen
exponential weights   =   with  = 1.65 as it has been proven to be efective by previous research
[
          <xref ref-type="bibr" rid="ref13">13</xref>
          ] [
          <xref ref-type="bibr" rid="ref4">4</xref>
          ].
        </p>
        <p>,′ = ∑ =+1    ∑=+    ⋅  ,, = ∑ =+1    ∑=+    ⋅  ,,
Where  is the frame block start index, ,  are pixel positions.</p>
        <p>
          An example is provided on Figure 2.
3.2. Quadrant rHash
rHash has been proposed in [
          <xref ref-type="bibr" rid="ref4">4</xref>
          ] as a more robust extension to the popular aHash. In rHash, each frame
is divided into non-overlapping blocks. A sequence of block mean intensity values  = ( 1,  2, … ,   )
is obtained. The sequence is then binarized using the following rule:
        </p>
        <p>{
ℎ( ) =
1, if   ≥ 
otherwise</p>
        <p>( )
ℎ( ) =
1, if   ≥</p>
        <p>otherwise</p>
        <p>We propose a further simple extension to rHash. In Quadrant rHash, the mean values sequence
is divided into four non-overlapping blocks  1,  2,  3,  4. The sequence is then binarized using
median values of corresponding blocks   . The binarization rule becomes:
{</p>
        <p>(  ), where   ∈ { 1,  2,  3,  4} s.t.   ∈  
An example is presented on Figure 3.</p>
        <p>Quadrant rHash vectors can be eficiently compared using Hamming distance. Computing them
involves only simple operations like computing the mean, so it’s extremely fast.</p>
      </sec>
      <sec id="sec-3-2">
        <title>3.3. FAISS index</title>
        <p>
          Faiss is a library for eficient approximate k-nearest neighboors search. It’s a good choice for
scalable systems because it has been developed with scale in mind [
          <xref ref-type="bibr" rid="ref6">6</xref>
          ]. It’s possible to utilize it’s GPU
parallelization capabilities for speeding up search.
        </p>
      </sec>
      <sec id="sec-3-3">
        <title>3.4. Video retrieval strategy</title>
        <p>Given query hash vectors, extracted from a query video, a matching reference video needs to be found.
We compared two strategies for video matching.</p>
        <p>Majority vote (MV). Majority vote is a simple matching strategy. The hypotheses behind this
method is formulated this way: if most frames of a query video are closest to a reference video,
then the query video is most likely a partial copy of that reference video. This method involves the
following steps:
1. For each query frame  , find the tuple ( , ,  ), where  is a reference video frame with minimal
Hamming distance from  ,  is the distance value, and  is the id of the video  belongs to.</p>
        <p>This is a very fast operation with Faiss.
2. Threshold the tuples with a distance threshold value   . For each tuple ( , ,  ), if  &gt;   ,
replace  with a special value −1, indicating lack of match.
3. Construct a sequence of retrieved video ids:  1,  2, … ,   . Where   is the video id retrieved
for   ,  is the query length.
4. Return the most frequent item in  1,  2, … ,   . It can either be a video id, or −1. −1 indicates,
that the query video didn’t match any videos in database.</p>
        <p>
          Longest sequence with candidate elimination (LS). This is a slightly more complex matching
strategy, similar to the approach used in [
          <xref ref-type="bibr" rid="ref4">4</xref>
          ]. It makes use of temporal information. It involves the
following steps:
1. For each query frame obtain  = 10 closest frames in database. Construct a sequence of tuples:
( 1,  1,  1,  1,  1), ( 1,  2,  2,  2,  2), … , (  ,   × ,   × ,   × ,   × )
. Here each tuple represents a potential matched frame,   is the index of a query frame,  is
reference video frame,  is the Hamming distance value,  is the time in seconds at which 
occurs in video,  is the video id. Threshold tuples on a distance threshold   .
        </p>
        <p>The first row contains an image (left) and it’s Quadrant rHash (right). The second row contains the same
image, but with noise, rotation, rescaling, blur applied. The third row contains a diferent image, included for
comparison. It’s evident that hashes of the first two images are visually similar, whilist a diferent image
produces a diferent hash.
2. Consider each two tuples (  ,   ,   ,   ,   ) and (  ,   ,   ,   ,   ). Eliminate  -th tuple if:   &gt;   and   &lt;
  . We remove those potential matches that are chronologically impossible: if query frame 
occurs after query frame  , matched frame   must occur after   .
3. For each unique video id   present in the sequence of potential matches, count the amount of
tuples that contain   . Return the video id with the largest count of matches.</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>4. Experiments and results</title>
      <p>We evaluated the performance of our method on the VCDB core dataset. VCDB is an annotated
dataset of real partial copies. It contains 519 videos, about 27 hours of video data in total. Partial
copies in this dataset are subject to compression artifacts, framerate alternations, picture-in-picture,
logo attacks, inserted frames, frame removal, noise, blur and more. It’s a very challenging dataset for
Matching strategy</p>
      <p>MV</p>
      <p>LS
baseline
a CBVR system.</p>
      <p>
        We have taken 9 videos as the reference videos, 106 partial copies of them as query videos and
the rest 413 videos as distraction queries. Our experiments have shown that best results are achieved
with resizing each video to 200x200, downsampling the framerate to 5 frames per second, extracting
TIRI from every  = 5 frames. Similar conclusions were made in [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ]. We extracted Quadrant rHashes
of dimensionality 100. The distance threshold value was chosen as   = 20 based on experiments.
      </p>
      <p>The traditional F-score measure was used for evaluation:
 =</p>
      <p>+</p>
      <p>=
 +</p>
      <p>× 
 = 2</p>
      <p>+</p>
      <p>All experiments were ran using Python 3.6, on an Ubuntu 18 Linux operation system, 2.50GHz Intel
Core i5-7200U CPU.</p>
      <p>
        The evaluation results are presented in Table 1. We compare obtained F-score results to [22], which
used traditional rHash and temporal network matching approaches. We compare obtained query
times to [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ] which used a special inverted file-based index.
      </p>
      <p>The longest sequence matching method achieves better performance at the cost of longer query
processing times. Surprisingly, majority voting achieves a result similar to more complex methods.
Our system is able to conduct very fast searches. LS provides more fine-grained results at the cost of
searching times.</p>
    </sec>
    <sec id="sec-5">
      <title>5. Conclusions</title>
      <p>Content-based video retrieval is a challenging task due to many disturbances that query videos might
be subject to. We propose a system architecture based on a modified image fingerprinting measure
called Quadrant rHash that is very fast to compute and robust against image distortions. Temporally
informative representative images (TIRI) are used during video preprocessing. In our approach, the
Faiss library is used to build an index on obtained binary vectors. The proposed system was
evaluated on the well-known VCDB core dataset. Two diferent video sequence matching strategies were
evaluated. Experimental results have shown a slight performance increase when compared to similar
approaches.
[20] P. Nagarkar, K. S. Candan, Pslsh: An index structure for eficient execution of set queries in
highdimensional spaces, in: Proceedings of the 27th ACM International Conference on Information
and Knowledge Management, ACM, 2018, pp. 477–486.
[21] G. Awad, P. Over, W. Kraaij, Content-based video copy detection benchmarking at trecvid, ACM
Trans. Inf. Syst. 32 (2014) 14:1–14:40. URL: http://doi.acm.org/10.1145/2629531. doi:10.1145/
2629531.
[22] L. Mengyang, L.-M. Po, Z. Chang, W. Y. Yuen, H.-K. Cheung, H. Peter, H.-T. Luk, K.-W. Lau,
Content-based video copy detection using binary object fingerprints, in: 2018 IEEE International
Conference on Signal Processing, Communications and Computing (ICSPCC), IEEE, 2018, pp. 1–
6.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>S. A.</given-names>
            <surname>Bhat</surname>
          </string-name>
          ,
          <string-name>
            <given-names>O. V.</given-names>
            <surname>Sardessai</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P. P.</given-names>
            <surname>Kunde</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S. S.</given-names>
            <surname>Shirodkar</surname>
          </string-name>
          ,
          <article-title>Overview of existing content based video retrieval systems</article-title>
          ,
          <source>International Journal of Advanced Engineering and Global Technology</source>
          <volume>2</volume>
          (
          <year>2014</year>
          )
          <fpage>476</fpage>
          -
          <lpage>483</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>E.</given-names>
            <surname>Esen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Ozkan</surname>
          </string-name>
          ,
          <string-name>
            <surname>I. Atil</surname>
          </string-name>
          ,
          <article-title>Large-scale video search with eficient temporal voting structure</article-title>
          ,
          <source>arXiv preprint arXiv:1607.07160</source>
          (
          <year>2016</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>X.</given-names>
            <surname>Wu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.-W.</given-names>
            <surname>Ngo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A. G.</given-names>
            <surname>Hauptmann</surname>
          </string-name>
          , H.
          <article-title>-</article-title>
          <string-name>
            <surname>K. Tan</surname>
          </string-name>
          ,
          <article-title>Real-time near-duplicate elimination for web video search with content and context</article-title>
          ,
          <source>IEEE Transactions on Multimedia</source>
          <volume>11</volume>
          (
          <year>2009</year>
          )
          <fpage>196</fpage>
          -
          <lpage>207</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>M.</given-names>
            <surname>Liu</surname>
          </string-name>
          , L.
          <string-name>
            <surname>-M. Po</surname>
            ,
            <given-names>Y. A. U.</given-names>
          </string-name>
          <string-name>
            <surname>Rehman</surname>
            ,
            <given-names>X.</given-names>
          </string-name>
          <string-name>
            <surname>Xu</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          <string-name>
            <surname>Li</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          <string-name>
            <surname>Feng</surname>
          </string-name>
          ,
          <article-title>Video copy detection by conducting fast searching of inverted files</article-title>
          ,
          <source>Multimedia Tools and Applications</source>
          <volume>78</volume>
          (
          <year>2019</year>
          )
          <fpage>10601</fpage>
          -
          <lpage>10624</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>G.</given-names>
            <surname>Amato</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Bolettieri</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Carrara</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Falchi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Gennaro</surname>
          </string-name>
          ,
          <article-title>Large-scale image retrieval with elasticsearch</article-title>
          ,
          <source>in: The 41st International ACM SIGIR Conference on Research &amp; Development in Information Retrieval</source>
          , ACM,
          <year>2018</year>
          , pp.
          <fpage>925</fpage>
          -
          <lpage>928</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>J.</given-names>
            <surname>Johnson</surname>
          </string-name>
          , M. Douze,
          <string-name>
            <given-names>H.</given-names>
            <surname>Jégou</surname>
          </string-name>
          ,
          <article-title>Billion-scale similarity search with gpus</article-title>
          ,
          <source>arXiv preprint arXiv:1702.08734</source>
          (
          <year>2017</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>T.</given-names>
            <surname>Deselaers</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Keysers</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Ney</surname>
          </string-name>
          ,
          <article-title>Features for image retrieval: an experimental comparison</article-title>
          ,
          <source>Information retrieval 11</source>
          (
          <year>2008</year>
          )
          <fpage>77</fpage>
          -
          <lpage>107</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>L.</given-names>
            <surname>Chen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Stentiford</surname>
          </string-name>
          ,
          <article-title>Video sequence matching based on temporal ordinal measurement</article-title>
          ,
          <source>Pattern Recognition Letters</source>
          <volume>29</volume>
          (
          <year>2008</year>
          )
          <fpage>1824</fpage>
          -
          <lpage>1831</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>B.</given-names>
            <surname>Yang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Gu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Niu</surname>
          </string-name>
          ,
          <article-title>Block mean value based image perceptual hashing</article-title>
          ,
          <source>in: 2006 International Conference on Intelligent Information Hiding and Multimedia</source>
          , IEEE,
          <year>2006</year>
          , pp.
          <fpage>167</fpage>
          -
          <lpage>172</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>M.</given-names>
            <surname>Malekesmaeili</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Fatourechi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Ward</surname>
          </string-name>
          ,
          <article-title>A robust and fast video copy detection system using content-based fingerprinting</article-title>
          ,
          <source>IEEE Transactions on Information Forensics and Security</source>
          <volume>6</volume>
          (
          <year>2011</year>
          )
          <fpage>213</fpage>
          -
          <lpage>226</lpage>
          . doi:
          <volume>10</volume>
          .1109/TIFS.
          <year>2010</year>
          .
          <volume>2097593</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>G.</given-names>
            <surname>Roth</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Laganière</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Lambert</surname>
          </string-name>
          , I. Lakhmiri,
          <string-name>
            <given-names>T.</given-names>
            <surname>Janati</surname>
          </string-name>
          ,
          <article-title>A simple but efective approach to video copy detection</article-title>
          ,
          <source>in: 2010 Canadian Conference on Computer and Robot Vision</source>
          , IEEE,
          <year>2010</year>
          , pp.
          <fpage>63</fpage>
          -
          <lpage>70</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>R. C.</given-names>
            <surname>Harvey</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Hefeeda</surname>
          </string-name>
          ,
          <article-title>Spatio-temporal video copy detection</article-title>
          ,
          <source>in: Proceedings of the 3rd Multimedia Systems Conference, ACM</source>
          ,
          <year>2012</year>
          , pp.
          <fpage>35</fpage>
          -
          <lpage>46</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <surname>M. M. Esmaeili</surname>
            ,
            <given-names>R. K.</given-names>
          </string-name>
          <string-name>
            <surname>Ward</surname>
          </string-name>
          ,
          <article-title>Robust video hashing based on temporally informative representative images</article-title>
          ,
          <source>in: 2010 Digest of Technical Papers International Conference on Consumer Electronics (ICCE)</source>
          , IEEE,
          <year>2010</year>
          , pp.
          <fpage>179</fpage>
          -
          <lpage>180</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <given-names>M.</given-names>
            <surname>Steinebach</surname>
          </string-name>
          , H. Liu,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Yannikos</surname>
          </string-name>
          , Forbild:
          <article-title>Eficient robust image hashing</article-title>
          ,
          <source>in: Media Watermarking, Security, and Forensics</source>
          <year>2012</year>
          , volume
          <volume>8303</volume>
          ,
          <string-name>
            <surname>International</surname>
            <given-names>Society</given-names>
          </string-name>
          <source>for Optics and Photonics</source>
          ,
          <year>2012</year>
          , p.
          <fpage>83030O</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <given-names>Y.-G.</given-names>
            <surname>Jiang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Jiang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <article-title>Vcdb: a large-scale database for partial copy detection in videos</article-title>
          ,
          <source>in: European conference on computer vision</source>
          , Springer,
          <year>2014</year>
          , pp.
          <fpage>357</fpage>
          -
          <lpage>371</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [16]
          <string-name>
            <given-names>L.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Bao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Fan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Luo</surname>
          </string-name>
          ,
          <article-title>Compact cnn based video representation for eficient video copy detection</article-title>
          , in: International conference on multimedia modeling, Springer,
          <year>2017</year>
          , pp.
          <fpage>576</fpage>
          -
          <lpage>587</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [17]
          <string-name>
            <given-names>Y.</given-names>
            <surname>Jiang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <article-title>Partial copy detection in videos: A benchmark and an evaluation of popular methods</article-title>
          ,
          <source>IEEE Transactions on Big Data</source>
          <volume>2</volume>
          (
          <year>2016</year>
          )
          <fpage>32</fpage>
          -
          <lpage>42</lpage>
          . doi:
          <volume>10</volume>
          .1109/TBDATA.
          <year>2016</year>
          .
          <volume>2530714</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          [18]
          <string-name>
            <surname>Spotify</surname>
          </string-name>
          , ANNOY library, https://github.com/spotify/annoy,
          <year>2017</year>
          . Accessed:
          <fpage>2017</fpage>
          -08-01.
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          [19]
          <string-name>
            <given-names>M.</given-names>
            <surname>Bawa</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Condie</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Ganesan</surname>
          </string-name>
          ,
          <article-title>Lsh forest: self-tuning indexes for similarity search</article-title>
          ,
          <source>in: Proceedings of the 14th international conference on World Wide Web, ACM</source>
          ,
          <year>2005</year>
          , pp.
          <fpage>651</fpage>
          -
          <lpage>660</lpage>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>