<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>A Visual Search of Multimedia Documents in ImageCLEF 2014</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Sana FAKHFAKH</string-name>
          <email>sanafakhfakh@gmail.com</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Belhassen AKROUT</string-name>
          <email>bakrout@gmail.com</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Mohamed Tmar</string-name>
          <email>mohamed.tmar@isimsf.rnu.tn</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Walid MAHDI</string-name>
          <email>walid.mahdi@isimsf.rnu.tn</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Laboratoire Miracl ISIMS Pole Technologique de Sfax BP 242-3021, Sakiet Ezzit Sfax. TUNISIE</institution>
        </aff>
      </contrib-group>
      <fpage>715</fpage>
      <lpage>723</lpage>
      <abstract>
        <p>The Image search by content is an area that is based on a set of low-level features such as histograms, textures, the distribution of colors, shapes an brightness. The structure element is a shortcut of the image may vary depending on the designer of the XML document and which may change according to application needs. The visual appearance of the image is a permanent factor that undergoes no change. Therefore, we present in this paper a new method of search by content based on Harris detector, Haar wavelet and Color histogram for the step of feature extraction to extract a binary signature for all multimedia documents.Finally We estimate the similarities between codes and search relevant images with Hamming distance.</p>
      </abstract>
      <kwd-group>
        <kwd>Future extraction</kwd>
        <kwd>Plant identification</kwd>
        <kwd>Hamming distance</kwd>
        <kwd>Signature extraction</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>The research of the image content was the subject of numerous studies in recent
years. Techniques developed aim to extract and represent effectively the content
of images. The image indexing can be done either on the basis of low-level
features or those of semantics.</p>
      <p>
        The need for methods of indexing and retrieval directly based on the content of
the image is no longer evident. The first prototype system has been proposed
in 1970 and this system has attracted the attention of many researchers. Some
systems become commercial systems such as QBIC (Query By Image Content)
IBM [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ] and CIRES [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ].
      </p>
      <p>
        The Frip (Finding Regions in the Pictures) system [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ] KoB05 proposes to
research by regions of interest to drawn by the user [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ] . RETIN system (REcherche
and Hunt Interactive) [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ] developed at the University of Cergy-Pontoise France,
selects a random set of pixels in each image to extract their color values. The
texture of these pixels is obtained by applying the method of Gabor filters [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ].
These values are grouped and classified using a neural network.
The comparison between the images is done by calculating the similarity
between their feature vectors [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ]. Some studies have proposed to change their field
of research, for example changing the color feature space [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ] .
      </p>
      <p>
        Sometimes the user wants to query such as ”Find all the images of the database
with certain regions similar to those of the query image” [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ]. Some work has
been done, or the system segments an image or a spatial structure is used. In
the first case, having a segmented image, the user selects some regions that we
want from the segmented image. In the second case, the spatial structure is used.
A structure is often used quadtree [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ]. This structure is used to store the visual
characteristics of different image regions and filter images by increasing
progressively as the level of detail.
      </p>
      <p>
        Some work has proceeded to minimize the search space using the technique of
calculating the nearest neighbors to group similar data in classes [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ]. Thus the
search for an image is done by searching a class. The disadvantage of these
systems is that the user does not always have an image that expresses real need,
which makes use of such a system difficult.
      </p>
      <p>
        One solution to this problem is the technique of vectorization, which allows to
find relevant images to a query and which are not returned by the initial system.
Its sequence requires the choice of a set of said reference material. The choice of
these references is still problematic for the construction of the vector space.
In this context, it is necessary to select a heterogeneous set of documents. The
number of references is also problematic. Indeed, the number of documents is the
new dimension of vector space. If we choose a small set of documents, we may
not be able to properly represent all documents. These references are randomly
selected by [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ] , or the first results of an initial search or through the selection
of the centers of gravity of homogeneous grouping images [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ].
2
      </p>
    </sec>
    <sec id="sec-2">
      <title>Proposed approach</title>
      <p>In this section, we propose a description of the proposed visual system. In fact,
the search for multimedia documents requires a critical step in the abstract
offline indexing. This step requires the feature extraction from the image to give
a general description of documents. Following this course indexing we can go in
search of the visual image based on the descriptors already studied. Figure 1
describes the steps mentioned by an explanatory diagram.
2.1</p>
      <sec id="sec-2-1">
        <title>Feature extraction from the image</title>
        <p>
          The objective of extracting descriptors is the set up the search method contained
in the images. One example is the search for a database containing images of
plants mentioned. Images are indexed from the local descriptors such as Harris
detector [
          <xref ref-type="bibr" rid="ref13">13</xref>
          ] to calculate point of interest, the Haar wavelet [
          <xref ref-type="bibr" rid="ref14">14</xref>
          ] to analyze the
        </p>
        <p>Color Image</p>
        <p>Feature extraction
Harris detector</p>
        <p>Haar wavelet
decomposition</p>
        <p>Color histogram
Principal component analysis
Extraction of binary signature
0101010101000101010101001
textures and the calculation of the histogram color to represent the distribution
of intensities.</p>
        <p>
          Harris detector For twenty years, several detectors points of interest have
been developed. Schmid has compared the performance of several of them [
          <xref ref-type="bibr" rid="ref13">13</xref>
          ]
[
          <xref ref-type="bibr" rid="ref15">15</xref>
          ] . The famous Harris detector was evaluated as the best point detector and
has a good reputation in the field of automatic extraction points of interest.
In fact, Harris and Stephens [
          <xref ref-type="bibr" rid="ref13">13</xref>
          ] defined a sensor that recognizes the first
derivatives on a window signal representing an image and provide better results in two
important criteria [
          <xref ref-type="bibr" rid="ref15">15</xref>
          ] . The first criterion is based on their robustness with
respect to rotation and scale changes of the camera and points of view. The second
is summed up in his power with respect to variations of illumination.
Harris detector was applicable at the beginning only the grayscale image.
Montesinos [
          <xref ref-type="bibr" rid="ref15">15</xref>
          ] generalized Harris detector for color images. Figure 2 shows an
example of detection of interest points in an image.
        </p>
      </sec>
      <sec id="sec-2-2">
        <title>The decomposition of the image by Haar wavelet The wavelet transform</title>
        <p>of the sinusoid replaces the Fourier transform by a family of translations and
dilations of a single function. The translation parameters and expansion of the
two arguments are the wavelet transform.</p>
        <p>Haar transform was introduced in 1910 and is the oldest wavelet transform.
It is based on the demonstration of Haar (equation 1).
(1)
h(j, k; x) =</p>
        <p>2j/2
2H(2jx</p>
        <p>k)
−</p>
        <p>With H(x) = 1 f or x in ]0, 1/2[, H(x) = 1 f or x in ]1/2, 1[ and H(x) =
0 ailleurs form a complete orthogonal set for s−pace L2(R). The Haar wavelet is
simple to study and implement. Figure 3 shows the result of decomposition of an
image texture to analyze and extract the diagonal details, the vertical details,
the horizontal details and finally the description.</p>
        <p>Image analysis by color histogram For a monochrome image, the histogram
is defined as a discrete function that maps each intensity value the number of
pixels taking this value. The determination of the histogram is made by counting
the number of intensity for each pixel of the image. The quantification, which
includes several intensity values in one class can better visualize the distribution
of image intensities.</p>
        <p>Our job is to build a histogram of a color image in the RGB space, the red, green
and blue. Each component is quantized into eight intervals. In each interval of
each Red, green and blue component; we calculate the number of pixels
normalized frequency. We can then represent the image by a color histogram as shown
in Figure 4 following.
2.2</p>
      </sec>
      <sec id="sec-2-3">
        <title>Principal component analysis and signature extraction</title>
        <p>The principal component analysis ( P CA), developed in France in the 1960s by
JP.Benzcri is an exploratory statistical method used to describe a wide array of
data type individuals / variables. When individuals are described with 5 large
numbers of variables, no simple graphical representation allows to visualize the
point cloud formed by the data.</p>
        <p>The PCA provides a representation in a space of reduced size, thereby allowing
to highlight any structures in the data. For this, we look for subspaces where
the projection of the cloud deforms the least possible initial cloud.
Applying the P CA to the descriptors, previously computed, A signal f (x), that
includes all the characteristic values, is obtained. This signal is then subjected
to an automatic thresholding (equation 2 ) to the transformed into a binary
signature C(x) (equation 3).</p>
        <p>m
Mf = X f (x)
n=1
(2)</p>
        <p>In this way each image is represented by a binary code which is stored in a
database for later use in the research phase.</p>
      </sec>
      <sec id="sec-2-4">
        <title>Estimate similarities between codes and search relevant images</title>
        <p>
          Once the signatures are extracted, we must think of a technique that
quantifies the distance between them, in order to identify similar images. The bit-wise
comparison of each two image codes A and B each detail is given by the
normalization of the Hamming distance (HD) [
          <xref ref-type="bibr" rid="ref16">16</xref>
          ] defined as following:
⊗
With
        </p>
        <p>is the boolean operator ( XOR) described by the equation 5 follows
The probability that two bits of arbitrary code images are equal or not, worth
P = 0.5. When the majority of codes A(j) and B(j) are equal , is therefore in
the case where the two codes compared are equivalent. In the opposite case, the
two codes are different. Generally, when HD tends to zero it is called relevant
images.</p>
        <p>Our group submitted just one run in your rst participation in the ImageCLEF
Plant task 2014. In this paper we described a plant species retrieval model based
on a extraction of unique signature for each image in collection. This signature
is composed by Haar and Harris features.</p>
        <p>(a)
(3)
(4)
(5)</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>Experiments and Results</title>
      <sec id="sec-3-1">
        <title>Evaluation Metric</title>
        <p>
          The metric evaluation S was related to the rank of the correct species in the list
of retrieved species as follows [
          <xref ref-type="bibr" rid="ref17">17</xref>
          ]:
        </p>
        <p>S =
1 XU 1 XPu 1 NXu,p Su,p,n
U u=1 Pu p=1 Nu,p n=1
(6)
Where U :is the number of users.</p>
        <p>Pu : number of individual plants observed by the uth user.</p>
        <p>Nu,p : number of pictures taken from the pth plant observed by the uth user.
Su,p,n : score between 1 and 0 equals to the inverse of the rank of the correct
species (for the nth picture taken from the pth plant observed by the uth user).
3.2</p>
      </sec>
      <sec id="sec-3-2">
        <title>The Results</title>
        <p>IVprocessing team has submitted a single automatic run to the multi-image plant
observation queries. In this run, we extracted visual descriptors from images,
used texture, color and interest points model for generating the feature vectors.
The result is showed in Table 1. At a closer look, the score values of our results
are rather low. We cannot fully explain the failure with these plant image as
there were no indications for the weak performance during the training stage.
One reasonable interpretation of the results is that the conducted learning runs
led to an over-fitting.</p>
        <p>A total of 10 participating groups submitted 27 runs focusing on plant
observations. 6 teams submitted 14 complementary runs on images in ImageCLEF
2014. The following graphic shows the scores obtained on the main task on
multi-image plant observation queries.
We presented in this paper a research method by which the content is based
on extracting a signature for a given image. This signature is composed by a
set of descriptors, such as the analysis of the histogram of the color image, the
extraction of the interest points and Haar wavelet. This signature undergoes
a step dimensionality reduction by PCA method. An experimental study on
ImageCLEF 2014 showed the effectiveness of our descriptors.</p>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <given-names>Myron</given-names>
            <surname>Flickner</surname>
          </string-name>
          , Harpreet Sawhney, Wayne Niblack, Jonathan Ashley, Qian Huang, Byron Dom, Monika Gorkani, Jim Hafner,
          <string-name>
            <given-names>Denis</given-names>
            <surname>Lee</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Dragutin</given-names>
            <surname>Petkovic</surname>
          </string-name>
          , David Steele, and
          <string-name>
            <given-names>Peter</given-names>
            <surname>Yanker</surname>
          </string-name>
          .
          <article-title>Query by image and video content: The qbic system</article-title>
          . volume
          <volume>28</volume>
          , pages
          <fpage>23</fpage>
          -
          <lpage>32</lpage>
          , Los Alamitos, CA, USA,
          <year>September 1995</year>
          . IEEE Computer Society Press.
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <given-names>Kashif</given-names>
            <surname>Iqbal</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Michael O.</given-names>
            <surname>Odetayo</surname>
          </string-name>
          , and
          <string-name>
            <given-names>Anne</given-names>
            <surname>James</surname>
          </string-name>
          .
          <article-title>Content-based image retrieval approach for biometric security using colour, texture and shape features controlled by fuzzy heuristics</article-title>
          . volume
          <volume>78</volume>
          , pages
          <fpage>1258</fpage>
          -
          <lpage>1277</lpage>
          , Orlando, FL, USA,
          <year>July 2012</year>
          . Academic Press, Inc.
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <given-names>ByoungChul</given-names>
            <surname>Ko</surname>
          </string-name>
          and
          <string-name>
            <given-names>Hyeran</given-names>
            <surname>Byun</surname>
          </string-name>
          .
          <article-title>Frip: a region-based image retrieval tool using automatic image segmentation and stepwise boolean and matching</article-title>
          . volume
          <volume>7</volume>
          , pages
          <fpage>105</fpage>
          -
          <lpage>113</lpage>
          ,
          <year>2005</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <given-names>Yves</given-names>
            <surname>Caron</surname>
          </string-name>
          , Pascal Makris, and
          <string-name>
            <given-names>Nicole</given-names>
            <surname>Vincent</surname>
          </string-name>
          .
          <article-title>Caract´erisation d'une r´egion d'int´erˆet dans les images</article-title>
          .
          <source>In EGC</source>
          , pages
          <fpage>451</fpage>
          -
          <lpage>462</lpage>
          ,
          <year>2005</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <given-names>J.</given-names>
            <surname>Fournier</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Cord</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Philipp-Foliguet</surname>
          </string-name>
          ,
          <article-title>and F Cergy pontoise Cedex. Retin: A content-based image indexing and retrieval system</article-title>
          .
          <year>2001</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Carlos</surname>
            Rivero-Moreno and
            <given-names>Stphane</given-names>
          </string-name>
          <string-name>
            <surname>Bres</surname>
          </string-name>
          . Les filtres de Hermite et de Gabor donnent
          <article-title>-ils des modles quivalents du systme visuel humain?</article-title>
          <source>In ORASIS</source>
          <year>2003</year>
          ,
          <article-title>Journes francophones des jeunes chercheurs en vision par ordinateur</article-title>
          , pages
          <fpage>423</fpage>
          -
          <lpage>432</lpage>
          , May
          <year>2003</year>
          .
          <article-title>Article apparu en confrence nationale avec actes de</article-title>
          congrs et comit de lecture.
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <given-names>Sabrina</given-names>
            <surname>Tollari</surname>
          </string-name>
          .
          <article-title>Indexation et recherche d'images par fusion d'informations textuelles et visuelles</article-title>
          .
          <source>PhD thesis</source>
          ,
          <year>2006</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <article-title>Jean pierre Braquelaire and Luc Brun. Comparison and optimization of methods of color image quantization</article-title>
          . volume
          <volume>6</volume>
          , pages
          <fpage>1048</fpage>
          -
          <lpage>1051</lpage>
          ,
          <year>1997</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <given-names>Mimoun</given-names>
            <surname>Malki</surname>
          </string-name>
          , Madjid Ayache, and Mustapha Kamal Rahmouni. R´
          <article-title>etro-ing´enierie des bases de donn´ees relationnelles : approche bas´ee sur l'analyse de formulaires</article-title>
          .
          <source>In INFORSID</source>
          , pages
          <fpage>55</fpage>
          -
          <lpage>74</lpage>
          ,
          <year>1999</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <surname>Sid-Ahmed</surname>
            <given-names>Berrani</given-names>
          </string-name>
          , Laurent Amsaleg, and
          <string-name>
            <given-names>Patrick</given-names>
            <surname>Gros</surname>
          </string-name>
          .
          <article-title>Recherche par similarit´es dans les bases de donn´ees multidimensionnelles: panorama des techniques d'indexation</article-title>
          . volume
          <volume>7</volume>
          , pages
          <fpage>9</fpage>
          -
          <lpage>44</lpage>
          ,
          <year>2002</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <string-name>
            <surname>Vincent</surname>
            <given-names>Claveau</given-names>
          </string-name>
          , Romain Tavenard, and
          <string-name>
            <given-names>Laurent</given-names>
            <surname>Amsaleg</surname>
          </string-name>
          .
          <article-title>Vectorisation des processus d'appariement document-requˆete</article-title>
          .
          <source>In CORIA</source>
          , pages
          <fpage>313</fpage>
          -
          <lpage>324</lpage>
          ,
          <year>2010</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          12.
          <string-name>
            <given-names>Hanen</given-names>
            <surname>Karamti</surname>
          </string-name>
          .
          <article-title>Vectorisation du mod`ele d'appariement pour la recherche d'images par le contenu</article-title>
          .
          <source>In CORIA</source>
          , pages
          <fpage>335</fpage>
          -
          <lpage>340</lpage>
          ,
          <year>2013</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          13.
          <string-name>
            <given-names>Chris</given-names>
            <surname>Harris</surname>
          </string-name>
          and
          <string-name>
            <given-names>Mike</given-names>
            <surname>Stephens</surname>
          </string-name>
          .
          <article-title>A combined corner and edge detector</article-title>
          .
          <source>In Proc. of Fourth Alvey Vision Conference</source>
          , pages
          <fpage>147</fpage>
          -
          <lpage>151</lpage>
          ,
          <year>1988</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          14. F. Javier D´ıaz, Angel M. Bur´on, and Jos´e
          <string-name>
            <given-names>M.</given-names>
            <surname>Solana</surname>
          </string-name>
          .
          <article-title>Haar wavelet based processor scheme for image coding with low circuit complexity</article-title>
          . volume
          <volume>33</volume>
          , pages
          <fpage>109</fpage>
          -
          <lpage>126</lpage>
          , Tarrytown,
          <string-name>
            <surname>NY</surname>
          </string-name>
          , USA, March
          <year>2007</year>
          . Pergamon Press, Inc.
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          15.
          <string-name>
            <given-names>P.</given-names>
            <surname>Montesinos</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Gouet</surname>
          </string-name>
          , and
          <string-name>
            <given-names>R.</given-names>
            <surname>Deriche.</surname>
          </string-name>
          <article-title>Differential invariants for color images</article-title>
          .
          <source>In Pattern Recognition</source>
          ,
          <year>1998</year>
          . Proceedings. Fourteenth International Conference on, volume
          <volume>1</volume>
          , pages
          <fpage>838</fpage>
          -
          <lpage>840</lpage>
          vol.
          <volume>1</volume>
          ,
          <string-name>
            <surname>Aug</surname>
          </string-name>
          <year>1998</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          16.
          <string-name>
            <given-names>B.</given-names>
            <surname>Akrout</surname>
          </string-name>
          ,
          <string-name>
            <given-names>I. Khanfir</given-names>
            <surname>Kallel</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Benamar</surname>
          </string-name>
          , and
          <string-name>
            <given-names>B. Ben</given-names>
            <surname>Amor</surname>
          </string-name>
          .
          <article-title>A new scheme of signature extarction for iris authentication</article-title>
          .
          <source>In Systems, Signals and Devices</source>
          ,
          <year>2009</year>
          . SSD '
          <volume>09</volume>
          . 6th International Multi-Conference on, pages
          <fpage>1</fpage>
          -
          <lpage>8</lpage>
          ,
          <year>March 2009</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          17. Herv´e Go¨eau, Alexis Joly, Pierre Bonnet, Jean-Fran¸cois Molino, Daniel Barth´el´emy, and Nozha Boujemaa.
          <source>Lifeclef plant identification task</source>
          <year>2014</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>