<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Medical Image Classi cation via 2D color feature based Covariance Descriptors</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Pol Cirujeda</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Xavier Binefa</string-name>
          <email>xavier.binefag@upf.edu</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Department of Information and Communication Technologies Universitat Pompeu Fabra</institution>
          ,
          <addr-line>Barcelona</addr-line>
          ,
          <country country="ES">Spain</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>In these notes we present an image classi cation method which has been submitted to the ImageCLEF 2015 Medical Classi cation challenge. The aim is to classify images from 30 heterogeneous classes ranging from diagnose images coming from di erent acquisition techniques, to various biomedical publication illustrations. The presented work is intended to be a proof of concept of how our method, which uses only visual information, performs in the modelling of such image classes. Our approach uses 1st and 2nd order color features obtained at a whole image level. These features are considered as samples of a multidimensional statistical distribution, and a distinctive signature of the represented image can be built in the form of a Covariance-matrix based descriptor. The Riemannian manifold structure of such descriptors can be exploited in order to formulate an image classi cation methodology. Despite the challenging task due to unbalanced classes and image homogeneity, the obtained results in the task place our method on the top of the most accurate ones using purely visual features. This asserts the feasibility of our methodology and proves that its performance can be on par with other methods which use also complementary textual features for complex image retrieval.</p>
      </abstract>
      <kwd-group>
        <kwd>Covariance descriptor</kwd>
        <kwd>Medical image</kwd>
        <kwd>classi cation</kwd>
        <kwd>retrieval</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>
        Medical image classi cation provides a challenge on the identi cation of similar
medical images: this is an interesting problem due to the subtle changes between
di erent image sources. For instance, inside the range of microscopy images there
exist di erent acquisition devices (light, electron, uorescence or transmission)
which are able to capture di erent tissue details. Despite of that, the resemblance
between image cues is high and poses a challenging problem from a classi cation
perspective [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ].
      </p>
      <p>
        The ImageCLEF Medical Classi cation challenge [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ] provides a benchmark
to test the impact of di erent image classi cation and feature selection methods
in retrieval, specially those using visual and/or textual information. We are
presenting our results in the medical sub gure classi cation task, which provides 30
di erent classes including diagnose images (radiology, visible light photography,
microscopy, etc.) and also generic biomedical illustrations. More details can be
found in [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ].
      </p>
      <p>In the Computer Vision research area, many pattern recognition methods
have been developed for image classi cation and retrieval. Most of them
include the development of content and feature selection functions, or the usage
of keypoint extractors and associated descriptors which can be later categorized
by supervised classi cation methods (Support Vector Machines, Boosting,
Neural Networks, etc). Our presented approach is based in our ongoing research in
Covariance-based descriptors, and we are specially motivated by the demanding
conditions found in the di erent images of the medical classi cation subtask.
Our method provides a simplistic formulation, which provides a discriminative
signature for a whole image according to the variation of di erent features at
its pixels. We are particularly interested in seeing if this proposed description,
using purely visual information, is discriminative enough.</p>
      <p>The following sections of this report present an overview of our
methodology, the results obtained on the train data and the own challenge, and a nal
discussion about some aspects of the presented approach and associated future
work.
2</p>
    </sec>
    <sec id="sec-2">
      <title>Methodology</title>
      <p>An inspection of the provided images of this medical classi cation task makes
evident that class separation from purely visual cues is not a trivial task. Di erent
image sources might share visual features, or su er from a lack of
discriminative salient cues (see Fig. 1). Nevertheless, this also yields to our rst intuition
of what should be taken into account. First of all, there are several
information cues that are equally important: not only texture patterns, but also color,
sparsity, structure features... And in a second place, even more important than
the features themselves: the modelling must take into account all the feature
interactions together. That is, a diagram gure in a medical publication can be
in grayscale just as an electron microscopy image, but structural features in a
diagram contain pure lines or geometrical shapes which are not present on a
biological tissue captured by the microscope. At the same time, di erent
microscopy devices might capture similar natural tissue patterns, but for instance
a visible light microscope can capture a di erent range of color spectra than a
transmission microscope. Therefore, in an analogy with a natural visual
perceptual system, our goal is to model the space of di erent visual cues and their joint
relationships, and correlate them to the wide range of image classes.
2.1</p>
      <p>Visual features Covariance-based descriptor
An ideal image representation must encode all images in a common compact, size
invariant notation regardless of the di erent image sizes. The description must
DRCO
DSEC
GCHE
GMAT
DRCT
DSEE
GFIG
GNCP
DRMR
DSEM
GFLO
GPLI
DRPE
DVDM
GGEL
GSCR
DRUS
DVEN
GGEN
GSYS
DRXR
DVOR
GHDR
GTAB
also be robust to intraclass spatial transformations, such as rotations, and if
possible it should not depend on computationally loading intermediate stages, such
as keypoint extractions. The intuition behind our method is not to use image
features themselves, but rather than that observe the features along the
complete image and consider them as unstructured samples of a multidimensional
statistical distribution, using their covariance as a descriptive signature.</p>
      <p>
        Covariance matrices were introduced as descriptors in the Computer Vision
domain by Tuzel et al. in [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ] where they presented an object recognition method
for 2D color images. In our ongoing research we have extended this framework
to other domains such as 3D object recognition in unstructured point clouds [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ],
gesture recognition in depth image sequences [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ] or also tissue classi cation in
3D CT medical images [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ]. By their construction, covariance-based descriptors
are robust to noisy inputs and lose structural information about the observed
features. Their representation capability is based on the statistical notion of
covariance as a measure of how several random variables change together {a
set of visual cues for any image in our case. Therefore, the proposed descriptor
characterizes a given distribution of feature variations along the image, rather
than using feature absolute values, which is independent of the number of used
samples (the image size). This provides invariance to size and spatial rigid
transformations such as rotations.
      </p>
      <p>In order to formally de ne this 2D color feature based Covariance Descriptors,
we denote a feature selection function (I) for a given image I as:
(I) = f x;y 8x; y 2 Ig ;
(1)
which provides a set of feature vectors x;y for each one of the pixel coordinates
fx; yg inside all the image I. These 11-dimensional feature vectors are expressed
as:
and include the pixel coordinates, the di erent RGB color values, rst and
second order image intensity derivatives and their magnitude and pixel
curvature. These cues provide information about the color distribution of a given
image class, as well as their texture patterns and visual structure {as found in
the rst and second order gradient and curvature features. Then, for a given
color image I the associated Covariance Descriptor can be obtained as:</p>
      <p>Cov ( (I)) = N 1 1 Xi=N1 ( x;y ) ( x;y )T ; (3)
where is the vector mean of the set of vectors f g within the image I.</p>
      <p>The resulting 11 11 matrix Cov is a symmetric matrix where the diagonal
entries will represent the variance of each feature channel, and the non-diagonal
elements represent their pairwise covariance, as seen in Fig. 2. This provides a
signature of how feature behave in a characteristic way for each one of the images
of the di erent classes.
Covariance Descriptors have the form of covariance matrices which, besides
providing a compact and exible representation, causes them to lie in the
Riemannian manifold of symmetric de nite positive matrices Symd+. This has a major
impact on their interest as descriptive units, as their spatial variety is
geometrically meaningful: samples of classes sharing similar feature characteristics will
remain under close areas in this descriptor space. Nevertheless, it is important to
bear in mind that this spatial distribution is non Euclidean and has to be treated
with its particular Riemannian metric in order to perform analytic operations
with the descriptors.</p>
      <p>
        According to [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ], the Riemannian manifold can be approximated in close
neighborhoods by the Euclidean metric in its tangent space, TY , where the
symmetric matrix Y is a reference projection point in the manifold. TY is formed
by a vector space of d d symmetric matrices, and the tangent mapping of a
manifold element X to x 2 TY is made by the point-dependent logY operation:
1 1 1 1
x = logY (X) = Y 2 log Y 2 XY 2 Y 2 :
      </p>
      <p>For computational simplicity in certain problems, the projection point can be
established to the Identity matrix, and therefore the tangent mapping becomes:
(4)
(5)
(6)
log(X) = U log(D)U 0;
where U and D are the elements of the single value decomposition (SVD) of
X 2 Symd+.</p>
      <p>In an analogous manner, the exponential mapping of a point y 2 TY returns
its original point representation Y in the Symd+ manifold:</p>
      <p>exp(y) = U exp(D)U 0;</p>
      <p>One property of the projected symmetric matrices in the tangent space TY is
that they contain only d(d+1)=2 independent coe cients, in their upper or lower
triangular parts. Therefore it is possible to apply the vectorization operation in
order to obtain a linear orthonormal space for the independent coe cients:
x^ = vect(x) = (x1;1; x1;2; :::; x1;d; x2;2; x2;3; :::; xd;d);
(7)
where x is the mapping of X 2 Symd+ to the tangent space, resulting from
Eq. (4). The obtained vector x^ will lie in the Euclidean space Rm, where m =
d(d + 1)=2 {R66 in the current approach.</p>
      <p>This set of operations is useful for data visualization, feature selection, and
for developing Machine Learning and classi cation techniques on top of the
particular geometric space of the proposed Covariance Descriptors, specially taking
into account the following Riemannian metric which expresses the geodesic
distance between two points X1 and X2 on Symd+:
(X1; X2) =
s</p>
      <p>1 1 2
T race log X1 2 X2X1 2
;
(8)
or more simply (X1; X2) =</p>
      <p>1 1
values of X1 2 X2X1 2 .</p>
      <p>qPd</p>
      <p>i=1 log( i)2, where i are the positive
eigen2.3</p>
      <p>
        Classi cation via a Manifold-regularized sparse representation
For the classi cation of the proposed 2D color feature based Covariance
Descriptors we propose a manifold-based sparse classi cation method which is part of
our research as presented in previous approaches [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]. We intend to test the
performance of this approach in the heterogeneous class distribution found in the
ImageCLEF Medical classi cation task, and see if it is on par with other
textualbased methods eventually presented to the challenge by other participants.
      </p>
      <p>
        The topological layout of the proposed Covariance Descriptor yields to focus
on a geometrically sensitive classi cation method which can exploit the
Riemannian manifold distribution. Sparse representation based methods [
        <xref ref-type="bibr" rid="ref9">9, 10</xref>
        ] have
shown a recent rise in the Machine Learning community in the context of face
recognition. In this application, two key concepts are very relevant: sparsity and
collaborativeness. They are related to the complexity of the model learning: not
only because a complete set of learning samples is hardly available, but also
because an unknown element can share characteristics from di erent classes. As
this also the case in medical image retrieval, where images from a particular class
might be scarce and the low-level visual cues provide a complex class de nition,
we propose a new sparse method formulation adapted to the manifold of 2D
color based Covariance Descriptors.
      </p>
      <p>The base intuition is that an unknown sample should be ideally represented,
as accurate as possible, by using the smallest group of most similar samples from
a learning set A. Then, a test sample in the form of a new vectorized Covariance
descriptor C 2 R66 can be expressed as a linear combination on top of the
tangent space T of the Symd+ manifold of the available set of training samples:
C = A .</p>
      <p>Let A be the whole set of n training samples, in its vectorized form
according to eq. 7, from K di erent classes: A = [A1; A2; :::; AK ] 2 R66 n, where each
Ai = fvect(i)g is the set of vectorized Covariance descriptors which form the
subset of training samples for the class i. And let = [ 1; 2; :::; K ] be a vector
of weights corresponding to each one of the training samples in A. Then, the
sparsity restriction on can be achieved via its L2 norm minimization,
proposing a manifold-aware minimization constraint which relaxes the computational
expense of the method and adds numerical stability:
. . .</p>
      <p>0
where D is a diagonal matrix of size n n which allows the imposition of prior
knowledge on the solution with respect to the training set, using the Riemannian
metric de ned in eq. (8). This term contributes also on making the least squares
solution stable, and on introducing forehand sparsity conditions to the vector ^
as well. D is de ned as:
(9)
(10)
where A0i and C0 are the unvectorized covariance descriptors for training and test
samples respectively. The solution to the sparse collaborative representation, ^,
can be calculated by the following derived expression according to [10]:
^ =</p>
      <p>AT A + DT D
1 AT C</p>
      <p>Finally, the classi cation label of the test sample C can be obtained by
observing the regularized reconstruction residuals from the resulting sparse vector
^:
class(C) = argmin
i
kC</p>
      <p>Ai ^ik2
k ^ik2
(11)
(12)
3</p>
    </sec>
    <sec id="sec-3">
      <title>Results</title>
      <p>The evaluation score used on the task performance assessment is the classi cation
accuracy ratio for all the classes, computed as the ratio of true positives and
negatives over the total number of samples. We collect the top results in Table
3, which are also publicly available on the challenge website 1.</p>
      <sec id="sec-3-1">
        <title>Method</title>
      </sec>
      <sec id="sec-3-2">
        <title>Features</title>
      </sec>
      <sec id="sec-3-3">
        <title>True positive ratio</title>
        <sec id="sec-3-3-1">
          <title>Participants 1 Visual + text</title>
        </sec>
        <sec id="sec-3-3-2">
          <title>Participants 1 Only visual</title>
        </sec>
        <sec id="sec-3-3-3">
          <title>Our method</title>
        </sec>
        <sec id="sec-3-3-4">
          <title>Only visual</title>
        </sec>
        <sec id="sec-3-3-5">
          <title>Participants 3 Only visual</title>
          <p>Before the submission of the task, we tested our method on the training
data set, using a 10-fold cross-validation. Each fold was adapted so at least
20% of samples of each class were kept in each subset. In classes with a very
low number of samples which would cause to have some folds without class
representation, some samples where duplicated. Therefore, classes with very few
samples where guaranteed to be balanced and represented on the training set of
our classi cation method. After iterating the cross-validation runs, we obtained
an average accuracy of 73.24 %. As we have commented in section 2.3, the
presented classi er arises as a method for expressing unknown samples as the
best sparse representation regarding to a learning set. Therefore, we explain this
1 http://www.imageclef.org/2015/medical
increase on the accuracy as a direct e ect of the balancing preprocessing of those
classes with very few elements.</p>
          <p>Once the groundtruth annotations of the testing set have been made publicly
available, we can analyse the di erent Precision and Recall values for each class
as presented in Table 2, and observe if there is a particular correlation between
these values and the di erent cardinality of each class or their visual nature.</p>
        </sec>
      </sec>
      <sec id="sec-3-4">
        <title>Class Class # Precision Recall</title>
        <p>Class
Class #
Precision
Recall
Class
Class #
Precision
Recall</p>
        <p>These results assert our hypothesis of a mandatory class balancing stage in
order to boost the accuracy performance of our proposed sparse classi er.
4</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>Conclusions and future work</title>
      <p>The presented approach provides two main outcomes: on one side, a
Covariancebased descriptor which uses only low-level visual features and requires very low
computational cost for its construction. On the other side, a classi cation method
which takes into account the geometric properties of such representation. All
together, the system provides an image retrieval method which is fast and has
demonstrated to be of similar accuracy levels to other methods using
complementary textual information.</p>
      <p>Despite of that, we rmly believe that this method can be further extended in
the future, in many directions. Descriptor features could be extended with a
codi cation of medical terms associated to di erent image classes. Thus, visual and
textual feature fusion would take place within the nature of our descriptor. On
the other side, after analysing the results and the available groundtruth
annotations, we have observed a major dependency of our method on class cardinality
due to its sparse representation formulation. Classes with minor representation
can lead to higher classi cation error as a consequence of the minimization
formulation of our method. Therefore, we have observed that this can be solved by
incorporating a class balancing stage before the sparse regularization.</p>
      <p>So far, the participation on the ImageCLEF Medical Classi cation task has
provided an interesting benchmark which has contributed to test our ongoing
research and identify some improvements for our methodology thanks to the
particular nature of the provided testing data.
10. D Zhang, Meng Yang, and Xiangchu Feng. Sparse representation or
collaborative representation: Which helps face recognition? In International Conference on</p>
      <sec id="sec-4-1">
        <title>Computer Vision, pages 471{478, 2011.</title>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <given-names>Vincent</given-names>
            <surname>Arsigny</surname>
          </string-name>
          , Pierre Fillard, Xavier Pennec, and Nicholas Ayache.
          <article-title>Logeuclidean metrics for fast and simple calculus on di usion tensors</article-title>
          .
          <source>Magnetic resonance in medicine</source>
          ,
          <volume>56</volume>
          (
          <issue>2</issue>
          ):
          <volume>411</volume>
          {
          <fpage>421</fpage>
          ,
          <year>2006</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <given-names>Pol</given-names>
            <surname>Cirujeda</surname>
          </string-name>
          and
          <string-name>
            <given-names>Xavier</given-names>
            <surname>Binefa</surname>
          </string-name>
          . 4DCov:
          <article-title>A nested covariance descriptor of spatiotemporal features for gesture recognition in depth sequences</article-title>
          .
          <source>In International Conference on 3D vision (3DV)</source>
          ,
          <year>2014</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <given-names>Pol</given-names>
            <surname>Cirujeda</surname>
          </string-name>
          , Yashin Dicente Cid, Xavier Mateo, and
          <string-name>
            <given-names>Xavier</given-names>
            <surname>Binefa</surname>
          </string-name>
          .
          <article-title>A 3d scene registration method via covariance descriptors and an evolutionary stable strategy game theory solver</article-title>
          .
          <source>International Journal of Computer Vision</source>
          , pages
          <volume>1</volume>
          {
          <fpage>24</fpage>
          ,
          <year>2015</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <given-names>Pol</given-names>
            <surname>Cirujeda</surname>
          </string-name>
          , Henning Muller, Daniel Rubin, Todd A.
          <string-name>
            <surname>Aguilera</surname>
          </string-name>
          , Billy W. Loo Jr.,
          <string-name>
            <surname>Maximilian</surname>
            <given-names>Diehn</given-names>
          </string-name>
          , Xavier Binefa, and
          <string-name>
            <given-names>Adrien</given-names>
            <surname>Depeursinge</surname>
          </string-name>
          .
          <article-title>3d riesz-wavelet based covariance descriptors for texture classi cation of lung nodule tissue in ct</article-title>
          .
          <source>In International Conference of the IEEE Engineering in Medicine and Biology Society</source>
          (to appear),
          <year>2015</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <given-names>Alba</given-names>
            <surname>Garc</surname>
          </string-name>
          a Seco de Herrera,
          <article-title>Henning Muller, and Stefano Bromuri. Overview of the ImageCLEF 2015 medical classi cation task</article-title>
          .
          <source>In Working Notes of CLEF</source>
          <year>2015</year>
          (
          <article-title>Cross Language Evaluation Forum)</article-title>
          ,
          <source>CEUR Workshop Proceedings</source>
          . CEURWS.org,
          <year>September 2015</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6. Henning Muller, Nicolas Michoux, David Bandon,
          <string-name>
            <given-names>and Antoine</given-names>
            <surname>Geissbuhler</surname>
          </string-name>
          .
          <article-title>A review of content-based image retrieval systems in medical applications-clinical benets and future directions</article-title>
          .
          <source>International Journal of Medical Informatics</source>
          ,
          <volume>73</volume>
          (
          <issue>1</issue>
          ):1{
          <fpage>23</fpage>
          ,
          <year>2004</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <given-names>Onzel</given-names>
            <surname>Tuzel</surname>
          </string-name>
          , Fatih Porikli, and
          <string-name>
            <given-names>Peter</given-names>
            <surname>Meer</surname>
          </string-name>
          .
          <article-title>Region covariance: A fast descriptor for detection and classi cation</article-title>
          .
          <source>European Conference on Computer Vision</source>
          , pages
          <volume>589</volume>
          {
          <fpage>600</fpage>
          ,
          <year>2006</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <given-names>Mauricio</given-names>
            <surname>Villegas</surname>
          </string-name>
          , Henning Muller, Andrew Gilbert, Luca Piras, Josiah Wang, Krystian Mikolajczyk, Alba Garc a Seco de Herrera, Stefano Bromuri,
          <string-name>
            <given-names>M. Ashraful</given-names>
            <surname>Amin</surname>
          </string-name>
          , Mahmood Kazi Mohammed, Burak Acar, Suzan Uskudarli, Neda B.
          <string-name>
            <surname>Marvasti</surname>
          </string-name>
          , Jose F. Aldana, and
          <article-title>Mar a del Mar Roldan Garc a</article-title>
          .
          <source>General Overview of ImageCLEF at the CLEF 2015 Labs. Lecture Notes in Computer Science</source>
          . Springer International Publishing,
          <year>2015</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9. John Wright, Yi Ma, Julien Mairal, Guillermo Sapiro, Thomas S Huang, and
          <string-name>
            <given-names>Shuicheng</given-names>
            <surname>Yan</surname>
          </string-name>
          .
          <article-title>Sparse representation for computer vision and pattern recognition</article-title>
          .
          <source>Proc. of the IEEE</source>
          ,
          <volume>98</volume>
          (
          <issue>6</issue>
          ):
          <volume>1031</volume>
          {
          <fpage>1044</fpage>
          ,
          <year>2010</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>