<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>IBM T.J. Watson Research Center, Multimedia Analytics: Modality Classi cation and Case-Based Retrieval tasks of ImageCLEF2012</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Liangliang Cao</string-name>
          <email>liangliang.cao@us.ibm.com</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Yuan-Chi Chang</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Noel Codella</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Michele Merler</string-name>
          <email>mimerler@us.ibm.com</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Quoc-Bao Nguyen</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>John R. Smith</string-name>
          <email>jsmith@us.ibm.com</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>19 Skyline Dr. Hawthorne NY 10532</institution>
          ,
          <country country="US">USA</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>In this paper we present the modeling strategies that were applied by the IBM T.J. Watson research team to the modality classi cation and case-based retrieval tasks of ImageCLEF 2012. The primary challenges of this year's medical modality classi cation task were as follows: 1) the supplied training data was extremely limited, with some categories having as few as 5 positive examples, leaving little room for internal testing, and 2) some modalities appeared to be visually similar. In order to address these challenges, we approached the task from two fronts: 1) we attempted to augment the training data with additional examples of each category, and 2) we experimented with a broad range of modeling strategies and feature extraction techniques. For the case based retrieval task, we employed a semantic similarity approach to measure the relatedness among medical concepts found in the text corpus. We believe the lack of using additional lexical database besides the UMLS-methathesaurus led to poor performance in relation to other approaches.</p>
      </abstract>
      <kwd-group>
        <kwd>SVM</kwd>
        <kwd>Multiclass</kwd>
        <kwd>Kernel Approximation</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>The ImageCLEF 2012 Medical Modality Classi cation Task is a standardized
benchmark for systems to automatically classify medical image modality from
PubMed journal articles. The 2012 dataset has changed from the previous year
in 3 signi cant ways: 1) there are more categories, 2) the number of training
examples is far fewer, and 3) some modalities are more similar.</p>
      <p>Our approach can be described as scaling up to utilize as many features and
data as possible. Our experiments demonstrate that increasing either axes tends
to boost performance. In addition, we present a method for kernel approximation
to help address the computational time costs of using a wide variety of methods.</p>
      <p>
        For data augmentation, we drew from several sources outside the
ImageCLEF2012 collection, such as a Bing web-crawl for each category, as well as
publicly available medical image datasets, such as The Cancer Imaging Archives
(TCIA), and Image Retrieval in Medical Applications (IRMA). For our
modeling approaches, we selected multiple features extracted from a set of image
granularities, such as SIFT variants [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ], GIST [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ], Local Binary Patterns (LBP)
[
        <xref ref-type="bibr" rid="ref3">3</xref>
        ], edge and color histograms, and Curvelets [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ]. In addition, we experimented
with a variety of feature fusion and learning approaches, including early, late,
and kernel fusion, kernel approximation, multiclass SVM, and one-vs-all. We
discovered that multiclass SVM using an early fusion of many features with an
augmented dataset yielded the best performance. Kernel approximation
methods were able to signi cantly increase the e ciency of modeling at a small cost
to performance.
      </p>
      <p>Modality Classi cation Task
Our experiments are performed on the following datasets:</p>
      <p>Dataset 1: The original ImageCLEF2012 training dataset.</p>
      <p>
        Dataset 2: A dataset augmented with up to 100 additional examples per
category. Augmenting data was collected from Bing Image Search, a
Cornell University Vision and Image Analysis Group &amp; International Early
Lung Cancer Action Program (VIA/I-ELCAP) Public CT datase1, The
Cancer Imaging Archives (TCIA)2, Image Retrieval in Medical Applications
(IRMA)3, and the Japanese Society of Radological Technology (JSRT) [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ].
1 http://www.via.cornell.edu/lungdb.html
2 Image data used in this research were obtained from The Cancer Imaging Archive
(http://cancerimagingarchive.net/) sponsored by the Cancer Imaging Program,
DCTD/NCI/NIH.
3 courtesy of TM Deserno, Dept. of Medical Informatics, RWTH Aachen, Germany
      </p>
      <p>Due to the low number of positive examples in some categories in the original
training dataset, we chose to construct an additional augmented dataset. The
number of examples per category for both is shown in Fig. 1.
2.2</p>
      <p>Feature Collections
All our experiments are run using 1 of 7 sets of low-level visual features. Here
we name and describe the contents of 7 sets of features that were used for
modeling. Each feature is described by name, with the granularity of extraction
in parenthesis. The granularities are as follows:</p>
      <p>Global: Feature extracted from entire image. Native feature dimensionality
preserved.</p>
      <p>Grid: 5x5 image grid, with feature vector extracted from each grid block
and concatenated. Increases dimensionality by factor of 25.</p>
      <p>Grid7: 7x7 image grid, with feature vector extracted from each grid block
and concatenated. Increases dimensionality by factor of 49.</p>
      <p>Layout: 5 image regions including the center and the 4 quarters. Increases
dimensionaltiy by a facor of 5.</p>
      <p>Pyramid: Spatial pyramid, with global as rst level, and 2x2 image grid as
second level. Increases dimensionaltiy by a facor of 5.</p>
      <p>The feature sets referenced later in the section are as follows:
Feature Set 1 Color Correlogram (grid), Edge Histogram (grid), Image
Type (grid), LBP histogram (grid7), Color SIFT AM Codebook Size 2000
(pyramid), SIFT AM Codebook Size 2000 (pyramid), HSV SIFT AM
Codebook Size 1000 (pyramid), Image Stats (grid), Gist (layout), Curvelet
Texture (layout).</p>
      <p>Feature Set 2 Feature Set 1, removing SIFT features, and adding the
following: Color Histogram (grid), Color Moments (grid), Dominant Colors
(global), Thumbnail Vector (global), Color SIFT AM Codebook Size 1000
(pyramid), SIFT AM Codebook Size 1000 (pyramid),
FourierOrientationVector (grid). FourierOrientationVector is a feature representing the average
of diameters in Fourier-Mellin space, across varying angles from 0 to 180
degrees.</p>
      <p>Feature Set 3: Feature Set 2, adding FourierPolarPyramid (layout).
FourierPolarPyramid is a pyramid constructed in polar coordinates of
FourierMellin space. 4 radial levels are employed (partitions of size 1, 2, 4, and 8),
with 6 angular levels, across 4 color channels (RGB and Grayscale).
Feature Set 4: Color Correlogram (grid), Color Histogram (grid), Edge
Histogram (layout), Edge Histogram (grid), Gist (layout), FourierPolarPyramid
(global), FourierPolarPyramid (layout), Image Type (grid), LBP Histogram
(grid7), SIFT Codebook Size 1000 (global),</p>
      <p>Feature Set 5: Feature Set 2, minus all SIFT variants.
Feature Set 6: Image Stats (grid), LBP Histogram (grid), Image Type
(grid), Edge Histogram (grid), Shape Moments (grid), Dominant Colors
(grid), image stats (global), LBP histogram (global), Image Type (global),
Shape Moments (global), Edge Histogram (global), Color Wavelet (global),
Dominant Colors (global), Color Moments (global), Color Correlogram (global),
Color Moments (global), Gist (global), Wavelet Texture (global), Tamura
Texture (grid).</p>
      <p>Feature Set 7: SIFT Codebook Size 1000 (pyramid), HSV SIFT AM
Codebook Size 1000 (pyramid).
2.3</p>
      <p>
        Multiclass SVM
We employed the LibSVM library [
        <xref ref-type="bibr" rid="ref17">17</xref>
        ] to perform Multiclass SVM classi cation.
The process involves learning a set of 1-vs-1 classi ers, one for each pair of classes
in the dataset. In order to classify a new example, each 1-vs-1 model is evaluated
on it and the most likely label is selected based on a majority voting scheme. No
data sampling was performed: the Multiclass SVM model was learned directly
from the whole training set (either original and augmented). As such, one of the
advantages of this learning strategy consists in explicitly modeling the priors of
the classes in the dataset. All the parameters of the models chosen for the nal
submissions to ImageCLEF 2012 were chosen according to 5-fold (for the original
dataset) and 3-fold (for the augmented one) cross validation performances. After
experimenting on the internal cross-validation splits, a set of the best performing
descriptors was selected for fusion. For the original dataset, Feature Set 1, as
described in Section 2.2, was chosen. Feature Sect 2 was instead adopted for the
extended dataset. Furthermore, the Chi-square kernel was selected (in preference
over linear, RBF and histogram intersection) for all the Multiclass SVM runs,
computed as
      </p>
      <p>K(x; y) = 1
yd)2</p>
      <p>;
X (xd
d 12 (xd + yd)
(1)</p>
      <p>Three types of feature fusion methods were experimented: early, late, and
kernel fusion.</p>
      <p>Early fusion: consists of a concatenation of di erent descriptors, before
SVM modeling
Kernel Fusion: consists of a point-wise pooling (max or average operator)
over the kernel matrices produced by each descriptor. The Multiclass SVM
is then learned on top of the aggregate matrix
Late Fusion: consists of a pooling (max, average or product) operator over
the predictions of the models learned from individual features for each test
image. For this type of fusion we employed the probabilistic output option
in each SVM, which converts the 1-vs-1 comparisons into class probabilities.
For each test image, each model produced a vector with N probabilities
(where N is the number of classes, 31 in our case). After the pooling was
applied in a point-wise manner over the prediction vectors of the models,
the class with the maximum aggregate probability was chosen as the nal
prediction.</p>
      <p>Each strategy is exempli ed in Figure 2. As reported in Section 2.6, the
Multiclass SVM trained from the augmented dataset with early fusion strategy
(Experiment 12) provided the best performance. Kernel fusion proved to be
equivalent, while late fusion performed worse than the other methods.
To explore di erent aspects of visual phenomenon, we employed 19 di erent
features, described in Feature Set 5. We used an early fusion strategy by
concatenating these features together and training a kernelized Supporting Vector
Machine (SVM). However, a practical problem of using so many features lies
in the computational cost, both in the training and testing stage. When the
number of images grows, or when the feature dimension increases, traditional
SVM solvers may not work well or take a very long time to compute the optimal
solution.</p>
      <p>Among all the kernels in practice, the Chi-square kernel often yields very
good performance compared with the others. Moreover, a large amount of our
features, including LBP histogram, edge histogram, color histogram, and SIFT
histogram, are in the form of histogram features. Chi-square kernel is arguable
regarded as the rst choice for histogram form features. In our work, we focus on
how to e ciently solve Chi-square kernel only. We do not consider the problem
of general kernels.
We consider the Chi-square kernel in the form of</p>
      <p>K(x; y) = X</p>
    </sec>
    <sec id="sec-2">
      <title>2xdyd ;</title>
      <p>d xd + yd
where x = [x1; x2; ; xd; , y = [y1; y2; ; yd; .</p>
      <p>
        It is easy to see that Eq.(2) is de ned as the additive sum of di erent
dimensions. Such a kernel is referred to as an additive kernel. As suggested by [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ],
such a group of kernels can be approximated by mapping the feature into a high
dimensional space. By the representer theorem [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ], the solution of classi cation
model can be written in the form of
      </p>
      <p>N
f (x) = X K(x; xi);</p>
      <p>i=1
where i denotes the index of training samples. For any positive de nite kernel,
there exists a mapping x ! (x) so that the nal classi cation model becomes
f (x) = wT (x) + b
where w denotes the weights of the linear model in the mapped space. In this
work, we will use Nystrom's approximation to construct the mapping function
explicitly.</p>
      <p>To make the representation simply we let
then we can see the kernel is
k(x; y) =</p>
    </sec>
    <sec id="sec-3">
      <title>2xdyd ;</title>
      <p>xd + yd
K(x; y) = X k(xd; yd):
d
Next we will discuss how to approximate k(x; y), which is a function on 1D space.</p>
      <p>To approximate k(x; y), we employ
8
p 0</p>
      <p>if j = 0
j(x) = &lt;&gt;&gt; q2 j+1 cos( j+21 Lx) if j &gt; 0 odd</p>
      <p>
        2
&gt;:&gt; q2 2j sin( 2j Lx) if j &gt; 0 even
where j and j constitute one pair of Fourier transform, L is the frequency
parameter, and we use L = 2 =15 in practice. Then we can convert the nonlinear
kernels with the linear model over j(x). For more details, please refer to [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ].
      </p>
      <p>
        So far we have discussed how to approximate the kernelized SVM using a
linear model. Now we can compare the computational complexity of both
methods. Suppose we have m features, 1 m M , and each feature is of dimension
dm. The number of training samples is N , and size of testing set is T . For each
(2)
(3)
(4)
dm features, we map it to the space of dimension 7dm. Note that the evaluation
of kernel SVM depends on the number of support vectors, which in practice is
proportional to the number of training examples. Also our linear approximation
requires the extra cost of feature mapping, which is 7dN for training and 7dT
for testing. Table 1 compares the computational complexity of the two methods.
It is easy to see our linear approximation is much more e cient in both training
and testing stage. Our linear approximation is even plausible for the scenario
with a lot of features.
To measure the e ectiveness of our kernel approximation method, we
compare how much di erence exists between Chi-square kernels and our
approximated kernels. Table 2 illustrates the speed up and percentage of kernel
approximation error using randomly-generated features. It is easy to see the
approximation error is low, while the speed up will be increasingly signi cant when the
number of training samples grows.
It is also interesting to see the classi cation accuracy after our kernel
approximation. We implement the Chi-square kernel with LibSVM, and also use
liblinear [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ] to train a linear classi er using our projected features. As Figure 3
shows, the linear model based on the high dimensional approximation is not
necessarily worse than the original model. In fact in some categories, the
liblinear model even works slightly better. Note that we do not tune the parameters
of both models in this toy experiments. In the future, we plan to improve our
model with heterogeneous kernel learning methods [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ].
2.5
      </p>
      <p>
        One-vs-All Ensemble SVM with Data Sampling
IBM Multimedia Analytics and Retrieval System (IMARS) [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ] has been
developed and applied previously for semantic classi cation of unstructured images
and videos. The general framework is depicted in Fig. 4, and is primarily based
on generating collections of 1-vs-All SVM classi ers.
      </p>
      <p>Using the IMARS learning paradigm, we explored several variants of the
pipeline to better understand the contribution of each to system performance:
Feature Fusion: Early (vector concatenation) or Late (Unit Model score
averaging).</p>
      <p>Data Sampling: Random subsamples of negative examples, or using the
entire dataset.</p>
      <p>Number of Bags: Changing the number of Unit Models that are trained
for a particular feature by choosing di erent subsamplings of example data.
Feature Sets: Varying sets of features used for modeling.</p>
      <p>Kernel Type: A single RBF kernel with kernel parameter -4, C=100, and
sigmoid feature normalization was used for most model learning; however,
we performed one experiment with a Chi2 kernel, similar parameters, to
understand any e ect the variation of kernels might have.</p>
      <p>1-vs-All scores of multiple Ensemble Models were converted to Multiclass
labels by choosing the max over all the classi er scores.
One-vs-All Ensemble SVM with Data Sampling 1-vs-ALL experiments
are depicted in Table 3. Experiment index number, feature fusion type (Early
or Late), number of examples for each of positive and negative categories per
bag, number of bags across data, number of bags across features, the feature set
used, the average dimensionality across feature bags, the proportion of data used
for Validation Split (VS), multiclass accuracy scores, and relative performance
scaled by computational complexity (P/C) are shown. Computational
complexity is computed simply as the running time required to compute a kernel matrix.</p>
      <p>Experiment 8 was reproduced using a Chi2 kernel, yielding MACC of 64.4%,
and was not listed in this table.</p>
      <p>In the rst experiment, we established a baseline with a very small data
sampling rate (100 examples in each of positive and negative) and only two
SIFT features utilized. In the second experiment, we examined the performance
of our other features, excluding SIFT. In the third, we used both sets, and
in the fourth we added the newer FourierPolarPyramid feature vector. In the
fth experiment, we examined the e ect of using late fusion, instead of early
fusion. In the sixth experiment, we began to study the e ects of increasing
the number of examplars. In the seventh and eigth experiments, we examined
the e ect of increasing data sampling either by using larger data bag sizes,
or a larger number of bags. The ninth experiment represented our submission
NCFC ORIG 2 EXTERNAL SUBMIT.txt; this dataset used a restricted
number of features and late fusion, due to time constraints. In the last two
experiments, we nally examine the e ects of using larger amounts of data, up to the
limit of our dataset.</p>
      <p>In summary, our experiments yield a number of notable observations:
1. With early fusion, adding more features improves performance, SIFT
contributing the most (see Exp. 1-4).
2. Adding additional data improves performance (Exp. 4, 6-11) .
3. Early fusion of features outperforms late fusion with averaging, and provides
a better P/C ratio (Exp. 4-5).
4. Bagging along the data dimension reduces overall performance, but improves</p>
      <p>P/C (Exp. 7-8).
5. Best performance is acheived using early fusion of all features and all data
(Exp. 11).
6. Chi2 kernels and searching over SVM parameters holds the potential to boost
performance further.</p>
      <p>Multiclass SVM Multiclass SVM experiments are depicted in Table 4. Three
di erent aspects were explored in the experiments: fusion type, dataset size, and
aggregation type. From the results in the Table emerges that:
1. Early and kernel fusion are comparable, and perform better than late fusion
2. a larger training set leads to a better classi cation performance (Dataset 2
is better than Dataset 1)
3. Averaging seems to be the best aggregation strategy.
Exp. Fusion Aggregation Feature Set Dataset Mean ACC
12
13
14
15
16
17
18
19
20
21</p>
      <p>Early
Early
Kernel
Kernel</p>
      <p>Late
Late
Late
Late
Late
Late
Avg
Avg
Avg
Prod
Max
Avg
Prod
Max
2
1
2
1
2
2
2
1
1
1
2
1
2
1
2
2
2
1
1
1
In Figure 5 are reported the o cial runs submitted to the ImageCLEF 2012
website, in comparison to all other o cial submissions (43 in total). IBM achieved
the top three visual only classi cation performances, as well as the best overall
accuracy.</p>
      <p>The submissions were all purely visual, and corresponded to the following
experiments (in decreasing order of performance): 12, 13 with Kernel
approximation fusion (as described in Section 2.4), 13, 11, 12 with Kernel approximation
fusion, 18, 21. After the submissions we further improved the worst performing
ones through better normalization, parameter selection, and further debugging,
yielding to the performances reported in Tables 3 and 4.</p>
      <p>Fig. 6(b) shows the confusion matrix of the best performing run. Overall the
matrix presents a strong diagonal. However, some clear mis-classi cations are
evident. In particular, \GSYS - System overview" resulted to be the hardest class
to categorize, being quite often confused with \GFLO - Flowcharts". Looking at
the appearance of the images in such classes, it is evident that based on visual
features alone they are in most cases indistinguishable. This confusion might
be mitigated by exploiting the textual information associated with, or in, the
images. We plan to follow this direction in future experiments.</p>
      <p>We expect that extracting textual information and combining it with our
strong visual modeling will boost classi cation performance, given the
complementary information of those two representations, and also looking at the
performances of other groups.</p>
      <p>In conclusion, appropriately modeling the visual appearance can provide
strong modality classi cation performance, even without text analysis. In our
experiments we found that adding more features and training data lead to
models with better classi cation performance. Early fusion and kernel fusion seem to
be the best combination strategies. Our kernel approximation provides a
principled and e cient framework to perform such fusion, while signi cantly increasing
the e ciency of the computation of the best performing CHI2 kernel.
In this section, we give an overview of the application of our methods to
casebased medical image retrieval and present the results of our submitted runs.
Techniques used for case-based retrieval task are mainly based on methods from
Information Retrieval (IR) and Natural Language processing (NLP), including
rule-based and machine-learning techniques. In order to facilitate the classi
cation of a large dataset and to retrieve the most relevant documents for a given
query, an IR system applies various NLP methods to construct a semantic view
of each document indexed via relational database or text index. This
semantic view is summarized by a set of relevant keywords (i.e., index terms) as a
signature of this document.</p>
      <p>In the biomedical domain, the terminology is very important because the
words used in the document are related to medical terms that can refer to the
same concept with di erent semantic interpretation (i.e., senses) based on the
textual context. In addition to the NLP techniques for reducing the size of the
relevant keywords by eliminating stopwords and stemming words, our system
also applies semantic similarity methods to improve the understanding of textual
terms and remedy potential ambiguity among medical concepts. The focus of
the semantic similarity is to nd the strength of the semantic relatedness or the
semantic connections between textual terms. The taxonomic proximity between
terms measures the degree of overlapping between contextual word vectors using
Information Content (IC) based measures. To reduce computational complexity,
the semantic relatedness is applied within an ordered window to nd relevant
terms between adjacent terms in the document.
To build our retrieval framework based on the approaches we described above,
we use the YTEX (Yale cTAKES extensions) system for computing the
semantic IC-based measures and the cTAKES (clinical Text Analysis and Knowledge
Extraction) as a NLP system based on the Unstructured Information
Management Architecture (UIMA) that combines rule-based and machine-learning. The
cTAKES system uses the OpenNLP Maximum Entropy package for sentence
detection, tokenization (words), part-of-speech (POS) tagging. It performs the
name entity recognition of biomedical from the Uni ed Medical Language
system (UMLS) Metathesaurus, and other biomedical source such as Systematized
Nomenclature of Medicine, Clinical Terms (SNOMED CT).</p>
      <p>To classify the medical articles, our system used the YTEX semantic
similarity to identify and disambiguate medical terms before storing the annotations on
the documents in the relational database, and indexing relevant medical concepts
that can facilitate the topic case query matching. For each case-based query, we
rst executed the NLP pipeline to extract relevant medical concepts and
generate an SQL expression by combining the concepts using logical OR operator
(meaning all of the concepts to be optional). To limit the number of query
results and select the most relevant annotations, we de ne the weight measures for
sorting and ranking them based on the cosine distance between medical concept
vectors.
Fig. 7 describes the processing ow of the system. The data processing splits the
whole article le into multiple individual les in order to distribute them across
multiple nodes. Each article is stored in the database as an annotation
including its most relevant medical concepts. The case-based NLP pipeline processes
each case-based le to nd relevant medical concepts. Finally, the search engine
composes a SQL logical expression and ranks the result set retrieved from the
annotation database.</p>
      <p>Due to time constraint, we only submitted one run for case-based retrieval
task, shown in Table 5. First, we experimented with the semantic similarity
approach, and found good correlation among medical concepts within the text
corpus with appropriate senses. However, by matching only the medical concepts,
the results are not as good as could be if we had used additional lexical databases.
3.4</p>
      <p>Conclusion
The IC-based measures applied to adjacent words in a de ned text window
improves the semantic relatedness performance and is less expensive than
computing all the word-pairs in the corpus. However, we also observed low system
performance when the SQL expressions are complex and the number of concepts
is high. In the future, we would like to improve the semantic similarity methods
by incorporating the medical concepts with the lexical semantic analysis. We
also want to have a better matching measures to improve the accuracy.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>D.</surname>
          </string-name>
          <article-title>Lowe: Distinctive image features from scale-invariant keypoints</article-title>
          .
          <source>International Journal of Computer Vision</source>
          ,
          <volume>60</volume>
          ,
          <issue>2</issue>
          , pages
          <fpage>91110</fpage>
          , (
          <year>2004</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <given-names>Aude</given-names>
            <surname>Oliva</surname>
          </string-name>
          and Antonio Torralba:
          <article-title>Modeling the shape of the scene: A holistic representation of the spatial envelope</article-title>
          .
          <source>Int. J. Comput. Vision</source>
          ,
          <volume>42</volume>
          (
          <issue>3</issue>
          ):145
          <fpage>175</fpage>
          , (
          <year>2001</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <given-names>T.</given-names>
            <surname>Ahonen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Hadid</surname>
          </string-name>
          , and
          <string-name>
            <surname>M.</surname>
          </string-name>
          <article-title>Pietikainen: Face recognition with local binary patterns</article-title>
          .
          <source>ECCV</source>
          , pages
          <fpage>469</fpage>
          -
          <lpage>481</lpage>
          , (
          <year>2004</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Cands</surname>
          </string-name>
          , Emmanuel and Demanet, Laurent and Donoho, David and Ying, Lexing: Fast Discrete Curvelet Transforms.
          <source>Multiscale Modeling and Simulation</source>
          ,
          <volume>5</volume>
          (
          <issue>3</issue>
          ). pp.
          <fpage>861</fpage>
          -
          <lpage>899</lpage>
          , (
          <year>2006</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Shiraishi</surname>
            <given-names>J</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Katsuragawa</surname>
            <given-names>S</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ikezoe</surname>
            <given-names>J</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Matsumoto</surname>
            <given-names>T</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kobayashi</surname>
            <given-names>T</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Komatsu</surname>
            <given-names>K</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Matsui</surname>
            <given-names>M</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Fujita</surname>
            <given-names>H</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kodera</surname>
            <given-names>Y</given-names>
          </string-name>
          , and
          <article-title>Doi</article-title>
          K.:
          <article-title>Development of a digital image database for chest radiographs with and without a lung nodule: Receiver operating characteristic analysis of radiologists' detection of pulmonary nodules</article-title>
          .
          <source>AJR</source>
          <volume>174</volume>
          ;
          <fpage>71</fpage>
          -
          <lpage>74</lpage>
          ,
          <year>2000</year>
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Yan</surname>
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Fleury</surname>
            <given-names>M.O.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Merler</surname>
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Natsev</surname>
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Smith</surname>
            <given-names>J.R.</given-names>
          </string-name>
          <string-name>
            <surname>Large-Scale Multimedia Semantic Concept</surname>
          </string-name>
          <article-title>Modeling using Robust Subspace Bagging and MapReduce</article-title>
          . ACM Multimedia 2009
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>Smith</surname>
            ,
            <given-names>T.F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Waterman</surname>
            ,
            <given-names>M.S.:</given-names>
          </string-name>
          <article-title>Identi cation of Common Molecular Subsequences</article-title>
          .
          <source>J. Mol. Biol</source>
          .
          <volume>147</volume>
          ,
          <issue>195</issue>
          {
          <fpage>197</fpage>
          (
          <year>1981</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>May</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ehrlich</surname>
            ,
            <given-names>H.C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Steinke</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          :
          <article-title>ZIB Structure Prediction Pipeline: Composing a Complex Biological Work ow through Web Services</article-title>
          . In: Nagel,
          <string-name>
            <given-names>W.E.</given-names>
            ,
            <surname>Walter</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W.V.</given-names>
            ,
            <surname>Lehner</surname>
          </string-name>
          , W. (eds.) Euro-Par
          <year>2006</year>
          . LNCS, vol.
          <volume>4128</volume>
          , pp.
          <volume>1148</volume>
          {
          <fpage>1158</fpage>
          . Springer, Heidelberg (
          <year>2006</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <surname>Foster</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kesselman</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          :
          <article-title>The Grid: Blueprint for a New Computing Infrastructure</article-title>
          . Morgan Kaufmann, San Francisco (
          <year>1999</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <surname>Czajkowski</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Fitzgerald</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Foster</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kesselman</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          :
          <article-title>Grid Information Services for Distributed Resource Sharing</article-title>
          .
          <source>In: 10th IEEE International Symposium on High Performance Distributed Computing</source>
          , pp.
          <volume>181</volume>
          {
          <fpage>184</fpage>
          . IEEE Press, New York (
          <year>2001</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <string-name>
            <surname>Foster</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kesselman</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Nick</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Tuecke</surname>
            ,
            <given-names>S.:</given-names>
          </string-name>
          <article-title>The Physiology of the Grid: an Open Grid Services Architecture for Distributed Systems Integration</article-title>
          .
          <source>Technical report, Global Grid Forum</source>
          (
          <year>2002</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>12. National Center for Biotechnology Information, http://www.ncbi.nlm.nih.gov</mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          13.
          <string-name>
            <surname>Vedaldi</surname>
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zisserman</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          :
          <article-title>E cient Additive Kernels via Explicit Feature Maps</article-title>
          ,
          <source>IEEE Trans. Pattern Analysis and Machine Intelligence</source>
          ,
          <volume>34</volume>
          (
          <issue>3</issue>
          ),
          <fpage>2012</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          14.
          <string-name>
            <surname>Kimeldorf</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Wahba</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          :
          <article-title>Some results on tchebychefan spline functions</article-title>
          .
          <source>Journal of Mathematical Analysis and Applications</source>
          <volume>33</volume>
          (
          <year>1971</year>
          )
          <fpage>82</fpage>
          -
          <lpage>95</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          15.
          <string-name>
            <surname>Cao</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Luo</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Liang</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Huang</surname>
            ,
            <given-names>T. S.:</given-names>
          </string-name>
          <article-title>Heterogeneous Feature Machines for Visual Recognition IEEE Proc</article-title>
          .
          <string-name>
            <surname>Int'l Conf</surname>
          </string-name>
          .
          <source>Computer Vision</source>
          (ICCV),
          <year>2009</year>
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          16.
          <string-name>
            <surname>Ho</surname>
          </string-name>
          ,
          <string-name>
            <surname>C.-H. Lin</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          -J.:
          <article-title>Large-scale Linear Support Vector Regression</article-title>
          , http://www.csie.ntu.edu.tw/~cjlin/liblinear/
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          17.
          <string-name>
            <surname>Chang</surname>
          </string-name>
          ,
          <article-title>Chih-Chung and Lin, Chih-Jen: LIBSVM: A library for support vector machines</article-title>
          ,
          <source>ACM Transactions on Intelligent Systems and Technology</source>
          ,
          <volume>2</volume>
          (
          <issue>3</issue>
          ), pages
          <fpage>1</fpage>
          -
          <lpage>27</lpage>
          , (
          <year>2011</year>
          )
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>