<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>A Mission-Oriented Citizen Science Platform for Efficient Flower Classification Based on Combination of Feature Descriptors</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Andréa Britto Mattos</string-name>
          <email>abritto@br.ibm.com</email>
          <email>{abritto,rhermann,kshigeno}@br.ibm.com</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Copyright c by the paper's authors. Copying permitted only for private</string-name>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Rogerio Schmidt Feris</string-name>
          <email>rsferis@us.ibm.com</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>IBM TJ Watson Research</institution>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Ricardo Guimarães Herrmann, Kelly Kiyumi Shigeno, IBM Research -</institution>
          <country country="BR">Brazil</country>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>and academic purposes., In: S. Vrochidis, K. Karatzas, A. Karpinnen, A. Joly (eds.): Proceedings of, the International Workshop on Environmental Multimedia Retrieval (EMR</institution>
          ,
          <addr-line>2014), Glasgow, UK, April 1, 2014, published at http://ceur-ws.org</addr-line>
        </aff>
      </contrib-group>
      <fpage>45</fpage>
      <lpage>52</lpage>
      <abstract>
        <p>This paper describes a citizen science system for ora monitoring that employs a concept of missions, as well as an automatic approach for ower species classi cation. The proposed method is fast and suitable for use in mobile devices, as means to achieve and maintain high user engagement. Besides providing a web-based interface for visualization, the system allows the volunteers to use their smartphones as powerful sensors for collecting biodiversity data in a fast and easy way. The classi cation accuracy is increased by a preliminary segmentation step that requires simple user interaction, using a modi ed version of the GrabCut algorithm. The proposed classi cation method obtains good performance and accuracy, by combining traditional color and texture features together with carefully designed features, including a robust shape descriptor to capture ne morphological structures of the objects to be classi ed. A novel weighting technique assigns di erent costs to each feature, taking into account the inter-class and intra-class variation between the considered species. The method is tested on the popular Oxford Flower Dataset, containing 102 categories and we achieve state-of-theart accuracy while proposing a more e cient approach than previous methods described in the literature.</p>
      </abstract>
      <kwd-group>
        <kwd>Computer vision</kwd>
        <kwd>ne-grained classi cation</kwd>
        <kwd>ora classi cation</kwd>
        <kwd>citizen science</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Categories and Subject Descriptors</title>
    </sec>
    <sec id="sec-2">
      <title>1. INTRODUCTION</title>
      <p>Citizen science is not a new concept: the idea of
conducting research by citizens, gathering crowdsourced data that is
analyzed for scienti c purposes, has been active for
generations, dating back to 1900, when the Christmas Bird Count
(CBC)1 project started. It was a form of mass collaboration
citizen science project that used paper forms sent by post to
the responsible society or research group of scientists who
requested the data.</p>
      <p>Nevertheless, it is noticeable that the power of
crowdsourced science is growing intensely nowadays and such
projects are becoming increasingly popular, even gaining the
attention of major news media. For instance, projects such
as eBird2 and Galaxy Zoo3 are able to engage hundreds of
thousands of volunteers and are being broadly used for
scienti c and educational purposes.</p>
      <p>The citizen science's popularity boost is given by the fact
that data collection and analysis tasks became much
easier to address. Even simple and low cost smartphones are
equipped with GPS and high-resolution cameras that allow
huge amounts of data to be collected with small e ort. Also,
the Internet brings together a high number of volunteers to
work on this data remotely and simultaneously. However,
there are still di culties that prevent the further growth of
this methodology.</p>
      <p>
        One challenge of citizen science projects is how to obtain
high and constant engagement of the users. Although some
volunteers are motivated by the scienti c contribution on
its own, it is possible to resort to \gami cation", the use
of game elements in non-game contexts [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ], to get others
to further engage with the community and contribute more
enthusiastically [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ].
      </p>
      <p>
        Project Noah4, for instance, is a large scale project in
which every report is assigned to a speci c mission, whose
goal is to monitor certain types of ora or fauna classes in
a speci c location. There are also many aspects that can
increase user motivation, as observed in [
        <xref ref-type="bibr" rid="ref28">28</xref>
        ]. For instance,
the volunteers must have con dence that the collected data
is being used: therefore, they should have easy access to the
visualization of the collected data. Also, it is important to
notice that some users are not willing to modify their daily
activities to report data, and providing training and
mentoring for volunteers to increase their skills can be fundamental
for valid data registration.
      </p>
      <p>This last point highlights an additional drawback of
proje1http://birds.audubon.org/christmas-bird-count
2http://www.ebird.org/
3http://www.galaxyzoo.org/
4http://www.projectnoah.org/
cts monitoring biodiversity: relying on the users having the
knowledge to classify, without further assistance, the
specimens being reported. This can be a di cult task when one
considers the large number of species that share similar
morphology. Misclassi ed inputs can compromise the
environmental research, and the impossibility to assure the quality
of the analysis made by non-experts may be the bottleneck
of such projects. In this regard, an algorithm to assist the
user would be very useful. A system that could
automatically recognize an input image with small e ort, or at least
retrieve a list of the most similar species, could make the user
feel more con dent, perhaps even engaging a larger number
of volunteers, preventing errors and accelerating the
identi cation process. If such a system could be deployed in a
mobile device, the bene ts would be even greater, once the
classi cation could be executed at collection time.</p>
      <p>
        To assist non-expert classi cation, automatic works in the
ora eld { which is the domain of focus of this paper {
are getting good attention, and recent studies were able to
achieve good results in leaf and ower classi cation [
        <xref ref-type="bibr" rid="ref1 ref18 ref19 ref3">1, 3,
18, 19</xref>
        ]. This popularity is evident as we see e orts such as
the recent Plant Identi cation Task, promoted by the widely
popular ImageCLEF challenge [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ].
      </p>
      <p>
        Besides robustness, automatic or semi-automatic methods
can reduce data collection time. For instance, as indicated
by Zou and Nagy [
        <xref ref-type="bibr" rid="ref33">33</xref>
        ], for a dataset of 102 ower species,
the time for a semi-automatic classi cation is much lower
than that made by humans alone.
1.1
      </p>
    </sec>
    <sec id="sec-3">
      <title>Related Work</title>
      <p>1.1.1</p>
      <sec id="sec-3-1">
        <title>Citizen Science in the Flora Domain</title>
        <p>
          Regarding automatic classi cation, the LeafSnap project5
o ers a good di erential with respect to other citizen science
projects, since it runs a shape-based algorithm to classify
plant species from their leaves automatically [
          <xref ref-type="bibr" rid="ref15">15</xref>
          ]. However,
the user needs to extract (i.e., cut) the leaf and place it on a
white background for the segmentation algorithm to work.
        </p>
        <p>The segmentation method proposed by LeafSnap might
be considered harmful in an environmental aspect, once it
requires the extraction of the leaves from their natural
surroundings. Also, lighting variations can a ect the
background color and interfere in the segmentation result.
Finally, it is hard to extend LeafSnap for additional domains
other than leaf classi cation: besides taking shape
information as the only feature for classi cation, a
segmentation method based on a xed background is only feasible
for static species, once it is clear that, when photographing
specimens that can move, there is no guarantee that they
will remain in place.</p>
        <p>Table 1 shows a list of various citizen science projects in
the botanical eld. To the best of our knowledge, these are
the most representative projects in this area. Their
limitations highlights the previously discussed issues regarding the
di culty of species identi cation or adaptation of the
volunteers routine, once the data upload requires a computer
and/or Internet connection.</p>
        <p>Our goal is to develop a system similar to LeafSnap, but
focusing on owers, and replacing the user e ort for
segmentation by a simpler and less intrusive action taken in his
own device. Since our study is part of the citizen science
context, user interaction is accepted, once we assume there
is always a volunteer operating the system in his smartphone
or tablet. Our platform is developed with the concern of
be5http://leafsnap.com/
ing attractive to users so it can reach the largest number of
volunteers as possible. Therefore, our ower classi cation
algorithm must have high accuracy and also must be fast
enough so that it can run in the web browser and also in
mobile devices. Besides, we use missions to engage and
entertain the volunteers, following gami cation principles. In
order to compare our classi cation results with a benchmark,
we use the challenging Oxford Flower Dataset6.
1.1.2</p>
      </sec>
      <sec id="sec-3-2">
        <title>Flower Classification</title>
        <p>
          In the ower classi cation eld, previous studies worked
on small datasets of 10 to 17 ower species [
          <xref ref-type="bibr" rid="ref10 ref14 ref20 ref25">10, 14, 20, 25</xref>
          ],
achieving accuracy rates of 81% - 96%. Most of these works
rely on contour analysis and SIFT features [
          <xref ref-type="bibr" rid="ref17">17</xref>
          ], which is
known to be a robust descriptor, but computationally
inefcient.
        </p>
        <p>
          Using larger datasets, consisting of 30 to 79 ower species,
accuracy rates vary from 63% - 94%, combining a variety
of approaches such as histograms in HSV color space,
contour features, co-occurrence matrix (GLCM) for texture
features, RBF and probabilistic classi cation [
          <xref ref-type="bibr" rid="ref29 ref30 ref6">6, 29, 30</xref>
          ]. Qi
[
          <xref ref-type="bibr" rid="ref26">26</xref>
          ] focuses on a fast environmental classi cation for owers
and leaves, using spatial co-occurrence features and a linear
SVM, but accuracy is not reported.
        </p>
        <p>
          Works on the Oxford Flower Dataset, that contains 102
ower species, reported accuracy rates no greater than 80%.
The dataset was introduced by Nilsback and Zisserman[
          <xref ref-type="bibr" rid="ref21">21</xref>
          ]
and is used by many authors. It is considered an extremely
challenging dataset, because it contains pictures with severe
variations in illumination and viewpoint, and also, due to
small inter-class variations and large intra-class variations
between the considered ower species. See some examples
of such cases in Fig. 1.
        </p>
        <p>(a) Spear Thistle (left) and (b) English Marigold (left)
Artichoke (right). and Barbeton Daisy (right).
(c) Spring Crocus.</p>
        <p>(d) Bougainvillea.</p>
        <p>
          Using the Oxford Dataset, studies working with
unsegmented images demonstrate lower accuracy in classi cation
and require powerful classi ers, that are computationally
expensive [
          <xref ref-type="bibr" rid="ref11 ref2">2, 11</xref>
          ]. Khan's work [
          <xref ref-type="bibr" rid="ref13">13</xref>
          ] relies on color and shape,
using a SIFT based approach. Kanan [
          <xref ref-type="bibr" rid="ref12">12</xref>
          ] applies salience
maps to extract a high-dimensional image representation
that is used in a probabilistic classi er. Both works require
high computation time due to their feature extraction
methods.
6http://www.robots.ox.ac.uk/~vgg/data/flowers/
        </p>
        <p>
          Using segmented images, the authors are able to achieve
a more robust classi cation for this dataset, though none of
the considered works is suitable for real time. Nilsback [
          <xref ref-type="bibr" rid="ref21 ref22">21,
22</xref>
          ] applies a multi-kernel SVM classi er using four di
erent features (color in HSV space, HOG, and SIFT in
foreground and boundary regions). Chai [
          <xref ref-type="bibr" rid="ref5">5</xref>
          ] proposes a robust
approach for co-segmentation (segmenting images with
similar background color distributions) and applies a SVM
classi er based on color features in the Lab space and SIFT in
foreground region. Angelova [
          <xref ref-type="bibr" rid="ref1 ref16">1, 16</xref>
          ] proposes a
segmentation approach, followed by the extraction of HOG features
at four di erent levels, encoded using LLC. The encoded
features are subject to a max pooling and a SVM classi er
is used. They mention future e orts in making a real time
approach.
1.2
        </p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>Contributions and Organization</title>
      <p>Our main goal is to develop a citizen science application
for ower monitoring whose collection is oriented by
missions. Also, to assist the volunteers, we propose an algorithm
for classi cation of ower species that meets the following
requirements:</p>
      <p>
        (i) Has high accuracy: We propose a novel approach that
relies on e cient histogram-based feature descriptors that
capture both global properties and ne shape information of
objects. This is made while leveraging a learning-based
distance measure to properly weight the feature contributions,
which is a critical step for increasing accuracy as
demonstrated in previous studies [
        <xref ref-type="bibr" rid="ref31 ref32 ref4">4, 31, 32</xref>
        ].
      </p>
      <p>(ii) Is fast: Mobile-based applications have a challenge
regarding connectivity: users can only submit requests and
receive answers if their devices have access to the Internet.
There are several locations in which this kind of service is
not a ordable or has poor quality, making image
transference prohibitive. Our algorithm is fast enough to run
directly in the devices (without transferring information to a
server), so the user does not have to rely upon a network
connection and wait too long for a response, considering
additional time latency for uploading the image and retrieving
the classi cation results.</p>
      <p>This paper is structured as follows. In Section 2, we
describe the overall structure of our citizen science platform
and, in Section 3, we describe the core approach for
negrained classi cation. In Section 4, we show and discuss the
obtained classi cation results on the considered dataset and
our conclusions and future work are described in Section 5.</p>
    </sec>
    <sec id="sec-5">
      <title>2. SYSTEM</title>
      <p>The system is comprised of two main parts: a web portal,
organized toward data visualization and community
organization, and a mobile application, aimed at data collection.
Both components arrange user-contributed data around the
concept of missions, which aggregate users with the common
goal of collecting structured and unstructured data about a
certain class of observations.</p>
      <p>The collected data is structured in a set of attributes,
whose pre-established values are lled in by the user and
refer to the domain being registered. This structure allows us
to provide query-by-example features, which helps to avoid
common input errors and establish a common vocabulary for
user-contributed data. The unstructured data, on the other
hand, is comprised of images only, since we are dealing with
plants.
(a) Image upload
(b) Free-hand user marker (c) Classi cation of the
(in white) segmented input
(d) Other results
(e) Screen for geo-analysis</p>
      <p>Missions are created by users and can be set as public or
private, as determined by the mission's owner, since some
of them may contain sensitive geo-spatial data about
observations belonging to endangered species. As the number of
missions in the system may be rather large, the portal ranks
them according to their activity, to make it easier for users
to engage with the rest of the community.</p>
      <p>It is currently possible to collect geo-referenced photos
with consumer-grade mobile devices, using their cameras
and GPS receivers, as mentioned in Section 1. The steps
the user has to go through in order to submit a contribution
are highlighted in Fig. 2. For the purpose of submitting a
report, the user rst needs to join one mission. After the
picture is taken, he must loosely delineate the region of interest
in the image. This is used for a segmentation process that
will be described int the next Section. This step can require
user con rmation for proceeding with the next task, which
is the automatic classi cation of the segmented input. The
system retrieves the top 5 matches that are more similar to
the uploaded image and the user can select the correct one
or browse within a list of all the registered classes.</p>
      <p>In situations where there is no connectivity, the algorithms
can be executed directly on the devices, as a signi cant range
of current mobile processors have multiple cores and some
of them have vector processing instructions, which are
already supported by popular libraries for mobile platforms,
such as OpenCV. After the report is submitted (and
synchronized with the server, in situations of intermittent or no
connectivity), other users can view and bookmark this
report. The feedback is provided to the users as bookmarked
reports. Additionally, the user may ll in a form
containing values for the observable raw morphological attributes
of the sample. The set of supported attributes is de ned
by the mission's owner at the moment it is created, where
he can de ne the required elds according to the mission's
needs. For the botanic eld, some examples of attributes
are stem diameter, leaf thickness and so on.</p>
      <p>In order to obtain more accurate classi cation results, we
let the user choose the category of the image he is uploading.
We currently only have trained classi ers for owers, but the
overall concept of the system is generic for accepting data
belonging to other categories. Since observations can be
arbitrary, we need to load the corresponding classi er that
was trained with the relevant data for the current category.
3.</p>
    </sec>
    <sec id="sec-6">
      <title>ALGORITHM</title>
      <p>This section describes the proposed algorithm for ower
recognition that is divided in segmentation and classi cation
steps. Both tasks are designed to be fast and precise and
are generic enough to be used for classifying categories other
than owers.
3.1</p>
    </sec>
    <sec id="sec-7">
      <title>Segmentation</title>
      <p>
        We use the GrabCut algorithm [
        <xref ref-type="bibr" rid="ref27">27</xref>
        ] for a semi-manual
segmentation using small user interaction. This method is
also used in Qi's work [
        <xref ref-type="bibr" rid="ref26">26</xref>
        ] due to its good performance. In
our method, instead of de ning a bounding box and control
points, the user draws a free hand marker, which is more
intuitive for a general user. The marker replaces the control
points for a more re ned boundary in a faster interaction.
It must involve the whole object but does not have to be
precise.
      </p>
      <p>A Gaussian Mixture Model is used to learn the pixel
distribution in the background (outside the marker) and
possible foreground regions and a graph is built from this pixel
distribution. Edges' weights are given according to pixel
similarity (a large di erence in pixel color generates a low
weight). Finally, graph and image are segmented by a
mincut algorithm. See an example in Fig. 2(b) of a user marker
and corresponding segmented output in Fig. 2(c).</p>
      <p>The advantages of this approach are its simplicity in a user
perspective, generation of good segmentations and that it is
faster than the segmentation approaches proposed in Section
1.1.2, requiring average time of 1.5 second.
3.2</p>
    </sec>
    <sec id="sec-8">
      <title>Classification</title>
      <p>We compare multiple features (i.e. color, texture and
shape-based features) with histogram matching, which is
fast and invariant to rotation, scale and partial occlusion.</p>
      <p>For each class, a weight is assigned to each feature and
(a) Two classes that have similar color for (b) Species that can have many di erent col- (c) Two classes with similar shape and color,
all samples, but variations in shape. ors, while having a similar shape. but with large variation in texture.
the segmented images are matched by comparing their
histograms, using a kNN classi er. The di erence between
two histograms is computed by a metric based on the
Bhattacharyya distance that is described as follows.</p>
      <p>Consider histograms H1 and H2, with n bins. Let H1(i)
denote the ith bin element of H1, i 2 1 : : : n. H2(i) is de ned
analogously. The distance between H1 and H2 is measured
as:
v
u
d(H1; H2) = uu1
u
t</p>
      <p>Pin=1 pH1(i)
s n</p>
      <p>P H1(i)
i=1</p>
      <p>H2(i)
n
P H2(i)
i=1</p>
      <p>Note that low scores indicate best matches. Our algorithm
compares histograms built on three features cues.
3.2.1</p>
      <sec id="sec-8-1">
        <title>Features</title>
      </sec>
      <sec id="sec-8-2">
        <title>Color.</title>
        <p>We build a histogram of the segmented image's colors in
the HSV space, which is known to be less sensitive to lighting
variations, using 30 bins for hue and 32 bins for saturation.
Image quantization is applied before computation and the
histograms are normalized in order to be comparable with
the proposed distance metric.</p>
      </sec>
      <sec id="sec-8-3">
        <title>Texture.</title>
        <p>
          The texture operator applied in the segmented image is
LBP (Local Binary Pattern) [
          <xref ref-type="bibr" rid="ref23">23</xref>
          ], that is commonly used in
real time applications due to its computational simplicity.
In LBP, pixels are represented as a binary number that is
product from thresholding its neighbors against it.
        </p>
        <p>
          We use the extended version of LBP [
          <xref ref-type="bibr" rid="ref24">24</xref>
          ], that considers
a circular neighborhood with variable radius and is able to
detect details in images with di erent scales. Our method
computes a single histogram with 255 bins, once the
considered textures are mainly uniform and spatial texture
information did not improve the classi cation in our tests.
        </p>
      </sec>
      <sec id="sec-8-4">
        <title>Shape.</title>
        <p>We propose the use of two simple and fast descriptors
that, when combined, are able to represent di erent shape
characteristics. The shape contour is partitioned in 72 bins
using xed angles, resulting of m contour points. For each
point pi, i 2 1 : : : m, we compute modular and angular
descriptors, and both vectors are later represented as separate
histograms.</p>
        <p>The modular descriptor is taken by computing the point
distance from the shape centroid, normalized by the major
distance dmax computed in this process. It measures the
contour relation according to the shape mass center, and
elongated shapes are easily distinguished from round shapes.
The angular descriptor computes the angles between p and
its neighbors pi 1 and pi+1, measuring the smoothness of
the contour. The proposed descriptor becomes powerful by
merging two simple features being able to represent, for
instance, petals length, density and symmetry.
3.2.2</p>
      </sec>
      <sec id="sec-8-5">
        <title>Metric Learning</title>
        <p>In order to achieve higher accuracy, we assign di erent
weights to the features when matching each class, as to
contemplate the situations described in Fig. 3. Our idea is to</p>
        <p>nd which features are more discriminative taking into
account each feature variation inside a same class, but also,
with respect to the global variation of all classes. We learn,
for each class, one weight for each of the four feature
descriptors. We estimate few weights due to the small number
of training samples per class.</p>
        <p>We consider N classes, each containing M training
samples, evaluated according to P features7.</p>
        <p>Let Hip;j denote, for a class i, the histogram of a sample
j regarding feature p. We compute the mean histogram of a
feature p for a class i as:</p>
        <p>M</p>
        <p>P Hip;j
Hip = j=1</p>
        <p>M</p>
        <p>Likewise, we de ne "ip as the mean distance per class i,
regarding feature p, computing the distance from all the
training samples to the mean histogram. This help us to
evaluate intra-class variations, once we are considering the
di erence between all samples with respect to a histogram
that estimates the overall structure of the class.</p>
        <p>Also, we compute the mean distance per feature p ("p) as
follows:</p>
        <p>M</p>
        <p>P d(Hip;j ; Hip)
"ip = j=1</p>
        <p>M
N
P "p</p>
        <p>i
"p = i=1</p>
        <p>N
and we use this information to estimate inter-class
variations between all types of species regarding each considered
feature. Finally, the weight attributed to each of the
considered features p in class i is given by:
7In our tests, we use N = 102, M = 20 and P = 4, since we
consider two histograms for shape information.</p>
        <p>(a) Input belonging to Bearded Iris (BI) class, with large variation in color.
(b) Input belonging to Bishop of Llanda (BoL) class, with small variation in color.
(c) Input belonging to Spring Crocus (SC) class, with large variation on texture.
(d) Input belonging to Blackberry Lily (BL) class, with small variation in texture.
and we are able to take into account intra-class and
interclass variations between all training samples, estimating how
each feature should be considered when evaluating each
considered class.</p>
        <p>After the weights computation, we employ a kNN classi er
in the test phase where the cost of matching an image I with
an image IC of class C is computed as:</p>
        <p>C(I; IC ) = XP (C; p)
p=1</p>
        <p>d(HI ; HIC )
maxi("ip)
and select the classes with the lowest costs. We choose the
kNN classi er because it is robust and performs well with
training sets whose dimension is similar to the ones we are
dealing with.</p>
      </sec>
    </sec>
    <sec id="sec-9">
      <title>RESULTS</title>
      <p>Our tests followed the speci cations for using the Oxford
Dataset as a benchmark, considering 20 training samples per
class and computing the mean-per-class accuracy. In this
dataset, the number of images in each class varies from 40
to 258. Fig. 4 shows the top 5 results for 4 di erent inputs.
Every image is segmented as described in Section 3.1 and
the used weights are given as described in our metric
learning method. Note how the weights a ect the classi cation
results.</p>
      <p>Table 2 describes the algorithm's accuracy for a variable
number of top matches, with and without using metric
learning. Once our platform returns the top 5 most likely classes,
we can see that the system is able to achieve a very high
accuracy rate for a very challenging dataset.</p>
      <p>
        E ciency can not be compared precisely, because this
information is not reported in all the previous works. However,
our average time for extracting and matching all features
proved to be 4 times faster than running SIFT in the
images' foreground region. The baseline work of Angelova and
Shenghuo [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ] takes about 5 seconds for segmentation and 2
seconds for classi cation. Our classi cation runs in less than
a second, and there is a tradeo for segmentation: ours is
semi-automatic, but requires 1.5 seconds on average.
      </p>
      <p>Finally, we replicated the training instances, and Table
4 shows that our method is still e cient for much larger
training sets. Approximate nearest-neighbor methods based
on hashing or kd-trees could also be used for obtaining even
higher e ciency.</p>
    </sec>
    <sec id="sec-10">
      <title>CONCLUSIONS</title>
      <p>This paper proposes a citizen science platform based on
a novel approach for ower recognition. We introduce a
strategy for comparing feature histograms for ne-grained
classi cation, a robust shape descriptor and a metric
learning approach that employs di erent weights to each feature,
that can improve classi cation accuracy signi cantly. Our
algorithm is extremely fast, being suitable for o ine
mobile applications and was able to outperform previous works
using the popular Oxford Dataset.</p>
      <p>Our system is organized around missions, which in
general will help us acquire more data due to gami cation
aspects. Also, besides engaging users, missions provide
additional data external to the image, that may be useful for
aiding classi cation in the future.</p>
      <p>Other future works include testing our algorithm in other
species that contain large variations in the considered
features (color, texture and shape), such as butter ies, shes,
birds and so on, and evaluating more e cient classi ers.</p>
      <p>We will also evaluate our system with real users
analyzing how their behavior is a ected with gami cation and
automatic classi cation techniques. Finally, we aim to use
crowdsourcing for labeling training images, in order to build
various datasets that should represent the local ora and
fauna for diverse locations.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>A.</given-names>
            <surname>Angelova</surname>
          </string-name>
          and
          <string-name>
            <given-names>S.</given-names>
            <surname>Zhu</surname>
          </string-name>
          .
          <article-title>E cient object detection and segmentation for ne-grained recognition</article-title>
          .
          <source>In 2013 IEEE Conference on Computer Vision</source>
          and Pattern Recognition,
          <source>CVPR'13</source>
          , pages
          <fpage>811</fpage>
          {
          <fpage>818</fpage>
          ,
          <year>2013</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>A.</given-names>
            <surname>Angelova</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Zhu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Lin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Wong</surname>
          </string-name>
          , and
          <string-name>
            <given-names>C.</given-names>
            <surname>Shpecht</surname>
          </string-name>
          .
          <article-title>Development and deployment of a large-scale ower recognition mobile app</article-title>
          .
          <source>In NEC Labs America Technical Report</source>
          ,
          <year>2012</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>E.</given-names>
            <surname>Aptoula</surname>
          </string-name>
          and
          <string-name>
            <given-names>B.</given-names>
            <surname>Yanikoglu</surname>
          </string-name>
          .
          <article-title>Morphological features for leaf based plant recognition</article-title>
          .
          <source>In IEEE International Conference on Image Processing, ICIP '13</source>
          , pages
          <fpage>1496</fpage>
          {
          <fpage>1499</fpage>
          ,
          <year>2013</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>H.</given-names>
            <surname>Cai</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Yan</surname>
          </string-name>
          , and
          <string-name>
            <given-names>K.</given-names>
            <surname>Mikolajczyk</surname>
          </string-name>
          .
          <article-title>Learning weights for codebook in image classi cation and retrieval</article-title>
          .
          <source>In Computer Vision and Pattern Recognition</source>
          , 2010 IEEE Conference on,
          <source>CVPRa^AZ10</source>
          , pages
          <fpage>2320</fpage>
          {
          <fpage>2327</fpage>
          ,
          <year>June 2010</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>Y.</given-names>
            <surname>Chai</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Lempitsky</surname>
          </string-name>
          ,
          <article-title>and</article-title>
          <string-name>
            <surname>A. Zisserman.</surname>
          </string-name>
          <article-title>BiCoS: A bi-level co-segmentation method for image classi cation</article-title>
          .
          <source>In IEEE International Conference on Computer Vision</source>
          , ICCV'
          <volume>11</volume>
          , pages
          <fpage>2579</fpage>
          {
          <fpage>2586</fpage>
          ,
          <year>2011</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>S.-Y.</given-names>
            <surname>Cho</surname>
          </string-name>
          .
          <article-title>Content-based structural recognition for ower image classi cation</article-title>
          .
          <source>In IEEE Conference on Industrial Electronics and Applications</source>
          , ICIEA'
          <volume>12</volume>
          , pages
          <fpage>541</fpage>
          {
          <fpage>546</fpage>
          ,
          <year>2012</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>S.</given-names>
            <surname>Deterding</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Dixon</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Khaled</surname>
          </string-name>
          , and
          <string-name>
            <given-names>L.</given-names>
            <surname>Nacke</surname>
          </string-name>
          .
          <article-title>From game design elements to gamefulness: De ning "gami cation"</article-title>
          .
          <source>In Proceedings of the 15th International Academic MindTrek Conference: Envisioning Future Media Environments, MindTrek '11</source>
          , pages
          <fpage>9</fpage>
          {
          <fpage>15</fpage>
          , New York, NY, USA,
          <year>2011</year>
          . ACM.
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>S.</given-names>
            <surname>Deterding</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Sicart</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Nacke</surname>
          </string-name>
          ,
          <string-name>
            <surname>K. O'Hara</surname>
            ,
            <given-names>and D.</given-names>
          </string-name>
          <string-name>
            <surname>Dixon</surname>
          </string-name>
          .
          <article-title>Gami cation. using game-design elements in non-gaming contexts</article-title>
          .
          <source>In CHI '11 Extended Abstracts on Human Factors in Computing Systems, CHI EA '11</source>
          , pages
          <fpage>2425</fpage>
          {
          <fpage>2428</fpage>
          , New York, NY, USA,
          <year>2011</year>
          . ACM.
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>H.</given-names>
            <surname>Go</surname>
          </string-name>
          <article-title>eau, A</article-title>
          . Joly,
          <string-name>
            <given-names>P.</given-names>
            <surname>Bonnet</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Bakic</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Barthelemy</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Boujemaa</surname>
          </string-name>
          , and
          <string-name>
            <given-names>J.-F.</given-names>
            <surname>Molino</surname>
          </string-name>
          .
          <article-title>The ImageCLEF plant identi cation task 2013</article-title>
          .
          <source>In Proceedings of the 2Nd ACM International Workshop on Multimedia Analysis for Ecological Data, MAED '13</source>
          , pages
          <fpage>23</fpage>
          {
          <fpage>28</fpage>
          , New York, NY, USA,
          <year>2013</year>
          . ACM.
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <surname>R.-G. Huang</surname>
            ,
            <given-names>S.-H.</given-names>
          </string-name>
          <string-name>
            <surname>Jin</surname>
          </string-name>
          , Y.-L. Han,
          <article-title>and</article-title>
          K.-S. Hong.
          <article-title>Flower image recognition based on image rotation and DIE</article-title>
          . In International Conference on Digital Content,
          <source>Multimedia Technology and its Applications</source>
          , IDC'
          <volume>10</volume>
          , pages
          <fpage>225</fpage>
          {
          <fpage>228</fpage>
          ,
          <year>2010</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>S.</given-names>
            <surname>Ito</surname>
          </string-name>
          and
          <string-name>
            <given-names>S.</given-names>
            <surname>Kubota</surname>
          </string-name>
          .
          <article-title>Object classi cation using heterogeneous co-occurrence features</article-title>
          .
          <source>In European Conference on Computer Vision</source>
          , ECCV'
          <volume>10</volume>
          , pages
          <fpage>27</fpage>
          {
          <fpage>30</fpage>
          ,
          <year>2010</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>C.</given-names>
            <surname>Kanan</surname>
          </string-name>
          and
          <string-name>
            <given-names>G.</given-names>
            <surname>Cottrell</surname>
          </string-name>
          .
          <article-title>Robust classi cation of objects, faces, and owers using natural image statistics</article-title>
          .
          <source>In IEEE Conference on Computer Vision</source>
          and Pattern Recognition,
          <source>CVPR'10</source>
          , pages
          <fpage>2472</fpage>
          {
          <fpage>2479</fpage>
          ,
          <year>2010</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <given-names>F. S.</given-names>
            <surname>Khan</surname>
          </string-name>
          , J. van de Weijer,
          <string-name>
            <given-names>A. D.</given-names>
            <surname>Bagdanov</surname>
          </string-name>
          , and
          <string-name>
            <given-names>M.</given-names>
            <surname>Vanrell</surname>
          </string-name>
          .
          <article-title>Portmanteau vocabularies for multi-cue image representation</article-title>
          .
          <source>In International Conference on Neural Information Processing Systems</source>
          ,
          <year>2011</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <surname>J.-H. Kim</surname>
            , R.-G. Huang,
            <given-names>S.-H.</given-names>
          </string-name>
          <string-name>
            <surname>Jin</surname>
            , and
            <given-names>K.-S.</given-names>
          </string-name>
          <string-name>
            <surname>Hong</surname>
          </string-name>
          .
          <article-title>Mobile-based ower recognition system</article-title>
          .
          <source>In International Conference on Intelligent Information Technology Application, IITA'09</source>
          , pages
          <fpage>580</fpage>
          {
          <fpage>583</fpage>
          ,
          <year>2009</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <given-names>N.</given-names>
            <surname>Kumar</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P. N.</given-names>
            <surname>Belhumeur</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Biswas</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D. W.</given-names>
            <surname>Jacobs</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W. J.</given-names>
            <surname>Kress</surname>
          </string-name>
          ,
          <string-name>
            <given-names>I. C.</given-names>
            <surname>Lopez</surname>
          </string-name>
          , and
          <string-name>
            <given-names>J. a. V. B.</given-names>
            <surname>Soares</surname>
          </string-name>
          . Leafsnap:
          <article-title>A computer vision system for automatic plant species identi cation</article-title>
          .
          <source>In European Conference on Computer Vision</source>
          , ECCV'
          <volume>12</volume>
          , pages
          <fpage>502</fpage>
          {
          <fpage>516</fpage>
          ,
          <year>2012</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [16]
          <string-name>
            <given-names>Y.</given-names>
            <surname>Lin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Zhu</surname>
          </string-name>
          ,
          <article-title>and</article-title>
          <string-name>
            <given-names>A.</given-names>
            <surname>Angelova</surname>
          </string-name>
          .
          <article-title>Image segmentation for large-scale subcategory ower recognition</article-title>
          .
          <source>In IEEE Workshop on Applications of Computer Vision</source>
          , WACV '
          <volume>13</volume>
          , pages
          <fpage>39</fpage>
          {
          <fpage>45</fpage>
          ,
          <year>2013</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [17]
          <string-name>
            <given-names>D.</given-names>
            <surname>Lowe</surname>
          </string-name>
          .
          <article-title>Object recognition from local scale-invariant features</article-title>
          .
          <source>In International Conference on Computer Vision</source>
          , ICCV '
          <volume>99</volume>
          , pages
          <fpage>1150</fpage>
          {
          <fpage>1157</fpage>
          ,
          <year>1999</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          [18]
          <string-name>
            <given-names>S.</given-names>
            <surname>Mouine</surname>
          </string-name>
          ,
          <string-name>
            <surname>I.</surname>
          </string-name>
          <article-title>Yahiaoui, and</article-title>
          <string-name>
            <given-names>A.</given-names>
            <surname>Verroust-Blondet</surname>
          </string-name>
          .
          <article-title>Plant species recognition using spatial correlation between the leaf margin and the leaf salient points</article-title>
          .
          <source>In IEEE International Conference on Image Processing, ICIP '13</source>
          ,
          <year>2013</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          [19]
          <string-name>
            <given-names>S.</given-names>
            <surname>Mouine</surname>
          </string-name>
          ,
          <string-name>
            <surname>I.</surname>
          </string-name>
          <article-title>Yahiaoui, and</article-title>
          <string-name>
            <given-names>A.</given-names>
            <surname>Verroust-Blondet</surname>
          </string-name>
          .
          <article-title>A shape-based approach for leaf classi cation using multiscaletriangular representation</article-title>
          .
          <source>In ACM Conference on International Conference on Multimedia Retrieval, ICMR '13</source>
          , pages
          <fpage>127</fpage>
          {
          <fpage>134</fpage>
          ,
          <year>2013</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          [20]
          <string-name>
            <surname>M.-E. Nilsback</surname>
            and
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Zisserman</surname>
          </string-name>
          .
          <article-title>A visual vocabulary for ower classi cation</article-title>
          .
          <source>In IEEE Computer Society Conference on Computer Vision</source>
          and Pattern Recognition,
          <source>CVPR'06</source>
          , pages
          <fpage>1447</fpage>
          {
          <fpage>1454</fpage>
          ,
          <year>2006</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          [21]
          <string-name>
            <surname>M.-E. Nilsback</surname>
            and
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Zisserman</surname>
          </string-name>
          .
          <article-title>Automated ower classi cation over a large number of classes</article-title>
          .
          <source>In Indian Conference on Computer Vision</source>
          , Graphics &amp; Image
          <string-name>
            <surname>Processing</surname>
          </string-name>
          ,
          <source>ICVGIP'08</source>
          , pages
          <fpage>722</fpage>
          {
          <fpage>729</fpage>
          ,
          <year>2008</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          [22]
          <string-name>
            <surname>M.-E. Nilsback</surname>
            and
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Zisserman</surname>
          </string-name>
          .
          <article-title>An automatic visual Flora - segmentation and classi cation of ower images</article-title>
          .
          <source>PhD thesis</source>
          , University of Oxford, UK,
          <year>2009</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          [23]
          <string-name>
            <given-names>T.</given-names>
            <surname>Ojala</surname>
          </string-name>
          , M. Pietikainen, and
          <string-name>
            <given-names>D.</given-names>
            <surname>Harwood</surname>
          </string-name>
          .
          <article-title>A comparative study of texture measures with classi cation based on featured distributions</article-title>
          .
          <source>Pattern Recognition</source>
          ,
          <volume>29</volume>
          (
          <issue>1</issue>
          ):
          <volume>51</volume>
          {
          <fpage>59</fpage>
          ,
          <year>1996</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          [24]
          <string-name>
            <given-names>T.</given-names>
            <surname>Ojala</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Pietikainen</surname>
          </string-name>
          , and
          <string-name>
            <given-names>T.</given-names>
            <surname>Maenpaa</surname>
          </string-name>
          .
          <article-title>Multiresolution gray-scale and rotation invariant texture classi cation with local binary patterns</article-title>
          .
          <source>Pattern Analysis and Machine Intelligence</source>
          , IEEE Transactions on,
          <volume>24</volume>
          (
          <issue>7</issue>
          ):
          <volume>971</volume>
          {
          <fpage>987</fpage>
          ,
          <year>2002</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref25">
        <mixed-citation>
          [25]
          <string-name>
            <given-names>W.</given-names>
            <surname>Qi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Liu</surname>
          </string-name>
          , and
          <string-name>
            <given-names>J.</given-names>
            <surname>Zhao</surname>
          </string-name>
          .
          <article-title>Flower classi cation based on local and spatial visual cues</article-title>
          .
          <source>In IEEE International Conference on Computer Science and Automation Engineering</source>
          , CSAE'
          <volume>12</volume>
          , pages
          <fpage>670</fpage>
          {
          <fpage>674</fpage>
          ,
          <year>2012</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref26">
        <mixed-citation>
          [26]
          <string-name>
            <given-names>X.</given-names>
            <surname>Qi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Xiao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Zhang</surname>
          </string-name>
          , C.-G. Li,
          <string-name>
            <given-names>and J.</given-names>
            <surname>Guo</surname>
          </string-name>
          .
          <article-title>A rapid ower/leaf recognition system</article-title>
          .
          <source>In ACM International Conference on Multimedia, MM '12</source>
          , pages
          <fpage>1257</fpage>
          {
          <fpage>1258</fpage>
          ,
          <year>2012</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref27">
        <mixed-citation>
          [27]
          <string-name>
            <given-names>C.</given-names>
            <surname>Rother</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Kolmogorov</surname>
          </string-name>
          ,
          <article-title>and</article-title>
          <string-name>
            <given-names>A.</given-names>
            <surname>Blake</surname>
          </string-name>
          . \
          <article-title>Grabcut": Interactive foreground extraction using iterated graph cuts</article-title>
          .
          <source>In ACM SIGGRAPH, SIGGRAPH'04</source>
          , pages
          <fpage>309</fpage>
          {
          <fpage>314</fpage>
          ,
          <year>2004</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref28">
        <mixed-citation>
          [28]
          <string-name>
            <given-names>H.</given-names>
            <surname>Roy</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Pocock</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Preston</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Roy</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Savage</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Tweddle</surname>
          </string-name>
          , and
          <string-name>
            <given-names>L.</given-names>
            <surname>Robinson</surname>
          </string-name>
          .
          <article-title>Understanding citizen science and environmental monitoring: nal report on behalf of UK Environmental Observation Framework</article-title>
          .
          <source>Technical report</source>
          , Wallingford, NERC/Centre for Ecology &amp; Hydrology,
          <year>2012</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref29">
        <mixed-citation>
          [29]
          <string-name>
            <given-names>T.</given-names>
            <surname>Saitoh</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Aoki</surname>
          </string-name>
          , and
          <string-name>
            <given-names>T.</given-names>
            <surname>Kaneko</surname>
          </string-name>
          .
          <article-title>Automatic recognition of blooming owers</article-title>
          .
          <source>In International Conference on Pattern Recognition, ICPR'04</source>
          , pages
          <fpage>27</fpage>
          {
          <fpage>30</fpage>
          ,
          <year>2004</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref30">
        <mixed-citation>
          [30]
          <string-name>
            <given-names>F.</given-names>
            <surname>Siraj</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Salahuddin</surname>
          </string-name>
          , and
          <string-name>
            <given-names>S.</given-names>
            <surname>Yusof</surname>
          </string-name>
          .
          <article-title>Digital image classi cation for malaysian blooming ower</article-title>
          .
          <source>In International Conference on Computational Intelligence, Modelling and Simulation</source>
          , pages
          <volume>33</volume>
          {
          <fpage>38</fpage>
          ,
          <year>2010</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref31">
        <mixed-citation>
          [31]
          <string-name>
            <given-names>J.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Kalousis</surname>
          </string-name>
          ,
          <article-title>and</article-title>
          <string-name>
            <given-names>A.</given-names>
            <surname>Woznica</surname>
          </string-name>
          .
          <article-title>Parametric local metric learning for nearest neighbor classi cation</article-title>
          .
          <source>In Neural Information Processing Systems (NIPS)</source>
          , pages
          <fpage>1610</fpage>
          {
          <fpage>1618</fpage>
          ,
          <year>2012</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref32">
        <mixed-citation>
          [32]
          <string-name>
            <given-names>K. Q.</given-names>
            <surname>Weinberger</surname>
          </string-name>
          and
          <string-name>
            <given-names>L. K.</given-names>
            <surname>Saul</surname>
          </string-name>
          .
          <article-title>Distance metric learning for large margin nearest neighbor classi cation</article-title>
          .
          <source>J. Mach. Learn. Res.</source>
          ,
          <volume>10</volume>
          :
          <fpage>207</fpage>
          {
          <fpage>244</fpage>
          ,
          <year>June 2009</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref33">
        <mixed-citation>
          [33]
          <string-name>
            <given-names>J.</given-names>
            <surname>Zou</surname>
          </string-name>
          and
          <string-name>
            <given-names>G.</given-names>
            <surname>Nagy</surname>
          </string-name>
          .
          <article-title>Evaluation of model-based interactive ower recognition</article-title>
          .
          <source>In International Conference on Pattern Recognition, ICPR '04</source>
          ,
          <year>2004</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>