<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Image Aesthetics and its Efects on Product Clicks in E-Commerce Search</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Alexandre Maros</string-name>
          <email>alexandremaros@dcc.ufmg.br</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Fabiano Belém</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Rodrigo Silva</string-name>
          <email>rmsilva@dcc.ufmg.br</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Sérgio Canuto Jussara M. Almeida</string-name>
          <email>jussara@dcc.ufmg.br</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Marcos A. Gonçalves</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Universidade Federal de Minas Gerais - Computer Science Department</institution>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2019</year>
      </pub-date>
      <abstract>
        <p>Product search engines are a key factor for online business. Retrieving relevant products in an E-Commerce (EC) website is of the utmost importance, as a single EC website can have millions of products with very similar features. One aspect that is not widely studied in this scenario is the efect that the image of the product (notably aesthetic properties) shown to the customer has on the customer's interest. Previous studies have been able to link certain characteristics of images to our innate interest. For instance, it is known that bright images with several colors are more likely to attract one's attention than dark ones. However, these issues have been understudied in the EC context. In this context, we conduct experiments on real-world EC to analyze the efects that the product's image aesthetic has on the user interest (expressed in clicks) in the product. Experimental results show that this relationship exists and that it is more visible in some categories of products.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>INTRODUCTION</title>
      <p>
        E-Commerce (EC) platforms have become a popular means to bring
greater shopping convenience to costumers around the globe. These
platforms bring a series of social and technical challenges, which
have not been extensively studied. Examples of these challenges
include problems related to the trust and familiarity that these
websites convey to the customers [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ] as well as new computational
challenges, specially in the area of information retrieval, such as
new ways to increase revenue given specific s earches a nd user
profiles [
        <xref ref-type="bibr" rid="ref19">19</xref>
        ].
      </p>
      <p>
        After stumbling upon a product in an EC website, as a result of
an explicit search or associated with an ad, the user’s decision to
click on such product (which, in turn, may lead to a purchase) may
be influenced by the image of the product shown to the user. Thus,
even though this is commonly overlooked or taken for granted,
the product’s image, alongside with its title, description and tags,
may be deciding factors between buying the product or not. The
quality of the product must be clearly conveyed through these
features in order to convince the customer to buy the product
[
        <xref ref-type="bibr" rid="ref10 ref2">2, 10</xref>
        ]. In such context, it is known in other domains that certain
image characteristics (e.g., brightness, colorfulness) are related to
one’s innate interest [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ]. However, this issue has not been fully
investigated in the context of EC search. Thus, we here investigate
the hypothesis that aesthetics properties of the product’s image shown
to the customers of EC websites does influence the product clicks in
searches, and possibly the amount of purchases, since clicking and
visualizing a product is an important step to guide the user decision
for purchasing the product or not.
      </p>
      <p>
        The challenge of evaluating the aesthetics, beauty, or quality of
an image, music or any artistic work has been studied for quite
some time [
        <xref ref-type="bibr" rid="ref1 ref6">1, 6</xref>
        ]. The main problem involving this field of study
is that beauty is often considered as personal; what may be pretty
for one person may not be liked by another. With product image
classification these factors can vary even more, e.g., the image does
not necessarily needs to be pretty, but it must be clearly visible.
However, our driving hypothesis is that products that have more
attractive images (have better brightness, contrast, saturation) may
have a higher probability of being clicked when they appear.
      </p>
      <p>
        Even though this is a complex and noisy problem to model, there
are few general characteristics of artistic works that determine their
visual quality. For instance, in photography, exposure, rule of thirds,
contrast, and other characteristics are often carefully planned and
chosen by the photographer. A few authors, such as [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ], define
beauty as a ratio between harmony and complexity of a work.
      </p>
      <p>The present study tackles the aforementioned hypothesis by
analyzing the relationship between the quality of a product’s image
and the probability of it being clicked when presented as a result
for a search query in a large EC website focused on crafts and
personalized products. Our goal is to verify whether the image
features, including features related to aesthetics, can explain, at
least partially, those clicks. Our experimental results show that
this relationship does exist, even though it seems to be noisy. The
applied methodology could be used to improve the search results
of EC, presenting more attractive products to the customers or to
guide sellers towards improving the attractiveness of their products.
2</p>
    </sec>
    <sec id="sec-2">
      <title>RELATED WORK</title>
      <p>
        There are several works focused on the definition and evaluation
of the aesthetics of an image. In one of the earliest studies [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ],
the aesthetic measure is formalized as the ratio between order
and complexity. In [
        <xref ref-type="bibr" rid="ref18">18</xref>
        ], the artistic process is modeled from an
Information Theory perspective to allow the quantification of these
properties. In particular, the Shannon Entropy and Kolmogorov
Complexity are used to estimate values for the order and complexity
of some famous paintings.
      </p>
      <p>
        A more recent work [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ] proposes a method to extract features
from images in the Photo.net dataset in order to use machine
learning techniques such as Support Vector Machines (SVMs) to classify
images into aesthetically good or bad categories. Some of these
features, originally proposed by [
        <xref ref-type="bibr" rid="ref11 ref5">5, 11</xref>
        ], include characteristics such
as brightness, contrast, saturation, central saturation and image ratio.
Other characteristics such as Bags of Visual Words (BOVW) [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ],
Fisher Vectors (FV) [
        <xref ref-type="bibr" rid="ref17">17</xref>
        ] and GIST descriptors [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ] have also been
used to improve learning accuracy. The BOVW and FV algorithms
work by clustering Scale-Invariant Feature Transform (SIFT)
vectors [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ] that capture local properties of an image (e.g., “Does this
patch contain sharp edges?”, “Is the color of this patch saturated?”).
GIST in turn uses histograms to capture information from images.
      </p>
      <p>
        The use of Convolutional Neural Networks (CNN) is explored in
[
        <xref ref-type="bibr" rid="ref13">13</xref>
        ]. The objective is to build a robust way to automatically extract
features and predict whether an image is considered good or bad.
The main advantage is that there is no need to propose nor calculate
predefined features, which can be expensive. However, there are
two main disadvantages: (a) the models must have a fixed-size input,
implying the need to scale the images, causing information loss
(this can be minimized by using random crops of the same image)
and; (b) the ability to interpret and understand the produced model
is afected. The CNNs are trained and tested using the AVA dataset 1,
containing over 1.5 million images.
      </p>
      <p>
        There are very few works that try to correlate the quality of the
image of a product with the product popularity (e.g., click rate).
In [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ], a study of the impact of the image on user clicks was
conducted. The authors used a limited number of image features and a
stochastic gradient boosting model to predict CTR of the randomly
selected products. The results indicated significant correlation
between the images features and CTR. In [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ], the authors find that
the “Perceived Product Quality”, whose definition relies directly on
information available to the consumer, such as the product images,
description, title and reviews directly afects the interest in a
product. In here, we take a diferent perspective based on the aesthetics
of the image associated with the product and on features that can
be automatically extracted from them based on this perspective.
      </p>
    </sec>
    <sec id="sec-3">
      <title>3 IMAGE FEATURES</title>
      <p>We use use three sets of of automatically extracted features to
capture the aesthetic of the product images. The first one corresponds
to state-of-the-art aesthetic base features that capture the most
fundamental aspects of the image (e.g. exposure and saturation). The
second one is the GIST descriptor that captures scene categorization
and image layout. The third and last set is the BOVW, a generic
content-based set of features which describe the distribution of
local patches within the image.
3.1</p>
    </sec>
    <sec id="sec-4">
      <title>Base Features</title>
      <p>
        The base features consists of 66 attributes extracted from the
product’s image. These attributes are directly related to the main aspects
1https://research.google.com/ava/
of the image, including features that capture the exposure, contrast,
saturation, ratio, depth, sharpness and composition [
        <xref ref-type="bibr" rid="ref15 ref3 ref5">3, 5, 15</xref>
        ]. Most
of these features were successfully exploited in aesthetic
classification in other domains [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ], thus being be a good starting point.
Table 1 presents these features with a brief description .
3.2
      </p>
    </sec>
    <sec id="sec-5">
      <title>GIST</title>
      <p>
        The GIST descriptor is a low-dimensional scene descriptor that
captures a set of characteristics, such as naturalness, roughness,
expansion and ruggedness. These characteristics are estimated using
spectral information and coarse localization. The image is
partitioned into a 4 × 4 regular grid and a histogram of gradients (with
20 bins) is computed for each of the 16 regions and 3 color channels.
Finally, all histograms are concatenated to form a 960D vector [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ].
3.3
      </p>
    </sec>
    <sec id="sec-6">
      <title>Bag of Visual Words (BOVW)</title>
      <p>
        BOVW represents an image by a histogram of local features [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ].
First, an unordered set of local patches are extracted and described
by SIFT descriptor. A visual vocabulary is learned by clustering
these descriptors with K-Means. The local features are then
extracted from an image by counting the number of local descriptors
assigned to each visual word in a fixed-length histogram. This
algorithm has been very successful in image classification. [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ].
4
      </p>
    </sec>
    <sec id="sec-7">
      <title>DATASET</title>
      <p>Our evaluation exploits real data from Elo7, the largest Brazilian
emarketplace focused on creative and personalized products2. In this
EC platform, sellers register their own products, uploading pictures
and providing descriptive text data (e.g. title, description, price,
tags). Product searches return a grid of products where the photos
are prominent and textual data include only the product’s title and
price. Since the sellers are responsible for this information, there
is a lot of heterogeneity in terms of quality of the products’ text
features and, specially, their pictures, making it an ideal scenario
for this study. The data contains over four million products in 42
Image Aesthetics and its Efects on Product Clicks in E-Commerce Search
diferent categories and a hundred thousand unique queries across
two months (October and September of 2018).</p>
      <p>In this work, departing from our main hypothesis, we associate
the interest of users in a product with the quality of its image.
More formally, given a search query q and a product p, we have the
number of times in which p was retrieved after a customer searched
for q (impressions) and how many times it was clicked. The interest
score of a product with relation to a query is then calculated as the
ratio between the number of clicks and impressions times the log
clicks
of the number of clicks: clickscore = impressions × log2 clicks.</p>
      <p>The addition of the logarithmic term is to give an additional
priority for products that have a high click count and are more
popular. We only considered products that have more than a thousand
impressions to try to eliminate eventual noise coming from
products that appeared very few times to the customers. For example,
a high click rate on a product that appeared very few times over
many queries may suggest a customer looking for a very specific
product that does not otherwise show up in other searches and
therefore is not a clear indication that the product has a good image
as perceived by larger set of users.</p>
      <p>Based on these scores, we then label images that are above the
80th percentile as “highly clicked” images. Images that are bellow
this percentile were considered “poorly clicked”. The idea is to
check if image features, including features related to aesthetics,
influence the amount of product clicks in e-commerce search.</p>
      <p>Figure 1 shows two examples of product’s images from the
website. Figure 2a and 2b are from products that have a high and low
score, respectively. The first image has more colors, better lighting
and is overall more attractive than the first one.
5</p>
    </sec>
    <sec id="sec-8">
      <title>EXPERIMENTS AND ANALYSIS</title>
      <p>Two experiments are conducted. First, we verify whether there is
a significant diference in the values of image features between
highly relevant (clicked) products and the less clicked products.
Second, we test the capacity of a machine learning model to predict
whether or not a product will be highly clicked based on its image
features.
(a) f 5 – Exposure
(b) f 8 – Average Intensity
5.1</p>
    </sec>
    <sec id="sec-9">
      <title>Feature Distributions</title>
      <p>To check whether there is a significant diference between
products that are more frequently clicked against their counterparts,
we selected two thousand queries that had products with over a
thousand impressions. For each query, we compare image features
of the highest and lowest ranked products (most and least
“relevant” products, respectively). For each feature, we performed the
Kolmogorov-Smirnov test to check whether the two distributions
(associated with the lowest and highest ranked products according
to their click scores) are significantly diferent. 42 out of the 66
features had p-values under the 0.05 threshold, meaning that they
are statistically diferent. Figure 2 shows the density plot of two
features with a p-value under 0.05.</p>
      <p>By looking at the density plot (Figure 2) of the features under
the two categories (most and least relevant) we can see that the
distributions are slightly diferent, corroborating the idea that
image features do influence users in the click decision. For instance,
in Figure 3a we can see that images that are more clicked have a
higher exposure, which makes sense, since brighter images tend to
attract more attention.
5.2</p>
    </sec>
    <sec id="sec-10">
      <title>Quality Prediction</title>
      <p>
        We performed a second experiment to check how accurately we can
predict the quality of the product’s image based on its click data.
Given that the products have a category, and since the images vary
vastly from category to category, we built diferent models for each
category. Thus, we divided the prediction into 10 sets of products
corresponding to the ten most common categories. A thousand
images labeled as “highly clicked” and a thousand image labeled as
“poorly clicked” were put into each one of these ten groups. Finally,
we use three sets of features (Base, Base + GIST and Base + BOVW)
to train an SVM [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ] with a 10-Fold cross validation procedure.
      </p>
      <p>Table 2 shows the metrics F1, Precision and Recall for the
predictions (in the test sets of the cross-validation procedure) of the ten
categories and the three sets of features. Firstly, we can see that,
although predictions are hard, as the categories have a mean F1 of
56.7% for the Base features (the worst being 53% and the best 61%),
for most categories they are still statistically better than random
(F1=0.5). Thus, there is indeed some predictive power in the set
of Base features. Secondly, we can see that the GIST features only
helped to improve the F1 score in four out of the ten categories,
making only the recall higher whilst reducing the precision. The
BOVW features improved the results in six of the ten categories,
increasing both, recall and precision. This means that local features
have an interesting potential in this application.</p>
      <p>Finally, we note that clearly some categories are easier to predict
than others. The categories “Candies” and “Children’s” had the
highest F1 while categories such as “Invitations” and “Baby” had
the lowest. This could be explained by the types of images that are
presented in these categories. In the “Candies” category, we usually
have products that have bright colors, with diferent shapes, sizes
and textures that can attract attention and clicks. In this category,
images that do not have many colors or that are darker may have
a lower click rate. However, in the “Invitations” category, where
we have very similar images to one another, mostly white paper
invitations with similar texture, the image may not be a deciding
factor for a product’s click rate.</p>
      <p>Although there are some categories that are easier to predict
than others, the precision is still low. This indicates that the image
alone is not completely responsible for the user’s decision to click
on a product. Other factors such as the price range and the title
of the product may influence customers actions. But, as we can
see from the results, a product’s image quality does have some
influence and could possibly be used to improve search results
for an EC platform. The question of how much the quality of a
product’s image influences its click rate as compared to its other
properties (such as title and price) is left as future work.
6</p>
    </sec>
    <sec id="sec-11">
      <title>CONCLUSIONS &amp; FUTURE WORK</title>
      <p>In this work we investigated how product images influence
product clicks in e-commerce platforms. The detection of aesthetically
“bad” or “good” images can be used to improve e-commerce search
engines, and, consequently, customers’ satisfaction and revenue
for e-commerce companies. They may also provide feedback to
help sellers to make their product more attractive to customers.
Our experiments show that image attributes such as brightness,
colorfulness and contrast can influence product clicks. First, we
analyzed how some of these features vary when comparing more
frequently clicked products with less clicked ones. We then used
machine learning to try to predict product clicks based on image
features. The performed experiments show that there is potential
in this technique, specially for some specific categories of products.</p>
      <p>As future work, we intend to test other machine learning
methods and to add other features, such as Fisher Vectors. We can also
propose specialized models for diferent categories, since the
influence of the images in product clicks seem to vary according to the
product category. We plan to study in more detail why and how
product categories difer in order to produce more accurate models
to better predict product clicks from image quality. Finally, we
intend to run additional comparative experiments (e.g., comparing
images with product’s titles and prices) to deepen our
understanding of a customer’s motivation to click on a product.</p>
    </sec>
    <sec id="sec-12">
      <title>ACKNOWLEDGEMENTS</title>
      <p>This work was partially supported by Elo7, the
FAPEMIG-PRONEXMASWeb project – Models, Algorithms and Systems for the Web,
process APQ-01400-14, as well as by the National Institute of
Science and Technology for the Web (INWEB), CNPq and FAPEMIG.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>G. D.</given-names>
            <surname>Birkhof</surname>
          </string-name>
          . Aesthetic measure, volume
          <volume>38</volume>
          . Harvard University Press Cambridge, 79 Garden St, Cambridge, MA 02138, USA,
          <year>1933</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>Z.</given-names>
            <surname>Chen</surname>
          </string-name>
          and
          <string-name>
            <given-names>A. J.</given-names>
            <surname>Dubinsky</surname>
          </string-name>
          .
          <article-title>A conceptual model of perceived customer value in e-commerce: A preliminary investigation</article-title>
          .
          <source>Psychology &amp; Marketing</source>
          ,
          <volume>20</volume>
          (
          <issue>4</issue>
          ),
          <year>2003</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>F.</given-names>
            <surname>Crete</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Dolmiere</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Ladret</surname>
          </string-name>
          , and
          <string-name>
            <given-names>M.</given-names>
            <surname>Nicolas</surname>
          </string-name>
          .
          <article-title>The blur efect: perception and estimation with a new no-reference perceptual blur metric</article-title>
          .
          <source>In Human vision and electronic imaging XII</source>
          , volume
          <volume>6492</volume>
          .
          <source>International Society for Optics and Photonics</source>
          ,
          <year>2007</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>G.</given-names>
            <surname>Csurka</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Dance</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Fan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Willamowski</surname>
          </string-name>
          , and
          <string-name>
            <given-names>C.</given-names>
            <surname>Bray</surname>
          </string-name>
          .
          <article-title>Visual categorization with bags of keypoints</article-title>
          .
          <source>In Workshop on Statistical Learning in Computer Vision</source>
          ,
          <year>2004</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>R.</given-names>
            <surname>Datta</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Joshi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>and J. Z.</given-names>
            <surname>Wang</surname>
          </string-name>
          .
          <article-title>Studying aesthetics in photographic images using a computational approach</article-title>
          .
          <source>In ECCV</source>
          ,
          <year>2006</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>L.</given-names>
            <surname>Frost</surname>
          </string-name>
          .
          <article-title>What makes a painting good</article-title>
          ?
          <source>PhD thesis</source>
          , Rhodes University Grahamstown,
          <year>1987</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>D.</given-names>
            <surname>Gefen</surname>
          </string-name>
          . E-commerce:
          <article-title>the role of familiarity and trust</article-title>
          .
          <source>Omega</source>
          ,
          <volume>28</volume>
          (
          <issue>6</issue>
          ),
          <year>2000</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>A.</given-names>
            <surname>Goswami</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Chittar</surname>
          </string-name>
          , and
          <string-name>
            <given-names>C. H.</given-names>
            <surname>Sung</surname>
          </string-name>
          .
          <article-title>A study on the impact of product images on user clicks for online shopping</article-title>
          .
          <source>In WWW</source>
          ,
          <year>2011</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>M. A.</given-names>
            <surname>Hearst</surname>
          </string-name>
          .
          <article-title>Support vector machines</article-title>
          .
          <source>IEEE Intelligent Systems</source>
          ,
          <volume>13</volume>
          (
          <issue>4</issue>
          ),
          <year>July 1998</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>S. E.</given-names>
            <surname>Kaplan</surname>
          </string-name>
          and
          <string-name>
            <given-names>R. J.</given-names>
            <surname>Nieschwietz</surname>
          </string-name>
          .
          <article-title>A web assurance services model of trust for b2c e-commerce</article-title>
          .
          <source>Int. J. of Accounting Information Systems</source>
          ,
          <volume>4</volume>
          (
          <issue>2</issue>
          ):
          <fpage>95</fpage>
          -
          <lpage>114</lpage>
          ,
          <year>2003</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>Y.</given-names>
            <surname>Ke</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Tang</surname>
          </string-name>
          , and
          <string-name>
            <given-names>F.</given-names>
            <surname>Jing</surname>
          </string-name>
          .
          <article-title>The design of high-level features for photo quality assessment</article-title>
          .
          <source>In CVPR '06</source>
          , pages
          <fpage>419</fpage>
          -
          <lpage>426</lpage>
          ,
          <year>2006</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>D. G.</given-names>
            <surname>Lowe</surname>
          </string-name>
          .
          <article-title>Distinctive image features from scale-invariant keypoints</article-title>
          .
          <source>Int. J. of Computer Vision</source>
          ,
          <volume>60</volume>
          (
          <issue>2</issue>
          ):
          <fpage>91</fpage>
          -
          <lpage>110</lpage>
          ,
          <year>2004</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <given-names>X.</given-names>
            <surname>Lu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Lin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Jin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Yang</surname>
          </string-name>
          , and
          <string-name>
            <given-names>J. Z.</given-names>
            <surname>Wang</surname>
          </string-name>
          .
          <article-title>Rating image aesthetics using deep learning</article-title>
          .
          <source>IEEE Transactions on Multimedia</source>
          ,
          <volume>17</volume>
          (
          <issue>11</issue>
          ):
          <fpage>2021</fpage>
          -
          <lpage>2034</lpage>
          ,
          <year>2015</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <given-names>L.</given-names>
            <surname>Marchesotti</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Perronnin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Larlus</surname>
          </string-name>
          , and
          <string-name>
            <given-names>G.</given-names>
            <surname>Csurka.</surname>
          </string-name>
          <article-title>Assessing the aesthetic quality of photographs using generic image descriptors</article-title>
          .
          <source>In ICCV</source>
          ,
          <year>2011</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <given-names>P.</given-names>
            <surname>Marziliano</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Dufaux</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Winkler</surname>
          </string-name>
          , and
          <string-name>
            <given-names>T.</given-names>
            <surname>Ebrahimi</surname>
          </string-name>
          .
          <article-title>A no-reference perceptual blur metric</article-title>
          .
          <source>In Image processing'02</source>
          , volume
          <volume>3</volume>
          ,
          <string-name>
            <surname>pages</surname>
            <given-names>III</given-names>
          </string-name>
          -III. IEEE,
          <year>2002</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [16]
          <string-name>
            <given-names>A.</given-names>
            <surname>Oliva</surname>
          </string-name>
          and
          <string-name>
            <given-names>A.</given-names>
            <surname>Torralba</surname>
          </string-name>
          .
          <article-title>Modeling the shape of the scene: A holistic representation of the spatial envelope</article-title>
          .
          <source>Int. J. of Computer Vision</source>
          ,
          <volume>42</volume>
          (
          <issue>3</issue>
          ):
          <fpage>145</fpage>
          -
          <lpage>175</lpage>
          ,
          <year>2001</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [17]
          <string-name>
            <given-names>F.</given-names>
            <surname>Perronnin</surname>
          </string-name>
          and
          <string-name>
            <given-names>C.</given-names>
            <surname>Dance</surname>
          </string-name>
          .
          <article-title>Fisher kernels on visual vocabularies for image categorization</article-title>
          .
          <source>In CVPR</source>
          , pages
          <fpage>1</fpage>
          -
          <lpage>8</lpage>
          ,
          <year>June 2007</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          [18]
          <string-name>
            <given-names>J.</given-names>
            <surname>Rigau</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Feixas</surname>
          </string-name>
          , and
          <string-name>
            <given-names>M.</given-names>
            <surname>Sbert</surname>
          </string-name>
          .
          <article-title>Informational aesthetics measures</article-title>
          .
          <source>IEEE Computer Graphics and Applications</source>
          ,
          <volume>28</volume>
          (
          <issue>2</issue>
          ),
          <year>2008</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          [19]
          <string-name>
            <given-names>L.</given-names>
            <surname>Wu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Hu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Hong</surname>
          </string-name>
          , and
          <string-name>
            <given-names>H.</given-names>
            <surname>Liu</surname>
          </string-name>
          .
          <article-title>Turning clicks into purchases: Revenue optimization for product search in e-commerce</article-title>
          .
          <source>In SIGIR</source>
          , pages
          <fpage>365</fpage>
          -
          <lpage>374</lpage>
          ,
          <year>2018</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>