<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Towards Explanations for Visual Recommender Systems of Artistic Images</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Vicente Dominguez</string-name>
          <email>vidominguez@uc.cl</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Christoph Traner</string-name>
          <email>christoph.trattner@uib.no</email>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Recommender systems, Artwork Recommendation, Explainable</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Pablo Messina</string-name>
          <email>pamessina@uc.cl</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Denis Parra</string-name>
          <email>dparra@ing.puc.cl</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>IMFD &amp; PUC Chile</institution>
          ,
          <addr-line>Santiago</addr-line>
          ,
          <country country="CL">Chile</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Interfaces</institution>
          ,
          <addr-line>Visual Features</addr-line>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>University of Bergen</institution>
          ,
          <addr-line>Bergen</addr-line>
          ,
          <country country="NO">Norway</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2018</year>
      </pub-date>
      <abstract>
        <p>Explaining automatic recommendations is an active area of research since it has shown an important eect on users' acceptance over the items recommended. However, there is a lack of research in explaining content-based recommendations of images based on visual features. In this paper, we aim to ll this gap by testing three dierent interfaces (one baseline and two novel explanation interfaces) for artistic image recommendation. Our experiments with N=121 users conrm that explanations of recommendations in the image domain are useful and increase user satisfaction, perception of explainability, relevance, and diversity. Furthermore, our experiments show that the results are also dependent on the underlying recommendation algorithm used. We tested the interfaces with two algorithms: Deep Neural Networks (DNN), with high accuracy but with dicult to explain features, and the more explainable method based on Aractiveness Visual Features (AVF). e beer the accuracy performance -in our case the DNN method- the stronger the positive eect of the explainable interface. Notably, the explainable features of the AVF method increased the perception of explainability but did not increase the perception of trust, unlike DNN, which improved both dimensions. ese results indicate that algorithms in conjunction with interfaces play a signicant role in the perception of explainability and trust for image recommendation. We plan to further investigate the relationship between interface explainability and algorithmic performance in recommender systems.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>INTRODUCTION</title>
      <p>
        Online artwork recommendation has received lile aention
compared to other areas such as movies [
        <xref ref-type="bibr" rid="ref1 ref10">1, 10</xref>
        ], music [
        <xref ref-type="bibr" rid="ref16 ref4">4, 16</xref>
        ] or
pointsof-interest [
        <xref ref-type="bibr" rid="ref25 ref28 ref29">25, 28, 29</xref>
        ]. e rst works in the area date from
20062007 such as the CHIP [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ] project, which implemented traditional
techniques such as content-based and collaborative ltering for
artwork recommendation at the Rijksmuseum, and the m4art
system by Van den Broek et al. [
        <xref ref-type="bibr" rid="ref26">26</xref>
        ], which used histograms of color
to retrieve similar artworks where the input query was a painting
image. More recently, deep neural networks (DNN) have been used
for artwork recommendation and are the current state-of-the-art
model [
        <xref ref-type="bibr" rid="ref12 ref7">7, 12</xref>
        ], which is rather expected considering that DNNs are
the top performing models for obtaining visual features for several
tasks, such as image classication [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ], and scene identication
[
        <xref ref-type="bibr" rid="ref23">23</xref>
        ]. However, no user study has been conducted to validate the
performance of DNNs versus other visual features. is aspect is
important since past works have shown that o-line results might
not always replicate when tested with actual users [
        <xref ref-type="bibr" rid="ref14 ref17">14, 17</xref>
        ].
Moreover, we provide evidence of the important value of explanations
in artwork recommender systems over several dimensions of user
perception. Visual features obtained from DNNs are still dicult
to explain to users, despite current eorts to understand them and
explain them [
        <xref ref-type="bibr" rid="ref20">20</xref>
        ]. In contrast, features of visual aractiveness
could be easily explained, based on color, brightness or contrast
[
        <xref ref-type="bibr" rid="ref21">21</xref>
        ]. Explanations in recommender systems have been shown to
have a signicant eect on user satisfaction [
        <xref ref-type="bibr" rid="ref24">24</xref>
        ], and, to the best of
our knowledge, no previous work has shown how to explain
recommendations of images based on visual features. Hence, there is no
study of the eect on users when explaining images recommended
by a Visual Content-based Recommender (Hereinaer, VCBR).
      </p>
      <p>Objective. In this paper, we research the eect of explaining
artistic image suggestions. In particular, we conduct a user study
on Amazon Mechanical Turk under three dierent interfaces and
two dierent algorithms. e three interfaces are: i) no
explanations, ii) explanations based on similar images, and iii) explanations
based on visual features. Moreover, the two algorithms are: Deep
Neural Networks (DNN) and Aractiveness Visual Features (AVF).
In our study, we used images provided by the online store UGallery
(hp://www.UGallery.com/).</p>
      <p>Research estions To drive our research, the following two
questions were dened:
a RQ1. Given three dierent types of interfaces, one baseline
interface without explanations and two with them, employing
similar image explanations and a feature bar chart, which one is
perceived as most useful?
a RQ2. Furthermore, based on the visual and content-based
recommender algorithm chosen, are there observable dierences in
how the three interfaces are perceived?
2</p>
    </sec>
    <sec id="sec-2">
      <title>RELATED WORK</title>
      <p>Relevant related research is collated in two sub-sections: First,
we review research on recommending artistic images to people.
Second we summarize studies on explaining recommender systems.
Both are important to our problem at hand. e nal paragraph
in this section highlights the dierences to previous work and our
contributions to the existing literature in the area.</p>
      <sec id="sec-2-1">
        <title>Recommendations of Artistic Images. e works of Aroyo</title>
        <p>
          et al. [
          <xref ref-type="bibr" rid="ref2">2</xref>
          ] with the CHIP project and Semeraro et al. [
          <xref ref-type="bibr" rid="ref22">22</xref>
          ] with
FIRSt (Folksonomy-based Item Recommender syStem) made early
contributions to this area using traditional techniques. More
complex methods were implemented recently by Benouaret et al. [
          <xref ref-type="bibr" rid="ref3">3</xref>
          ],
using context obtained through a mobile application, that makes
a museum tour recommendation. Finally, the work of He et al.
addresses digital artwork recommendations based on pre-trained
deep neural visual features [
          <xref ref-type="bibr" rid="ref12">12</xref>
          ], and the work of Dominguez et
al. [
          <xref ref-type="bibr" rid="ref7">7</xref>
          ] and Messina et al. [
          <xref ref-type="bibr" rid="ref18">18</xref>
          ] compared neural against traditional
visual features. None of the aforementioned works performed a
user study under explanation interfaces to generalize their results.
        </p>
        <p>
          Explaining Recommender Systems. ere are some related
works on explanations for recommender systems [
          <xref ref-type="bibr" rid="ref24">24</xref>
          ]. ough a
good amount of research has been published in the area, to the best
of our knowledge, no previous research has conducted a user study
to understand the eect of explaining recommendation of artwork
images based on dierent visual features. e closest works in
this aspect are researches oriented to automatically add caption to
images [
          <xref ref-type="bibr" rid="ref19 ref9">9, 19</xref>
          ] or to explain image classications [
          <xref ref-type="bibr" rid="ref13">13</xref>
          ], but they are
not directly related to personalized recommender systems.
        </p>
      </sec>
      <sec id="sec-2-2">
        <title>Dierences to Previous Research &amp; Contributions. Although</title>
        <p>we focus on artistic images, to the best of our knowledge this is
the rst work which studies the eect of explaining
recommendations of images based on visual features. Our contributions are
two-fold: i) we analyze and report the positive eect of explaining
artistic recommendations especially for the VCBR based on neural
features, and ii) by a user study we validate o-line results stating
the superiority of neural visual features compared to aractiveness
visual features over several dimensions, such as users’ perception
of explainability, relevance, trust and general satisfaction.
3</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>METHODS</title>
      <p>In the following section we describe in detail our study methods.
First, we introduce the dataset chosen for the purpose of our study.
Second we introduce the three dierent explainable visual interfaces
implemented which we evaluate. ird the two algorithms chosen
for our study are revealed. Finally, the user study procedure is
explained.
3.1</p>
    </sec>
    <sec id="sec-4">
      <title>Materials</title>
      <p>
        For the purpose of our study we rely on a dataset provided by the
online web store UGallery, which has been selling artwork for more
than 10 years [
        <xref ref-type="bibr" rid="ref27">27</xref>
        ]. ey support emergent artists by helping them
sell their artwork online. For our research, UGallery provided us
with an anonymized dataset of 1,371 users, 3,490 items and 2,846
purchases (transactions) of artistic artifacts, where all users have
made at least one transaction. On average, each user bought 2-3
items over recent years .
3.2
      </p>
      <p>e Explainable Recommender Interfaces
In our study we explore the eect of explanations in visual
contentbased artwork recommender systems. As such, our study contains
conditions depending on how recommendations are displayed: i)
no explanations, as shown in Figure 1, ii) explanations given by
text and based on the top-3 most similar images a user liked in the
past, as shown in Figure 2, and iii ) explanations employing a visual
aractiveness bar chart and showing the most similar image of the
user’s item prole, as presented in Figure 3.</p>
      <p>In all three cases the interfaces are vertically scrollable. While
Interface 1 (baseline) is able to show 5 images in a row at the same
time, interfaces 2 and 3 are capable of showing one recommended
image at the same time in one row to the user.
3.3</p>
    </sec>
    <sec id="sec-5">
      <title>Visual Recommendation Approaches</title>
      <p>As mentioned earlier in this paper, we make use of two
dierent content-based visual recommender approaches in our work.
e reason for choosing content-based methods over collaborative
ltering-based methods is grounded in the fact that once an item is
sold via the UGallery store, it is not available anymore (every item
preference
elicitation</p>
      <p>Algorithm: Within subjects
(repeated measures)</p>
      <p>Swap order of
algorithm randomly</p>
      <p>DNN AVF
Interface 1 No explanation No explanation
Interface 2 Explanation based on Explanation based on top3 Interface:</p>
      <p>top3 similar images similar images Bsuetbwjeecetsn
Interface 3 Explanation based on Explanation based on</p>
      <p>top3 similar images barchart of visual features
pre-study
survey
post-DNN
survey
post-AVF
survey
is unique) and hence traditional collaborative ltering approaches
do not apply.</p>
      <p>
        DNN Visual Feature (DNN) Algorithm. e rst
algorithmic approach we employed was based on image similarity, itself
based on features extracted with a deep neural network. e output
vector representing the image is usually called an image’s visual
embedding. e visual embedding in our experiment was a vector
of features obtained from an AlexNet, a convolutional deep neural
network developed to classify images [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ]. In particular, we use an
AlexNet model pre-trained with the ImageNet dataset [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ]. Using
the pre-trained weights, for every image a vector of 4,096
dimensions was generated with the Cae (hp://cae.berkeleyvision.org/)
framework. We resized every image to a 227x 227 image. is is the
standard pre-processing needed to use the AlexNet.
      </p>
      <p>
        Attractiveness Visual Features (AVF) Algorithm. e
second content-based algorithmic recommender approach employed
was a method based on visual aractiveness features. San Pedro
and Siersdorfer in [
        <xref ref-type="bibr" rid="ref21">21</xref>
        ] proposed several explainable visual features
that to a great extent, can capture the aractiveness of an image
posted on Flickr. Following their procedure, for every image in
our UGallery dataset we calculated: (a) average brightness, (b)
saturation, (c) sharpness, (d) RMS-contrast, (e) colorfulness and (f)
naturalness. In addition, we added (g) entropy, which is a good way
to characterize and measure the texture of an image [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ]. ese
metrics have also been used in another study [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ], where we show
how to nudge people with aractive images to take up more healthy
recipe recommendations. To compute these features, we used the
original size of the images and did not pre-process them.
      </p>
      <p>
        Due space constrains, the details to calculate the features are
described in the article by Messina et al. [
        <xref ref-type="bibr" rid="ref18">18</xref>
        ]
      </p>
      <p>
        Computing Recommendations. Given a user u who has
consumed a set of artworks Pu , a constrained prole size K , and an
arbitrary artwork i from the inventory, the score of this item i to
be recommended to u is:
score u; i X
minrK;¶Pu ¶x
&lt;
r 1
max r rsim ViX ; VjX x
jϵ Pu
minrK; ¶Pu ¶x
;
(1)
where VzX is a feature vector of item z obtained with method X ,
where X can be either a pre-trained AlexNet (DNN) or aractiveness
visual features (AVF). max r denotes the r -th maximum value,
e.g., if r = 1 it is the overall maximum, if r = 2 it is the second
maximum, and so on. We compute the average similarity of the
top-K most similar images because as shown in Messina et al. [
        <xref ref-type="bibr" rid="ref18">18</xref>
        ],
for dierent users, the recommendations match beer using smaller
subsets of the entire user prole. Users do not always look to buy a
painting similar to one they bought before, but they look for one
that resembles a set of artworks that they liked. sim Vi ; Vj denotes
a similarity function between vectors Vi and Vj . In this particular
case, the similarity function used was cosine similarity:
sim Vi ; Vj
cos Vi ; Vj
      </p>
      <p>Vi Vj
½Vi ½½Vj ½
(2)</p>
      <p>
        Both methods use the same formula to calculate the
recommendations. e dierence is in the origin of the visual features. For
the DNN method, the features were extracted with the AlexNet
[
        <xref ref-type="bibr" rid="ref15">15</xref>
        ], and in the case of AVF, the features were extracted based on
San Pedro et al. [
        <xref ref-type="bibr" rid="ref21">21</xref>
        ].
3.4
      </p>
    </sec>
    <sec id="sec-6">
      <title>User Study Procedure</title>
      <p>To evaluate the performance of our explainable interfaces we
conducted a user study in Amazon Mechanical Turk using a 3x2 mixed
design: 3 interfaces (between-subjects) and 2 algorithms
(withinsubjects, DNN and AVF). e interface conditions were: Interface
1: interface without explanations, as in Figure 1; Interface 2: each
item recommendation is explained based on the top 3 most similar
images in the user prole, as in Figure 2; and Interface 3: only for
AVF, based on a bar chart of visual features, as in Figure 3. Notice
that in the condition Interface 3, for DNN we used the explanation
based on top 3 most similar images, because the neural embedding
of 4,096 dimensions has no human-interpretable features to show
in a bar chart.</p>
      <p>To compute the recommendations for each of the three interface
conditions two recommender algorithms were chosen: one based
on DNN visual features, and the other based on aractiveness visual
features (AVF). e order in which the algorithms were presented
was chosen at random to diminish the chance of a learning eect.</p>
      <p>
        e full study procedure is shown in Figure 4. Participants
accepted the study on Mechanical Turk (hps://www.mturk.com)
and were redirected to a web application. Aer accepting a consent
form, they are redirected to the pre-study survey, which collects
demographic data (age, gender) and a subject’s previous knowledge
of art based on the test by Chaerjee et al. [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ].
      </p>
      <p>Following this, they had to perform a preference elicitation task.
In this step, the users had to “like” at least ten paintings, using
a Pinterest-like interface. Next, they were randomly assigned to
one interface condition. In each condition, they again provided
feedback (rating with 1-5 scale to each image) to top ten
recommendations of images with employing either the DNN or the AVF
algorithm (also assigned at random as discussed before). Finally,
the participants were asked to next answer a post-algorithm survey.
e dimensions evaluated in the post-algorithm survey are the
same for DNN and AVF algorithms, and they are shown in Table
1. is process is repeated for the second algorithm as well. Once
the participants nished answering the second post study survey,
they were redirected to the nal view, where they received a survey
code for later payment in Amazon Mechanical Turk.
4</p>
      <p>RESULTS
e study was nished by in total 200 users out of which 121 were
able to answer our validation questions successfully and hence were
included in the results. In total, we had two validation questions set
to check for aention of our study participants. Filtering out users
not responding properly to these questions allowed us to include 41
users for the Interface 1 condition, 41 users for Interface 2 condition
and 39 users for Interface 3 condition. In total, participants were
paid an amount of 0.40 USD per study, which took them around 10
minutes to complete.</p>
      <p>Our subjects were between 18 to over 60 years old. 36% were
between 25 to 32 years old, and 29% between 32 to 40 years old.
Females made up 55.4% . 12% just nished high school, 31% had
a some college degree, 57% had a bachelor’s, master’s or Ph.D.
degree. Only 8% reported some visual impairment. W.r.t. their
understanding about art, 20% had null experience, 48% had aended
1 or 2 lessons, and 32% reported to have aended 3 or more at high
school level or above. 20% of our subjects also reported that they
had almost never visited a museum or an art gallery; 36% do this
once a year; and 44% do this once every 1 or 6 months.</p>
      <p>Dierences between Interfaces. Table 2 summarizes the
results of the user study. First we compared interface performance
and then we looked at the algorithmic performance. e explainable
interfaces (Interface 2 and 3) signicantly improved the perception
of explainability compared to Interface 1 under both algorithms.
ere is also a signicant improvement over Interface 1 in terms
of relevance and diversity, but this is only achieved by the DNN
method when this is compared against the AVF method using the
interface 3. Interestingly, this is the condition where the interface
is more transparent, since it explains exactly what is used to
recommend (brightness, saturation, sharpness, etc.). People report that
they understand why the images are recommended (70.4), but since
the relevance is rather insucient (56.2), the perception of trust is
reported as low (55.4).</p>
      <sec id="sec-6-1">
        <title>Dierences between Algorithms. With the only exception of</title>
        <p>the dimension Diverse where AVF was signicantly beer, DNN
was perceived more positively than AVF at large. In interfaces
2 and 3, the DNN method was perceived signicantly beer in 5
dimensions (explainability, relevance, interface satisfaction, interest
for eventual use, and trust), as well as higher average rating.</p>
        <p>Overall, the results indicate that the explainable interface based
on top 3 similar images works beer than an interface without
explanation. Moreover, this eect is enhanced by the accuracy of
the algorithm, so even if the algorithm has no explainable features
(DNN) it could induce more trust if the user perceives a larger
predictive preference accuracy.</p>
      </sec>
    </sec>
    <sec id="sec-7">
      <title>5 CONCLUSIONS &amp; FUTURE WORK</title>
      <p>In this paper, we have studied the eect of explaining
recommendation of images employing three dierent recommender interfaces,
as well as interactions with two dierent visual content-based
recommendation algorithms: one with high predictive accuracy but
with unexplainable features (DNN), and another with lower
accuracy but with higher potential for explainable features (AVF).</p>
      <p>e rst result, which answers RQ1, shows that explaining the
images recommended has a positive eect vs. no explanation.
Moreover, the explanation based on top 3 similar images presents the
best results, but we need to consider that the alternative method,
explanations based on visual features, was only used with the AVF.
is result is preliminary and opens a path of research in terms of
new interfaces which could help to explain the features learned by
a deep neural network of images.</p>
      <p>Regarding RQ2, we see that the algorithm used plays an
important role in conjunction with the interface. DNN is perceived
beer than AVF in most dimensions evaluated, showing that further
research should focus on the interaction between algorithm and
explainable interfaces. In the future we will expand this work to
other datasets, beyond artistic images, to generalize our results.
6</p>
    </sec>
    <sec id="sec-8">
      <title>ACKNOWLEDGEMENTS</title>
      <p>e authors from PUC Chile were funded by Conicyt, Fondecyt
grant 11150783, as well as by the Millennium Institute for
Foundational Research on Data (IMFD).</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>Xavier</given-names>
            <surname>Amatriain</surname>
          </string-name>
          .
          <year>2013</year>
          .
          <article-title>Mining large streams of user data for personalized recommendations</article-title>
          .
          <source>ACM SIGKDD Explorations Newsleer 14</source>
          ,
          <issue>2</issue>
          (
          <year>2013</year>
          ),
          <fpage>37</fpage>
          -
          <lpage>48</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>LM</given-names>
            <surname>Aroyo</surname>
          </string-name>
          ,
          <string-name>
            <surname>Y Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R</given-names>
            <surname>Brussee</surname>
          </string-name>
          , Peter Gorgels, LW Rutledge, and
          <string-name>
            <given-names>N</given-names>
            <surname>Stash</surname>
          </string-name>
          .
          <year>2007</year>
          .
          <article-title>Personalized museum experience: e Rijksmuseum use case</article-title>
          .
          <source>In Proceedings of Museums and the Web.</source>
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>Idir</given-names>
            <surname>Benouaret</surname>
          </string-name>
          and
          <string-name>
            <given-names>Dominique</given-names>
            <surname>Lenne</surname>
          </string-name>
          .
          <year>2015</year>
          .
          <article-title>Personalizing the Museum Experience through Context-Aware Recommendations</article-title>
          .
          <source>In Systems, Man, and Cybernetics (SMC)</source>
          ,
          <source>2015 IEEE International Conference on. IEEE</source>
          ,
          <fpage>743</fpage>
          -
          <lpage>748</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>Oscar</given-names>
            <surname>Celma</surname>
          </string-name>
          .
          <year>2010</year>
          .
          <article-title>Music recommendation</article-title>
          .
          <source>In Music Recommendation and Discovery</source>
          . Springer,
          <fpage>43</fpage>
          -
          <lpage>85</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>Anjan</given-names>
            <surname>Cha</surname>
          </string-name>
          erjee, Page Widick, Rebecca Sternschein,
          <string-name>
            <given-names>William</given-names>
            <surname>Smith</surname>
          </string-name>
          <string-name>
            <surname>II</surname>
          </string-name>
          , and
          <string-name>
            <given-names>Bianca</given-names>
            <surname>Bromberger</surname>
          </string-name>
          .
          <year>2010</year>
          .
          <article-title>e Assessment of Art Aributes</article-title>
          .
          <volume>28</volume>
          (
          <issue>07</issue>
          <year>2010</year>
          ),
          <fpage>207</fpage>
          -
          <lpage>222</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>Jia</given-names>
            <surname>Deng</surname>
          </string-name>
          , Wei Dong, Richard Socher,
          <string-name>
            <surname>Li-Jia</surname>
            <given-names>Li</given-names>
          </string-name>
          ,
          <string-name>
            <given-names>Kai</given-names>
            <surname>Li</surname>
          </string-name>
          , and
          <string-name>
            <surname>Li</surname>
          </string-name>
          Fei-Fei.
          <year>2009</year>
          .
          <article-title>Imagenet: A large-scale hierarchical image database</article-title>
          .
          <source>In Computer Vision and Paern Recognition</source>
          ,
          <year>2009</year>
          .
          <article-title>CVPR 2009</article-title>
          .
          <article-title>IEEE Conference on</article-title>
          . IEEE,
          <fpage>248</fpage>
          -
          <lpage>255</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>Vicente</given-names>
            <surname>Dominguez</surname>
          </string-name>
          , Pablo Messina, Denis Parra, Domingo Mery, Christoph Traner, and
          <string-name>
            <given-names>Alvaro</given-names>
            <surname>Soto</surname>
          </string-name>
          .
          <year>2017</year>
          .
          <article-title>Comparing Neural and Aractiveness-based Visual Features for Artwork Recommendation</article-title>
          .
          <source>In Proceedings of the Workshop on Deep Learning for Recommender Systems, co-located at RecSys</source>
          <year>2017</year>
          . DOI: http://dx.doi.org/10.1145/3125486.3125495 arXiv:arXiv:
          <fpage>1706</fpage>
          .
          <fpage>07515</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>David</given-names>
            <surname>Elsweiler</surname>
          </string-name>
          , Christoph Traner, and Morgan Harvey.
          <year>2017</year>
          .
          <article-title>Exploiting food choice biases for healthier recipe recommendation</article-title>
          .
          <source>In Proceedings of the 40th international acm sigir conference on research and development in information retrieval. ACM</source>
          ,
          <volume>575</volume>
          -
          <fpage>584</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>Hao</given-names>
            <surname>Fang</surname>
          </string-name>
          , Saurabh Gupta, Forrest Iandola, Rupesh Srivastava, Li Deng, Piotr Dolla´r, Jianfeng Gao, Xiaodong He, Margaret Mitchell, John Pla, and others.
          <source>2015</source>
          .
          <article-title>From captions to visual concepts and back</article-title>
          . (
          <year>2015</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <surname>Carlos</surname>
            <given-names>A</given-names>
          </string-name>
          <string-name>
            <surname>Gomez-Uribe</surname>
            and
            <given-names>Neil</given-names>
          </string-name>
          <string-name>
            <surname>Hunt</surname>
          </string-name>
          .
          <year>2016</year>
          .
          <article-title>e netix recommender system: Algorithms, business value, and innovation</article-title>
          .
          <source>ACM Transactions on Management Information Systems (TMIS) 6</source>
          ,
          <issue>4</issue>
          (
          <year>2016</year>
          ),
          <fpage>13</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <surname>Rafael</surname>
            <given-names>C Gonzalez</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Steven L Eddins</surname>
          </string-name>
          , and Richard E Woods.
          <year>2004</year>
          .
          <article-title>Digital Image Publishing Using MATLAB</article-title>
          . Prentice Hall.
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <surname>Ruining</surname>
            <given-names>He</given-names>
          </string-name>
          , Chen Fang,
          <string-name>
            <given-names>Zhaowen</given-names>
            <surname>Wang</surname>
          </string-name>
          , and
          <string-name>
            <surname>Julian McAuley</surname>
          </string-name>
          .
          <year>2016</year>
          .
          <article-title>Vista: A Visually, Socially, and Temporally-aware Model for Artistic Recommendation</article-title>
          .
          <source>In Proceedings of the 10th ACM Conference on Recommender Systems (RecSys '16)</source>
          . ACM, New York, NY, USA,
          <fpage>309</fpage>
          -
          <lpage>316</lpage>
          . DOI:http://dx.doi.org/10.1145/2959100. 2959152
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <given-names>Lisa</given-names>
            <surname>Anne</surname>
          </string-name>
          <string-name>
            <surname>Hendricks</surname>
          </string-name>
          , Zeynep Akata, Marcus Rohrbach, Je Donahue,
          <string-name>
            <given-names>Bernt</given-names>
            <surname>Schiele</surname>
          </string-name>
          , and
          <string-name>
            <given-names>Trevor</given-names>
            <surname>Darrell</surname>
          </string-name>
          .
          <year>2016</year>
          .
          <article-title>Generating visual explanations</article-title>
          .
          <source>In European Conference on Computer Vision</source>
          . Springer,
          <fpage>3</fpage>
          -
          <lpage>19</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <surname>Joseph</surname>
            <given-names>A</given-names>
          </string-name>
          <string-name>
            <surname>Konstan and John Riedl</surname>
          </string-name>
          .
          <year>2012</year>
          .
          <article-title>Recommender systems: from algorithms to user experience</article-title>
          .
          <source>User Modeling and User-Adapted Interaction 22</source>
          ,
          <fpage>1</fpage>
          -
          <lpage>2</lpage>
          (
          <year>2012</year>
          ),
          <fpage>101</fpage>
          -
          <lpage>123</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <surname>Alex</surname>
            <given-names>Krizhevsky</given-names>
          </string-name>
          , Ilya Sutskever, and Georey
          <string-name>
            <given-names>E</given-names>
            <surname>Hinton</surname>
          </string-name>
          .
          <year>2012</year>
          .
          <article-title>Imagenet classication with deep convolutional neural networks</article-title>
          .
          <source>In Advances in neural information processing systems</source>
          .
          <volume>1097</volume>
          -
          <fpage>1105</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [16]
          <string-name>
            <surname>Pa</surname>
          </string-name>
          <article-title>ie Maes and others</article-title>
          .
          <source>1994</source>
          .
          <article-title>Agents that reduce work and information overload</article-title>
          .
          <source>Commun. ACM 37</source>
          ,
          <issue>7</issue>
          (
          <year>1994</year>
          ),
          <fpage>30</fpage>
          -
          <lpage>40</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [17]
          <string-name>
            <surname>Sean</surname>
            <given-names>M McNee</given-names>
          </string-name>
          ,
          <string-name>
            <given-names>Nishikant</given-names>
            <surname>Kapoor</surname>
          </string-name>
          , and Joseph A Konstan.
          <year>2006</year>
          .
          <article-title>Don't look stupid: avoiding pitfalls when recommending research papers</article-title>
          .
          <source>In Proceedings of the 2006 20th anniversary conference on Computer supported cooperative work. ACM</source>
          ,
          <volume>171</volume>
          -
          <fpage>180</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          [18]
          <string-name>
            <surname>Pablo</surname>
            <given-names>Messina</given-names>
          </string-name>
          , Vicente Dominguez, Denis Parra, Christoph Traner, and
          <string-name>
            <given-names>Alvaro</given-names>
            <surname>Soto</surname>
          </string-name>
          .
          <year>2018</year>
          .
          <article-title>Content-Based Artwork Recommendation: Integrating Painting Metadata with Neural and Manually-Engineered Visual Features. User Modeling</article-title>
          and
          <string-name>
            <surname>User-Adapted Interaction</surname>
          </string-name>
          (
          <year>2018</year>
          ). DOI:http://dx.doi.org/10.1007/ s11257-018-9206-9
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          [19] Margaret Mitchell, Xufeng Han, Jesse Dodge, Alyssa Mensch, Amit Goyal, Alex Berg, Kota Yamaguchi, Tamara Berg, Karl Stratos, and Hal Daume´, III.
          <year>2012</year>
          .
          <article-title>Midge: Generating Image Descriptions from Computer Vision Detections</article-title>
          .
          <source>In Proceedings of the 13th Conference of the European Chapter of the Association for Computational Linguistics (EACL '12)</source>
          .
          <article-title>Association for Computational Linguistics</article-title>
          , Stroudsburg, PA, USA,
          <fpage>747</fpage>
          -
          <lpage>756</lpage>
          . http://dl.acm.org/citation.cfm?id=
          <volume>2380816</volume>
          .
          <fpage>2380907</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          [20]
          <string-name>
            <surname>Chris</surname>
            <given-names>Olah</given-names>
          </string-name>
          , Alexander Mordvintsev, and
          <string-name>
            <given-names>Ludwig</given-names>
            <surname>Schubert</surname>
          </string-name>
          .
          <year>2017</year>
          .
          <string-name>
            <given-names>Feature</given-names>
            <surname>Visualization</surname>
          </string-name>
          .
          <source>Distill</source>
          (
          <year>2017</year>
          ). DOI:http://dx.doi.org/10.23915/distill.00007 hps://distill.pub/2017/feature-visualization.
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          [21] Jose San Pedro and
          <string-name>
            <given-names>Stefan</given-names>
            <surname>Siersdorfer</surname>
          </string-name>
          .
          <year>2009</year>
          .
          <article-title>Ranking and Classifying Aractiveness of Photos in Folksonomies</article-title>
          .
          <source>In Proceedings of the 18th International Conference on World Wide Web (WWW '09)</source>
          . ACM, New York, NY, USA,
          <fpage>771</fpage>
          -
          <lpage>780</lpage>
          . DOI:http://dx.doi.org/10.1145/1526709.1526813
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          [22]
          <string-name>
            <surname>Giovanni</surname>
            <given-names>Semeraro</given-names>
          </string-name>
          , Pasquale Lops, Marco De Gemmis, Cataldo Musto, and
          <string-name>
            <given-names>Fedelucio</given-names>
            <surname>Narducci</surname>
          </string-name>
          .
          <year>2012</year>
          .
          <article-title>A folksonomy-based recommender system for personalized access to digital artworks</article-title>
          .
          <source>Journal on Computing and Cultural Heritage (JOCCH) 5</source>
          ,
          <issue>3</issue>
          (
          <year>2012</year>
          ),
          <fpage>11</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          [23]
          <string-name>
            <given-names>Ali</given-names>
            <surname>Sharif</surname>
          </string-name>
          <string-name>
            <surname>Razavian</surname>
          </string-name>
          , Hossein Azizpour, Josephine Sullivan, and
          <string-name>
            <surname>Stefan Carlsson.</surname>
          </string-name>
          <year>2014</year>
          .
          <article-title>CNN features o-the-shelf: an astounding baseline for recognition</article-title>
          .
          <source>In Proceedings of the IEEE Conference on Computer Vision</source>
          and Pa
          <source>ern Recognition Workshops</source>
          .
          <fpage>806</fpage>
          -
          <lpage>813</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          [24]
          <string-name>
            <given-names>Nava</given-names>
            <surname>Tintarev</surname>
          </string-name>
          and Judith Mastho.
          <year>2015</year>
          .
          <article-title>Explaining recommendations: Design and evaluation</article-title>
          .
          <source>In Recommender Systems Handbook</source>
          . Springer,
          <fpage>353</fpage>
          -
          <lpage>382</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref25">
        <mixed-citation>
          [25] Christoph Traner, Alexander Oberegger, Lukas Eberhard, Denis Parra, Leandro Marinho, and others.
          <source>2016</source>
          .
          <article-title>Understanding the Impact of Weather for POI Recommendations</article-title>
          .
          <source>Proceedings of RecTour Workshop</source>
          , co-located at ACM RecSys (
          <year>2016</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref26">
        <mixed-citation>
          [26]
          <string-name>
            <surname>Egon L van den</surname>
            <given-names>Broek</given-names>
          </string-name>
          , ijs Kok, eo
          <string-name>
            <given-names>E</given-names>
            <surname>Schouten</surname>
          </string-name>
          , and
          <string-name>
            <given-names>Eduard</given-names>
            <surname>Hoenkamp</surname>
          </string-name>
          .
          <year>2006</year>
          .
          <article-title>Multimedia for art retrieval (m4art)</article-title>
          .
          <source>In Multimedia Content Analysis, Management, and Retrieval</source>
          <year>2006</year>
          , Vol.
          <volume>6073</volume>
          . International Society for Optics and Photonics,
          <year>60730Z</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref27">
        <mixed-citation>
          [27]
          <string-name>
            <given-names>Deborah</given-names>
            <surname>Weinswig</surname>
          </string-name>
          .
          <year>2016</year>
          .
          <article-title>Art Market Cooling, But Online Sales Booming</article-title>
          . https://www.forbes.com/sites/deborahweinswig/2016/05/13/art-market
          <article-title>-coo ling-but-online-sales-booming/</article-title>
          . (
          <year>2016</year>
          ). [Online; accessed 21-March-2017].
        </mixed-citation>
      </ref>
      <ref id="ref28">
        <mixed-citation>
          [28]
          <string-name>
            <surname>Mao</surname>
            <given-names>Ye</given-names>
          </string-name>
          , Peifeng Yin,
          <string-name>
            <surname>Wang-Chien Lee</surname>
          </string-name>
          , and
          <string-name>
            <surname>Dik-Lun Lee</surname>
          </string-name>
          .
          <year>2011</year>
          .
          <article-title>Exploiting geographical inuence for collaborative point-of-interest recommendation</article-title>
          .
          <source>In Proceedings of the 34th international ACM SIGIR conference on Research and development in Information Retrieval. ACM</source>
          ,
          <volume>325</volume>
          -
          <fpage>334</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref29">
        <mixed-citation>
          [29]
          <article-title>an Yuan, Gao Cong</article-title>
          , Zongyang Ma, Aixin Sun, and Nadia Magnenat almann.
          <year>2013</year>
          .
          <article-title>Time-aware point-of-interest recommendation</article-title>
          .
          <source>In Proceedings of the 36th international ACM SIGIR conference on Research and development in information retrieval. ACM</source>
          ,
          <volume>363</volume>
          -
          <fpage>372</fpage>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>