<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <journal-title-group>
        <journal-title>ICTAI.</journal-title>
      </journal-title-group>
    </journal-meta>
    <article-meta>
      <title-group>
        <article-title>Your Click Matters: Enhancing Click-based Image Retrieval performance through Collaborative Filtering</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Deepanwita Datta</string-name>
          <email>ddatta.rs.cse13@itbhu.ac.in</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Manajit Chakraborty</string-name>
          <email>chakrm@usi.ch</email>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Aveek Biswas</string-name>
          <email>a4biswas@ucsd.edu</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>NTNU</institution>
          ,
          <addr-line>Trondheim</addr-line>
          ,
          <country country="NO">Norway</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>UCSD</institution>
          ,
          <addr-line>California</addr-line>
          ,
          <country country="US">USA</country>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>Universita` della Svizzera italiana</institution>
          ,
          <addr-line>Lugano</addr-line>
          ,
          <country country="CH">Switzerland</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2015</year>
      </pub-date>
      <volume>45</volume>
      <fpage>234</fpage>
      <lpage>241</lpage>
      <abstract>
        <p>Image retrieval has been an active research area since the early days of computing. While ensemble, multimodal and hybrid methods coupled with machine learning has seen an upward surge replacing unimodal, heuristic-based methods; a rather new offshoot has been to identify new features associated with images on the web. One such feature is the 'click count' based on the clicks an image or its corresponding text gets in response to a query. Previous state-of-the-art methods have tried to exploit this feature by using its raw count and machine learning. In this paper, we build on this idea and propose a new collaborative filtering based technique to employ the click-log of users from the web to better identify and associate images in response to either a text or an image query. Experiments performed on a large scale publicly available standard dataset having genuine click logs from actual users corroborate the efficacy and significant increase in efficiency of our approach.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1 Introduction</title>
      <p>
        Cross-media retrieval has proven to be an effective
solution to search through enormous multi-varied
datasets. A fairly common example of such
crossmedia retrieval is when an image is searched
using a text query. Here the textual description of
the image content acts as the text query. However,
it is not a trivial task to illustrate a non-textual
visual content using only text. Hence, a semantic
gap is introduced between the user needs and the
given (existing) descriptions. Although in
literature, a significant amount of work has been
devoted to correlating textual and visual
information to bridge the semantic gap between the
highlevel information needs of users and commonly
employed low-level features, it continues to be a
major challenge. The existing state-of-the-art
solutions to this challenge is two-pronged. Few of
the existing works
        <xref ref-type="bibr" rid="ref1 ref9">(Feng et al., 2014; Zhen and
Yeung, 2012)</xref>
        stress on learning mapping functions
whereas rest of the works explore the high-level
semantic representation of modalities
        <xref ref-type="bibr" rid="ref3 ref6 ref7">(Karpathy
and Fei-Fei, 2015; Reed et al., 2016)</xref>
        . Among
these, semantic representation based approaches
and deep learning based approaches have gained
reasonable success. Deep Convolutional Neural
Networks (CNNs) are used to learn the latent
features and these learned features are utilized as
visual and textual semantic representation in these
models.
      </p>
      <p>
        In ACM Multimedia 2015 MSR-Bing Image
Retrieval Challenge1, it was argued that the massive
amount of click data from commercial search
engines provide a data set that is unique in bridging
the semantic and intent gap. Millions of click data
i.e. clicked image-query pairs, generated from
search engines, were collected and released
publicly as a new large-scale real-world image click
data (Clickture) to investigate how to effectively
leverage this click-count based data to mitigate the
semantic gap. This click data is stored in a large
table with multiple rows and three tuples (I, Q, C),
indicating that the image I was clicked C times
against the search results of a given textual query
Q. Wu et al.
        <xref ref-type="bibr" rid="ref8">(Wu et al., 2016a)</xref>
        view the entire
dataset as a bipartite graph which has two types
of vertices, queries and images respectively, and
the edge’s weight is assigned according to the
total number of clicks from all the users. The
authors learn a common representation for both
image and text query from the perspective of
encod
      </p>
      <sec id="sec-1-1">
        <title>1http://press.liacs.nl/mmgrand/microsoft.pdf</title>
        <p>ing the explicit/implicit relevance relationship
between the vertices in the click graph. The common
representation is obtained as well as any unseen
query or image is dealt with by reducing the
truncated random walk loss and the distance between
the learned representation of vertices and their
corresponding deep neural network output.</p>
        <p>Thus, the relevance relationship between a text
query and its resulting image is obtained purely by
measuring the distance between their
corresponding learned high-level feature representation. In
other words, the cross-modal retrieval on the
unseen queries and images is dealt with here in a
content-based fashion. As these content-based
systems operate solely on feature representations,
the definition of similarity in these systems is
frequently ad-hoc and not explicitly optimized or
generalized for any particular task i.e. here
crossmodal retrieval. Frequently, the optimization of
similarity for ranking affects the quantity of
interest. Thus the retrieved items often become
coarsely abstracted or potentially irrelevant. To
overcome this shortfall, we try to capture relevant
similarity information expressed by collaborative
filtering.</p>
        <p>
          The motivation behind using collaborative
filtering (CF) is that this method produces user-specific
recommendations of items based on patterns of
ratings or usage without the need for exogenous
information about either items or users
          <xref ref-type="bibr" rid="ref3 ref7">(Koren and
Bell, 2015)</xref>
          . Recommendation by collaborative
filtering relies on explicit or implicit feedback from
the user or indirectly obtained through observing
user behavior respectively. In our scenario, for any
unseen query, aside from relying on the similarity
score of the learned features, we can exploit some
previous knowledge i.e. implicit feedback of the
user. We consider click counts as the implicit
feedback from the user. Stemming from this
observation, we predict the similarity structure encoded
by collaborative filtering data. Finally, we use
collaborative filtering for generating a ranked list for
cross-modal retrieval. To the best of our
knowledge, ours is the first attempt in using
collaborative filtering in conjunction with Deep Learning
towards cross-modal image retrieval. A rigorous
experiment is carried out over the Clickture dataset
and the experimental results validate our claim.
Our method outperforms the current state of the
art on the learning and content-based methods.
A fundamental image retrieval technique is to
search for images by textual queries. The
conventional image search engines leverage the
benefits of associated or surrounding text which are
generally collected from the data like image
captions, tags, comments etc. To train such systems
by labeled text-image pairs human intervention is
necessary. However, such human labeling is
expensive, time-consuming and quite cumbersome.
These labeled data is unreliable as quite often they
suffer from noise. Expressing an image entirely
through a concise set of keywords, keyphrases
or free text is a non-trivial task even for humans
let alone a system. The problem is compounded
when the user has limited to no knowledge of how
the search system or IR work. To alleviate these
problems of cross-view learning, the use of
clickthrough data have gained momentum
          <xref ref-type="bibr" rid="ref5">(Pan et al.,
2014)</xref>
          . Cross-View learning creates a common
latent subspace where the data from different
modalities like text, image etc., can be compared with
each other easily.
        </p>
        <p>
          On the other hand, the click-through data is
available in huge amount and is relatively easy to
access. Also, this click-through data helps in
better understanding a query. In the work by Pan et
al.
          <xref ref-type="bibr" rid="ref5">(Pan et al., 2014)</xref>
          , the distance between
mappings of query and image in the latent subspace
is reduced and the inherent structure is preserved
back to each original space. Once the mapping is
done, and the latent representations are acquired,
the next step is to compute the distance among
these representations. Hence, choosing an
appropriate similarity function becomes crucial as
it is the key to make the cross-modal similarity
tractable. He et al.
          <xref ref-type="bibr" rid="ref2">(He et al., 2016)</xref>
          propose
a deep and bidirectional representation learning
model to address the issue of imagetext
crossmodal retrieval. The authors adopt two
convolutional deep neural networks to extract
semantic representation from both raw image and text
data and calculate cosine distance among those.
Subsequently, a bidirectional network is learned
from the matched and unmatched image-text pairs
for training to capture the property of the
crossmodal retrieval. This learning framework uses
maximum likelihood criterion and optimizes the
network through backpropagation and stochastic
gradient descent.
        </p>
        <p>
          Similarly, in the paper by Wang et al.
          <xref ref-type="bibr" rid="ref7">(Wang et al.,
2015)</xref>
          , the authors propose a supervised
framework based on a deep neural network which
captures the intra-modal and inter-modal relationships
efficiently. The proposed model requires only a
little prior knowledge to exploring high-level
semantic correlation and also it can tackle the
situation if any modality is missing. While most of
the recent works focus on learning semantic
representation, the work by Wu et al.
          <xref ref-type="bibr" rid="ref8">(Wu et al.,
2016b)</xref>
          concentrates on distance metric learning
which is essential to improve similarity search for
content-based retrieval. Usually, single modal
distance metric learning methods suffer from some
critical issues such as choosing the dominant
feature from diverse feature representation, learning a
distance metric on the combined high-dimensional
feature space which is very time-consuming etc..
To overcome these issues, the authors proposed
a multi-modal distance metric learning scheme
called online multi-modal distance metric learning
(OMDML), which can learn an optimized distance
metric on each individual feature space and learn
to find an optimal combination of diverse types of
features.
        </p>
        <p>
          Our work in this paper adopts the approach of
using convolutional neural networks as suggested by
He et al.
          <xref ref-type="bibr" rid="ref2">(He et al., 2016)</xref>
          . The reason for
choosing this over other learning methods is that while
training stage might take longer than some other
methods, CNN usually supersedes others when it
comes to the accuracy of learning. It should be
noted that we have modified the settings of CNN
to fit our problem and adapted it to our needs.
3
        </p>
      </sec>
    </sec>
    <sec id="sec-2">
      <title>Methodology</title>
      <p>Our model consists of two phases, training and
testing, as is the case with any learning based
technique. A click graph is used as labeled training
data where the number of click count is treated as
a label between a text query-image pair i.e. if any
click is present between any text query and image,
the image must be relevant to the text query. This
assumption stems from the fact that a user usually
clicks on an image against a text query only if she
finds it relevant and useful. Here, the number of
click counts reinforces how strongly relevant the
image is against the query or vice versa. Thus
a labeled query-image pair is learned through the
click count. In the testing phase, relevant images
from the test set are retrieved against any given
test query and ranked. Hence, we perform
crossmodal ranking over the new images and queries
that are not involved in the training click graph.
3.1</p>
      <sec id="sec-2-1">
        <title>Obtaining feature vector representations of query and documents</title>
        <p>Multimodal objects from different feature spaces
are present in the click graph. So, the first step of
our model consists of projecting a feature vector
representation of multimodal data into a common
dimensional space. Here, image and text are the
two different sources of information where the
dimension of an image depends on the pixel
intensity, and the number of pixels present in the
image and vocabulary size of bag-of-words denotes
the dimension of the text. Let us say if M is the
dimension of image feature and N is the
dimension of text feature then our objective is to come
up with a common latent vector space of
dimension D for both the image and the text.</p>
        <p>We obtain a common representation through
Convolutional Neural Network (CNN). Some
pretrained model such as inception-v3 model2 can be
used to learn the proper representation of input
data. The learned representations account for the
variations associated with the features. Any
established CNN model consists of layers like
convolutional filtering, local contrast normalization,
maxpooling and finally fully connected neural network
layers. For our model, we eliminate the last fully
connected layer of the CNN (the output layer) and
retrieve the image vectors from the penultimate
layer. This is done since we are not interested in
the classification of the images but instead in the
generated embeddings of the images. The raw
image features (embeddings) are directly fed into the
model to get a latent representation. However, the
text queries are represented in vector space model
(bag-of-words). So, all words present in the
vocabulary are inserted into a vector lookup table.
Finally, a D dimensional representation is learned
for each word from the lookup table. Test query
usually consists of multiple words. So, the entire
query can be represented by summing up all their
corresponding word vectors. The obtained latent
representations are used for learning purpose as
depicted in the next subsection.
3.2</p>
      </sec>
      <sec id="sec-2-2">
        <title>Learning from labeled click data</title>
        <p>In Recommendation System, Collaborative
Filtering (CF) models capture the interaction between
2https://www.kaggle.com/google-brain/inception-v3
users and item based on the rating. A rating
indicates the preference of an individual user towards
a particular item. High values of rating indicate a
stronger preference of the user towards the
particular item. The rating values are by nature either
implicit or explicit feedback provided by the user
or collected from user behavior or history.
Perceiving some resemblance with the inherent nature
of collaborative filtering with the characteristics
of our dataset, we hypothesize that learning from
clicked data through collaborating filtering may
increase the retrieval performance. In this
scenario, the text query and the images corresponding
to the query play the role of user and item
respectively. We treat click-count of each query-image
pair as an implicit rating and train our model from
these labeled query-image pairs and the rating
matrix.
3.3</p>
      </sec>
      <sec id="sec-2-3">
        <title>Prediction of click count for unseen query</title>
        <p>Model-based collaborative filter predicts users’
rating of unrated items. CF engines are more
versatile, in the sense that they can be applied to any
domain, and with some care could also provide
cross-domain recommendations. Also, CF works
best when the user space is large, which is the
case for image searching where thousands of users
are looking for images over the web every second.
Taking a cue from this fact, we choose a
modelbased collaborative approach to predict the click
count for an unseen query. The collaborative filter
tries to predict ratings or click counts by
characterizing both the query and image. Let us consider
that the learned latent vector for any query q is Vq
and the learned latent vector for any image i is Vi
such that for a given query q, R measure the
extent of relevance the query has with images that are
highly relevant. The interaction R between query
q and image i can be captured by the following
Equation 1:</p>
        <p>R = VqT Vi
where, the dot product between two vectors x, y ∈
Rf is defined as in Equation 2.
(1)
(2)
xT y =
f
X xkyk
k=1
Thus, the predicted click count becomes R which
is calculated using the Equation 1.
3.4</p>
      </sec>
      <sec id="sec-2-4">
        <title>Calculating similarity score</title>
        <p>By incorporating predicted click count between
the unseen query and each image in the dataset,
we calculate the similarity between each pair of
the images. Let us consider, the predicted click
count for the two images iu and iv against the nth
query qn as Rn,u and Rn,v respectively. Then the
similarity measure between any two images, Su,v,
can be calculated by using the following Equation
3:</p>
        <p>Pn∈I (Rn,u − R¯n)(Rn,v − R¯n)
Su,v = qP
n∈I (Rn,u − R¯n)2qPn∈I (Rn,v − R¯n)2
(3)
where R¯n is the average click-count for those
images and I denotes the entire image set. It can be
observed from the above equation that the
similarity measure depends on how much the click-count
for a pair of images deviates from the average
rating for those images. So, the similarity measure
is purely dependent on the predicted click-count.
As stated earlier, we calculate all the similarities
between each pair of the images present in the
dataset and based on the similarity score we rank
the images against the each query. Thus a final
ranked list is prepared and we select top-most
images as the most relevant retrieved ones.
4</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>Experimental Setup</title>
      <p>
        In this section we list the possible requirements for
the experiment. To run Convolutional Neural
Network for learning the latent representations over
the large set of images, we have used cloud
computing services of Google Cloud TPU 3 through 10
different instances. The learning is done by a
pretrained model through ImageNet4 i.e. inception-v3
model5. The last fully connected layer of the
Convoluted Neural Network, i.e. the penultimate layer
of the CNN, is extracted using TensorFlow6. The
dimension of all the learnt image vectors are kept
to 2048. The other libraries which aid this process
are NumPy7, SciPy8, scikit-learn9, pickle10 etc.
3https://cloud.google.com/tpu/
4www.image-net.org/
5https://cloud.google.com/tpu/docs/inception-v3advanced
6https://www.tensorflow.org/
7https://www.numpy.org/
8https://www.scipy.org/
9https://scikit-learn.org/stable/
10https://docs.python.org/3/library/pickle.html
Dataset Our experiments are performed over an
established real world dataset, Clickture 2014
        <xref ref-type="bibr" rid="ref4">(Microsoft, 2014)</xref>
        , released by Microsoft as part of
an Image Retrieval Challenge in 2015.
Commercial image search engines like Google, Bing
record clicks against queries to capture the user
behaviour. Insightful usage of the recorded
clicklogs may lead to better cross-modal retrieval. The
dataset comprises of two parts: (a) the training
dataset and (b) the testing dataset or Dev dataset.
The training dataset, consists of 1 million images
and 11.7 million unique queries, is a sample of
user click log which is a large table consisting of
text queries, its associated images and number of
of clicks for each query-image pair. An example
of a subgraph of the Clickture dataset is depicted
in Figure 1.
      </p>
      <p>The click count between an image and a query is
calculated from different users at different times.
There are at least 23.1 million query-image pairs
which have click count equal or more than 1.
The Dev Dataset, which has 79, 926 query-image
pairs generated from 1, 000 queries, is composed
to have consistent query distribution, judgment
guidelines and quality of a test dataset. For
performance evaluation, manually annotated relevance
measurement which is purely qualitative (labeled
as Excellent, Good and Bad) is provided with the
dataset.
5</p>
    </sec>
    <sec id="sec-4">
      <title>Results and Analysis</title>
      <p>In this work, each image is ranked by its
respective Discounted Cumulated Gain (DCG) measure
against the test queries. To calculate DCG, we sort
the images against each query based on the final
similarity score, obtained from our process. DCG
for each query is computed as depicted in the
following Equation 4. This metric was the official
metric for the MSR- Bing Image Retrieval
Challenge 2014 and 201511.</p>
      <p>
        25 2reli − 1
DCG25 = Z25 X
i=1 log2(i + 1)
(4)
where reli = {Excelent = 3; Good = 2; Bad =
0} is the manually judged relevance for each
image with respect to the query, and Z25 = 0.01757
is a normalizer to make the score for 25 Excellent
results. Here, we report the final evaluation metric
as the average of for all queries present in the test
set. We choose the comparative methods
(baselines) against our proposed system from the base
paper
        <xref ref-type="bibr" rid="ref8">(Wu et al., 2016a)</xref>
        . The comparative
methods are as follows:
      </p>
      <sec id="sec-4-1">
        <title>1. Bag-of-Words similarity method (BoWDNN-R) based ranking</title>
        <p>11http://press.liacs.nl/mmgrand/microsoft.pdf
3. Passive-Aggressive Model for Image
Retrieval (PAMIR)</p>
      </sec>
      <sec id="sec-4-2">
        <title>4. Polynomial Semantic Indexing (PSI)</title>
        <p>5. Cross-Model Ranking Neural Network
(CM</p>
        <p>RNN) and
6. Multimodal Random Walk Neural Network
(MRW-NN) respectively.</p>
        <p>From Table 1, we can conclude that there is
marked improvement in terms of retrieval
performance when compared to other state-of-the-art
techniques. The gain in terms of DCG is also
statistically significant (indicated with a superscript
p). Hence, we can safely conclude that
considering click-counts as ratings and formulating the
problem of image retrieval as item
recommendation yields significantly better performance. The
possible reason why our system performs better
than others could be attributed to the fact that we
did not rely on either latent semantic
representation based learning or collaborative filtering
individually. Instead, we proposed a new model
that incorporates the learned feature
representations from CNN as user and items, which possibly
negated the shortfall of both techniques.
6</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>Conclusion and Future Work</title>
      <p>Image retrieval has been one of the focal points
of information retrieval systems since the early
days of computing. Recent techniques have
focused on various learning techniques to minimize
the semantic gap between the query intent of users
and the actual information retrieved by IRs. The
same applies to image retrieval as well. While
hybrid and multi-modal systems have shown
superior performance when compared to unimodal
retrieval systems, the problem of capturing the user
information need ideally continues to be a
challenge. Click counts offer a new dimension to aid in
better understanding user’s information need
concerning images and when used judiciously can
significantly improve the corresponding IR’s
performance. In this paper, we have applied the
knowledge embedded within the clicks by using a
collaborative filtering technique as an implicit
feedback mechanism to enhance the latent
representation based similarity computation. Our proposed
technique performs superlatively against the
stateof-the-art baseline over a real-world dataset.
As part of our future work, we would like to
address the irregularities in the retrieval mechanism
in the absence of modalities and would like to
explore and suggest techniques to handle
common problems associated with collaborative
filtering methods. We would also like to study
our method’s effectiveness and scalability for
realtime data. This work is a part of a larger project
where we aim to integrate the retrieval model with
image classification and automatic image
annotation techniques proposed by us in our earlier
works.
Semantic Alignments for Generating Image
Descriptions. In The IEEE Conference on Computer
Vision and Pattern Recognition (CVPR).</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          <string-name>
            <given-names>Fangxiang</given-names>
            <surname>Feng</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Xiaojie</given-names>
            <surname>Wang</surname>
          </string-name>
          , and
          <string-name>
            <given-names>Ruifan</given-names>
            <surname>Li</surname>
          </string-name>
          .
          <year>2014</year>
          .
          <article-title>Cross-modal Retrieval with Correspondence Autoencoder</article-title>
          .
          <source>In Proceedings of the 22Nd ACM International Conference on Multimedia. ACM</source>
          , New York, NY, USA, MM '
          <volume>14</volume>
          , pages
          <fpage>7</fpage>
          -
          <lpage>16</lpage>
          . https://doi.org/10.1145/2647868.2654902.
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          <string-name>
            <given-names>Y.</given-names>
            <surname>He</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Xiang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Kang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Wang</surname>
          </string-name>
          , and
          <string-name>
            <given-names>C.</given-names>
            <surname>Pan</surname>
          </string-name>
          .
          <year>2016</year>
          .
          <article-title>Cross-Modal Retrieval via Deep and Bidirectional Representation Learning</article-title>
          .
          <source>IEEE Transactions on Multimedia</source>
          <volume>18</volume>
          (
          <issue>7</issue>
          ):
          <fpage>1363</fpage>
          -
          <lpage>1377</lpage>
          . https://doi.org/10.1109/TMM.
          <year>2016</year>
          .
          <volume>2558463</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          <string-name>
            <given-names>Andrej</given-names>
            <surname>Karpathy</surname>
          </string-name>
          and
          <string-name>
            <given-names>Li</given-names>
            <surname>Fei-Fei</surname>
          </string-name>
          .
          <year>2015</year>
          .
          <article-title>Deep VisualYehuda Koren</article-title>
          and
          <string-name>
            <given-names>Robert</given-names>
            <surname>Bell</surname>
          </string-name>
          .
          <year>2015</year>
          . Advances in Collaborative Filtering,
          <string-name>
            <surname>Springer</surname>
            <given-names>US</given-names>
          </string-name>
          , Boston, MA, pages
          <fpage>77</fpage>
          -
          <lpage>118</lpage>
          . https://doi.org/10.1007/978-1-
          <fpage>4899</fpage>
          - 7637-6
          <fpage>3</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          <string-name>
            <surname>Microsoft</surname>
          </string-name>
          .
          <year>2014</year>
          .
          <article-title>Clickture project</article-title>
          . https:// www.microsoft.com/en-us/research/ project/clickture/. Accessed:
          <fpage>2019</fpage>
          -04-22.
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          <string-name>
            <given-names>Yingwei</given-names>
            <surname>Pan</surname>
          </string-name>
          , Ting Yao, Tao Mei,
          <string-name>
            <given-names>Houqiang</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <surname>Chong-Wah Ngo</surname>
            , and
            <given-names>Yong</given-names>
          </string-name>
          <string-name>
            <surname>Rui</surname>
          </string-name>
          .
          <year>2014</year>
          .
          <article-title>Clickthrough-based Cross-view Learning for Image Search</article-title>
          .
          <source>In Proceedings of the 37th International ACM SIGIR Conference on Research &amp;#38; Development in Information Retrieval. ACM</source>
          , New York, NY, USA, SIGIR '
          <volume>14</volume>
          , pages
          <fpage>717</fpage>
          -
          <lpage>726</lpage>
          . https://doi.org/10.1145/2600428.2609568.
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          <string-name>
            <given-names>Scott</given-names>
            <surname>Reed</surname>
          </string-name>
          , Zeynep Akata,
          <string-name>
            <given-names>Honglak</given-names>
            <surname>Lee</surname>
          </string-name>
          ,
          <string-name>
            <given-names>and Bernt</given-names>
            <surname>Schiele</surname>
          </string-name>
          .
          <year>2016</year>
          .
          <article-title>Learning Deep Representations of Fine-Grained Visual Descriptions</article-title>
          .
          <source>In The IEEE Conference on Computer Vision and Pattern Recognition (CVPR).</source>
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          <string-name>
            <given-names>C.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Yang</surname>
          </string-name>
          , and
          <string-name>
            <given-names>C.</given-names>
            <surname>Meinel</surname>
          </string-name>
          .
          <year>2015</year>
          .
          <article-title>Deep Semantic Mapping for Cross-Modal Retrieval</article-title>
          . In Fei Wu, Xinyan Lu, Jun Song, Shuicheng Yan, Zhongfei Mark Zhang, Yong Rui, and
          <string-name>
            <given-names>Yueting</given-names>
            <surname>Zhuang</surname>
          </string-name>
          .
          <year>2016a</year>
          .
          <article-title>Learning of Multimodal Representations With Random Walks on the Click Graph</article-title>
          .
          <source>IEEE transactions on image processing : a publication of the IEEE Signal Processing Society</source>
          <volume>25</volume>
          (
          <issue>2</issue>
          ):
          <fpage>630642</fpage>
          . https://doi.org/10.1109/tip.
          <year>2015</year>
          .
          <volume>2507401</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          <string-name>
            <surname>P. Wu</surname>
            ,
            <given-names>S. C. H.</given-names>
          </string-name>
          <string-name>
            <surname>Hoi</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          <string-name>
            <surname>Zhao</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          <string-name>
            <surname>Miao</surname>
            , and
            <given-names>Z. Y.</given-names>
          </string-name>
          <string-name>
            <surname>Liu</surname>
          </string-name>
          . 2016b.
          <article-title>Online Multi-Modal Distance Metric Learning with Application to Image Retrieval</article-title>
          .
          <source>IEEE Transactions on Knowledge and Data Engineering</source>
          <volume>28</volume>
          (
          <issue>2</issue>
          ):
          <fpage>454</fpage>
          -
          <lpage>467</lpage>
          . https://doi.org/10.1109/TKDE.
          <year>2015</year>
          .
          <volume>2477296</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          <string-name>
            <given-names>Yi</given-names>
            <surname>Zhen</surname>
          </string-name>
          and
          <string-name>
            <surname>Dit-Yan Yeung</surname>
          </string-name>
          .
          <year>2012</year>
          .
          <article-title>A Probabilistic Model for Multimodal Hash Function Learning</article-title>
          .
          <source>In Proceedings of the 18th ACM SIGKDD International Conference on Knowledge Discovery and Data Mining. ACM</source>
          , New York, NY, USA, KDD '
          <volume>12</volume>
          , pages
          <fpage>940</fpage>
          -
          <lpage>948</lpage>
          . https://doi.org/10.1145/2339530.2339678.
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>