<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <issn pub-type="ppub">1613-0073</issn>
    </journal-meta>
    <article-meta>
      <title-group>
        <article-title>Knowledge Distillation for a Domain-Adaptive Visual Recommender System</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="editor">
          <string-name>Knowledge Distillation, Object Detection, Computer Vision, Visual Search</string-name>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Computer Science Department, University of Turin</institution>
          ,
          <country country="IT">Italy</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Inferendo srl</institution>
          ,
          <addr-line>Alessandria</addr-line>
          ,
          <country country="IT">Italy</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>In the last few years large-scale foundational models have shown remarkable performance in computer vision tasks. However, deploying such models in a production environment poses a significant challenge, because of their computational requirements. Furthermore, these models typically produce generic results and they often need some sort of external input. The concept of knowledge distillation provides a promising solution to this problem. By leveraging the teacher-student framework, the smaller ”student” model learns to mimic the larger ”teacher” model. In this paper, we focus on the challenges faced in the application of such techniques in the task of augmenting an object detection dataset used in a commercial Visual Recommender System that needs to detect items in various e-commerce websites, encompassing a wide range of product categories. We also present a simple solution to the problems we identified and propose a possible direction of future works.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>CEUR
ceur-ws.org</p>
    </sec>
    <sec id="sec-2">
      <title>1. Introduction</title>
      <p>
        Visual Recommender Systems have emerged as a powerful tool in the field of e-commerce,
providing personalized product recommendations based on visual similarity and user preferences.
As described in 1[], while traditional recommender systems primarily rely on user-item
interactions, Visual Recommender Systems leverage image similarity and visual search techniques to
enhance the recommendation process. The fundamental building blocks entailing these systems
are:
• image similarity and feature extraction: at the core of a visual recommender or any
Content Based Instance Retrieval system (CBIR) is the ability to quantify and compare
visual content[
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]. This involves extracting meaningful features from images, which can
be extracted by means of statistical analysis such as a simple color histogram or can be
produced by complex deep learning models as in the case we are studying.
• visual search: it is a crucial component of Visual Recommender Systems; it involves the
detection of objects or items within user-uploaded images. This task is carried out by
deep learning object detection models, enabling the system to identify products or items
within the images with varying levels of accuracy. Most similar items to the one detected
are then searched and retrieved.
      </p>
      <p>In summary, to build a proper Visual Recommender System two models are needed:
1. image embedding model: it is responsible for producing embeddings of images that
are capable of describing them in enough detail to perform a successful similarity search,
e.g. embeddings of similar items must be as close as possible in the latent space.
2. object detection model: it must be able to recognize relevant items in images of products
extracted from an e-commerce picture or from user-uploaded content.</p>
      <p>While the use of a large pre-trained model such as CLI3P][has empirically proven to be more
than enough to produce accurate vector embeddings of the images, the dataset composition on
which the object detection model is trained plays a pivotal role in determining the system’s
capability to recognize a broad array of objects. More importantly, the classes within the dataset
correspond to the actual items within a potential e-commerce platform that would utilize the
recommender system, making it a critical factor influencing the system’s capacity to identify
and classify diverse items.</p>
      <p>In the present work, we aim at discussing a solution for a Visual Recommender System
operating as a service, known as Recommendations as a Service (RaaS). Traditionally, recommendation
systems were embedded within specific applications, requiring significant engineering efort
and resources to implement and maintain. With the emergence of cloud computing and
microservices architectures, the concept of RaaS has gained prominence.</p>
      <p>The research project we are pursuing is being developed in the context of Visi4d]e,aa[
commercial product that aims to ofer a wide range of recommender modalitites as a service, so
that developers and businesses can add recommendation capabilities to their platforms through
APIs, allowing seamless integration into various applications and websites. Visidea focuses on
ofering recommendations for the e-commerce sector and thus needs to be able to handle as
many items as possible, due to the ever-varying nature of the markets.</p>
      <p>While a Visual Recommender System ofers significant benefits, it also faces a unique
challenge known as domain adaptation. In this context, domain adaptation refers to the ability to
accommodate a wide range of e-commerce websites that query the system, each one with its
own set of products and classes of objects. To ensure accurate and relevant recommendations
across diferent domains, the system needs to be able to quickly adapt and add new classes of
objects to its object detection dataset.</p>
      <p>E-commerce platforms, especially in the fashion industry, frequently introduce new clothing
styles, accessories, or product categories to attract customers. For example, a new fashion
e-commerce platform might start selling ”Jumpsuits” a category of product that the existing
object detection model may not be able to detect, because the original dataset on which it has
been trained does not have that class in its labels. Traditional object detection models rely on
extensive training data for each object class they are supposed to detect. When a new class
appears, there is often a shortage of labeled training data, making it challenging to fine-tune
the model efectively.</p>
      <p>To tackle this issue, an automatic method is employed to add new classes to the object detection
dataset. This method leverages techniques like transfer learning and knowledge distillati5o]n [
to eficiently transfer the knowledge from the large foundational models to smaller and faster
models that can handle the new classes.</p>
    </sec>
    <sec id="sec-3">
      <title>2. Knowledge Distillation through Autodistill</title>
      <p>Distillation-based processes, similar to those used in natural language processing, have been
applied to create more compact models that derive their knowledge from larger, pre-existing
models. An illustrative instance of this approach is the Stanford Alpaca model, introduced in
March 2023. The developers employed OpenAI’s text-davinci-003 model to generate 52,000
instructions using an initial dataset as its reference poi6n]t.[The Autodistill Python package
developed by Roboflow company ofers automated image labeling using foundational models
that have undergone training on an extensive dataset, drawing on millions of images and
significant computational resources invested by major industry leaders such as Meta, Google,
and Amazon. The resultant dataset can subsequently serve as a valuable resource for training
cutting-edge models, harnessing their eficiency and reliability for deployment in production
environments. Autodistill provides a selection of base models and target models. One of
the base models can generate a dataset in the precise format needed for the chosen target
model. Consequently, training the target model on this generated dataset represents the actual
”distillation” of knowledge</p>
      <p>
        Amongst the pletora of base models ofered by autodistill, we choose to employ
GroundingDINO [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ], an extension of the DINO 8[] model developed precisely for the purpose of
zero-shot object detection. DINO is a self-supervised learning method for visual representation
learning. It introduces a novel training objective that encourages visual representations to
emerge with consistent semantic properties. DINO achieves this by contrasting multiple views
of the same image and optimizing a similarity-based loss function.
      </p>
      <p>GroundingDINO is an extension of DINO that focuses on grounding visual representations
with textual descriptions. It leverages the contrastive learning framework of DINO to learn
representations that capture the semantics of both images and text. GroundingDINO performs
joint training with image-text pairs and learns to align the visual and textual modalities by
maximizing their similarity.</p>
      <p>Autodistill needs as input both the images to label and a description, called “ontology”, of the
objects to detect inside those images. The ontology comes in the following form{a"tp:rompt":
"label"} where the prompt is a natural language description of the object to detect and the
label is simply the associated class word. According to Autodistill an ontology can include
several objects, each with its own prompt, that will be detected at the same time in each image.
While the primary objective of autodistill is to facilitate the labeling of objects within images,
an unexpected outcome was observed when dealing with ontologies containing more than one
class. In these scenarios, the distillation process resulted in the labeling of all objects within the
images with every possible class specified in the ontology. We identified two main issues:
• label duplication: several objects within the image were labeled with not only the
appropriate class but also with almost every other class present in the ontology. Consequently,
the correct label appeared multiple times, making the labeling output redundant.
• erroneous labelling: objects within the image were often mislabeled with classes that
did not correspond to them. This meant that the distillation process was not only overly
inclusive in assigning labels but also misinterpreted parts of the image as objects belonging
to classes unrelated to the actual target.</p>
      <p>These issues posed a significant obstacle in achieving precise and reliable object labeling,
particularly when dealing with images containing multiple objects belonging to diferent classes.</p>
      <p>Moreover, another important aspect is the sensitivity of the provided prompts. Even minor
variations in the wording of prompts could have a profound impact on the quality and accuracy
of the produced dataset. This sensitivity is notable when generating labels for specific objects,
as it directly influences the model’s ability to correctly identify and categorize objects within
images.</p>
      <p>To illustrate this issue, let’s consider a practical example. Suppose the objective is to identify
swimsuits within images. Initially, we provided the prompt ”a picture of a swimsuit” to guide
the model. However, this often resulted in the model not only correctly identifying the swimsuit
itself but also erroneously associating the ”swimsuit” class with the entire person wearing the
swimsuit. In other words, the labeling extended beyond the target object to encompass the
broader context.</p>
      <p>Minor alterations of the prompt such as: ”a single piece of a swimsuit.” led to significantly
improved results as shown in Figur1e. The model’s ability to distinguish between the swimsuit
as the target object and the person wearing it as the contextual background became notably
more precise.</p>
    </sec>
    <sec id="sec-4">
      <title>3. One-class-at-the-time solution</title>
      <p>To take under control the issue related to the “multi-class identification” (i.e. the identification
of objects not directly related to what one is actually searching for), a solution focusing on a
single-class labeling was devised. This involves a one-class-at-a-time labeling strategy, where
images of a single class, such as swimsuits, were collected from the web. The ontology provided
for this approach contained solely the prompt and label for the specific class under consideration,
omitting references to other objects. The prompt was manually created after some empirical
testing on a subset of the downloaded images.</p>
      <p>This one-class-at-a-time labeling approach yielded promising results, with the majority of
images correctly labeled. The model demonstrated competence in identifying and associating
the label with the intended class. Certain errors persisted, particularly when clothing is worn
by humans. In such instances, the model occasionally struggled to diferentiate between the
garment, which was the intended target object, and the person wearing it, considered as part of
the contextual background. This issue highlighted the complexities involved in recognizing
objects within a contextual setting and remarks the need to find better methodologies to refine
these foundational models, that frequently end up in being too generic to be actually used in a
real world commercial application.</p>
    </sec>
    <sec id="sec-5">
      <title>4. Conclusion and Future Works</title>
      <p>Moving forward our research endeavors aim to overcome the challenges encountered in the
single-class labeling approach within Autodistill. While this method has shown promising
results in isolating and labeling objects, it still requires significant manual data collection eforts
for each class and struggles in complex contextual scenarios. Our vision for the future involves
the development of a more advanced and eficient solution that harnesses spatial and relational
knowledge. This solution seeks to infer accurate object labels by analyzing the interplay
between objects within an image. By leveraging sophisticated techniques for recognizing object
relationships and spatial configurations, we aspire to enhance the quality and precision of image
labeling in multi-class environments.</p>
    </sec>
    <sec id="sec-6">
      <title>Acknowledgments</title>
      <p>My reasearch is conducted as “Dottorato in Alto Apprendistato”Ininferendo, an innovative
start-up, spin-of of the University of Piemonte Orientale. Funding as been provided by Regione
Piemonte. I want to deeply thank my tutors Luigi Portinale (University of Piemonte Orientale)
and Roberto Esposito (Univeristy of Torino) for their guidance and support throughout my
academic journey.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>A.</given-names>
            <surname>Abluton</surname>
          </string-name>
          ,
          <article-title>Visual recommendation and visual search for fashion e-commerce</article-title>
          ,
          <source>in: International Conference on Similarity Search and Applications</source>
          , Springer,
          <year>2022</year>
          , pp.
          <fpage>299</fpage>
          -
          <lpage>304</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>W.</given-names>
            <surname>Chen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Liu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E. M.</given-names>
            <surname>Bakker</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Georgiou</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Fieguth</surname>
          </string-name>
          , L. Liu,
          <string-name>
            <given-names>M. S.</given-names>
            <surname>Lew</surname>
          </string-name>
          ,
          <article-title>Deep learning for instance retrieval: A survey</article-title>
          ,
          <source>IEEE Transactions on Pattern Analysis and Machine Intelligence</source>
          <volume>45</volume>
          (
          <year>2023</year>
          )
          <fpage>7270</fpage>
          -
          <lpage>7292</lpage>
          .
          <year>doi1</year>
          :
          <fpage>0</fpage>
          .1109/TPAMI.
          <year>2022</year>
          .
          <volume>3218591</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>A.</given-names>
            <surname>Radford</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. W.</given-names>
            <surname>Kim</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Hallacy</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Ramesh</surname>
          </string-name>
          , G. Goh,
          <string-name>
            <given-names>S.</given-names>
            <surname>Agarwal</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Sastry</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Askell</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Mishkin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Clark</surname>
          </string-name>
          , et al.,
          <article-title>Learning transferable visual models from natural language supervision</article-title>
          ,
          <source>in: International conference on machine learning, PMLR</source>
          ,
          <year>2021</year>
          , pp.
          <fpage>8748</fpage>
          -
          <lpage>8763</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <surname>Visidea</surname>
          </string-name>
          , ???? URL: https://visidea.ai/.
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>M. S. e. a. Gou J.</given-names>
            ,
            <surname>Yu</surname>
          </string-name>
          <string-name>
            <surname>B.</surname>
          </string-name>
          ,
          <article-title>Knowledge distillation: A survey</article-title>
          ,
          <source>International Journal of Computer Vision</source>
          <volume>129</volume>
          (
          <year>2021</year>
          )
          <fpage>1789</fpage>
          -
          <lpage>1819</lpage>
          .
          <year>doi1</year>
          :
          <fpage>0</fpage>
          .1007/s11263-021-01453-z.
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>R.</given-names>
            <surname>Taori</surname>
          </string-name>
          , I. Gulrajani,
          <string-name>
            <given-names>T.</given-names>
            <surname>Zhang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Dubois</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Guestrin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Liang</surname>
          </string-name>
          , T. B.
          <string-name>
            <surname>Hashimoto</surname>
          </string-name>
          , Stanford alpaca:
          <article-title>An instruction-following llama model</article-title>
          , https://github.com/tatsu-lab/stanford_alpaca,
          <year>2023</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>S.</given-names>
            <surname>Liu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Zeng</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Ren</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Zhang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Yang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Yang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Su</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Zhu</surname>
          </string-name>
          , et al.,
          <article-title>Grounding dino: Marrying dino with grounded pre-training for open-set object detection</article-title>
          ,
          <source>arXiv preprint arXiv:2303.05499</source>
          (
          <year>2023</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>M.</given-names>
            <surname>Caron</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Touvron</surname>
          </string-name>
          , I. Misra,
          <string-name>
            <given-names>H.</given-names>
            <surname>Jégou</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Mairal</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Bojanowski</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Joulin</surname>
          </string-name>
          ,
          <article-title>Emerging properties in self-supervised vision transformers</article-title>
          ,
          <source>in: Proceedings of the IEEE/CVF international conference on computer vision</source>
          ,
          <year>2021</year>
          , pp.
          <fpage>9650</fpage>
          -
          <lpage>9660</lpage>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>