<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <journal-title-group>
        <journal-title>Ital-IA</journal-title>
      </journal-title-group>
    </journal-meta>
    <article-meta>
      <title-group>
        <article-title>meets Industry: New Challenges and Opportunities at AImageLab</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Rita Cucchiara</string-name>
          <email>rita.cucchiara@unimore.it</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Lorenzo Baraldi</string-name>
          <email>lorenzo.baraldi@unimore.it</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Simone Calderara</string-name>
          <email>simone.calderara@unimore.it</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Marcella Cornia</string-name>
          <email>marcella.cornia@unimore.it</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Matteo Fabbri</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Costantino Grana</string-name>
          <email>costantino.grana@unimore.it</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Angelo Porrello</string-name>
          <email>angelo.porrello@unimore.it</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Roberto Vezzani</string-name>
          <email>roberto.vezzani@unimore.it</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>GoatAI S.r.l.</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>AImageLab, University of Modena and Reggio Emilia</institution>
          ,
          <addr-line>Modena</addr-line>
          ,
          <country country="IT">Italy</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Artificial Intelligence</institution>
          ,
          <addr-line>Innovation, Industry 5.0, Video Surveillance, Generative AI, Continual Learning</addr-line>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>Natural Language Processing, Medical Imaging</institution>
          ,
          <addr-line>Robot-</addr-line>
        </aff>
        <aff id="aff3">
          <label>3</label>
          <institution>sented by the AImageLab laboratory</institution>
          ,
          <addr-line>an internationally</addr-line>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2023</year>
      </pub-date>
      <volume>3</volume>
      <fpage>29</fpage>
      <lpage>31</lpage>
      <abstract>
        <p>The application of Artificial Intelligence (AI) is becoming increasingly ubiquitous in industrial fields, posing some of the major issues and challenges of the future decades. In this respect, the old dichotomy Industry vs. Research appears tired: instead, a strong synergy between these two worlds represents a key aspect toward the most breakthrough incoming innovations. This manuscript presents an example of such a connection: namely, some of the activities carried out at the AImageLab laboratory, one of the most active research lab in Italy and Europe in the fields of Computer Vision and Artificial Intelligence.</p>
      </abstract>
      <kwd-group>
        <kwd>successful case of research/industry integration</kwd>
        <kwd>repre-</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <sec id="sec-1-1">
        <title>Artificial Intelligence (AI) has emerged as the next break</title>
        <p>through technology and is going to shape the innovation
landscape of various fields, in ways not yet fully
underthat dates back to the early 1950s (with the definition of
the former idealized artificial neurons).</p>
      </sec>
      <sec id="sec-1-2">
        <title>In this respect, the landscape has radically changed</title>
        <p>in the last two decades: the major research fields falling
under the umbrella term AI – i.e. Deep Learning and</p>
      </sec>
      <sec id="sec-1-3">
        <title>Computer Vision – have crossed the chasm, switching</title>
        <p>from frontier academic topics to mature sources of
operative and foundational business tools. Indeed, there is a
disruptive change of perspective in how industry players
look at AI: not just as a cutting-edge technology, but as a
as already holds for electricity and water.</p>
      </sec>
      <sec id="sec-1-4">
        <title>In such a context, the boundaries between academic research and the industry field have been blurring: to</title>
        <p>nEvelop-O
recognized research lab of the Department of Engineering
”Enzo Ferrari” located within the University of Modena
and Reggio Emilia. The research covers topics of
Computer Vision, Pattern Recognition &amp; Machine Learning,</p>
      </sec>
      <sec id="sec-1-5">
        <title>Human Interaction, and Multimedia applied to optical</title>
        <p>images and videos as well as data from diferent sensors.</p>
      </sec>
      <sec id="sec-1-6">
        <title>Since its foundation – which dated more than 25 years ago – AImageLab has joined activities for industrial research with many manufacturers and health companies, recently as a partner of the Artificial Intelligence</title>
      </sec>
      <sec id="sec-1-7">
        <title>Technopole. The projects carried out and the topics cov</title>
        <p>DESIGNER
ered have achieved a strict connection with the economic
fabric of the Emilia-Romagna region, tied principally to
manufacturing production. However, in the last years,
the AImageLab has been evolving to meet the growing
demands of the service industry: more and more startups
and established IT service companies are active in
producing digital solutions for local industry and for export
in Italy and the world.</p>
        <p>Therefore, we kindly refer the reader to the next
sections, each of which is presenting a cross-section of some
of the industry-oriented research activities carried out
within the AImageLab laboratory.
sensitive information useful for data analytics purposes.</p>
        <p>Featuring cutting-edge AI technology, GoatEye is able
to deliver powerful video analysis capabilities while
ensuring that user privacy is always protected. With
GoatEye, customers enjoy all the benefits of AI without
worrying about people privacy being compromised. It utilizes
advanced algorithms to perceive humans in real-time,
2. GoatEye: the Privacy Camera assessing position, posture, and even action performed
with high accuracy. What sets GoatEye apart from the
Among the industry-oriented activities of AImageLab, competition is its commitment to user privacy. It uses
the one carried out along with the startup GoatAI is with techniques to ensure that personal information is never
no doubt the more emblematic. GoatAI has been con- collected or shared without users’ consent. This means
ceived within AImageLab and, as such, fully embraces that customers are able to use GoatEye with confidence,
a mindset where research and industry are strictly en- knowing that the privacy of people is always respected.
twined. GoatAI purpose is to apply the most modern AI
techniques for Human Behavior Understanding in health
care, fitness, retail, safety and security. 3. AI for the Ceramic Industry</p>
        <p>As AI continues to advance, it is crucial that it is used
in ways that respect individuals and their personal infor- Image enhancement In the design industry, the
aumation. That’s why GoatAI is on a mission to develop tomation of color management and image enhancement
innovative solutions that not only provide benefits, but are challenging tasks, with several applications related to
also prioritize and safeguard the privacy of those who use the creation and the print of design surfaces. As outlined
them. Nowadays and even more in the future, applying in the next paragraphs, automated tools would speed up
the wrong data privacy strategy can cost an organization the work of graphic designers in the creation of novel
billion in fees and damages. ceramic tiles, with particular focus on the editing and the</p>
        <p>Additionally, with the advent of the AI Act, several preparation of images fed to the printing pipeline.
measures will be enforced to protect individual rights Indeed, one of the most time-consuming activities of
and privacy when processing sensitive data through AI designers is to modify the image so that, once printed,
models, intended to ensure that AI systems will be used its quality and level of details are consistent with what
in a responsible and ethical manner. The AI Act will customers see on the monitor (or on paper). Such an
have a significant impact on those developing AI systems, operation – currently carried out manually through
adespecially for Human Behavior Understanding. vanced image editing tools – is not only complex and</p>
        <p>That is why GoatAI wants to change the way people time-consuming, but also hard to formalize, as it is based
see cameras with GoatEye: the Privacy Camera. GoatEye on the experience and sensitivity of designers.
wants to provide a new layer of abstraction by producing The AImageLab team has recently proposed novel AI
anonymized video material while maintaining all the non- methods – based on Convolutional Neural Networks – to
enhance the visual quality of industrial design surfaces,
RefereRnceeferenTcrey-OnTry-OnTry-OnTry-On
ModelModGealrmeGntasrmenRtsesult Result</p>
        <p>ReferenRceeferenTcrey-OnTry-OnTry-OnTry-On
ModelModGealrmeGntasrmenRtsesult Result</p>
        <p>ReferenRceeferenTcrey-OnTry-OnTry-OnTry-On
ModelModGealrmeGntasrmenRtsesult Result
R</p>
        <p>F</p>
        <p>F</p>
      </sec>
    </sec>
    <sec id="sec-2">
      <title>4. Generative AI for Fashion</title>
      <sec id="sec-2-1">
        <title>With the advent of e-commerce, the variety and avail</title>
        <p>
          according to specific properties dictated by color profiles ability of online garments have become increasingly
(for a visual of the obtained results, see Fig. 2). These overwhelming for the final user. Consequently,
userapproaches were developed in collaboration with Digital oriented services and applications such as virtual
tryDesign Srl, a leading design company located in Fiorano on [
          <xref ref-type="bibr" rid="ref7 ref8 ref9">2, 3, 4</xref>
          ], customer-to-shop garment retrieval, and
Modenese (Modena) specialized in creating high quality vision-and-language interactions [
          <xref ref-type="bibr" rid="ref10 ref11">5, 6</xref>
          ] are increasingly
design surfaces for diferent applications, in particular important for online shopping, helping fashion
compaceramic tile printing. nies to tailor the e-commerce experience and maximize
Generative design In order to provide further support customer satisfaction.
to designers, a team of researchers at AImageLab devel- In this context, the task of image-based virtual
tryoped a new approach to synthesize novel ceramic tiles. on aims at synthesizing an image of a reference person
By generating new patterns and textures, the approach wearing a given try-on garment while preserving the
allows to generate the diferent tile faces that make up person’s intrinsic information such as body shape and
an entire ceramic surface. Such a case study is a practical pose. Existing datasets for the task are either
propriexample of how AI technologies can support the creation etary and therefore not publicly available, or feature a
and evolution of new products, which represent the main very limited number of images which are usually low
goals of generative design. The latter is an important mile- resolution [
          <xref ref-type="bibr" rid="ref7">2</xref>
          ]. To tackle these problems, AImageLab
stone in industrial production, especially for processes researchers recently presented Dress Code [
          <xref ref-type="bibr" rid="ref9">4</xref>
          ]: a new
involving an important creative component. dataset of high-resolution images (1024 × 768)
contain
        </p>
        <p>
          To do so, the researchers have leveraged Generative ing more than 50k image pairs of try-on garments and
Adversarial Networks (GANs): these models, by virtue corresponding catalog images where each item is worn
of the extremely high-quality and high-fidelity images by a model. Diferently from existing public datasets,
provided, have seen a significant growing interest for a which contain only upper-body clothes, Dress Code
feavariety of applications, ranging from super-resolution tures upper-body, lower-body, and full-body clothes, as
to image-to-image translation, from image inpainting to well as full-body images of human models.
the field of image synthesis [
          <xref ref-type="bibr" rid="ref6">1</xref>
          ]. In addition to the dataset, the researchers have also
in
        </p>
        <p>The idea of supporting designers and engineers by troduced a novel image-based virtual try-on architecture
using a generative AI-based system opens new ways in that can anchor the given garment to the right portion
the so-called Industry 5.0 paradigm, where humans and of the body. As a consequence, it is possible to perform
AI-based engines should collaborate for the industry of a “complete” try-on over a given person by selecting
tomorrow. However, the proposed system has already diferent garments (Fig 4).
been patented and is currently in use in the ceramic tile
industry, whose designers, who supervise both the training 5. AI for Interior Design
and the generation process, are exploiting with
considerable satisfaction the flexibility and creative capabilities
of our generative model.</p>
      </sec>
      <sec id="sec-2-2">
        <title>An additional research and industrial innovation field on which AImageLab works is that of empowering aug</title>
        <p>DEPDLEOPYLMOYEMNTENTONDOEDE
ADAADPATPT</p>
        <p>DEPDLEOPYLMOYEMNTENTONDOEDE</p>
        <p>DODNOONTOT
TRATNRASMNSITMIT</p>
        <p>ADAADPATPT
TRATIRNAININGINNGONDOEDE</p>
        <p>TRATIRNAININGINNGONDOEDE</p>
        <p>Autonomous Learning</p>
        <p>DEPDLEOPYLMOYEMNTENTONDOEDE</p>
        <p>CLCTLRTARIANIN
MM
E E
R R
G G
E E</p>
        <p>SMSAMLLALL</p>
        <p>WOWRKOIRNKGING
TRATIRNAININGINNGONDOEDE MEMEOMRYORY
TRATIRNAININGING
DATDAASTEATSET</p>
        <p>TRATIRNAININGING</p>
        <p>DATDAASTEATSET
RE-RTER-ATIRNAMINOMDOELDEL</p>
        <p>RE-RTER-ATIRNAMINOMDOELDEL</p>
        <p>RE-RTER-ATIRNAMINOMDOELDEL
mented and virtual reality application for interior design. 6. Continual Learning
In this context, a fundamental task which is addressed
by our research activities is that of automatically parsing Human intelligence is distinguished from Artificial
Neuand understanding pictures of indoor scenes. The goal of ral Networks (ANNs) by the former’s ability to acquire
the task is that of providing detailed information about knowledge incrementally and retain long-term
memthe objects in a scene, the layout of the space, and how ory. Conversely, ANNs sufer sudden performance
deobjects interact with each other. terioration, termed catastrophic forgetting, in response</p>
        <p>One of the core subtasks which need to be solved in to changes in training data distribution [9]. Continual
this context is that of performing a semantic segmen- Learning (CL) [10] is a rapidly growing area of machine
tation over the input image. The research on semantic learning which focuses on bridging this gap by devising
segmentation models has focused on the introduction of technical solutions that allow AI systems to overcome
either fully convolutional networks or Vision Transform- forgetting. In recent years, the AImageLab team has
ers which leverage upsampling operations to increase the actively contributed to CL research, focusing especially
output resolution. Although this architectural choice is on the development of novel CL methods belonging to
necessary to encode contextual information and deal with performant rehearsal-based category [11, 12].
objects at large scales, it also leads to feature smoothing While apparently abstract, this emerging discipline has
across object boundaries, and thus to a degraded quality significant implications for MLOps practices. Today, a
in the final result. model’s life cycle typically begins with training on a
ded</p>
        <p>
          With the aim of improving the quality of semantic icated training node (e.g., a high-compute server), with
segmentation in indoor scenarios, especially in boundary the result being frozen and used for inference on a
separegions, we have investigated the design of boundary- rate deployment node (a machine that operates on the
aware losses for the optimization of semantic segmen- edge, close to the source of operational data). Herein, we
tation architectures, both in CNN-based and ViT-based examine three scenarios (see also Fig. 5) that demonstrate
architectures. We started from two recently proposed the potential of CL algorithms to alter this paradigm.
loss functions, namely the Boundary loss and the Ac- Drift Prevention. In-deployment model re-training is
tive Boundary loss, and designed two improved versions required when significant diferences occur between
inthat can significantly increase the overall quality of the ference and training data distributions. As this process
insegmentation at boundary level [
          <xref ref-type="bibr" rid="ref12">7</xref>
          ]. volves recording new data, transmitting it to the training
        </p>
        <p>
          Our semantic segmentation solutions have been ap- node and learning a new model from scratch, the model’s
plied to the recognition of walls and floors in indoor response to new data may be delayed or unreliable. By
images, and to the recognition of particular surfaces (e.g. applying a CL algorithm on the deployment node, the
counter-tops or steps). In addition to that, we also work- model can adapt to incoming data gracefully controlling
ing towards the recognition of long-tail objects on the its performance degradation until an update is available
scene, and to the generation of natural language descrip- (Fig. 5a). Recent advancements permit this procedure
tions from indoor images [
          <xref ref-type="bibr" rid="ref13">8</xref>
          ]. even with limited supervision on new data [13].
Decoupled Adaptation. Due to their physical
separation, it is likely that some novel data-points recorded on
the deployment might not be transmitted back to the
training node (e.g., due to security constraints or
technical limitations). In such a scenario (Fig. 5b), CL allows the
model to fit these additional data-points without them
leaving the node. When a re-trained model (unaware of
secure data) becomes available, the CL learner cannot be
trivially replaced, but must follow an adequate procedure
to allow knowledge merging [14].
        </p>
      </sec>
      <sec id="sec-2-3">
        <title>Autonomous Learning. As CL methods improve, it</title>
        <p>becomes increasingly feasible to let them operate over
very long periods of time. Ideally, this will allow the
in-deployment model to fully adapt to upcoming data
dynamically by only leveraging a small working memory,
removing the need of re-training.</p>
        <p>While such a degree of resilience has not yet been
achieved, AImageLab has developed a strong focus on
removing bias from continual learners, allowing them to
operate for longer timespans [15, 16].
In the context of Industry 4.0, experts agree that the
cooperation between humans and intelligent robots [17],
rather than the complete removal of human operators,
will be the best solution for the advancement in
manufacturing. Therefore, safe interaction is a crucial element,
especially regarding social and physical coordination
between coworkers. In this scenario, the ability to predict
the 3D pose of the agents in a collaborative environment
is an enabling technology for real-world safety
applications, e.g. collision avoidance and anomaly detection.</p>
        <p>Thinking about a surveillance system installed in an
industrial setting, vision-based pose estimation approaches
can be exploited to retrieve the 3D pose of an articulated
object with respect to the camera viewpoint. However,
a model that jointly predicts human and robot poses is
still an open problem in the research community. Thus,
the current approach is to consider humans and robots
as separate agents. Since the human pose task has been
extensively investigated in the computer vision
community [18], the focus of our work is on the robotic scenario.</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>7. Theft Recognition</title>
      <p>Robot Pose Estimation (RPE). In robotics, the
common approach to estimate the absolute pose of a robot
The evolution of self-checkout methods in retail stores with respect to the camera is the Hand-Eye Calibration.
presents new challenges that computer vision and deep However, with recent advances in deep learning, many
learning technologies are able to address. As there no works have been proposed to estimate the
camera-tocashier are involved in scanning the items, it is important robot pose using CNNs. Recently, Lee et al. [19]
demonto develop a system able to monitor the customer actions strate that learning-based approaches could replace
clasand detect anomalous behaviours such as shoplifting at- sic marker-based calibration also for standard
manipulatempts or even good-faith mistakes. Modern embedded tors (e.g. Rethink Baxter and Franka Emika Panda). Using
systems allow such solutions to work on the edge, pro- synthetic data for training, they feed RGB images into
viding useful feedback to on-site staf. an encoder-decoder network that predicts the 2D
loca</p>
      <p>To achieve this goal, researchers of AImageLab devel- tion of the robot joints. Assuming the camera intrinsics
oped an approach to track the movement of customers and the configuration of the joint angles are known, the
at the self-checkout station and used this data to clas- camera-to-robot transform is computed via PnP.
sify their actions, distinguishing between malevolent and Going beyond learning-based methods, recent works
benevolent behaviours. This system can be divided into propose approaches based on rendering. Labbe et al. [20]
three main parts: (a) human behaviour detector, (b) ac- pave the way to this field of research presenting the
tion classifier, ( c) retrieval network. ifrst method for robot pose estimation based on the
ren</p>
      <p>The role of the first part of the system is to extract from der&amp;compare paradigm. This optimization algorithm
itvisual data useful information that are able to describe the eratively refines an initial robot state defined as the joint
type of actions of the person involved in the self-checkout. angles configuration and the pose of an anchor part with
Once a good descriptor of the events has been computed, respect to the camera.
it is fed to a more complex architecture, able to link each
scanning action to a set of pre-defined behaviors. Finally, SPDH Pose Representation. In contrast with the
disthe latter module attempts at bridging the gap between cussed works, our work aims to regress the
camera-tothe camera and the traditional laser bar code scanner: robot pose using a novel heatmap-based representation
briefly, it is carried out by learning a neural network of the 3D pose. In this way, we can exploit any
CNNspecialized on a retrieval task, i.e., identifying objects and based architecture trained to predict heatmaps and adapt
ascertaining whether its visual appearance corresponds it to predict our pose representation. Moreover, since
with the reading of its bar code. deep learning applied to robotics requires a lot of data</p>
      <p>The system has shown promising results in identify- and it is not feasible to record real data covering all the
ing diferent types of theft, demonstrating that research possible scenarios, we exploited depth images to close
results can earn a relevant place in the retail sector. the domain gap between the synthetic domain of the
sim</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          <article-title>ulator and the real world. An overview of the proposed memory transformer for image captioning, in: CVPR, approach is depicted in Figure 6</article-title>
          .
          <year>2020</year>
          .
          <article-title>Regarding the pose representation</article-title>
          , we propose the [9]
          <string-name>
            <given-names>M.</given-names>
            <surname>McCloskey</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N. J.</given-names>
            <surname>Cohen</surname>
          </string-name>
          ,
          <article-title>Catastrophic interference in Semi-Perspective Decoupled Heatmaps (SPDH) [21] that connectionist networks: The sequential learning probrely on projections of the 3D space in two 2D spaces:  lem, in: Psychology of learning and motivation, voland  space. The  space, i.e.the camera image plane</article-title>
          ,
          <source>ume 24</source>
          ,
          <year>1989</year>
          , pp.
          <fpage>109</fpage>
          -
          <lpage>165</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          <article-title>and the  space, composed by quantized Z-values</article-title>
          and [10]
          <string-name>
            <surname>M. De Lange</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          <string-name>
            <surname>Aljundi</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          <string-name>
            <surname>Masana</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          <string-name>
            <surname>Parisot</surname>
            ,
            <given-names>X.</given-names>
          </string-name>
          <article-title>Jia, the u dimension.A pose estimation algorithm is trained A</article-title>
          .
          <string-name>
            <surname>Leonardis</surname>
            , G. Slabaugh,
            <given-names>T.</given-names>
          </string-name>
          <string-name>
            <surname>Tuytelaars</surname>
          </string-name>
          ,
          <article-title>A Continual to generate two heatmaps for each joint, one for each Learning Survey: Defying Forgetting in Classification space. The setting is a Sim2Real scenario, so the method Tasks</article-title>
          ,
          <source>IEEE Trans. PAMI</source>
          <volume>44</volume>
          (
          <year>2022</year>
          )
          <fpage>3366</fpage>
          -
          <lpage>3385</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          <article-title>is trained on synthetic data and tested directly on real [11] A. Robins, Catastrophic forgetting, rehearsal and pseudata. Thus, we presented SimBa, a dataset containing dorehearsal</article-title>
          ,
          <source>Connection Science</source>
          <volume>7</volume>
          (
          <year>1995</year>
          )
          <fpage>123</fpage>
          -
          <lpage>146</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          <article-title>RGB-D synthetic and real sequences with a Baxter robot</article-title>
          [12]
          <string-name>
            <given-names>P.</given-names>
            <surname>Buzzega</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Boschini</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Porrello</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Calderara</surname>
          </string-name>
          ,
          <article-title>Reperforming pick-n-place movements. We proved that thinking experience replay: a bag of tricks for continual using depth maps as input reduces the domain gap ob- learning</article-title>
          , in: ICPR,
          <year>2021</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          <article-title>taining promising results for the RPE task</article-title>
          . [13]
          <string-name>
            <given-names>M.</given-names>
            <surname>Boschini</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Buzzega</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Bonicelli</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Porrello</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Calderara</surname>
          </string-name>
          ,
          <article-title>Continual semi-supervised learning through contrastive interpolation consistency</article-title>
          ,
          <source>PRL 162 References</source>
          (
          <year>2022</year>
          )
          <fpage>9</fpage>
          -
          <lpage>14</lpage>
          . [14]
          <string-name>
            <given-names>M.</given-names>
            <surname>Boschini</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Bonicelli</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Porrello</surname>
          </string-name>
          , G. Bellitto, M. Pen-
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>P.</given-names>
            <surname>Shamsolmoali</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Zareapoor</surname>
          </string-name>
          , E. Granger, H. Zhou, nisi, S. Palazzo,
          <string-name>
            <given-names>C.</given-names>
            <surname>Spampinato</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Calderara</surname>
          </string-name>
          ,
          <string-name>
            <surname>Transfer R. Wang</surname>
            ,
            <given-names>M. E.</given-names>
          </string-name>
          <string-name>
            <surname>Celebi</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          <string-name>
            <surname>Yang</surname>
          </string-name>
          ,
          <article-title>Image synthesis with without forgetting</article-title>
          ,
          <source>in: ECCV</source>
          ,
          <year>2022</year>
          .
          <article-title>adversarial networks: A comprehensive survey</article-title>
          and case [15]
          <string-name>
            <given-names>M.</given-names>
            <surname>Boschini</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Bonicelli</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Buzzega</surname>
          </string-name>
          ,
          <string-name>
            <surname>A</surname>
          </string-name>
          . Porrello, studies,
          <source>Information Fusion</source>
          <volume>72</volume>
          (
          <year>2021</year>
          )
          <fpage>126</fpage>
          -
          <lpage>146</lpage>
          . S. Calderara,
          <article-title>Class-incremental continual learning into</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>X.</given-names>
            <surname>Han</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Wu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Wu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Yu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L. S.</given-names>
            <surname>Davis</surname>
          </string-name>
          ,
          <article-title>Viton: An the extended der-verse</article-title>
          ,
          <source>IEEE Trans. PAMI</source>
          (
          <year>2022</year>
          )
          <fpage>1</fpage>
          -
          <lpage>16</lpage>
          .
          <article-title>image-based virtual try-on network</article-title>
          ,
          <source>in: CVPR</source>
          ,
          <year>2018</year>
          . [16]
          <string-name>
            <given-names>L.</given-names>
            <surname>Bonicelli</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Boschini</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Porrello</surname>
          </string-name>
          , C. Spampinato,
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>M.</given-names>
            <surname>Fincato</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Cornia</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Landi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Cesari</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Cucchiara</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Calderara</surname>
          </string-name>
          ,
          <article-title>On the Efectiveness of Lipschitz-Driven Transform, Warp, and Dress: A New Transformation- Rehearsal in Continual Learning</article-title>
          , in: NeurIPS,
          <year>2022</year>
          .
          <article-title>guided Model for Virtual Try-on</article-title>
          ,
          <source>ACM TOMM 18</source>
          (
          <year>2022</year>
          ) [17]
          <string-name>
            <given-names>A.</given-names>
            <surname>Weiss</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Buchner</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Tscheligi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Fischer</surname>
          </string-name>
          , Exploring 1-24.
          <article-title>human-robot cooperation possibilities for semiconductor</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>D.</given-names>
            <surname>Morelli</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Fincato</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Cornia</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Landi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Cesari</surname>
          </string-name>
          , manufacturing, in: CTS,
          <year>2011</year>
          . R. Cucchiara, Dress Code:
          <string-name>
            <surname>High-Resolution</surname>
            Multi- [18]
            <given-names>W.</given-names>
          </string-name>
          <string-name>
            <surname>Liu</surname>
            ,
            <given-names>Q.</given-names>
          </string-name>
          <string-name>
            <surname>Bao</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          <string-name>
            <surname>Sun</surname>
          </string-name>
          , T. Mei,
          <source>Recent advances of monocCategory Virtual Try-On, in: ECCV</source>
          ,
          <year>2022</year>
          .
          <article-title>ular 2d and 3d human pose estimation: a deep learning</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>M.</given-names>
            <surname>Stefanini</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Cornia</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Baraldi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Cascianelli</surname>
          </string-name>
          , G. Fi- perspective,
          <source>ACM Computing Surveys</source>
          <volume>55</volume>
          (
          <year>2022</year>
          )
          <fpage>1</fpage>
          -
          <lpage>41</lpage>
          . ameni, R. Cucchiara, From Show to Tell: A Survey on [19]
          <string-name>
            <given-names>T. E.</given-names>
            <surname>Lee</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Tremblay</surname>
          </string-name>
          , T. To, J. Cheng, T. Mosier,
          <string-name>
            <surname>O.</surname>
          </string-name>
          <article-title>KroeDeep Learning-based Image Captioning</article-title>
          ,
          <source>IEEE Trans. mer</source>
          , D. Fox,
          <string-name>
            <given-names>S.</given-names>
            <surname>Birchfield</surname>
          </string-name>
          ,
          <string-name>
            <surname>Camera-</surname>
          </string-name>
          to-
          <source>robot pose estimaPAMI 45</source>
          (
          <year>2022</year>
          )
          <fpage>539</fpage>
          -
          <lpage>559</lpage>
          .
          <article-title>tion from a single image</article-title>
          ,
          <source>in: ICRA</source>
          ,
          <year>2020</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>N.</given-names>
            <surname>Moratelli</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Barraco</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Morelli</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Cornia</surname>
          </string-name>
          , L. Baraldi, [20]
          <string-name>
            <given-names>Y.</given-names>
            <surname>Labbé</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Carpentier</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Aubry</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Sivic</surname>
          </string-name>
          ,
          <string-name>
            <surname>Single-view R. Cucchiara</surname>
          </string-name>
          ,
          <article-title>Fashion-Oriented Image Captioning with robot pose and joint angle estimation via render &amp; comExternal Knowledge Retrieval and Fully Attentive Gates, pare</article-title>
          , in: CVPR,
          <year>2021</year>
          . Sensors 23 (
          <year>2023</year>
          )
          <fpage>1286</fpage>
          . [21]
          <string-name>
            <given-names>A.</given-names>
            <surname>Simoni</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Pini</surname>
          </string-name>
          , G. Borghi,
          <string-name>
            <given-names>R.</given-names>
            <surname>Vezzani</surname>
          </string-name>
          , Semi-
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>P.</given-names>
            <surname>Bruno</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Amoroso</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Cornia</surname>
          </string-name>
          , S. Cascianelli,
          <article-title>perspective decoupled heatmaps for 3d robot pose estimaL</article-title>
          . Baraldi,
          <string-name>
            <given-names>R.</given-names>
            <surname>Cucchiara</surname>
          </string-name>
          ,
          <article-title>Investigating bidimensional tion from depth maps</article-title>
          ,
          <source>IEEE RA-L</source>
          <volume>7</volume>
          (
          <year>2022</year>
          )
          <fpage>11569</fpage>
          -
          <lpage>11576</lpage>
          .
          <article-title>downsampling in vision transformer models</article-title>
          ,
          <source>in: Image Analysis and Processing-ICIAP</source>
          <year>2022</year>
          ,
          <year>2022</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>M.</given-names>
            <surname>Cornia</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Stefanini</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Baraldi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Cucchiara</surname>
          </string-name>
          , Meshed-
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>