<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Measurement Calculation, Motion Tracking, Gesture Recognition, and Upper Limb Segmentation</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Gilda Manfredi</string-name>
          <email>gilda.manfredi@unibas.it</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Nicola Capece</string-name>
          <email>nicola.capece@unibas.it</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Ugo Erra</string-name>
          <email>ugo.erra@unibas.it</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Monica Gruosso</string-name>
          <email>gruosso.monica@gmail.com</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>At the same time, the Anthropometric Measurement Sys-</institution>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Body Tracking and Anthropometric Measurement Sys-</institution>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>University of Basilicata</institution>
          ,
          <addr-line>via dell'Ateneo Lucano 10, Potenza, 85100</addr-line>
          ,
          <country country="IT">Italy</country>
        </aff>
        <aff id="aff3">
          <label>3</label>
          <institution>unconstrained real-world scenarios. The laboratory has</institution>
        </aff>
      </contrib-group>
      <abstract>
        <p>This contribution discusses the potential of Artificial Intelligence (AI) to enhance Human-Computer Interaction (HCI) methods. Researchers at the Laboratory of Computer Graphics and Parallel Computing at the University of Basilicata have developed several AI-based systems for HCI applications. These systems include an upper limb segmentation system, an XR gesture recognition system, and a virtual dressing room that utilizes Body Tracking and Anthropometric Measurement Systems. The systems use deep learning algorithms to accurately track body movements and interpret hand gestures in real-time, creating a more natural and intuitive interaction with XR environments. The virtual dressing room enables users to create a 3D model of themselves and try on virtual clothing and accessories, ensuring a perfect fit through Anthropometric Measurement System calculations. These AI-based systems have significant potential to enhance user experience and interaction in the HCI field.</p>
      </abstract>
      <kwd-group>
        <kwd>Tracking</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>human-computer interaction, motion tracking, gesture recognition, deep learning, extended reality
1. Introduction
mative technology [1, 2] as it afects nearly every aspect
of human life. One of the areas where AI is showing
ing. In particular, the development of eXtended Reality
(XR) applications, which encompass virtual, augmented,
and mixed-reality experiences, has led to a search for
new and innovative ways to engage and immerse users.</p>
    </sec>
    <sec id="sec-2">
      <title>Starting from the development of XR applications, re</title>
      <p>searchers are exploring new ways of interaction with
these platforms that go beyond traditional HCI methods,
such as mouse and keyboard input. Such researches are
due because traditional HCI methods limit the ability of
users to immerse themselves in XR environments [3].</p>
    </sec>
    <sec id="sec-3">
      <title>Therefore, to enhance the user experience, researchers are turning to AI to develop more natural and intuitive forms of interaction.</title>
    </sec>
    <sec id="sec-4">
      <title>This contribution presents an overview of some of the main research activities focused on AI for the HCI ifeld, conducted by the Laboratory of Computer Graphics and Parallel Computing of the University of Basilicata.</title>
      <p>nized by CINI, May 29–31, 2023, Pisa, Italy
∗Corresponding author.
and interaction in the field of HCI.</p>
      <sec id="sec-4-1">
        <title>2. Egocentric Upper Limb</title>
      </sec>
      <sec id="sec-4-2">
        <title>Segmentation</title>
        <p>One promising area of research in HCI is the development
of egocentric vision-based approaches that enable users
to control their virtual avatars using their body
movements. Many applications involving the use of hands are
based on hand segmentation, which is usually used as
human-robot interaction, hand gesture recognition, and
mixed reality. With the increasing popularity of wearable
and forearms with diferent skin tones, lighting, and
obdevices, there has been a growing interest in egocentric
ject occlusions. The EgoCam dataset, which we manually
or first-person vision (FPV) systems for hand
segmentalabelled, shows four male and female people in simple
tion. However, most existing approaches focused only on
and cluttered environments, indoor and outdoor real-life
the hand up to the wrist or bare arm. They were limited
scenes, and inter-hand occlusions. A subset of the first
in their ability to deal with occlusions, varied lighting
two datasets was used, and data with labelling errors
conditions, and dynamic camera and wearer movement.
were discarded. The images were cropped and resized
This study extended the hand segmentation task to fo- to 360 × 360 to accelerate the training. The upper limb
cus on upper limb segmentation in egocentric vision
segmentation dataset was divided into training and test
and unconstrained real-world scenarios. We trained an
subsets. Additionally, we evaluated the network’s
genencoder-decoder deep convolutional neural network us- eralization level on challenging cases using another test
ing the DeepLabv3+ [4] architecture to overcome the
set named EgoGestureSeg. This subset was taken from a
limitations of the existing methods. The DeepLabv3+
benchmark dataset for egocentric hand gesture
recogniencoder consists of a backbone network, followed by an
atrous spatial pyramid pooling module (ASPP) [5] and a</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>1 × 1 convolutional layer. We selected atrous rates of 6,</title>
      <p>12, and 18 for the ASPP atrous convolutions. We utilized
convolutional and bilinear upsampling operations for the
decoder to obtain spatial information from the encoder
features and refine the segmentation result, resulting in
detailed object boundaries.</p>
      <p>Although various neural networks are available as the
network backbone, we opted for the Xception model [6]
due to its favourable qualitative and quantitative results
tion called EgoGesture and consisted of 235 images
manually labelled by Gonzalez-Sosa et al. [10]. The images
were captured in challenging indoor/outdoor scenarios
and showed clothed and bare limbs, natural or artificial
light, and various occlusions and motion blur.</p>
    </sec>
    <sec id="sec-6">
      <title>We trained the network using our upper limb segmen</title>
      <p>tation train set, and a training method similar to Chen
et al. [11]. We utilized the stochastic gradient descent
optimization algorithm with momentum set at 0.9, the
base learning rate  0 set to 0.0001, and a batch size of 8.</p>
      <p>We also used a cross-entropy loss function and a
polynofor image classification tasks, surpassing prior networks
mial learning rate policy, which was more efective than
like VGG-16, ResNet-152, and Inception V3 while still
other policies and resulted in faster convergence [12, 5].
convolution into a depthwise convolution (spatial convo- (set at 90 for our training phase), and  represents the
hand-to-object occlusions, and a variable amount of mo- scenarios, significantly outperforming the
state-of-thedataset includes indoor and outdoor video frames show- promising results in demanding situations, although it
ing diferent lighting conditions and a user’s bare limb
during real-life actions. TEgO is a large dataset of
highresolution indoor images showing two subjects’ hands
was not particularly developed for egocentric vision
segmentation. The results from HGR-Net were the worst (as
shown in the fourth row of Figure 1), with a poor
classifiachieving a fast computation time. Our experimental
testing revealed that it was the most efective model for our
case study [7]. Specifically, we utilized the Xception-65
model adapted by Chen et al. [4] for semantic
segmentation tasks. It comprises 65 layers and replaces the
original max-pooling layers with atrous depthwise separable
convolutions. These convolutions factorize a standard
lution carried out independently for each channel) with
atrous convolution, followed by a pointwise (1 × 1)
convolution. Additionally, batch normalization and ReLU
were included after each</p>
      <p>3 × 3 depthwise convolution.</p>
      <p>Our dataset includes about 46, 000 varied RGB images
with accurate labels, which enables our model to learn a
wide range of realistic activities without any fine-tuning
or domain adaptation. The RGB images used in this
study were obtained from various sources and captured
in unconstrained real-world scenarios, showing various
situations such as diferent indoor and outdoor
environments, lighting conditions, skin tones, hand-to-hand and
tion blur. The images were well-annotated and collected
from an egocentric perspective. The upper limb
segmentation task dataset was compiled from three
diferent sources: EDSH [8], TEgO [9], and EgoCam. EDSH
according to the following equation:</p>
      <sec id="sec-6-1">
        <title>This policy adjusts the learning rate   during training</title>
        <p>=  0 × (1 −

 
)
(1)</p>
      </sec>
      <sec id="sec-6-2">
        <title>Here,   represents the learning rate at the current iter</title>
        <p>ation step  ,  represents the total number of iterations
power value, set to 0.9.</p>
      </sec>
    </sec>
    <sec id="sec-7">
      <title>During network training, we utilized pre-trained</title>
      <p>weights from the ImageNet [13] and MS-COCO [14]
datasets and accelerated the process with one Nvidia
Titan Xp GPU with 12GB memory. We employed Python
3.6 and the tensorflow [ 15] machine learning library,
which was tested on Microsoft Windows 10 Pro. Finally,
we applied data augmentation by randomly flipping
images and labels left/right during training to prevent model
overfitting. Our trained network achieved impressive
results for both whole upper limb and hand-only
segmentation tasks in egocentric view and unconstrained real-life
art (SOTA). We assessed our outcomes against the SOTA
arm and hand segmentation techniques in egocentric
vision, namely Ego2Hands [16] and EgoArm [10].
Moreover, we also evaluated HGR-Net [17], which presented
has become an increasingly popular way to enhance the
user experience. However, these devices are not always
accessible due to their high cost and usability issues, but
new technologies have emerged that allow HGR through
general-purpose low-cost devices. In this context, we
propose a deep learning approach that enables HGR to use a
simple, low-cost RGB camera. This approach promises to
democratize access to HGR and enhance the user
experience of a wide range of HCI systems. To create a system
for HGR independent of camera features, we developed a
pipeline that uses landmarks predicted by the MediaPipe
Hands solution as input to a feed-forward neural network
(FFNN). Our FFNN is trained on a dataset of 15 static and
dynamic gestures and can predict corresponding hand
gestures based on the input landmarks. These gestures
involve a range of combinations of open and closed
finFigure 1: Visual examples of the Upper Limb Segmentation gers and hands. The input for the model is represented by
test set, including input images and corresponding ground- 21 hand landmarks obtained from MediaPipe, and each
truth (GT) segmentation masks. The first two images are from landmark is made up of  ,  , and  coordinates. However,
EgoCam, and the last three are from TEgO. only the  and  coordinates were used for this model,
and the  coordinate was ignored. The FFNN is
characterized by a single hidden layer of 32 units that uses
the pixel space of the hand landmark model as input. A
cation of the limbs evident from the per-class metrics in rectified linear unit (ReLU) activation function was used
Table 1. It is possible that HGR-Net was not specifically on the hidden layer, while a Softmax [ 19] activation
funcdesigned to segment limbs captured from an egocentric tion was used on the last FFNN layer. Dynamic gestures
perspective, and a general-purpose approach may not be were handled diferently in this model; instead of using
enough to produce optimal results. EgoArm performed recurrent neural networks (RNNs) [20] and other data
well in identifying a large part of the limb but had dif- sequences-based approaches, a simple FFNN was used.
ifculty with background pixels and tended to classify This decision was made due to the real-time performance
objects incorrectly as limbs. Ego2Hands also struggled requirement of the dynamic gestures, as XR collaborative
with accurate segmentation (as seen in the last image of applications are well-suited for synchronous dynamic
Figure 1). In contrast, our network performed exception- gestures. The proposed gestures were divided into two
ally well in various scenarios, including diferent lighting templates; static gestures and dynamic gestures. Static
conditions, skin tones, and occlusions caused by objects, gestures were represented by a fixed hand pose that the
as illustrated in the fith row of Figure 1. FFNN directly predicts. Dynamic gestures were
char</p>
      <p>Our work is the first to evaluate and prove the efec- acterized by moving hands and can be further divided
tiveness of a deep learning model for upper limb segmen- into single and combo gestures based on the number of
tation in such cases. It provides a promising direction for tracked hands. A dataset of 130, 000 hand pose samples
future research in egocentric vision-based HCI. was manually labelled and used to train the FFNN model.</p>
      <p>The study case was published in the Virtual Reality The FFNN was trained for 2000 epochs using the Adam
journal with the title “Egocentric upper limb segmentation optimizer [21] algorithm with a learning rate of 0.0001,
in unconstrained real-life scenarios” [18]. resulting in a prediction accuracy of 98%. Testing was
conducted in real-time using RGB sensors from the Intel
3. XR Hand Gesture Recognition RealSense D455 camera, the 40MP Huawei P30 back
camera, and the 1080p MacBook Pro (M1 Pro) camera. The
System HGR system proposed in this study was designed for
collaborative use in a multi-user XR 3D environment. Each
In recent years, there has been a growing interest in devel- gesture was mapped to a specific action in the XR 3D
enoping HCI systems that provide users with more intuitive vironment, and the approach was validated in a Unity 3D
and natural ways to interact with technology. Hand ges- game engine environment. The Netcode mid-level
netture recognition (HGR) has emerged as one of the most working library was utilized to enable multi-user scene
promising techniques for achieving this goal. With the authoring. The application was tested with two clients
latest HMDs, such as Oculus Quest 2 and Vive Focus equipped with a simple RGB camera for hand landmark
3, incorporating onboard hand-tracking sensors, HGR tracking and gesture prediction. Users could interact with
dressing rooms, as it can learn from vast amounts of data
and adapt to diferent body shapes and clothing styles.</p>
      <p>This approach can lead to more accurate simulations and
a more realistic virtual shopping experience.</p>
      <p>Our research aims to address the on-the-market
solutions’ limitations, such as the absence of dress animations,
the presence of artefacts due to the incorrect tracking
of body measurements, and clothes that do not adapt
to the user’s body. We developed a 3D virtual dressing
room application called TryItOn using Unreal Engine
4.27 (UE4), known for its ability to create hyper-realistic
Figure 2: An example of our HGR system in use. This Figure environments that provide an immersive virtual reality
shows how to move objects in a virtual scene using hand experience. With TryItOn, users can try on digital
gargestures and showing their tracking in AR. If necessary, it is ments and choose from various sizes. A single RGB-D
ipnosAsRib. le to view and interact with the elements of the scene camera system accurately captures the user’s body
measurements and tracks their movements in real-time. This
information is utilized to create a 3D model of the user
that is as realistic as possible. Additionally, a third-party
visible gameObjects using the designed hand gestures, plugin called uDraper1 is used for modelling and
simand a set of user interaction actions were implemented ulating garment movement based on the specific
physand associated with the hand gestures. The scene could ical characteristics of the fabric. Using deep learning
be viewed using various devices and modes, including in our system allows for accurate anthropometric
meadesktop mode, virtual, AR, and MR, using a smartphone, surements and tracking of the user’s body movements.
HMD, or Google Cardboard. The HGR system was also The pipeline for TryItOn consists of two types of
optrained to recognize egocentric hand tracking. Figure 2 erations: one-time and real-time. The former type of
shows an example of the scene in desktop mode, with operation is executed solely during the modelling phase.
hand movement visualization for debugging purposes. This includes the creation of the base character before</p>
      <p>The proposed work was presented at the 2022 IEEE the release of the application, as well as the modelling
International Conference on Metrology for Extended of new garments during the creation and updating of
Reality, Artificial Intelligence and Neural Engineer- the clothes catalogue. Meanwhile, real-time operations
ing (MetroXRAINE) with the title “An easy Hand Ges- are executed during the application runtime. As
previture Recognition System for XR-based collaborative pur- ously stated, TryItOn is built on the UE4 platform. Hence,
poses” [22]. the models that were created during non-runtime
operations are imported into the UE4 project, and real-time
4. Virtual Dressing Room with operations are executed within the game engine. In the
group of one-time operations, a base 3D model character
Body Tracking is created with realistic body measurements. This was
created using the MB-Lab2 add-on for Blender, which
Virtual Dressing Rooms (VDRs) are an emerging technol- ofers a character editor that can generate a realistic 3D
ogy that allows users to try on clothing virtually without rigged character model. This can be morphed through
physically wearing the clothes. The system uses com- the Blender shape keys, with each corresponding to a
puter vision and deep learning to track the user’s body specific anthropometric measurement parameter. Our
and simulate the clothing in real-time. This technology methodology involves exporting the shape keys, also
has the potential to revolutionize the retail industry, pro- known as “morph targets”, in an FBX file, alongside the
viding customers with a more personalized shopping mesh and skeleton, which can be imported into a UE4
experience and reducing the need for physical inventory. project. To ensure proper sizing and alignment, we
creTo improve the usability of VDRs, there have been ef- ated a Python script that converts the Blender and UE4
forts to incorporate HCI principles and anthropometric coordinate systems and includes a collision mesh with
measurement systems. HCI aims to improve the interac- lower LOD to optimize performance. The collision mesh
tion between humans and computers, making the virtual includes relative morph targets, allowing the garment
dressing room more intuitive and user-friendly. On the physics engine to interact with complex meshes while
other hand, anthropometric measurement systems use remaining hidden from the user’s view. To ensure easy
body measurements better to simulate the fit of clothing
on a specific individual. In this context, deep learning
has been shown to be a promising technology for virtual</p>
    </sec>
    <sec id="sec-8">
      <title>1https://udraper.com/</title>
      <p>2https://mb-lab-community.github.io/MB-Lab.github.io/
integration of the base avatar with existing plugins in as keypoints. The ZED camera can provide 2D and 3D
UE4, we replaced the skeleton generated by MB-Lab with information on each detected keypoint and local rotation
the standard mannequin skeleton included in UE4, which between neighbouring bones. This information,
includinvolved renaming and adjusting bone orientations and ing each person’s 3D position and velocity, is shared in
eliminating unnecessary bones. Finally, the modified the outputs. To detect keypoints, the body tracking
modskeleton was set to the same starting pose as the UE4 ule also employs a neural network. It utilizes the depth
mannequin to maintain consistency within the virtual en- and positional tracking of the ZED SDK module to obtain
vironment. The one-time operations group also includes the final 3D position of each keypoint. The Azure Kinect
garment modelling, which can be accomplished using the body tracking system uses deep learning for detecting
uDraper modelling software. This software enables the and tracking human bodies. The system begins by
accreation of a 3D garment by starting from a 2D pattern, quiring depth and infrared images through the Azure
which can also be designed using other software. After Kinect SDK. The infrared image is then passed through a
designing the 2D pattern, the diferent sections must be convolutional neural network, which extracts the users’
stitched together and wrapped around the 3D character 2D joint coordinates and silhouette. Each pixel in the 2D
to create the 3D garment. image is assigned the corresponding depth value from</p>
      <p>The real-time operations group includes the anthro- the depth frame, which provides its position in 3D space.
pometric measurement calculation, the body tracking, The results are then post-processed to produce accurate
and the physically-based garment simulation. To begin human body skeletons. The MediaPipe Pose estimation
the calculation of a customer’s anthropometric measure- system uses a convolutional neural network to detect
ments, advanced deep learning techniques for computer the joints in the image and estimate their 3D positions.
vision tasks are employed using the FrankMocap frame- The system is trained on large labelled image datasets
work [23, 24]. This involves analyzing 2D images of the to ensure accurate joint detection and tracking. Once
customer captured through an RGB camera. FrankMocap the joints are detected, and their 3D positions are
estiemploys deep learning models trained to reconstruct a hu- mated, BTSP maps them to the UE4 mannequin skeleton
man 3D mesh and a corresponding 3D skeleton, complete to animate the avatar.
with body joints, in real-time. Anthropometric Measure- Regarding garment simulation, we used the physics
ment Calculation (AMC), a Python algorithm, has been simulation method of uDraper plugin to create realistic
developed to compute the anthropometric measurements, movement and deformation of virtual garments,
allowwhich are classified as linear or circular. AMC calculates ing them to interact with the avatar’s body and other
linear measurements by directly measuring the distance objects in the scene. This approach is diferent from
between relevant joints on the skeleton. For circular traditional keyframe animation techniques. The plugin
measurements, AMC uses FrankMocap’s body joints as calculates deformation based on internal and external
landmarks to detect the points on the 3D mesh for ac- forces and requires pre-modeling and pre-simulation of
curate measurements. These measurements are used to the virtual garment on the avatar’s collision mesh. To
morph the customer’s 3D avatar within the VICO-DR address the limitation of pre-modeling, we developed a
application. Certain guidelines must be followed to en- solution that involves implementing a base avatar that
sure accurate measurements, such as being at the correct can be deformed at runtime, allowing garments to adapt
distance from the camera, wearing form-fitting clothing, to the avatar’s shape during simulation.
and having proper lighting conditions. The system in- Figure 3 shows the interface of the TryItOn application.
cludes instructions, prompts, and real-time feedback to The proposed system is a work-in-progress project
ensure precise anthropometric measurement calculation, that was presented at the Extended Reality: First
Internadelivering personalized and engaging virtual avatars for tional Conference, XR Salento 2022, with the title
“TryIcustomers. As we needed the 3D character model to fol- tOn: A Virtual Dressing Room with Motion Tracking and
low the user’s movements, we developed a Body Track- Physically Based Garment Simulation” [25].
ing System Plugin (BTSP) for UE4. This plugin supports
three types of cameras: the Azure Kinect camera, the
ZED camera, and a simple RGB camera. We utilized their References
proprietary APKs for body tracking for the Azure Kinect
and ZED cameras. However, for the simple RGB camera, [1] R. Gruetzemacher, J. Whittlestone, The
transformawe use the MediaPipe Pose estimation system. Both the tive potential of artificial intelligence, Futures 135
camera APKs and MediaPipe utilize deep learning to ana- (2022) 102884.
lyze the images captured by the cameras and estimate the [2] N. Capece, F. Banterle, P. Cignoni, F. Ganovelli,
3D position of the joints in real-time. The ZED module R. Scopigno, U. Erra, Deepflash: Turning a flash
for body tracking focuses on detecting and tracking a per- selfie into a studio portrait, Signal Processing:
Imson’s bones, represented by two endpoints, also known age Communication 77 (2019) 28–39.
[13] O. Russakovsky, J. Deng, H. Su, J. Krause,</p>
      <p>S. Satheesh, S. Ma, Z. Huang, A. Karpathy,
A. Khosla, M. Bernstein, et al., Imagenet large scale
visual recognition challenge, International journal
of computer vision 115 (2015) 211–252.
[14] T.-Y. Lin, M. Maire, S. Belongie, J. Hays, P. Perona,</p>
      <p>D. Ramanan, P. Dollár, C. L. Zitnick, Microsoft coco:
Common objects in context, in: Computer Vision–
ECCV 2014: 13th European Conference, Zurich,
Switzerland, September 6-12, 2014, Proceedings,
Figure 3: Interface of the TryItOn application. Part V 13, Springer, 2014, pp. 740–755.
[15] M. Abadi, A. Agarwal, P. Barham, E. Brevdo,</p>
      <p>Z. Chen, C. Citro, G. S. Corrado, A. Davis, J. Dean,
[3] G. Caggianese, N. Capece, U. Erra, L. Gallo, M. Ri- M. Devin, et al., Tensorflow: Large-scale
manaldi, Freehand-steering locomotion techniques for chine learning on heterogeneous distributed
sysimmersive virtual environments: A comparative tems, arXiv preprint arXiv:1603.04467 (2016).
evaluation, International Journal of Human–Com- [16] F. Lin, T. Martinez, Ego2hands: A dataset for
puter Interaction 36 (2020) 1734–1755. egocentric two-hand segmentation and detection,
[4] L.-C. Chen, Y. Zhu, G. Papandreou, F. Schrof, arXiv preprint arXiv:2011.07252 (2020).</p>
      <p>H. Adam, Encoder-decoder with atrous separable [17] A. Dadashzadeh, A. T. Targhi, M. Tahmasbi,
convolution for semantic image segmentation, in: M. Mirmehdi, Hgr-net: a fusion network for hand
Proceedings of the European conference on com- gesture segmentation and recognition, IET
Computer vision (ECCV), 2018, pp. 801–818. puter Vision 13 (2019) 700–707.
[5] L.-C. Chen, G. Papandreou, I. Kokkinos, K. Murphy, [18] M. Gruosso, N. Capece, U. Erra, Egocentric
upA. L. Yuille, Deeplab: Semantic image segmentation per limb segmentation in unconstrained real-life
with deep convolutional nets, atrous convolution, scenarios, Virtual Reality (2022) 1–13.
and fully connected crfs, IEEE transactions on pat- [19] C. M. Bishop, N. M. Nasrabadi, Pattern recognition
tern analysis and machine intelligence 40 (2017) and machine learning, volume 4, Springer, 2006.
834–848. [20] F. Yang, Y. Wu, S. Sakti, S. Nakamura, Make
[6] F. Chollet, Xception: Deep learning with depth- skeleton-based action recognition model smaller,
wise separable convolutions, in: Proceedings of the faster and better, in: Proceedings of the ACM
mulIEEE conference on computer vision and pattern timedia asia, 2019, pp. 1–6.</p>
      <p>recognition, 2017, pp. 1251–1258. [21] D. P. Kingma, J. Ba, Adam: A method for
stochas[7] M. Gruosso, N. Capece, U. Erra, Exploring up- tic optimization, arXiv preprint arXiv:1412.6980
per limb segmentation with deep learning for aug- (2014).</p>
      <p>mented virtuality (2021). [22] N. Capece, G. Manfredi, V. Macellaro, P. Carratù,
[8] C. Li, K. M. Kitani, Pixel-level hand detection in An easy hand gesture recognition system for
xrego-centric videos, in: Proceedings of the IEEE con- based collaborative purposes, in: 2022 IEEE
Interference on computer vision and pattern recognition, national Conference on Metrology for Extended
2013, pp. 3570–3577. Reality, Artificial Intelligence and Neural
Engineer[9] K. Lee, H. Kacorri, Hands holding clues for object ing (MetroXRAINE), IEEE, 2022, pp. 121–126.
recognition in teachable machines, in: Proceedings [23] Y. Rong, T. Shiratori, H. Joo, Frankmocap: A
monocof the 2019 CHI Conference on Human Factors in ular 3d whole-body pose estimation system via
reComputing Systems, 2019, pp. 1–12. gression and integration, in: IEEE International
[10] E. Gonzalez-Sosa, P. Perez, R. Tolosana, R. Kachach, Conference on Computer Vision Workshops, 2021.</p>
      <p>A. Villegas, Enhanced self-perception in mixed re- [24] H. Joo, N. Neverova, A. Vedaldi, Exemplar
fineality: Egocentric arm segmentation and database tuning for 3d human pose fitting towards
in-thewith automatic labeling, IEEE Access 8 (2020) wild 3d human pose estimation, 3DV (2021).
146887–146900. [25] G. Manfredi, N. Capece, U. Erra, G. Gilio, V. Baldi,
[11] L.-C. Chen, G. Papandreou, F. Schrof, H. Adam, Re- S. G. Di Domenico, Tryiton: A virtual dressing
thinking atrous convolution for semantic image seg- room with motion tracking and physically based
mentation, arXiv preprint arXiv:1706.05587 (2017). garment simulation, in: Extended Reality: First
[12] W. Liu, A. Rabinovich, A. C. Berg, Parsenet: International Conference, XR Salento 2022, Lecce,
Looking wider to see better, arXiv preprint Italy, July 6–8, 2022, Proceedings, Part I, Springer,
arXiv:1506.04579 (2015). 2022, pp. 63–76.</p>
    </sec>
  </body>
  <back>
    <ref-list />
  </back>
</article>