<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <journal-title-group>
        <journal-title>Ital-IA</journal-title>
      </journal-title-group>
    </journal-meta>
    <article-meta>
      <title-group>
        <article-title>Algorithms for Industry</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Giovanni Maria Farinella</string-name>
          <email>giovanni.farinella@unict.it</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Antonino Furnari</string-name>
          <email>antonino.furnari@unict.it</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Figure 1: Main areas of research of the</institution>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Wearable and Mobile Systems, Egocentric Vision, Human Behavior Understanding, Human Behavior Anticipation</institution>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2023</year>
      </pub-date>
      <volume>3</volume>
      <fpage>29</fpage>
      <lpage>31</lpage>
      <abstract>
        <p>Integrating artificial intelligence and computer vision on wearable devices in industrial environments can increase productivity, eficiency, and safety in the workplace. Despite the availability of wearable devices such as Microsoft HoloLens, Magic Leap, and nreal, the application of artificial intelligence algorithms on wearable devices equipped with cameras is an open research topic. To address this gap, the FPV@IPLAB group at the University of Catania has conducted research on the construction of machine learning and computer vision algorithms for portable devices. The research has focused on three main areas: localization and navigation, user-object interaction understanding, and user-object interaction anticipation. The work conducted by the FPV@IPLAB group aims to enhance the use wearable devices and to develop artificial intelligence techniques that can improve workplace eficiency and safety.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <sec id="sec-1-1">
        <title>Artificial intelligence can be used in industrial environ</title>
        <p>ments to increase eficiency and safety in the workplace.
Wearable devices which acquire and analyze images and
videos in the surrounding environment can be used to
develop intelligent systems able to assist workers during
their activities. In this context, wearable devices can
allow to overlay virtual elements on the observed scene
through augmented reality, and provide services based
on artificial intelligence via to the analysis of images and
videos acquired by the user. Moreover, due to their
intrinsic mobility, wearable devices tend to be naturally
exposed to large amounts of data specific to the user’s
visual experience, which can in principle provide an
important source of knowledge for training and adapting
machine learning and artificial intelligence algorithms.</p>
      </sec>
      <sec id="sec-1-2">
        <title>Even if the market ofers today devices such as Microsoft</title>
        <p>HoloLens 1, Magic Leap 2, and Nreal 3, which are be
nEvelop-O
(A. Furnari)
explored field.</p>
      </sec>
      <sec id="sec-1-3">
        <title>This article presents research conducted by the FPV@IPLAB group at the University of Catania on the construction of Machine Learning and Computer Vision</title>
        <p>CEUR
Workshop
Proce dings
htp:/ceur-ws.org
ISN1613-073</p>
        <p>CEUR</p>
        <p>Workshop Proceedings (CEUR-WS.org)
https://www.microsoft.com/en-us/hololens/</p>
        <sec id="sec-1-3-1">
          <title>Localization and</title>
        </sec>
        <sec id="sec-1-3-2">
          <title>Navigation</title>
        </sec>
        <sec id="sec-1-3-3">
          <title>User-Object</title>
        </sec>
        <sec id="sec-1-3-4">
          <title>Interaction</title>
        </sec>
        <sec id="sec-1-3-5">
          <title>Understanding</title>
        </sec>
        <sec id="sec-1-3-6">
          <title>User-Object</title>
        </sec>
        <sec id="sec-1-3-7">
          <title>Interaction</title>
        </sec>
        <sec id="sec-1-3-8">
          <title>Anticipation</title>
          <p>and navigation based on images acquired by portable
devices, user-object interaction understanding, and
userobject interaction anticipation (see Figure 1). For further
information on the research conducted by the IPLAB
laboratory in the field of First Person Vision, please visit the
web page http://iplab.dmi.unict.it/fpv/.</p>
        </sec>
      </sec>
    </sec>
    <sec id="sec-2">
      <title>2. Localization and Navigation</title>
      <sec id="sec-2-1">
        <title>The ability to identify the position of workers within a</title>
        <p>and guide them to a destination, can be achieved through
the processing of images acquired from wearable devices.</p>
        <p>While outdoor localization generally relies on GPS
systion of artificial intelligence algorithms in the context of
suitable for use in industrial environments, the applica- to the use of these technologies in industrial
environments. In particular, research conducted in three areas
wearable devices equipped with vision is still an under- relevant to the industry will be presented: localization
algorithms for portable devices, with particular reference
© 2022 Copyright for this paper by its authors. Use permitted under Creative Commons License building, provide them with contextualized information
Attribution 4.0 International (CC BY 4.0).</p>
        <p>g
n
iir
n
a
T
t
s
e
T
Real-worldenvironment</p>
        <p>Realroboticplatform
obserRveaatilo-wnofrroldmthe Mid-levelrepresentations</p>
        <p>robot extraction
2.2. Camera pose estimation</p>
        <p>
          We have also studied the problem of localization through
camera position estimation. In particular, the
localization of shopping carts inside a supermarket was studied
tems, traditional approaches to indoor localization are using cameras mounted on the carts themselves [6, 7].
based on radio-frequency technologies, such as Wi-Fi This allows the development of intelligent systems
capaand BLE, which requires the installation of ad hoc in- ble of guiding customers inside the store and studying
frastructures and cannot always guarantee a satisfactory their behavior to ofer personalized services [ 8].
Localizaaccuracy. On the other hand, image-based localization tion is performed using image retrieval techniques and a
allows to obtain more accurate results without the need metric built through deep metric learning.5 These same
for dedicated infrastructures. technologies can be used to develop systems capable of
localizing operators and guiding them inside a warehouse
2.1. Context-Based Localization or other industrial environment [9]. In addition,
camera pose estimation from wearable devices has also been
IPLAB has worked on recognizing the environment in studied considering simulated data generated from a 3D
which the user is located, called “personal location”, model of a real building [10]. The generation of synthetic
which corresponds to specific places where the user data allows obtaining large amounts of labeled data
suitperforms certain activities, such as an ofice or a work- able for the development of localization algorithms. The
bench [
          <xref ref-type="bibr" rid="ref1 ref2">1, 2</xref>
          ]. The set of relevant personal locations varies use of domain adaptation techniques also allows using
from user to user, and the developed algorithms allow synthetic data to train localization models that can work
to identify the environments from a set of examples pro- on real data [11].
vided by the user themselves, acquiring a 30-second video
for each environment. The algorithm is able to recog- 2.3. Navigation
nize the personal locations of interest and discard other
environments, even those not specified during system Beyond localizing workers in an industrial site, wearable
configuration. This approach allows for real-time lo- systems should be able to navigate them towards a
descalization and temporal segmentation of the video into tination. A navigation system may also guide workers
coherent temporal units based on the user’s context. Fig- to follow secure paths, e.g., avoiding dangerous areas
ure 2 illustrates the developed system, which allows to and suspended loads. Navigation algorithms can also
automatically index the video in order to easily navigate be used to enable robots to move within the industrial
it, segment it into coherent clips, and estimate the time environment and support the workers by escorting them
spent in each personal location. These applications can or retrieving tools for them. The research activity of the
be useful in industrial contexts for work analysis and FPV@IPLAB group in this area has focused on techniques
staf training. The localization system has also been used for embodied visual navigation in virtual replicas of real
for environment recognition in a museum for visitor lo- environments [13], adaptation techniques for transfer
calization [
          <xref ref-type="bibr" rid="ref3">3, 4</xref>
          ].4 Similar approaches have been exploited to real scenarios [12] (see Figure 3), and human-aware
in the context of natural sites [5]. robot navigation [14].
        </p>
      </sec>
      <sec id="sec-2-2">
        <title>4Demonstration video: https://youtu.be/VYZ6Awqy1ko</title>
      </sec>
      <sec id="sec-2-3">
        <title>5Demonstration video: https://youtu.be/BxbdgWxFhgc</title>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>3. User-Object Interaction Understanding</title>
      <p>3.1. Object Detection and Tracking
Recognizing the interactions between the user and the
objects through wearable devices can be useful for a
variety of applications in industrial scenarios, ranging from
providing additional information on objects of interest impact of visual object tracking algorithms in first-person
through augmented reality, to monitoring user behavior vision [28]. Visual tracking algorithms allow to keep a
and assessing the correct execution of procedures. reference to a given object of interest in each frame of a
video, which can be useful for the analysis of user-object
interactions and for anticipation.</p>
      <p>As a first step towards user-object interaction
understanding, we focused on the detection of objects seen from
a first-person perspective, including synthetic-to-real
domain adaptation for object detection and
segmentation [15, 16, 17, 18], panoptic segmentation in industrial
environments [19] and safety monitoring in construction
sites [20]. Other works have considered the problem of
recognizing objects in the scene and estimating which of
them are currently being observed by the user [21, 22]. 6
This type of analysis allows for acquiring behavioral
information on users by inferring which points of interest
have been observed and for how long. Moreover,
artificial intelligence algorithms can use such information to
provide suggestions on the next things to see or to use.
3.2. Interaction Understanding
The FPV@IPLAB group has also focused on algorithms
for understanding user-object interactions from synthetic
images [23] (see Figure 4), and by leveraging the software
layer provided by augmented reality devices [24]. The
group is currently investigating the integration of natural
language processing and object recognition to develop
systems able to provide assistance to the user on the
execution of specific procedures [ 25]. The problem of
temporal segmentation of videos based on actions performed by
users has also been studied in [26] and later in [27] in the
form of temporal action detection. The output of these
algorithms can be used as input for advanced artificial
intelligence systems capable of analyzing user actions,
inferring the next relevant actions, or determining any
critical points in the workflow. We also investigated the</p>
      <sec id="sec-3-1">
        <title>6Demo video: https://youtu.be/nBkYOdKYu0s</title>
        <p>3.3. Datasets to Study User-Behavior</p>
        <p>Understanding
To facilitate the study of user-object interactions from
ifrst-person vision, the FPV@IPLAB group has
contributed by collecting and labeling diferent datasets of
egocentric videos. The MECCANO dataset [29, 30] is
a multimodal dataset of egocentric videos collected in
an industrial-like procedural scenario where subjects
were asked to assemble a toy model of a motorbike. The
dataset is provided with gaze signals, depth maps, and
RGB videos acquired simultaneously with a custom
headset, explicitly labeled for fundamental tasks in the context
of human behavior understanding from a first-person
view, such as recognizing and anticipating human-object
interactions.7 Figure 5 illustrates the parts involved in
the assembly of the toy model. We collaborated in the
creation of EPIC-KITCHENS [31, 32], a large dataset of
ifrst-person-view videos for action recognition, action
anticipation, and object recognition. The dataset was
acquired from 32 subjects in 3 diferent countries (Italy, UK,
and Canada) and contains 55 hours of video, annotations
for approximately 40,000 actions, and 500,000 objects. 8
An extension of the dataset including more videos,
labels and benchmark tasks has been subsequently
proposed [27]. The FPV@IPLAB group has also participated
in the definition, collection, labeling and benchmarking
of EGO4D [33], a large-scale egocentric video dataset that
ofers a vast amount of daily-life activity videos captured
by 931 individuals from 74 diferent locations and 9
countries, accompanied by audio, 3D meshes, eye gaze, stereo</p>
      </sec>
      <sec id="sec-3-2">
        <title>7Dataset: https://iplab.dmi.unict.it/MECCANO/ 8Dataset: https://epic-kitchens.github.io/</title>
        <p>Short-term
ActionAnticipation
Region</p>
        <p>Trajectory
ALcotniogn-teCramptFiountuinreg InteraPcrteiodnicHtioontspots Tra(jCecatmoreyraFWoreeacraesrt)ing TrajectorHyaFnodrsecasting</p>
        <p>NextPArecdtiivcetioOnbject FPurteudreicGtioazne (OTrtahjeercPtoerrysoFnorse/Ocabsjeticntgs)</p>
        <p>Tasks not aimed to future predictions
ActionRecognition ActionTeSmegpmoreanltation</p>
        <p>Level of prediction</p>
        <p>Anticipation
Prospection
Expectation</p>
        <p>Forecasting
and synchronized videos. Together with the dataset, new
benchmark challenges for understanding the first-person
visual experience in the past, present and future are
presented. The work has been done in collaboration with a
consortium of 13 universities around the world.</p>
        <p>U-LSTM
R-LSTM</p>
        <p>U-LSTM
R-LSTM</p>
        <p>U-LSTM
R-LSTM</p>
        <p>U-LSTM
R-LSTM</p>
        <p>U-LSTM</p>
        <p>U-LSTM
S  1,1</p>
        <p>2,1
U-LSTM
U-LSTM
S  1,2</p>
        <p>2,2
U-LSTM
S  1,3
 2,3</p>
        <p>1,1
U-LSTM L (  = 0.75 )
×  1
× + (  = 0.75 )
 2,1
U-LSTM L (  = 0.75 )</p>
        <p>1,2
L (  = 0.5 )</p>
        <p>2
×× + (  = 0.5 )
 2,2
L (  = 0.5 )</p>
        <p>1,3
L (  = 0.25 )
×  3
× + (  = 0.25 )
 2,3</p>
        <p>L (  = 0.25 )
  Input video snippets L Linear transformation</p>
        <p>Message passing</p>
        <p>S  1, MNoedtawliotyrkA(tMteAnTtiTo)n S SoftMax</p>
        <p>2,
from first-person videos [ 36]. A novel architecture to
tackle the problem based on recurrent networks has been
4. Object Interaction Anticipation proposed in [37, 38] and extended in [39]. The
developed algorithms allow predicting the set of likely next
A desirable feature for a wearable device equipped with actions based on the observation of videos before they
artificial intelligence is the ability to anticipate what will occur.12 Applications of this approach to the domain of
happen in the scene in advance. This allows building sys- personal health have also been explored in [40]. The
analtems that can guide the user through complex workflows ysis of the problem has later been extended considering
and notify them if an incorrect or dangerous action is untrimmed input videos [41] and real-time computation
about to be taken. Our research group has investigated constraints [42].
algorithms to predict which objects in the scene the user
will interact with in the short term. We surveyed the
main tasks related to the prediction of the future form 5. Conclusion
egocentric videos in [34] (see Figure 6) and investigated
approaches to tackle specific prediction tasks. We presented the research conducted by the FPV@IPLAB</p>
        <p>The FPV@IPLAB group investigated algorithms to pre- group in the development of artificial intelligence
algodict which objects in the scene will be used by the user rithms for wearable vision devices. The problems
adin the short term. In particular, the studies conducted dressed have potential applications in industrial contexts
in [35] have highlighted how the analysis of trajectories and revolve around three main themes related to
localizaof objects identified from first-person videos allows ob- tion and navigation, user-object interaction understanding,
taining information about the next objects that will be and user-object interaction anticipation.
used by the user in a dynamic context. 9 We later
explored the task in the context of procedural videos using References
object-detection based approaches in [29, 30]. The task
has then been formalized as the multi-task problem of
recognizing objects, and predicting future actions and
time-to-contact for each of them in [33], which has also
lead to the definition of the “short-term object interaction
anticipation” challenge.10</p>
        <p>The topic of anticipating interactions with objects has
also been investigated through the definition of a
challenge on egocentric action anticipation 11 related to the
EPIC-KITCHENS dataset [31] and through the study of
architectures and evaluation measures suitable for
addressing the problem of anticipated prediction of actions
9Video: http://iplab.dmi.unict.it/NextActiveObjectPrediction/
10Challenge: https://eval.ai/web/challenges/challenge-page/1623/</p>
        <p>overview
11Challenge: https://codalab.lisn.upsaclay.fr/competitions/707
12Demo video: https://youtu.be/buIEKFHTVIg
sites, Journal on Computing and Cultural Heritage through habitat, in: International Conference on
(2019). Pattern Recognition (ICPR), 2020. URL: https://iplab.
[4] F. Ragusa, A. Furnari, S. Battiato, G. Signorello, G. M. dmi.unict.it/EmbodiedVN/.</p>
        <p>Farinella, EGO-CH: Dataset and fundamental tasks [14] R. Möller, A. Furnari, S. Battiato, A. Härmä, G. M.
for visitors behavioral understanding using egocen- Farinella, A survey on human-aware robot
navigatric vision, Pattern Recognition Letters (2020). URL: tion, Robotics and Autonomous Systems (2021).
https://iplab.dmi.unict.it/EGO-CH/. [15] F. Ragusa, D. D. Mauro, A. Palermo, A. Furnari, G. M.
[5] F. L. Milotta, A. Furnari, S. Battiato, G. Signorello, Farinella, Semantic object segmentation in cultural
G. M. Farinella, Egocentric visitors localization sites using real and synthetic data, in: International
in natural sites, Journal of Visual Communica- Conference on Pattern Recognition (ICPR), 2020.
tion and Image Representation (2019) 102664. URL: [16] G. Pasqualino, A. Furnari, G. Signorello, G. M.
https://iplab.dmi.unict.it/EgoNature/. doi:h t t p s : / / Farinella, An unsupervised domain
adaptad o i . o r g / 1 0 . 1 0 1 6 / j . j v c i r . 2 0 1 9 . 1 0 2 6 6 4 . tion scheme for single-stage artwork
recogni[6] E. Spera, A. Furnari, S. Battiato, G. M. Farinella, tion in cultural sites, Image and Vision
Egocentric shopping cart localization, in: Computing (2021). URL: https://iplab.dmi.unict.it/
International Conference on Pattern Recog- EGO-CH-OBJ-UDA/.
nition, 2018. URL: http://iplab.dmi.unict.it/ [17] G. Pasqualino, A. Furnari, G. M. Farinella, A multi
EgocentricShoppingCartLocalization/. camera unsupervised domain adaptation pipeline
[7] E. Spera, A. Furnari, S. Battiato, G. M. Farinella, for object detection in cultural sites through
adEgocart: a benchmark dataset for large-scale in- versarial learning and self-training, Computer
door image-based localization in retail stores, IEEE Vision and Image Understanding (CVIU) (2022)
Transactions on Circuits and Systems for Video 103487. URL: https://iplab.dmi.unict.it/OBJ-MDA/.
Technology 31 (2021) 1253–1267. URL: https://iplab. doi:h t t p s : / / d o i . o r g / 1 0 . 1 0 1 6 / j . c v i u . 2 0 2 2 . 1 0 3 4 8 7 .
dmi.unict.it/EgocentricShoppingCartLocalization/. [18] G. Pasqualino, A. Furnari, G. M. Farinella,
Unsuper[8] V. Santarcangelo, G. M. Farinella, A. Furnari, vised multi-camera domain adaptation for object
S. Battiato, Market basket analysis from ego- detection in cultural sites, in: International
Concentric videos, Pattern Recognition Letters ference on Image Analysis and Processing (ICIAP),
112 (2018) 83–90. URL: http://iplab.dmi.unict. 2022. URL: https://iplab.dmi.unict.it/OBJ-MDA/.
it/vmba15. doi:h t t p s : / / d o i . o r g / 1 0 . 1 0 1 6 / j . p a t r e c . [19] C. Quattrocchi, D. D. Mauro, A. Furnari, G. M.
2 0 1 8 . 0 6 . 0 1 0 . Farinella, Panoptic segmentation in industrial
en[9] F. Ragusa, A. Furnari, A. Lopes, M. Moltisanti, E. Ra- vironments using synthetic and real data, in:
Ingusa, M. Samarotto, L. Santo, N. Picone, L. Scarso, ternational Conference on Image Analysis and
ProG. M. Farinella, Enigma: Egocentric navigator for cessing (ICIAP), 2022. URL: https://iplab.dmi.unict.
industrial guidance, monitoring and anticipation, it/ENIGMA_SEG/.
in: International Conference on Computer Vision [20] C. Quattrocchi, D. D. Mauro, A. Furnari, A. Lopes,
Theory and Applications (VISAPP), 2023. M. Moltisanti, G. M. Farinella, Put your ppe on:
[10] S. Orlando, A. Furnari, G. M. Farinella, Egocen- A tool for synthetic data generation and related
tric visitor localization and artwork detection incul- benchmark in construction site scenarios, in:
Intertural sites using synthetic data, Pattern Recogni- national Conference on Computer Vision Theory
tion Letters (2020). URL: https://iplab.dmi.unict.it/ and Applications (VISAPP), 2023.</p>
        <p>SimulatedEgocentricNavigations/. [21] F. Ragusa, A. Furnari, S. Battiato, G. Signorello, G. M.
[11] D. D. Mauro, A. Furnari, G. Signorello, G. M. Farinella, Egocentric point of interest recognition in
Farinella, Unsupervised domain adaptation for cultural sites, in: International Conf. on Computer
6dof indoor localization, in: International Con- Vision Theory and Applications, 2019.
ference on Computer Vision Theory and Applica- [22] M. Mazzamuto, F. Ragusa, A. Furnari, G. M.
tions - VISAPP, 2021. URL: https://iplab.dmi.unict. Farinella, Weakly supervised attended object
deit/EGO-CH-LOC-UDA/. tection using gaze data as annotations, in:
Interna[12] M. Rosano, A. Furnari, L. Gulino, C. Santoro, G. M. tional Conference on Image Analysis and
ProcessFarinella, Image-based navigation in real-world ing (ICIAP), 2022.
environments via multiple mid-level representa- [23] R. Leonardi, F. Ragusa, A. Furnari, G. M. Farinella,
tions: Fusion models, benchmark and eficient eval- Egocentric human-object interaction detection
uation, CoRR abs/2202.01069 (2022). URL: https: exploiting synthetic data, in: International
//arxiv.org/abs/2202.01069. a r X i v : 2 2 0 2 . 0 1 0 6 9 . Conference on Image Analysis and Processing
[13] M. Rosano, A. Furnari, L. Gulino, G. M. Farinella, On (ICIAP), 2022. URL: https://iplab.dmi.unict.it/EHOI_
embodied visual navigation in real environments SYNTH/.
[24] M. Mazzamuto, F. Ragusa, A. Resta, G. M. Farinella, S. Kottur, A. Kumar, F. Landini, C. Li, Y. Li, Z. Li,
A. Furnari, A wearable device application for K. Mangalam, R. Modhugu, J. Munro, T.
Murhuman-object interactions detection., in: Inter- rell, T. Nishiyasu, W. Price, P. R. Puentes, M.
Ranational Conference on Computer Vision Theory mazanova, L. Sari, K. Somasundaram, A.
Southerand Applications (VISAPP), 2023. land, Y. Sugano, R. Tao, M. Vo, Y. Wang, X. Wu,
[25] C. Bonanno, F. Ragusa, R. Leonardi, A. Furnari, G. M. T. Yagi, Y. Zhu, P. Arbelaez, D. Crandall, D. Damen,
Farinella, Hero: An artificial conversational assis- G. M. Farinella, B. Ghanem, V. K. Ithapu, C. V.
Jawatant to support humans in industrial scenarios, in: har, H. Joo, K. Kitani, H. Li, R. Newcombe, A. Oliva,
International Conference on Signal Processing and H. S. Park, J. M. Rehg, Y. Sato, J. Shi, M. Z. Shou,
Multimedia Applications (SIGMAP), 2022. A. Torralba, L. Torresani, M. Yan, J. Malik, Around
[26] A. Furnari, S. Battiato, G. M. Farinella, How shall we the World in 3,000 Hours of Egocentric Video, in:
evaluate egocentric action recognition?, in: ICCV IEEE/CVF International Conference on Computer
Workshops, 2017. Vision and Pattern Recognition, 2022.
[27] Damen, Doughty, Farinella, Furnari, Kazakos, Ma, [34] I. Rodin, A. Furnari, D. Mavroedis, G. M. Farinella,
Moltisanti, Munro, Perrett, Price, Wray, Rescaling Predicting the future from first person
(egocenegocentric vision: Collection, pipeline and chal- tric) vision: A survey, Computer Vision and
Imlenges for epic-kitchens-100, International Journal age Understanding 211 (2021) 103252. doi:h t t p s :
on Computer Vision (IJCV) 130 (2022) 33–55. URL: / / d o i . o r g / 1 0 . 1 0 1 6 / j . c v i u . 2 0 2 1 . 1 0 3 2 5 2 .
http://epic-kitchens.github.io/2020-100. [35] A. Furnari, S. Battiato, K. Grauman, G. M.
[28] M. Dunnhofer, A. Furnari, G. M. Farinella, C. Mich- Farinella, Next-active-object prediction from
egoeloni, Visual object tracking in first person vi- centric videos, Journal of Visual Communication
sion, International Journal of Computer Vision and Image Representation 49 (2017) 401 – 411.
(IJCV) (2022). URL: https://machinelearning.uniud. [36] A. Furnari, S. Battiato, G. M. Farinella, Leveraging
it/datasets/trek150/. uncertainty to rethink loss functions and
evalua[29] F. Ragusa, A. Furnari, S. Livatino, G. M. Farinella, tion measures for egocentric action anticipation, in:
The meccano dataset: Understanding human- ECCV Workshop on Egocentric Perception,
Interobject interactions from egocentric videos in an action and Computing (EPIC), 2018.
industrial-like domain, in: IEEE Winter Confer- [37] A. Furnari, G. M. Farinella, What would you
exence on Application of Computer Vision (WACV), pect? anticipating egocentric actions with
rolling2021. URL: https://iplab.dmi.unict.it/MECCANO. unrolling lstms and modality attention., in:
Intera r X i v : 2 0 1 0 . 0 5 6 5 4 . national Conference on Computer Vision (ICCV),
[30] F. Ragusa, A. Furnari, G. M. Farinella, Meccano: A 2019, pp. 6252–6261.</p>
        <p>multimodal egocentric dataset for humans behavior [38] A. Furnari, G. M. Farinella, Rolling-unrolling lstms
understanding in the industrial-like domain, 2022. for action anticipation from first-person video, IEEE
a r X i v : 2 2 0 9 . 0 8 6 9 1 . Transactions on Pattern Analysis and Machine
In[31] D. Damen, H. Doughty, G. M. Farinella, S. Fidler, telligence (PAMI) 43 (2021) 4021–4036. doi:1 0 . 1 1 0 9 /
A. Furnari, E. Kazakos, D. Moltisanti, J. Munro, T P A M I . 2 0 2 0 . 2 9 9 2 8 8 9 .</p>
        <p>T. Perrett, W. Price, M. Wray, Scaling egocentric [39] G. Camporese, P. Coscia, A. Furnari, G. M. Farinella,
vision: The epic-kitchens dataset, in: European L. Ballan, Knowledge distillation for action
anticiConference on Computer Vision, 2018. pation via label smoothing, in: International
Con[32] D. Damen, H. Doughty, G. M. Farinella, S. Fidler, ference on Pattern Recognition (ICPR), 2020.</p>
        <p>A. Furnari, E. Kazakos, D. Moltisanti, J. Munro, [40] I. Rodin, A. Furnari, G. M. Farinella, D. Mavroeidis,
T. Perrett, W. Price, M. Wray, The epic-kitchens Egocentric action anticipation for personal health,
dataset: Collection, challenges and baselines, IEEE in: IEEE International Conference on Acoustics,
Transactions on Pattern Analysis and Machine In- Speech, and Signal Processing (ICASSP), 2023.
telligence (PAMI) 43 (2021) 4125–4141. [41] I. Rodin, A. Furnari, D. Mavroedis, G. M. Farinella,
[33] K. Grauman, A. Westbury, E. Byrne, Z. Chavis, Untrimmed action anticipation, in: International
A. Furnari, R. Girdhar, J. Hamburger, H. Jiang, Conference on Image Analysis and Processing
M. Liu, X. Liu, M. Martin, T. Nagarajan, I. Ra- (ICIAP), 2022.
dosavovic, S. K. Ramakrishnan, F. Ryan, J. Sharma, [42] A. Furnari, G. M. Farinella, Towards streaming
M. Wray, M. Xu, E. Z. Xu, C. Zhao, S. Bansal, egocentric action anticipation, in: International
D. Batra, V. Cartillier, S. Crane, T. Do, M. Doulaty, Conference on Pattern Recognition (ICPR), 2022.
A. Erapalli, C. Feichtenhofer, A. Fragomeni, Q. Fu,
C. Fuegen, A. Gebreselasie, C. Gonzalez, J. Hillis,
X. Huang, Y. Huang, W. Jia, W. Khoo, J. Kolar,</p>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>A.</given-names>
            <surname>Furnari</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Battiato</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G. M.</given-names>
            <surname>Farinella</surname>
          </string-name>
          ,
          <article-title>Personallocation-based temporal segmentation of egocentric video for lifelogging applications</article-title>
          ,
          <source>Journal of Visual Communication and Image Representation</source>
          <volume>52</volume>
          (
          <year>2018</year>
          )
          <fpage>1</fpage>
          -
          <lpage>12</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>A.</given-names>
            <surname>Furnari</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G. M.</given-names>
            <surname>Farinella</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Battiato</surname>
          </string-name>
          ,
          <article-title>Recognizing personal locations from egocentric videos</article-title>
          ,
          <source>IEEE Transactions on Human-Machine Systems</source>
          <volume>47</volume>
          (
          <year>2017</year>
          )
          <fpage>6</fpage>
          -
          <lpage>18</lpage>
          . URL: http://iplab.dmi.unict.it/ PersonalLocations/.
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>F.</given-names>
            <surname>Ragusa</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Furnari</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Battiato</surname>
          </string-name>
          , G. Signorello,
          <string-name>
            <given-names>G. M.</given-names>
            <surname>Farinella</surname>
          </string-name>
          , Egocentric visitors localization in cultural
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>