<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <journal-title-group>
        <journal-title>Ital-IA</journal-title>
      </journal-title-group>
    </journal-meta>
    <article-meta>
      <title-group>
        <article-title>Wearable Visual Intelligence to Assist Humans in Workplaces</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Giovanni Maria Farinella</string-name>
          <email>gmfarinella@nextvisionlab.it</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Antonino Furnari</string-name>
          <email>afurnari@nextvisionlab.it</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Francesco Ragusa</string-name>
          <email>fragusa@nextvisionlab.it</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="editor">
          <string-name>Artificial Intelligence, Egocentric Vision, Wearable devices</string-name>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Next Vision s.r.l. - Spin-of of the University of Catania</institution>
          ,
          <country country="IT">Italy</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>gies: NAIROBI, an Artificial Intelligence assistant able to</institution>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2023</year>
      </pub-date>
      <volume>3</volume>
      <fpage>29</fpage>
      <lpage>31</lpage>
      <abstract>
        <p>Wearable devices equipped with a camera and a display allow to develop human-centric applications providing specific services able to improve worker's productivity and increase safety in industrial environments. Despite there are diferent wearable devices in the market, current solutions are mainly focused only on passive augmented reality and real-time remote assistance. This contribution presents Artificial Intelligence technologies for wearable devices provided by NEXT VISION s.r.l. - Spin-of of the University of Catania. In particular, we present three technologies developed by NEXT VISION able to 1) localize humans in unfamiliar environments and support them to navigate the space to reach a specific destination, 2) understanding human-object interactions to provide support to the workers in industrial workplaces during complex procedures of maintenance and 3) interact with humans thanks to a conversational intelligent agent exploiting natural language and the artificial vision to answer questions regarding the surrounding environment.</p>
      </abstract>
      <kwd-group>
        <kwd>exploiting Artificial Intelligence and Computer Vision</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <sec id="sec-1-1">
        <title>Artificial Intelligence technologies allow to support hu</title>
        <p>mans in industrial environments improving workers’
safety and their productivity. Several vision systems
sually using fixed cameras (e.g., surveillance cameras) to
observe the surrounding environment from third person
point of view.</p>
        <p>More recently, diferent wearable devices equipped
troduced in both consumers (e.g., Microsoft Hololens</p>
      </sec>
    </sec>
    <sec id="sec-2">
      <title>2. Indoor Visual Navigation</title>
      <p>
        The ability to localize workers in indoor environments
is important to support humans in performing complex
tasks and to improve safety in industrial environments
(e.g., for rescue applications). For example, a worker
could need to move from side to side in a big factory
avoiding dangerous areas with suspended loads as well
vices can provide artificial intelligence services analyzing
real world through Augmented Reality. Also these de- covered by two patents [
        <xref ref-type="bibr" rid="ref6 ref7">6, 7</xref>
        ]. Please, refer to our web
page https://www.nextvisionlab.it/ for further
informaimages and videos acquired from the user’s point of view. tion about the technologies developed by NEXT VISION.
is able to localize humans, then it can advice workers can be integrated into new wearable device equipped
informing them in case are on dangerous areas. In case with a camera. Currently, it has been developed on
Miof emergencies (e.g., fires), the system can localize and crosoft Hololens 2 device and on iOS and Android
smartguide workers to the closest fire extinguisher or the clos- phones. For both settings, the user is able to receive
est emergency exit helping to get to safety. When an augmented reality information to navigate the unknown
operator is working in an unfamiliar environment, his environment.
ability to navigate it is initially reduced, indeed, this can
have negative afect to his productivity and decrease his
safety. A system able to localize and suggest paths to the 3. Human-Object Interaction
user can help to navigate the new unfamiliar environ- Understanding
ments.
      </p>
      <p>
        To support workers in these complex scenarios, NEXT Human-object interactions algorithms are fundamental
VISION developed NAIROBI (Navigating Autonomously to understand how humans relate with the surrounding
Indoor Routes by Observing Building Information), an environments while achieving their goals. In
particuArtificial Intelligence assistant able to guide the user to lar, in industrial places, the usage of specific objects and
reach specific areas (e.g., a specific workbench), or the machines need an initial training phase, often with the
closest points of interest (eg., fire extinguisher or emer- cooperation of another worker which has more
experigency exit). NAIROBI is able to localize the user in the ence, to transfer specific skills to the new workers. For
industrial environment through computational vision example, a testing procedure of a machine, may require
allowing the implementation of rescue procedures in the usage of an electric panel composed of several
butdangerous situations for the operator. The intelligent as- tons that must be pressed following a specific sequence.
sistant uses Computer Vision and Artificial Intelligence These complex operations often require people which
algorithms to localize the worker inside the building from have an expertise on the field, indeed a company has
sigimages, compute the best route and to guide the user with nificant costs in terms of time and money for providing
a path shown in the display of the wearable device using specialized training courses. Considering a complex
prothe Augmented Reality. An example of navigation pro- cedure, even an expert operator could consult a manual
vided by NAIROBI in an indoor environment is shown in during the execution of the procedure, which can slow
Figure 1. NAIROBI is also able to provide a map (i.e., “You down the task or it can lead to making mistakes.
are here” map shown through holograms on the display To support workers during the execution of complex
of the wearable device) of the indoor environment show- procedures, NEXT VISION proposes NAOMI (Next Active
ing the real-time user’s position to help users navigate a Object for Monitoring Interactions) [
        <xref ref-type="bibr" rid="ref8 ref9">8, 9</xref>
        ], an Artificial
building (Figure 2). Intelligence assistant able to understand and monitor
in
      </p>
      <p>NAIROBI is multiplatform and not device specific. It teractions between humans and the objects present in the
surrounding environment. NAOMI is able to recognize for example to check the wear of a specific industrial
the objects and the interactions with them and to guide tool in order to carry out preventive maintenance, or if
the user during the execution of complex procedures sug- a specific machine has been used correctly and safely.
gesting the diferent steps to perform and understanding Furthermore, the obtained information can be used to
if each step has been performed correctly. Moreover, it check the quality of procedures.
is able to notify alarms if an interaction is dangerous Object recognition and interaction monitoring are
ad(e.g., electrick risk) avoiding mistakes which could cause dressed thanks to the artificial intelligence algorithms
a machine breakdown or sending information to IoT sys- developed by Next Vision. The communication with the
tems present in the environments (e.g., to immediately users through the Augmented Reality is done thanks to
turn of the electric power for safety). All the informa- the Microsoft Hololens 2 capabilities (Figure 3).
tion about the human-object interactions, predicted by
the artificial intelligence and computer vision algorithms
composing NAOMI, can be aggregated and processed to 4. Human-Machine Interaction
provide statistics about the interacted objects (e.g., usage
time, maintenance date, electric consumption,
temperature). These information are useful to understand how
the worker interacts with the surrounding environment,
Artificial Intelligence allows to build systems which
enable automatic conversation between humans and
machines using natural language. These conversational
agents could be integrated in more complex systems
with the aim to support humans where they live and
work. In particular, user interactions with the
surrounding environment and the objects present in an industrial
workplace require knowledge of specific information and
procedures obtained through the training of operators
with specific courses or cooperation with an expert figure.</p>
      <p>Considering a complex environment, specific
information such as “what is the testing procedure of the electric
board”, “what is the next step of this maintenance
procedure”, “how to calibrate the oscilloscope”, or “how to use
this specific item”, may not be immediately available and
require continuous interactions with laboratory
technicians or qualified workers. These information are also
more complex to obtain when a visual representation
of the surrounding environment is needed to correctly
answer to the worker’s question avoiding language
ambiguity or diversity (for example “what is the object in
front of me?”).</p>
      <p>
        NEXT VISION developed an artificial intelligence
technologies to support operators which need to obtain
information during procedures. HERO (Human Expertise
Replication from Observation) [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ] is an artificial
assistant that through the interaction with a conversational
agent equipped with artificial vision algorithms, is able
to receive questions by workers using natural language,
images and videos acquired from the user’s point of view
through wearable devices. HERO is able to receive the
human input, which can be represented by text or audio
signals, understand the questions processing the input
with the artificial intelligence algorithms as well as the
surrounding environments analyzing the images/videos
acquired from a camera (e.g., head-mounted camera or
smartphone camera) to disambiguate the questions asked
by the user. HERO provides the most reasonable
answers to the questions performed by users. An example
is shown in Figure 4 where the user asked a question
to HERO using the chat. The text has been generated
using a speech-to-text module. Note that to disambiguate
the question “How to use this objects?”, HERO asked to
the user to send an image of the surrounding
environment, understanding that the user is referring to the iron
soldering. Then, HERO provided the answer with the
information on how to use that specific object.
      </p>
      <p>The conversational assistant is able to cooperate to the
other intelligent assistants developed by NEXT VISION
and discussed in previous paragraphs. For example, the
user can ask “guide me to the closest fire extinguisher” or
“start the wizard for the object I’m looking at” to receive
support from NAIROBI and NAOMI.</p>
    </sec>
    <sec id="sec-3">
      <title>5. Conclusions</title>
      <sec id="sec-3-1">
        <title>This manuscript presented technologies developed by NEXT VISION - Spin-of of the University of Catania.</title>
        <p>The proposed technologies make use of Artificial
Intelligence algorithms and wearable devices equipped with a
camera to improve workers productivity and to increase
their safety in industrial workplaces. In particular, three
artificial intelligent assistants have been presented to
support workers in diferent manners. NAIROBI is able
to localize the user and provide routes to reach a specific
point of interest, NAOMI understands human-object
interactions providing instructions on how to perform a
procedure of maintenance, and HERO, a conversational
assistant able to interact with workers using the natural
language and artificial vision.</p>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          <source>[1] Microsoft hololens 2</source>
          , https://www.microsoft.com/ en-us/hololens, Last accessed on 2023-
          <volume>04</volume>
          -01.
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <surname>Nreal</surname>
            <given-names>light</given-names>
          </string-name>
          , https://www.nreal.ai/light/,
          <source>Last accessed on 2023-04-01.</source>
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <surname>Vuzix</surname>
            <given-names>blade</given-names>
          </string-name>
          , https://www.vuzix.com/products/ vuzix-blade-2
          <string-name>
            <surname>-</surname>
          </string-name>
          smart-glasses,
          <source>Last accessed on 2023- 04-01.</source>
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <surname>Realwear</surname>
          </string-name>
          , https://www.realwear.com/,
          <source>Last accessed on 2023-04-01.</source>
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <surname>Vuzix</surname>
            <given-names>m400</given-names>
          </string-name>
          , https://www.vuzix.com/products/ m400-smart-glasses?variant=41517448659110,
          <string-name>
            <surname>Last</surname>
            <given-names>accessed</given-names>
          </string-name>
          <source>on 2023-04-01.</source>
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>G. M.</given-names>
            <surname>Farinella</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A. F. F.</given-names>
            <surname>Ragusa</surname>
          </string-name>
          , E. Ragusa,
          <string-name>
            <given-names>G.</given-names>
            <surname>Sorbello</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Lopes</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Santo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Samarotto</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Scarso</surname>
          </string-name>
          , E. Scarso,
          <article-title>Metodo di assistenza virtuale relativo dispositivo e sistema</article-title>
          ,
          <year>2022</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>G. M.</given-names>
            <surname>Farinella</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Signorello</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Furnari</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Battiato</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Scuderi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Lopes</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Santo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Samarotto</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Distefano</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D. G.</given-names>
            <surname>Marano</surname>
          </string-name>
          ,
          <article-title>Metodo integrato con kit indossabile per analisi comportamentale e visione aumentata</article-title>
          ,
          <year>2020</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>R.</given-names>
            <surname>Leonardi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Ragusa</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Furnari</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G. M.</given-names>
            <surname>Farinella</surname>
          </string-name>
          ,
          <article-title>Egocentric human-object interaction detection exploiting synthetic data</article-title>
          ,
          <source>in: International Conference on Image Analysis and Processing (ICIAP)</source>
          ,
          <year>2022</year>
          . URL: https://iplab.dmi.unict.it/EHOI_ SYNTH/.
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>M.</given-names>
            <surname>Mazzamuto</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Ragusa</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Resta</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G. M.</given-names>
            <surname>Farinella</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Furnari</surname>
          </string-name>
          ,
          <article-title>A wearable device application for human-object interactions detection</article-title>
          .,
          <source>in: International Conference on Computer Vision Theory and Applications (VISAPP)</source>
          ,
          <year>2023</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>C.</given-names>
            <surname>Bonanno</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Ragusa</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Leonardi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Furnari</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G. M.</given-names>
            <surname>Farinella</surname>
          </string-name>
          ,
          <string-name>
            <surname>HERO:</surname>
          </string-name>
          <article-title>An artificial conversational assistant to support humans in industrial scenarios</article-title>
          ,
          <source>in: International Conference on Signal Processing and Multimedia Applications (SIGMAP)</source>
          ,
          <year>2022</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>