<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Using Voice Assistants as HMI for Robots in Smart Production Systems</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Javad Ghofrani</string-name>
          <email>javad.ghofrani@gmail.com</email>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Dirk Reichelt</string-name>
          <email>dirk.reichelt@htw-dresden.de</email>
        </contrib>
      </contrib-group>
      <fpage>62</fpage>
      <lpage>65</lpage>
      <abstract>
        <p>Smart voice assistant systems are widely used in the smart home and entertainment domain. Despite of the fact that touch screens are used in industry for giving commands to the robots and production machines, the use of voice or video assistants is still limited to smart home and technologies like Apple Siri or Amazon' Alexa. In this paper, we discuss the feasibility of applying existing digital assistant systems in industrial contexts and various aspects of their usage S. Kolb, C. Sturm (Eds.): 11th ZEUS Workshop, ZEUS 2019, Bayreuth, Germany, 14-15 February 2019, published at http://ceur-ws.org/Vol-2339</p>
      </abstract>
      <kwd-group>
        <kwd>Robotic</kwd>
        <kwd>industry 4</kwd>
        <kwd>0</kwd>
        <kwd>IIoT</kwd>
        <kwd>digital assistant</kwd>
        <kwd>voice recognition</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Introduction</title>
      <p>
        Using sound activated systems to control various devices such as smart home,
BMW cars and also unmanned aerial vehicles [
        <xref ref-type="bibr" rid="ref1 ref5">1, 5</xref>
        ] is more and more intended.
In comparison to mouse and keyboard, touch screen are more flexible since they
occupy less space while making it possible to use diferent layouts in one place.
Nevertheless, working with touch screens distracts the user from his main task.
Sound based assistant systems compensate this shortage. In these systems, the
natural language interfaces create a layer of abstraction and try to hide big
amount of complexity to work with system interfaces. The aim of this paper is to
propose the application of smart digital assistant for communicating with robots
and to prepare a research project for the use of digital, voice-controlled assistance
systems for communication with robots in smart production systems.
Problem of processing natural languages is the first era in which many works have
been done. Ortiz [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ] investigates and discusses the challenges which the developers
should consider when they developing a new interface. Furthermore, Milhorat
et al. [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ] address the existing challenges in the way of implementing the personal
digital assistant which works with voice. Both of these papers mentions the
understanding the language, handling the knowledge, and reasoning as important
aspects of this field. In the work of Park et al. [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ] a speech recognition system is
used to send the movement directives to a robot in the unsafe area (fire). They
mention the problem of natural language processing, noise of the environment and
network resilience as problems which can occur during their experiment. Solorio
et al. [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ] categorize the facing challenges in the implementation of IoT into
standards, data management, security, privacy, and the lack of flawless technology
in speech recognition. Polyakov et al. [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ] investigates the implementation of a
voice assistant system which is not equipped with cloud based services but locally
exploiting the deep neural networks for analyzing the voice commands. They
mention the lack of compatibility of middle-wares, security sphere of use cases,
as well as unrealistic thinking about ubiquitous voice assistants as the problems
in applying the voice assistant systems without cloud computing. However, cloud
computing is not the only possible solution to solve the speech recognition
problems. Other solutions are developed to overcome the latency and security
issues of solutions based on cloud computing. Fog computing and edge with a lot
of smart, connected endpoints are some of these solutions.
      </p>
      <p>
        Replacing the Human Machine Interface (HMI) with speech commands are
the next challenge which is considered in publications and scientific works. Kawai
et. al [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ] proposes a voice-activated system as a replacement of conventional
inputs (mouse and keyboard) to get and put some objects in the 3D environment.
It works like a calling a menu and selecting an item from it. While Green [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]
investigates the various possible way to talk with the equipment, installed in
smart home. Fernández et al. [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ], similar to Meszaros et al. [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ], investigates the
Speech Command Interaction for giving command to the flying devices. They
implemented a feedback mechanism which acknowledges the performing the
commanded action. In the opinion of Kep¨uska and Gamal [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ] a combination of
diferent input channels such as voice, text, image, gesture and etc. are the way
that the next generation of smart assistant systems are going to work.
3
      </p>
    </sec>
    <sec id="sec-2">
      <title>Smart Digital Assistants as Interface for Controlling Robots</title>
      <p>As mentioned in previous sections, using alternative channels, especially voice,
to communicate with robots is not a new field of work. Various works have
investigated the application of voice recognition techniques and converting them to
the proper commands to communicate with various parts of smart environments,
e.g. smart home or smart cars. Hence, the usage of such channels in industrial
environments is an attention point by researchers and practitioners. Recent
advances in the personal digital assistants and their usage in smart home and
driving cars make them a proper candidates to be applied to the industrial
environments for controlling robots, especially collaborative robots (cobots).
Using smart digital assistant systems, the complexities of voice recognition and
natural language processing are pushed to the side of the companies which provide
these technologies. Furthermore, the developers are capable of implementing
their own skills according to their needs. It bring more flexibility for the case of
applying them in industrial environments and solutions. Nevertheless, applying
voice recognition in the industry has additional challenge compared to smart</p>
      <p>Javad Ghofrani and Dirk Reichelt
home. Since in production environments, especially involving robots, there is a
lot of noise that disturbs the interaction using voice. Furthermore, in industrial
environments, there is a need for distributing and orchestrating a larger number of
assistants to cover a machine shop or production line. Considering these problems
and the problems that we are still not aware of, we are going to analyze feasibility
of applying smart digital assistants to the industrial robots. For this purpose,
we are going to emphasize potential problems that could occur if we apply such
assistant systems in smart production systems. We plan to collect data from
various sources such as literature studies, surveys, conducting experiments and
interviews in first step. We are going to propose the following research questions:
(i) which factors prevent the application of smart assistant systems within the
industrial context? (ii) is there any solution, best practice and work-around to
overcome these factors? (iii) which points should be considered in implementation
of such systems at most? Answering these research questions enables us to
determine the specifications of our system.</p>
      <p>We propose our notion of a smart assistant systems for controlling the cobots
in an smart production system similar to industry 4.0 testbed on HTW Dresden1.
In such scenarios, there is workplace in which a human works together with a
cobot. The cobots are able to perform some tasks more precisely than human,
e.g. gluing. In the scenario without a smart voice assistant system, the human
should put a certain part in a certain place on the desk and then use flexpendant
to select the task–here gluing which is located under tasks menu–and then cobot
starts to perform the task by taking the part from specific area of the desk. In
other scenario with a smart sound assistant, the user can talk to his micro and tell
the robot “YuMi, take the piece from left side of desk and glue it”. The keyword
“YuMi” will activate the smart assistant system. It processes the sentences after
the keyword and sends them to a program that is responsible for controling the
cobot. This program should be developed within our project. This program could
be similiar to skills of Amazon Alexa that will be activated by some phrases such
as “glue”. It generates the proper control signals for performing the task and
send them to the cobot.</p>
      <p>In this project, we will be able to measure the quality of communication in the
industrial environments between human and robot. In addition, we are going to
investigate the diferences in these communication ways beside their advantages
and disadvantages. After implementing our solutions, we are going to test it
under various circumstances such as noisy place or with some accent of the user.
The efect of using various equipment such as Amazon Alexa or IBM Watson
will be considered as well. The results of these tests will be used to recognize
the challenges and to provide a comparison between diferent configurations
and solutions. In this regard, our work will cover the acceptance of such smart
assistant systems in the industrial environments using empirical methods such as
qualitative surveys.</p>
      <p>1 https://www.htw-dresden.de/industrie40</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Fernández</surname>
            ,
            <given-names>R.A.S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sanchez-Lopez</surname>
            ,
            <given-names>J.L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sampedro</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bavle</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Molina</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Campoy</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          :
          <article-title>Natural user interfaces for human-drone multi-modal interaction</article-title>
          .
          <source>In: Unmanned Aircraft Systems (ICUAS)</source>
          , 2016 International Conference on. pp.
          <fpage>1013</fpage>
          -
          <lpage>1022</lpage>
          . IEEE (
          <year>2016</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Green</surname>
          </string-name>
          , A.:
          <article-title>C-roids: Life-like characters for situated natural language user interfaces</article-title>
          .
          <source>In: Robot and Human Interactive Communication</source>
          ,
          <year>2001</year>
          .
          <source>Proceedings. 10th IEEE International Workshop on</source>
          . pp.
          <fpage>140</fpage>
          -
          <lpage>145</lpage>
          . IEEE (
          <year>2001</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Kawai</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Higashiyama</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Koyama</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Okada</surname>
            ,
            <given-names>M.:</given-names>
          </string-name>
          <article-title>A fundamental study on a natural-language-based 3d cg modeling</article-title>
          .
          <source>In: Systems, Man, and Cybernetics</source>
          ,
          <year>1999</year>
          .
          <source>IEEE SMC'99 Conference Proceedings. 1999 IEEE International Conference on. vol. 5</source>
          , pp.
          <fpage>714</fpage>
          -
          <lpage>719</lpage>
          . IEEE (
          <year>1999</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Këpuska</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bohouta</surname>
          </string-name>
          , G.:
          <article-title>Next-generation of virtual personal assistants (microsoft cortana, apple siri, amazon alexa and google home)</article-title>
          .
          <source>In: Computing and Communication Workshop and Conference (CCWC)</source>
          ,
          <source>2018 IEEE 8th Annual</source>
          . pp.
          <fpage>99</fpage>
          -
          <lpage>103</lpage>
          . IEEE (
          <year>2018</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Meszaros</surname>
            ,
            <given-names>E.L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chandarana</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Trujillo</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Allen</surname>
            ,
            <given-names>B.D.</given-names>
          </string-name>
          :
          <article-title>Speech-based natural language interface for uav trajectory generation</article-title>
          .
          <source>In: Unmanned Aircraft Systems (ICUAS)</source>
          , 2017 International Conference on. pp.
          <fpage>46</fpage>
          -
          <lpage>55</lpage>
          . IEEE (
          <year>2017</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Milhorat</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Schlogl</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chollet</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Boudy</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Esposito</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Pelosi</surname>
          </string-name>
          , G.:
          <article-title>Building the next generation of personal digital assistants</article-title>
          .
          <source>In: Advanced Technologies for Signal and Image Processing (ATSIP)</source>
          ,
          <year>2014</year>
          1st International Conference on. pp.
          <fpage>458</fpage>
          -
          <lpage>463</lpage>
          . IEEE (
          <year>2014</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>Ortiz</surname>
            ,
            <given-names>C.L.</given-names>
          </string-name>
          :
          <article-title>The road to natural conversational speech interfaces</article-title>
          .
          <source>IEEE Internet Computing (2)</source>
          ,
          <fpage>74</fpage>
          -
          <lpage>78</lpage>
          (
          <year>2014</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>Park</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kim</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Matson</surname>
            ,
            <given-names>E.T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Jang</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lee</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Park</surname>
            ,
            <given-names>W.:</given-names>
          </string-name>
          <article-title>An intuitive interaction system for fire safety using a speech recognition technology</article-title>
          .
          <source>In: Automation, Robotics and Applications (ICARA)</source>
          ,
          <year>2015</year>
          6th International Conference on. pp.
          <fpage>388</fpage>
          -
          <lpage>392</lpage>
          . IEEE (
          <year>2015</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <surname>Polyakov</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Mazhanov</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rolich</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Voskov</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kachalova</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Polyakov</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          :
          <article-title>Investigation and development of the intelligent voice assistant for the internet of things using machine learning</article-title>
          .
          <source>In: Electronic and Networking Technologies (MWENT)</source>
          , 2018 Moscow Workshop on. pp.
          <fpage>1</fpage>
          -
          <lpage>5</lpage>
          . IEEE (
          <year>2018</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <surname>Solorio</surname>
            ,
            <given-names>J.A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Garcia-Bravo</surname>
            ,
            <given-names>J.M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Newell</surname>
            ,
            <given-names>B.A.</given-names>
          </string-name>
          :
          <article-title>Voice activated semi-autonomous vehicle using of the shelf home automation hardware</article-title>
          .
          <source>IEEE Internet of Things Journal</source>
          (
          <year>2018</year>
          )
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>