<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Auditory indication system for object-finding in remote collaborative assistance</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Takayuki KOMODA</string-name>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Fumina UTSUMI</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Takayoshi YAMADA</string-name>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Keiichi ZEMPO</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>College of Engineering Systems, University of Tsukuba</institution>
          ,
          <addr-line>1-1-1 Tennodai, 3058573</addr-line>
          ,
          <country country="JP">Japan</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Faculty of Engineering</institution>
          ,
          <addr-line>Information and Systems</addr-line>
          ,
          <institution>University of Tsukuba</institution>
          ,
          <addr-line>1-1-1 Tennodai, 3058573</addr-line>
          ,
          <country country="JP">Japan</country>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>Graduate School of Science and Technology, University of Tsukuba</institution>
          ,
          <addr-line>1-1-1 Tennodai, 3058573</addr-line>
          ,
          <country country="JP">Japan</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>In this paper, we propose a remote collaboration system to assist the person with visually impaired in object-finding. The system consists of a 360-degree image centering on a person with visually impaired and presented as a panoramic image on the PC screen of a supporter in a remote location. By clicking on the PC screen, the supporter can present the AR audio (auditory indicator) superimposed on the real space that the person with visually impaired perceives, using spatial sounds. Auditory indicator enables the person with visually impaired to understand the location of the object intuitively. We conducted experiments to clarify the efect of the proposed system on the performance time of the object-finding task and the phrases of the supporters. The results of the experiment showed that the auditory indicator enabled the supporter to guide the simulated person with visually impaired by using demonstrative pronouns such as “this” and “here”.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;Augmented reality audio</kwd>
        <kwd>Human Augmentations</kwd>
        <kwd>Assistive technology</kwd>
        <kwd>Remote collaboration</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>
        in language, sighted people describe the route based on
their location. In contrast, PVI describes the route based
According to 2012 World Health Organization (WHO) on the locations of landmarks. Furthermore, for route
report, there are approximately 285 million person with descriptions between specific points, audio navigation
visually impaired (PVI) worldwide1. PVI sufer many in- created by PVI is subjectively more satisfactory for PVI
conveniences in their daily lives due to their inability to than that created by sighted people [8]. The reason was
recognize visual information. Various assistive technolo- that PVI felt more secure, oriented, and clear when the
gies have been studied, such as navigation aids [
        <xref ref-type="bibr" rid="ref1 ref2">1, 2</xref>
        ] and direction of travel was explained using landmarks. In
object-finding aids [
        <xref ref-type="bibr" rid="ref3 ref4 ref5">3, 4, 5</xref>
        ]. Assistive technologies have addition, in conversation among sighted people, they can
made progress in assisting PVI. On the other hand, there point to an arbitrary area using demonstrative pronouns
are still many problems before assistive technology can such as "there" and communicate without redundant
exbe widely used in daily life, such as system error rates pressions [9]. On the other hand, the PVI are less likely
and communication speed. Remote sighted assistance to use demonstrative pronouns in conversation because
(RSA) [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ] has received lots of attention in addressing they cannot recognize visual information. They tend to
these issues. RSA combines assistive technologies with communicate more verbosely compared to the sighted [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ].
remote assistance from sighted people, and systems such Redundant communication has been suggested to be a
as VizWiz [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ] and Be My Eyes2 have been developed. stress factor in remote collaboration [10].
      </p>
      <p>
        However, these remote collaboration systems do not According to Kraut et al. [11], in a remote
collaboraconsider the diferences in spatial perception between tion among sighted people, sharing the local user’s visual
PVI and sighted people. According to Tsai et al. [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ], PVI space with the supporter in a remote location caused
and sighted people perceive space diferently. For ex- the supporter to utter demonstrative pronouns. Gupta et
ample, when describing a route between specific points al. [12] also showed that sharing the local user’s visual
space and the remote supporter’s gaze shortens the local
APMAR’22: The 14th Asia-PacificWorkshop on Mixed and Augmented user’s task performance time in a remote collaboration
Reality, Dec. 02-03, 2022, Yokohama, Japan among sighted people. The supporter used more
demon* Corresponding author. strative pronouns than when only the visual space was
$ z0e0m00p-o00@0i2i-t4.t0s1u2k-u4b4a1.7ac(.Tjp.Y(KA.MZAEDMAP)O;0)000-0003-2339-5298 shared.
(K. ZEMPO) In this paper, we propose an auditory indication
sys© 2022 Copyright for this paper by its authors. Use permitted under Creative Commons License tem for remote collaboration with PVI in object-finding,
CPWrEooUrckReshdoinpgs IhStpN:/c1e6u1r3-w-0s.o7r3g ACttEribUutRion W4.0oInrtekrnsahtioonpal (PCCroBYce4.0e).dings (CEUR-WS.org) which utilizes AR audio superimposed on the real space
1https://www.emro.who.int/control-and-preventions-of-blindness- that PVI perceives. We conducted experiments to clarify
iamndp-adiremafennets.hs/tamnlnouncements/global-estimates-on-visual- the proposed system’s efect on the performance time
2https://www.bemyeyes.com/ of the object-finding task and the phrases of supporters.
The contributions of this paper are as follows
• Proposal of an interface using AR audio.
• AR audio technology for remote collaboration.
• Suggestion of a new interaction between PVI and
      </p>
      <p>sighted people using AR audio.
2. Related work
remote location. A 360-degree camera is used to present
a panoramic image of the environment around PVI on the
screen of the supporter’s PC. When the supporter clicks
any point in the panoramic image, auditory indicators
are presented to PVI from the direction corresponding
to the panoramic image. By presenting auditory
indicators with spatial sounds, PVI can intuitively perform
object-finding.</p>
      <p>
        This chapter describes studies on object-finding assis- 3.2. Interface
tance for PVI. Kaul et al. [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ] developed a mobile
application that combines an object detection framework In this section, we describe the configuration of the
inand spatial sound to assist PVI in navigation and object- terface. PVI wears a 360-degree camera (RICOH THETA
ifnding. The application recognizes objects around a PVI Z1) on the top of the head and presents a panoramic
imusing a smartphone’s camera function and provides audi- age of the surrounding environment to the screen of the
tory feedback by spatial sound to explain the scene. The supporter’s PC. PVI wears open-ear headphones (Sony
user evaluation of the auditory feedback was favorable. LinkBuds) to listen to auditory indicators without
interHowever, some issues related to the system, such as low fering with the environmental sounds. The supporter
object detection accuracy and narrow detection range, ifnds the object (target) from the presented images of the
were mentioned. surroundings of PVI and clicks on the screen. The sound
      </p>
      <p>
        In order to improve the reliability and usefulness of source is placed on the computer’s three-dimensional
assistive technology, Bigham et al. developed VizWiz [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ], space corresponding to the target’s position in real space
which combines the assistance of sighted people. PVI by clicking on the screen. Based on the positional
relauses a smartphone’s camera function to take pictures, tionship between PVI and the sound source in the
comask questions to online supporters, and receive voice puter’s three-dimensional space, spatial sounds are
genresponses. The system does not make object detection erated and presented to PVI.
errors, but PVI requires some recognition of the object’s
location to be found. Therefore, if the object’s location is 3.3. Auditory indicator
unknown, it is not easy to use the system. The proposed system presents auditory indicators as
spa
      </p>
      <p>
        Be My Eyes2 is a mobile application that allows PVI to tial sounds to PVI.
ask for assistance from a remote location using a video Spatial sound means that meta-information such as
chat on their smartphones. PVI does not need to be distance, direction, and spatial extent is represented in
aware of the object’s location in advance to receive as- the sound reproduction. The perceived information, such
sistance. However, the field of view that can be shared as distance and direction, is called sound image
localizawith the user is limited due to smartphone camera use. tion. Sound image localization can be applied to a sound
Jones et al. [13] show that in remote collaboration us- source by giving volume, time, and frequency response
ing smartphones, users who receive video sharing from diferences to the left and right voices. In this paper, to
their smartphones intentionally use information from present spatial sound, sound image localization is
conthe camera images when asking questions. However, the volved with a sound source using Unity and Steam Audio.
narrow field of view and the inability to control the
direction of the camera was found to be stress factors. In
Be My Eyes, since the person being assisted is PVI, the 4. Experiment
supporter cannot easily convey the information obtained
from the visual images, which may cause redundant com- In order to clarify the efect of using the proposed system
munication. Wang et al. [10] suggest that redundant on the performance time of the object-finding task and
communication is a stress factor in remote collaboration. the phrases of the supporter, we conducted an
objectifnding task experiment with a simulated person with
visually impaired (SPVI) and a participant playing the
3. System design role of a supporter in a remote location, based on the
experiment by Kual et al. [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]
3.1. System overview Eight Japanese university students were randomly
paired, with one participant as SPVI and the other as the
supporter. SPVI wore an eye mask. Of the four groups of
eight participants in the experiment, two groups of four
      </p>
      <sec id="sec-1-1">
        <title>The system proposed is shown in Figure 1. The proposed system consists of PVI, who receives support, and a supporter (sighted people) who provides support from a</title>
        <p>used the proposed system, which enables the sharing of
360-degree images, auditory indicators, and voice
instructions (Condition proposed.) The remaining two groups
of four shared the 360-degree images and performed the
task by voice instructions without auditory indicators
(Condition control.).</p>
        <p>In the laboratory, desks are arranged in four directions
around the SPVI, and seven objects (targets) are placed on
the desks. Figure 1 shows an overview of the laboratory.</p>
        <p>In the experiment, SPVI is instructed by the experiment
supervisor on what to find (targets). Although there are
seven possible targets, SPVI is only informed of them
once the experiment supervisor instructs SPVI to look
for them. SPVI cooperates with the supporter through the
system to find the target and carry it to the designated
position. Three trials were conducted per pair. Since
the eye mask blocked SPVI’s vision, the experiment was
conducted with SPVI sitting on a swivel chair with wheels
to avoid the risk of falling.</p>
      </sec>
    </sec>
    <sec id="sec-2">
      <title>5. Results</title>
      <sec id="sec-2-1">
        <title>5.1. Performance time</title>
        <sec id="sec-2-1-1">
          <title>The task performance time was compared between the conditions in which spatial sounds were presented and those in which they were not. The results are shown in Table 1.</title>
          <p>Participants in Pairs #1 and #2 were presented with
auditory indicators, while those in Pairs #3 and #4 were
not presented with auditory indicators.</p>
        </sec>
        <sec id="sec-2-1-2">
          <title>The remote collaboration could not be performed cor</title>
          <p>rectly in the third trial of pair #1 due to the problem with
the output of the panoramic image. Therefore, the task
performance time was excluded as an outlier. In addition,
the task of the third trial of pair #3 was also excluded
from matching the number of tasks.</p>
          <p>As a result of the experiment, the task performance
time tends to be shortened when auditory indicators are
presented.</p>
        </sec>
      </sec>
      <sec id="sec-2-2">
        <title>5.2. Efect on speech</title>
        <sec id="sec-2-2-1">
          <title>As a result of the experiment, it was confirmed that when the auditory indicator was not presented, the supporters used directional phrases such as "right/left/straight" to support SPVI.</title>
          <p>When auditory indicators were presented, phrases
such as "right/left/straight" were also confirmed. On the
other hand, we could confirm the use of demonstrative
pronouns such as "this" and "here". In addition, it was
confirmed that SPVI responded correctly to the
direction of the target when demonstrative pronouns were
used. Table 2 shows the number of times demonstrative
pronouns, the phrases "left," "right," and "straight" were
used.</p>
        </sec>
      </sec>
      <sec id="sec-2-3">
        <title>5.3. Discussion</title>
        <sec id="sec-2-3-1">
          <title>The experiment results showed that the use of auditory indicators shortened the performance time of the objectifnding task and enabled the use of demonstrative pronouns by the supporters.</title>
          <p>Pro.</p>
          <p>Left
Right
Straight
Table 1 were presented. In addition, it was also confirmed that the
Performance time. #1, #2 are using auditory indicator, #3, #4 auditory indicator made the supporter use more
demonare not using auditory indicator. strative pronouns such as "this" and "here". We concluded</p>
          <p>Performance time[s] that these results are because the presentation of the
tarPair 1st 2nd 3rd Ave get direction by auditory indicators functioned in the
#1 56.3 41.0 - 48.7 same way as gaze sharing in the remote collaboration
#2 72.6 57.5 40.7 56.9 between sighted people. The results of this paper
sug#3 52.0 74.3 - 63.2 gest that demonstrative pronouns can be used in remote
#4 99.6 55.3 52.8 69.2 collaboration with PVI, indicating a new interaction
between PVI and sighted people using AR audio. However,
Table 2 one limitation of this paper is that the participants of
Number of times demonstrative pronouns, the phrases “left”, the experiment were SPVI. Since SPVI is a blindfolded
“right” and “straight” is used. Condition proposed: Using audi- sighted person, it is not representative of PVI, and future
tory indicator, condition control: Not using auditory indicator. studies should be conducted with PVI as a participant.</p>
          <p>In the future, it is necessary to investigate whether the
presentation of auditory indicators has the same role as
gaze sharing and to clarify the mechanism by which the
demonstrative pronouns were used.</p>
          <p>We conclude that the reason for the shorter task
performance time is that the auditory indicator enables an
intuitive presentation of the target direction. In the remote
collaboration between sighted people, task eficiency was
improved by sharing the gaze of the supporter by
pointing [12]. We consider that a similar mechanism is
responsible for shorting task performance time. We also
note that demonstrative pronouns were used. Previous
work [12, 11] confirmed that sharing the visual space
enables supporters to use demonstrative pronouns,
improving task eficiency. The results of this paper are
consistent with these results.</p>
          <p>The use of demonstrative pronouns is considered to be
because the auditory indicators functioned in the same
way as pointing for the sighted people, and a
pseudovisual space was shared between SPVI and the supporters.</p>
          <p>This is concluded from the fact that in the study of remote
collaboration among sighted people by Kraut et al. [11],
the supporter started to use demonstrative pronouns after
the visual space was shared.</p>
        </sec>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>6. Conclusion and future works</title>
      <p>In this paper, we propose a system that uses auditory
indicators to provide directions to assist PVI in
objectifnding through remote collaboration. In order to clarify
the efectiveness of the proposed system, we conducted
an object-finding task experiment. We investigated the
efect of the auditory indicator on the task performance
time and the phrases of supporters. As a result of the
experiment, it was confirmed that the task performance
time tended to be shortened when auditory indicators
spatial voice navigation for people with visual
impairment and without visual impairment, in:
Proceedings of the 2019 7th International Conference
on Information and Education Technology, 2019,
pp. 295–300.
[8] Y.-H. Hung, K.-Y. Tsai, E. Chang, R. Chen, Voice
navigation created by vip improves spatial
performance in people with impaired vision, International
journal of environmental research and public health
19 (2022) 4138.
[9] Y. Hato, S. Satake, T. Kanda, M. Imai, N. Hagita,
Pointing to space: modeling of deictic interaction
referring to regions, in: 2010 5th ACM/IEEE
International Conference on Human-Robot Interaction
(HRI), IEEE, 2010, pp. 301–308.
[10] B. Wang, Y. Liu, J. Qian, S. K. Parker, Achieving
efective remote working during the covid-19
pandemic: A work design perspective, Applied
psychology 70 (2021) 16–59.
[11] R. E. Kraut, D. Gergle, S. R. Fussell, The use of
visual information in shared visual spaces: Informing
the development of virtual co-presence, in:
Proceedings of the 2002 ACM conference on Computer
supported cooperative work, 2002, pp. 31–40.
[12] K. Gupta, G. A. Lee, M. Billinghurst, Do you see
what i see? the efect of gaze tracking on task space
remote collaboration, IEEE transactions on
visualization and computer graphics 22 (2016) 2413–2422.
[13] B. Jones, A. Witcraft, S. Bateman, C. Neustaedter,
A. Tang, Mechanics of camera work in mobile video
collaboration, in: Proceedings of the 33rd Annual
ACM Conference on Human Factors in Computing
Systems, 2015, pp. 957–966.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>M. H. A.</given-names>
            <surname>Wahab</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A. A.</given-names>
            <surname>Talib</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H. A.</given-names>
            <surname>Kadir</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Johari</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Noraziah</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R. M.</given-names>
            <surname>Sidek</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A. A.</given-names>
            <surname>Mutalib</surname>
          </string-name>
          ,
          <article-title>Smart cane: Assistive cane for visually-impaired people</article-title>
          ,
          <source>arXiv preprint arXiv:1110.5156</source>
          (
          <year>2011</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>A.</given-names>
            <surname>Helal</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S. E.</given-names>
            <surname>Moore</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Ramachandran</surname>
          </string-name>
          ,
          <article-title>Drishti: An integrated navigation system for visually impaired and disabled</article-title>
          ,
          <source>in: Proceedings fifth international symposium on wearable computers, IEEE</source>
          ,
          <year>2001</year>
          , pp.
          <fpage>149</fpage>
          -
          <lpage>156</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>O. B.</given-names>
            <surname>Kaul</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Behrens</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Rohs</surname>
          </string-name>
          ,
          <article-title>Mobile recognition and tracking of objects in the environment through augmented reality and 3d audio cues for people with visual impairments</article-title>
          ,
          <source>in: Extended Abstracts of the 2021 CHI Conference on Human Factors in Computing Systems</source>
          ,
          <year>2021</year>
          , pp.
          <fpage>1</fpage>
          -
          <lpage>7</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>M.</given-names>
            <surname>Eckert</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Blex</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C. M.</given-names>
            <surname>Friedrich</surname>
          </string-name>
          , et al.,
          <article-title>Object detection featuring 3d audio localization for microsoft hololens</article-title>
          ,
          <source>in: Proc. 11th Int. Joint Conf. on Biomedical Engineering Systems and Technologies</source>
          , volume
          <volume>5</volume>
          ,
          <year>2018</year>
          , pp.
          <fpage>555</fpage>
          -
          <lpage>561</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>J. P.</given-names>
            <surname>Bigham</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Jayant</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Ji</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Little</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Miller</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R. C.</given-names>
            <surname>Miller</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Miller</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Tatarowicz</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>White</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>White</surname>
          </string-name>
          , et al.,
          <article-title>Vizwiz: nearly real-time answers to visual questions</article-title>
          ,
          <source>in: Proceedings of the 23nd annual ACM symposium on User interface software and technology</source>
          ,
          <year>2010</year>
          , pp.
          <fpage>333</fpage>
          -
          <lpage>342</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>S.</given-names>
            <surname>Lee</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Reddie</surname>
          </string-name>
          ,
          <string-name>
            <surname>C.-H. Tsai</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          <string-name>
            <surname>Beck</surname>
            ,
            <given-names>M. B.</given-names>
          </string-name>
          <string-name>
            <surname>Rosson</surname>
            ,
            <given-names>J. M.</given-names>
          </string-name>
          <string-name>
            <surname>Carroll</surname>
          </string-name>
          ,
          <article-title>The emerging professional practice of remote sighted assistance for people with visual impairments</article-title>
          ,
          <source>in: Proceedings of the 2020 CHI Conference on Human Factors in Computing Systems</source>
          ,
          <year>2020</year>
          , pp.
          <fpage>1</fpage>
          -
          <lpage>12</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <surname>K.-Y. Tsai</surname>
            ,
            <given-names>Y.-H.</given-names>
          </string-name>
          <string-name>
            <surname>Hung</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          <string-name>
            <surname>Chen</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          <string-name>
            <surname>Chang</surname>
          </string-name>
          , Indoor
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>