<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Collision prediction with oncoming pedestrians on Braille blocks</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Tomoya Ishii</string-name>
          <email>ishii.tomoya@image.iit.tsukuba.ac.jp</email>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Hidehiko Shishido</string-name>
          <email>shishido.hidehiko@image.iit.tsukuba.ac.jp</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Yoshinari Kameda</string-name>
          <email>kameda@ccs.tsukuba.ac.jp</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Center for Computational Sciences, University of Tsukuba</institution>
          ,
          <addr-line>1-1-1 Tennoudai, Tsukuba, Ibaraki, 305-8573</addr-line>
          ,
          <country country="JP">Japan</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Currently at Soka University</institution>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>Master's program in Intelligent and Mechanical Interaction Systems, University of Tsukuba</institution>
          ,
          <addr-line>1-1-1 Tennoudai, Tsukuba, Ibaraki, 305-8573</addr-line>
          ,
          <country country="JP">Japan</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>This research targets blind people walking on braille blocks. Some non-blind pedestrians walk on braille blocks while they watch their smartphones on the street. Blind people may collide with these oncoming pedestrians. When the oncoming pedestrian does not notice the blind people, a collision will likely occur. We propose a new method for predicting collisions in such situations. We use a smartphone's camera to predict collisions. To predict the collision, we set two conditions. The first condition is whether the oncoming pedestrian is on the collision path. The second condition is whether the oncoming pedestrian notices the blind people.</p>
      </abstract>
      <kwd-group>
        <kwd>1 Collision prediction</kwd>
        <kwd>distance estimation</kwd>
        <kwd>path estimation</kwd>
        <kwd>gaze estimation</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>
        This research targets blind people walking on
braille blocks. Blind people rely on braille blocks
when they go out. Some non-blind pedestrians
walk on braille blocks while they watch their
smartphones on the street. It is difficult for blind
people to avoid these pedestrians and may collide
with them. In a potential collision situation,
oncoming non-blind pedestrians should give way.
Braille blocks are installed to help blind people
walk safely in Japan [
        <xref ref-type="bibr" rid="ref1 ref2">1, 2</xref>
        ]. When oncoming
pedestrians notice blind people, they are asked to
give their way to avoid collisions.
      </p>
      <p>We propose a new method for predicting
collisions with oncoming pedestrians using a
smartphone’s camera. To predict collision, we set
two conditions. The first condition is whether the
oncoming pedestrian is on the collision path. The
second condition is whether the oncoming
pedestrian notices the blind people. We foresee
that a collision will occur when the oncoming
pedestrian is on the collision path and does not
notice the blind people.</p>
      <p>Once the collisions can be predicted, a loud
sound signal can warn both people. The warning
can make spare time for blind people to protect
themselves in case of a collision. The warning can
also ask the oncoming pedestrian to avoid a
collision. As the loud sound signal causes a big
stress on all the people on the street, the collision
prediction should be accurate, and the collision
prediction system calls the warning at the last
moment when the collision is inevitable.</p>
    </sec>
    <sec id="sec-2">
      <title>2. Related work</title>
    </sec>
    <sec id="sec-3">
      <title>2.1. Collision avoidance for blind people</title>
      <p>
        Various types of obstacle avoidance systems
for blind people have been proposed. These
include the systems to detect obstacles by
attaching an ultrasonic sensor [
        <xref ref-type="bibr" rid="ref3 ref4">3, 4</xref>
        ] or LiDAR [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ]
to a white cane, which blind people use daily. The
ultrasonic sensor can detect obstacles up to 4
meters away, while LiDAR can measure up to 10
meters away.
      </p>
      <p>
        Collision avoidance systems using a
suitcaseshaped device have been proposed [
        <xref ref-type="bibr" rid="ref6">6-8</xref>
        ]. The
suitcase has a stereo camera, LiDAR [
        <xref ref-type="bibr" rid="ref6">6, 8</xref>
        ], and a
laptop computer. The BBeep system [7] predicts
collision by estimating the future location of
oncoming pedestrians based on their walking
trajectories. The system beeps and asks
pedestrians to give their way to avoid a collision.
Vibrators [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ] and levers [8] are equipped on the
suitcase handles to indicate the path for blind
people. These tactile feedbacks enable them to
avoid a collision on their own. These are
promising approaches in case the equipment can
accompany blind people.
      </p>
      <p>Collision avoidance systems using
smartphones [9, 10] and wearable devices [11]
have been proposed. These systems use LiDAR
on smartphones and stereo cameras equipped to
the chest of blind people to measure the distance
to the object. These systems focus on obstacles
and pedestrians at a short distance and do not
consider oncoming pedestrians walking from a
distance.</p>
      <p>
        Many studies have been proposed on
appearance-based gaze estimation using deep
learning [
        <xref ref-type="bibr" rid="ref10 ref11 ref12 ref7 ref8 ref9">21-28</xref>
        ]. Appearance-based gaze
estimation requires datasets containing a variety
of environments, subjects, and targets [
        <xref ref-type="bibr" rid="ref7 ref9">21, 23, 25,
27 28</xref>
        ]. Zhang et al. provided a gaze image dataset
of participants acquired while they watched a
laptop computer in their daily life and proposed a
gaze estimation method [21]. Sugano et al.
estimated gaze on a public display without
individual user calibration [22]. Recasens et al.
estimated which objects the user looked at in the
image [
        <xref ref-type="bibr" rid="ref7">23</xref>
        ]. Chong et al. extended it to video and
correctly detected the gaze target even outside the
image [
        <xref ref-type="bibr" rid="ref8">24</xref>
        ].
      </p>
      <p>
        Kellnhofer et al. provided a dataset captured at
a wide range of head postures and distances to
achieve 3D gaze estimation [
        <xref ref-type="bibr" rid="ref9">25</xref>
        ], where the eyes
may not be visible by cameras such as
surveillance cameras due to occlusion. 3D gaze
estimation [
        <xref ref-type="bibr" rid="ref10">26</xref>
        ] and target object estimation [
        <xref ref-type="bibr" rid="ref11">27</xref>
        ]
have been proposed in situations where only the
back of the head is visible. Bermejo et al.
approximated the gaze by the posture of the head
[
        <xref ref-type="bibr" rid="ref10">26</xref>
        ]. Nonaka et al. estimated the gaze by
considering the body orientation [
        <xref ref-type="bibr" rid="ref12">28</xref>
        ].
      </p>
    </sec>
    <sec id="sec-4">
      <title>2.2. Collision avoidance autonomous mobile robots</title>
      <sec id="sec-4-1">
        <title>As for obstacle collision avoidance, systems</title>
        <p>on autonomous mobile robots have been proposed.
Ultrasonic sensors [12, 13] and LiDAR [14, 15]
detect obstacles. The systems are large because
they use specific sensors to measure the distance
of distant obstacles. It is challenging to adopt
these methods directly for humans.</p>
        <p>Collision avoidance with dynamic obstacles
[16] and moving pedestrians [17] have been
proposed too. The robot is set to go away from the
original planned path or make a curve to avoid
collisions. We should not instruct blind people to
change their paths because they may lose their
way once they are off the braille blocks.
2.3.</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>Gaze estimation</title>
      <p>The gaze plays an important role in human
interaction. People select collision-free paths by
considering the gaze of other pedestrians and the
direction they walk [18-20].
in</p>
    </sec>
    <sec id="sec-6">
      <title>3. Collision</title>
      <p>blocks
avoidance
on
braille</p>
      <sec id="sec-6-1">
        <title>Braille blocks are installed in a straight line [2].</title>
        <p>Once the blind people become on braille blocks,
they follow the blocks. When oncoming
pedestrians notice blind people, they should give
their way to avoid collisions because the block is
installed to support the safe walking of blind
people. As shown in Figure 1, a collision occurs
when an oncoming pedestrian walks on braille
blocks without noticing the blind people.</p>
        <p>In our proposal, blind people wear their
smartphones at chest height, as shown in Figure 2.
The smartphone's camera is tilted downward from
the horizontal. This tilt allows the camera to
capture the walking area in front of the user. The
setup of the smartphone in this way does not
interfere with their walking style.</p>
        <p>The proposed system uses only a smartphone
to cover the process from video acquisition to
collision prediction.</p>
        <p>Collision prediction system
implemented on a smartphone</p>
      </sec>
    </sec>
    <sec id="sec-7">
      <title>4. Oncoming pedestrians</title>
      <p>collision path</p>
    </sec>
    <sec id="sec-8">
      <title>4.1. Detection on the</title>
      <sec id="sec-8-1">
        <title>Pedestrian detection methods from various</title>
        <p>
          camera viewpoints have been proposed [
          <xref ref-type="bibr" rid="ref13 ref14 ref15 ref16 ref17">29-33</xref>
          ].
Pedestrians are detected for tracking [
          <xref ref-type="bibr" rid="ref13 ref14">29, 30</xref>
          ],
counting [
          <xref ref-type="bibr" rid="ref15">31</xref>
          ], and autonomous driving systems
[
          <xref ref-type="bibr" rid="ref16">32</xref>
          ]. Static obstacle detection methods have been
proposed from a pedestrian’s viewpoint [
          <xref ref-type="bibr" rid="ref18 ref19">34, 35</xref>
          ].
We can apply such methods to static objects. In
this research, we focus on the detection of
oncoming pedestrians.
        </p>
        <p>A detector with high-speed detection is
required to achieve real-time detection. We need
a new detector that finds pedestrians on braille
blocks. Note that the braille blocks are square.</p>
        <p>
          The system should first find braille blocks. We
utilize YOLOv7 [
          <xref ref-type="bibr" rid="ref20">36</xref>
          ] as it can run fast. We use the
braille block dataset [
          <xref ref-type="bibr" rid="ref21">37</xref>
          ] to train YOLOv7. We
split 2000 images 4:1 for training and validation.
The batch size is 16. The number of epochs is 150.
The image size was resized from 1024 × 1024 to
512 × 512 for training.
        </p>
        <p>Once the trained YOLOv7 find more than two
braille blocks, the braille block region can be
found. The region is defined by two sidelines of
the braille blocks as the blocks form a straight line
on the street. The left and right vertices of the
bottom edge of the detected blocks are used to
estimate the sidelines of the braille block region.
The sidelines are estimated by the least-squares
method. The region is the collision path.</p>
        <p>We also use YOLOv7 to detect pedestrians.
The center of the bottom edge of the detected
pedestrian rectangle is counted as the pedestrian's
foot position. If the foot position is inside the
braille block region, we detect the pedestrian as an
oncoming pedestrian on the collision path.</p>
        <p>In case of finding less than two braille blocks,
the braille block region found in the last frame is
used to detect the oncoming pedestrian on the
collision path.</p>
        <p>Figure 3 shows the detection results of
oncoming pedestrians. The light blue bounding
boxes indicate the braille blocks detected by the
trained YOLOv7. Note that the bottom edge of the
braille block rectangle is marked by blue color.
The red crossing lines indicate the sidelines of the
braille block region, inside of which is the
collision path.</p>
        <p>Figure 3 (a) shows a situation where an
oncoming pedestrian is on the collision path. The
person wearing black is on the collision path and
is detected. It is marked with a red bounding box.
As the oncoming pedestrian approaches, we can
detect the smartphone they hold, as indicated by
the green bounding box by the YOLOv7.</p>
        <p>Figure 3 (b) shows a situation where the
oncoming pedestrian is not on the collision path.
The person wearing white is not on the collision
path and is marked by the blue bounding box.</p>
        <p>Oncoming pedestrians appear larger in the
image as they approach. Oncoming pedestrians at
close locations likely hide most of the braille
blocks. In such a case, the braille block region
detected in the previous frame will be used.</p>
        <p>
          Braille blocks may not be installed in a straight
line. Based on the guideline [
          <xref ref-type="bibr" rid="ref2">2</xref>
          ], braille blocks are
rarely curved in Japan. Therefore, we assume the
sidelines of the braille block region can be
approximated almost as straight lines.
(a) A situation where an oncoming pedestrian is
on the collision path.
(b) A situation where an oncoming pedestrian is
not on the collision path.
        </p>
        <p>Figure 3: Detection of oncoming pedestrians.
4.2.</p>
      </sec>
    </sec>
    <sec id="sec-9">
      <title>Distance estimation</title>
      <p>
        We need to estimate the distance to oncoming
pedestrians to predict a collision with oncoming
pedestrians. Methods for estimating the distance
to an obstacle using a camera have been proposed
[
        <xref ref-type="bibr" rid="ref22">38-40</xref>
        ]. Detected facial features [
        <xref ref-type="bibr" rid="ref22">38</xref>
        ] and
rectangle size [39, 40] are used. Chen et al. used
the camera’s focal length, angle of view, and
information on its orientation relative to the
ground [
        <xref ref-type="bibr" rid="ref18">34</xref>
        ]. Combined with the assumption that
the road surface is horizontal, the distance to the
obstacle can be estimated.
      </p>
      <p>
        In this research, we use the method [
        <xref ref-type="bibr" rid="ref18">34</xref>
        ] and
recognize an oncoming pedestrian as an obstacle.
We estimate the distance from the blind people to
the position where the oncoming pedestrian
stands, as shown in Figure 4.
4.3.
      </p>
    </sec>
    <sec id="sec-10">
      <title>Future position estimation</title>
      <sec id="sec-10-1">
        <title>Blind people walk on braille blocks installed in</title>
        <p>a straight line. If oncoming pedestrians also walk
on braille blocks, they walk straight toward blind
people. We must decide whether the oncoming
pedestrian walks straight toward the blind people.
We estimate the future position of the oncoming
pedestrian. The distance of the collision with the
oncoming pedestrian is about 60 cm, which is
within the reach of a white cane. Therefore, the
distance and future position estimations must be
accurate enough to meet this requirement.
Foot position of
an oncoming pedestrian</p>
        <p>
          Some research proposed a method for
estimating the trajectory of pedestrians from an
egocentric video [41, 42]. Yagi et al. estimate the
position of a pedestrian's waist using the
pedestrian's skeletal information and the camera's
pose information [41]. Qiu et al. estimate the
future position by referring to the property of a
rectangle instead of a point [42]. This method can
be combined with the distance estimation method
[
          <xref ref-type="bibr" rid="ref18">34</xref>
          ] to estimate the walking trajectory.
        </p>
      </sec>
    </sec>
    <sec id="sec-11">
      <title>5. The decision of whether oncoming pedestrians notice blind people</title>
      <p>Suppose the oncoming pedestrian is on the
collision path. In that situation, the remaining
problem is to decide whether the oncoming
pedestrian notices the blind people and intends to
avoid the collision. Our proposed system
estimates the gaze direction of the oncoming
pedestrians. It checks whether the gaze direction
is toward the location of the blind people.</p>
      <p>
        Lee et al. have proposed a system to decide
whether an oncoming pedestrian is looking at
blind people [
        <xref ref-type="bibr" rid="ref22">38</xref>
        ]. Their system is trained by the
annotations of the pedestrian's face images. The
direction of gaze has been estimated as a 3D
vector [
        <xref ref-type="bibr" rid="ref10 ref12 ref9">25, 26, 28, 43</xref>
        ]. As shown in Figure 5,
Zhang et al. achieve gaze estimation even when
the resolution of the cropped face image is low
[43].
      </p>
      <p>We plan to adapt the method [43] to gaze
estimation of the oncoming pedestrians to decide
whether they notice blind people. Gaze estimation
should start at the moment when the oncoming
pedestrian is detected, up to the time when the
collision occurs.</p>
    </sec>
    <sec id="sec-12">
      <title>6. Collision prediction with oncoming pedestrians</title>
      <sec id="sec-12-1">
        <title>To predict collision, we set two conditions.</title>
        <p>The first condition is whether the oncoming
pedestrian is on the collision path. The second
condition is whether the oncoming pedestrian
notices the blind people. Even when the oncoming
pedestrian keeps the collision path, the system
does not make a warning once it detects the
oncoming pedestrian gaze the blind people just at
a frame. The system calls the warning of collision
if the oncoming pedestrian comes within the
hazardous distance of the blind people without
even a glance at the blind people in their front.</p>
      </sec>
    </sec>
    <sec id="sec-13">
      <title>7. Conclusion</title>
      <p>We proposed a new method for predicting
collisions with oncoming pedestrians using a
smartphone’s camera. To predict collision, we set
two conditions. The first condition is whether the
oncoming pedestrian is on the collision path. The
second condition is whether the oncoming
pedestrian notices the blind people. We developed
a system for the first condition and showed the
snapshots of the results.</p>
      <p>We plan to incorporate the procedure of the
second condition into our system. Implementing
the total system on a smartphone will help blind
people to avoid collisions with oncoming
pedestrians.</p>
      <p>Part of this research is supported by JSPS
Kaken 22K19803.</p>
    </sec>
    <sec id="sec-14">
      <title>8. References</title>
      <p>Robot for Blind People, in: Proceedings of Conference on Computer Vision and Pattern
the 21st International ACM SIGACCESS Recognition (CVPR), New Orleans, LA,
Conference on Computers and Accessibility USA, 2022, pp. 19606-19616.
(ASSETS '19), ACM, New York, NY, USA, doi:10.1109/CVPR52688.2022.01902.
2019, pp. 68-82. [14] D. Hutabarat, M. Rivai, D. Purwanto, H.
doi:10.1145/3308561.3353771. Hutomo, Lidar-based Obstacle Avoidance
[7] S. Kayukawa, K. Higuchi, J. Guerreiro, S. for the Autonomous Mobile Robot, in: 2019
Morishima, Y. Sato, K. Kitani, C. Asakawa, 12th International Conference on
BBeep: A Sonic Collision Avoidance System Information and Communication
for Blind Travellers and Nearby Pedestrians, Technology and System (ICTS), Surabaya,
in: Proceedings of the 2019 CHI Conference Indonesia, 2019, pp. 197-202.
on Human Factors in Computing Systems doi:10.1109/ICTS.2019.8850952.
(CHI '19), ACM, New York, NY, USA, 2019, [15] Y. Han, I. H. Zhan, W. Zhao, J. Pan, Z.
Paper 52, pp. 1-12. Zhang, Y. Wang, Y. J. Liu, Deep
doi:10.1145/3290605.3300282. reinforcement learning for robot collision
[8] S. Kayukawa, T. Ishihara, H. Takagi, S. avoidance with self-state-attention and
Morishima, C. Asakawa, BlindPilot: A sensor fusion, in: IEEE Robotics and
Robotic Local Navigation System that Leads Automation Letters, 7(3), 2022, pp.
6886Blind People to a Landmark Object, in: 6893. doi:10.1109/LRA.2022.3178791.
Extended Abstracts of the 2020 CHI [16] T. Xu, S. Zhang, Z. Jiang, Z. Liu, H. Cheng,
Conference on Human Factors in Computing Collision Avoidance of High-Speed
Systems (CHI EA '20), ACM, New York, Obstacles for Mobile Robots via
MaximumNY, USA, 2020, pp. 1-9. Speed Aware Velocity Obstacle Method, in:
doi:10.1145/3334480.3382925. IEEE Access, 8, 2020, pp. 138493-138507.
[9] A. R. See, B. G. Sasing, W. D. Advincula, A doi:10.1109/ACCESS.2020.3012513.</p>
      <p>Smartphone-Based Mobility Assistant Using [17] L. Zeng, G. M. Bone, Mobile Robot
Depth Imaging for Visually Impaired and Collision Avoidance in Human
Blind, in: Applied Sciences, 12(6), 2802, Environments, in: International Journal of
2022. doi:10.3390/app12062802. Advanced Robotic Systems, 10(1), 2013, pp.
[10] M. Kuribayashi, S. Kayukawa, H. Takagi, C. 1-14. doi:10.5772/54933.</p>
      <p>Asakawa, S. Morishima, LineChaser: A [18] A. Colombi, M. Scianna, Modelling human
Smartphone-Based Navigation System for perception processes in pedestrian dynamics:
Blind People to Stand in Lines, in: A hybrid approach, in: Royal Society Open
Proceedings of the 2021 CHI Conference on Science, 4(3), 2017, 160561.
Human Factors in Computing Systems (CHI doi:10.1098/rsos.160561.
'21), ACM, New York, NY, USA, 33, 2021, [19] M. Dicks, C. Clashing, L. O’Reilly, C. Mills,
pp. 1-13. doi:10.1145/3411764.3445451. Perceptual-motor behaviour during a
[11] A. A. Díaz-Toro, S. E. C. Bastidas, E. C. simulated pedestrian crossing, in: Gait &amp;
Bravo, Vision-Based System for Assisting Posture, 49, 2016, pp. 241-245.
Blind People to Wander Unknown doi:10.1016/j.gaitpost.2016.07.003.
Environments in a Safe Way, in: Journal of [20] H. Murakami, T. Tomaru, C. Feliciani, Y.
Sensors, 2021, pp. 1-18. Nishiyama, Spontaneous behavioral
doi:10.1155/2021/6685686. coordination between avoiding pedestrians
[12] L. Rusli, B. Nurhalim, R. Rusyadi, Vision- requires mutual anticipation rather than
based vanishing point detection of mutual gaze, in: iScience, 25(11), 2022,
autonomous navigation of mobile robot for 105474. doi:10.1016/j.isci.2022.105474.
outdoor applications, in: Journal of [21] X. Zhang, Y. Sugano, M. Fritz, A. Bulling,
Mechatronics, Electrical Power, and Appearance-based gaze estimation in the
Vehicular Technology, 12(2), 2021, pp. 117- wild, in: 2015 IEEE Conference on
125. doi:10.14203/j.mev.2021.v12.117-125. Computer Vision and Pattern Recognition
[13] Z. Q. Cheng, Q. Dai, H. Li, J. Song, X. Wu, (CVPR), Boston, MA, USA, 2015, pp.
4511A. G. Hauptmann, Rethinking Spatial 4520. doi:10.1109/CVPR.2015.7299081.
Invariance of Convolutional Networks for [22] Y. Sugano, X. Zhang, A. Bulling,
Object Counting, in: 2022 IEEE/CVF AggreGaze: Collective Estimation of
Audience Attention on Public Displays, in:
2009 Third International Conference on
Multimedia and Ubiquitous Engineering,
Qingdao, China, 2009, pp. 137-141.</p>
      <p>doi:10.1109/MUE.2009.34.
[39] S. Duman, A. Elewi, Z. Yetgin, Distance</p>
      <p>Estimation from a Monocular Camera Using
Face and Body Features, in: Arabian Journal
for Science and Engineering, 47(2), 2022, pp.
1547–1557.doi:10.1007/s13369-021-06003w.
[40] K. Lee, D. Sato, S. Asakawa, C. Asakawa, H.</p>
      <p>Kacorri, Accessing Passersby Proxemic
Signals through a Head-Worn Camera:
Opportunities and Limitations for the Blind,
in: Proceedings of the 23rd International
ACM SIGACCESS Conference on
Computers and Accessibility (ASSETS '21).</p>
      <p>ACM, New York, NY, USA, 8, 2021, pp. 1–
15. doi:10.1145/3441852.3471232.
[41] T. Yagi, K. Mangalam, R. Yonetani, Y. Sato,</p>
      <p>Future Person Localization in First-Person
Videos, in: 2018 IEEE/CVF Conference on
Computer Vision and Pattern Recognition
(CVPR), Salt Lake City, UT, USA, 2018, pp.</p>
      <p>7593-7602. doi:10.1109/CVPR.2018.00792.
[42] J. Qiu, F. P.-W. Lo, X. Gu, Y. Sun, S. Jiang,</p>
      <p>B. Lo, Indoor Future Person Localization
from an Egocentric Wearable Camera, in:
2021 IEEE/RSJ International Conference on
Intelligent Robots and Systems (IROS),
Prague, Czech Republic, 2021,
pp.85868592.</p>
      <p>doi:10.1109/IROS51168.2021.9635868.
[43] M. Zhang, Y. Liu, F. Lu, GazeOnce:
Real</p>
      <p>Time Multi-Person Gaze Estimation, in:
2022 IEEE/CVF Conference on Computer
Vision and Pattern Recognition (CVPR),
New Orleans, LA, USA, 2022, pp.
41874196. doi:10.1109/CVPR52688.2022.00416.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>Road</given-names>
            <surname>Bureau</surname>
          </string-name>
          ,
          <article-title>Ministry of Land, Infrastructure, Transport and Tourism, Guidelines for the Development of Roadway Mobility Facilitation</article-title>
          , URL: https://www.mlit.go.jp/road/road/traffic/bf/k ijun/pdf/all.pdf.
          <article-title>(published in Japanese)</article-title>
          .
          <source>Accessed 22 Jun</source>
          .
          <year>2023</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>K.</given-names>
            <surname>Tokuda</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Mizuno</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Nishidate</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Arai</surname>
          </string-name>
          , and
          <string-name>
            <given-names>M.</given-names>
            <surname>Aoyagi</surname>
          </string-name>
          ,
          <article-title>Guidebook for the Proper Installation of Tactile Ground Surface Indicator (Braille Blocks): Common Installation Errors</article-title>
          .
          <source>International Association of Traffic and Safety Sciences</source>
          , Tokyo, Japan,
          <year>2008</year>
          . URL: https://www.iatss.or.jp/common/pdf/researc h/h966_e.pdf
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>M.</given-names>
            <surname>Apostolos</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Filios</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Llorente</surname>
          </string-name>
          ,
          <article-title>Reliable Ultrasonic Obstacle Recognition for Outdoor Blind Navigation</article-title>
          , in: Technologies,
          <volume>10</volume>
          (
          <issue>3</issue>
          ),
          <fpage>54</fpage>
          ,
          <year>2022</year>
          . doi:
          <volume>10</volume>
          .3390/technologies10030054.
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          <article-title>[4] WeWALK</article-title>
          . URL:https://wewalk.io/en/.
          <source>Accessed 22 Jun</source>
          .
          <year>2023</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>P.</given-names>
            <surname>Slade</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Tambe</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. J.</given-names>
            <surname>Kochenderfer</surname>
          </string-name>
          ,
          <article-title>Multimodal sensing and intuitive steering assistance improve navigation and mobility for people with impaired vision</article-title>
          , in: Science Robotics,
          <volume>6</volume>
          (
          <issue>59</issue>
          ),
          <year>2021</year>
          . doi:
          <volume>10</volume>
          .1126/scirobotics.abg6594
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>J.</given-names>
            <surname>Guerreiro</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Sato</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Asakawa</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Dong</surname>
          </string-name>
          ,
          <string-name>
            <surname>K. M. Kitani</surname>
          </string-name>
          , C. Asakawa,
          <source>CaBot: Designing and Evaluating an Autonomous Navigation Proceedings of the 29th Annual Symposium on User Interface Software and Technology (UIST '16)</source>
          , ACM, New York, NY, USA,
          <year>2016</year>
          , pp.
          <fpage>821</fpage>
          -
          <lpage>831</lpage>
          . doi:
          <volume>10</volume>
          .1145/2984511.2984536.
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [23]
          <string-name>
            <given-names>A.</given-names>
            <surname>Recasens</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Khosla</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Vondrick</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Torralba</surname>
          </string-name>
          ,
          <article-title>Where are they looking?</article-title>
          ,
          <source>in: Proceedings of the 28th International Conference on Neural Information Processing Systems - Volume 1 (NIPS '15)</source>
          ., MIT Press, Cambridge, MA,
          <year>2015</year>
          , pp.
          <fpage>199</fpage>
          -
          <lpage>207</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [24]
          <string-name>
            <given-names>E.</given-names>
            <surname>Chong</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Ruiz</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. M.</given-names>
            <surname>Rehg</surname>
          </string-name>
          ,
          <article-title>Detecting attended visual targets in video</article-title>
          ,
          <source>in: 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)</source>
          , Seattle, WA, USA,
          <year>2020</year>
          , pp.
          <fpage>5395</fpage>
          -
          <lpage>5405</lpage>
          . doi:
          <volume>10</volume>
          .1109/CVPR42600.
          <year>2020</year>
          .
          <volume>00544</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [25]
          <string-name>
            <given-names>P.</given-names>
            <surname>Kellnhofer</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Recasens</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Stent</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W.</given-names>
            <surname>Matusik</surname>
          </string-name>
          ,
          <string-name>
            <surname>A. Torralba,</surname>
          </string-name>
          <article-title>Gaze360: Physically unconstrained gaze estimation in the wild</article-title>
          ,
          <source>in: 2019 IEEE/CVF International Conference on Computer Vision</source>
          (ICCV), Seoul, Korea (South),
          <year>2019</year>
          , pp.
          <fpage>6911</fpage>
          -
          <lpage>6920</lpage>
          . doi:
          <volume>10</volume>
          .1109/ICCV.
          <year>2019</year>
          .
          <volume>00701</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [26]
          <string-name>
            <given-names>C.</given-names>
            <surname>Bermejo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Chatzopoulos</surname>
          </string-name>
          , P. Hui, EyeShopper: Estimating Shoppers'
          <article-title>Gaze using CCTV Cameras</article-title>
          ,
          <source>in: Proceedings of the 28th ACM International Conference on Multimedia (MM '20)</source>
          , ACM, New York, NY, USA,
          <year>2020</year>
          , pp.
          <fpage>2765</fpage>
          -
          <lpage>2774</lpage>
          . doi:
          <volume>10</volume>
          .1145/3394171.3413683.
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [27]
          <string-name>
            <given-names>H.</given-names>
            <surname>Tomas</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Reyes</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Dionido</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Ty</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Mirando</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Casimiro</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Guinto</surname>
          </string-name>
          ,
          <article-title>Goo: A dataset for gaze object prediction in retail environments</article-title>
          , in: 2021 IEEE/CVF Conference on
          <article-title>Computer Vision and Pattern Recognition Workshops (CVPRW), Nashville</article-title>
          ,
          <string-name>
            <surname>TN</surname>
          </string-name>
          , USA,
          <year>2021</year>
          , pp.
          <fpage>3119</fpage>
          -
          <lpage>3127</lpage>
          . doi:
          <volume>10</volume>
          .1109/CVPRW53098.
          <year>2021</year>
          .
          <volume>00349</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [28]
          <string-name>
            <given-names>S.</given-names>
            <surname>Nonaka</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Nobuhara</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Nishino</surname>
          </string-name>
          ,
          <article-title>Dynamic 3d gaze from afar: Deep gaze estimation from temporal eye-head-body coordination</article-title>
          ,
          <source>in: 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)</source>
          , New Orleans, LA, USA,
          <year>2022</year>
          , pp.
          <fpage>2182</fpage>
          -
          <lpage>2191</lpage>
          . doi:
          <volume>10</volume>
          .1109/CVPR52688.
          <year>2022</year>
          .
          <volume>00223</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [29]
          <string-name>
            <given-names>Y.</given-names>
            <surname>Zhang</surname>
          </string-name>
          , P. Sun,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Jiang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Yu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Weng</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Yuan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Luo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W.</given-names>
            <surname>Liu</surname>
          </string-name>
          ,
          <string-name>
            <surname>X. Wang,</surname>
          </string-name>
          <article-title>ByteTrack: Multi-Object Tracking by Associating Every Detection Box</article-title>
          ,
          <source>in: Proceedings of the European Conference on Computer Vision</source>
          (ECCV),
          <source>Tel Aviv, Israel</source>
          ,
          <year>2022</year>
          , pp.
          <fpage>1</fpage>
          -
          <lpage>21</lpage>
          . doi:
          <volume>10</volume>
          .1007/978-3-
          <fpage>031</fpage>
          - 20047-
          <issue>2</issue>
          _
          <fpage>1</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [30]
          <string-name>
            <given-names>N.</given-names>
            <surname>Aharon</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Orfaig</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B. Z.</given-names>
            <surname>Bobrovsky</surname>
          </string-name>
          , BoTSORT: Robust Associations MultiPedestrian Tracking,
          <source>in: arXiv preprint 2206.14651</source>
          ,
          <year>2022</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [31]
          <string-name>
            <given-names>Z. Q.</given-names>
            <surname>Cheng</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Q.</given-names>
            <surname>Dai</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Song</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Wu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A. G.</given-names>
            <surname>Hauptmann</surname>
          </string-name>
          ,
          <article-title>Rethinking Spatial Invariance of Convolutional Networks for Object Counting</article-title>
          , in: 2022 IEEE/CVF Conference on
          <article-title>Computer Vision and Pattern Recognition (CVPR), New Orleans</article-title>
          , LA, USA,
          <year>2022</year>
          , pp.
          <fpage>19606</fpage>
          -
          <lpage>19616</lpage>
          . doi:
          <volume>10</volume>
          .1109/CVPR52688.
          <year>2022</year>
          .
          <year>01902</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [32]
          <string-name>
            <given-names>A. H.</given-names>
            <surname>Khan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. S.</given-names>
            <surname>Nawaz</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Dengel</surname>
          </string-name>
          ,
          <article-title>Localized Semantic Feature Mixers for Efficient Pedestrian Detection in Autonomous Driving</article-title>
          ,
          <source>in: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)</source>
          .
          <year>2023</year>
          , pp.
          <fpage>5476</fpage>
          -
          <lpage>5485</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [33]
          <string-name>
            <given-names>I.</given-names>
            <surname>Hasan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Liao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S. U.</given-names>
            <surname>Akram</surname>
          </string-name>
          , L. Shao, Generalizable Pedestrian Detection:
          <article-title>The Elephant in the Room</article-title>
          , in: 2021 IEEE/CVF Conference on
          <article-title>Computer Vision and Pattern Recognition (CVPR), Nashville</article-title>
          ,
          <string-name>
            <surname>TN</surname>
          </string-name>
          , USA,
          <year>2021</year>
          , pp.
          <fpage>11323</fpage>
          -
          <lpage>11332</lpage>
          . doi:
          <volume>10</volume>
          .1109/CVPR46437.
          <year>2021</year>
          .
          <volume>01117</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          [34]
          <string-name>
            <given-names>Q.</given-names>
            <surname>Chen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Wu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Chen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Lin</surname>
          </string-name>
          , S. Cheng, Z. Wu,
          <article-title>Smartphone Based Outdoor Navigation and Obstacle Avoidance System for the Visually Impaired</article-title>
          ,
          <source>in: Multidisciplinary Trends in Artificial Intelligence</source>
          ,
          <volume>11909</volume>
          ,
          <year>2019</year>
          , pp.
          <fpage>26</fpage>
          -
          <lpage>37</lpage>
          . doi:
          <volume>10</volume>
          .1007/978-3-
          <fpage>030</fpage>
          -33709-
          <issue>4</issue>
          _
          <fpage>3</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          [35]
          <string-name>
            <given-names>B.-S.</given-names>
            <surname>Lin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.-C.</given-names>
            <surname>Lee</surname>
          </string-name>
          , P.-Y. Chiang,
          <article-title>Simple Smartphone-Based Guiding System for Visually Impaired People</article-title>
          , in: Sensors,
          <volume>17</volume>
          (
          <issue>6</issue>
          ),
          <fpage>1371</fpage>
          ,
          <year>2017</year>
          . doi: doi:10.3390/s17061371.
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          [36]
          <string-name>
            <given-names>C. Y.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Bochkovskiy</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Yuan</surname>
          </string-name>
          ,
          <string-name>
            <surname>M.</surname>
          </string-name>
          <article-title>Liao, YOLOv7: Trainable bag-of-freebies sets new state-of-the-art for real-time object detectors</article-title>
          ,
          <source>in: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)</source>
          ,
          <year>2023</year>
          , pp.
          <fpage>7464</fpage>
          -
          <lpage>7475</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          [37]
          <string-name>
            <given-names>S.</given-names>
            <surname>Nakamura</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Shishido</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Kameda</surname>
          </string-name>
          , Braille Bock Detection at Shortest Distance by Mobile Devices, in: International Workshop on Advanced Image
          <source>Technology (IWAIT)</source>
          <year>2023</year>
          ,
          <year>2023</year>
          , 6 pages.
          <source>doi:10.1117/12</source>
          .2666662.
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          [38]
          <string-name>
            <given-names>K. A.</given-names>
            <surname>Rahman</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. S.</given-names>
            <surname>Hossain</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. A.</given-names>
            <surname>Bhuiyan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Zhang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Hasanuzzaman</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Ueno</surname>
          </string-name>
          , Person to Camera
          <source>Distance Measurement Based on Eye-Distance</source>
          , in:
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>