<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <journal-title-group>
        <journal-title>IEEE Transactions on Robotics</journal-title>
      </journal-title-group>
    </journal-meta>
    <article-meta>
      <title-group>
        <article-title>Vehicle Telepresence Maneuvering with Live Video and 3D Point Cloud without Perceptible Communication Delays★</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Shota Matsuo</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Monica Perusquía-Hernández</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Naoya Isoyama</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Hideaki Uchiyama</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Kiyoshi Kiyokawa</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Nara Institute of Science and Technology</institution>
          ,
          <addr-line>Nara-ken</addr-line>
          ,
          <country country="JP">Japan</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2007</year>
      </pub-date>
      <volume>23</volume>
      <issue>2007</issue>
      <fpage>927</fpage>
      <lpage>941</lpage>
      <abstract>
        <p>Telepresence is a technology that allows users to experience a real-time experience as if they were in a remote location. However, there is often a delay in communication before the user sees the image of the control feedback. Such communication delays significantly reduce the operability at the remote site so is the user's sense of agency in control of, for example, a remote vehicle. To reduce perceptible communication delays, we propose a remote vehicle maneuvering telepresence system that displays a predicted view by using odometry and 3D point clouds while showing live video when the odometry catches up with the vehicle's current pose. As a result of evaluation experiments, we confirmed that the proposed system significantly improves subjective operability and workload.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;Telepresence</kwd>
        <kwd>Latency</kwd>
        <kwd>Odometry</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>awareness; or using predictive techniques.</p>
      <p>
        There are several approaches to predictive technology
Teleoperation is the electronic remote control of a machine to reduce the impact of delay. These approaches are
cateor vehicle, and remotely operated vehicles (ROVs) have gorized as dynamic system models and free models. Free
a wide range of applications on the ground, underwater, model approaches include information superposition
modin the air, and in space [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. Telepresence refers to “the els [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ], 3D graphic models [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ], and image processing [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ].
feeling of being physically present at a remote or simulated Information superposition [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ] and 3D graphic models [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ]
location.” Teleoperation with telepresence is expected to can significantly reduce work time. However, they tend
improve eficiency and reduce operator workload. to require sophisticated algorithms, expensive equipment,
      </p>
      <p>
        Remote control presents multiple challenges, one of and large amounts of information about the environment
which is delay in operation. In this study, we define it and the vehicle. Image processing-based approaches [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ]
as the delay between the operator’s input action (control can improve operator performance because it modifies
command) and the corresponding video display [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]. Tele- delayed video to mimic the motion and the environment in
operation in a delayed environment is dificult and very real-time. When extensive information about the vehicle
stressful for the operator. Also, it is observed as a high and its environment is not available, or when there is
cognitive workload [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ] and poor performance [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]. For no opportunity to use expensive equipment, image
proexample, an increased task completion time and reduced cessing can improve operator performance eficiently and
accuracy [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ]. Overcoming the detrimental efects of delay inexpensively.
in teleoperation can be accomplished by increasing the Based on existing image processing based techniques,
level of automation (i.e., reducing human control); provid- we propose a method to adaptively switch between
preing information that increases the operator’s situational dictive display and normal RGB images (delayed live
video) by applying position and orientation coordinate
APMAR’23: The 15th Asia-Pacific Workshop on Mixed and Aug- transformation to point cloud images generated from depth
mented Reality, Aug. 18-19, 2023, Taipei, Taiwan information acquired by a stereo camera mounted on a
∗Corresponding author. vehicle, depending on the vehicle status. This predictive
($M.shPoenrugswqauiítaa-iHl@ergnmánaidle.czo);mis(oSy.aMmaa@tsuiso.)n;amis.tp.jepru(Nsq.uIisao@yaism.naa);ist.jp display is based on stereo cameras and odometry, which
hideaki.uchiyama@is.naist.jp (H. Uchiyama); kiyo@is.naist.jp can be applied to various vehicle configurations. This
(K. Kiyokawa) study focuses on operator performance and subjectively
 https://carelab.info/en/kiyoshi-kiyokawa/ (K. Kiyokawa) experienced workload while using the predictive display.
      </p>
      <p>
        0000-0002-0486-1743 (M. Perusquía-Hernández); Therefore, we investigate whether a simple predictive
(0H00.0U-c0h0i0y2a-m65a3);50-8040309-0(0N0.3-I2so2y6a0m-1a7);0700(0K0.-0K0i0y2o-k6a1w1a9)-1184 display can improve operator performance and reduce
sub© 2023 Copyright for this paper by its authors. Use permitted under Creative Commons License jective workload during remote control. An experiment
CPWrEooUrckReshdoinpgs IhStpN:/c1e6u1r3-w-0s.o7r3g ACttEribUutiRon 4W.0 Iontrekrnsathioonapl(CPCrBoYc4e.0e).dings (CEUR-WS.org) was set up to investigate changes in operator performance
and workload when piloting a vehicle under the following
three diferent video display conditions. The conditions
were: 1) RGB video (delayed live video), 2) point cloud
video (a naïve predictive display), and 3) automatic
switching between RGB and point cloud video. Participants were
asked to perform two simple navigation tasks: a driving
task and an observation task using a vehicle captured
by a first-person camera. Data collected included task
performance, perceived workload, and demographics. A
Likert scale questionnaire was administered to  = 15
participants to test two hypotheses related to task
performance and subjective workload using ANOVA [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ]. The
results suggest that the proposed system actually improves
subjective operability and workload.
      </p>
    </sec>
    <sec id="sec-2">
      <title>2. Related research</title>
      <p>The challenges of telepresence, delays, and their
detrimental efects are described from two perspectives: operator
performance and subjective workload. Means to
compensate for delays, especially forecasting techniques, are
described.</p>
      <sec id="sec-2-1">
        <title>2.1. Issues in Telepresence</title>
        <p>
          Draper et al. [
          <xref ref-type="bibr" rid="ref1">1</xref>
          ] defined telepresence as the perception
of presence in a physically remote or simulated location.
According to this definition, teleoperation is a subclass of
telepresence [
          <xref ref-type="bibr" rid="ref9">9</xref>
          ]. It is hypothesized that telepresence is
beneficial to mission execution and, furthermore, can lead
to increased eficiency and reduced operator workload.
Chen et al. [
          <xref ref-type="bibr" rid="ref2">2</xref>
          ] reviewed 150 articles that investigated
factors related to telepresence and how they afect operator
performance and related issues. They found eight main
factors: field of view (FOV), orientation, camera
viewpoint, depth perception, video quality, frame rate, time
delay, and motion.
remember the command given until he sees the vehicle
perform the desired action in the video [
          <xref ref-type="bibr" rid="ref3">3</xref>
          ]. Furthermore,
it must be psychologically connected to the previously
entered command (the previous state of the vehicle), and
then a new command must be entered based on this
combination of information when new information appears in
the video [
          <xref ref-type="bibr" rid="ref13">13</xref>
          ]. Thus, delays can degrade the operator’s
performance [
          <xref ref-type="bibr" rid="ref2">2</xref>
          ] and increase the subjectively perceived
workload [
          <xref ref-type="bibr" rid="ref13">13</xref>
          ].
        </p>
      </sec>
      <sec id="sec-2-2">
        <title>2.3. Approaches to Mitigate Delays</title>
        <p>
          There are multiple approaches to mitigating the detrimental
efects of delays. The first option is to increase the level
of automation (LOA) to reduce operator workload and
improve safety [
          <xref ref-type="bibr" rid="ref14 ref15 ref16 ref17">14, 15, 16, 17</xref>
          ]. The second option is to
present the operator with previously given maneuvering
commands to increase situational awareness, resulting in
higher performance and reduced subjective workload [
          <xref ref-type="bibr" rid="ref18 ref19 ref2">2,
18, 19</xref>
          ]. The third option is predictive technology, which
is a display, control algorithm, or graphical model that
attempts to predict the state of the vehicle based on the
current state of the vehicle and the commands entered
by the operator. Chen et al. [
          <xref ref-type="bibr" rid="ref2">2</xref>
          ] conclude that the third
option is the most promising solution when it is not
possible to eliminate delay from the system and emphasize
that predictive displays have been shown to reduce task
completion time by up to 50%.
        </p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>3. Proposed System</title>
      <sec id="sec-3-1">
        <title>3.1. Prerequisite</title>
        <p>We assume a situation in which a remote vehicle-type
telepresence is remotely piloted. The user, hereinafter
referred to as the “operator”, is supposed to operate the
vehicle and view the camera mounted on the vehicle by
keyboard input while watching the PC screen.</p>
        <p>
          In general, there is a delay between the time when a
2.2. Type of Delay control command is inputted and when it is reflected in
The literature [
          <xref ref-type="bibr" rid="ref2">2</xref>
          ] focused on delays in maneuvering com- the image that the operator sees. This is due to the fact
mands, or operator input actions; and time delays, or the that data is transferred between the time when the input
delay between the input action and the display of the control command arrives at the remote vehicle and the
corresponding image. There is an important distinction time when the control command is executed and returned
between these two delays, with diferent efects on perfor- to the operator as a video image.
mance [
          <xref ref-type="bibr" rid="ref10 ref11 ref12">10, 11, 12</xref>
          ]. The maneuvering command delays Therefore, the proposed system sends the operator’s
are outside the scope of this study; we consider only time input to the remote location and at the same time reflects
delays. We focus on the perceived delay, i.e., the time it in the local environment to present a video image that
between the operator’s input of a maneuver command and immediately reflects the maneuver. To achieve this, a
the visual perception of the vehicle’s response in the video. stereo camera and an odometry calculation module are
The delay creates a gap between the command given and provided. The stereo camera can generate a 3D point cloud
the visual feedback showing the vehicle’s response. This of the remote location by acquiring depth information in
causes inconsistencies in the operator’s perception. To addition to RGB information. The odometry computation
correct this during maneuvering, the operator needs to module is provided as standard when the remote control is
        </p>
        <sec id="sec-3-1-1">
          <title>Windows</title>
          <p>DS-700AC</p>
        </sec>
        <sec id="sec-3-1-2">
          <title>Router ZED Mini</title>
        </sec>
        <sec id="sec-3-1-3">
          <title>Pioneer 3-DX</title>
          <p>Predicted Image
using Point Cloud
&amp; Odometry</p>
          <p>Distance PPoriendticCtleodudIm&amp;aOgedoumsientgry Vehicle Movement</p>
          <p>Transferred</p>
          <p>Delay Delayed TImimaege
realized by ROS. For example, a single command to move
forward causes the vehicle to move 10 cm forward, and
so on. The odometry computation module estimates the
current position and attitude of the vehicle by integrating
the angle of rotation of the wheels. By performing this
estimation in the local environment as well, it can estimate
how the remote vehicle should behave according to the
input maneuvering commands, and render a 3D point cloud
from the estimated position and orientation to present the
operator with a realistic reproduced image that immediately
responds to the maneuvering commands without waiting
for the actual image to arrive. The realistic reproduced
image that immediately responds to the control command
is presented to the operator without waiting for the actual 3.3. Software
image to arrive. If the local environment and the remote To render the 3D point cloud using the RGB and
vehicle’s position and orientation are judged to match depth information acquired by the ZED Mini, we used
based on the amount of delay and odometry information, depth_sensing.py, available on GitHub from
Stereothe delayed RGB image (live video) is presented as is, labs, the distributor of the ZED Mini. The point cloud
instead of the 3D point cloud. This prevents deterioration rendering part of this program was modified to render
of the video quality at the remote location due to inaccurate based on the vehicle odometry information. The
moveor missing depth information when the vehicle is stopped. ment of the virtual camera in the 3D point cloud and the
movement of the vehicle in the real world as a result of
3.2. Hardware the control is shown in Figure 3.</p>
          <p>We implemented the system so that the control
commands entered by the operator are sent to the vehicle
via a socket communication. The program1 available on
The hardware used for the prototype is shown in Figure 1
and the communication of the system is shown in Figure 2.</p>
          <p>A ZED Mini was used for the stereo camera and a Pioneer
3-DX was used for the vehicle. To transfer the images from
the stereo camera to the operator’s PC, we used a USB
1https://github.com/kumahika/delivery_navigation
GitHub that can control the Pioneer 3-DX was modified
to control the vehicle via keyboard input.</p>
        </sec>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>4. Evaluation Experiment</title>
      <sec id="sec-4-1">
        <title>4.1. Overview</title>
        <p>
          To investigate the impact of the proposed method on the
performance and workload when maneuvering a vehicle,
we conducted an evaluation experiment comparing three
feedback conditions. Participants were asked to perform
a driving task in which they drove a vehicle on a route
with repeated curves in all the diferent conditions, and
an observation task in which they answered questions
based on images from three monitors facing in various
directions during the driving task. These tasks were
designed because they involve multiple types of tasks,
especially lateral movement tasks. The benefits of the
prediction technique are highly task-dependent [
          <xref ref-type="bibr" rid="ref2">2</xref>
          ]. Since
the maximum speed of the vehicle produces a ceiling
efect, it is likely that the prediction display is more useful
for tasks that involve lateral movement than for tasks that
primarily involve long forward movement.
        </p>
        <p>The following hypotheses are tested to investigate the
efectiveness of the proposed method.</p>
        <p>H1 An improvement in worker performance with a hybrid
method when compared to single RGB and point
cloud methods.</p>
        <p>H2 An improvement in the subjective maneuverability
of the worker with a hybrid method method when
compared to single RGB and point cloud methods.</p>
        <p>Figure 4 shows what the operator sees in the video
while the vehicle is moving forward, stopping, and turning
in each of the three conditions. In RGB, the images
are presented as captured by the camera in Figure 4(a)
forward, Figure 4(d) stop, and Figure 4(g) turning. On
the PC, the images are presented as a 3D point cloud
in all cases shown in Figure 4(b) forward, Figure 4(e)
stop, and Figure 4(h) turning. By immediately reflecting
the maneuvering commands input to the virtual camera
in the 3D point cloud, the reproduced image seen when
the maneuver is reflected is presented without waiting
time. However, the reproduced image using the 3D point
cloud image has many missing parts, and therefore, the
image quality is inferior to the RGB condition. In the
proposed HB method, the camera presents the image as it
captures it only when the vehicle is stopped (Figure 4(f)),
and presents the image as a 3D point cloud when the
vehicle is moving (e.g., forward (Figure 4(c)) or turning
(Figure 4(i))). As soon as a maneuvering command is
inputted, the system switches to presenting the image in
the 3D point cloud, and when the vehicle catches up with
the virtual camera, the system switches to presenting the
image as it was captured by the camera.</p>
        <p>To avoid order efects and learning efects, the order in
which the conditions were implemented was randomized
using the 3 × 3 Latin square method. Participants were
not informed of the characteristics and functions of each
condition, and their ability to use them was left to their
intuitive comprehension.</p>
      </sec>
      <sec id="sec-4-2">
        <title>4.3. Participants</title>
        <p>= 15 volunteers participated (age range: 22-28 years,
average age: 23.5 ± 1.5 years, 3 females). All experiments
4.2. Experimental Conditions in this research were approved by the ethical review board
The three conditions are RGB (delayed live video), PC of the Nara Institute of Science and Technology with
(Point Cloud), and HB (Hybrid, i.e., the proposed method). review code 2022-I-41.</p>
        <p>As previously described, the intentionally added delay
was 2 seconds for all conditions, meaning that the total 4.4. Task
efective delay was about 2.5 seconds.</p>
        <p>In general, experiments on forecasting technology measure
RGB This is a naïve remote control method that presents performance using two metrics: course completion time
RGB images (live video) with large delay as they and task score. The former measures the time required to
are. complete a course for the task of moving a vehicle along
a predefined path. In the latter case, the task is to move
PC The 3D point cloud is drawn from the vehicle’s esti- toward a given target and complete the indicated task. In
mated current position and orientation based on this experiment, the participants were asked to look over
odometry information, and a reproduced image is a fence and answer formulas displayed on three monitors
presented that immediately reflects the maneuver- facing in diferent directions. The formulas included
ing of the vehicle. addition with a decimal point (e.g., 12 + 3.4 =?), and the
HB The proposed method combines RGB and PC. Nor- task was designed in such a way that the participants would
mally, a PC is used, and is switched to RGB make a mistake if they overlooked the decimal point.
adaptively when the real vehicle catches up with
the position and orientation of the virtual camera
in the 3D point cloud.</p>
        <p>(a) RGB: Moving forward
(b) PC: Moving forward
(c) HB: Moving forward
(d) RGB: Standstill
(e) PC: Standstill
(f) HB: Standstill
(g) RGB: While turning
(h) PC: While turning
(i) HB: While turning</p>
      </sec>
      <sec id="sec-4-3">
        <title>4.5. Experimental environment</title>
        <p>A 17.3-inch laptop computer with an Intel Core i9-9900K
CPU 3.60 GHz and Windows 10 was used as the operator’s
remote control computer. The laptop screen functioned
as a monitor, and the WASD and QEZXC keys on the
keyboard were used to control the vehicle. The M key
was used to record the time stamps that the participants
themselves pressed at the beginning and end of each of the
two tasks. The keyboard and mouse were used to answer
the questionnaire.</p>
        <p>A photograph of the test course is shown in Figure 5(a)
and its description in Figure 5(b). The green area in
Figure 5(b) was the driving task route, and the red area
was judged to be out of course. No time limit was set for
the driving task, and participants were instructed to drive
as safely as possible. If the participants went of course,
the number of times they went of course was recorded
while the driving task and the experiment continued. The
course line of the test course consisted of plastic chains so
that the experimenter would be able to recognize by sound
when the vehicle went of the course or rode up on the
course. The participants, who were the operators, were
not aware of the sound of the plastic chain because they
were wearing noise-canceling headphones playing white
noise during the maneuver.</p>
        <p>We used three monitors facing diferent directions at the
two observation points (Figure 6) placed in the gray area
shown in Figure 5(b). Each of the three monitors displayed
a diferent formula which flowed from right to left. A fence
was placed in front of the monitors so that the monitors’
how it worked. Therefore, the use of the prediction display
was left to the intuitive understanding of the individual
participant.</p>
      </sec>
      <sec id="sec-4-4">
        <title>4.7. Measurements</title>
        <p>Participants’ subjective delay experience in each condition
was investigated, and a perceived delay time question was
asked in the questionnaire to compare each condition. By
asking about perceived delay, we hoped to provide an
indication of the efectiveness of the predictive display in
reducing the perceived delay of the system.</p>
        <p>As evaluation indices, four items were used: time, the
number of times each maneuver key was entered, the
number of times each maneuver key was released, and the
vehicle’s xyz coordinates, quaternion number, and
movement speed were recorded from the time the participant
started the experiment until the end. These indices were
continuously collected during the execution of the task in
each condition. After each condition, the participants were
asked to respond to a questionnaire created with Google
Forms. The questionnaire consisted of 15 questions: 10
Likert-type questions in which participants were asked
to select one of seven levels from 1 (strongly disagree)
to 7 (strongly agree), and five open questions in which
they were asked to write freely about the content of the
questions.</p>
        <p>See Figure 8 for the text of each question.
screens are visible only from the front. The operator
moved the vehicle within the observation task area, the
yellow area in Figure 5(b), and maneuveres it to a position
where the monitor screen was visible. The purpose of this
arrangement was to prevent the operator from completing
the observation task without maneuvering the vehicle,
and to encourage the operator to maneuver the vehicle to
the position in front of each of the three monitors. The
observation task was terminated by having the operator
verbally answer all of the formulas displayed on each of
the three monitors. The number of incorrect answers and
their causes were recorded.</p>
      </sec>
      <sec id="sec-4-5">
        <title>4.6. Experimental procedure</title>
      </sec>
      <sec id="sec-4-6">
        <title>4.8. Result</title>
        <p>
          Participants were first briefed on the experimental flow
using slide materials after reviewing the experimental envi- The task completion time for each of the driving and
ronment to understand the situation and the issues involved. observation tasks in each condition is presented in a
boxParticipants were seated in a chair at a desk with a laptop and-whisker diagram. Tests were conducted to confirm
computer, with their backs to the area where the vehicle if there were any significant diferences, as shown in
Figwas moving. During the experiment, participants did not ure 7. A one-way repeated measures analysis of variance
have direct visual contact with the moving vehicle and (ANOVA) was conducted to determine if the diferences
wore headphones playing white noise to avoid the auditory in the three conditions were statistically significant. The
perception of the vehicle. After agreeing to participate Holm method was used to correct for multiple comparisons.
in the experiment, participants completed a demographic For the driving task, there was a significant diference
bequestionnaire. The experiment was described in terms tween RGB and HB (p&lt;0.05), and the proposed condition
of the experimental procedure, the test course, the types reduced the task completion time. For the observation
of tasks, the content of each task, and the maneuvering task, a trend toward significance was found between PC
method. At the beginning of each condition, there was a and HB (p&lt;0.10). Also, there was no significant
difer30-second practice period to familiarize themselves with ence between RGB and HB. Therefore, we confirmed that
the conditions. Thereafter, the participants performed the the proposed condition was as eficient as RGB for the
driving task and the observation task consecutively, with observation task. From these results, H1 “Significant
the end of the observation task marking the end of one con- improvement in worker performance” was supported.
dition. In the driving task, the experiment was continued The Likert scale ratings in each condition are presented
even if the vehicle went of the course, and the number of in a box-and-whisker diagram for each question, and
times this occurred was recorded by the experimenter. In tests were conducted to confirm significant diferences,
the observation task, participants were asked to answer a as shown in Figure 8. Simply treating the ratings as an
total of three questions on three monitors, and the number interval scale [
          <xref ref-type="bibr" rid="ref8">8</xref>
          ], a one-way repeated measures analysis
of incorrect answers was also recorded by the experimenter. of variance (ANOVA) was conducted to determine if
Participants were not informed of the predictive display or the diferences per question in the three conditions were
statistically significant. The Holm method was used to
correct for multiple comparisons. Of the 10 Likert-scale
questions, PC and HB were found to be significantly
diferent from RGB in nine of the questions, respectively,
supporting the alternative hypothesis. The remaining
question was about “how easy it was to see the image”.
        </p>
        <p>Significant diferences were found between RGB and PC,
and HB and PC, respectively. These results supported the
alternative hypothesis.</p>
        <p>Significant diferences between RGB and PC, and
between RGB and HB were confirmed for impression, sense
of security, comfort, discomfort, operability, and delay, at
p&lt;0.001. A significant diference of p&lt;0.01 was found
between RGB and PC, and between RGB and HB in the
sense of being able to control by one’s own will.</p>
        <p>For the question of visibility, a significant diference
of p&lt;0.001 was confirmed between RGB and PC, and
between PC and HB.</p>
        <p>A significant diference of p&lt;0.001 was found between
RGB and PC, and between PC and HB in the question of
whether the pilot was able to maneuver as intended. A
significant diference of p&lt;0.01 was found between RGB
and HB. A significant diference of p&lt;0.01 was found
between the RGB and PCs in the sense of local control. A
significant diference of p&lt;0.001 was found between RGB
and HB.</p>
        <p>From these results, we confirmed that H2 “Significant
improvement in subjective maneuverability of the worker”
was supported.</p>
      </sec>
      <sec id="sec-4-7">
        <title>4.9. Comments per condition</title>
        <p>Below we list representative participants’ comments
obtained for each condition per question. Overall, the
comments were more positive for the HB condition than for
the others.
4.9.1. RGB
• “There was a considerable delay, and it took some
time for the intended operation to be reflected,
during which time I felt uneasy.” (Male, 23 years
old),
• “The delay was so large that I felt insecure about
whether the input was correct or not.” (Male, 22
years old),
• “I felt a considerable delay, which made it dificult
to make fine adjustments and made me uneasy.”
(Male, 22 years old),
• “I felt uneasy because I did not know how much I
had to press a button to move forward or backward.”
(Female, 24 years old)
• “If there had been no delay, I would have been
able to control the vehicle as intended.” (Male, 23
years old).
4.9.2. PC
• “The robot responded quickly, so it seemed to stop
when the timing was right.” (Male, 25 years old),
• “The robot’s position was easy to grasp without
any delay.” (Male, 22 years old),
• “There was no stuttering or switching of images,
and I felt that the robot could be operated
comfortably.” (Male, 23 years old),
• “I was able to control the robot as I expected.”
(Male, 23 years old),
• “There was no problem in driving, but the image
was unclear when performing observation tasks.”
(Male, 23 years old),
• “The image was rough and there was no image
when turning.” (Male, 24 years old), and
• “When the visibility of the image was extremely
poor, I strongly felt that it was a remote control.”
(Male, 24 years old).
7
6
ltrscea45
ikLe32
1
7
6
ltrscea45
ikeL32
1
***</p>
        <p>***
**</p>
        <p>**
PointCloud
Condition
PointCloud
Condition
PointCloud
Condition</p>
        <p>Ours
Ours
Ours</p>
        <p>RGB</p>
        <p>Ours</p>
        <p>RGB
(a) Good impression of the maneuver.
(b) Secure in maneuvering.
(c) Comfort in maneuvering.</p>
        <p>Ours</p>
        <p>RGB</p>
        <p>Ours</p>
        <p>RGB
(d) Comfortable with the controls.
(e) Maneuvering as intended.
(f) Operability was good.</p>
        <p>Ours</p>
        <p>RGB</p>
        <p>Ours</p>
        <p>RGB</p>
        <p>Ours
(g) The video was easy to view.
(h) Not notice any major delays.</p>
        <p>(i) Piloting of my own volition.
***
***
***
***
PointCloud</p>
        <p>Condition
***</p>
        <p>**
PointCloud</p>
        <p>Condition
***</p>
        <p>***
PointCloud</p>
        <p>Condition
**
***
7
6
ltrscea54
ikeL32
1
7
6
ltrscea54
ikeL32
1
7
6
ltrscea54
ikeL32
1</p>
        <p>RGB
RGB
***</p>
        <p>***
***</p>
        <p>***
PointCloud
Condition
PointCloud</p>
        <p>Condition
***</p>
        <p>***
RGB</p>
        <p>PointCloud
Condition
7
6
ltrscae45
ikeL32
1
7
6
ltrscea45
ikLe23
1
7
6
ltrscea45
ikLe32
1</p>
        <p>RGB</p>
        <p>Ours
PointCloud</p>
        <p>Condition
(j) Feeling that piloting the vehicle
locally, not remotely.
4.9.3. HB (Proposed Method)
• “It was the easiest to maneuver.” (Male, 24 years
old),
• “It was easy to operate with little delay.” (Female,
24 years old),
• “The front/rear maneuvering was almost no delay
at all.” (Male, 24 years old),
• “The observation task was quite comfortable,
although the impression was not so diferent from
other conditions when driving.” (Male, 23 years
old),
• “The response time was fast, so it seemed that the
camera stopped when the user wanted it to stop.”
(Male, 25 years old),
• “There was some discomfort in the switching
between low and high-quality images.” (Male, 23
years old),
• “The field of view was narrow and it was dificult
to operate when turning, although I did not feel
much delay when moving straight ahead.” (Male,
23 years old),
• “It was stressful to have to wait until the delay was
resolved because the image did not catch up with
the rotation.” (Male, 24 years old), and
• “While turning, there was a delay while waiting
for the image to be displayed.” (Male, 24 years
old).</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>5. Conclusions</title>
      <p>Dealing with delays is one of the major challenges in
telepresence. Therefore, we proposed a telepresence
system that automatically switches between predictive
images and live images using 3D point clouds according
to the situation. It also presents images that show the
user’s maneuvering immediately so that the user does
not perceive any delay in the images. As a result of
evaluation experiments, we confirmed that the proposed
system actually improves subjective operability and ease
of viewing. In the future, we aim to resolve issues such as
the phenomenon of the edges of the screen being blurred
and many parts not being rendered when turning on PCs
and HBs. This can be achieved, by switching between
RGB images and 3D point cloud-reproduced images and
by devising a new method for rendering point clouds.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>J. V.</given-names>
            <surname>Draper</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D. B.</given-names>
            <surname>Kaber</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. M.</given-names>
            <surname>Usher</surname>
          </string-name>
          , Telepresence, Human factors
          <volume>40</volume>
          (
          <year>1998</year>
          )
          <fpage>354</fpage>
          -
          <lpage>375</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>J. Y.</given-names>
            <surname>Chen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E. C.</given-names>
            <surname>Haas</surname>
          </string-name>
          ,
          <string-name>
            <surname>M. J. Barnes,</surname>
          </string-name>
          <article-title>Human performance issues and user interface design for teleoperated robots</article-title>
          ,
          <source>IEEE Transactions on Systems, Man, and Cybernetics</source>
          , Part C (
          <article-title>Applications</article-title>
          and Reviews)
          <volume>37</volume>
          (
          <year>2007</year>
          )
          <fpage>1231</fpage>
          -
          <lpage>1245</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>A.</given-names>
            <surname>Matheson</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Donmez</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Rehmatullah</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Jasiobedzki</surname>
          </string-name>
          , H.
          <article-title>-</article-title>
          <string-name>
            <surname>K. Ng</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          <string-name>
            <surname>Panwar</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          <string-name>
            <surname>Li</surname>
          </string-name>
          ,
          <article-title>The efects of predictive displays on performance in driving tasks with multi-second latency: Aiding tele-operation of lunar rovers</article-title>
          ,
          <source>in: Proceedings of the Human Factors and Ergonomics Society Annual Meeting</source>
          , volume
          <volume>57</volume>
          ,
          <string-name>
            <given-names>SAGE</given-names>
            <surname>Publications Sage</surname>
          </string-name>
          <string-name>
            <surname>CA</surname>
          </string-name>
          : Los Angeles, CA,
          <year>2013</year>
          , pp.
          <fpage>21</fpage>
          -
          <lpage>25</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>J. C.</given-names>
            <surname>Lane</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C. R.</given-names>
            <surname>Carignan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B. R.</given-names>
            <surname>Sullivan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D. L.</given-names>
            <surname>Akin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Hunt</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Cohen</surname>
          </string-name>
          ,
          <article-title>Efects of time delay on telerobotic control of neutral buoyancy vehicles</article-title>
          ,
          <source>in: Proceedings 2002 IEEE international conference on robotics and automation (Cat. No. 02CH37292)</source>
          , volume
          <volume>3</volume>
          , IEEE,
          <year>2002</year>
          , pp.
          <fpage>2874</fpage>
          -
          <lpage>2879</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>T. B.</given-names>
            <surname>Sheridan</surname>
          </string-name>
          ,
          <article-title>Space teleoperation through time delay: Review and prognosis</article-title>
          ,
          <source>IEEE Transactions on robotics and Automation</source>
          <volume>9</volume>
          (
          <year>1993</year>
          )
          <fpage>592</fpage>
          -
          <lpage>606</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>A. K.</given-names>
            <surname>Bejczy</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W. S.</given-names>
            <surname>Kim</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S. C.</given-names>
            <surname>Venema</surname>
          </string-name>
          ,
          <article-title>The phantom robot: predictive displays for teleoperation with time delay</article-title>
          ,
          <source>in: Proceedings., IEEE International Conference on Robotics and Automation</source>
          , IEEE,
          <year>1990</year>
          , pp.
          <fpage>546</fpage>
          -
          <lpage>551</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>H.</given-names>
            <surname>Dybvik</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Løland</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Gerstenberg</surname>
          </string-name>
          ,
          <string-name>
            <surname>K. B. Slåttsveen</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          <string-name>
            <surname>Steinert</surname>
          </string-name>
          ,
          <article-title>A low-cost predictive display for teleoperation: Investigating efects on human performance and workload</article-title>
          ,
          <source>International Journal of Human-Computer Studies</source>
          <volume>145</volume>
          (
          <year>2021</year>
          )
          <fpage>102536</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <surname>J. F. de Winter</surname>
          </string-name>
          , D. Dodou,
          <article-title>Five-point likert items: t test versus mann-whitney-wilcoxon (addendum added october</article-title>
          <year>2012</year>
          ), Practical Assessment, Research, and Evaluation 15 (
          <year>2010</year>
          )
          <fpage>11</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>T. B.</given-names>
            <surname>Sheridan</surname>
          </string-name>
          ,
          <article-title>Teleoperation, telerobotics and telepresence: A progress report</article-title>
          ,
          <source>Control Engineering Practice</source>
          <volume>3</volume>
          (
          <year>1995</year>
          )
          <fpage>205</fpage>
          -
          <lpage>214</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>J.</given-names>
            <surname>Davis</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Smyth</surname>
          </string-name>
          ,
          <string-name>
            <surname>K.</surname>
          </string-name>
          <article-title>McDowell, The efects of time lag on driving performance and a possible mitigation</article-title>
          ,
          <source>IEEE Transactions on Robotics</source>
          <volume>26</volume>
          (
          <year>2010</year>
          )
          <fpage>590</fpage>
          -
          <lpage>593</lpage>
          . doi:
          <volume>10</volume>
          .1109/TRO.
          <year>2010</year>
          .
          <volume>2046695</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>S.</given-names>
            <surname>Neumeier</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Wintersberger</surname>
          </string-name>
          ,
          <string-name>
            <surname>A.-K. Frison</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Becher</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          <string-name>
            <surname>Facchi</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Riener</surname>
          </string-name>
          ,
          <article-title>Teleoperation: The holy grail to solve problems of automated driving? sure, but latency matters</article-title>
          ,
          <source>in: Proceedings of the 11th International Conference on Automotive User Interfaces and Interactive Vehicular Applications</source>
          ,
          <year>2019</year>
          , pp.
          <fpage>186</fpage>
          -
          <lpage>197</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>R.</given-names>
            <surname>Oboe</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Fiorini</surname>
          </string-name>
          ,
          <article-title>A design and control environment for internet-based telerobotics</article-title>
          ,
          <source>The International journal of robotics research 17</source>
          (
          <year>1998</year>
          )
          <fpage>433</fpage>
          -
          <lpage>449</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <given-names>B.</given-names>
            <surname>Ricks</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Nielsen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Goodrich</surname>
          </string-name>
          ,
          <article-title>Ecological displays for robot interaction: a new perspective</article-title>
          ,
          <source>in: 2004 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS) (IEEE Cat. No.04CH37566)</source>
          , volume
          <volume>3</volume>
          ,
          <year>2004</year>
          , pp.
          <fpage>2855</fpage>
          -
          <lpage>2860</lpage>
          vol.
          <volume>3</volume>
          . doi:
          <volume>10</volume>
          .1109/IROS.
          <year>2004</year>
          .
          <volume>1389842</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <given-names>G.</given-names>
            <surname>Dorais</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R. P.</given-names>
            <surname>Bonasso</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Kortenkamp</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Pell</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Schreckenghost</surname>
          </string-name>
          ,
          <article-title>Adjustable autonomy for humancentered autonomous systems</article-title>
          ,
          <source>in: Working notes of the sixteenth international joint conference on artificial intelligence workshop on adjustable autonomy systems</source>
          ,
          <year>1999</year>
          , pp.
          <fpage>16</fpage>
          -
          <lpage>35</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <given-names>M. R.</given-names>
            <surname>Endsley</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D. B.</given-names>
            <surname>Kaber</surname>
          </string-name>
          ,
          <article-title>Level of automation efects on performance, situation awareness and workload in a dynamic control task</article-title>
          ,
          <source>Ergonomics</source>
          <volume>42</volume>
          (
          <year>1999</year>
          )
          <fpage>462</fpage>
          -
          <lpage>492</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [16]
          <string-name>
            <given-names>M. A.</given-names>
            <surname>Goodrich</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D. R.</given-names>
            <surname>Olsen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. W.</given-names>
            <surname>Crandall</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T. J.</given-names>
            <surname>Palmer</surname>
          </string-name>
          ,
          <article-title>Experiments in adjustable autonomy</article-title>
          , in: Proceedings of ĲCAI Workshop on autonomy,
          <article-title>delegation and control: interacting with intelligent agents</article-title>
          , Seattle, WA,
          <year>2001</year>
          , pp.
          <fpage>1624</fpage>
          -
          <lpage>1629</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [17]
          <string-name>
            <given-names>J. P.</given-names>
            <surname>Luck</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P. L.</given-names>
            <surname>McDermott</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Allender</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D. C.</given-names>
            <surname>Russell</surname>
          </string-name>
          ,
          <article-title>An investigation of real world control of robotic assets under communication latency</article-title>
          ,
          <source>in: Proceedings of the 1st ACM SIGCHI/SIGART conference on Human-robot interaction</source>
          ,
          <year>2006</year>
          , pp.
          <fpage>202</fpage>
          -
          <lpage>209</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          [18]
          <string-name>
            <given-names>D. P.</given-names>
            <surname>Miller</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Machulis</surname>
          </string-name>
          ,
          <article-title>Visual aids for lunar rover tele-operation</article-title>
          ,
          <source>in: proceedings of 8th International Symposium on Artificial Intelligence</source>
          , Robotics and Automation in Space, edited by R. Battrick, ESA Publishing, Noordwĳk, Netherlands,
          <year>2005</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          [19]
          <string-name>
            <given-names>C. W.</given-names>
            <surname>Nielsen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. A.</given-names>
            <surname>Goodrich</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R. W.</given-names>
            <surname>Ricks</surname>
          </string-name>
          ,
          <article-title>Ecological interfaces for improving mobile robot teleop-</article-title>
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>