<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Better Faulty than Sorry : Investigating Social Recovery Strategies to Minimize the Impact of Failure in Human-Robot Interaction</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Sara Engelhardt</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Emmeli Hansson</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Iolanda Leite</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Department of Robotics, Perception and Learning KTH Royal Institute of Technology</institution>
        </aff>
      </contrib-group>
      <abstract>
        <p>Failure happens in most social interactions, possibly even more so in interactions between a robot and a human. This paper investigates di erent failure recovery strategies that robots can employ to minimize the negative e ect on people's perception of the robot. A between-subject Wizard-of-Oz experiment with 33 participants was conducted in a scenario where a robot and a human play a collaborative game. The interaction was mainly speech-based and controlled failures were introduced at speci c moments. Three types of recovery strategies were investigated, one in each experimental condition: ignore (the robot ignores that a failure has occurred and moves on with the task), apology (the robot apologizes for failing and moves on) and problem-solving (the robot tries to solve the problem with the help of the human). Our results show that the apology-based strategy scored the lowest on measures such as likeability and perceived intelligence, and that the ignore strategy lead to better perceptions of perceived intelligence and animacy than the employed recovery strategies.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>
        Social interactions are not always successful, but humans have developed social
norms to deal with such cases. Sunstein de ned social norms as \social
attitudes of approval and disapproval, specifying what ought to be done and what
ought not to be done"[
        <xref ref-type="bibr" rid="ref6">6</xref>
        ]. Considering that even interactions between humans
fail at times, it is not surprising that Human-Robot interaction (HRI) might
inevitably fail too, especially because of some robot malfunction. These failures
can be critical because they might require costly human intervention and, more
importantly, they can cause users to lose trust and interest in the robot.
      </p>
      <p>
        Giuliani et al. found two types of failures in HRI: technical failures and social
norm violations [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ]. While technical failures are often a result of execution
failures (i.e., an appropriate action that was carried out incorrectly), social norm
violations are de ned as \a deviation from the social script or the usage of the
wrong social signals" [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ]. Social norm violations often occur due to planning
failures or actions that are executed correctly but are inappropriate to the
situation. An example of a planning failure would be the robot asking the user
the same question several times even though an appropriate answer has been
given. Inappropriate social signals can occur, for example, when the robot is not
looking at the person it is talking to [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ].
      </p>
      <p>This paper will investigate the impact of social norm failures committed by
a robot and how that impacts people's perception of the robot. We limit the
scope of our study to verbal failures during human-robot conversation because
we anticipate that speech recognition failures will be one of the main causes of
disruption of the natural course of the interaction once social robots are deployed
in real world environments. In a between-subject pilot study, we investigated
people's perceptions of a robot that employed one of out three types of failure
recovery strategies (ignore, apology or problem-solving).</p>
    </sec>
    <sec id="sec-2">
      <title>Related Work</title>
      <sec id="sec-2-1">
        <title>Human Perception of Robot Failure</title>
        <p>
          Earlier work has studied how a robot is perceived by the user in failure situations.
In a study by Lee et al. [
          <xref ref-type="bibr" rid="ref5">5</xref>
          ], where the interaction consisted of a human asking
the robot to get a drink and receiving either the right or wrong one, it was found
that robot failure decreased all ratings of the robot compared to the successful
interaction, except how much they liked the robot. On the other hand, Bajones,
Weiss, and Vincze [
          <xref ref-type="bibr" rid="ref1">1</xref>
          ] found a tendency for only a small negative impact on
perceived intelligence, likability and robot contribution when a robot malfunctions,
in their study about how to mitigate robot failure with the help of the user. They
attributed that to the robot's recovery strategies that made it able to ful ll its
tasks in the end. The importance of task success is also found in a study by
Foster et al. [
          <xref ref-type="bibr" rid="ref3">3</xref>
          ], where a robot is used as a bartender. They found that dialogue
e ciency and task success had the biggest impact on the subjective measures,
as well as perceived intelligence and likeability, which showed a generally
positive outcome [
          <xref ref-type="bibr" rid="ref3">3</xref>
          ]. However, Bajones, Weiss, and Vincze [
          <xref ref-type="bibr" rid="ref1">1</xref>
          ] noted that repeated
demands for help became an annoyance. Torrey, Fussell, and Kiesler [
          <xref ref-type="bibr" rid="ref7">7</xref>
          ] studied
how a robot is perceived when o ering advice. Using hedges and discourse
markers improved how considerate, controlling and likable the participant perceived
the robot. Their results also indicate that robots using politeness might have an
even bigger positive e ect on the interaction than humans doing the same. Lee
et al. [
          <xref ref-type="bibr" rid="ref5">5</xref>
          ] also concluded that the importance of politeness is apparent (politeness
ratings of the robot increased with every recovery strategy used).
2.2
        </p>
      </sec>
      <sec id="sec-2-2">
        <title>Failure Recovery Strategies in HRI</title>
        <p>
          When a robot fails, just like when a human fails, there are di erent ways to
recover from that failure. Earlier studies have tested di erent recovery strategies,
for example, the robot stating what the problem is and how the participants
could help x it when it malfunctions [
          <xref ref-type="bibr" rid="ref1">1</xref>
          ]. This recovery strategy has a
problemsolving approach, where the robot asks the users for help when necessary. The
aforementioned study by Lee et al. [
          <xref ref-type="bibr" rid="ref5">5</xref>
          ] compared four di erent recovery
strategies: forewarning, where the robot warns that it might fail at the start of the
interaction; apology, where the robot apologizes for the failure; compensation,
where the robot o ers some kind of compensation for the failure; and option,
where the robot suggests di erent ways to try to solve the failure. In their
control condition, the robot's only response to the failure was to say \OK". Overall,
the apology strategy scored best for them. However, people with low relational
or high utilitarian orientation liked compensation best, and actually preferred
no recovery strategy over both apology and options, suggesting that recovery
strategies need to be tailored to a persons orientation to services, and di erent
scenarios might bene t from di erent strategies [
          <xref ref-type="bibr" rid="ref5">5</xref>
          ]. It is important to note that
the ndings by Lee et al. were obtained in the context of service robots. It
remains unknown whether similar results will hold for other types of human-robot
interaction, like the collaborative scenario described in this paper.
3
        </p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>Method</title>
      <p>The research question we seek to answer with this pilot study is the following:
which is the optimal strategy that robots can use to minimize the negative impact
of failure in social collaborative interactions with a human?
3.1</p>
      <sec id="sec-3-1">
        <title>Scenario</title>
        <p>We investigated social interaction failure in a collaborative task between a person
and a humanoid robot Nao1 as displayed in Figure 1. Twelve cards were placed
on a table, facing down, in spots labeled from A to L. The goal of the task is to
nd all the four queen cards, the robot already having the rst one. Because the
robot is not capable of turning the cards itself, it asks for the human's help to
turn certain cards (one at a time, using the labels) and say which card is where,
in order to nd the hidden queens.
1 https://www.ald.softbankrobotics.com/en/cool-robots/nao</p>
        <p>Fig. 1. The experiment setup with the robot, Nao, and a participant.</p>
        <p>The failure consisted of the robot interpreting human speech incorrectly when
asked about which card is in a speci c location. To make it clear that the robot
fails, the robot always repeats the card it just heard and asks the participant to
con rm or deny. When the participant answers \no" to this question, di erent
recovery strategies will be used depending on the experimental conditions. Each
recovery strategy will have its own protocol of how to handle and recover from
the failure (more details in Section 3.2).</p>
        <p>Since the main focus of the study was to investigate failures, making sure
the failures were controlled and the same for every experiment was important.
Therefore, a Wizard-of-Oz-technique was used for the verbal aspects of the
interaction because the speech recognition for the robot is not reliable enough yet.
The non-verbal aspects, such as face tracking and gestures of the robot were
fully autonomous and not synchronized to the game.
3.2</p>
      </sec>
      <sec id="sec-3-2">
        <title>Conditions</title>
        <p>
          The three strategies for mitigating the negative impact of failure in social
contexts that we want to study are based on the strategies used previously in the
work of Lee et al. [
          <xref ref-type="bibr" rid="ref5">5</xref>
          ]:
- Ignore: the robot ignores that it has failed and simply keeps going by saying
\OK". This can be considered our control condition.
- Apology: the robot apologizes (for example, by saying \I'm sorry, sometimes
        </p>
        <p>I don't interpret speech correctly") for its failure and then moves on.
- Problem-solving: the robot tries to solve its failure with the help of the
human by asking him/her to repeat the card. After hearing the card one more
time, the robot acknowledges that it understood which card the participant
is referring to.
3.3</p>
      </sec>
      <sec id="sec-3-3">
        <title>Procedure</title>
        <p>Participants were guided to a room by an experimenter and instructed to sit at
a table in front of the robot. The robot then provided the instructions for the
game ( nd the remaining queens). When 9 of the 12 cards were ipped the game
would nish. The outcome of the game would always be the same, for example, all
queens were found. Also, to make sure the outcome of the game was not a ected,
all failures occurred in cards other than queens. Each participant experienced
three failures, three queens found and heard correctly and three other cards that
were heard correctly. For all conditions, failures occurred on cards number 3, 6
and 7, while the other 6 plays were normal interactions (i.e., no failure).</p>
        <p>The average length of the interaction with the robot across all conditions
was about 3.0 minutes, with the average for the ignore condition slightly shorter
(M=2:49, SD=15sec) and for problem-solving slightly longer (M=3:13, SD=32sec).
The apology condition had the same average as the overall time (M=3:01,
SD=17sec). At the end of the interaction, the robot thanked the participant
for playing, and the participant was asked by the experimenter to ll in a
questionnaire.</p>
      </sec>
      <sec id="sec-3-4">
        <title>Measures</title>
        <p>At the end of the interaction, participants lled in a survey consisting of two
parts: some general questions about the participant and perceptual measures
commonly used in HRI experiments. The general questions included participant's
age, gender, occupation, as well as experience with programming and experience
with robots on a ve-point Likert scale for Godspeed and nine-point for RoSAS.</p>
        <p>
          The perceptual measures were taken from the Godspeed Questionnaire [
          <xref ref-type="bibr" rid="ref8">8</xref>
          ]
and the The Robotics Social Attributes Scale (RoSAS) [
          <xref ref-type="bibr" rid="ref2">2</xref>
          ]. The Godspeed
Questionnaire Series is an established way to measure peoples perception of robots.
There are ve measures in total: Anthropomorphism, Animacy, Likeability,
Perceived Intelligence and Perceived Safety. We chose to only use Animacy,
Likeability and Perceived Intelligence since we believe these are the most
relevant measures to our study. The Robotics Social Attributes Scale (RoSAS) builds
upon the Godspeed Questionnaire Series, but seeks to improve the cohesiveness
of the measurements [
          <xref ref-type="bibr" rid="ref2">2</xref>
          ]. We used Competence and Discomfort, but chose to
exclude the warmth metric because it seemed less relevant to this study. All the
di erent items of each measure were randomized in the survey.
3.5
        </p>
      </sec>
      <sec id="sec-3-5">
        <title>Participants</title>
        <p>We recruited 33 adults participants for our study (11 participants per condition).
All participants were computer science undergraduate students from the same
university in Sweden and their median age was 22 years old. There is a majority
of male students in computer science programs, which was re ected in the gender
distribution of our participants. Participants were randomly distributed between
conditions. In the ignore condition we had 64% males (7) and 36% females
(4); in the apology condition we had 82% males (9) and 18% other (2), and
in the problem-solving condition, we had 73% males (8) and 27% females (3).
Participants reported that their previous experience with robots was low, with
an average of 1.8 on a scale from 1 to 5.
4
4.1</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>Results</title>
      <sec id="sec-4-1">
        <title>Likeability</title>
        <p>Generally the participants in all three conditions rated the robot high on
likeability. Participants in the apology condition rated the robot slightly lower
(M = 4:3; SD = 0:4) than in the ignore (M = 4:5; SD = 0:7) and
problemsolving condition (M = 4:5; SD = 0:2).
4.2</p>
      </sec>
      <sec id="sec-4-2">
        <title>Perceived Intelligence</title>
        <p>In general, the robot was perceived as fairly competent and intelligent, with
average scores mostly between 3 and 4. The ignore condition scored highest
(M = 3:7; SD = 0:6), followed by problem-solving (M = 3:4; SD = 0:4), and
apology (M = 3:1; SD = 0:6).
Overall, the robot was seen as fairly animated, with scores slightly above the
middle of the scale. The ignore condition generated a slightly higher rating (M =
3:4; SD = 0:7) than apology (M = 3:2; SD = 0:6) and problem-solving (M =
3:2; SD = 0:4).
4.4</p>
      </sec>
      <sec id="sec-4-3">
        <title>Competence</title>
        <p>Generally, the apology condition (M = 5:3; SD = 1:3) rated the robot lower on
perceived competence, with the ignore condition (M = 6:2; SD = 0:9) scoring
slightly higher than problem-solving (M = 6:0; SD = 1:2).
4.5</p>
      </sec>
      <sec id="sec-4-4">
        <title>Discomfort</title>
        <p>The robot scored low on discomfort, with the averages for all three conditions
below 3 on the 9-point scale. Ignore (M = 3:0; SD = 1:5) scored slightly higher
than apology (M = 2:7; SD = 0:9) and problem solving (M = 2:6; SD = 0:9).
The low scores on discomfort and high scores on robot likeability can be
explained by several factors such as the robot's appearance, the non-critical nature
of the task and also the participants' background (computer science students).
An interest in technology, experience with programming and some experience
with robots could all reduce the feeling of discomfort in the presence of a robot,
as well as increase the interest for, and liking of, robots.</p>
        <p>
          An important aspect of all the participants in our study is that their
education focuses heavily on problem-solving, which might attract students with a
problem-solving mindset. This could be the reason why the participants in our
study preferred the problem-solving strategy over the apology strategy. As Lee
et al. [
          <xref ref-type="bibr" rid="ref5">5</xref>
          ] found, participant individual traits can a ect which recovery strategy
is preferred.
        </p>
        <p>
          The ignore condition was designed to be less responsive and not intended to
recover from failure. In the study by Lee et al. [
          <xref ref-type="bibr" rid="ref5">5</xref>
          ], a similar behavior resulted in
a lower score for perceived competence compared to problem-solving and
apologizing strategies. However, in our study, the apology condition scored lowest on
intelligence and competence scores, while the ignore condition scored highest on
responsiveness. This suggests that the ignore condition was not perceived as
expected. We believe this might have happened due to an unexpected pattern in the
participants behavior that emerged during the experiments; many participants
instinctively repeated the card that the robot misheard. The ignore condition
could be seen as acknowledging this when it says \OK", while the other
conditions clearly ignored it. Not only does this make the robot seem more responsive
in the ignore condition, but it can be perceived as the robot having corrected
the mistake, similar to what happened in the problem-solving condition. That
would leave apology the only one that does not correct the mistake, which could
a ect its ratings negatively.
        </p>
        <p>In conclusion, based on previous research, we expected that the apology
strategy would do best, followed by problem-solving and then ignore. However,
our results show that the apology condition resulted in less positive perceptions
of the robot in most of our measures of interest, and that the ignore condition
did as well as, or better than, problem-solving in many cases.</p>
      </sec>
      <sec id="sec-4-5">
        <title>Limitations and Future Work</title>
        <p>Our participant pool was a very homogeneous group: all participants were
between 20 and 27 years old and studied computer science at the same university.
The rather small sample size entails that wizard or participant errors might
have a larger impact on the results, but the homogeneousness of the group is an
advantage when drawing conclusions about this particular age group.</p>
        <p>There is still a lot of research to be done in this area. While the results of this
pilot study were quite informative, a large sample size would enable us to apply
statistical analysis to determine whether the trends we found are statistically
signi cant. In the future, we should also consider including a condition where
the robot does not fail at all, to see how much the failure itself in uences the
participants perception of the robot.</p>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Bajones</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Weiss</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Vincze</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          :
          <article-title>Help, anyone? a user study for modeling robotic behavior to mitigate malfunctions with the help of the user (</article-title>
          <year>2016</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Carpinella</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Wyman</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Perez</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Stroessner</surname>
            ,
            <given-names>S.:</given-names>
          </string-name>
          <article-title>The robotic social attributes scale (rosas): Development and validation</article-title>
          .
          <source>In: Proceedings of the 2017 ACM/IEEE International Conference on human-robot interaction</source>
          . pp.
          <volume>254</volume>
          {
          <fpage>262</fpage>
          . HRI '17,
          <string-name>
            <surname>ACM</surname>
          </string-name>
          (March
          <year>2017</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Foster</surname>
            ,
            <given-names>M.E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gaschler</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Giuliani</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Isard</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Pateraki</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Petrick</surname>
          </string-name>
          , R.:
          <article-title>Two people walk into a bar: dynamic multi-party social interaction with a robot agent</article-title>
          .
          <source>In: Proceedings of the 14th ACM international conference on multimodal interaction</source>
          . pp.
          <volume>3</volume>
          {
          <fpage>10</fpage>
          . ICMI '12,
          <string-name>
            <surname>ACM</surname>
          </string-name>
          (
          <year>October 2012</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Giuliani</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Mirnig</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Stollnberger</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Stadler</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Buchner</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Tscheligi</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          :
          <article-title>Systematic analysis of video data from di erent human-robot interaction studies: a categorization of social signals during error situations</article-title>
          .
          <source>Frontiers in psychology 6</source>
          (
          <year>2015</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Lee</surname>
            ,
            <given-names>M.K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kielser</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Forlizzi</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Srinivasa</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rybski</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          :
          <article-title>Gracefully mitigating breakdowns in robotic services</article-title>
          .
          <source>In: Proceedings of the 5th ACM/IEEE international conference on human-robot interaction</source>
          . pp.
          <volume>203</volume>
          {
          <fpage>210</fpage>
          . HRI '10, IEEE Press (
          <year>2010</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Sunstein</surname>
            ,
            <given-names>C.R.</given-names>
          </string-name>
          :
          <article-title>Social norms and social roles</article-title>
          .
          <source>Columbia Law Review</source>
          <volume>96</volume>
          (
          <issue>4</issue>
          ),
          <volume>903</volume>
          {
          <fpage>968</fpage>
          (
          <year>1996</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>Torrey</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Fussell</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kiesler</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          :
          <article-title>How a robot should give advice</article-title>
          .
          <source>In: Proceedings of the 8th ACM/IEEE international conference on human-robot interaction</source>
          . pp.
          <volume>275</volume>
          {
          <fpage>282</fpage>
          . HRI '13, IEEE Press (
          <year>2013</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>Weiss</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bartneck</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          :
          <article-title>Meta analysis of the usage of the godspeed questionnaire series</article-title>
          . pp.
          <volume>381</volume>
          {
          <fpage>388</fpage>
          .
          <string-name>
            <surname>IEEE</surname>
          </string-name>
          (
          <year>August 2015</year>
          )
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>