<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Dimensions Impact Automation Preferences with a Conversational Task Assistant</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Jessica He</string-name>
          <email>jessicahe@ibm.com</email>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>David Piorkowski</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Michael Muller</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Kristina Brimijoin</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Stephanie Houde</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Justin D. Weisz</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>IBM Research</string-name>
          <email>djp@ibm.com</email>
          <email>kbrimij@us.ibm.com</email>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Seattle WA US</string-name>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>IBM Research</institution>
          ,
          <addr-line>Cambridge MA</addr-line>
          <country country="US">US</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>IBM Research</institution>
          ,
          <addr-line>Yorktown Heights NY</addr-line>
          <country country="US">US</country>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>ing Automation Experiences</institution>
        </aff>
      </contrib-group>
      <abstract>
        <p>Organizations have recently begun to deploy conversational task assistants that collaborate with business users to partially automate their work tasks. These assistants are becoming more intelligent: users initiate automated task support through natural language, and the system can dynamically orchestrate new task sequences accordingly. As these tools become more intelligent and automated, they sometimes shift control away from users to increase process eficiency at the cost of consequences for users' preferences and productivity. Particularly in high stakes work environments, this shift raises questions of when automation is suitable or unsuitable and how to delegate agency such that users feel suficiently in control of their tasks. We explored these questions through two studies comprised of interviews and co-design activities with business users and identified four task dimensions along which their automation and interaction preferences vary: process consequence, social consequence, task familiarity, and task complexity. These dimensions are useful for understanding when, why, and how to delegate control between users and conversational task assistants.</p>
      </abstract>
      <kwd-group>
        <kwd>Conversational Task</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>
        1. Introduction
Recent innovations in AI automation have enabled the
adoption of AI assistants in partially automating business
workflows for knowledge workers. Partial automation in
high stakes work environments raises questions of when
automation is suitable and unsuitable, and accordingly,
when users should retain task control and when they can
delegate it to the AI. In this paper, we explore these
questions in the context of conversational task assistants that
interact with business users through a natural language
tems [
        <xref ref-type="bibr" rid="ref12">1, 2, 3</xref>
        ]. One example of such a system is Watson
      </p>
    </sec>
    <sec id="sec-2">
      <title>Orchestrate, which automates repetitive tasks for business users in a variety of domains [4].</title>
      <p>Such systems are built on multi-agent orchestration
technologies, where a front-end dialogue-manager
transforms natural language utterances (what the user says or</p>
    </sec>
    <sec id="sec-3">
      <title>Wiberg and Bergqvist [6] posit that automation poses</title>
      <p>earlier work [7, 8] in allocation of function to humans and</p>
    </sec>
    <sec id="sec-4">
      <title>AIs through their Engaging Interaction through Automa</title>
      <p>tion scale, which outlines a spectrum from ”no
automation of interaction” to “full automation of interaction.”</p>
    </sec>
    <sec id="sec-5">
      <title>Other prior workshop papers have discussed the notion</title>
      <p>that within a workflow, only certain steps or sub-tasks
may be suitable for automation [9, 10].</p>
      <p>Given the wide variety of tasks that conversational
task assistants can support and the broad spectrum of
users interacting with them, we sought to explore factors
that inform a suitable division of automation. In line with
principles of human-centered AI, we explored this from
the perspective of human preferences. Through two
studies comprised of user interviews and co-design activities,
we identify four dimensions of tasks that can help
designers of conversational task assistants determine when
trol. These dimensions are: consequences of task errors
© 2023 Copyright for this paper by its authors. Use permitted under Creative Commons License to automate interactions vs. when to provide user
con(referred to as process consequence), social consequence, developers, researchers, designers, managers, and
salesusers’ familiarity with a task, and task complexity. people) took part in Study 1. 15 new participants, also</p>
      <p>Researchers have previously characterized task sup- with varied job functions (including design, research, and
port in terms of multiple task dimensions based on clas- business), took part in Study 2. Both studies were
comsifications and trade-of dimensions. Tasks have been prised of one-hour sessions with one researcher and one
analyzed along dimensions of retrospective vs. prospec- participant. Participants were compensated the
equivative [11], informative vs. actionable [12], reminding vs. lent of $25 USD. We refer to participants from Study 1 as
being-reminded [13], visible vs. invisible [14, 15, 16], P1xx, and participants from Study 2 as P2xx.
content-oriented vs. relationship-oriented [17, 13, 18], In the first study, sessions were comprised of an
interand holistic [19] vs. itemized [20, 21, 22, 23]. view and a co-design analysis aimed at understanding</p>
      <p>In this paper, we build on this work through two stud- current task practices and eliciting considerations
releies that help define requirements for designing an AI vant to the design of a conversational task assistant. For
task assistant. We show the importance of four new task the co-design activity, we used a more structured version
dimensions based on human-AI interaction that are use- of the CARD method [25], adapted for remote
particiful for understanding when, why, and how to delegate pation via Mural (Figure 1a). Participants were asked
control between users and conversational task assistants. to identify a work task suitable for partial automation
(with input from the researcher). They mapped out their
current process and pain points for this task in Mural,
2. Methods then discussed the level of automation they would be
comfortable with for each step and why, along with any
information and tooling needs they envisioned.</p>
      <p>In the second study, we extended our inquiry into how
automation preferences should be incorporated into the
design of a conversational task assistant by conducting
a participatory design study using an online version of
paper-prototyping [26, 27]. We began each session by
introducing the conversational task assistant’s capabilities
and limitations and showing a brief video demo.
Participants were provided with a blank, low-fidelity version of</p>
    </sec>
    <sec id="sec-6">
      <title>We conducted two interview studies. Adapting a “bifo</title>
      <p>cal” approach [24], the first study was a high-level task
analysis, and the second study was a more detailed
participatory design based on the task analysis. Participants for
both studies were recruited from several company-wide
Slack channels for business users of a large international
technology company. Through a brief survey that asked
them to describe their day-to-day job, we screened for
people who had experience with workflows suitable for
automation. 13 participants with varying roles (including
• Process consequence: The user’s perceived cost of</p>
      <p>failure when the assistant makes a mistake.
• Social consequence: The user’s perceived risks in
al</p>
      <p>lowing an assistant to represent them to others.
• Familiarity: The user’s knowledge of the system</p>
      <p>and/or the task.
• Complexity: The overall dificulty or efort required</p>
      <p>for a user to complete a task.
the assistant’s user interface, along with UI components users understand the system’s behavior and give them
that they could drag and drop and tools to design their the control to mediate potential risks.
own (Figure 1b). Using these components, they were Social consequence. Some participants chose tasks
asked to co-design a series of interactions to support the involving others, such as teammates or clients, as
canpartial automation of two tasks: scheduling a meeting didates for automation. For example, P212 wanted the
and a task of their choice. This activity served as the assistant to “email everyone on [a]... project,” and P104
basis for discussions on participants’ automation needs. wanted to schedule customer calls. Regarding
automa</p>
      <p>We collected 28 hours of video interview data. Us- tion preferences, P104 emphasized that initial contact
ing participants’ responses and video transcriptions, we with a new customer should be handled by a human “to
conducted a thematic analysis [28] to analyze the data establish the... relationship with the customer.” Similarly,
for factors that afected participants’ automation pref- P111 felt that customer follow-ups should be
humanerences. We describe these factors as characteristics of driven, and P106 wanted “to be able to control how widely
tasks, which formed the basis of what we termed task di- communication [goes].” These observations indicate that
mensions. We then analyzed how users preferred to work social consequence is another task dimension along which
with the assistant in the context of these dimensions to information needs and desired control vary.
understand design implications for task assistants. The Where interpersonal skills and emotional
sensitivfollowing section presents results across both studies. ity are required, participants felt that the task assistant
lacked the emotional intelligence to handle the task
independently, an insight consistent with Gofman’s
empha3. Task dimensions sis that impression-management is important to people
We identified four dimensions of tasks along which au- in their organizational lives [30]. These findings suggest
tomation preferences varied: that automation involving other people should largely
remain under user control. Some participants did
consider certain tasks less consequential despite involving
other people (e.g. creating an HR ticket). Such tasks may
be suitable for automation but should first request user
approval given the diversity of concerns.</p>
      <p>Familiarity. Participants also varied by familiarity,
either with a task or with the system. Both types were
determinants of how much automation, transparency,
and guidance participants wanted from the system.</p>
      <p>Non-experts of a task asked for transparency into
sys</p>
      <p>Process consequence. When asked for tasks to au- tem actions and guidance on the task process. Task
fatomate, participants described a mix of tasks with mini- miliarity is dynamic—people’s inactivity in a particular
mal and significant consequences of error. Our studies domain may transform them from expert to non-expert,
revealed several instances where significant process con- and unfamiliar tasks can also become more familiar over
sequences made users hesitant to automate a task or step. time. Participants expected the system to recognize the
For example, P109 described an onboarding task that was latter and reduce its support accordingly. After P212’s
high-consequence due to the sensitive information re- first few expense approvals, they said they would no
quested of them. They wanted to know how the system longer require step-by-step guidance on required inputs
would use their information: “I would want to know if and instead would initiate the conversation with
“Subthat information is secure, if it’s going anywhere after that, mit an expense report for &lt;event&gt;. Here are the dates and
or if they just delete everything.” Thus, we find that the locations. Here’s a folder of the receipts.”
process consequences associated with a task is a dimension Similarly, transparency can orient non-experts of a
that afects users’ information and agency needs. system to its capabilities. Several participants expressed</p>
      <p>These findings imply that only low-consequence tasks an initial distrust of the system. For instance, P107 spoke
would be considered for automation, but further prob- of submitting an expense report—a task that could have
ing revealed that additional user oversight and control ifnancial ramifications if done incorrectly: “if I hadn’t
gotmay help participants automate work without substan- ten that trust yet, then I’d probably ask the system, ‘prepare
tial risks. Among these controls were abilities to preview the expense report for my review’ rather than letting the
to-be-automated actions, verify outputs, and see con- system submit on its own.” P105 also commented that to
sequential steps of an upcoming task. These granular automate work with the assistant, they wanted to “watch
insights would increase user comfort and control of the it first” to calibrate trust (similar to [31, 32]). A review
automation (similar to Park et al. [29]). Such interac- step that ofers users insight into the automation and
tions can draw attention to the risks within a task to help allows them to verify outputs supports non-experts in
understanding a system’s capabilities. view and confirmation steps that provide users with</p>
      <p>Complexity. In the co-design activities of Study 2, decision-making authority, and richer interaction
modalparticipants adapted their designs to reflect the complex- ities. These afordances support user eficiency and
comity of tasks and steps. Participants showed that they pre- fort by transforming rigid workflows into human-AI
colferred chat interactions for simple workflows, but they laborative processes with users in control. Thus, far from
preferred traditional graphical user interfaces (GUIs) for the rigid and predetermined control structures of the
almore complex, information-rich workflows. One reason location models [36, 7, 8], we propose that participants
for this diference was the perceived richness of these positioned themselves as co-creators of their task
modalities. Namely, natural language-based interactions automation experience.
were more eficient when participants had small amounts Beyond these afordances, participants also expected
of information to convey to the assistant. In contrast, for task assistant interactions to be adaptive to the their
tasks that they considered more complex, the richer inter- background and needs and highly personalized to their
actions of GUIs aforded participants with more control unique ways of working. Becoming more eficient in
(e.g. direct manipulation of a list of files). the parts of their jobs that they personally cared about</p>
      <p>Participants suggested that the additional details and was the primary motivator that participants cited for
user control provided by a richer interface could make working with a conversational task assistant. Identifying
automation of complex tasks more desirable. For a meet- such priorities through user studies with specific
working scheduling task that was more complex than usual, groups can help prioritize development of task assistance
P203 said, “If the meeting is with more people... proba- that focuses on functionality that matters most.
bly I would prefer another kind of interaction, maybe a Personalization plays a larger role as AI assistants are
traditional one where I can see the schedule of the people.” adopted in more contexts with broader user groups. We
Hence, we saw that completing complex work with a task saw that users expected to control how they interact
assistant requires interactions beyond a chat interface with the assistant on a per-task basis both in terms of the
for participants to feel comfortable with automation. interaction and how the task was represented. We
expect that as such assistants become more intelligent and
widespread, the need for personalization will similarly
4. Discussion increase and the design of these assistants should take
this into account.</p>
      <p>Although some models propose that humans and AI
assistants may be equally capable of doing certain tasks
[33, 34, 35], current task assistant architectures are based 5. Conclusion
on more asymmetric principles in which the AI
assistant retains execution capability for many operations, We conducted two studies to understand business users’
as described in the allocation models [36, 7, 8]. Partici- automation preferences and needs for working
comfortpants’ comments and designs encourage us to re-examine ably and eficiently with a conversational task assistant.
whether and when the assistant should retain control. The This work identified important user considerations in
four task dimensions provide guidance on when to del- the context of task dimensions for the design of
converegate control to whom in automation, and the features sational task assistants:
described by participants in the context of these
dimensions are examples of how to delegate it. • We identified four task dimensions along which
automa</p>
      <p>Participants’ concerns about possible process and so- tion preferences varied: process consequence, social
cial consequences, along with concerns about working consequence, familiarity, and complexity. These
on complex tasks using natural language, made them task dimensions provide a human-centered
perspecwary of automation. Low familiarity with the assistant tive into when, why, and how to delegate control
also raised concerns about how much they could trust between system and user.
the assistant to automate consequential tasks. These task • Along these dimensions, we elicited several afordances
dimensions can provide insight on where along Wiberg to put users more in control: transparency into system
and Bergqvist’s [6] Engaging Interaction through Au- actions, decision-making authority on if and how a step
tomation scale to design, and address questions of what or task gets done, and richer interaction modalities.
kinds of tasks or steps are suitable for automation [9, 10]. • The dynamic nature of these dimensions means that</p>
      <p>Despite these concerns, participants described sev- users’ automation preferences will change over time.
eral afordances across dimensions that could provide Therefore, conversational task assistants should allow
them with more control and hence help them feel more users to situationally adapt their task controls for a
comfortable with automation: transparency into the comfortable and eficient automation experience.
task assistant’s capabilities and step-by-step process,
re</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          <source>in computing systems</source>
          ,
          <year>1992</year>
          , pp.
          <fpage>455</fpage>
          -
          <lpage>462</lpage>
          . [27]
          <string-name>
            <given-names>B.</given-names>
            <surname>Still</surname>
          </string-name>
          ,
          <string-name>
            <surname>J. Morris,</surname>
          </string-name>
          <article-title>The blank-page technique: Rein-</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          <volume>53</volume>
          (
          <year>2010</year>
          )
          <fpage>144</fpage>
          -
          <lpage>157</lpage>
          . [28]
          <string-name>
            <given-names>G.</given-names>
            <surname>Terry</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Hayfield</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Clarke</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Braun</surname>
          </string-name>
          , The-
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          <source>research in psychology 2</source>
          (
          <year>2017</year>
          )
          <fpage>17</fpage>
          -
          <lpage>37</lpage>
          . [29]
          <string-name>
            <given-names>S.</given-names>
            <surname>Park</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A. X.</given-names>
            <surname>Zhang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L. S.</given-names>
            <surname>Murray</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D. R.</given-names>
            <surname>Karger</surname>
          </string-name>
          ,
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          <article-title>need-finding study</article-title>
          ,
          <source>in: Proceedings of the 2019</source>
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          <string-name>
            <surname>Systems</surname>
          </string-name>
          ,
          <year>2019</year>
          , pp.
          <fpage>1</fpage>
          -
          <lpage>12</lpage>
          . [30]
          <string-name>
            <given-names>E.</given-names>
            <surname>Gofman</surname>
          </string-name>
          , et al.,
          <article-title>The presentation of self in every-</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          <source>day life</source>
          .
          <year>1959</year>
          ,
          <string-name>
            <surname>Garden</surname>
            <given-names>City</given-names>
          </string-name>
          , NY
          <volume>259</volume>
          (
          <year>1959</year>
          /
          <year>2002</year>
          ). [31]
          <string-name>
            <given-names>J. D.</given-names>
            <surname>Weisz</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Muller</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Houde</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Richards</surname>
          </string-name>
          ,
          <string-name>
            <surname>S. I.</surname>
          </string-name>
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          <article-title>code translation</article-title>
          , in: 26th International Conference
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          <source>on Intelligent User Interfaces</source>
          ,
          <year>2021</year>
          , pp.
          <fpage>402</fpage>
          -
          <lpage>412</lpage>
          . [32]
          <string-name>
            <given-names>J.</given-names>
            <surname>Drozdal</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Weisz</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Dass</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Yao</surname>
          </string-name>
          ,
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          <source>ceedings of the 25th International Conference on</source>
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          <string-name>
            <given-names>Intelligent</given-names>
            <surname>User Interfaces</surname>
          </string-name>
          ,
          <year>2020</year>
          , pp.
          <fpage>297</fpage>
          -
          <lpage>307</lpage>
          . [33]
          <string-name>
            <given-names>I.</given-names>
            <surname>Grabe</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>González-Duque</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Risi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Zhu</surname>
          </string-name>
          , To-
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          <article-title>terns in co-creative gan applications</article-title>
          ,
          <source>HAIGEN</source>
          <year>2022</year>
          :
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          3rd Workshop on Human-AI
          <string-name>
            <surname>Co-Creation with</surname>
          </string-name>
          Gen-
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          <string-name>
            <surname>erative Models</surname>
          </string-name>
          (
          <year>2022</year>
          ). [34]
          <string-name>
            <given-names>M.</given-names>
            <surname>Muller</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. D.</given-names>
            <surname>Weisz</surname>
          </string-name>
          , W. Geyer, Mixed initiative
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          <article-title>work for generative ai applications</article-title>
          ,
          <year>2020</year>
          . URL:
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          <fpage>cocreative</fpage>
          -iccc20/papers/Future_of_co-creative_
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          <source>systems_185</source>
          .pdf. [35]
          <string-name>
            <given-names>A.</given-names>
            <surname>Spoto</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Oleynik</surname>
          </string-name>
          , Library of mixed-
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          <source>initiative creative interfaces</source>
          ,
          <year>2017</year>
          . URL:
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          http://mici.codingconduct.cc/. [36]
          <string-name>
            <surname>P. M. Fitts</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          <string-name>
            <surname>Viteles</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          <string-name>
            <surname>Barr</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          <string-name>
            <surname>Brimhall</surname>
          </string-name>
          , G. Finch,
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          <article-title>and trafic-control system, and appendixes 1 thru 3</article-title>
          ,
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          <source>dation Columbus</source>
          ,
          <year>1951</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>