<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Gestural Inputs as Control Interaction for Generative Human-AI Co-Creation</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>John Joon Young Chung</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Minsuk Chang</string-name>
          <email>minsuk.chang@navercorp.com</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Eytan Adar</string-name>
          <email>eadar@umich.edu</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="editor">
          <string-name>Helsinki, Finland</string-name>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>CEUR Workshop Proceedings</institution>
          ,
          <addr-line>CEUR-WS.org</addr-line>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Naver AI Lab</institution>
          ,
          <addr-line>Seongnam</addr-line>
          ,
          <country>Republic of Korea</country>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>University of Michigan</institution>
          ,
          <addr-line>Ann Arbor, MI</addr-line>
          ,
          <country country="US">USA</country>
        </aff>
        <aff id="aff3">
          <label>3</label>
          <institution>Workshop Proce dings</institution>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2022</year>
      </pub-date>
      <abstract>
        <p>While AI-powered generative systems ofer new avenues for art-making, directing these algorithms remains a central challenge. Current methods for steering have focused on conventional interaction techniques (widgets, examples, etc.). This position paper argues that the intersection of user needs in creative contexts and algorithmic capabilities requires re-thinking our interactions with generative AI. We propose that rough gestural inputs, such as hand gestures or sketching, can enhance the experience of human-AI co-creation-even for text. First, the undetermined and ambiguous nature of gestural inputs corresponds to the purpose and the capabilities of generative systems. Second, rough gestural can be intuitive and expressive, facilitating iterative co-creation. We discuss design dimensions for inputs of artifact-creating systems, then characterize existing and proposed input interactions with those dimensions. We highlight how gestural inputs can expand the control interaction for generative systems by analyzing existing tools and describing speculative input designs. Our hope is that gestural inputs become actively studied and adopted to support user intentions and maximize the perceived eficacy of and preference. Most commonly, produced text can sud- tasks. For example, numerical inputs imply an 'exact' to re-run the algorithm and iterate until they get the de- expectation for the user: the system will produce exbest controls go beyond steering the behavior of the al- for other reasons. Creative work often requires iterarough gestural 'sketches' coupled with abstract represen- lenged to find or create new examples or prompts. The generative algorithms.</p>
      </abstract>
      <kwd-group>
        <kwd>generation</kwd>
        <kwd>controllability</kwd>
        <kwd>gestural input</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>
        denly turn from what the author wants. Users often have
1. Introduction
Technologies such as Generative adversarial network
(GAN) [
        <xref ref-type="bibr" rid="ref2">1</xref>
        ], and pretrained language models (PLM) [
        <xref ref-type="bibr" rid="ref3">2</xref>
        ]
have the potential to enable human-AI co-creation. These
algorithms are attractive in creative contexts as the AI can
generate novel creations–something the human hadn’t
considered. However, this process can backfire when
novelty and surprise misalign with the user’s intention
sired results. Without control, users can only hope that
the next generation will be better than the last. Thus,
controllability becomes key to efective iteration. The
gorithm. They also manage the user’s expectations of
what the algorithm will produce. There are many
conventional ways to provide interactive control but without
addressing these goals efectively. This paper proposes that
tations of content (i.e., information visualizations) can
facilitate control interaction for generative algorithms.
      </p>
    </sec>
    <sec id="sec-2">
      <title>Our proposal is strongly motivated by limitations in</title>
      <p>
        https://minsukchang.com/ (M. Chang); http://cond.org/ (E. Adar)
CEUR
htp:/ceur-ws.org
ISN1613-073
existing control interactions for generative algorithms.
Current control interactions range from inputting a a
simple number (e.g., have a violin play with the
maximum amount of vibrato, by setting the parameter value of
1.0 [
        <xref ref-type="bibr" rid="ref4">3</xref>
        ]) to using natural language prompts (e.g., produce
an image of a dragon sitting on a castle [
        <xref ref-type="bibr" rid="ref1 ref5">4</xref>
        ]) to providing
examples (e.g., make this photograph look like this
example from Picasso [
        <xref ref-type="bibr" rid="ref6">5</xref>
        ]). We argue that these approaches
are limited in diferent ways, particularly in co-creative
level of control. This over-promises and sets a very high
actly what was specified. Unfortunately, this does not
often match algorithmic capabilities. Natural language
prompts and example-driven interfaces are problematic
tion and experimentation with alternatives. Prompts and
examples do not readily support this iteration. Users
may not understand why the algorithm did what it did,
how the results can be corrected, or may simply be
chalcost of iterative practice may make generative algorithms
unappealing in practice.
      </p>
      <p>In answer to many challenges for generative tools, we
propose that rough gestural inputs, such as sketching, can
be a sweet spot for human-AI co-creation. First, gestural
input conveys imprecise and ambiguous intentions [6,</p>
    </sec>
    <sec id="sec-3">
      <title>7], which corresponds to the nondeterministic nature of generative algorithms. Second, because impreciseness is allowed (e.g., simple brush strokes [8]), the interaction of specification would be easier. With easier interactions,</title>
      <p>the end-user may not need to think carefully about the and imprecise nature of sketching corresponds to the
examples or prompts they generate, thus allowing for user’s ambiguous intentions and the algorithm’s
uncermore rapid iteration. tainty. This example also demonstrates how one input</p>
      <p>
        The idea of gestural or sketching inputs in the context modality and representation (i.e., visual) can be used to
of generation has some history. For example, low fidelity guide a diferent output modality (i.e., textual).
sketches created by the end-users can guide the genera- In this paper, we expand on this idea. We first
introtion of photorealistic images [
        <xref ref-type="bibr" rid="ref10">9</xref>
        ]. Here the sketch is the duce the design dimensions for inputs of artifact-creation
input and is in the same modality as the output (e.g., take systems. These include the types of support one wants
the visual dragon I scribbled and make a visual photo- with an AI tool, considerations of algorithmic uncertainty,
realistic version). The input indicates “what generation the precision of input, and ease of iteration on algorithms
should be done.” Control is implemented through more and inputs. Using design dimensions, we characterize
standard interactive approaches (e.g., by adjusting this diferent existing input types for generative human-AI
slider, I am indicating how to bias color selection for the co-creation. We specifically discuss how our example
sysdragon). This is not to imply a clear separation between tem, TaleBrush [
        <xref ref-type="bibr" rid="ref11">10</xref>
        ], adopts sketched inputs to facilitate
input and control, as they are often inexorably connected iterative human-AI co-creation in story writing.
Conas mechanisms to have a system produce the desired out- sidering sketching and gestural inputs for control will
put. However, our specific suggestion is that the input enable new ways to support human-AI co-creation.
interactions–and sketching and gesture, in particular–
can be expanded to also control “how the generation
should be done.’’ 2. Design Dimensions of
      </p>
      <p>
        One example of a‘sketch-as-control’ interaction is our Artifact-Creation Support
controllable story generation system, TaleBrush
(Figure 1). Here, TaleBrush leverages abstract visual rep- We first consider possible dimensions for designing
conresentation of the character’s fortune to control the story trol interactions for creation support. We scope “creation
generation. The canvas is a 2D plane that allows for the support” to systems that help creatives directly
implespecification of the protagonist’s fortune (  -axis) and the ment artifacts. We exclude those tools that serve a more
story’s progression ( -axis). In this interface, the control indirect role, such as critiquing the created artifact. This
interaction is as simple as a single stroke of a line. This boundary is something we have previously considered
approach has several benefits. First, it is easier to interact in surveying the range of tools in the creative space [
        <xref ref-type="bibr" rid="ref12">11</xref>
        ].
with than alternatives (e.g., having multiple sliders for We propose a focus on three aspects: 1) type of support,
diferent story parts). Most importantly, the ambiguous 2) algorithm, and 3) input (summarized in Figure 2).
generation, use variants of generative algorithms. In
contrast to the augmentation category, the end-user is ceding
some creative control to the tool. Though, of course, the
human maintains ultimate control over what makes it
into the final artifact.
      </p>
      <p>
        Transfer tools turn one artifact into another. A
common feature is that they receive some ‘original’ artifact
(e.g., a picture, a piece of text, a sketch, etc.) as input. The
tool will then act on this input to generate a variant–often
some alteration of the original input. A wide range of
tools fall into this category, and they are often
modalityspecific. For example, in the visual design/art space, we
see systems that transfer one visual art piece’s style to
another image [
        <xref ref-type="bibr" rid="ref6">5</xref>
        ]. Other tools in the space will transform
rough sketches into photorealistic images [
        <xref ref-type="bibr" rid="ref10">9</xref>
        ]. As with
image-based style transfer, we find similar approaches
for text where the software can transform the written
input to the style of a particular author [
        <xref ref-type="bibr" rid="ref13">12</xref>
        ]. Though
tarFigure 2: Summary of the design dimensions of artifact- get styles are commonly required, not all transfer tools
creation support in relation to designing generative co- need them. In music, for example, there are tools that
creation tools. Orange items (transfer and generation) are two transform some input piece of music by adding efects
sgurepepnor(tnsopnr-odveitdeerdmwiniitshtigce,ndeirreactti,veancdo-icnrdeiarteicotn) taoroelds.eIstiegmnselien- like delay or compression [13].
ments for generative co-creation tools (our focus is on indirect Finally, we observe tools focused on generation. With
inputs). In indirect input design, the requirements for genera- these, the algorithm generates content from incomplete
tive co-creation tools include: 1) easy and fast iteration and inputs or those of a diferent modality. For example, an
2) algorithmic uncertainty matching the user’s expectation. input might be some previous part of the music, story,
These align with gestural inputs, in considering the ease of or some portion of drawings. The algorithm’s purpose
interaction and their ability to express the user’s ambiguous is not to change this initial input, but rather to add to
intentions. Note that the two-dimensional diagram in charac- them. In music and text, these algorithms continue the
teristics of input is drawn based on the qualitative analysis of from the user-provided ‘start’ or ‘infill’ when given some
diferent input approaches. start and end states [14, 15, 16, 17]. In visual arts, we
most often find this type of algorithm in systems that can
ifll empty spaces in an image [ 18, 19]. Note that many
2.1. Type of Support tools sit somewhere between transfer and generation
and may depend on how the underlying task is defined.
      </p>
      <p>For example, we might have a tool that automatically
colors a part of an image. From the perspective of the
whole image, this may be transfer (especially if the input
is some color palette or color model). However, because
we are also generating new colors, we might treat the
colorization task as generative.</p>
      <p>Diferent types of tools will require diferent types of
controls. However, there are similarities in user needs
and expectations (e.g., surprise and novelty but also a
willingness to cede some creative control to the software).</p>
      <p>This is in contrast to non-creative applications (e.g.,
predictive form filling) where ambiguity and surprise and
undesirable. As we argue below, gestural and sketched
inputs hold promise here.</p>
    </sec>
    <sec id="sec-4">
      <title>CSTs are an extremely diverse and broad category, even</title>
      <p>
        when restricted to direct influence [
        <xref ref-type="bibr" rid="ref12">11</xref>
        ]. Within this
category we see tools that can augment, transfer, or generate.
Though the last two categories are most relevant to our
proposal, augmentation is also worth considering.
      </p>
      <p>CSTs that provide augmentation support often enhance
a task the creative is already doing through
computational means. Many direct manipulation tools fall into
this category. The least ‘intelligent’ of these replicate
existing tools in a digital format. For example, a
digital painting canvas has various types of digital brushes.
Other augmentation tools provide some limited
automation. For example, a bucket tool will flood-fill a closed
area in a sketch. Most augmentation tools are highly
deterministic. They are “predictable” and more naturally
correspond to the user’s mental model of what the
system will do. When using augmentation-focused tools,
the user is firmly in control over both the idea and style
of the final artifact.</p>
      <p>The second and third types of support, transfer and
2.2. Algorithm</p>
      <sec id="sec-4-1">
        <title>2.2.1. Type: Algorithmic Uncertainty</title>
        <p>the artifact are iterated on. For example, with the direct
manipulation of the box, only the position is changing,
but not the color or size. Style transfer algorithms [21] are
at the other extreme. With every run of these algorithms,
the entire image (or many parts of it) will change.</p>
        <p>Iteration naturally connects back to algorithmic
uncertainty. If the user better understands the algorithm’s
behavior (e.g., what the transfer algorithm changes and
how), iteration may become easier. With high
uncertainty, the user may need to iterate many times to get
the efect or artifact they want.</p>
        <p>A tool’s algorithmic pipeline can difer depending on
how certain we are of the pipeline’s output. Deterministic
algorithms are one extreme in that users can predict the
result when using these algorithms. Direct manipulation
implementations are, naturally, one example. When a
box is dragged with a mouse cursor, the end-user knows
where it will end up. Automated algorithms with clear
rules are also deterministic. For example, with
floodifll (e.g., a bucket tool), the user knows the system will
ifll closed areas. If something goes wrong, the user can 2.3. Input
quickly isolate the problem.</p>
        <p>
          On the other extreme are non-deterministic algorithms 2.3.1. Type: Input Directness
which represent many machine learning (ML) algorithms. An input method targets the artifact directly or
indiThough powerful, the inferences made by these algo- rectly [
          <xref ref-type="bibr" rid="ref12">11</xref>
          ]. With direct input, the end-user indicates
rithms lead to increased uncertainty and failures. For the artifact or subject ‘target.’ Because of this directness,
example, in comic colorization, ‘flatting’ is the process of inputs are usually in the same medium as the target
artiautomatically creating colored polygons under diferent fact. In some situations, a portion of the artifact can also
parts of the linear art (e.g., one for the face, one for the be used as a direct input. For example, we can select a
shirt, etc.). The algorithm for automated flatting makes portion of the image or the story. At the other extreme
inferences about shapes even when they are not ‘closed’ are those inputs that do not directly impact the artifact
in an expected way. For example, creases in a drawing but may give broad instructions on how the tool should
for a shirt may lead a poorly designed algorithm to make implement something. The simplest example might be a
too many polygons or not connect them appropriately. slider control for some parameters. The user isn’t
touchIdeally, the system will produce one polygon that en- ing the artifact directly (i.e., the story or image) but the
capsulates the entire shirt. However, current flatting change in the slider guides the tool. The modality of
software is imperfect and can make the wrong inference. indirect input can be far from the medium (e.g., visual
The algorithm’s uncertainty in what makes up the object arts as artifacts and numbers as inputs). As with our
can lead to unexpected bleeding [20]–a failure case. introductory example, abstract visual encodings can also
        </p>
        <p>However, in the creative setting, and specifically for be used for indirect inputs. In that example, the end-user
generative algorithms, uncertainty can be a feature (rather drew the character’s fortune to produce text.
than a failure). Or, more precisely, the line between a
novel, desirable result, and an error are not necessarily
clear cut. There is rarely a single gold standard for what 2.3.2. Characteristics: Input Precision
should be generated, and the user might subjectively
decide whether the output fits their goal.</p>
      </sec>
      <sec id="sec-4-2">
        <title>2.2.2. Characteristics: Ease of Algorithmic</title>
      </sec>
      <sec id="sec-4-3">
        <title>Iteration</title>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>Iterative design is important in creating artifacts. This</title>
      <p>is mainly due to the explorative nature of the task. How
easy it is to iterate depends on the algorithm’s properties.
First among these is latency–the time taken by the
algorithm for each cycle. The lower the latency, the easier
the iteration. For example, when moving a box with a
mouse cursor, the iteration is real-time as the box’s
position instantly updates with the user’s movement. On the
other hand, many generative algorithms take significant
time to generate artifacts, significantly slowing down
iteration.</p>
      <p>A second algorithmic aspect that impacts iteration is
scope. Here, we define scope as relating to what parts of</p>
    </sec>
    <sec id="sec-6">
      <title>While there are numerous input approaches for artifact</title>
      <p>creation systems, they vary on the spectrum of precision.</p>
      <p>These varying levels are helpful in diferent contexts. The
most traditional type of widgets receives one specific
value. Examples include a number in the slider or a
category in a dropdown box. With this precise control,
users will expect the output to react precisely.</p>
      <p>Not all inputs need to be precise. Natural language
prompt is one example and can handle a wider range of
input precision [22, 23]. Roughly specified language would
be imprecise, but at the same time, allow a high degree
of freedom in how it can be interpreted. For example,
asking for a “rough texture” can mean many things—
anything from Jackson Pollock’s chaotic style to Van
Gogh’s impressionism. However, language can support
ifner control. For example, if we say “move the selected
square 3 pixels left,” this does not leave much room for
misinterpretation.</p>
      <p>
        At the imprecise end, we often find Examples as in- or examples, can increase iterative costs. This is mainly
puts [
        <xref ref-type="bibr" rid="ref14">24, 21</xref>
        ]. While they are often used as direct material due to the vast space of options for these modalities. For
for transfer (e.g., source of style in visual style transfer), natural language prompts, the user needs to come up with
it is up to the algorithm to determine, if it can, which better wordings or more specific details on the prompts.
aspects of the input should be followed closely and which This can be tricky if the user is to express diferences in
are only suggestions. For example, when transferring degree (e.g., how would one use language to express the
the style of Van Gogh’s The Starry Night, it may not be level of roughness of the texture in a painting?).
Simiclear whether the user wants the colors or textures to larly, iterating with examples is dificult because the user
be transferred. Adding more examples might make the needs to search for more or better examples. If such an
target clearer. However, it may be hard for the user to de- example can’t be easily found or created, the user will
termine which attributes overlap between the examples struggle to iterate.
and which are ambiguous. The interaction with uncer- Gestural or sketch inputs can help with iteration. While
tain algorithms makes this problem even more complex these input modalities come at the cost of precision,
gesas it is not obvious if the issue is with the input or the tural input is flexible, intuitive with lowered cognitive
inherent ambiguity of the system. demands. These properties can reduce iteration time. For
      </p>
      <p>As with language prompts, Gestural inputs, such as example, the user can erase and redraw a portion of the
sketches or hand gestures, can also have a wide range of sketches to quickly change the specifications.
precision. For example, gestural inputs for direct
manipulation require outputs to follow the given input exactly.</p>
      <p>
        When resizing a box in graphics editors, users expect 3. Designing Generative
the box to follow the cursor they are moving. However, Co-Creation Tools
sketches can be used for low-precision input. For
example, sketches can express flexible and lightweight ideas The design dimensions above represent a large design
with their roughness, ambiguity, and uncertainty [
        <xref ref-type="bibr" rid="ref7 ref8">7, 6</xref>
        ]. space. However, we can begin to consider points in the
Similarly, hand gestures have been used to provide im- space that are either required, or are more suitable, for
precise but intuitive and flexible inputs, such as serving co-creation tools.
as rough scafolds in 3D modeling [
        <xref ref-type="bibr" rid="ref15">25</xref>
        ]. As we see in
these examples, gestural inputs can be designed to pro- 3.1. Requirements for Generative
vide high intuitiveness and flexibility and be traded of Co-creation Tools
against precision.
      </p>
      <p>
        We note that input precision is often related to input As we argued above, generative algorithms are usually
dificulty. As we know from psychophysical properties used to support transfer or generation. Additionally,
such as Fitts’s Law [
        <xref ref-type="bibr" rid="ref16">26</xref>
        ], certain input precision comes at these systems have increasingly leaned towards
machinethe cost of time or dificulty. Lower precision interactions, learning-based approaches. Thus, we are largely working
such as gestures, can often lower interaction dificulty. in the non-deterministic algorithmic space and this
implies a couple of key requirements for tools.
2.3.3. Characteristics: Ease of Input Iteration First, iteration should be easy and fast. In
creative tasks, iteration and exploration are necessary as
they expose the artist to more options and, eventually,
a concretization of ‘direction’ [
        <xref ref-type="bibr" rid="ref17 ref18">27, 28</xref>
        ]. Thus, users of
creative tools often want to be able to iterate, which is
well-aligned with the reality that to use non-deterministic
tools, they need to iterate. Unfortunately, sometimes the
cost of iteration becomes high. Thus, tools should either
act to speed up the number of iterations and, if that is
not possible, to reduce them. In both situations, reducing
the iteration cost is critical.
      </p>
      <p>Second, algorithmic uncertainty should match the
user’s expectations. With standard algorithms, we
would only need to worry about the user’s expectations
in how their input and deterministic output relate. For
example, dragging an icon into the trash would lead to
it being deleted. However, with non-deterministic
algorithms, instead of a specific output, they would need to</p>
    </sec>
    <sec id="sec-7">
      <title>Just as we consider the iterative cost at the algorithmic</title>
      <p>level, it is worth considering it at the input level as well.
Though these two might be tied, a tool might have
relatively small back-end iterative costs but widely diverging
front-end costs. For example, the algorithm itself might
run quickly but generating good example inputs may
take a long time. Thus, diferent input approaches vary
in how well they support iteration.</p>
      <p>Traditional input widgets, such as numerical values
on sliders, are relatively easy. With a single slider, the
control options given to the user are tightly restricted,
and a change in value does not require much efort.
However, even with simple slider widgets, the user needs to
decide if a control should be changed and then make the
actual change to the correct value. As the number of
input controls grows, so does iteration cost.</p>
      <p>Other types of inputs, such as natural language prompts</p>
    </sec>
    <sec id="sec-8">
      <title>To achieve our requirements, we argue that interaction is a critical factor. We focus on possible input approaches. These will naturally range based on the type of input directness.</title>
      <sec id="sec-8-1">
        <title>3.3.1. Direct Inputs: A Small Space of Design</title>
        <p>model a range of possible outputs. Without this under- 3.3. Input Design
standing, end-users are likely to be dissatisfied with the
results. They will also find it dificult to model how their
input choices will lead to a better, or more certain, output.</p>
        <p>Our advantage in creative tools is that some degree of
uncertainty is actually a desired property. Our goal is
not necessarily to make the tool appear deterministic as
creativity often requires ‘surprise.’ Thus, users both want
and expect some level of (controlled) uncertainty. A user
may be willing to make a rough specification. At the
same time, they would understand and expect that the
tool will have some degrees of freedom within that space.</p>
        <p>Note that none of this is to say that we need to force
the algorithms to match the user’s expectations. In some
cases, users might not have well-defined expectations. In
other cases, we may change their expectations.</p>
      </sec>
    </sec>
    <sec id="sec-9">
      <title>Direct inputs are usually made in the same medium as the</title>
      <p>artifact. For example, we might use low-fidelity sketches
on the drawing surface when the target artifact is visual.</p>
      <p>
        These sketches will then be transferred to high-fidelity
images on (essentially) the same surface/encoding [
        <xref ref-type="bibr" rid="ref10">9</xref>
        ]. In
other cases, the algorithm may simply append elements
to the sketch [18]. This type of input serves as the
‘material’ for the generation—where the transfer is applied or
what the generative algorithms build upon.
      </p>
      <p>While direct inputs may depend on the application
domain, their specific type may be largely constrained
to a small design space. This is largely because the
representation depends on the target artifact’s medium (e.g.,
a drawing canvas). Additionally, the interactions are
constrained by the underlying algorithm. For example,
we may train an algorithm to produce a photo-realistic
image given a low-resolution sketch. Such datasets are
more readily available and easier to produce. The
userfacing input modality and form are thus constrained to
something that looks like the training data. Finally,
direct inputs create a set of expectations for the end-user
that need to be maintained in the interactive controls.</p>
      <p>Because of these constraints, which may limit our design
space options, we move to consider indirect inputs.</p>
      <p>There are various ways to approach the iteration and
uncertainty problems on the algorithmic side. For
example, by adding extensive controllability features, we can
provide the user with fine-grained controls for steering
the behavior of the generative algorithms. However, this
requires building algorithms that can actually accept all
these controls.</p>
      <p>On the positive side, detailed control may reduce the
number of iterations, at least from the algorithmic
perspective. That is, fine-grained controls would reduce the
ambiguity of the input and enable the generative system
to produce a more targeted response. Detailed controls
also work to ‘teach’ the end-user how to model and direct
the underlying algorithm. Their expectations of system
capabilities would come to be more in alignment with 3.3.2. Indirect Inputs: Approaches and Their
reality with fewer iterations. Of course, reducing the Limitations
latency of the algorithm would also facilitate the ease of
iteration. Clever designs, such as using smaller models Indirect inputs serve as instructions for both what and
before executing more costly larger ones may help here. how to generate. Unlike direct inputs, they are not
depen</p>
      <p>However, shifting the responsibility of satisfying our dent on the artifact’s medium. One can work in abstract
requirements to the algorithmic side exclusively is not spaces or through abstract representations. Thus, there
realistic. Regardless of the algorithm, many bottlenecks is often more freedom to design with indirect inputs. The
for iteration are from the interaction side. Increasing consequence of freeing ourselves from the constraints of
the number of controls may be cognitively costly for the the domain also enables us to consider additional
algoend-user. This is not to say that fewer controls or simpler rithm types. This flexibility further afords a better
abilinputs, such as examples or prompts, reduce cognitive ity to match the end user’s high-level intentions rather
cost. The cognitive cost of figuring out how to change than forcing them to work within their algorithm and
or create an example can be equally bad. When coupled interface constraints. However, this is not to say that all
with the specific demands of creative applications–that indirect inputs are good ones. A novel indirect
interacwe want some iteration and some surprise–achieving tion might be further from the target artifact’s modality
a ‘sweet-spot’ through algorithmic means alone seems and thus might be harder to master when the mapping
implausible. Put another way, simply changing the algo- is complex. A poorly designed indirect interaction can
rithm can’t solve our problem if the interface costs are also increase cognitive costs and reduce the ability to
high or the user’s requirements for a creative tool are iterate. All together, indirect inputs open up a vast space
unmet. of possibilities but introduce various pitfalls.</p>
      <sec id="sec-9-1">
        <title>3.2. Algorithmic Design</title>
        <p>
          To better understand which aspects may help or hin- encodings. Finally, as we have argued before, gestural
der, we focus on three types of inputs: traditional input inputs convey the sense of being rough and flexible. This
widgets, natural language prompts, and gestural inputs. strongly aligns with the non-determinism of the
algoTraditional input widgets, such as sliders for numerical rithm and the ambiguity of the user’s intent. Moreover,
inputs, represent the simplest option. If we are able to use it can convey the ‘unfinished’ nature of the generative
these in the interface, it often means that we can directly process.
map the user’s expectation to the actual behavior of the As a demonstration of the feasibility of this approach
system. In reality, this depends on how the end-user un- we describe our system TaleBrush [
          <xref ref-type="bibr" rid="ref11">10</xref>
          ] (Figure 1).
Talederstands the construct represented with the widget. If Brush is a human-AI story co-creation tool that
generthe label on the slider is ambiguous (e.g., this will control ates story sentences according to the specifications of the
the ‘brightness’ of the text) or novel (e.g., this will control protagonist’s fortune. For example, if we were
describthe ‘certitudeness’ of the text), the user may struggle with ing Cinderella’s fortune, we might say that: her fortune
the control. Clearly, increased experience with the tool started low (with he stepmother and sisters), improved
will improve as the user calibrates to the system. More greatly as she went to the ball, collapsed as she was forced
critical, however, are situations where a user only has a to flee, and then improved again when the prince found
rough idea of what they want to generate. Here, a stan- her.
dard input widget may be insuficient. As critically, the TaleBrush allows the user first to input a portion of the
algorithms themselves may not deliver on the precision story (direct input) in the text box (Figure 1A1). Then,
of the input. Thus, the interface is over-promising. The they can sketch out the protagonist’s fortune in a
2ease of iteration with input widgets largely depends on dimensional line sketch as in Figure 1A2. This is roughly
the complexity of the interface. A single slider is simple, a standard time series with the  and  axes standing for
but many controls and interactions will naturally become sequence position and fortune levels, respectively.
Usmore challenging. ing this sketched line (which is actually represented as a
        </p>
        <p>Recent generative co-creation systems have enabled sketch rendering), TaleBrush will generate a story
(Figindirect natural language prompts as input. For exam- ure 1B1). Because the underlying algorithm is ambiguous
ple, natural language prompts can steer vision-language and may not precisely match the desired fortune sketch,
models to generate visual images [22, 23]. As these ap- the best matching generation is also displayed in the
viproaches can be used with imprecision or ambiguity, they sualization (Figure 1B2). Technically, this sketch-based
are useful for giving high-level specifications on gener- control is powered by steering a big pretrained language
ations. However, they would be dificult to iterate with model with a smaller language model that receives sketch
due to the vast space of inputs. position as control code. More details can be found in</p>
        <p>
          Surprisingly, few human-AI co-creation tools have the Chung et al. [
          <xref ref-type="bibr" rid="ref11">10</xref>
          ]
used gestural or sketch interactions for indirect control. With TaleBrush, the benefits of gestural inputs hold.
We argue that this is a missed opportunity as there are First, iteration on the generation is easy. The user only
several benefits to this approach. In the next section, we needs to redraw parts of the sketch. Additionally, a single
expand on this possibility and why it may be appropri- drawn line expresses two dimensions simultaneously: (1)
ate. where in the story and (2) at what fortune level should the
sentence be. Notably, the first (position) is a direct input,
whereas the fortune level represents an indirect one. In
4. Gestural Indirect Inputs for reality, we also use the speed at which the sketched line is
Generative Co-Creation drawn to indicate how much ambiguity the user will
tolerate in the generated result. This is visually represented
We propose that gestural or sketch-based interfaces for in the thickness of the line. A thinner line indicates the
indirect specification satisfy our requirements for co- user wants a better match. Internally, this is implemented
generation tools. At the very least, this approach may by regenerating the sentences multiple times and
findcomplement other input controls. First, simple gestural ing the one that best matches that desired fortune level.
interactions (e.g., producing a rough sketch) are easy to This visualized boundary further emphasizes the
ambiiterate with. This characteristic can complement more guity and non-determinism of the algorithm. Note that
efortful controls such as prompts or examples. More- a ‘sketch’ does not necessarily mean a ‘sketchy
appearover, the multi-dimensional characteristics of sketches ance.’ However, we have opted to use this rendering
and gestures can reduce the efort to interact with multi- aesthetic to further lower the user’s expectations that the
ple attributes simultaneously. For example, 2D sketching algorithm should be precise [
          <xref ref-type="bibr" rid="ref7 ref8">7, 6</xref>
          ].
coupled with pressure and speed recognition can be used
to encode multiple parameters simultaneously. This
flexibility also means that we can work in abstract visual
        </p>
      </sec>
      <sec id="sec-9-2">
        <title>4.1. Design Approaches for Gestural</title>
      </sec>
      <sec id="sec-9-3">
        <title>Indirect Inputs</title>
      </sec>
    </sec>
    <sec id="sec-10">
      <title>Implementing TaleBrush has given us some insight about what may work well (and poorly) for indirect gestural inputs.</title>
      <sec id="sec-10-1">
        <title>4.1.1. Ease of Iteration on Input</title>
        <p>Controlling the precision When using gestural
inputs, the user’s intention regarding the precision might
vary. For example, in TaleBrush, the user might have
wanted the algorithm to follow the generation more
tightly when they draw the line with more care. The
system can be designed to leverage other input
dimensions to control the precision to reflect these intentions
better. In TaleBrush, the sketching speed was used to
decide how tightly generation should be done. That is,
if the user drew slowly, we assumed this indicated that
they wanted a better fit (represented as a thinner error
envelope). In this way, gestural interactions and
representations can be used to align input precision with
system capabilities.</p>
      </sec>
      <sec id="sec-10-2">
        <title>Combine direct and indirect input if possible For</title>
        <p>some generation tools, indirect inputs may be suficient.</p>
        <p>For example, if the tool generates any character
biographies based on the good-evil nature of the character, then
it might not require direct inputs. However, as with
TaleBrush, certain tasks require control with direct input.</p>
        <p>The user needs to be able to indicate, “where the
generation should be done’’ (e.g., where in the story a certain 5. Conclusion
fortune level should exist or what does the start of the
story look like?). In some cases, as we did with TaleBrush, In this position paper, we have explored where generative
the indirect and direct controls can be combined into a algorithms sit in the overall design space of co-creative
single gestural sketch. That is, with a single brushstroke, tools. We have further isolated those properties that are
the sequential position ( position–the ‘where’) and the desirable and potentially required for supporting
humanlevel of the protagonist’s fortune ( position–the ‘how’) AI co-creation. Our focus was on how inputs (both the
are both specified. ‘what’ and the ‘how’) can interact with underlying
algorithms. Our focus on enabling iteration and managing
Complement hard-to-iterate control inputs (lan- expectations allowed us to consider the pros and cons of
guage, examples) Spatial positions by themselves do diferent input types. Ultimately, we argued that
gestunot necessarily convey meaning. They are meaningful ral and sketch-based interactions would work well for
when combined with semantic structures that can be put the control of generative algorithms. We showcased the
on a continuous scale. For example, TaleBrush takes a benefits of this approach with TaleBrush. We believe
restricted set of numerical semantics: whether the char- that there are significant possibilities opened up by using
acter’s fortune is good or ill. However, this design can be abstract visual representations when coupled with novel
extended to receive qualitative inputs as the endpoints interaction types.
of the axes. For example, the user can give natural
language prompts or examples on each end and explore the
confined space with gestural inputs. This complements Acknowledgments
the limitations and features of diferent input approaches.</p>
        <p>Language prompts and examples lack the ease of iter- We thank our reviewers for providing helpful feedback
ation, which is the strength of gestural inputs. On the on this work.
other hand, gestural inputs lack semantics, which
language prompts and examples can convey.</p>
      </sec>
      <sec id="sec-10-3">
        <title>Match the algorithmic precision with the input pre</title>
        <p>cision To have better expectations of how the
algorithm will behave, the user should ideally be aware of
the precision of the algorithm. Gestural inputs can be
designed to convey this information. For example, in
TaleBrush, this level of precision is reflected in the width
of the sketched line. This was designed to match the
median error from the test dataset we used during
development. Thus, the interaction and representation can be
used in ways that reduce ambiguity and help to match
(and manage) expectations.</p>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>4.1.2. Matching Input Precision with Algorithmic Uncertainty</mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>I.</given-names>
            <surname>Goodfellow</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Pouget-Abadie</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Mirza</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Xu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Warde-Farley</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Ozair</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Courville</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Bengio</surname>
          </string-name>
          ,
          <article-title>Generative adversarial nets</article-title>
          , in: Z.
          <string-name>
            <surname>Ghahramani</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          <string-name>
            <surname>Welling</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          <string-name>
            <surname>Cortes</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          <string-name>
            <surname>Lawrence</surname>
            ,
            <given-names>K. Q.</given-names>
          </string-name>
          <string-name>
            <surname>Weinberger</surname>
          </string-name>
          (Eds.),
          <source>Advances in Neural Information Processing Systems</source>
          , volume
          <volume>27</volume>
          ,
          <string-name>
            <surname>Curran</surname>
            <given-names>Associates</given-names>
          </string-name>
          , Inc.,
          <year>2014</year>
          . URL: https://proceedings.neurips.cc/paper/2014/file/ 5ca3e9b122f61f8f06494c97b1afccf3-Paper.pdf.
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>T.</given-names>
            <surname>Brown</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Mann</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Ryder</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Subbiah</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. D.</given-names>
            <surname>Kaplan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Dhariwal</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Neelakantan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Shyam</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Sastry</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Askell</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Agarwal</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Herbert-Voss</surname>
          </string-name>
          , G. Krueger,
          <string-name>
            <given-names>T.</given-names>
            <surname>Henighan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Child</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Ramesh</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Ziegler</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Wu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Winter</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Hesse</surname>
          </string-name>
          , M. Chen, [13]
          <string-name>
            <given-names>C. J.</given-names>
            <surname>Steinmetz</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. D.</given-names>
            <surname>Reiss</surname>
          </string-name>
          , Steerable discovery of E. Sigler,
          <string-name>
            <given-names>M.</given-names>
            <surname>Litwin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Gray</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Chess</surname>
          </string-name>
          , J. Clark, neural audio efects,
          <year>2021</year>
          . arXiv:
          <volume>2112</volume>
          .02926.
          <string-name>
            <surname>C. Berner</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          <string-name>
            <surname>McCandlish</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Radford</surname>
            ,
            <given-names>I. Sutskever</given-names>
          </string-name>
          , [14]
          <string-name>
            <given-names>R.</given-names>
            <surname>Louie</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Coenen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C. Z.</given-names>
            <surname>Huang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Terry</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C. J. D.</given-names>
            <surname>Amodei</surname>
          </string-name>
          ,
          <article-title>Language models are few-shot learn- Cai, Novice-ai music co-creation via ai-steering ers</article-title>
          , in: H.
          <string-name>
            <surname>Larochelle</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          <string-name>
            <surname>Ranzato</surname>
          </string-name>
          , R. Hadsell,
          <article-title>tools for deep generative models</article-title>
          , in:
          <string-name>
            <surname>Proceedings M. F. Balcan</surname>
          </string-name>
          , H. Lin (Eds.),
          <source>Advances in Neu- of the 2020 CHI Conference on Human Factors in ral Information Processing Systems</source>
          , volume
          <volume>33</volume>
          ,
          <string-name>
            <surname>Computing</surname>
            <given-names>Systems</given-names>
          </string-name>
          , CHI '20, Association for ComCurran Associates, Inc.,
          <year>2020</year>
          , pp.
          <fpage>1877</fpage>
          -
          <lpage>1901</lpage>
          . puting Machinery, New York, NY, USA,
          <year>2020</year>
          , p. URL: https://proceedings.neurips.cc/paper/2020/ 1-
          <fpage>13</fpage>
          . URL: https://doi.org/10.1145/3313831.3376739. file/1457c0d6bfcb4967418bfb8ac142f64a-Paper.pdf.
          <source>doi:10.1145/3313831</source>
          .3376739.
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>Y.</given-names>
            <surname>Wu</surname>
          </string-name>
          , E. Manilow,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Deng</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Swavely</surname>
          </string-name>
          , K. Kast- [15]
          <string-name>
            <surname>C.-J. Chang</surname>
            ,
            <given-names>C.-Y.</given-names>
          </string-name>
          <string-name>
            <surname>Lee</surname>
            ,
            <given-names>Y.-H.</given-names>
          </string-name>
          <string-name>
            <surname>Yang</surname>
            , Variable-length ner, T. Cooijmans,
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Courville</surname>
          </string-name>
          , C.
          <article-title>-Z. A. Huang, music score infilling via xlnet and musically specialJ. Engel, Midi-ddsp: Detailed control of musi- ized positional encoding</article-title>
          ,
          <year>2021</year>
          . arXiv:
          <volume>2108</volume>
          .05064. cal performance via hierarchical modeling,
          <year>2021</year>
          . [16]
          <string-name>
            <given-names>P.</given-names>
            <surname>Ammanabrolu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W.</given-names>
            <surname>Cheung</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W.</given-names>
            <surname>Broniec</surname>
          </string-name>
          ,
          <string-name>
            <surname>M. O.</surname>
          </string-name>
          <year>arXiv</year>
          :
          <volume>2112</volume>
          .09312.
          <string-name>
            <surname>Riedl</surname>
          </string-name>
          , Automated storytelling via causal, common-
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>A.</given-names>
            <surname>Ramesh</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Pavlov</surname>
          </string-name>
          , G. Goh,
          <string-name>
            <given-names>S.</given-names>
            <surname>Gray</surname>
          </string-name>
          , C. Voss, sense plot ordering, CoRR abs/
          <year>2009</year>
          .00829 (
          <year>2020</year>
          ). A.
          <string-name>
            <surname>Radford</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          <string-name>
            <surname>Chen</surname>
            ,
            <given-names>I. Sutskever</given-names>
          </string-name>
          , Zero-shot text- URL: https://arxiv.org/abs/
          <year>2009</year>
          .00829.
          <article-title>to-image generation</article-title>
          , in: M.
          <string-name>
            <surname>Meila</surname>
            , T. Zhang [17]
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Fan</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          <string-name>
            <surname>Lewis</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          <string-name>
            <surname>Dauphin</surname>
          </string-name>
          , Strategies for (Eds.),
          <source>Proceedings of the 38th International Con- structuring story generation</source>
          ,
          <source>in: Proceedings of ference on Machine Learning</source>
          , volume
          <volume>139</volume>
          of
          <article-title>Pro- the 57th Annual Meeting of the Association for ceedings of Machine Learning Research</article-title>
          , PMLR,
          <year>2021</year>
          ,
          <string-name>
            <given-names>Computational</given-names>
            <surname>Linguistics</surname>
          </string-name>
          , Association for Compp.
          <volume>8821</volume>
          -
          <fpage>8831</fpage>
          . URL: https://proceedings.mlr.press/ putational Linguistics, Florence, Italy,
          <year>2019</year>
          , pp.
          <source>v139/ramesh21a.html</source>
          .
          <volume>2650</volume>
          -
          <fpage>2660</fpage>
          . URL: https://aclanthology.org/P19-1254.
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>L. A.</given-names>
            <surname>Gatys</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A. S.</given-names>
            <surname>Ecker</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Bethge</surname>
          </string-name>
          , Image style doi:
          <volume>10</volume>
          .18653/v1/
          <fpage>P19</fpage>
          - 1254.
          <article-title>transfer using convolutional neural networks</article-title>
          , in: [18]
          <string-name>
            <given-names>J. E.</given-names>
            <surname>Fan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Dinculescu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Ha</surname>
          </string-name>
          ,
          <string-name>
            <surname>Collabdraw:</surname>
          </string-name>
          <article-title>An 2016 IEEE Conference on Computer Vision and environment for collaborative sketching with an Pattern Recognition (CVPR), IEEE</article-title>
          , USA,
          <year>2016</year>
          , pp.
          <source>artificial agent</source>
          ,
          <source>in: Proceedings of the 2019 on 2414-2423</source>
          . doi:
          <volume>10</volume>
          .1109/CVPR.
          <year>2016</year>
          .
          <volume>265</volume>
          . Creativity and Cognition, C&amp;C '19, Association for
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>M. D.</given-names>
            <surname>Gross</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E. Y.</given-names>
            <surname>Do</surname>
          </string-name>
          ,
          <article-title>Ambiguous intentions: A Computing Machinery</article-title>
          , New York, NY, USA,
          <year>2019</year>
          , paper
          <article-title>-like interface for creative design</article-title>
          , in: ACM p.
          <fpage>556</fpage>
          -
          <lpage>561</lpage>
          . URL: https://doi.org/10.1145/3325480. Symposium on User
          <source>Interface Software and Tech- 3326578. doi:10.1145/3325480</source>
          .3326578. nology, ACM,
          <year>1996</year>
          , pp.
          <fpage>183</fpage>
          -
          <lpage>192</lpage>
          . [19]
          <string-name>
            <given-names>Y.</given-names>
            <surname>Lin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Guo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Chen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Yao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Ying</surname>
          </string-name>
          , It is your
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>J.</given-names>
            <surname>Landay</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Myers</surname>
          </string-name>
          ,
          <article-title>Sketching interfaces: toward turn: Collaborative ideation with a co-creative more human interface design</article-title>
          ,
          <source>Computer</source>
          <volume>34</volume>
          (
          <year>2001</year>
          )
          <article-title>robot through sketch</article-title>
          ,
          <source>in: Proceedings of the 56-64. doi:10.1109/2</source>
          .910894. 2020 CHI Conference on Human Factors in Com-
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>M.</given-names>
            <surname>Eitz</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Hays</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Alexa</surname>
          </string-name>
          ,
          <article-title>How do humans sketch puting Systems</article-title>
          , CHI '20,
          <string-name>
            <surname>Association</surname>
          </string-name>
          for Comobjects?,
          <source>ACM Trans. Graph</source>
          .
          <source>(Proc. SIGGRAPH) 31 puting Machinery</source>
          , New York, NY, USA,
          <year>2020</year>
          , p. (
          <year>2012</year>
          )
          <volume>44</volume>
          :
          <fpage>1</fpage>
          -
          <lpage>44</lpage>
          :
          <fpage>10</fpage>
          .
          <fpage>1</fpage>
          -
          <lpage>14</lpage>
          . URL: https://doi.org/10.1145/3313831.3376258.
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>S.-Y.</given-names>
            <surname>Chen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W.</given-names>
            <surname>Su</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Gao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Xia</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Fu</surname>
          </string-name>
          , Deep- doi:10.1145/3313831.3376258. FaceDrawing:
          <article-title>Deep generation of face images from</article-title>
          [20]
          <string-name>
            <given-names>C.</given-names>
            <surname>Yan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. J. Y.</given-names>
            <surname>Chung</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Yoon</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Gingold</surname>
          </string-name>
          , E. Adar, sketches, ACM Transactions on
          <string-name>
            <surname>Graphics (Proceed- S. R. Hong</surname>
          </string-name>
          ,
          <source>FlatMagic: Improving Flat Colorization ings of ACM SIGGRAPH</source>
          <year>2020</year>
          )
          <volume>39</volume>
          (
          <year>2020</year>
          )
          <volume>72</volume>
          :
          <fpage>1</fpage>
          -
          <lpage>72</lpage>
          :
          <fpage>16</fpage>
          .
          <article-title>through AI-driven Design for DigitalComic Pro-</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>J. J. Y.</given-names>
            <surname>Chung</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W.</given-names>
            <surname>Kim</surname>
          </string-name>
          ,
          <string-name>
            <surname>K. M. Yoo</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          <string-name>
            <surname>Lee</surname>
            , E. Adar, fessionals, Association for Computing Machinery,
            <given-names>M.</given-names>
          </string-name>
          <string-name>
            <surname>Chang</surname>
          </string-name>
          , TaleBrush: Sketching Stories with Gen- New York, NY, USA,
          <year>2022</year>
          . erative Pretrained Language Models, Association [21]
          <string-name>
            <given-names>L.</given-names>
            <surname>Sheng</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Lin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Shao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <article-title>Avatar-net: for Computing Machinery</article-title>
          , New York, NY, USA, Multi
          <article-title>-scale zero-shot style transfer by feature dec2022. oration, in: Computer Vision</article-title>
          and Pattern Recog-
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>J. J. Y.</given-names>
            <surname>Chung</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>He</surname>
          </string-name>
          ,
          <string-name>
            <surname>E. Adar,</surname>
          </string-name>
          <article-title>The intersection of nition (CVPR</article-title>
          ),
          <source>2018 IEEE Conference on, 2018</source>
          , pp.
          <article-title>users, roles, interactions, and technologies in cre- 1-9. ativity support tools</article-title>
          , in: Conference on Designing [22]
          <string-name>
            <given-names>A.</given-names>
            <surname>Nichol</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Dhariwal</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Ramesh</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Shyam</surname>
          </string-name>
          ,
          <string-name>
            <surname>Interactive</surname>
            <given-names>Systems</given-names>
          </string-name>
          , ACM,
          <year>2021</year>
          , pp.
          <fpage>1817</fpage>
          -
          <lpage>1833</lpage>
          . P. Mishkin,
          <string-name>
            <given-names>B.</given-names>
            <surname>McGrew</surname>
          </string-name>
          ,
          <string-name>
            <surname>I. Sutskever</surname>
          </string-name>
          , M. Chen,
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>B.</given-names>
            <surname>Syed</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Verma</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B. V.</given-names>
            <surname>Srinivasan</surname>
          </string-name>
          ,
          <string-name>
            <surname>A</surname>
          </string-name>
          . Natara- Glide:
          <article-title>Towards photorealistic image generation jan, V. Varma, Adapting language models for and editing with text-guided difusion models, 2021. non-parallel author-stylized rewriting</article-title>
          ,
          <year>2020</year>
          . arXiv:
          <volume>2112</volume>
          .10741. arXiv:
          <year>1909</year>
          .
          <volume>09962</volume>
          . [23]
          <string-name>
            <given-names>F.</given-names>
            <surname>Huang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. F.</given-names>
            <surname>Canny</surname>
          </string-name>
          , Sketchforme:
          <article-title>Composing sketched scenes from text descriptions for interactive applications</article-title>
          ,
          <source>in: Proceedings of the 32nd Annual ACM Symposium on User Interface Software and Technology, UIST '19</source>
          ,
          <string-name>
            <surname>Association</surname>
          </string-name>
          for Computing Machinery, New York, NY, USA,
          <year>2019</year>
          , p.
          <fpage>209</fpage>
          -
          <lpage>220</lpage>
          . URL: https://doi.org/10.1145/3332165. 3347878. doi:
          <volume>10</volume>
          .1145/3332165.3347878.
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [24]
          <string-name>
            <given-names>E.</given-names>
            <surname>Frid</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Gomes</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Jin</surname>
          </string-name>
          ,
          <article-title>Music creation by example</article-title>
          ,
          <source>in: Proceedings of the 2020 CHI Conference on Human Factors in Computing Systems, CHI '20</source>
          ,
          <string-name>
            <surname>Association</surname>
          </string-name>
          for Computing Machinery, New York, NY, USA,
          <year>2020</year>
          , p.
          <fpage>1</fpage>
          -
          <lpage>13</lpage>
          . URL: https://doi.org/10.1145/ 3313831.3376514. doi:
          <volume>10</volume>
          .1145/3313831.3376514.
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [25]
          <string-name>
            <given-names>Y.</given-names>
            <surname>Kim</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.-G.</given-names>
            <surname>An</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. H.</given-names>
            <surname>Lee</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.-H.</given-names>
            <surname>Bae</surname>
          </string-name>
          ,
          <article-title>Agile 3D Sketching with Air Scafolding, Association for Computing Machinery</article-title>
          , New York, NY, USA,
          <year>2018</year>
          , p.
          <fpage>1</fpage>
          -
          <lpage>12</lpage>
          . URL: https://doi.org/10.1145/3173574. 3173812.
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [26]
          <string-name>
            <surname>I. S. MacKenzie</surname>
          </string-name>
          ,
          <article-title>Fitts' law as a research and design tool in human-computer interaction</article-title>
          ,
          <source>Hum.- Comput. Interact</source>
          .
          <volume>7</volume>
          (
          <year>1992</year>
          )
          <fpage>91</fpage>
          -
          <lpage>139</lpage>
          . URL: https:// doi.org/10.1207/s15327051hci0701_3. doi:
          <volume>10</volume>
          .1207/ s15327051hci0701_
          <fpage>3</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [27]
          <string-name>
            <surname>T. M. Amabile</surname>
          </string-name>
          ,
          <article-title>The social psychology of creativity: A componential conceptualization</article-title>
          .,
          <source>Journal of personality and social psychology 45</source>
          (
          <year>1983</year>
          )
          <fpage>357</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          [28]
          <string-name>
            <surname>T. M. Amabile</surname>
          </string-name>
          ,
          <source>Componential theory of creativity</source>
          (
          <year>2012</year>
          ).
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>