<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <journal-title-group>
        <journal-title>March</journal-title>
      </journal-title-group>
      <issn pub-type="ppub">1613-0073</issn>
    </journal-meta>
    <article-meta>
      <title-group>
        <article-title>ExpressEdit: Video Editing with Natural Language and Sketching</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Juho Kim</string-name>
          <email>juhokim@kaist.ac.kr</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Bekzat Tilekbay</string-name>
          <email>tilekbay@kaist.ac.kr</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Saelyne Yang</string-name>
          <email>saelyne@kaist.ac.kr</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Michal Lewkowicz</string-name>
          <email>michal.lewkowicz@yale.edu</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Alex Suryapranata</string-name>
          <email>alextio@kaist.ac.kr</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>School of Computing, KAIST</institution>
          ,
          <addr-line>Daejeon</addr-line>
          ,
          <country>Republic of Korea</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Yale University</institution>
          ,
          <addr-line>New Haven, Connecticut</addr-line>
          ,
          <country country="US">USA</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2024</year>
      </pub-date>
      <volume>1</volume>
      <fpage>8</fpage>
      <lpage>21</lpage>
      <abstract>
        <p>Informational videos serve as a crucial source of conceptual and procedural knowledge for many people. While it is important to make informational videos instructive and engaging, editing such videos (e.g., trimming, overlaying text/image, etc.) can be dificult and time-consuming. Especially for novice video editors, who often struggle with expressing and executing their editing ideas. We present ExpressEdit, a system that facilitates editing informational videos via natural language text and sketching directly on the video frame by interpreting multimodal editing commands and suggesting applicable edits. Powered by a multimodal technical pipeline, the system interprets (1) temporal, (2) spatial, and (3) operational references in an NL editing command and spatial references from sketching. This work ofers insights into building multimodal interfaces for video editing.</p>
      </abstract>
      <kwd-group>
        <kwd>video editing</kwd>
        <kwd>human-AI interaction</kwd>
        <kwd>multimodal input</kwd>
        <kwd>LLM</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>CEUR
ceur-ws.org</p>
    </sec>
    <sec id="sec-2">
      <title>1. Introduction</title>
      <p>Informational videos are videos that introduce, explain, or demonstrate conceptual or
procedural knowledge [1, 2, 3]. They encompass a broad range of topics such as cooking, health,
programming, and craft, and can be produced in various formats (e.g., lecture, tutorial, q&amp;a,
demonstration, etc.) [4, 5, 6, 7]. They have become a popular source of knowledge for many
people due to their rich and engaging content [5].</p>
      <p>However, editing informational videos is a tedious task that involves carefully organizing the
footage, removing unnecessary parts, and finding and incorporating additional media assets
[4]. While popular commercial tools for video editing ofer all the necessary instruments to
implement a variety of edits, for novices, these tools are dificult to learn and use, as they require
great manual efort and have steep learning curves [ 8, 9].</p>
      <p>We investigate how multimodality – natural language (NL) and sketching – can be leveraged
in the informational video editing scenario. We conducted a formative study with 10 video
editors with diverse levels of expertise and collected 176 expressions of video editing requests in
nEvelop-O
CEUR
Workshop
Proceedings
the form of NL texts, sketches, and media assets. Focusing on edit expressions that initiate edits
rather than revise or adjust applied edits, we found that editors feel comfortable expressing
their general editing requests through NL text and use sketching on top of the frame to indicate
specific locations or regions of interest.</p>
      <p>
        Based on the findings from the formative study, we built ExpressEdit, a multimodal interactive
system for editing informational videos. It supports the expression of video editing requests
through NL text and sketching on top of a frame, and is powered by a computer vision and large
language model-based technical pipeline that comprehends and executes the edits by extracting
and interpreting three types of references from the multimodal command (Figure 1d,e): (
        <xref ref-type="bibr" rid="ref1">1</xref>
        )
temporal location (e.g., “whenever he mentions laptop”), (
        <xref ref-type="bibr" rid="ref2">2</xref>
        ) spatial location within the frame
(e.g., “near the head”), and (
        <xref ref-type="bibr" rid="ref3">3</xref>
        ) references to edit operations and their parameters (e.g., “put a
text with the mentioned specifications of the laptop” ). ExpressEdit also provides the breakdown
of the command into aforementioned types of references, as well as gives reasoning for each
generated edit. From a user study with three novice video editors, we found ExpressEdit makes
video editing more efective and allows the participants to try out several edit ideas eficiently.
      </p>
    </sec>
    <sec id="sec-3">
      <title>2. Formative Study</title>
      <p>
        To learn about (
        <xref ref-type="bibr" rid="ref1">1</xref>
        ) the role of natural language (NL) text and sketching in describing video editing
requests and (
        <xref ref-type="bibr" rid="ref2">2</xref>
        ) their use cases in editing informational videos, we conducted a formative study
with 10 video editors, where they were asked to express their edit requests that would improve
the informativeness and engagement of a given video. We call these expressions video editing
commands. We focused on editing expressions that initiate edits rather than revise or adjust
applied edits. We believe future work can build on top of the initial investigation.
      </p>
      <p>We recruited five novices who had edited at least 2 videos and five experienced editors who
had edited at least 20 videos and 5 informational videos to cover a diverse range of editing
commands since edit expressions such as attention to detail and vocabulary used can vary
depending on the participant’s editing expertise. We chose five archived informational live
streams as raw footage for the study as they are usually unedited and closely resemble a
continuous stream of raw footage, which allows for tasks closer to real-world video editing
settings. To allow participants to describe their edit commands in both text and sketch, we used
Google Slides1, a popular slide authoring tool. We chose the tool because of its functionalities of
adding text, images, and shapes, which could be used in expressing edit commands. Participants
were also allowed to take a screenshot of a frame of the video and sketch over it. After the main
task, we conducted a short semi-structured interview to learn about the participant’s experience
performing the given task.</p>
      <p>As a result, the collected 176 multimodal editing commands and the interview results revealed
several patterns of multimodal video editing request expressions and the role of each modality
provided within the tool. We found that the participants comfortably described video editing
intents using various modalities: NL text, sketch, image, and graphics. In almost all the
commands, the participants referenced moments (e.g., timestamps, visual content, verbal
content) in the video where edit should be applied using NL text. Furthermore, they
used both NL text and sketching on top of the frame to reference the spatial location
of the edits within the frame. The participants used NL text to describe edit operations
and their parameters. They mentioned the exact names of the operation (e.g., cut, text, image,
etc.) or mentioned the main purpose or intended efect of the edit (e.g., highlight, emphasize,
focus). For parameters of the edit operations, they gave either precise numbers or descriptions
(e.g., large text, slow zoom). Lastly, participants frequently iterated on their commands to make
them clearer by revisiting and redefining their expressions, for example, to make them more
precise or to keep the consistency between several editing commands.</p>
    </sec>
    <sec id="sec-4">
      <title>3. ExpressEdit: Interface</title>
      <p>Based on formative study findings, we designed ExpressEdit, a multimodal video editing tool
for editing informational videos. Our system interprets the user’s editing command in the form
of NL&amp;S and suggests a set of edits with temporal location (where in the video timeline), spatial
location (where in the video frame), and edit operation &amp; parameters (which edit and how). To
better understand the generated edits and iterate on the NL&amp;S command, users can examine the
breakdown of the parsing results as well as the reasoning for the temporal &amp; spatial location of
each generated edit. Additionally, users can manually adjust the generated edits by the system
or create their own edits.</p>
      <p>(a) The examine panel
(b) The edit result panel
(c) The summary panel</p>
      <sec id="sec-4-1">
        <title>3.1. User Scenario</title>
        <p>To illustrate the envisioned user scenario of ExpressEdit, let’s follow Lia, a businesswoman
and a YouTube creator who wants to edit her video about entrepreneurship. She recorded
a talking-head video that she wants to make more concise, engaging, and informative using
ExpressEdit.</p>
        <sec id="sec-4-1-1">
          <title>3.1.1. Creating a new edit</title>
          <p>To start editing the video, Lia uploads her recorded footage to ExpressEdit and comes up with
the first edit she wants to implement. She presses the Add Tab button on the Tabs list and
creates a new layer on top of the video where she can apply edits. Edits within a single layer
will be of a single edit operation and cannot temporally intersect with each other.</p>
        </sec>
        <sec id="sec-4-1-2">
          <title>3.1.2. Describing the edit with Natural Language &amp; Sketch</title>
          <p>
            Lia decides to add text captions whenever she mentions valuable advice or tips. In the Edit
description (Figure 1d), she types “whenever there is a mention of advice or a tip, put it as a
big white text with a transparent background on the bottom part of the frame”. Additionally, Lia
specifies the exact part of the frame where the text should appear using the Sketch function
(Figure 1e) and draws the bounding box on the bottom half of the frame. She presses Enter to
process the NL&amp;S request and ExpressEdit provides (
            <xref ref-type="bibr" rid="ref1">1</xref>
            ) the breakdown of her command (Figure
2a), (
            <xref ref-type="bibr" rid="ref2">2</xref>
            ) a list of edits with respective reasoning and previews (Figure 1f) for each, and (
            <xref ref-type="bibr" rid="ref3">3</xref>
            ) a
summary of the processing results with the checklist of generated edits (Figure 2b). Along with
the summary, the system provides Get More Edits button that will generate more edits for the
same video editing command, Move to New Tab button which moves the set of edits to a new
tab (i.e., to a new layer on top of the video), and Previous Summary button which navigates
to the summary of the previous request (if one exists).
          </p>
        </sec>
        <sec id="sec-4-1-3">
          <title>3.1.3. Examining the results</title>
          <p>To determine if the system understood her NL&amp;S editing command, Lia examines the parsing
results for the NL part. She looks at Parsing Results (Figure 2a) and sees that the parts of the NL
command have colored backgrounds. To see the breakdown of the parsing results, she presses
the Reference Breakdown and ExpressEdit shows Temporal (when) (i.e., “whenever there is
a mention of advice or a tip”), Spatial (where) (i.e., “bottom part of the frame”), Edit (what)
(i.e., “put it as a big white text”), and Parameters (how) (i.e., “big white text” and “transparent
background”) references within the NL command, which assures her that the parsing was
accurate.</p>
        </sec>
        <sec id="sec-4-1-4">
          <title>3.1.4. Applying the generated edits</title>
          <p>To decide on generated edits to apply, Lia quickly glances over the generated edits. For each
generated edit, she looks at the Reasoning (Temporal (when) and Spatial (where) aspects) and
Preview (Figure 2b) to decide if she wants to apply the edit. Additionally, the system allows
her to test the edit by quickly “turning it on and of” by toggling the radio button. After making
all the decisions on generated edits, Lia can press the Get More Edits button to ask for more
edits for the same command or manually adjust/add her own edits.</p>
        </sec>
      </sec>
      <sec id="sec-4-2">
        <title>3.2. Edit operations</title>
        <p>
          ExpressEdit supports seven edit operations: (
          <xref ref-type="bibr" rid="ref1">1</xref>
          ) text overlay, (
          <xref ref-type="bibr" rid="ref2">2</xref>
          ) image overlay, (
          <xref ref-type="bibr" rid="ref3">3</xref>
          ) shape overlay
(i.e., circle, rectangle, star), (
          <xref ref-type="bibr" rid="ref4">4</xref>
          ) cutting out segments of the video, (
          <xref ref-type="bibr" rid="ref5">5</xref>
          ) zooming in/out, (6) cropping
the video, and (7) blurring the video. We decided to focus only on visual edit operations as they
cover important categories of parameters (temporal, spatial, edit-specific) common to other
edit operations (e.g., audio-related edits, coloring edits). We believe this set of edit operations
efectively demonstrates the feasibility of implementing various edit operations based on NL&amp;S
commands.
        </p>
      </sec>
      <sec id="sec-4-3">
        <title>3.3. Implementation</title>
        <p>ExpressEdit is implemented as a Web-based React2 application. The backend server is based on
Flask3, which hosted the videos along with their transcript and processed users’ requests. We
obtained all the videos and transcripts from YouTube using the youtube-dl package 4.
2React
3Flask
4youtube-dl</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>4. ExpressEdit: Pipeline</title>
      <p>We designed our pipeline to interpret natural language and sketch (NL&amp;S) descriptions of an edit
(Figure 3). In order to facilitate the real-time interactions, we perform pre-processing to extract
frame-level and clip-level metadata from the video that ExpressEdit uses to reason about video
context. The metadata consists of segmentations of frames (based on Segment Anything model
[10]) for every 1-second of the video and textual description of each 10-second clip (based on
InternVideo [11] and BLIP-2 [12]) that are then summarized using GPT-3.5 [13]. The overview
of the pipeline can be found in Appendix A.</p>
      <p>
        We first parse the NL request and divide the language command into (
        <xref ref-type="bibr" rid="ref1">1</xref>
        ) temporal reference
(i.e., any information in the NL command that could refer to a segment of the video), (
        <xref ref-type="bibr" rid="ref2">2</xref>
        ) spatial
reference (i.e., any information in the NL command that could refer to location or region in the
video frame), (
        <xref ref-type="bibr" rid="ref3">3</xref>
        ) edit operation reference (i.e., any information in the NL command that could
indicate an edit operation to use), and (
        <xref ref-type="bibr" rid="ref4">4</xref>
        ) edit parameter reference (i.e., any information in the
NL command that could refer to specific parameters of edit operation that was determined)
with GPT-4 [14].
      </p>
      <p>We interpret the temporal references using GPT-4 by decomposing them into “positional”,
“transcript-based”, and “video-based” references, compiling all segments of the video that match
these references, and pass the candidate segments (i.e., edits) further along the pipeline.</p>
      <p>To interpret the spatial references in NL we first categorize them based on their dependency
on the visual content of the video using GPT-4. For visual content-dependent references such
as NL references to a specific object in the video, we extract representative frames for each
candidate edit and then obtain the candidate spatial locations based on the segmentation that
has the highest cosine similarity with the given NL text reference and sketch (if provided) using
CLIP [15]. If no visual content-dependent references are detected and no sketch is given, we
designate the top-left corner of the frame as the candidate spatial location. Then, we refine
and resize each candidate spatial location with the visual content-independent spatial references
(e.g., left, top, etc.) using GPT-4.</p>
      <p>Lastly, using GPT-4, we identify the edit operations that are most suitable for the request based
on edit operation references and our system’s available edit operations: “text”, “image”, “shape”,
“blur”, “cut”, “crop”, or “zoom”. Then, we identify the video editing parameters corresponding to
each predicted edit operation for candidate edits based on identified edit parameter references.
To get text and images that are more appropriate to the video for “text” and “image” operations,
we provide video context to guide the GPT-4 generation process of text and search query,
respectively.</p>
    </sec>
    <sec id="sec-6">
      <title>5. Pilot Study</title>
      <p>
        We conducted a pilot user study with three novice video editors to (
        <xref ref-type="bibr" rid="ref1">1</xref>
        ) preliminarily evaluate the
potential of ExpressEdit in supporting video editing workflows of novices and (
        <xref ref-type="bibr" rid="ref2">2</xref>
        ) obtain early
feedback on the design of the system. The participants were allowed to freely use ExpressEdit for
25 minutes and edit a talking-head video about entrepreneurship5 to make it more informative
and engaging. We followed the think-aloud protocol and asked the participants to share their
thoughts while editing the video.
      </p>
      <p>During the pilot studies, we observed that novice video editors could intuitively use
ExpressEdit to generate multiple edits based on their natural language and sketch (NL&amp;S)
command. When requesting new edits, the participants first confirmed that the NL&amp;S editing
command was understood correctly with examine panel (Figure 2a) that appears before any
generated edits. Then, the edit result panel (Figure 2b) and ‘turning on and of’ feature allowed
novices to quickly judge and add edits to the video. Furthermore, they were able to manually
revise the generated edits by adjusting the edit operation and parameters. Additionally, the
participants appreciated the convenience of applying edits with NL&amp;S which efectively alleviated
the manual eforts of making edits from scratch and allowed them to try out several edit ideas
eficiently.</p>
      <p>Since the pilot study lasted only 25 minutes, the participants made a few NL&amp;S requests and
spent the rest of the time applying the generated edits. Thus, we did not observe any notable
use of the summary panel (Figure 2c), probably, due to its role of facilitating the management
of many edits and the navigation between several NL&amp;S requests. Additionally, because of
the iterative nature of video editing, the participants had to constantly scroll through multiple
panels whenever they wanted to turn on or of a specific generated edit. Since this behavior
5The video used for pilot study: How To SURVIVE As An Entrepreneur
made them feel less eficient, they wished for more efective support for navigation between
generated edits.</p>
    </sec>
    <sec id="sec-7">
      <title>6. Conclusion</title>
      <p>
        We propose ExpressEdit, a multimodal video editing system that allows users to edit videos
using natural language (NL) text and sketching on top of a video frame. The design of our
system is motivated by findings from the formative study and the analysis of 176 multimodal
expressions of edit requests. ExpressEdit can comprehend video editing commands and generate
edits based on a technical pipeline comprised of CV and large language models that extracts
and interprets (
        <xref ref-type="bibr" rid="ref1">1</xref>
        ) temporal, (
        <xref ref-type="bibr" rid="ref2">2</xref>
        ) spatial, and (
        <xref ref-type="bibr" rid="ref3">3</xref>
        ) operational references in an NL editing command
and spatial references from sketching. Early feedback from the pilot study (N=3) suggests that
ExpressEdit has the potential to greatly facilitate the expression of video editing intents for
novice video editors.
      </p>
    </sec>
    <sec id="sec-8">
      <title>Acknowledgments</title>
      <p>This work was supported by Institute of Information &amp; Communications Technology Planning
&amp; Evaluation (IITP) grant funded by the Korea government (MSIT) (No.2021-0-01347,Video
Interaction Technologies Using Object-Oriented Video Modeling).
Machinery, New York, NY, USA, 2014, pp. 4017–4026. URL: https://dl.acm.org/doi/10.1145/
2556288.2556986. doi:10.1145/2556288.2556986.
[6] P. J. Guo, J. Kim, R. Rubin, How video production afects student engagement: an empirical
study of MOOC videos, in: Proceedings of the first ACM conference on Learning @ scale
conference, L@S ’14, Association for Computing Machinery, New York, NY, USA, 2014,
pp. 41–50. URL: https://dl.acm.org/doi/10.1145/2556325.2566239. doi:10.1145/2556325.
2566239.
[7] M. Bétrancourt, K. Benetos, Why and when does instructional video facilitate learning? A
commentary to the special issue “developments and trends in learning with instructional
video”, Computers in Human Behavior 89 (2018) 471–475. URL: https://www.sciencedirect.
com/science/article/pii/S0747563218304102. doi:10.1016/j.chb.2018.08.035.
[8] T. Jokela, K. Mäkelä, M. Karukka, Empirical observations on video editing in the
mobile context, in: Proceedings of the 4th international conference on mobile technology,
applications, and systems and the 1st international symposium on Computer human
interaction in mobile technology, Mobility ’07, Association for Computing Machinery, New
York, NY, USA, 2007, pp. 482–489. URL: https://dl.acm.org/doi/10.1145/1378063.1378140.
doi:10.1145/1378063.1378140.
[9] G. Chandler, Cut by cut: editing your film or video, Michael Wiese Productions, Studio</p>
      <p>City, CA, 2004.
[10] A. Kirillov, E. Mintun, N. Ravi, H. Mao, C. Rolland, L. Gustafson, T. Xiao, S. Whitehead,
A. C. Berg, W.-Y. Lo, P. Dollár, R. Girshick, Segment Anything, 2023. URL: http://arxiv.org/
abs/2304.02643. doi:10.48550/arXiv.2304.02643, arXiv:2304.02643 [cs].
[11] Y. Wang, K. Li, Y. Li, Y. He, B. Huang, Z. Zhao, H. Zhang, J. Xu, Y. Liu, Z. Wang, S. Xing,
G. Chen, J. Pan, J. Yu, Y. Wang, L. Wang, Y. Qiao, InternVideo: General Video Foundation
Models via Generative and Discriminative Learning, 2022. URL: https://arxiv.org/abs/2212.
03191v2.
[12] J. Li, D. Li, S. Savarese, S. Hoi, BLIP-2: Bootstrapping Language-Image Pre-training with
Frozen Image Encoders and Large Language Models, 2023. URL: http://arxiv.org/abs/2301.
12597. doi:10.48550/arXiv.2301.12597, arXiv:2301.12597 [cs].
[13] T. B. Brown, B. Mann, N. Ryder, M. Subbiah, J. Kaplan, P. Dhariwal, A. Neelakantan,
P. Shyam, G. Sastry, A. Askell, S. Agarwal, A. Herbert-Voss, G. Krueger, T. Henighan,
R. Child, A. Ramesh, D. M. Ziegler, J. Wu, C. Winter, C. Hesse, M. Chen, E. Sigler, M. Litwin,
S. Gray, B. Chess, J. Clark, C. Berner, S. McCandlish, A. Radford, I. Sutskever, D. Amodei,
Language Models are Few-Shot Learners, 2020. URL: https://arxiv.org/abs/2005.14165v4.
[14] OpenAI, GPT-4 Technical Report, 2023. URL: http://arxiv.org/abs/2303.08774. doi:10.</p>
      <p>48550/arXiv.2303.08774, arXiv:2303.08774 [cs].
[15] A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell,
P. Mishkin, J. Clark, G. Krueger, I. Sutskever, Learning Transferable Visual Models From
Natural Language Supervision, 2021. URL: http://arxiv.org/abs/2103.00020. doi:10.48550/
arXiv.2103.00020, arXiv:2103.00020 [cs].</p>
    </sec>
    <sec id="sec-9">
      <title>A. Overview of the pipeline</title>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>L.</given-names>
            <surname>Fiorella</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R. E.</given-names>
            <surname>Mayer</surname>
          </string-name>
          ,
          <article-title>What works and doesn't work with instructional video</article-title>
          ,
          <source>Computers in Human Behavior</source>
          <volume>89</volume>
          (
          <year>2018</year>
          )
          <fpage>465</fpage>
          -
          <lpage>470</lpage>
          . URL: https://www.sciencedirect.com/science/article/ pii/S0747563218303376. doi:
          <volume>10</volume>
          .1016/j.chb.
          <year>2018</year>
          .
          <volume>07</volume>
          .015.
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>P. T.</given-names>
            <surname>Hove</surname>
          </string-name>
          ,
          <article-title>Characteristics of instructional videos for conceptual knowledge development, 2014</article-title>
          . URL: https://www. semanticscholar.org/paper/Characteristics-of
          <article-title>-instructional-videos-for-</article-title>
          <string-name>
            <surname>Hove</surname>
          </string-name>
          /
          <year>c377da3ea8c08dbe79cd36927b25154ecb51cb48</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>A.</given-names>
            <surname>Truong</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Berthouzoz</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <surname>M.</surname>
          </string-name>
          <article-title>Agrawala, QuickCut: An Interactive Tool for Editing Narrated Video</article-title>
          ,
          <source>in: Proceedings of the 29th Annual Symposium on User Interface Software and Technology, UIST '16</source>
          ,
          <string-name>
            <surname>Association</surname>
          </string-name>
          for Computing Machinery, New York, NY, USA,
          <year>2016</year>
          , pp.
          <fpage>497</fpage>
          -
          <lpage>507</lpage>
          . URL: https://dl.acm.org/doi/10.1145/2984511.2984569. doi:
          <volume>10</volume>
          .1145/ 2984511.2984569.
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>P.-Y.</given-names>
            <surname>Chi</surname>
          </string-name>
          , J. Liu,
          <string-name>
            <given-names>J.</given-names>
            <surname>Linder</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Dontcheva</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <surname>B. Hartmann,</surname>
          </string-name>
          <article-title>DemoCut: generating concise instructional videos for physical demonstrations, in: Proceedings of the 26th annual ACM symposium on User interface software and technology</article-title>
          ,
          <source>UIST '13</source>
          ,
          <string-name>
            <surname>Association</surname>
          </string-name>
          for Computing Machinery, New York, NY, USA,
          <year>2013</year>
          , pp.
          <fpage>141</fpage>
          -
          <lpage>150</lpage>
          . URL: https://dl.acm.org/ doi/10.1145/2501988.2502052. doi:
          <volume>10</volume>
          .1145/2501988.2502052.
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>J.</given-names>
            <surname>Kim</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P. T.</given-names>
            <surname>Nguyen</surname>
          </string-name>
          , S. Weir,
          <string-name>
            <given-names>P. J.</given-names>
            <surname>Guo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R. C.</given-names>
            <surname>Miller</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K. Z.</given-names>
            <surname>Gajos</surname>
          </string-name>
          ,
          <article-title>Crowdsourcing step-by-step information extraction to enhance existing how-to videos</article-title>
          ,
          <source>in: Proceedings of the SIGCHI Conference on Human Factors in Computing Systems, CHI '14</source>
          ,
          <string-name>
            <surname>Association</surname>
          </string-name>
          for Computing
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>