<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <journal-title-group>
        <journal-title>March</journal-title>
      </journal-title-group>
    </journal-meta>
    <article-meta>
      <title-group>
        <article-title>Human-AI Co-Creation of Worked Examples for Programming Classes</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Mohammad Hassany</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Peter Brusilovsky</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Jiaze Ke</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Kamil Akhuseyinoglu</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Arun Balajiee Lekshmi Narayanan</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Carnegie Mellon University</institution>
          ,
          <addr-line>Pittsburgh, PA, 15213</addr-line>
          ,
          <country country="US">USA</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>University of Pittsburgh</institution>
          ,
          <addr-line>Pittsburgh, PA, 15260</addr-line>
          ,
          <country country="US">USA</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2024</year>
      </pub-date>
      <volume>1</volume>
      <fpage>8</fpage>
      <lpage>21</lpage>
      <abstract>
        <p>Worked examples (solutions to typical programming problems presented as a source code in a certain language and are used to explain the topics from a programming class) are among the most popular types of learning content in programming classes. Most approaches and tools for presenting these examples to students are based on line-by-line explanations of the example code. However, instructors rarely have time to provide line-by-line explanations for a large number of examples typically used in a programming class. In this paper, we explore and assess a human-AI collaboration approach to authoring worked examples for Java programming. We introduce an authoring system for creating Java worked examples that generates a starting version of code explanations and presents it to the instructor to edit if necessary. We also present a study that assesses the quality of explanations created with this approach.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;Code Examples</kwd>
        <kwd>Authoring Tool</kwd>
        <kwd>Human-AI Collaboration</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>
        Program code examples play a crucial role in learning how to program [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. Instructors use
examples extensively to demonstrate the semantics of the programming language being taught
and to highlight the fundamental coding patterns. Programming textbooks also pay a lot of
attention to examples, with a considerable textbook space allocated to program examples and
associated comments [
        <xref ref-type="bibr" rid="ref2 ref3">2, 3</xref>
        ]. A typical worked example presents a code for solving a specific
programming problem and explains the role and function of code lines or code chunks. In
textbooks, these explanations are usually presented as comments in the code or as explanations
on the margins. While informative, this approach focused on passive learning, which is known
for its low eficiency. Recognizing this problem, several research teams developed learning tools
that ofered more interactive and engaging ways to learn from examples [
        <xref ref-type="bibr" rid="ref4 ref5 ref6 ref7 ref8">4, 5, 6, 7, 8</xref>
        ].
      </p>
      <p>
        The example-focused learning tools demonstrated their efectiveness in classroom studies, but
their use by programming instructors is still limited due to the insuficient number of worked
examples ofered by these tools. Although the authors of these tools usually provide a good
set of worked examples that can be presented through their tools, many instructors prefer to
use their own favorite code examples. The instructors are usually happy to broadly share the
code of examples they created (usually providing it on the course web page), but they rarely
have time or patience to augment examples with explanations and add their examples to an
example-focused interactive system. Indeed, producing a single explained example could take
30 minutes or more, since it requires typing an explanation for each code line [
        <xref ref-type="bibr" rid="ref4 ref8">4, 8</xref>
        ] or creating
a screencast in a specific format [
        <xref ref-type="bibr" rid="ref5 ref7">5, 7</xref>
        ].
      </p>
      <p>
        This issue has been recognized by several research teams that have ofered several ways to
address the lack of content. Among the approaches explored are learner-sourcing, that is, engaging
students in creating and reviewing explanations for instructor-provided code [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ] and automatic
extraction of information content from available sources, such as lecture recordings [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ]. In this
paper, we present an alternative approach to address the lack of worked examples based on
human-AI collaboration. With this approach, the instructor provides the code of one of their
favorite examples along with the statement of the programming problem it is solving. The AI
engine based on large language models (LLM) examines the code and generates explanations for
each code line. The explanations could be reviewed and edited by the instructor. To support and
explore this authoring approach, we created an authoring system, which radically decreases the
time to create a new interactive worked example. The examples created by the system could be
uploaded to an example-exploration system such as WebEx [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ] or PCEX [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ] or exported in a
reusable format. To assess the quality of the resulting examples, we performed a user study in
which TAs and students compared code explanations created by experts through a traditional
process with examples created by AI to contribute to human-AI collaborative process.
      </p>
      <p>The remainder of the paper is structured as following. We start by reviewing related work,
introduce the example authoring system that implements the proposed collaborative approach,
and explain how specific design decisions were made through several rounds of internal
evaluation. Next, we explain the design of our user study and review its results. We conclude with a
summary of the work and plans for future research.</p>
    </sec>
    <sec id="sec-2">
      <title>2. Related Work</title>
      <sec id="sec-2-1">
        <title>2.1. Worked Examples in Programming</title>
        <p>
          Code examples are important pedagogical tools for learning programming. Not surprisingly,
considerable eforts have been devoted to the development of learning materials and tools to
support students in studying code examples. For many years, the state-of-the-art approach
for presenting worked code examples in online tools was simply code text with comments
[
          <xref ref-type="bibr" rid="ref1 ref10 ref11">1, 10, 11</xref>
          ]. More recently, this approach has been enhanced with multimedia by adding audio
narrations to explain the code [
          <xref ref-type="bibr" rid="ref12">12</xref>
          ] or by showing video fragments of code screencasts with
the instructor’s narration being heard while watching code in slides or an editor window [
          <xref ref-type="bibr" rid="ref5 ref6">5, 6</xref>
          ].
Both ways, however, support passive learning, which is the least eficient approach from the
prospect of the ICAP framework [
          <xref ref-type="bibr" rid="ref13">13</xref>
          ]1
        </p>
        <p>
          An attempt to make learning from program construction examples active was made in the
WebEx system, which allowed students to interactively explore instructor-provided
line-byline comments for program examples via a web-based interface [
          <xref ref-type="bibr" rid="ref4">4</xref>
          ]. More recently, several
projects [
          <xref ref-type="bibr" rid="ref6 ref7 ref8">6, 7, 8</xref>
          ] augmented examples with simple problems and other constructive activities
to elevate the example study process to the interactive and constructive levels of the ICAP
framework, known as the most pedagogically eficient.
        </p>
        <p>
          A good example of a modern interactive tool for studying code examples is the PCEX
system [
          <xref ref-type="bibr" rid="ref8">8</xref>
          ]. PCEX (Program Construction EXamples) was created in the context of an NSF
Infrastructure project (https://cssplice.org) with a focus on broad reuse and has been used by
several universities in the US and Europe in the context of Java, Python, and SQL courses. PCEX
interface (Figure 1) provides interactive access to traditionally organized worked examples, i.e.,
code lines augmented with instructor’s explanations. Separating explanations (Figure 1-3) from
the code (Figure 1-2), allows students to selectively study explanations for code lines they want.
Explanations are provided on several levels of detail, so more details could be requested if the
brief explanation is not suficient (Figure 1-3).
        </p>
        <p>
          Since line-by-line multi-level example explanations ofered by PCEX is currently the most
detailed approach for explaining worked examples, we selected the code example structure
implemented by PCEX as the target model for our authoring tool presented in this paper. The
tool produces code augmented with line-by-line explanations on several levels of detail. The
resulting example could be directly uploaded to PCEX or exported in a system-independent
format to be uploaded to other example exploration systems like WebEx [
          <xref ref-type="bibr" rid="ref4">4</xref>
          ].
        </p>
      </sec>
      <sec id="sec-2-2">
        <title>2.2. Use of LLMs for Code Explanations</title>
        <p>
          Several research teams explored the use of LLM for code explanations using GPT-3 [
          <xref ref-type="bibr" rid="ref14">14, 15, 16</xref>
          ],
GPT-3.5 [15, 17, 18], GPT-4 [17], OpenAI Codex [19, 20, 15], and GitHub Copilot [18]. LLMs were
used to generate explanations at diferent levels of abstraction (line-by-line, step-by-step, and
high-level summary). Sarsa et al. [19] observed that ChatGPT can generate better explanations at
low-level (lines). Explanations and summaries generated by these LLMs were mostly evaluated
by authors [19], students [15, 16], and tool users [18]. Sarsa et al. [19] reported a high correct
ratio for generated explanations with minor mistakes that can be resolved by the instructor
or teaching assistant. Students rated LLM-generated explanations as being useful, easier, and
more accurate than learner–sourced explanations [16].
        </p>
        <p>
          Since prompts directly influence the LLM’s performance, several studies focused on exploring
diferent prompting strategies [ 21, 22]. Tian et al [20] reported that a verbose prompt will
limit the LLM’s ability to utilize its knowledge [20]. Iterative prompts are proven to perform
well [
          <xref ref-type="bibr" rid="ref14">14</xref>
          ]. Zamfirescu-Pereira et al. [
          <xref ref-type="bibr" rid="ref14">14</xref>
          ] observed that non-experts have misconceptions about
LLMs and struggle to come up with a well-formed prompt. Researchers believe that LLMs can
be beneficial in environments where humans and AI can work together, where the human can
perform the expert evaluation and tune the responses generated by the AI while the AI performs
the time-consuming manual tasks [22].
1The ICAP framework diferentiates four modes of engagement, behaviorially exhibited by learners: passive, active,
constructive and interactive.
        </p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>3. The Feasibility Studies</title>
      <p>To assess the feasibility of Human-AI co-creation of worked examples, we performed three
rounds of preliminary studies. The purpose of these studies was to develop an approach for
producing LLM code explanations of reasonable quality, compare the explanations produced by
LLMs with the explanations produced by humans, and assess whether the LLM explanations
are considered satisfactory by instructors and students.</p>
      <p>In the first study [ 23] guided by earlier work on LLM code explanations reviewed above,
we explored a range of prompts and performed an evaluation of the quality of explanations
generated by the prompts to select the best-performing prompt for the next rounds of our work.</p>
      <p>In the second study [24], we used a dataset of explanations produced by two experts and 60
students for the same four Java code examples with 33 explainable lines to compare ChatGPT
explanations with explanations produced by experts and students using several formal metrics.
To make this comparison, we generated ChatGPT explanations using our selected prompt for
the 33 explainable lines four times, using temperature 0 once and temperature 1 three times.
To calculate all comparison metrics, we merged all line explanations generated by each source
(i.e, each expert, each student, and each round of ChatGPT generation) into a single source
document. As the data shows (Table 1), the explanations produced by ChatGPT have comparable
length (measured by the number of tokens) and lexical density with the explanations produced
by experts, while the explanations produced by students were more than twice as short and
more lexically dense than the explanations produced by the other two sources. Surprisingly
(given the length diference) the readability of explanations produced by experts is very similar
to the readability of student explanations, while ChatGPT explanations are much less readable.
Expert explanations are also much more similar than ChatGPT explanations to the explanations
produced by students (Table 2). This data could be partially explained by the considerably larger
vocabulary used by ChatGPT even in comparison to experts.</p>
      <p>Source</p>
      <p>Experts
ChatGPT*
Students</p>
      <p>N
2
4
60</p>
      <p>In the third study [23], we conducted a comparative evaluation of explanations produced by
experts and ChatGPT from the point of view of human users. We used two types of human
users: authors (instructors and TAs) who are expected to use ChatGPT-generated explanations
as the starting point in the co-creation process, and students who are the target users of the
co-created product. Explanations were compared in pairs, each explanation in a pair has to be
judged by completeness, and the best explanation in the pair has to be selected. A pair included
an expert and a ChatGPT explanation, and the judges were not aware of which source produced
each explanation. The study results indicated strong preferences for ChatGPT in both groups of
judges (Table 3). In general, ChatGPT explanations were rated as more complete and judged
to be better in the majority of cases. However, it was not a clear win. In a substantial number
of cases (15.05% for students and 27.41% for authors), expert explanations were selected as the
best option in a pair.</p>
      <p>Taking the results of these two studies together, we could conclude that producing
explanations for code examples is a promising application area for Human-AI co-creation. On the
one hand, the LLM-generated explanations are lagging behind expert explanations in several
aspects. ChatGPT explanations have higher reading dificulty than expert explanations, and
they are further away from the students’ own explanations, as measured by most similarity
metrics. The vocabulary data hints that ChatGPT tends to use terms, which might not be easy
for the students to understand, while experts have experience in phrasing their explanations
closer to the students’ active vocabulary. On the other hand, the explanations produced by
ChatGPT were generally rated higher than the expert explanations by both instructors and
students. These data hint that presenting ChatGPT explanations directly to students might not
be a perfect solution, but they can serve as an excellent starting point for instructors in shaping
their own explanations. Following that, we decided to structure the Human-AI collaboration in
creating working examples as follows. Instructors have the ultimate control over producing
explanations. Depending on the context (such as example complexity), they can either choose
to explain example lines themselves or request AI (LLM) help in producing explanations for
specific lines. In the latter case, LLM generates the initial line explanations leaving it to the
instructor to accept or reject it and, if accepted, to further edit the explanation text to satisfaction.
The Human-AI co-creation interface presented in the next section is based on this model of
collaboration.</p>
    </sec>
    <sec id="sec-4">
      <title>4. The Human-AI Co-Creation Interface Design</title>
      <p>
        On the basis of our feasibility studies, we developed a Worked Example Authoring Tool (WEAT).
WEAT enables instructors to create worked code examples for PCEX system, [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ] through the
human-AI co-creation interface. In this co-creation process, the main task of a human author
is to provide the code of the example and the statement of the problem that the code solves.
The main task of ChatGPT is to generate the bulk of code line explanations on several levels of
detail. As an option, a human author could edit and refine the text produced by ChatGPT to
adapt it to the class goals and target students. As in any productive collaboration, each side
does what it is best suited to do, leaving the rest to the partner.
      </p>
      <p>In the main part of the WEAT interface, the problem (Figure 2-1) and the code (Figure 2-2)
have to be provided by the instructor, while the explanations for each line (Figure 2-3) can
be created by the instructor or generated by ChatGPT. The generated explanations could be
further edited by the instructor. While we expect that co-creation of code explanations will be
the preferred way to use WEAT, the system supports the whole range of options from using AI
explanations without human editing to creating the whole example from scratch, without the
help of AI. Authors who want to start by creating explanations themselves could simply select a
code line to explain (Figure 2-2) and add one or more explanation fragments to this line (Figure
2-3). The order of the fragments is important: the first fragment is displayed in PCEX when
the line is clicked, while the remaining fragments can be accessed by clicking the “Additional
Details” button (Figure 1-3).</p>
      <p>To generate ChatGPT explanations for the provided example code and problem description,
the author has to click the “Generate Explanations” button to open the ChatGPT dialog (Figure
3). In this dialog, the explanations could be generated by clicking “Generate” button and added to
the example by clicking “Use Explanations” button. Experienced authors have the opportunity to
tune the default prompt before generating explanations and review the generated explanations
before using them. Reviewing the generated explanations can be done line by line: selecting
one of the explained lines (marked by “?") in the code box (Figure 3-3) will display all generated
explanations for this line in the explanation box (Figure 3-4). The explanation could be accepted
or rejected by clicking the checkbox next to the “Include this line” prompt.</p>
      <p>To support the review at the finer grain level, WEAT divides the explanations into fragments
that can be independently accepted or rejected by clicking the small green check mark icon next
to the fragment (Figure 3-4a). The author can also click on the small gray thumb-up icon (Figure
3-4b) to provide positive feedback on the explanation fragment. Once the “Use Explanations”
button is clicked, all accepted explanation fragments are added to the corresponding example
lines and can be further edited in the main interface (Figure 2).</p>
    </sec>
    <sec id="sec-5">
      <title>5. Evaluation</title>
      <p>To assess how well WEAT supports co-creation of worked examples, we engaged five instructors
(A1-A5) teaching Java of Python classes and asked them to create one or more worked examples
for PCEX from real examples they use in their classes. To explain the tool to the instructor, we
provided a video tutorial and integrated textual help into WEAT. Their interactions and usage
of the tool were recorded through logs and used for the analysis presented below.</p>
      <p>The instructors used the tool to create 12 examples in total (Table 4). The ChatGPT dialog
was used 21 times, and in 13 cases (A1=6, A2=2, A3=3, A4=1, and A5=1), instructors added
generated explanations to the example by clicking the “Use Explanations” button. As discovered
from an interview with instructors, in several cases they closed and reopened the ChatGPT
dialog to access the main interface blocked by the dialog. Analyzing the interaction logs, we
observed this has been done at least 5 times (3 times with the close-reopen interval of 5 seconds
and 2 with 12 seconds interval) leaving only 16 cases where explanations had a chance to be
examined. In total, 269 explanation fragments were generated for 119 lines of code with an
average of 2.26 fragments per line. In 13 cases where ChatGPT explanations were added to
the example by instructors, ChatGPT generated 237 explanations for 99 lines of code (Table 4).
We found no cases in which the entire set of explanations generated for the line was excluded
by the instructors in its entirety, and among the 237 generated fragments, only 24 (10.12%)
237 were excluded. The interview revealed that in some cases the generated fragments were
rejected not because they were unsatisfactory, but because they were incorrect (Figure 4). On
the other hand, instructors liked 15 (6.32%) explanations.</p>
      <p>After adding explanations to the example, instructors still didn’t remove the explanations for
any line entirely, but removed 23 (9.7%) ChatGPT generated explanation fragments. Instructor A5
reported that he removed several fragments when merging two or more explanation fragments.
Since the tool did not provide support for merging fragments, it did so by copying the explanation
from one fragment to the end of the other fragment and removing the obsolete fragment. In
only 10 cases, instructors attempted to create new explanations from scratch, but in the end
Examples Created
Generated Explanations
Lines of Code being Explained by ChatGPT
Explanations Excluded
Explanations Liked
Explanations Edited
Explanations Removed
these explanations were removed. In other words, all remaining explanation fragments were
originally generated by ChatGPT with some of them being edited later by the instructors.
Apparently, the instructors preferred to edit the explanation fragments rather than create them
from scratch. In total, the instructors edited 66 (27.84%) of ChatGPT generated explanation
fragments, on average 1.4 times (stdev=0.55). Feedback from instructors indicated that most of
their edits involved summarizing, adding missing details, or removing unnecessary parts. Table
4 shows that almost half of the generated fragments were used without being touched, saving a
noticeable amount of instructor time.</p>
      <p>ChatGPT Edited Explanations
ChatGPT Explanation Edits
Average Levenshtein Ratio across all</p>
      <p>Final and Original ChatGPT Explanations 0.435</p>
      <p>A1
29
42</p>
      <p>A2
0
0
1</p>
      <p>A3
11
12
0.833</p>
      <p>A4
0
0
1</p>
      <p>A5
26
39
0.412</p>
      <p>The average Levenshtein edit ratio for ChatGPT-generated explanations (edited and unedited)
is 0.73 (Table 5), indicating a high acceptance rate for generated explanations. This indicator,
however, is somewhat misleading since a portion of ChatGPT-generated explanations were
edited because the first version of the tool evaluated in the study didn’t provide direct support
for reordering and merging the explanations, resulting in copy-pasting the explanations (as
reported by A5 for whom the ratio dropped to 0.412). The table also points out that WEAT was
able to support diferent editing approaches pursued by instructors. Some instructors spent
more time reviewing the generated explanations before adding them to the example (A1), some
prefer adding them to the examples and then evaluating and editing them (A3), while some
used the generated explanations without changes.</p>
    </sec>
    <sec id="sec-6">
      <title>6. Conclusion</title>
      <p>In this paper, we introduce a worked code example authoring tool WEAT that supports
humanAI co-creation in the process of developing such examples. WEAT supports human authors
by using ChatGPT for the generation of line-by-line code explanations and by providing an
interface to integrate this functionality into a balanced authoring process. To the best of our
knowledge, this is the first attempt to develop an authoring tool that produces worked examples
through human-AI collaboration.</p>
      <p>To develop WEAT, we performed several rounds of feasibility studies. These studies supported
the need for a human-AI co-creation in authoring worked examples. As the studies showed, in
the majority of cases, the explanations generated by ChatGPT with a carefully tuned prompt
were positively evaluated by authors and students. However, in a good fraction of cases they
were inferior to the explanations provided by experts. The study also revealed that on average
experts can create explanations that are more easily readable and closer to the explanations
generated by the students themselves. With this data, we hypothesized that human-AI
cocreation could ofer the “best of both worlds” solution where good explanations could be simply
accepted by authors, while inferior or hard-to-understand explanations could be improved.</p>
      <p>An evaluation of WEAT system with five course instructors supported these expectations
and provided strong evidence in favor of co-creation. As the log analysis demonstrated, in many
cases, instructors choose to accept generated explanations without changes, which should have
decreased the time and efort required for example creation. Yet in other cases, the instructor
rejected or edited the generated explanation to achieve the desired quality. In some cases,
explanations were rejected by being simply incorrect, which stresses the importance of human
presence in the authoring process. The interview with authors revealed several cases where
authors acted ineficiently due to specific interface issues, such as blocking the main edit window
by the generation dialog or the lack of tools to move or merge fragments. Now we are using
these observations to develop an improved version of WEAT.</p>
      <p>As the first step towards this important goal, our work has limitations. Most importantly,
the scale of our evaluation is relatively small. Since we targeted real instructors as users in
our evaluation process, we were able to recruit only five qualified subjects. Additionally, since
the study was done at the beginning of the semester when instructors were busy setting up
their classes, they created only 12 examples using this tool. To obtain more reliable data, we
plan a larger-scale semester-long study by engaging instructors to create a variety of worked
examples of varying dificulty and use them in their classes. Such a study will also enable us to
assess the quality of explanations produced through human-AI collaboration and their value for
students in introductory programming classes.
Conference on Human Factors in Computing Systems, CHI ’23, Association for Computing
Machinery, New York, NY, USA, 2023.
[15] S. MacNeil, A. Tran, A. Hellas, J. Kim, S. Sarsa, P. Denny, S. Bernstein, J. Leinonen,
Experiences from using code explanations generated by large language models in a web
software development e-book, in: Proceedings of the 54th ACM Technical Symposium on
Computer Science Education V. 1, SIGCSE 2023, Association for Computing Machinery,
New York, NY, USA, 2023, p. 931–937.
[16] J. Leinonen, P. Denny, S. MacNeil, S. Sarsa, S. Bernstein, J. Kim, A. Tran, A. Hellas,
Comparing code explanations created by students and large language models, 2023.
[17] J. Li, S. Tworkowski, Y. Wu, R. Mooney, Explaining competitive-level programming
solutions using llms, 2023.
[18] E. Chen, R. Huang, H.-S. Chen, Y.-H. Tseng, L.-Y. Li, Gptutor: A chatgpt-powered
programming tool for code explanation, in: N. Wang, G. Rebolledo-Mendez, V. Dimitrova,
N. Matsuda, O. C. Santos (Eds.), Artificial Intelligence in Education. Posters and Late
Breaking Results, Workshops and Tutorials, Industry and Innovation Tracks, Practitioners,
Doctoral Consortium and Blue Sky, Springer Nature Switzerland, Cham, 2023, pp. 321–327.
[19] S. Sarsa, P. Denny, A. Hellas, J. Leinonen, Automatic generation of programming exercises
and code explanations using large language models, in: Proceedings of the 2022 ACM
Conference on International Computing Education Research - Volume 1, ICER ’22, Association
for Computing Machinery, New York, NY, USA, 2022, p. 27–43.
[20] H. Tian, W. Lu, T. O. Li, X. Tang, S.-C. Cheung, J. Klein, T. F. Bissyandé, Is chatgpt the
ultimate programming assistant – how far is it?, 2023.
[21] D. Zhou, N. Scharli, L. Hou, J. Wei, N. Scales, X. Wang, D. Schuurmans, O. Bousquet, Q. Le,
E. H. hsin Chi, Least-to-most prompting enables complex reasoning in large language
models, ArXiv (2022).
[22] J. White, Q. Fu, S. Hays, M. Sandborn, C. Olea, H. Gilbert, A. Elnashar, J. Spencer-Smith,
D. C. Schmidt, A prompt pattern catalog to enhance prompt engineering with chatgpt,
2023.
[23] M. Hassany, P. Brusilovsky, J. Ke, K. Akhuseyinoglu, A. B. Lekshmi Narayanan,
Authoring Worked Examples for Java Programming with Human-AI Collaboration, Report
arXiv:2312.02105, arXiv, 2023. URL: https://doi.org/10.48550/arXiv.2312.02105.
[24] A.-B. Lekshmi-Narayanan, P. Oli, J. Chapagain, M. Hassany, R. Banjade, P. Brusilovsky,
V. Rus, Explaining code examples in introductory programming courses: Llm vs humans,
in: Workshop on AI for Education - Bridging Innovation and Responsibility at AAAI 2024„
2024.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>M. C.</given-names>
            <surname>Linn</surname>
          </string-name>
          ,
          <string-name>
            <surname>M. J. Clancy,</surname>
          </string-name>
          <article-title>The case for case studies of programming problems</article-title>
          ,
          <source>Commun. ACM</source>
          <volume>35</volume>
          (
          <year>1992</year>
          )
          <fpage>121</fpage>
          -
          <lpage>132</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>H. M.</given-names>
            <surname>Deitel</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P. J.</given-names>
            <surname>Deitel</surname>
          </string-name>
          , C How to Program,
          <source>2nd Edition</source>
          , Prentice Hall, New York,
          <year>1994</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>A.</given-names>
            <surname>Kelley</surname>
          </string-name>
          , I. Pohl,
          <string-name>
            <surname>C by</surname>
          </string-name>
          <article-title>Dissection : The Essentials of C Programming, Addison-</article-title>
          <string-name>
            <surname>Wesley</surname>
          </string-name>
          , New York,
          <year>1995</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>P.</given-names>
            <surname>Brusilovsky</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. V.</given-names>
            <surname>Yudelson</surname>
          </string-name>
          ,
          <string-name>
            <given-names>I.-H.</given-names>
            <surname>Hsiao</surname>
          </string-name>
          ,
          <article-title>Problem solving examples as first class objects in educational digital libraries: Three obstacles to overcome</article-title>
          ,
          <source>Journal of Educational Multimedia and Hypermedia</source>
          <volume>18</volume>
          (
          <year>2009</year>
          )
          <fpage>267</fpage>
          -
          <lpage>288</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>R.</given-names>
            <surname>Sharrock</surname>
          </string-name>
          , E. Hamonic,
          <string-name>
            <given-names>M.</given-names>
            <surname>Hiron</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Carlier</surname>
          </string-name>
          ,
          <string-name>
            <surname>Codecast:</surname>
          </string-name>
          <article-title>An innovative technology to facilitate teaching and learning computer programming in a c language online course</article-title>
          ,
          <source>Proceedings of the Fourth</source>
          (
          <year>2017</year>
          ) ACM Conference on Learning @ Scale (
          <year>2017</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>K.</given-names>
            <surname>Khandwala</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P. J.</given-names>
            <surname>Guo</surname>
          </string-name>
          ,
          <article-title>Codemotion: expanding the design space of learner interactions with computer programming tutorial videos</article-title>
          ,
          <source>Proceedings of the Fifth Annual ACM Conference on Learning at Scale</source>
          (
          <year>2018</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>J.</given-names>
            <surname>Park</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y. H.</given-names>
            <surname>Park</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Kim</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Cha</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Kim</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A. H.</given-names>
            <surname>Oh</surname>
          </string-name>
          ,
          <article-title>Elicast: embedding interactive exercises in instructional programming screencasts</article-title>
          ,
          <source>Proceedings of the Fifth Annual ACM Conference on Learning at Scale</source>
          (
          <year>2018</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>R.</given-names>
            <surname>Hosseini</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Akhuseyinoglu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Brusilovsky</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Malmi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Pollari-Malmi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Schunn</surname>
          </string-name>
          , T. Sirkiä,
          <article-title>Improving engagement in program construction examples for learning python programming</article-title>
          ,
          <source>International Journal of Artificial Intelligence in Education</source>
          <volume>30</volume>
          (
          <year>2020</year>
          )
          <fpage>299</fpage>
          -
          <lpage>336</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>I.-H.</given-names>
            <surname>Hsiao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Brusilovsky</surname>
          </string-name>
          ,
          <article-title>The role of community feedback in the student example authoring process: an evaluation of annotex</article-title>
          ,
          <source>British Journal of Educational Technology</source>
          <volume>42</volume>
          (
          <year>2011</year>
          )
          <fpage>482</fpage>
          -
          <lpage>499</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>A.</given-names>
            <surname>Davidovic</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. R.</given-names>
            <surname>Warren</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Trichina</surname>
          </string-name>
          ,
          <article-title>Learning benefits of structural example-based adaptive tutoring systems</article-title>
          ,
          <source>IEEE Trans. Educ</source>
          .
          <volume>46</volume>
          (
          <year>2003</year>
          )
          <fpage>241</fpage>
          -
          <lpage>251</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>B. B.</given-names>
            <surname>Morrison</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L. E.</given-names>
            <surname>Margulieux</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Ericson</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Guzdial</surname>
          </string-name>
          ,
          <article-title>Subgoals help students solve parsons problems</article-title>
          ,
          <source>Proceedings of the 47th ACM Technical Symposium on Computing Science Education</source>
          (
          <year>2016</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>B.</given-names>
            <surname>Ericson</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Guzdial</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B. B.</given-names>
            <surname>Morrison</surname>
          </string-name>
          ,
          <article-title>Analysis of interactive features designed to enhance learning in an ebook</article-title>
          ,
          <source>Proceedings of the eleventh annual International Conference on International Computing Education Research</source>
          (
          <year>2015</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <surname>M. T. H. Chi</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          <string-name>
            <surname>Adams</surname>
            ,
            <given-names>E. B.</given-names>
          </string-name>
          <string-name>
            <surname>Bogusch</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          <string-name>
            <surname>Bruchok</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          <string-name>
            <surname>Kang</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          <string-name>
            <surname>Lancaster</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          <string-name>
            <surname>Levy</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          <string-name>
            <surname>Li</surname>
            ,
            <given-names>K. L.</given-names>
          </string-name>
          <string-name>
            <surname>McEldoon</surname>
            ,
            <given-names>G. S.</given-names>
          </string-name>
          <string-name>
            <surname>Stump</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          <string-name>
            <surname>Wylie</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          <string-name>
            <surname>Xu</surname>
            ,
            <given-names>D. L.</given-names>
          </string-name>
          <string-name>
            <surname>Yaghmourian</surname>
          </string-name>
          ,
          <article-title>Translating the icap theory of cognitive engagement into practice</article-title>
          ,
          <source>Cognitive Science 42</source>
          (
          <year>2018</year>
          )
          <fpage>1777</fpage>
          -
          <lpage>1832</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <given-names>J.</given-names>
            <surname>Zamfirescu-Pereira</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R. Y.</given-names>
            <surname>Wong</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Hartmann</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Q.</given-names>
            <surname>Yang</surname>
          </string-name>
          ,
          <article-title>Why johnny can't prompt: How non-ai experts try (and fail) to design llm prompts</article-title>
          ,
          <source>in: Proceedings of the 2023 CHI</source>
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>