<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>ming⋆</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Qianou Christina Ma</string-name>
          <email>qianouma@cmu.edu</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Sherry Tongshuang Wu</string-name>
          <email>sherryw@cs.cmu.edu</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Ken Koedinger</string-name>
          <email>koedinger@cmu.edu</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Carnegie Mellon University (CMU)</institution>
          ,
          <addr-line>Pittsburgh, PA</addr-line>
          ,
          <country country="US">United States</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2023</year>
      </pub-date>
      <abstract>
        <p>The emergence of large-language models (LLMs) that excel at code generation and commercial products such as GitHub's Copilot has sparked interest in human-AI pair programming (referred to as “pAIr programming”) where an AI system collaborates with a human programmer. While traditional pair programming between humans has been extensively studied in both industry and education, it remains uncertain whether its findings can be applied to human-AI pair programming. We compare interaction, measures, benefits, and challenges of human-human and human-AI pair programming. We find that the efectiveness of both approaches is mixed in the literature (the measures used for pAIr programming are not as comprehensive). We summarize moderating factors on the success of human-human pair programming, which provide opportunities for pAIr programming. For example, mismatched expertise makes pair programming less productive, therefore well-designed AI programming assistants may adapt to diferences in expertise levels. Finally, we discuss the potential of using LLMs to provide efective pAIr programming learning for students at scale.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>
        Pair programming describes the practice of two programmers working together on the same
task using a single computer, keyboard, and mouse. One programmer in the pair, the “driver,”
performs the coding (typing) and implements the task, while the other programmer, the
“navigator,” aids in planning, reviewing, debugging, and suggesting improvements and alternatives.
Now, pair programming is used in a wide range of settings, including education, industry,
and open-source software development [
        <xref ref-type="bibr" rid="ref1">1, 2</xref>
        ]. For example, in education, human-human pair
programming has been adopted from K12 [3], CS1 [4] to higher-level project-based courses [5].
      </p>
      <p>Recent advances in code-generating large-language models (LLMs) have led to the widespread
popularity of commercial AI-powered programming assistance tools such as GitHub Copilot
[6], which advertises itself as “your AI pair programmer.” Instead of two humans working on a
single computer, it is the programmer and the LLM-based AI that work together on the same
task. The shift in the paradigm raises the questions: Is the AI programming partner comparable
to a human pair programmer? Can they achieve similar or better performance, and should people
interact with them in the same way?</p>
      <p>The question of whether AI can serve as a better programming partner is crucial.
Understanding the comparative performance of AI and human programmers in a pair programming
context can guide developers and educators in utilizing the most efective collaboration methods.
Understanding the potential role and design of AI in pair programming can help educators
design more scalable pedagogical approaches to promote student learning and engagement
in programming. Furthermore, identifying the strengths and weaknesses of human and AI
programming partners may in turn contribute to the refinement and development of better AI
programming tools that augment human programmers’ capabilities.</p>
      <p>Based on our readings of four existing meta-analysis papers on human-human pair
programming and over 50 studies on the topic of pair programming or AI-assisted programming, we
dive into comparisons of measurements of success (Section 2), as well as moderators, e.g., pair
compatibility factors like expertise (Section 3). We find that (1) prior work on both pair
programming paradigms has observed mixed results in quality, productivity, satisfaction, learning,
and cost, (2) human-AI pair programming has yet to develop comprehensive measurements, and
(3) key factors to pAIr’s success have been largely unexplored.</p>
      <p>Building on our exploration, we elaborate on future opportunities for developing best practices
and guidelines for pAIr programming (Section 4). First, we argue that moderating factors that
bring challenges to human-human pair programming (e.g., compatibility and communication)
unveil opportunities to improve human-AI pair programming. It can be promising to exploit
the diferences between a human and an AI partner (e.g., more customizable expertise level
and more adaptable communication styles) to design for more successful pAIr programming
experiences. Second, we encourage future research to explore the best deployment environment
for pAIr programming, such as education. We hope this paper can inspire better evaluations
and designs of code-generating LLMs as a pAIr programmer, especially for students.</p>
    </sec>
    <sec id="sec-2">
      <title>2. Mixed Outcomes</title>
      <p>
        Literature reviews have suggested various benefits as well as mixed efects of human-human pair
programming [
        <xref ref-type="bibr" rid="ref1">29, 1, 2</xref>
        ]. According to Alves De Lima Salge and Berente [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ], pair programming
improves code quality, productivity, and learning outcomes. However, according to Hannay
et al. [29], pair programming improves quality and shortens duration, but it increases efort,
higher quality comes at the expense of considerably greater efort, and reduced completion time
comes with lower quality. In the education context, pair programming brings benefits including
higher quality software, student confidence in solutions, increased assignment grades, exam
scores, success/passing rates in introductory courses, and retention [
        <xref ref-type="bibr" rid="ref7">2, 30, 19</xref>
        ]. All the reviews
acknowledged that even though meta-analyses can show a significant efect size, individual
studies could report contradictory outcomes (see examples in Table 1).
      </p>
      <p>For human-AI pair programming, existing works mainly focus on quality, productivity, and
satisfaction, and already demonstrated mixed results in quality and productivity [9, 31, 12] (see
examples in Table 1). Additionally, some measures are arguably too simplified as evaluation
metrics. For example, Imai [9] used the number of lines of added code as the measure of
Outcomes
Quality
Productivity
Satisfaction
Learning
Cost
Moderators
Task Types
&amp; Complexity
Communication
Collaboration</p>
      <p>Logistics
Comparison of Outcome Variables and Moderators for Human-Human Pair Programming vs. Human-AI
pAIr Programming</p>
      <p>Human-Human vs. Human Solo</p>
      <p>Human-AI (Copilot)
significantly lower defect density for complex
code [7]
no diference for simpler code [ 7]
significantly higher percentage of test cases
passed [8]</p>
      <p>significantly fewer lines of code per person
hour writing simpler code [7]</p>
      <p>no significant diference writing more
complex code [7]</p>
      <p>
        29% shorter time to complete task (pair speed
advantage = 1.4) [13]
higher self-ratings of satisfaction [
        <xref ref-type="bibr" rid="ref2">14</xref>
        ]
students with greater self-confidence and
self-eficacy less enjoy the pair programming
experience [
        <xref ref-type="bibr" rid="ref3">15</xref>
        ]
      </p>
      <p>
        higher grades, exam scores [
        <xref ref-type="bibr" rid="ref6">18</xref>
        ], and
retention [
        <xref ref-type="bibr" rid="ref7">19</xref>
        ]
      </p>
      <p>
        significantly higher gains in exam
performance in female students than male students
[
        <xref ref-type="bibr" rid="ref8">20</xref>
        ]
      </p>
      <p>
        increased management workload to match,
schedule a pair, resolve collaboration conflict,
assess individual contributions, etc. [
        <xref ref-type="bibr" rid="ref9">21</xref>
        ]
      </p>
      <p>reduced teaching staf workload (grading one
assignment from a pair) [8]
Human-Human vs. Human Solo</p>
      <p>vs. Human-Human: more lines of code deleted in next
session (lower quality) [9]</p>
      <p>vs. Human Solo: significantly improve correctness score
and reduce encountered errors for novice students [10]</p>
      <p>vs. Human Solo: no significant diference in task success
[11] or task success rate in given time [12]
vs. Human-Human: more lines of added code [9]
vs. Human Solo: 55.8% reduction in completion time [11]
vs. Human Solo: significantly increase task completion and
reduce task completion time for novice students [10]</p>
      <p>vs. Human Solo: no significant diference in the task
completion rate in given time [12]</p>
      <p>vs. Human Solo: higher self-ratings of satisfaction [12, 16,</p>
      <p>
        vs. Human Solo: no significant diference in immediate
and retention post-test performance of novices, students with
more prior experiences have more learning gains from AI code
generator [10]
No experiment yet. Vaithilingam et al. [12], Bird et al. [
        <xref ref-type="bibr" rid="ref4">16</xref>
        ]
hypothesized that human-AI may lead to more unnecessary
debugging vs. Human Solo
Complex task improve quality, simple one does not [7]; debugging is perceived as less
enjoyable or efective than comprehension or refactoring [
        <xref ref-type="bibr" rid="ref10">22</xref>
        ]
Compatibility Random pairing led to incompatible partners and conflicts during work [
        <xref ref-type="bibr" rid="ref6">18</xref>
        ]. Expertise:
(E.g., Expertise) improve quality more efectively if pair is similarly skilled [
        <xref ref-type="bibr" rid="ref2">14</xref>
        ]; less-skilled students
learn more and enjoy more [
        <xref ref-type="bibr" rid="ref10 ref8">20, 22</xref>
        ]; if knowledge gap is large, less-skilled programmers
may tend to be more passive and disengaged [
        <xref ref-type="bibr" rid="ref11">23</xref>
        ]
Conversations with intermediate-level details contribute to pair programming success
[
        <xref ref-type="bibr" rid="ref12">24</xref>
        ]; diferent types of discourse lead to more attempts or more debug success [ 25]
Over-reliance leads to conflicts and impedes satisfaction and learning, as work is
entirely burdened on one partner [
        <xref ref-type="bibr" rid="ref6">4, 18</xref>
        ]; educators recommend regular role-switching
to ensure equitable learning in collaboration [2]
Scheduling dificulties [ 26], teaching &amp; evaluating individual responsibility and
accountability are important to collaboration success [27], but can lead to increased
management costs [
        <xref ref-type="bibr" rid="ref9">21, 28</xref>
        ]
Human-AI (Copilot)
N/A
N/A
N/A
N/A
N/A
productivity; however, the nature of interaction with Copilot (tab to accept suggestions) is
likely to contribute to more added lines in the human-Copilot condition, and how valid would
it represent the notion of productivity is questionable.
      </p>
      <p>Since there is not enough research for a comprehensive review of human-AI pair programming,
we cannot reach any conclusion on pAIr efectiveness yet. It is also hard to compare the
human-human and human-AI pair programming literature, as they difer in what outcomes and
measurements they adopt. Therefore, in the top rows of Table 1, we listed the most common
outcome variables in both literature (quality, productivity, satisfaction, learning, and cost) and
some sample works to demonstrate mixed outcomes and example measures that operationally
define the outcomes. For more details, please refer to our more elaborated work [32].</p>
      <p>
        Note that researchers can use diferent words to characterize similar constructs. For example,
duration, efort, and productivity are all types of “eficiency” outcomes that involve time and
accomplishment. Productivity can be measured in terms of the number of completed tasks in
a fixed unit of time, duration can be measured as the amount of elapsed or total time used to
complete a fixed number of tasks to a certain standard, and efort can be measured as twice the
duration, the person-hours required, etc. [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. We use productivity as an aggregated outcome
variable of diferent measures, for consistency with the human-AI literature.
      </p>
    </sec>
    <sec id="sec-3">
      <title>3. Moderators</title>
      <p>
        In search of the explanations of the cost-benefit of human-human pair programming experiences,
researchers have found moderators such as task type &amp; complexity [29], compatibility factors
like expertise [27, 33], communication [
        <xref ref-type="bibr" rid="ref11 ref12">34, 24, 23</xref>
        ], collaboration factors like over-reliance and
role-switching [
        <xref ref-type="bibr" rid="ref2">4, 35, 14</xref>
        ], and logistics dificulties including scheduling and training [ 26, 29] (as
shown in the bottom rows of Table 1). For human-AI pair programming’s moderators, much
was unexplored – we do not know what could make human-AI pair programming more or less
efective. Therefore, in this section, we discuss the key moderators that are examined in the
human-human pair programming literature, and individual examples of moderating efects are
provided in Table 1.
      </p>
      <sec id="sec-3-1">
        <title>3.1. Task Types &amp; Complexity</title>
        <p>
          For task type and task complexity, Chaparro et al. [
          <xref ref-type="bibr" rid="ref10">22</xref>
          ] found that debugging tasks lead to less
satisfaction and perceived eficacy compared to comprehension and refactoring tasks. Hannay
et al. [29] found that the duration is shorter for low complexity tasks, at the expense of lower
quality results, and quality is higher when complexity is higher, but it requires considerably
greater efort. Arisholm et al. [33] found that the moderating efect of complexity also depends
on the expertise of the pair, where “benefits of correctness on complex system apply mainly to
juniors, whereas the reductions in duration to perform the tasks correctly on the simple system
apply mainly to intermediates and seniors.”
        </p>
      </sec>
      <sec id="sec-3-2">
        <title>3.2. Compatibility</title>
        <p>
          Salleh et al. [
          <xref ref-type="bibr" rid="ref2">14</xref>
          ] listed multiple factors for pair compatibility, such as personality, perceived
skills, actual skills (expertise), self-esteem, gender, and work ethic. Thomas et al. [
          <xref ref-type="bibr" rid="ref3">15</xref>
          ] found that
paired students with similar self-confidence levels produce their best work. Hannay et al. [35]
found that Big Five personality traits only have modest predictive value on pair programming
performance, in comparison to expertise, task complexity, and country. There also seems to be
evidence that women benefit from pair programming more than men [ 27, 30].
        </p>
        <p>
          Expertise as a compatibility factor has been extensively studied. For example, researchers
found that a student pair performs the best when their expertise is similar [
          <xref ref-type="bibr" rid="ref2">14</xref>
          ] and students
preferred to be paired with similarly skilled partners [
          <xref ref-type="bibr" rid="ref10">22</xref>
          ]. However, in industry, Jensen [36]
reported that when both members were near the same capability level and strongly opinionated,
the collaboration was counter-productive and troublesome.
        </p>
        <p>
          In the introductory programming context, Lui and Chan [37] found that pairing up novices
results in a larger improvement in productivity than pairing up experts. However, there are
concerns about “the blind leading the blind” if they don’t have an expert to consult with [
          <xref ref-type="bibr" rid="ref9">21</xref>
          ].
Researchers also found that less-skilled students learn and enjoy more than more-skilled students
in pair programming [
          <xref ref-type="bibr" rid="ref10 ref8">22, 20</xref>
          ]. However, when the knowledge gap is too large, students can be
less satisfied and the benefits of quality may be smaller [ 13]. Chong and Hurlbutt [
          <xref ref-type="bibr" rid="ref11">23</xref>
          ] reported
that a novice programmer collaborating with an expert may become disengaged, have lower
self-esteem, and be afraid of slowing down or annoying their more-skilled partner [
          <xref ref-type="bibr" rid="ref9">21</xref>
          ].
        </p>
      </sec>
      <sec id="sec-3-3">
        <title>3.3. Communication</title>
        <p>
          According to Freudenberg et al. [
          <xref ref-type="bibr" rid="ref12">24</xref>
          ], “the key to the success of pair programming [is] the
proliferation of talk at an intermediate level of detail in pair programmers’ conversations.”
Researchers found that pair programming eliminates distracting activity and enables programmers
to focus on productive activity [38], which could be why engaging communications contribute
to successful pair programming. Murphy et al. [25] used transactive analysis to break down
communication by diferent types of transactions and found that attempting more problems
associated with more completion and debugging success correlated with more critique
transactions. Some other works pointed out the social support aspect of communication [
          <xref ref-type="bibr" rid="ref11">23</xref>
          ] and an
explanation efect where the verbalization of the thought process makes thinking clearer [
          <xref ref-type="bibr" rid="ref4">16</xref>
          ].
        </p>
        <p>In human-human pair programming, programmers spend about 1/3 of the time primarily
focusing on communication [34], which forces them to concentrate, rationalize, and explain
their thoughts [38, 29]. In human-AI pair programming, Mozannar et al. [39] has shown that
an analogous 1/3 amount of time is spent communicating with Copilot, such as thinking and
verifying (22.4%) Copilot’s suggestion, which may be replicating the self-explanation efects in
some ways, and prompt crafting, which takes 11.56% of the time. These activities are arguably
eforts to understand and communicate with Copilot. However, there is no other human
to co-verify the answers, and there is no study that evaluate the communicative nature of
human-Copilot interaction as human-human pair programming.</p>
      </sec>
      <sec id="sec-3-4">
        <title>3.4. Collaboration</title>
        <p>
          Collaboration can fail in various ways in a human-human pair. For example, the free-rider
problem, where the entire workload is on one partner while the other remains a marginal player,
can result in less satisfaction and learning [
          <xref ref-type="bibr" rid="ref6">4, 18</xref>
          ]. In human-AI pair programming, educators
are worried that easily available code-generation tools may lead to cheating, and over-reliance
on AI may hinder students learning [40]. However, no study has formally evaluated it yet.
        </p>
        <p>
          For human-human pair programming, there is a suggested collaboration pattern of
roleswitching – two software developers periodically and regularly switch between writing code
(driver) and suggesting code (navigator), aiming to ensure that both are engaged in the task
and alleviate the physical and cognitive load borne by the driver [
          <xref ref-type="bibr" rid="ref1">1, 34</xref>
          ]. Some researchers
Freudenberg et al. [
          <xref ref-type="bibr" rid="ref12">24</xref>
          ] argue that the success of pair programming should be attributed to
communication rather than “the diferences in behavior or focus between the driver and navigator,”
as they found both driver and navigator worked on similar levels of abstraction.
Nevertheless, instructors still recommend drivers and navigators to regularly alternate roles to ensure
equitable learning experiences [2].
        </p>
        <p>In human-AI interaction, given Copilot’s amazing capability to write code in diferent
languages, some have argued that Copilot can take on the role of the “driver” in pair programming,
allowing a solo programmer to take on the role of the “navigator” and focus on understanding
the code at a higher level [9]. However, while it is possible for humans to ofload some API
lookup and syntax details to Copilot, humans still need to jump back into the driver’s seat
frequently and fluidly switch between the thinking and writing activities [ 39]. It is ultimately
the human programmer’s sole responsibility to understand the code at the statement level [41].</p>
      </sec>
      <sec id="sec-3-5">
        <title>3.5. Logistics</title>
        <p>
          Logistical challenges, including scheduling dificulties, teaching and evaluating collaboration
for the pair, and figuring out individual accountability and responsibility [ 26, 27], can add to
the management cost of human-human pair programming [
          <xref ref-type="bibr" rid="ref9">21, 28</xref>
          ].
        </p>
        <p>In human-AI pair programming, some may argue that the human is solely responsible in the
human-AI pair [41], but the accountability of these LLM-based generative AI is still under debate
[40]. There may be new logistics issues for the human-AI pair, such as teaching humans how
to best collaborate with Copilot. There could also be unique challenges as in every human-AI
interaction scenario, such as bias, trust, and technical limitations – much to be explored. More
study would be needed to empirically and experimentally verify the moderating efects of
diferent variables in human-AI pair programming.</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>4. Discussion and Future Work</title>
      <sec id="sec-4-1">
        <title>4.1. LLM, A Better pAIr Programmer?</title>
        <p>As reviewed in Section 2, previous literature has explored a variety of measures to evaluate
diferent aspects of human-human pair programming, while the current exploration in
humanAI pair programming is quite limited. Murillo and D’Angelo [42] have proposed evaluation
metrics for LLM-based creative code writing assistants in software engineering. More works
could use more valid measures in the human-human pair programming literature to explore how
to best help humans and LLM-based AI programming assistant collaborate together. It would
also be interesting to have a study setup with three conditions – human-human, human-AI, and
human solo – working on the same task.</p>
        <p>
          Note that in this paper, we mostly covered studies using the VSCode Extension Copilot. Tools
like ChatGPT may support the communication aspect better than Copilot [43], and there are
also Bard developed by Google [44] and an experimental version of Copilot Labs by Github [45],
which support more functionalities such as fix bug, clean, and customizable prompts. Those
tools may already improve the human-AI pair programming interaction in some ways, so future
studies could also compare across a variety of LLM-based programming tools.
pair programming may yield opportunities to explore in human-AI pair programming (Table 2).
For example, self-eficacy can lead to a diference in satisfaction [
          <xref ref-type="bibr" rid="ref3">15</xref>
          ] and gender can lead to
a diference in learning [
          <xref ref-type="bibr" rid="ref8">20</xref>
          ], do these compatibility moderators influence pAIr too? Can we
improve pAIr outcomes using insights derived from human-human literature (e.g., simulate an
AI partner with similar self-eficacy levels and the same gender)?
        </p>
        <sec id="sec-4-1-1">
          <title>We discuss more details on</title>
          <p>how and why might LLM be used in ways presented in Table 2 in our later work [32].</p>
          <p>Therefore, in general, we can ask the following questions for future works: could these
factors be implemented for human-AI pair programming; would they make human-AI pair
programming more efective, less efective, or have no influence, and why?</p>
        </sec>
      </sec>
      <sec id="sec-4-2">
        <title>4.2. LLM, Students’ pAIr Programmer?</title>
        <p>Most current studies that evaluate the eficacy of Copilot are conducted with experienced
software developers. If we estimate Copilot’s problem-solving abilities as an average student
in introductory programming classes, evaluating its performance when pairing up with a
professional software developer with much more expertise may not bring enough benefit to the
professional. Therefore, working with LLM’s current capabilities, it seems like a student-AI
pair programming setup would be the most promising to explore, so the next question is: how
should we best support student-AI pair programming?
Re-prioritize programming skills.</p>
        <sec id="sec-4-2-1">
          <title>First of all, co-working with AI requires a special skill</title>
          <p>
            set, and future work could explore how to support students to better develop these crucial skills.
Bird et al. [
            <xref ref-type="bibr" rid="ref4">16</xref>
            ] argued that the popularity of LLM-based programming assistants will result in
the growing importance of reviewing code as a skill for developers. Nonetheless, in Perscheid
et al. [46]’s interview, none of the professional developers remembered training on debugging
at school. There is already rich literature on debugging and testing instructions [47, 48, 49], but
logistical challenges like the lack of instructional time still exist [49, 50], and educators need to
better prepare students with debugging and testing skills needed to work with unreliable AI.
Integrate AIEd frameworks. Holstein et al. [51] developed a framework to map ways to
mutually augment humans and AI in education, for example, by augmenting interpretation,
action, scalability, and capacity. Future works can use existing theories in the AI education
space to improve the design of the AI pAIr programming partner, and further investigate if
LLMs bring new focus and afordances to previous human-AI education frameworks.
Support explanation and communication with students. Previous attempts of using
AI agent as pair programming partner have shown some preliminary success in knowledge
transfer and retention [52, 53], and the limitation discussed was the lack of discussion and
explanation [54]. Nowadays, as an LLM-based agent can support more natural interaction and
provide good quality explanations in the introductory programming context [55], it would be
interesting to explore if LLM-based AI could resolve some limitations mentioned in pedagogical
and conversational agent works before. Self-reflection and explanation techniques may also be
adopted to make up for the communication aspect as in human-human pair programming.
Match expertise with students. Furthermore, as discussed in Section 3, matching expertise
is a tricky problem. Lui and Chan [37] found that expert-expert pair may not gain as much of
an advantage over an expert solo programmer, in comparison to novice-novice pair vs. a solo
novice. Meanwhile, pairing two novices together raise concerns of “the blind leading the blind,”
but pairing a novice with an expert may lead to lower self-esteem of the novice [
            <xref ref-type="bibr" rid="ref9">21</xref>
            ]. Given
all these complexities, when it comes to a student-AI pair and when we only care about the
student’s learning gains, there are a lot of research questions to ask. If we have full control of
the perceived skill level of the AI partner, should we configure it to be similar to the student,
slightly higher skilled, or a lot better? Would it be beneficial to have both a peer AI agent but
also a tutor AI agent to assist if students get stuck?
Avoid over-helping students. Note that for programming learners, it would be important
to configure the LLM-based programming assistant to avoid over-help. In the few studies that
examined novice interaction with Copilot [56] or a customized programming environment based
on LLM-based code generation model Codex [10]. Prather et al. [56] found that novices do have
unique interaction patterns with Copilot and a tendency to rely on and trust the generated code
too much. Kazemitabaar et al. [10] discussed design implications including control over-use
and support complete novices. There have also been concerns about academic integrity and
changing perception of learning when LLM-based programming tools become easily accessible
to students [40, 57, 56], which are issues to further explore for student-AI pair programming.
Boost students’ self-confidence. Pair programming has been shown to benefit students
with lower self-eficacy and self-confidence levels [
            <xref ref-type="bibr" rid="ref3">15</xref>
            ] and women [
            <xref ref-type="bibr" rid="ref8">20</xref>
            ] more, which could make
it a pedagogical tool to engage more vulnerable or underrepresented populations in CS. When
an AI is introduced in pair programming, would the same benefits retain? How should we
present the AI diferently to make it compatible with students with diferent confidence levels?
How do we mitigate the risks of unreliable but seemingly authoritative AI?
Address critical risks in using LLM. Last but not least, LLMs may be an opportunity to
address some existing challenges that student-student pair programming has (as summarized in
Table 2), but there are critical risks associated with using generative AI in education, such as bias,
trust, and transparency [58]. Evaluation on existing benchmarks shows that ChatGPT exhibits
ethnically risky behaviors and presents inaccurate information [59]. Therefore, to ethnically
and responsibly use AI as a student’s pair programming partner, besides keep improving model
designs, we also need to keep human instructors’ supervision and education students about
AI’s limitations.
          </p>
        </sec>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>5. Conclusion</title>
      <p>This paper has discussed the concept of human-AI pair programming (pAIr programming).
Research has yet to pinpoint which of these supposed advantages of human-AI pair programming
yields the largest benefits in eficiency and learning. Human-human pair programming literature
yield insights on what outcomes and measures should researchers use to evaluate their pAIr
programming work (e.g., use more valid quality and productivity measurements, and further
investigate cost), and what moderators should researchers consider to further analyze and
improve pAIr programming’s process and design (e.g, compatibility, communication, etc.).</p>
      <p>In conclusion, more valid and comprehensive measurements are needed to evaluate pAIr
programming, more comparisons can be drawn between human-human vs. human-AI pair
programming, and more works can explore how to best support LLM-assisted programming
with insights from the rich literature on human-human pair programming.</p>
    </sec>
    <sec id="sec-6">
      <title>Acknowledgments</title>
      <p>Thanks to Ken’s lab members for giving feedback on this work. Thanks to Stephen MacNeil for
coming up with the creative “pAIr” keyword for this project.
[2] K. Umapathy, A. D. Ritzhaupt, A Meta-Analysis of Pair-Programming in computer
programming courses: Implications for educational practice, ACM Trans. Comput. Educ. 17
(2017) 1–13. URL: https://doi.org/10.1145/2996201. doi:1 0 . 1 1 4 5 / 2 9 9 6 2 0 1 .
[3] X. Wei, L. Lin, N. Meng, W. Tan, S.-C. Kong, Others, The efectiveness of partial
pair programming on elementary school students’ computational thinking skills
and self-eficacy, Comput. Educ. 160 (2021) 104023. URL: https://www.sciencedirect.
com/science/article/pii/S0360131520302219?casa_token=BCC-5YvaAgAAAAAA:
lIOmfQPTBGpeL296DBtelqwPs7pNeHKuugy8gsVK0Nn6bha9pdJIMm2G3JLUNXTwTUZavVA2Fw.
[4] L. Williams, E. Wiebe, K. Yang, M. Ferzli, C. Miller, In support of pair programming in
the introductory computer science course, Comput. Sci. Educ. 12 (2002) 197–212. URL:
http://www.tandfonline.com/doi/abs/10.1076/csed.12.3.197.8618. doi:1 0 . 1 0 7 6 / c s e d . 1 2 . 3 .
1 9 7 . 8 6 1 8 .
[5] S. Xu, V. Rajlich, Pair programming in graduate software engineering course projects,
in: Proceedings Frontiers in Education 35th Annual Conference, IEEE, 2006. URL: http:
//ieeexplore.ieee.org/document/1612027/. doi:1 0 . 1 1 0 9 / f i e . 2 0 0 5 . 1 6 1 2 0 2 7 .
[6] GitHub, Your AI pair programmer: Copilot, https://github.com/features/copilot, 2021. URL:
https://github.com/features/copilot, accessed: 2022-10-5.
[7] R. Sison, Investigating the efect of pair programming and software size on software quality
and programmer productivity, in: 2009 16th Asia-Pacific Software Engineering Conference,
2009, pp. 187–193. URL: http://dx.doi.org/10.1109/APSEC.2009.71. doi:1 0 . 1 1 0 9 / A P S E C . 2 0 0 9 .
7 1 .
[8] L. Williams, R. L. Upchurch, In support of student pair-programming, in:
Proceedings of the thirty-second SIGCSE technical symposium on Computer Science Education,
ACM, New York, NY, USA, 2001. URL: https://collaboration.csc.ncsu.edu/laurie/Papers/
WilliamsUpchurch.pdf. doi:1 0 . 1 1 4 5 / 3 6 4 4 4 7 . 3 6 4 6 1 4 .
[9] S. Imai, Is GitHub copilot a substitute for human pair-programming? an empirical study,
in: 2022 IEEE/ACM 44th International Conference on Software Engineering: Companion
Proceedings (ICSE-Companion), ieeexplore.ieee.org, 2022, pp. 319–321. URL: http://dx.doi.
org/10.1145/3510454.3522684. doi:1 0 . 1 1 4 5 / 3 5 1 0 4 5 4 . 3 5 2 2 6 8 4 .
[10] M. Kazemitabaar, J. Chow, C. K. T. Ma, B. J. Ericson, D. Weintrop, T. Grossman, Studying the
efect of AI code generators on supporting novice learners in introductory programming
(2023). URL: http://arxiv.org/abs/2302.07427. a r X i v : 2 3 0 2 . 0 7 4 2 7 .
[11] S. Peng, E. Kalliamvakou, P. Cihon, M. Demirer, The impact of AI on developer
productivity: Evidence from GitHub copilot (2023). URL: http://arxiv.org/abs/2302.06590.
a r X i v : 2 3 0 2 . 0 6 5 9 0 .
[12] P. Vaithilingam, T. Zhang, E. L. Glassman, Expectation vs. experience: Evaluating the
usability of code generation tools powered by large language models, in: Extended
Abstracts of the 2022 CHI Conference on Human Factors in Computing Systems, number
Article 332 in CHI EA ’22, Association for Computing Machinery, New York, NY, USA,
2022, pp. 1–7. URL: https://doi.org/10.1145/3491101.3519665. doi:1 0 . 1 1 4 5 / 3 4 9 1 1 0 1 . 3 5 1 9 6 6 5 .
[13] V. V. K. Padmanabhuni, H. P. Tadiparthi, S. M. Muralidhar Yanamadala,
Effective pair programming practice-an experimental study, Journal of
Emerging Trends in Computing and Information Sciences 3 (2012) 471–479. URL:
http://www.agilemethod.csie.ncu.edu.tw/agileMethod/download/2012papers/2012%
[25] L. Murphy, S. Fitzgerald, B. Hanks, R. McCauley, Pair debugging: a transactive discourse
analysis, in: Proceedings of the Sixth international workshop on Computing education
research, ICER ’10, Association for Computing Machinery, New York, NY, USA, 2010, pp.
51–58. URL: https://doi.org/10.1145/1839594.1839604. doi:1 0 . 1 1 4 5 / 1 8 3 9 5 9 4 . 1 8 3 9 6 0 4 .
[26] A. Begel, N. Nagappan, Pair programming: what’s in it for me?, in: Proceedings of the
Second ACM-IEEE international symposium on Empirical software engineering and
measurement, ACM, New York, NY, USA, 2008. URL: https://dl.acm.org/doi/10.1145/1414004.
1414026. doi:1 0 . 1 1 4 5 / 1 4 1 4 0 0 4 . 1 4 1 4 0 2 6 .
[27] D. Preston, Using collaborative learning research to enhance pair programming pedagogy,
SIGITE Newsl. 3 (2006) 16–21. URL: https://doi.org/10.1145/1113378.1113381. doi:1 0 . 1 1 4 5 /
1 1 1 3 3 7 8 . 1 1 1 3 3 8 1 .
[28] W. Sun, G. Marakas, The true cost of pair programming: Development of a comprehensive
model and test, Americas Conference on Information Systems (2009). URL: https://www.
semanticscholar.org/paper/647fc48650e4f19962c8a6feb87f3bdedde9dd04.
[29] J. E. Hannay, T. Dybå, E. Arisholm, D. I. K. Sjøberg, The efectiveness of pair
programming: A meta-analysis, Information and Software Technology 51 (2009) 1110–1122.
URL: https://www.sciencedirect.com/science/article/pii/S0950584909000123. doi:1 0 . 1 0 1 6 /
j . i n f s o f . 2 0 0 9 . 0 2 . 0 0 1 .
[30] B. Hanks, S. Fitzgerald, R. McCauley, L. Murphy, C. Zander, Pair programming in education:
a literature review, Comput. Sci. Educ. 21 (2011) 135–173. URL: https://www.tandfonline.
com/doi/full/10.1080/08993408.2011.579808. doi:1 0 . 1 0 8 0 / 0 8 9 9 3 4 0 8 . 2 0 1 1 . 5 7 9 8 0 8 .
[31] S. Barke, M. B. James, N. Polikarpova, Grounded copilot: How programmers interact with</p>
      <p>Code-Generating models (2022). URL: http://arxiv.org/abs/2206.15000. a r X i v : 2 2 0 6 . 1 5 0 0 0 .
[32] Q. Ma, T. Wu, K. Koedinger, Is AI the better programming partner? Human-Human pair
programming vs. Human-AI pAIr programming (2023). URL: http://arxiv.org/abs/2306.
05153. a r X i v : 2 3 0 6 . 0 5 1 5 3 .
[33] E. Arisholm, H. Gallis, T. Dyba, D. I. K. Sjoberg, Evaluating pair programming with respect
to system complexity and programmer expertise, IEEE Trans. Software Eng. 33 (2007)
65–86. URL: http://dx.doi.org/10.1109/TSE.2007.17. doi:1 0 . 1 1 0 9 / T S E . 2 0 0 7 . 1 7 .
[34] L. Plonka, J. Segal, H. Sharp, J. van der Linden, Collaboration in pair programming: Driving
and switching, in: Agile Processes in Software Engineering and Extreme Programming
- 12th International Conference, XP 2011, Madrid, Spain, May 10-13, 2011. Proceedings,
volume 77, unknown, 2011, pp. 43–59. URL: https://www.researchgate.net/publication/
221592723_Collaboration_in_Pair_Programming_Driving_and_Switching. doi:1 0 . 1 0 0 7 /
9 7 8 - 3 - 6 4 2 - 2 0 6 7 7 - 1 \ _ 4 .
[35] J. E. Hannay, E. Arisholm, H. Engvik, D. I. K. Sjoberg, Efects of personality on pair
programming, IEEE Trans. Software Eng. 36 (2010) 61–80. URL: http://dx.doi.org/10.1109/
TSE.2009.41. doi:1 0 . 1 1 0 9 / T S E . 2 0 0 9 . 4 1 .
[36] R. W. Jensen, A pair programming experience, ACCU - professionalism in programming</p>
      <p>Overload 13 (2005). URL: https://accu.org/journals/overload/13/65/jensen_254/.
[37] K. M. Lui, K. C. C. Chan, Pair programming productivity: Novice–novice vs. expert–
expert, Int. J. Hum. Comput. Stud. 64 (2006) 915–925. URL: https://linkinghub.elsevier.
com/retrieve/pii/S1071581906000644. doi:1 0 . 1 0 1 6 / j . i j h c s . 2 0 0 6 . 0 4 . 0 1 0 .
[38] A. Sillitti, G. Succi, J. Vlasenko, Understanding the impact of pair programming on
developers attention: A case study on a large industrial experimentation, in: 2012 34th
International Conference on Software Engineering (ICSE), IEEE, 2012, pp. 1094–1101. URL:
http://dx.doi.org/10.1109/ICSE.2012.6227110. doi:1 0 . 1 1 0 9 / I C S E . 2 0 1 2 . 6 2 2 7 1 1 0 .
[39] H. Mozannar, G. Bansal, A. Fourney, E. Horvitz, Reading between the lines: Modeling user
behavior and costs in AI-assisted programming, ArXiv (2022). URL: http://dx.doi.org/10.
48550/ARXIV.2210.14306. doi:1 0 . 4 8 5 5 0 / A R X I V . 2 2 1 0 . 1 4 3 0 6 .
[40] B. A. Becker, P. Denny, J. Finnie-Ansley, A. Luxton-Reilly, J. Prather, E. A. Santos,
Programming is hard - or at least it used to be: Educational opportunities and challenges
of AI code generation, in: Proceedings of the 54th ACM Technical Symposium on
Computer Science Education V. 1, SIGCSE 2023, Association for Computing
Machinery, New York, NY, USA, 2023, pp. 500–506. URL: https://doi.org/10.1145/3545945.3569759.
doi:1 0 . 1 1 4 5 / 3 5 4 5 9 4 5 . 3 5 6 9 7 5 9 .
[41] A. Sarkar, A. D. Gordon, C. Negreanu, C. Poelitz, S. S. Ragavan, B. Zorn, What is it
like to program with artificial intelligence? (2022). URL: http://arxiv.org/abs/2208.06213.
a r X i v : 2 2 0 8 . 0 6 2 1 3 .
[42] A. Murillo, S. D’Angelo, An engineering perspective on writing assistants for productivity
and creative code, The Second Workshop on Intelligent and Interactive Writing Assistants
(2023). URL: https://cdn.glitch.global/d058c114-3406-43be-8a3c-d3afff35eda2/paper1_2023.
pdf.
[43] H. H. Thorp, ChatGPT is fun, but not an author, Science 379 (2023) 313. URL: http:
//dx.doi.org/10.1126/science.adg7879. doi:1 0 . 1 1 2 6 / s c i e n c e . a d g 7 8 7 9 .
[44] Google, Bard, https://bard.google.com/, ???? URL: https://bard.google.com/, accessed:
2023-5-19.
[45] Github, GitHub copilot labs, https://githubnext.com/projects/copilot-labs/, ???? URL:
https://githubnext.com/projects/copilot-labs/, accessed: 2023-5-19.
[46] M. Perscheid, B. Siegmund, M. Taeumel, R. Hirschfeld, Studying the advancement in
debugging practice of professional software developers, Software Quality Journal 25 (2017)
83–110. URL: https://doi.org/10.1007/s11219-015-9294-2. doi:1 0 . 1 0 0 7 / s 1 1 2 1 9 - 0 1 5 - 9 2 9 4 - 2 .
[47] W. Ahrendt, R. Bubel, R. Hähnle, Integrated and Tool-Supported teaching of testing,
debugging, and verification, in: Teaching Formal Methods, Springer Berlin
Heidelberg, 2009, pp. 125–143. URL: http://dx.doi.org/10.1007/978-3-642-04912-5_9. doi:1 0 . 1 0 0 7 /
9 7 8 - 3 - 6 4 2 - 0 4 9 1 2 - 5 \ _ 9 .
[48] J. Smith, J. Tessler, E. Kramer, C. Lin, Using peer review to teach software testing, in:
Proceedings of the ninth annual international conference on International computing
education research, ICER ’12, Association for Computing Machinery, New York, NY,
USA, 2012, pp. 93–98. URL: https://doi.org/10.1145/2361276.2361295. doi:1 0 . 1 1 4 5 / 2 3 6 1 2 7 6 .
2 3 6 1 2 9 5 .
[49] R. McCauley, S. Fitzgerald, G. Lewandowski, L. Murphy, B. Simon, L. Thomas, C. Zander,
Debugging: A review of the literature from an educational perspective, Computer Science
Education 18 (2008) 67–92. URL: http://www.informaworld.com/openurl?genre=article&amp;id=doi:
10.1080/08993400802114581. doi:1 0 . 1 0 8 0 / 0 8 9 9 3 4 0 0 8 0 2 1 1 4 5 8 1 .
[50] S. Fitzgerald, R. McCauley, B. Hanks, L. Murphy, B. Simon, C. Zander, Debugging from the
student perspective, IEEE Trans. Educ. 53 (2010) 390–396. URL: http://dx.doi.org/10.1109/
TE.2009.2025266. doi:1 0 . 1 1 0 9 / T E . 2 0 0 9 . 2 0 2 5 2 6 6 .
[51] K. Holstein, V. Aleven, N. Rummel, A conceptual framework for Human–AI hybrid
adaptivity in education, Artificial Intelligence in Education 12163 (2020) 240. URL: https:
//www.ncbi.nlm.nih.gov/pmc/articles/PMC7334162/. doi:1 0 . 1 0 0 7 / 9 7 8 - 3 - 0 3 0 - 5 2 2 3 7 - 7 \ _ 2 0 .
[52] P. Robe, S. K. Kuttal, Designing PairBuddy—A conversational agent for pair programming,
ACM Trans. Comput.-Hum. Interact. 29 (2022) 1–44. URL: https://doi.org/10.1145/3498326.
doi:1 0 . 1 1 4 5 / 3 4 9 8 3 2 6 .
[53] K.-W. Han, E. Lee, Y. Lee, The impact of a Peer-Learning agent based on pair programming
in a programming course, IEEE Trans. Educ. 53 (2010) 318–327. URL: http://dx.doi.org/10.
1109/TE.2009.2019121. doi:1 0 . 1 1 0 9 / T E . 2 0 0 9 . 2 0 1 9 1 2 1 .
[54] S. K. Kuttal, B. Ong, K. Kwasny, P. Robe, Trade-ofs for substituting a human with an agent
in a pair programming context: The good, the bad, and the ugly, in: Proceedings of the
2021 CHI Conference on Human Factors in Computing Systems, number Article 243 in
CHI ’21, Association for Computing Machinery, New York, NY, USA, 2021, pp. 1–20. URL:
https://doi.org/10.1145/3411764.3445659. doi:1 0 . 1 1 4 5 / 3 4 1 1 7 6 4 . 3 4 4 5 6 5 9 .
[55] J. Leinonen, P. Denny, S. MacNeil, S. Sarsa, S. Bernstein, J. Kim, A. Tran, A. Hellas,
Comparing code explanations created by students and large language models (2023). URL:
http://arxiv.org/abs/2304.03938. a r X i v : 2 3 0 4 . 0 3 9 3 8 .
[56] J. Prather, B. N. Reeves, P. Denny, B. A. Becker, J. Leinonen, A. Luxton-Reilly, G. Powell,
J. Finnie-Ansley, E. A. Santos, “it’s weird that it knows what I want”: Usability and
interactions with copilot for novice programmers (2023). URL: http://arxiv.org/abs/2304.
02491. a r X i v : 2 3 0 4 . 0 2 4 9 1 .
[57] B. Puryear, G. Sprint, Github copilot in the classroom: learning to code with AI assistance,
J. Comput. Sci. Coll. 38 (2022) 37–47. URL: https://dl.acm.org/doi/pdf/10.5555/3575618.
3575622.
[58] D. Mhlanga, Open AI in education, the responsible and ethical use of ChatGPT towards
lifelong learning, 2023. URL: https://papers.ssrn.com/abstract=4354422. doi:1 0 . 2 1 3 9 / s s r n .
4 3 5 4 4 2 2 .
[59] T. Y. Zhuo, Y. Huang, C. Chen, Z. Xing, Red teaming ChatGPT via jailbreaking: Bias,
robustness, reliability and toxicity (2023). URL: https://www.researchgate.net/publication/
368476294. doi:1 0 . 4 8 5 5 0 / A R X I V . 2 3 0 1 . 1 2 8 6 7 . a r X i v : 2 3 0 1 . 1 2 8 6 7 .</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>C.</given-names>
            <surname>Alves De Lima Salge</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Berente</surname>
          </string-name>
          ,
          <article-title>Pair programming vs. solo programming: What do we know after 15 years of research?</article-title>
          ,
          <source>in: 2016 49th Hawaii International Conference on System Sciences (HICSS)</source>
          ,
          <year>2016</year>
          , pp.
          <fpage>5398</fpage>
          -
          <lpage>5406</lpage>
          . URL: http://dx.doi.org/10.1109/HICSS.
          <year>2016</year>
          .
          <volume>667</volume>
          .
          <source>doi:1 0 . 1 1</source>
          <volume>0</volume>
          <fpage>9</fpage>
          <string-name>
            <surname>/ H I C S</surname>
          </string-name>
          <article-title>S . 2 0 1 6 . 6 6 7</article-title>
          . 20Effective%20Pair%
          <fpage>20Programming</fpage>
          %
          <fpage>20Practice</fpage>
          -%
          <source>20An%20Experimental%20Study/ Effective%20Pair%20Programming%20Practice-%20An%20Experimental%20Study.pdf.</source>
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [14]
          <string-name>
            <given-names>N.</given-names>
            <surname>Salleh</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Mendes</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Grundy</surname>
          </string-name>
          ,
          <article-title>Empirical studies of pair programming for CS/SE teaching in higher education: A systematic literature review</article-title>
          ,
          <source>IEEE Trans. Software Eng</source>
          .
          <volume>37</volume>
          (
          <year>2011</year>
          )
          <fpage>509</fpage>
          -
          <lpage>525</lpage>
          . URL: http://dx.doi.org/10.1109/TSE.
          <year>2010</year>
          .
          <volume>59</volume>
          .
          <source>doi:1 0 . 1 1 0 9 / T S E . 2 0</source>
          <volume>1 0 . 5</volume>
          <fpage>9</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [15]
          <string-name>
            <given-names>L.</given-names>
            <surname>Thomas</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Ratclife</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Robertson</surname>
          </string-name>
          ,
          <article-title>Code warriors and code-a-phobes: a study in attitude and pair programming</article-title>
          ,
          <source>SIGCSE Bull</source>
          .
          <volume>35</volume>
          (
          <year>2003</year>
          )
          <fpage>363</fpage>
          -
          <lpage>367</lpage>
          . URL: https://doi.org/10. 1145/792548.612007.
          <source>doi:1 0 . 1 1</source>
          <volume>4 5 / 7 9 2 5 4 8 . 6 1 2 0 0 7 .</volume>
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [16]
          <string-name>
            <given-names>C.</given-names>
            <surname>Bird</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Ford</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Zimmermann</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Forsgren</surname>
          </string-name>
          , E. Kalliamvakou,
          <string-name>
            <given-names>T.</given-names>
            <surname>Lowdermilk</surname>
          </string-name>
          ,
          <string-name>
            <surname>I. Gazit</surname>
          </string-name>
          ,
          <article-title>Taking flight with copilot: Early insights and opportunities of AI-powered pair-programming tools</article-title>
          ,
          <source>Queueing Syst</source>
          .
          <volume>20</volume>
          (
          <year>2023</year>
          )
          <fpage>35</fpage>
          -
          <lpage>57</lpage>
          . URL: https://doi.org/10.1145/3582083.
          <source>doi:1 0 . 1 1</source>
          <volume>4 5 / 3 5 8 2 0 8 3 .</volume>
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [17]
          <string-name>
            <given-names>E.</given-names>
            <surname>Kalliamvakou</surname>
          </string-name>
          , Research: quantifying
          <article-title>GitHub copilot's impact on developer productivity and happiness</article-title>
          , https://github.blog/ 2022-09-07-research
          <article-title>-quantifying-github-copilots-impact-on-developer-</article-title>
          <string-name>
            <surname>productivity-</surname>
          </string-name>
          and-happiness/,
          <year>2022</year>
          . URL: https://github.blog/2022-09-07-research
          <article-title>-quantifying-github-copilots-impact-on-developer-product accessed:</article-title>
          <year>2022</year>
          -10-13.
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [18]
          <string-name>
            <given-names>N.</given-names>
            <surname>Nagappan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Williams</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Ferzli</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Wiebe</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Yang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Miller</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Balik</surname>
          </string-name>
          ,
          <article-title>Improving the CS1 experience with pair programming</article-title>
          ,
          <source>in: Proceedings of the 34th SIGCSE technical symposium on Computer science education, ACM</source>
          , New York, NY, USA,
          <year>2003</year>
          . URL: https: //dl.acm.org/doi/10.1145/611892.612006.
          <source>doi:1 0 . 1 1</source>
          <volume>4 5 / 6 1 1 8 9 2 . 6 1 2 0 0 6 .</volume>
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [19]
          <string-name>
            <surname>C. McDowell</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          <string-name>
            <surname>Werner</surname>
            ,
            <given-names>H. E.</given-names>
          </string-name>
          <string-name>
            <surname>Bullock</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          <string-name>
            <surname>Fernald</surname>
          </string-name>
          ,
          <article-title>Pair programming improves student retention, confidence, and program quality</article-title>
          ,
          <source>Commun. ACM</source>
          <volume>49</volume>
          (
          <year>2006</year>
          )
          <fpage>90</fpage>
          -
          <lpage>95</lpage>
          . URL: https: //dl.acm.org/doi/10.1145/1145287.1145293.
          <source>doi:1 0 . 1 1</source>
          <volume>4 5 / 1 1 4 5 2 8 7 . 1 1 4 5 2 9 3 .</volume>
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [20]
          <string-name>
            <given-names>P.</given-names>
            <surname>Maguire</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Maguire</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Hyland</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Marshall</surname>
          </string-name>
          ,
          <article-title>Enhancing collaborative learning using pair programming: Who benefits?</article-title>
          ,
          <source>AISHE-J</source>
          <volume>6</volume>
          (
          <year>2014</year>
          ). URL: https://ojs.aishe.org/index.php/ aishe-j/article/view/141.
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [21]
          <string-name>
            <given-names>M.</given-names>
            <surname>Ally</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Darroch</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Toleman</surname>
          </string-name>
          ,
          <article-title>A framework for understanding the factors influencing pair programming success</article-title>
          ,
          <source>in: Extreme Programming and Agile Processes in Software Engineering, Lecture notes in computer science</source>
          , Springer Berlin Heidelberg, Berlin, Heidelberg,
          <year>2005</year>
          , pp.
          <fpage>82</fpage>
          -
          <lpage>91</lpage>
          . URL: http://link.springer.
          <source>com/10.1007/11499053_10. doi:1 0 . 1 0</source>
          <volume>0 7 / 1 1 4 9 9 0 5 3 \ _ 1</volume>
          <fpage>0</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [22]
          <string-name>
            <given-names>E. A.</given-names>
            <surname>Chaparro</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Yuksel</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Romero</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Bryant</surname>
          </string-name>
          ,
          <article-title>Factors afecting the perceived efectiveness of pair programming in higher education</article-title>
          , Annual Workshop of the Psychology of Programming Interest Group (
          <year>2005</year>
          ). URL: https://www.semanticscholar.org/paper/ c095f0d9b17cd9c2851000534740e7cc087253fa.
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [23]
          <string-name>
            <given-names>J.</given-names>
            <surname>Chong</surname>
          </string-name>
          , T. Hurlbutt,
          <article-title>The social dynamics of pair programming</article-title>
          ,
          <source>in: 29th International Conference on Software Engineering (ICSE'07)</source>
          , ieeexplore.ieee.org,
          <year>2007</year>
          , pp.
          <fpage>354</fpage>
          -
          <lpage>363</lpage>
          . URL: http://dx.doi.org/10.1109/ICSE.
          <year>2007</year>
          .
          <volume>87</volume>
          .
          <source>doi:1 0 . 1 1 0 9 / I C S E . 2 0</source>
          <volume>0 7 . 8</volume>
          <fpage>7</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [24]
          <string-name>
            <given-names>S.</given-names>
            <surname>Freudenberg</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Romero</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B. Du</given-names>
            <surname>Boulay</surname>
          </string-name>
          ,
          <article-title>Talking the talk: Is intermediate-level conversation the key to the pair programming success story?</article-title>
          ,
          <source>in: AGILE</source>
          <year>2007</year>
          , unknown,
          <year>2007</year>
          , pp.
          <fpage>84</fpage>
          -
          <lpage>91</lpage>
          . URL: https://www.researchgate.net/publication/4270516_Talking_
          <article-title>the_talk_ Is_intermediate-level_conversation_the_key_to_the_pair_programming_success_story</article-title>
          .
          <source>doi:1 0 . 1 1 0</source>
          <string-name>
            <given-names>9</given-names>
            <surname>/ A G I L E .</surname>
          </string-name>
          <article-title>2 0 0 7 . 1</article-title>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>