<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <journal-title-group>
        <journal-title>Journal of Personality and Social Psychology 51 (1986)
649-660. doi:10.1037/0022</journal-title>
      </journal-title-group>
    </journal-meta>
    <article-meta>
      <article-id pub-id-type="doi">10.1609/hcomp.v10i1.21989</article-id>
      <title-group>
        <article-title>Do Interpersonal Skills Afect Human-AI Collaboration Performance? A Study with ChatGPT</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Ryuki Nishioka</string-name>
          <email>ryuki.240ka@gmail.com</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
          <xref ref-type="aff" rid="aff4">4</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Shoko Wakamiya</string-name>
          <email>wakamiya@is.naist.jp</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
          <xref ref-type="aff" rid="aff4">4</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Nobuyuki Shimizu</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Sumio Fujita</string-name>
          <email>sufujita@lycorp.co.jp</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Eiji Aramaki</string-name>
          <email>aramaki@is.naist.jp</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
          <xref ref-type="aff" rid="aff4">4</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>AutomationXP25: Hybrid Automation Experiences</institution>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Centered AI, Human-AI communication</institution>
          ,
          <addr-line>Human-AI teaming</addr-line>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>Human-AI collaboration, Human-AI Interaction, Large Language Model (LLM)</institution>
          ,
          <addr-line>Interpersonal skill, Human-</addr-line>
        </aff>
        <aff id="aff3">
          <label>3</label>
          <institution>LY Corporation</institution>
          ,
          <addr-line>1-3 Kioi-cho, Chiyoda-ku, Tokyo 102-8282</addr-line>
          ,
          <country country="JP">Japan</country>
        </aff>
        <aff id="aff4">
          <label>4</label>
          <institution>Nara Institute of Science and Technology</institution>
          ,
          <addr-line>8916-5 Takayama-cho, Ikoma-shi, Nara 630-0192</addr-line>
          ,
          <country country="JP">Japan</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2025</year>
      </pub-date>
      <volume>11</volume>
      <fpage>127</fpage>
      <lpage>139</lpage>
      <abstract>
        <p>Collaboration between humans and artificial intelligence (AI) has demonstrated the potential to achieve performance surpassing that of AI alone. As AI becomes more integrated into society, human-AI collaboration is expected to emerge as a new form of teamwork. The recent advancements in large language models (LLMs) have accelerated research on human-LLM collaboration across various domains. While previous studies have focused on improving the performance of LLMs and methods for efective collaboration, little is known about how user-specific traits, such as interpersonal skills, influence collaboration outcomes with LLMs. This study addresses this gap by focusing on the role of interpersonal skills in human-AI interaction to deepen understanding of human-AI collaboration. The experimental results showed that participants with lower interpersonal skills were more likely to accept AI-generated responses, suggesting that they benefit more from AI. These findings suggest that interpersonal skills could influence how users critically assess with AI-generated content.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>A</p>
    </sec>
    <sec id="sec-2">
      <title>1. Introduction</title>
      <p>
        Collaboration between humans and artificial intelligence (AI) has attracted attention because of its
potential to achieve outcomes beyond that of AI alone. This phenomenon, known as the “centaur
phenomenon,” originates from a freestyle chess tournament in which a human-AI team named “Centaur”
outperformed an AI alone [
        <xref ref-type="bibr" rid="ref1 ref2">1, 2</xref>
        ]. This suggests that human intervention plays a crucial role in the
efective use of AI. Recently, the collaboration between humans and large language models (LLMs) has
garnered growing interest and is being actively explored in writing [
        <xref ref-type="bibr" rid="ref3 ref4">3, 4</xref>
        ], education [
        <xref ref-type="bibr" rid="ref5 ref6">5, 6</xref>
        ], and various
other fields [
        <xref ref-type="bibr" rid="ref7 ref8 ref9">7, 8, 9</xref>
        ].
      </p>
      <p>
        However, synergy in human-AI collaboration is not always observed, as it is influenced by the type
of task and psychological factors [
        <xref ref-type="bibr" rid="ref10 ref11">10, 11, 12</xref>
        ]. To address this issue, researchers are examining the
potential for human-AI collaboration from various perspectives such as complementarity [13, 14],
trust [15, 16], and teamwork [17, 18]. In addition, researchers are exploring approaches to control
LLM strategies for collaborating with humans and other LLMs, informed by psychology and cognitive
science theories [19, 20, 21]. However, the impact of the ability of humans to interact with LLMs on
the performance of human-AI collaboration remains underexplored, and the specific skills required
to collaborate with LLMs efectively have yet to be identified. We believe that it is crucial to evaluate
the overall performance of human-AI collaboration, considering both AI capabilities and human skills.
For example, some users may be better at extracting high-quality answers from LLMs, whereas others
may struggle to engage in conversations with them efectively. Thus, the performance of human-LLM
collaboration likely varies depending on the communication and interpersonal skills of the user. If
human skills influence collaboration with LLMs, this insight could enable applications such as skills
training with LLMs or personalized adjustments based on user skills. Conversely, if human skills have
      </p>
      <p>CEUR
Workshop</p>
      <p>ISSN1613-0073</p>
      <sec id="sec-2-1">
        <title>Measuring the interpersonal skills</title>
        <p>Self-assessment of six
interpersonal skills
!
H1
H2
・・・</p>
        <p>✔</p>
        <sec id="sec-2-1-1">
          <title>High skill</title>
        </sec>
        <sec id="sec-2-1-2">
          <title>Low skill</title>
          <p>…
?</p>
        </sec>
      </sec>
      <sec id="sec-2-2">
        <title>Measuring the enhancement score</title>
        <p>A Comparison of human alone (w/o AI)
and human-AI collaboration (w/ AI)</p>
        <p>in four tasks
e.g., How can 11 oranges be distributed equally
among three people?</p>
        <p>3toanedac2h/.3 +4 Make juice.</p>
        <p>H1
H2
Distribute as
3, 4, and 4.
little efect, LLMs can help to mitigate the skill disparities among users. Therefore, investigating the
impact of human skills on collaboration with LLMs contributes to human-AI collaboration.</p>
        <p>This study investigates the impact of human interpersonal skills, namely the ability to interact with
other humans efectively, on the performance of human-AI collaboration. Furthermore, given that
generative AI functions as an interactive conversational agent, interpersonal skills are particularly
critical in determining collaborative outcomes. To explore this, we analyze the relationship between
interpersonal skills and the quality of interaction in human-AI collaboration. Specifically, this study
experimentally investigates whether individuals with strong interpersonal skills are more efective
in collaborating with AI. Fig. 1 presents an overview of this study. We designed four types of tasks
for collaboration with AI, namely knowledge, language, problem-solving, and debate tasks, and six
interpersonal skills based on psychological scales. A total of 24 combinations (four tasks and six skills)
were examined, and the performance of the participants and ChatGPT individually was compared with
their combined performance. Our study enhances the understanding of human-AI collaboration by
focusing on the role of interpersonal skills. These findings provide valuable insights for designing
personalized AI systems that can adapt to diverse user skill levels.</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>2. Method</title>
      <p>First, we define interpersonal skills and the enhancement score, as measured through experiments with
60 participants. We then investigate the correlations between interpersonal skills and the enhancement
scores.</p>
      <p>Interpersonal skills include social skills for building relationships [22, 23] and communication
skills [24, 25]. Some conceptual overlap exists among these scales, which makes it dificult to define them
as distinct interpersonal skills on their own. Therefore, this study uses the validated questionnaire-based
scale ENDCOREs [26]. This scale defines communication skills in a hierarchical structure
comprising three basic skills (self control, expressiveness, and decipher ability) and three interpersonal skills
(assertiveness, other acceptance, and regulation of interpersonal relationships). These six skills are
defined as main skills, each of which has four sub-skills. In this study, the six main skills are defined as
interpersonal skills and are referred to as self-control, expressive, deciphering, assertive, other
acceptance, and regulation of interpersonal relationship skills. Each interpersonal skill is evaluated
using the average of the scores for the four sub skills of each main skill.</p>
      <p>We designed four tasks based on a case involving the exchange of information and opinions, which is
a typical activity in human interactions: knowledge, language, problem-solving, and debate tasks.
Each task consists of two questions.</p>
      <p>Knowledge task: This task is designed to test general knowledge. The participants were asked to
arrange the four events (one of them is fictitious) in chronological order. The  (∈ 1, 2, 3, 4)-th oldest
event was answered by selecting from the five options: 1, 2, 3, 4, or N/A. Answers to this task were
scored one point for each correct choice. Therefore, the maximum score for each question is 4 points.</p>
      <p>Language task: This task is designed to test the comprehension of word meanings and concepts. An
example question is ‘In what ways are “desks” and “chairs” similar?’ Answers are open-ended and the
participants were asked to list three answers that they were confident in. We manually grouped answers
that were based on the same underlying idea, and each answer group was scored for uniqueness and
similarity on a 1 (low)–4 (high) scale via crowdsourcing. The scoring for this task was calculated as the
sum of the scores from these two axes. Therefore, the maximum score for each question is 24 points
(eight points × three answers).</p>
      <p>Problem-solving task: This task is designed with reference to a book about lateral thinking [27]
to test the skill of thinking creatively. For example, ‘List three ways to distribute 11 oranges evenly
among 3 people’. Answers are open-ended and the participants were asked to list three answers that
they were confident in. We manually grouped answers that were based on the same underlying idea,
and each answer group was scored for uniqueness and efectiveness on a 1 (low)–4 (high) scale via
crowdsourcing. The scoring for this task was calculated as the sum of the scores from these two axes.
Therefore, the maximum score for each question is 24 points (eight points × three answers).</p>
      <p>Debate task: This task is designed to test comprehension and analytical skills. An example question
is ‘List opinions in favor of hosting the 2025 Osaka Expo.’ Answers are open-ended and the participants
were asked to list three answers that they were confident in. We manually grouped answers that
were based on the same underlying idea, and each answer group was scored for uniqueness and
convincingness on a 1 (low)–4 (high) scale via crowdsourcing. The scoring for this task was calculated
as the sum of the scores from these two axes. Therefore, the maximum score for each question is 24
points (eight points × three answers).</p>
      <p>Specifically,   is calculated as follows:
  =

+ −   + 1

max −   + 1
, where   denotes the enhancement score for the task  .   and  + denote the total scores for all questions
in task  when answered by humans alone and when answered in collaboration with AI, respectively.
max denotes the maximum value that the score can reach in task  . To prevent division by zero, 1 is
added to both the numerator and the denominator.   represents the extent to which the use of AI
brings the score closer to the full mark, with a value closer to 1 indicating better performance.</p>
      <p>We evaluated the correlation between each interpersonal skill and the enhancement score   . The
interpersonal skills consist of six skills, and the enhancement score   is defined for each of the four
tasks, resulting in 24 correlation coeficients (four tasks
× six skills).</p>
    </sec>
    <sec id="sec-4">
      <title>3. Experiments</title>
      <p>We conducted in-person experiments to investigate the impact of interpersonal skills on interaction
with AI. This study was reviewed and approved by our institutional review board (2023-I-45).
3.1. Setup and procedure
A total of 60 Japanese laypeople (30 males and 30 females) in their 20s participated in the experiments
from July 23, 2024, to August 9, 2024. Each participant was paid 5,000 yen as remuneration. The
participants answered the questionnaire of the ENDCOREs psychometric scale with a total of eight
questions (two questions for each task). In the task, participants first answered each question on their
own (w/o AI) and then answered it again using ChatGPT (w/ AI). They repeated the process for the
eight questions and answered a total of 16 times. To ensure uniform understanding, participants looked
Scores of ChatGPT-4o-mini alone and the mean and standard deviation of  and  + for each task. ChatGPT alone
demonstrated superior performance in the knowledge and debate tasks, whereas humans alone outperformed in
the language and problem-solving tasks.</p>
      <p>Knowledge
Language
Problem-solving
Debate</p>
      <p>ChatGPT
4.0
3.2. Results and discussion
3.2.1. Does ChatGPT enhance human performace?
The analysis focused on the results from 58 out of 60 participants who used ChatGPT 4o-mini. No
significant correlation was observed between the frequency of ChatGPT use and the enhancement scores.
The correlation coeficient was also near zero, suggesting that the participants exhibited consistent
enhancement scores regardless of how often they engaged with AI. The results for each task are shown
in Table 1. As a result, human performance was enhanced in tasks where ChatGPT alone outperformed
humans. However, in tasks where ChatGPT performed poorly on its own, collaboration led to a decline
in scores.
3.2.2. Correlation between skills and enhancement scores
We calculated the Pearson correlation coeficients between the six skills in ENDCOREs and the
enhancement scores for each task (Table 2). The results demonstrated significant negative correlations
between the knowledge task and expressive skill ( = −0.35 ,  &lt; 0.01 ), knowledge task and deciphering
skill ( = −0.27 ,  &lt; 0.05 ), knowledge task and assertive skill ( = −0.26 ,  &lt; 0.05 ) and language
task and self-control skill ( = −0.27 ,  &lt; 0.05 ), whereas significant positive correlations between
problem-solving task and expressive skill ( = 0.31 ,  &lt; 0.05 ). The results showed significant negative
correlations between some skills, particularly in the knowledge task, which may indicate that people
with lower interpersonal skills benefit more from AI.</p>
      <p>k
s
a
t
e
edg Et
l
w
o
n
K
k
s
a
t
e
agu Et
g
n
a
L
k
s
a
t
g
n
i
v
loS Et
m
e
l
b
o
r
P
k
s
a
t
te t
a E
b
e
D</p>
      <p>Low High</p>
      <p>Low High</p>
      <p>Low High</p>
      <p>Low High</p>
      <p>Low High
3.2.3. Enhancement scores of Low/High skill groups
The relationship between interpersonal skills and enhancement scores was examined, as illustrated
in Fig. 2. The participants were categorized into two groups based on their interpersonal skill scores:
“low skill” (scores of 4 or below, representing poor or neither poor nor good) and “high skill” (scores
above 4, indicating high proficiency). In terms of the combinations that showed a correlation, when
the results for humans alone were lower than those for ChatGPT alone, such as the knowledge task,
individuals with lower interpersonal skills tended to have higher enhancement scores, with a more
concentrated distribution. Conversely, when the results for humans alone were higher than those for
ChatGPT alone, such as the problem-solving task, those with higher interpersonal skills tended to have
higher enhancement scores. Overall, the results indicated that individuals with lower interpersonal
skills benefited more from AI.</p>
      <p>We analyzed the correlation between interpersonal skills and the number of answers that difered
before and after using ChatGPT. In this analysis, the participants were divided into two groups:
highconfidence, believing their performance was superior to ChatGPT, and low-confidence, believing
ChatGPT was superior. A changed answer was defined as a answer that difered between not using and
using ChatGPT. For the knowledge task, we counted the number of reordered events. For the other
tasks, we grouped similar answers and counted the changes in the number of groups. As a result, a
significant negative correlation was observed, particularly in the low-confidence group: between the
problem-solving task and expressive skill ( = −0.50 ,  &lt; 0.05 ), deciphering skill ( = −0.50 ,  &lt; 0.05 ),
other acceptance skill ( = −0.64 ,  &lt; 0.01 ), and regulation of interpersonal relationship skill ( = −0.51 ,
 &lt; 0.05 ).</p>
      <p>In addition, we analyzed the correlation between interpersonal skills and the number of dialogue
steps as well as between interpersonal skills and the total number of words in the prompts, but no
significant correlations were observed. Furthermore, there were also no significant diferences between
interpersonal skills and the words used in the prompts.</p>
      <p>These findings suggest that people with lower interpersonal skills are more likely to accept the AI’s
responses when they believe that it is superior. Therefore, in the knowledge task in which ChatGPT
alone performed relatively well, the score with ChatGPT improved. Conversely, in the problem-solving
task, where the performance of ChatGPT was relatively lower, it is likely that the score with ChatGPT
decreased because participants adopted answers generated by ChatGPT instead of their own. However,
a negative correlation was also observed in the other acceptance skill, which refers to the skill of
accepting others’ positions and opinions. These findings require further investigation to clarify of the
underlying reasons.
3.3. Limitation
The study was limited to experiments using ChatGPT-4o mini, which restricts the generalizability of
the findings to other models. Also, this study specifically focused on tasks emphasizing search-like use
cases, leaving creative and domain-specific tasks, such as those in medicine or artistic content creation,
unexplored and limiting the understanding of human-AI collaboration in these contexts. In addition, the
validity of the designed tasks, the order efects of the tasks, and other such confounding variables must be
critically evaluated to ensure their appropriateness for assessing human-AI collaboration. Furthermore,
the participants were limited to Japanese in their 20s. The results are based on a correlation analysis,
and more detailed qualitative analysis and statistical testing are required to further our understanding.</p>
    </sec>
    <sec id="sec-5">
      <title>4. Conclusion</title>
      <p>In this study, we investigated the impact of human skills on collaboration with AI, focusing on the
interpersonal skills of users. Our results showed that people with lower interpersonal skills were more
likely to accept AI-generated responses, suggesting that they benefit more from AI. However, if AI’s
task performance is lower than that of humans, there is a risk that collaborative performance may also
decline. These findings suggest that interpersonal skills could influence how users critically assess
AI-generated content. Furthermore, the extent to which AI enhances human task performance varies
depending on users’ interpersonal skills, implying that there is compatibility between humans and LLMs.
It is important to design AI systems to accommodate diverse user skills to achieve better human-AI
collaborations.</p>
      <p>Finally, there is much room for future studies. Although this study compared individual performance
in simple with/without AI settings, real-world settings often involve more complex teaming, such as a
human joining an AI team and vice versa. Considering such potential settings, this study presented the
basic framework for an AI-integrated teaming evaluation.</p>
    </sec>
    <sec id="sec-6">
      <title>Acknowledgments</title>
      <p>This study was supported by collaborative research funding from LY Corporation and Cross-ministerial
Strategic Innovation Promotion Program (SIP)” on “Integrated Health Care System” (grant JPJ012425).</p>
    </sec>
    <sec id="sec-7">
      <title>Declaration on Generative AI</title>
      <p>During the preparation of this work, the authors used ChatGPT, Grammarly in order to: Grammar
and spelling check, Improve writing style, Text Translation, Paraphrase and reword. After using this
tool/service, the authors reviewed and edited the content as needed and takes full responsibility for the
publication’s content.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>N.</given-names>
            <surname>Case</surname>
          </string-name>
          , How To Become A Centaur,
          <source>Journal of Design and Science</source>
          (
          <year>2018</year>
          ). https://jods.mitpress.mit.edu/pub/issue3-case.
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>E.</given-names>
            <surname>Muller</surname>
          </string-name>
          ,
          <article-title>How ai-human symbiotes may reinvent innovation and what the new centaurs will mean for cities</article-title>
          ,
          <source>Technology and Investment</source>
          <volume>13</volume>
          (
          <year>2022</year>
          )
          <fpage>1</fpage>
          -
          <lpage>19</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>Z.</given-names>
            <surname>Ding</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Smith-Renner</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W.</given-names>
            <surname>Zhang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Tetreault</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Jaimes</surname>
          </string-name>
          ,
          <article-title>Harnessing the power of LLMs: Evaluating human-AI text co-creation through the lens of news headline generation</article-title>
          ,
          <source>in: Findings of the Association for Computational Linguistics: EMNLP</source>
          <year>2023</year>
          ,
          <year>2023</year>
          , pp.
          <fpage>3321</fpage>
          -
          <lpage>3339</lpage>
          . URL: https: //aclanthology.org/
          <year>2023</year>
          .findings-emnlp.
          <volume>217</volume>
          . doi:
          <volume>10</volume>
          .18653/v1/
          <year>2023</year>
          .findings- emnlp.217.
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>Z.</given-names>
            <surname>Zhao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Song</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Duah</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Macbeth</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Carter</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. P.</given-names>
            <surname>Van</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N. S.</given-names>
            <surname>Bravo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Klenk</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Sick</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A. L. S.</given-names>
            <surname>Filipowicz</surname>
          </string-name>
          ,
          <article-title>More human than human: Llm-generated narratives outperform human-llm interleaved narratives</article-title>
          ,
          <source>in: Proceedings of the 15th Conference on Creativity and Cognition</source>
          , CC '
          <fpage>23</fpage>
          ,
          <year>2023</year>
          , pp.
          <fpage>368</fpage>
          -
          <lpage>370</lpage>
          . URL: https://doi.org/10.1145/3591196.3596612. doi:
          <volume>10</volume>
          .1145/3591196.3596612.
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>M.</given-names>
            <surname>Kazemitabaar</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Chow</surname>
          </string-name>
          ,
          <string-name>
            <surname>C. K. T. Ma</surname>
            ,
            <given-names>B. J.</given-names>
          </string-name>
          <string-name>
            <surname>Ericson</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          <string-name>
            <surname>Weintrop</surname>
          </string-name>
          , T. Grossman,
          <article-title>Studying the efect of ai code generators on supporting novice learners in introductory programming</article-title>
          ,
          <source>in: Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems, CHI '23</source>
          ,
          <string-name>
            <surname>Association</surname>
          </string-name>
          for Computing Machinery, New York, NY, USA,
          <year>2023</year>
          . URL: https://doi.org/10.1145/3544548.3580919. doi:
          <volume>10</volume>
          .1145/3544548.3580919.
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>J.</given-names>
            <surname>Liu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Yao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>An</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Q.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <article-title>Peergpt: Probing the roles of llm-based peer agents as team moderators and participants in children's collaborative learning</article-title>
          ,
          <source>in: Extended Abstracts of the CHI Conference on Human Factors in Computing Systems, CHI EA '24</source>
          ,
          <string-name>
            <surname>Association</surname>
          </string-name>
          for Computing Machinery, New York, NY, USA,
          <year>2024</year>
          . URL: https://doi.org/10.1145/3613905.3651008. doi:
          <volume>10</volume>
          .1145/3613905.3651008.
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>I.</given-names>
            <surname>Mondal</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Yuan</surname>
          </string-name>
          ,
          <string-name>
            <surname>A. N</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Garimella</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Ferraro</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Blair-Stanek</surname>
          </string-name>
          ,
          <string-name>
            <surname>B. Van Durme</surname>
          </string-name>
          ,
          <string-name>
            <surname>J. BoydGraber</surname>
          </string-name>
          , ADAPTIVE IE:
          <article-title>Investigating the complementarity of human-AI collaboration to adaptively extract information on-the-fly</article-title>
          , in: O.
          <string-name>
            <surname>Rambow</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          <string-name>
            <surname>Wanner</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          <string-name>
            <surname>Apidianaki</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          <string-name>
            <surname>Al-Khalifa</surname>
            ,
            <given-names>B. D.</given-names>
          </string-name>
          <string-name>
            <surname>Eugenio</surname>
          </string-name>
          , S. Schockaert (Eds.),
          <source>Proceedings of the 31st International Conference on Computational Linguistics</source>
          , Association for Computational Linguistics, Abu Dhabi,
          <string-name>
            <surname>UAE</surname>
          </string-name>
          ,
          <year>2025</year>
          , pp.
          <fpage>5870</fpage>
          -
          <lpage>5889</lpage>
          . URL: https://aclanthology.org/
          <year>2025</year>
          .coling-main.
          <volume>392</volume>
          /.
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>X.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Kim</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Rahman</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Mitra</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Miao</surname>
          </string-name>
          ,
          <article-title>Human-llm collaborative annotation through efective verification of llm labels</article-title>
          ,
          <source>in: Proceedings of the 2024 CHI Conference on Human Factors in Computing Systems, CHI '24</source>
          ,
          <string-name>
            <surname>Association</surname>
          </string-name>
          for Computing Machinery, New York, NY, USA,
          <year>2024</year>
          . URL: https://doi.org/10.1145/3613904.3641960. doi:
          <volume>10</volume>
          .1145/3613904.3641960.
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>C.</given-names>
            <surname>Rastogi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. Tulio</given-names>
            <surname>Ribeiro</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>King</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Nori</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Amershi</surname>
          </string-name>
          ,
          <article-title>Supporting human-ai collaboration in auditing llms with llms</article-title>
          ,
          <source>in: Proceedings of the 2023 AAAI/ACM Conference on AI</source>
          ,
          <string-name>
            <surname>Ethics</surname>
          </string-name>
          , and Society, AIES '
          <volume>23</volume>
          ,
          <year>2023</year>
          , p.
          <fpage>913</fpage>
          -
          <lpage>926</lpage>
          . URL: https://doi.org/10.1145/3600211.3604712. doi:
          <volume>10</volume>
          .1145/ 3600211.3604712.
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>M.</given-names>
            <surname>Vaccaro</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Almaatouq</surname>
          </string-name>
          , T. Malone,
          <article-title>When combinations of humans and ai are useful: A systematic review and meta-analysis</article-title>
          ,
          <source>Nature Human Behaviour</source>
          (
          <year>2024</year>
          ). doi:https://doi.org/ 10.1038/s41562- 024- 02024- 1.
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>V.</given-names>
            <surname>Vats</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. B.</given-names>
            <surname>Nizam</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Liu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Ho</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. S.</given-names>
            <surname>Prasad</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Titterton</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S. V.</given-names>
            <surname>Malreddy</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Aggarwal</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Xu</surname>
          </string-name>
          , et al.,
          <article-title>A survey on human-ai teaming with large pre-trained models</article-title>
          ,
          <source>arXiv preprint arXiv:2403.04931</source>
          (
          <year>2024</year>
          ).
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>