<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <journal-title-group>
        <journal-title>Workshop: Towards the Future of AI-Augmented Human Tutoring in Math Learning, July</journal-title>
      </journal-title-group>
    </journal-meta>
    <article-meta>
      <title-group>
        <article-title>Human Tutors⋆</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Jionghao Lin</string-name>
          <email>Jionghao@cmu.edu</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Danielle R. Thomas</string-name>
          <email>Drthomas@cmu.edu</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Feifei Han</string-name>
          <email>feifei.han@mail.utoronto.ca</email>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Shivang Gupta</string-name>
          <email>Shivang@cmu.edu</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Wei Tan</string-name>
          <email>wei.tan2@monash.edu</email>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Ngoc Dang Nguyen</string-name>
          <email>dan.nguyen2@monash.edu</email>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Kenneth R. Koedinger</string-name>
          <email>Koedinger@cmu.edu</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Carnegie Mellon University</institution>
          ,
          <addr-line>Pittsburgh, PA 15213</addr-line>
          ,
          <country country="US">USA</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Large Language Models, Named Entity Recognition</institution>
          ,
          <addr-line>Tutor Training, Explanatory Feedback, Natural</addr-line>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>Monash University</institution>
          ,
          <addr-line>Clayton, VIC 3800</addr-line>
          ,
          <country country="AU">Australia</country>
        </aff>
        <aff id="aff3">
          <label>3</label>
          <institution>University of Toronto</institution>
          ,
          <addr-line>Toronto, ON M5S 1A1</addr-line>
          <country country="CA">Canada</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2023</year>
      </pub-date>
      <volume>07</volume>
      <issue>2023</issue>
      <abstract>
        <p>Research demonstrates learners engaging in the process of producing explanations to support their reasoning, can have a positive impact on learning. However, providing learners real-time explanatory feedback often presents challenges related to classification accuracy, particularly in domain-specific environments, containing situationally complex and nuanced responses. We present two approaches for supplying tutors real-time feedback within an online lesson on how to give students efective praise. This work-in-progress demonstrates considerable accuracy in binary classification for corrective feedback of efective, or efort-based (</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>1
= 0.811), and inefective, or outcome-based (  1
= 0.350), praise
responses. More notably, we introduce progress towards an enhanced approach of providing explanatory
feedback using large language model-facilitated named entity recognition, which can provide tutors
feedback, not only while engaging in lessons, but can potentially suggest real-time tutor moves. Future
work involves leveraging large language models for data augmentation to improve accuracy, while also
developing an explanatory feedback interface.</p>
    </sec>
    <sec id="sec-2">
      <title>1. Introduction</title>
      <p>
        Tutoring is among the most highly adaptable and consistently successful interventions to
increase student learning [
        <xref ref-type="bibr" rid="ref1 ref2">1, 2</xref>
        ]. However, despite the known positive impacts of tutoring on
achievement, there is a lack of qualified and skilled tutors outside of private, high-income
communities, ready to provide content and socio-motivational support to students [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. Due
to the shortage of professional tutors, often certified teachers and paraprofessionals, the focus
has shifted to preparing novice tutors, such as community volunteers, retired adults, and
college students [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]. The demand for professional development personalized to meet the needs
of nonprofessional and novice tutors is high [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ], with training on social-emotional learning,
Japan
relationship building, and attending to student motivation and self-eficacy as common topics
requested among unskilled tutors [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]. Online, scenario-based lessons on these topics have
been developed to provide situational experiences to inexperienced tutors [21] and preservice
teachers [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ]. The ability to administer real-time explanatory feedback within
constructedresponse questions dealing with common tutoring scenarios (e.g., a student struggling with
motivation) is powerful. Immediate feedback on errors, similar to the feedback received while
engaging in the deliberate practice of responding to situational judgment tests, is described as a
“favorable learning condition,” supporting learning [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ].
      </p>
      <p>
        We present a method of providing tutors real-time explanatory feedback harnessing large
language models (LLMs). Our approach employs a template-based strategy leveraging named
entity recognition (NER), a subtask of natural language processing, that classifies similar pieces
of information [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ]. By tagging similar pieces of information, called named entities (NEs),
for efective and inefective tutor responses, NER becomes a suitable and viable method for
delivering tutors feedback. For example, classifying desired and less-desired tutor responses
on how to efectively praise students yields the following NEs: praising for efort, or
processfocused praise (Efort ); ability- or outcome-focused praise (Outcome); and person-based praise
(Person). Using NER, segments of tutor responses can be systematically identified aligning with
the appropriate NEs. For instance, in Figure 1 the tutor response “Good job! You got the right
answer, and you stuck with it” tags tutor utterances to produce the following NEs: “Good job”
(Outcome) and “stuck with it” (Efort) . Tagged NEs can be used to create the corresponding
templated feedback: “Saying [insert Efort] is a nice example of process-focused praise, which
praises students for their efort.” Conversely, templated feedback for a less-desired response
could be: “Saying [insert Outcome] is praising students for the outcome. You should focus on
praising the students for their efort and process towards learning. Do you want to try responding
again?” The research recommended approach, representing the desired tutor response and,
commonly observed, less-desired responses can be tagged to corresponding NEs. By tagging
pieces of information, aligning with tutoring approaches for responding to the given scenarios,
NER becomes a suitable method for generating templated feedback to tutors.
      </p>
      <p>
        Automatic short answer grading, or the process of automatically scoring learner answers to
constructed-responses questions (often applying several machine learning models), has received
notable attention due to advances in AI-based technologies [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ]. Most automatic short answer
grading methods follow a two-step approach: 1) using a representation, or training set, of
learner responses to train the model, using natural language processing methods, and 2) labeling
responses via a machine learning classifier to predict the learner’s score or performance [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ].
Presently, advanced approaches using LLMs to process learner responses are taking precedence
over traditional, human-identified feature analysis [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ]. Despite the advantages to using LLMs,
there are several limitations: the model not being well-adapted for the nuanced and varied
responses among the learner population; the requirement of having to train one model per
question, with related or follow-up questions being treated as mutually exclusive [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ]; and, the
need for a large number of tutor responses in the representation dataset. This workshop paper
presents a method of providing corrective and explanatory feedback to tutors participating
in an online lesson on giving students efective praise. This work-in-progress introduces an
ongoing efort to enhance approaches towards automatic short answer grading using LLMs,
enabling NER for identifying relevant components of a tutor response. The primary research
questions addressed, include:
RQ1: Can we apply a binary classification method for efectively labeling tutor responses, as
efective or inefective, to provide corrective feedback?
RQ2: How can we enhance past approaches of providing explanatory feedback using
LLMfacilitated named entity recognition to administer templated feedback identifying relevant
parts of the tutor response?
      </p>
    </sec>
    <sec id="sec-3">
      <title>2. Related Work</title>
      <sec id="sec-3-1">
        <title>2.1. Feedback Generation</title>
        <p>
          Feedback can have a profound impact on learning and achievement; however, its influence may
be beneficial or detrimental to learning depending on type and delivery [ 8, 9, 10]. Feedback is
most beneficial within the learning context, delivered after the learner has engaged with the
initial instruction, and when addressing misconceptions or faulty reasoning [8]. Immediate,
explanatory feedback, or feedback detailing the reasoning why a response is desired or not,
assists learners with participating in deliberate practice. The online lesson has tutors engaging
in the deliberate practice of responding to a common tutoring scenario (i.e., a student struggling
to stay motivated) by asking them how to best respond. Tutors then explain their reasoning
and observe the most-desired approach receiving feedback on their chosen selected response
option [
          <xref ref-type="bibr" rid="ref3">3, 11</xref>
          ]. An expansion of this previous work is to provide explanatory feedback to tutors
on their textual replies to the constructed-response questions. Generating explanatory feedback
to tutors using enhanced approaches, such as using LLMfacilitated NER, shows promise as a
method of providing accurate and timely feedback to tutors. The creation of templated feedback,
including specific references to desired and less-desired elements of the tutor responses, is
influenced by earlier results on the efectiveness of having a rich, data-driven error diagnosis
taxonomy driving template-based feedback [12].
        </p>
      </sec>
      <sec id="sec-3-2">
        <title>2.2. Named Entity Recognition</title>
        <p>
          A named entity (NE) is a word or phrase distinct from a set of words that have similar
attributes [
          <xref ref-type="bibr" rid="ref6">6</xref>
          ]. For example, in the text “John said that Pittsburgh is wonderful in the winter”,
“John”, “Pittsburgh”, and “winter” are considered NEs, which represent a person, location, and
time, respectively. Named entity recognition (NER) is a fundamental task in natural language
processing, which aims to automatically locate NE in the text and classify them into diferent
categories such as person, organization, and location [
          <xref ref-type="bibr" rid="ref6">6</xref>
          ]. In the example, to identify the NE
“Pittsburgh”, a NER model first locates the position of “Pittsburgh” in the text and then classifies
the entity “Pittsburgh” into the category of location. As discussed by [
          <xref ref-type="bibr" rid="ref6">6</xref>
          ], there are two primary
categories of name entities: (i) generic (e.g., person and organization) and (ii) domain-specific
(e.g., enzymes and genes). Since the present work aims to investigate the potential of the NER
model in providing explanatory feedback based on learning principles in the previous research
[
          <xref ref-type="bibr" rid="ref3">3</xref>
          ], we focus on the domain-specific NER model scheme. In the educational domain, researchers
have conducted NER for automatic text assessment [13]. However, the use of the NER model is
rarely used in feedback generation. Therefore, our approach aims to employ a NE recognition
model to highlight the NEs within tutors’ responses, which can be used to create templated
explanatory feedback to tutors to increase tutor learning.
        </p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>3. Method</title>
      <sec id="sec-4-1">
        <title>3.1. Dataset</title>
        <p>
          The dataset consisted of tutor responses to constructed-response questions within the Giving
Efective Praise lesson, comprising a total of 65 volunteer tutors. The tutor demographics
are: 52% White, 18% Asian, 52% male, and slightly more than half are reportedly 50 years
of age or older. The development of Giving Efective Praise involved collaboration between
the tutoring organization’s director and researchers to ensure accurate operationalization of
efective praise strategies within the tutoring environment. Lesson scenarios were chosen to
enhance face validity by aligning with typical situations encountered by tutors in the field [
          <xref ref-type="bibr" rid="ref3">3</xref>
          ].
Giving Efective Praise aims to support tutors with increasing student motivation by providing
efective praise, identifying its key features, and employing strategies to deliver praise and
feedback. In accordance with the lesson’s construct, efective praise should be: (1) sincere,
earned, and truthful; (2) specific by giving details of a student’s strengths; (3) immediate, with
praise given right after the student’s action; (4) authentic, avoiding repetitive phrases like “great
job” which diminishes meaning and becomes predictable, and (5) focused on the learning process
rather than innate ability [
          <xref ref-type="bibr" rid="ref3">3</xref>
          ].
        </p>
        <p>Based on characteristics of praise types in the literature, tutor praise statements can be
categorized into three diferent types: efort-based ( Efort ), outcomebased (Outcome), and
person-based (Person). Efort-based praise is a researchshown productive praise type, focusing
on the learning process (e.g., “I like how you worked hard to...” ). Outcome-based praise showcases
student’s achievements, such as getting an A on an assignment or getting a problem correct, and
is often, but not always, associated with unproductive praise (e.g., “Great job!” ). Person-based
praise suggests student’s success is caused from fixed qualities outside of student’s control (e.g.,
“You are so talented.” ) and, similar to outcome-based praise, is often associated with unproductive
praise [14].</p>
        <p>The Giving Efective Praise dataset contains 129 tutor responses categorized by praise type.
Because only one person-based praise statement (i.e., “You are very smart” ) was identified from
the dataset, person-based praise was not included in the analysis. It should be noted that a
tutor’s response can include more than one praise type. For example, the statement, “Great job!
I like how you worked hard on completing that task,” encompasses outcome- and efortbased
praise statements. Similarly, a tutor response may not contain any praise types such as, “Let’s
work together.”</p>
      </sec>
      <sec id="sec-4-2">
        <title>3.2. RQ 1: Binary Classification for Corrective Feedback</title>
        <p>
          To accurately identify diferent types of praise, we aim to conduct multi-label classification.
Relying on the praise framework proposed by [
          <xref ref-type="bibr" rid="ref3">3</xref>
          ]. we recruited an educational expert to annotate
each type of praise for the tutor’s responses in a binary form. The distribution of annotated
praise in tutor responses are as follows: 52 responses contained efort-based praise only; 29
responses contained both efort- and outcome-based praise; 26 contained outcome-based praise
only; and, the remaining 22 responses lacked any mention of neither efort or outcomebased
praise. Then, we trained the annotated responses on classifiers. Inspired by the efectiveness of
the BERT model on the educational classification tasks such as tutoring dialogue classification
[15, 16], we employed the BERT model to identify each type of praise in the tutors’ responses.
To train and evaluate the BERT model, we randomly split the dataset (i.e., annotated tutor
responses) into training, validation, and testing set in the ratio of 70%, 10%, 20%, respectively, as
suggested by [17]. The classification performance of the BERT model was measured by accuracy
and F1 score.
        </p>
      </sec>
      <sec id="sec-4-3">
        <title>3.3. RQ 2: Named Entity Recognition to Generate Explanatory Feedback</title>
        <p>
          In order to generate explanatory feedback to tutors, firstly, the categorization of relevant parts
of responses need to be identified through use of NEs. We refer to the annotation scheme by
Thomas et al. [
          <xref ref-type="bibr" rid="ref3">3</xref>
          ] introduced previously and annotate the NEs representing attributes associated
with Efort and Outcome, for 129 tutor responses. In line with the NER annotation in previous
works [
          <xref ref-type="bibr" rid="ref6">6, 18</xref>
          ], we apply the same BIO-tagging scheme to this present work, that is, B represents
the beginning position of the NE in the text, I represents the inside position of the NE in the
text, and O represents outside of NE. For example, when annotating praise NEs for a tutor’s
praise “You are doing a great job”, the word “great” is identified as the beginning (i.e., Bout
of the NE Outcome and “job” is identified inside (i.e., Iout) of the NE. The remaining text in
the response is identified as the outside (i.e., O) of the NE. After annotating the NEs for each
tutor’s response, we employed the BERT model to identify the NE from the tutors’ responses.
The dataset (annotated with NEs) was also divided into training, validation, and testing set
in the ratio of 70%:10%:20%, respectively. The statistics of NER annotation data are shown in
Table 1 which presents O as the major tag in our dataset. Informed by the previous study [18],
predicting O would not enhance the evaluation score of the NER model, our study also did not
take accurate predictions on O when calculating the performance score. To measure the NER
model performance, we used the F1 score in line with the recent works on NER task [
          <xref ref-type="bibr" rid="ref6">6, 18</xref>
          ].
        </p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>4. Results</title>
      <p>O</p>
      <sec id="sec-5-1">
        <title>4.1. RQ1: Identifying the correct type of praise</title>
        <p>First, a multi-label classification was implemented by using the case-sensitive BERT base model 1
[19] to identify the efective type (i.e., Efort and Outcome) of praise from tutor responses. To
minimize the potential impact of random variation, the model was trained on 10 diferent
random seeds and performance was evaluated using the classification of identifying each type
of praise. Table 2 illustrated the efectiveness of the BERT model in accurately tagging Efort ,
demonstrating notably high performance with an average classification accuracy of 0.731 and F1
score of 0.811. The results indicated that the BERT model could efectively tag Efort , which could
further help the provision of corrective feedback to inform the novice tutors on providing
efortbased praise. However, the BERT model’s performance in tagging Outcome was less successful.
The average F1 score for recognizing Outcome was 0.350 with a standard deviation of 0.235.
As described the distribution of annotated praise in Section 3.2, the number of tutor responses
tagging Outcome might be inadequate for the BERT model to identify the responses containing
Outcome accurately. Additionally, the standard deviation of classification performance for
tagging Outcome (SD = 0.235) was four times larger compared to the standard deviation of
tagging Efort (SD = 0.046). The possible explanation for this could be the inadequate number
of Outcome instances in the test set. Thus, future studies should annotate more tutor responses
labeling Outcome to improve the BERT model’s performance.</p>
      </sec>
      <sec id="sec-5-2">
        <title>4.2. RQ 2: Identifying and labeling praise statements in tutor responses</title>
        <p>The BERT model was employed in using the NER approach. To mitigate random variation, 10
diferent random seeds were used to evaluate performance of the NER model, ensuring reliable
estimations of the model’s performance. The average F1 score of the model was 0.202 and
the standard deviation was 0.039, with the model efectively identifying certain praise entities.
Table 3 presented examples of tutor responses with labeled utterances associated with the
corresponding NE, displaying Efort , (highlighted in blue) and Outcome (highlighted in red).
Case 1 showed that the model could accurately identify the location of the praise in the text (i.e.,
the text highlighted in blue) and predict the accurate entity type (i.e., Efort ). Then, the model
failed to annotate the NEs for some responses (e.g., Case 2 in Table 3). It should be noted that
the classification performance of the NER model still had space to improve the performance.
One of the major reasons was that the annotated dataset was limited or low-resourced [16, 18].
The model might not have a suficient dataset to train and test the model performance. In
Section 5.2, we summarized two solutions to enhance the NER model’s performance: i) data
augmentation approaches; ii) and AUC Maximization approaches.</p>
        <p>Through evaluating and analyzing the results of NER, we noted that there is a need for a
more nuanced measure that can acknowledge: partial overlap (e.g., Case 3); and true negative
prediction results (e.g., Case 4). In Table 3, Case 3, the model’s prediction exhibits a degree
of accuracy (i.e., “Good job” tagged as Outcome, “you stuck with it” tagged as Efort ) but lacks
complete correctness (i.e., “I’m proud of what you have done” mislabeled as Efort ). Nevertheless,
the former accurate labeling of NEs can still be used for guiding tutors in providing efective
praise. Thus, the prediction for Case 3 is deemed partially accurate. Future research endeavors
should focus on developing a measure to calculate partial accuracy, such as computing the
intersection over union of the number of tokens in the predicted text and the desired text.
Additionally, the NER model could also make true negative predictions where the tutor response
did not contain any praise entities and was identified as having none of NEs (i.e., Case 4 in
Table 3). As discussed in Section 3.3, there lacked accurate predictions on the O tag when
calculating the classification performance score. However, it is also important to identify the
tutor’s responses that only contain the O tag (i.e., none of the NEs) since it could indicate that
the tutors might not understand how to deliver correct praise. In Table 3, the tutor response
in Case 4 did not contain any praise entities and this response was not related to any type of
praise. The NER model could successfully identify that the response did not contain any type of
praise. Based on the model prediction, feedback can be generated to guide tutors on providing
praise that corresponds with the Efort and Outcome named entities.</p>
      </sec>
    </sec>
    <sec id="sec-6">
      <title>5. Discussion and Conclusion</title>
      <p>The construction of automatic short answer grading with the capability of providing explanatory
feedback is a longstanding task towards delivering timely, specific, and personalized feedback
to learners. This study employed large language models to facilitate the provision of corrective
and explanatory feedback to tutors, with the main findings summarized in two folds: (1) Large
language models (e.g., BERT) have the potential to identify the efort-based praise, which can be
used to provide corrective feedback to novice tutors on the appropriate use of efort-based praise
to students. (2) Large language models-facilitated named entity recognition (NER) can highlight
the key terms associated with praise types from tutors’ responses. The highlighted terms can
then be integrated into template-based feedback, which can provide real-time explanatory
feedback to tutors to enhance tutor learning.</p>
      <sec id="sec-6-1">
        <title>5.1. Implications</title>
        <p>Incorporation of a binary classifier can provide automatic corrective feedback. The
developed classifier can be used to determine the correctness of novice tutors in providing
diferent types of praise. The predicted classifier results can be further integrated into the
provision of corrective feedback, which is essential in the learning process since corrective
feedback can assist the feedback recipients in identifying errors and enhancing understanding
[20]. Through the integration of the classifier within the system, we aim to provide automatic
corrective feedback to tutors in dispensing various forms of praise.</p>
        <p>Providing automatic templated feedback enhances tutor learning. To better facilitate
the provision of corrective feedback, this study further investigated the potential of NER in
identifying the words within the tutors’ responses that correspond to the correct types of praise
(i.e., Efort and Outcome). The words identified as correct praise in the tutors’ responses can
be integrated into a system of providing explanatory feedback. Figure 1, within the Section
of Introduction, illustrates an example of providing a tutor templated explanatory feedback
using an integrated NER interface. Referencing the interface, when a tutor composes praise that
includes efort- and/or outcome-based praise, the system will label the Efort and/or Outcome
NEs and further provide explanatory feedback to the tutor. As informed by the suggestions
of efective feedback [ 20], incorporating explanations into feedback can help the tutor better
understand the lesson objectives and content. To this end, we believe that integrating the NER
model into our system could support the tutor’s learning process.</p>
      </sec>
      <sec id="sec-6-2">
        <title>5.2. Limitations and Future Work</title>
        <p>Managing low confidence prediction using the feedback interface. The confidence level
of the model’s predictions is a critical aspect to consider in real-world applications. The model
confidence level could afect people’s belief in the model’s accuracy [ 21]. Thus, when the model
presents low confidence in predicting an instance, it poses a challenge. In such situations, it
would be beneficial to design the feedback interface that presents the uncertainties to the learner.
For example, when a model’s confidence level on a prediction is below a certain threshold, our
template-based feedback could provide hedged responses, such as “Saying “you are committed”
might be an example of praising efort. Do you want to explain your reasoning?” . This approach
not only helps uphold the credibility of the system but also invites learners to engage critically
with the predictions. Future work entails providing hedged feedback responses ofering learners
the opportunity to explain their reasoning, along with other strategies for efectively managing
low confidence predictions in the feedback interface.</p>
        <p>Enhancing the evaluation metrics for NER. As indicated by Case 3 and Case 4 in Table 3,
respectively, the mode’s prediction may demonstrate partial correctness, for incidences where
the predicted text only partially matches the desired text or, true negative predictions, where the
tutor’s response is accurately predicted to contain no praise entities. We argue that both partial
correctness and the true negative predictions are useful in providing explanatory feedback and
thus, both types of predictions should be credited. However, the traditional NER measure (F1
score) might not fully account for partial correctness and true negative predictions. Therefore,
future work should explore the development of measures, such as the degree of dissimilarity
between sets via calculation of intersection over union [22] to account for these cases, thereby
leading to a more comprehensive evaluation of the model’s performance.</p>
        <p>Improving NER performance through data augmentation. By examining the model of
NER for identifying the praise entity from the lesson of Give efective praise , we found that
our annotated NEs might not be suficient to train the model, which is under a low-resource
data scenario [18]. To address this issue, we aim to collect more real-world data and explore
widely-used data augmentation approaches (e.g., oversampling and synonyms replacement)
[23] and ChatGPT-generated training instances [24] to improve NER model performance.
Robust models for enhancing the performance of NER. An alternative solution to
lowresource dataset is to employ robust machine learning models. As indicated in the statistic of our
dataset (Table 1), more than 70% of annotations were annotated as the O tag, which was highly
imbalanced. To achieve satisfactory performance under the low-resource and imbalance settings,
[18] proposed to use AUC Maximization approaches for the NER task in the biomedical field,
which efectively overcome challenges among low-resources and imbalanced class distribution.
Thus, we aim to further examine the eficacy of AUC Maximization approaches on recognizing
the praise entities.</p>
        <p>
          Generalizability across tutor lessons. Our ultimate goal is to provide automatic feedback
to all novice tutors who participate in our training sessions and assist them to understand the
efective ways to teach students. Thus, a qualified tutor should be able to comprehend all the
training lessons. Though our study examined the potentials of labeling tutor responses and
providing explanatory feedback for Giving Efective Praise lessons, it is necessary to investigate
our proposed methods in other lessons such as Responding to Student’s Errors and Learning
What Students Know discussed in previous work [
          <xref ref-type="bibr" rid="ref3">3</xref>
          ].
        </p>
      </sec>
    </sec>
    <sec id="sec-7">
      <title>6. Acknowledgments.</title>
      <p>This work is supported with funding from the Richard King Mellon Foundation (Grant #10851)
and the Heinz Endowments (E6291). Any opinions, findings, and conclusions expressed in this
material are those of the authors. We would like to thank Dr. Ralph Abboud for his guidance
and recommendations regarding the use of large language models and the application of named
entity recognition.
[8] J. Hattie, H. Timperley, The power of feedback-review of educational research, American</p>
      <p>Education Research Association and SAGE (2011) 86.
[9] T. Ryan, M. Henderson, K. Ryan, G. Kennedy, Designing learner-centred text-based
feedback: a rapid review and qualitative synthesis, Assessment &amp; Evaluation in Higher
Education 46 (2021) 894–912.
[10] J. Lin, W. Dai, L.-A. Lim, Y.-S. Tsai, R. F. Mello, H. Khosravi, D. Gasevic, G. Chen,
Learnercentred analytics of feedback content in higher education, in: LAK23: 13th International
Learning Analytics and Knowledge Conference, 2023, pp. 100–110.
[11] D. R. Chine, P. Chhabra, A. Adeniran, S. Gupta, K. R. Koedinger, Development of
scenariobased mentor lessons: an iterative design process for training at scale, in: Proceedings of
the Ninth ACM Conference on Learning@ Scale, 2022, pp. 469–471.
[12] V. Aleven, O. Popescu, K. R. Koedinger, Towards tutorial dialog to support self-explanation:
Adding natural language understanding to a cognitive tutor, in: Proceedings of Artificial
Intelligence in Education, 2001, pp. 246–255.
[13] C. Walter, Increasing teachers’ trust in automatic text assessment through named-entity
recognition, in: International Conference on Artificial Intelligence in Education, Springer,
2022, pp. 191–194.
[14] M. L. Kamins, C. S. Dweck, Person versus process praise and criticism: implications for
contingent self-worth and coping., Developmental psychology 35 (1999) 835.
[15] J. Lin, S. Singh, L. Sha, W. Tan, D. Lang, D. Gašević, G. Chen, Is it a good move? mining
efective tutoring strategies from human–human tutorial dialogues, Future Generation
Computer Systems 127 (2022) 194–207.
[16] J. Lin, W. Tan, N. D. Nguyen, D. Lang, L. Du, W. Buntine, R. Beare, G. Chen, D. Gašević,
Robust educational dialogue act classifiers with low-resource and imbalanced datasets,
in: International Conference on Artificial Intelligence in Education, Springer, 2023, pp.
114–125.
[17] A. Gholamy, V. Kreinovich, O. Kosheleva, Why 70/30 or 80/20 relation between training
and testing sets: A pedagogical explanation (2018).
[18] N. D. Nguyen, W. Tan, L. Du, W. Buntine, R. Beare, C. Chen, Auc maximization for
lowresource named entity recognition, in: Proceedings of the AAAI Conference on Artificial
Intelligence, volume 37, 2023, pp. 13389–13399.
[19] J. Devlin, M.-W. Chang, K. Lee, K. Toutanova, Bert: Pre-training of deep bidirectional
transformers for language understanding, in: Proceedings of the 2019 Conference of the
NAACL-HLT, Volume 1, 2019, pp. 4171–4186.
[20] A. C. Butler, N. Godbole, E. J. Marsh, Explanation feedback is better than correct answer
feedback for promoting transfer of learning., Journal of Educational Psychology 105 (2013)
290.
[21] A. Rechkemmer, M. Yin, When confidence meets accuracy: Exploring the efects of
multiple performance indicators on trust in machine learning models, in: Proceedings of
the 2022 CHI conference, 2022, pp. 1–14.
[22] M. Levandowsky, D. Winter, Distance between sets, Nature 234 (1971) 34–35.
[23] S. Y. Feng, V. Gangal, J. Wei, S. Chandar, S. Vosoughi, T. Mitamura, E. Hovy, A survey of
data augmentation approaches for nlp, in: Findings of the ACL: ACL-IJCNLP 2021, 2021,
pp. 968–988.
[24] D. Thomas, S. Gupta, K. Koedinger, Comparative analysis of learnersourced human-graded
and ai-generated responses for autograding online tutor lessons, in: Artificial Intelligence
in Education. 24th International Conference, AIED 2023, Tokyo, Japan July 3–7, 2023,
Springer, 2023.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>M. A.</given-names>
            <surname>Kraft</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Falken</surname>
          </string-name>
          ,
          <article-title>A blueprint for scaling tutoring and mentoring across public schools</article-title>
          ,
          <source>AERA Open 7</source>
          (
          <year>2021</year>
          )
          <fpage>1</fpage>
          -
          <lpage>21</lpage>
          . URL: https://journals.sagepub.com/doi/full/10.1177/ 23328584211042858.
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>V. Q.</given-names>
            <surname>Andre Joshua</surname>
          </string-name>
          <string-name>
            <surname>Nickow</surname>
          </string-name>
          , Philip Oreopoulos,
          <article-title>The impressive efects of tutoring on prek-12 learning: A systematic review and meta-analysis of the experimental evidence</article-title>
          ,
          <year>2020</year>
          . URL: http://www.edworkingpapers.com/ai20-
          <fpage>267</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>D.</given-names>
            <surname>Thomas</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Yang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Gupta</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Adeniran</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Mclaughlin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Koedinger</surname>
          </string-name>
          ,
          <article-title>When the tutor becomes the student: Design and evaluation of eficient scenario-based lessons for tutors</article-title>
          ,
          <source>in: LAK23: 13th International Learning Analytics and Knowledge Conference</source>
          ,
          <year>2023</year>
          , pp.
          <fpage>250</fpage>
          -
          <lpage>261</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>M.</given-names>
            <surname>Thompson</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Owho-Ovuakporie</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Robinson</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y. J.</given-names>
            <surname>Kim</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Slama</surname>
          </string-name>
          , J. Reich,
          <article-title>Teacher moments: A digital simulation for preservice teachers to approximate parent-teacher conversations</article-title>
          ,
          <source>Journal of Digital Learning in Teacher Education</source>
          <volume>35</volume>
          (
          <year>2019</year>
          )
          <fpage>144</fpage>
          -
          <lpage>164</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>K. R.</given-names>
            <surname>Koedinger</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P. F.</given-names>
            <surname>Carvalho</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Liu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E. A.</given-names>
            <surname>McLaughlin</surname>
          </string-name>
          ,
          <article-title>An astonishing regularity in student learning rate</article-title>
          ,
          <source>Proceedings of the National Academy of Sciences</source>
          <volume>120</volume>
          (
          <year>2023</year>
          )
          <article-title>e2221311120</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>J.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Sun</surname>
          </string-name>
          , J. Han,
          <string-name>
            <given-names>C.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <article-title>A survey on deep learning for named entity recognition</article-title>
          ,
          <source>IEEE Transactions on Knowledge and Data Engineering</source>
          <volume>34</volume>
          (
          <year>2020</year>
          )
          <fpage>50</fpage>
          -
          <lpage>70</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>M.</given-names>
            <surname>Zhang</surname>
          </string-name>
          , S. Baral,
          <string-name>
            <given-names>N.</given-names>
            <surname>Hefernan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Lan</surname>
          </string-name>
          ,
          <article-title>Automatic short math answer grading via in-context meta-learning</article-title>
          ,
          <source>arXiv preprint arXiv:2205.15219</source>
          (
          <year>2022</year>
          ).
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>