<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <journal-title-group>
        <journal-title>Workshop: Towards the Future of AI-Augmented Human Tutoring in Math Learning, July</journal-title>
      </journal-title-group>
    </journal-meta>
    <article-meta>
      <title-group>
        <article-title>Learning⋆</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Beverly Woolf</string-name>
          <email>bev@umass.edu</email>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Margrit Betke</string-name>
          <email>betke@bu.edu</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Hao Yu</string-name>
          <email>haoyu@bu.edu</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Sarah Adel Bargal</string-name>
          <email>sarah.bargal@georgetown.edu</email>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Ivon Arroyo</string-name>
          <email>arroyo@umass.edu</email>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>John Magee</string-name>
          <email>jmagee@clarku.edu</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Danielle Allessio</string-name>
          <email>allessio@umass.edu</email>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>William Rebelsky</string-name>
          <email>wrebelsky@umass.edu</email>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Boston University</institution>
          ,
          <addr-line>Massachusetts, MA 02215</addr-line>
          ,
          <country country="US">USA</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Clark University</institution>
          ,
          <addr-line>Worcester, MA 01610</addr-line>
          ,
          <country country="US">USA</country>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>Georgetown University</institution>
          ,
          <addr-line>Washington, D.C. 20057</addr-line>
          ,
          <country country="US">USA</country>
        </aff>
        <aff id="aff3">
          <label>3</label>
          <institution>University of Massachusetts-Amherst</institution>
          ,
          <addr-line>Massachusetts, MA 01003</addr-line>
          ,
          <country country="US">USA</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2023</year>
      </pub-date>
      <volume>07</volume>
      <issue>2023</issue>
      <abstract>
        <p>The future of AI-assisted individualized learning includes computer vision to inform intelligent tutors and teachers about student afect, motivation and performance. Facial expression recognition is essential in recognizing subtle diferences when students ask for hints or fail to solve problems. Facial features and classification labels enable intelligent tutors to predict students' performance and recommend activities. Videos can capture students' faces and model their efort and progress; machine learning classifiers can support intelligent tutors to provide interventions. One goal of this research is to support deep dives by teachers to identify students' individual needs through facial expression and to provide immediate feedback. Another goal is to develop data-directed education to gauge students' pre-existing knowledge and analyze real-time data that will engage both teachers and students in more individualized and precision teaching and learning. This paper identifies three phases in the process of recognizing and predicting student progress based on analyzing facial features: Phase I: Collecting datasets and identifying salient labels for facial features and student attention/engagement; Phase II: Building and training deep learning models of facial features; and Phase III: Predicting student problem-solving outcome.</p>
      </abstract>
      <kwd-group>
        <kwd>facial expression recognition</kwd>
        <kwd>intelligent tutors</kwd>
        <kwd>machine learning</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>
        As students engage with online learning technologies, they experience a variety of emotions
(confusion, excitement, frustration, anxiety) and various levels of engagement, depending on a
combination of motivation, mood, and background knowledge. Students’ afective states and
engagement are tightly correlated with learning gains [
        <xref ref-type="bibr" rid="ref1 ref2">1, 2</xref>
        ]. Having afective and engagement
information accessible to teachers (or digital tutors) can aid in understanding students’ progress
and suggest when and which students need further assistance.
Japan
      </p>
      <p>This paper describes the design and evaluation of a suite of tools for facial expression
recognition called FaceReaders, or tools that detect users’ faces and gestures for the purpose
of identifying and predicting engagement, motivation, and future behavior. For example, if an
intelligent tutor can predict that a student’s future behavior will be to “Give up”, the system might
provide an intervention (example problem, formula, hints or easier problem). Ethnographic
surveys are used to first identify activities and concrete questions that teachers ask in real-time:
Who needs my help most right now? Who is wheel-spinning right now? Is the class ready to
move to the next topic? How often is Arjun skipping, guessing or giving up? Answers to these
questions help teachers strategize responses, adapt class pedagogy and provide interventions.
This paper presents a survey of research activities addressed by our laboratories towards the
future of computer vision-augmented tutoring in math learning. One goal is to design, develop
and evaluate these tools. Specifically, we describe Phase I: Collecting datasets and identifying
labels for faces and gestures; Phase II: Identifying students’ attention in math learning, and
Phase III: Predicting problem-solving outcome.</p>
    </sec>
    <sec id="sec-2">
      <title>2. Related and Prior Work</title>
      <p>
        Intelligent Tutoring Systems. Intelligent Tutoring Systems (ITS) produce learning gains with
efects close to one letter grade improvement [
        <xref ref-type="bibr" rid="ref4 ref5">4, 5</xref>
        ]. Students using these tutors outperform
students from conventional classes in 92 percent of the controlled evaluations with performance
measured twice as high as for students using typical (non-intelligent systems) [
        <xref ref-type="bibr" rid="ref6 ref7 ref8">6, 7, 8</xref>
        ]. One
meta-analysis of findings from controlled ITS evaluations shows that test scores increased by
0.66 standard deviations over conventional levels, or from the 50th to the 75th percentile [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ]. In
an emotion-sensitive ITS, student emotion is automatically detected through facial expressions,
body posture and gestures, speech, text, or physiological metrics. Measuring physiological
signals is the rarest metric as it requires an intrusive learning experience [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ]. Strain and D’Mello
[
        <xref ref-type="bibr" rid="ref10">10</xref>
        ] studied the role of emotion in ITS engagement, task persistence, and learning gain. D’Mello
et al. used gaze prediction based on natural language dialogues. Their system responded to
students’ boredom and tried to engage students; boredom is one of the most frequent states a
student experiences during learning and negatively correlates with learning gain.
      </p>
      <p>
        Visual Facial Action Units. A correlational analysis found evidence for relationships
between visual facial Action Unit (AU) factors and self-reported traits such as academic efort,
study habits, and interest in subjects [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ]. Detected facial cues gave insight into the learner’s
mental state, but potential cues to predict learning did not ofer a consistent signal. Behavior
prediction can support improved learning by tailoring the interventions of the ITS to the
predicted actions of the student. Our work focuses on using predicted deep afect embeddings
learned from a large facial afect dataset to improve behavior prediction in an ITS [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ].
      </p>
      <p>
        Transfer Learning in Facial Analysis. Prior research in transfer learning for facial analysis
applications mostly focuses on transfer learning within the same application to improve results
or bridge domain gaps, e.g., personalize a prediction system to specific individuals [
        <xref ref-type="bibr" rid="ref12 ref13 ref14">13, 14, 12</xref>
        ],
ifne-tune neural networks pre-trained on external datasets for a similar prediction task [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ],
or pre-training on a related facial analysis task [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ]. In contrast, our work tackles transfer
learning across domains and tasks, which is a form of transductive transfer learning [
        <xref ref-type="bibr" rid="ref17">17</xref>
        ]. We
explore transfer learning from the facial analysis problem of in-the-wild afect recognition of
afect to a webcam video behavior prediction problem. Work exists to explore transfer learning
from facial analysis to behavior analysis such as AutoRate that uses VGGFace facial recognition
embeddings to improve predictions of driver attention scores [
        <xref ref-type="bibr" rid="ref18">18</xref>
        ].
      </p>
      <p>
        Interventions in Online Tutors. Afective messages delivered by avatars and empathetic
messages respond to students’ recent emotions [
        <xref ref-type="bibr" rid="ref19">19</xref>
        ]. Interventions in MathSpring ITS led to
improved grades in state standardized exams [
        <xref ref-type="bibr" rid="ref20">20</xref>
        ] as well as influencing students’ perceptions of
themselves as learners [
        <xref ref-type="bibr" rid="ref21">21</xref>
        ]. Empathetic characters generate superior results to improve student
interactions with the system and address negative student emotions [
        <xref ref-type="bibr" rid="ref22">22, 23, 24</xref>
        ].
      </p>
      <p>MathSpring Intelligent Tutoring System. MathSpring.org, is a freely available game-like
system [25]; students grow gardens that visually represent their mathematics progress. It is
multimedia in that it provides both audio and visual support, intelligent in that it builds an
internal model of students, as would a good teacher, and personalized in that it provides remedial
tutoring when needed. For instance, MathSpring might be set to teach Grade 7 mathematics and
will seamlessly move back to grade 6 and 5 material as needed, in a way that is unnoticeable
to the student. Animated learning companions (LCs) provide emotional support and build
students’ socio-emotional skills, instilling a growth mindset [26], encouraging students to
consider mistakes as a natural part of learning and stating that intelligence is malleable. LCs
support students’ learning processes and use well established instructional strategies. Students’
emotions are assessed regularly and the tutor ofers support when students become frustrated
or anxious [25]. MathSpring provides a positive afective impact on students in the USA and
Argentina. In controlled studies, students showed an increase in mathematics and reading
comprehension. The combination of the character and role design of pedagogical agents makes
a significant positive impact on student learning and behavior (e.g., [ 27, 28, 29, 30, 31]).</p>
    </sec>
    <sec id="sec-3">
      <title>3. Phase I: Collecting Datasets and Identifying Labels for Face and Gestures</title>
      <p>Students experience a variety of emotions while working online, e.g., boredom, frustration,
interest, and surprise [32] and these displayed emotions correlate well with students’ achievement in
the learning task [33]. Equipping an intelligent tutor with the ability to interpret such afective
signals could potentially enable it to monitor students’ progress, provide timely interventions
and present appropriate afective reactions via a virtual tutor. For example, machine learning
classifiers can be trained to recognize the subtle diferences in facial behavior between when
a student requires hints to solve a problem, see Figures 1- 2, so that the tutor can intervene
accordingly.</p>
      <p>
        During Phase I, we collected and annotated databases of facial afect videos of students
interacting with MathSpring, an intelligent tutor, Figure 3. Considering the dearth of large-scale,
publicly available afect video datasets in learning and education settings, we made these datasets
and annotations public [
        <xref ref-type="bibr" rid="ref12 ref22 ref3">22, 3, 12</xref>
        ]. The video datasets consist of college students solving math
problems with a front facing camera collecting visual feedback of student gestures. Datasets
consist of video clips which were obtained by trimming the raw videos based on problem start
and end times recorded in MathSpring’s log file.
      </p>
      <p>
        In the initial dataset collection [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ], we collected 1596 video clips of 30 diferent students
solving math problems. Each video clip is automatically annotated by MathSpring’s learning
log data. The labels used to annotate the video clips are: ATT (student did not see any hints
but solved the question after 1 incorrect attempt), GIVEUP (student performed some action
but did not solve the problem at all), GUESS (student did not see hints, but solved the question
after greater than 1 incorrect attempts), NOTR (student performed some action, but the first
action was too rapid for him to have read the problem), SHINT (student eventually got the
correct answer after seeing one or more hints), SKIP (student skipped problem with no action)
and SOF (student answered correctly in first attempt, without seeing any hints). In the next
iteration [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ], we collected and annotated an extended version of [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ], resulting in a dataset of
2749 video clips of math problems solved by 54 students. Next, we provide additional labels
for video frames for a subset of 400 video clips of 19 diferent students. The labels used to
annotate the extracted video frames are: “looking at their screen”, “looking at their paper”, or
“wandering”. This resulted in 18,721 annotated frames. An interface was presented to MTurk
workers for labeling to indicate whether students’ attention was engaged or wandering. Each
of the 18,721 frames was assigned to three diferent crowdworkers and we processed 56,163
(18,721 * 3) results.
      </p>
      <p>
        Our contributions are summarized as follows:
• Introduced a unique video dataset of 1596 student interactions labeled for problem
outcome, extracted from more than 30 hours of raw video data [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ];
• Augmented the dataset of [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ] to include 2749 student interactions labeled for problem
outcome [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ];
• Provided annotations for a subset of 400 student interactions labeled for
attention/engagement [
        <xref ref-type="bibr" rid="ref22">22</xref>
        ];
• Made these datasets publicly accessible to encourage and foster research in the intersection
of the education and computer vision communities; and
• Provided a set of baseline results predicting student learning outcomes and
attention/engagement solely from facial afect signals.
      </p>
      <p>Analysis of the Outcome Classes. We provided an exploratory analysis of the diferent
problem outcome classes that result when students interact with MathSpring, using typical
facial action unit activations to analyze students’ faces. We developed baseline models to predict
students’ problem outcome labels (e.g., ask for hints, solve problem) and discussed how early
problem outcome labels can be forecasted to provide possible interventions. Each data instance
in the data set consists of a video clip of a student working on a problem and its corresponding
label of the student’s problem-solving behavior. We researched baseline models to investigate
the problem of directly predicting the learning outcome of students solely from afect signals.
To visually illustrate the prediction of problem outcomes and to understand student behavior,
we present visual examples of an eighth grade student using MathSpring, see Figure 2. The
student used MathSpring for one session of around 20 minutes and consented to have his face
and screen recorded. Figure 2 shows the evolution of student expressions and gestures, and their
corresponding problem outcomes. When the student successfully solves the problem on the
ifrst attempt (SOF), we observe that he focused tightly on the problem during the period (first
row). When he finally solved the problem correctly, he clenched his fist which may indicate his
excitement and passion (second row). When asked for hints, the student looked confused.</p>
      <p>To visually illustrate the prediction of problem outcomes and to understand student behavior,
we present visual examples of an eighth grade student using MathSpring, see Figure 2. The
student used MathSpring for one session of around 20 minutes and consented to have his face
and screen recorded. Figure ?? shows the evolution of student expressions and gestures, and
their corresponding problem outcomes. When the student successfully solves the problem on
the first attempt (SOF), we observe that he focused tightly on the problem during the period
(first row). When he finally solved the problem correctly, he clenched his fist which may indicate
his excitement and passion (second row). When asked for hints, the student looked confused
scratching his head but still engaged and actively attempted to solve the problem (rows 3–4). For
the last problem (GIVEUP), the student gradually became distracted and presented frustration
and boredom (rows 5–6). These observations are consistent with our assumption that facial
expressions and gestures provide important cues for inferring students’ learning outcomes.</p>
      <p>Because of the interpretability of Facial Action Units (AUs), we visualized how often and with
what intensity various action units occurred on average for the diferent efort classes of the
entire dataset. For each data instance, we aggregated AU presence values weighted by their
respective intensity values and normalized them by the total number of frames in which the
face was detected. The input to our baseline models consists of variable-length webcam video
clips of participants working on MathSpring problems. For each frame of all the videos in the
dataset, 18 AU presence and 17 AU intensity values, along with head-pose and eye-gaze vectors,
are extracted using OpenFace [34]. In order to compute an aggregate feature representation, we
used statistics (mean, standard deviation, min and max) for each feature as well as statistics for
their derivatives to produce a uniform length 376-dimensional feature representation. These
features are then used as input to machine learning models trained in Phase III to predict
problem outcome labels. Other features that have been computed on our datasets are based on
transfer learning and are described in phase III.</p>
      <p>
        Long Term Goals to Identify Labels for Face and Gestures. One goal of Phase I is to
improve the performance obtained by facial action units and baseline models. For example, a
multi-modal model that utilizes signals from all streams of information in the dataset including
the mouse movements and clicks, as well as the video stream of the screen activity will likely
result in better predictive performance. Moreover, training models that explicitly utilize the
temporal dynamics of how facial behavior evolves over the duration of the student’s interaction
with the tutor could potentially yield further improvements in model performance. Finally, the
biggest challenge in recognizing afect is to utilize afect-sensitive models to provide appropriate
and efective interventions that quantifiably improve the learning experience. Some researchers
ventured in this direction [
        <xref ref-type="bibr" rid="ref22">22</xref>
        ]. In future work, we plan to provide personalized interventions
in MathSpring based on the proposed afect analysis models, and to conduct experiments to
validate the efectiveness of the interventions. Lessons learnt from this initial analysis will also
inform future data collection strategies. We intend to use richer data sets to investigate whether
the system can predict changes in student learning behaviors and strategies.
      </p>
    </sec>
    <sec id="sec-4">
      <title>4. Phase II: Identifying Students’ Attention in Math Learning</title>
      <p>
        When students are bored or distracted online, they might disengage and wander, leading to a
decline in the learning process. Currently, few online systems account for the context-sensitive
nature of learning, i.e., motivation, social and emotional learning, and climate as well as complex
interactions among these factors. In Phase II, we used computer vision to identify student
engagement and emotion, which are correlated with learning gains [
        <xref ref-type="bibr" rid="ref1 ref2">1, 2</xref>
        ]; emotion drives
attention and attention drives learning [35]. Computer vision-enhanced research can assist in
supporting students’ emotion and maintaining their engagement by recognizing students’ head
orientation and gaze expression.
      </p>
      <p>We summarize our contributions as follows:
• Demonstrated the use of computer vision with live video data to infer afect as one
indicator of students’ motivation;
• Developed a deep learning-based computer vision model to identify head pose as an
indicator of engagement vs. distraction; and
• Communicated this information to teachers by showing facial expressions.</p>
      <p>
        Research Approach and Results. Within the collected video dataset described in Phase I,
annotations indicated whether students’ attention at specific frames was engaged or wandering
[
        <xref ref-type="bibr" rid="ref22">22</xref>
        ], see Figure 2. In addition, we trained baselines for a computer vision module that determined
the extent of student engagement during remote learning. Baselines include state-of-the-art
deep-learning image classifiers and traditional conditional and logistic regression for head pose
estimation. We then incorporated a gaze baseline into the MathSpring learning platform and
evaluated its performance with the currently implemented approach.
      </p>
      <p>
        Development and early evaluation of this technology monitored student engagement in
real-time, detected waning attention and distraction, and assessed which interventions led to
more productive learning. We used pre-process and crowdsourced label frames of the videos to
propose a publicly available dataset that aids researchers in automated student engagement
prediction [
        <xref ref-type="bibr" rid="ref22">22</xref>
        ].
      </p>
      <p>The model was trained on benchmark datasets that were curated to help tutors solve such tasks.
We incorporated one of our baselines in the MathSpring tutor. Figure 3 exhibits an example
problem presented to students on MathSpring. The tutor targets sensing and interpreting
facial signals relevant to student emotions and provides students with real-time classroom
interventions that can aid their progress, suggesting when and who needs further assistance,
and identifying which interventions are working. The implemented computer vision module
alerts wandering students to regain their attention.</p>
      <p>Given the collected, annotated, and balanced dataset of students solving mathematical
problems, we considered state-of-the-art deep learning architectures that classified a student’s
gesture into “looking at their screen”, “looking at their paper”, or “wandering”. We compared
these to baselines that rely on head pose estimation.</p>
      <p>We fine-tuned diferent convolutional architectures that are pre-trained on ImageNet [ 36] to
classify video frames into the three classes. We also estimate head poses (i.e., yaw, pitch and roll)
of students using a deep neural network FSA-Net [37]. The predicted head poses were used to
classify video frames into the three aforementioned classes. We then compare the performance
of the deep convolutional networks to the performance of the head pose estimator approach.</p>
      <p>Pilot Study. We conducted a Pilot Study in which the head pose estimator was integrated
into MathSpring. A student’s head pose was computed in real-time and used with real students
during Summer 2021. The tutor detects whether a student is looking of-screen by analyzing
the pose angle values and considers a student facing straight at the screen as being in a neutral
state (i.e., the pose angle is 0°) and infers of-screen poses when the angle values exceed certain
thresholds. Real-time interventions, e.g., showing a focus circle, an animated character, or a
message, were delivered. Such interventions target re-engaging a wandering student. The
realtime detection and automatic responses help students sustain and efectively allocate attentional
resources on learning tasks, which is critical for efective learning [ 38].</p>
      <p>Results. All convolutional neural network architectures performed significantly better than
the head pose estimation strategies. We presented the per-class accuracy for the best deep
learning (94%) and head pose (60%) estimation models.</p>
      <p>Long-Term Goals to Identify Students’ Attention. One long-term goal of Phase II is to
evaluate students’ visual feedback in real classrooms through the head pose detector’s
performance, which provides a coarse estimation of where students are looking. The intelligent tutor
will acquire more information when students’ gaze direction can be detected and engagement
is inferred. In this case, the head pose detector’s intervention (e.g., animated character, verbal
message) is used as a learning companion for maintaining students’ level of engagement.
During learning or problem-solving, it is quite common for students to keep relatively fixed head
positions but the gaze direction moves frequently, which makes it insuficient to detect emotion
from head poses only. Therefore, in Phase II we focused on methodology for inferring gaze
direction while students interact in real-time.</p>
      <p>Future research will provide evidence about whether head pose interventions are successful
in reorienting student attention towards learning and which deep learning models demonstrated
superior classification performance. We also seek to determine which interventions are most
efective in promoting learning gains compared to the non-pose-reactive tutor. Also of interest
is whether individual student diferences (e.g.,in prior knowledge, aptitude, afective
predispositions) moderate the efects of computer vision-enhanced interventions (for the teacher or
student).</p>
    </sec>
    <sec id="sec-5">
      <title>5. Phase III: Predicting Problem-Solving Outcome</title>
      <p>
        In Phase III, we propose deep learning models that predict problem-solving outcomes for the
video clips collected and annotated during Phase I. Predicting these outcomes allows tutoring
systems to adapt interventions to enhance student learning. We first trained a classifier using
traditional facial analysis features such as head pose, gaze and facial action units (AUs) to predict
the exercise outcome. The multi-class model achieved a mean accuracy of 0.54 and a mean
F-score of 0.27 for predicting one of seven possible outcome classes [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]. To improve prediction
performance, we further developed a video-based transfer learning approach to predict problem
outcomes of students by analyzing their facial expressions and gestures [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ]. Our transfer
learning challenge involved designing a representation for facial expression analysis using
images from the Internet and transferring this knowledge to predict student behavior in webcam
videos of students in a classroom setting. We introduced a novel facial afect representation and
a user-personalized training scheme to harness the potential of this representation. Additionally,
we developed various recurrent neural network variants that model the temporal structure of
video sequences. Our final model, named ATL-BP for “Afect Transfer Learning for Behavior
Prediction,” outperformed the previous work on the dataset, achieving a 50% relative increase
in the mean F-score as well as an absolute 11 percentage point increase in accuracy.
      </p>
      <p>Ideally, an afect-sensitive model should be able to accurately predict the efort label of the user
as early as possible, in order to enable quick and efective interventions by the teacher or tutor.
Therefore, our team is currently working on predicting the outcome of student performance
using early visual and tabular cues demonstrating the eficacy of our approach and the potential
impact of early outcome prediction for the development of better intelligent tutors. We will
evaluate our classification models when only a fraction of the data is observed during test
time. We are also incorporating tabular cues, e.g., timestamps of students performing specific
actions. Again we are using a video-based transfer learning approach for predicting problem
outcomes by analyzing students’ faces and gestures, and combining them with tabular data.
The transfer-learning challenge is to design a representation in the source domain of images
obtained from the internet for facial expression analysis and transfer this learned representation
for human behavior prediction in the domain of webcam videos of students in a classroom
environment.</p>
    </sec>
    <sec id="sec-6">
      <title>6. Discussion and Conclusion</title>
      <p>This paper presented a survey of research activities and challenges for the future of computer
vision-augmented tutoring in math learning. The suite of computer vision tools that we
developed, called FaceReaders, uses facial expression recognition to identify and predict student
engagement, motivation, afect, and future behavior early in students’ interaction with online
learning, specifically while students spend a brief time working on an exercise. We trained
classifiers to directly predict the success or failure of a student’s attempt to answer questions,
based on features extracted from video streams. We extracted timing information from student
log data, which includes the exact time students take for actions, e.g., asking for a hint or
attempting to answer the exercise. Such information provides complementary insights into
students’ learning process and can be used to better understand their behavior and afective
states.</p>
      <p>To the best of our knowledge, no prior research combines visual afective analysis with
student log data in the context of predicting student learning outcomes. One goal is to create
and evaluate facial expression recognition tools with intelligent tutors.</p>
      <p>Real-time teachers need answers for many questions, e.g., Who needs my help most right
now? Is the class ready to move to the next topic? Answers to these questions will help teachers
strategize responses, adapt class pedagogy and provide interventions. We expect this research
to have a significant impact on development of better intelligent tutors. It should improve the
diagnostic and predictive power of online learning by accurately predicting student exercise
outcomes in the early stages.
[23] K. He, X. Zhang, S. Ren, J. Sun, Deep residual learning for image recognition, in:
Proceedings of the IEEE conference on computer vision and pattern recognition, 2016, pp.
770–778.
[24] H. Yu, A. Gupta, W. Lee, I. Arroyo, M. Betke, D. Allesio, T. Murray, J. Magee, B. P. Woolf,
Measuring and integrating facial expressions and head pose as indicators of engagement
and afect in tutoring systems, in: International Conference on Human-Computer
Interaction, Springer, 2021, pp. 219–233.
[25] I. Arroyo, B. P. Woolf, W. Burelson, K. Muldner, D. Rai, M. Tai, A multimedia
adaptive tutoring system for mathematics that addresses cognition, metacognition and afect,
International Journal of Artificial Intelligence in Education 24 (2014) 387–426.
[26] C. S. Dweck, Messages that motivate: How praise molds students’ beliefs, motivation, and
performance (in surprising ways), in: Improving academic achievement, Elsevier, 2002, pp.
37–60.
[27] I. Arroyo, B. P. Woolf, D. G. Cooper, W. Burleson, K. Muldner, The impact of animated
pedagogical agents on girls’ and boys’ emotions, attitudes, behaviors and learning, in:
2011 IEEE 11th International Conference on Advanced Learning Technologies, IEEE, 2011,
pp. 506–510.
[28] R. Azevedo, S. A. Martin, M. Taub, N. V. Mudrick, G. C. Millar, J. F. Grafsgaard, Are
pedagogical agents’ external regulation efective in fostering learning with intelligent
tutoring systems?, in: Intelligent Tutoring Systems: 13th International Conference, ITS
2016, Zagreb, Croatia, June 7-10, 2016. Proceedings 13, Springer, 2016, pp. 197–207.
[29] A. L. Baylor, S. Kim, Designing nonverbal communication for pedagogical agents: When
less is more, Computers in Human Behavior 25 (2009) 450–457.
[30] M. C. Dufy, R. Azevedo, Motivation matters: Interactions between achievement goals
and agent scafolding for self-regulated learning within an intelligent tutoring system,
Computers in Human Behavior 52 (2015) 338–348.
[31] S. Lallé, N. V. Mudrick, M. Taub, J. F. Grafsgaard, C. Conati, R. Azevedo, Impact of individual
diferences on afective reactions to pedagogical agents scafolding, in: Intelligent Virtual
Agents: 16th International Conference, IVA 2016, Los Angeles, CA, USA, September 20–23,
2016, Proceedings 16, Springer, 2016, pp. 269–282.
[32] S. D’Mello, R. W. Picard, A. Graesser, Toward an afect-sensitive autotutor, IEEE Intelligent</p>
      <p>Systems 22 (2007) 53–61.
[33] R. Pekrun, T. Goetz, L. M. Daniels, R. H. Stupnisky, R. P. Perry, Boredom in achievement
settings: Exploring control–value antecedents and performance outcomes of a neglected
emotion., Journal of educational psychology 102 (2010) 531.
[34] T. Baltrušaitis, P. Robinson, L.-P. Morency, Openface: an open source facial behavior
analysis toolkit, in: 2016 IEEE winter conference on applications of computer vision
(WACV), IEEE, 2016, pp. 1–10.
[35] R. J. Jagers, D. Rivas-Drake, B. Williams, Transformative social and emotional learning
(sel): Toward sel in service of educational equity and excellence, Educational Psychologist
54 (2019) 162–184.
[36] J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, L. Fei-Fei, Imagenet: A large-scale hierarchical
image database, in: 2009 IEEE conference on computer vision and pattern recognition,
Ieee, 2009, pp. 248–255.
[37] T.-Y. Yang, Y.-T. Chen, Y.-Y. Lin, Y.-Y. Chuang, Fsa-net: Learning fine-grained structure
aggregation for head pose estimation from a single image, in: Proceedings of the IEEE/CVF
conference on computer vision and pattern recognition, 2019, pp. 1087–1096.
[38] S. K. D’Mello, Gaze-based attention-aware cyberlearning technologies, Mind, Brain and
Technology: Learning in the Age of Emerging Technologies (2019) 87–105.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>R. S.</given-names>
            <surname>Baker</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>K. D'Mello</surname>
          </string-name>
          ,
          <string-name>
            <surname>M. M. T. Rodrigo</surname>
            ,
            <given-names>A. C.</given-names>
          </string-name>
          <string-name>
            <surname>Graesser</surname>
          </string-name>
          ,
          <article-title>Better to be frustrated than bored: The incidence, persistence, and impact of learners' cognitive-afective states during interactions with three diferent computer-based learning environments</article-title>
          ,
          <source>International Journal of Human-Computer Studies</source>
          <volume>68</volume>
          (
          <year>2010</year>
          )
          <fpage>223</fpage>
          -
          <lpage>241</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>S. D</given-names>
            <surname>'Mello</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Lehman</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Pekrun</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Graesser</surname>
          </string-name>
          ,
          <article-title>Confusion can be beneficial for learning</article-title>
          ,
          <source>Learning and Instruction</source>
          <volume>29</volume>
          (
          <year>2014</year>
          )
          <fpage>153</fpage>
          -
          <lpage>170</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>A.</given-names>
            <surname>Joshi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Allessio</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Magee</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Whitehill</surname>
          </string-name>
          ,
          <string-name>
            <given-names>I.</given-names>
            <surname>Arroyo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Woolf</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Sclarof</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Betke</surname>
          </string-name>
          ,
          <article-title>Afect-driven learning outcomes prediction in intelligent tutoring systems</article-title>
          ,
          <source>in: 2019 14th IEEE international conference on automatic face &amp; gesture recognition (FG</source>
          <year>2019</year>
          ), IEEE,
          <year>2019</year>
          , pp.
          <fpage>1</fpage>
          -
          <lpage>5</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>S. D</given-names>
            <surname>'Mello</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Olney</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Williams</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Hays</surname>
          </string-name>
          ,
          <article-title>Gaze tutor: A gaze-reactive intelligent tutoring system</article-title>
          ,
          <source>International Journal of human-computer studies 70</source>
          (
          <year>2012</year>
          )
          <fpage>377</fpage>
          -
          <lpage>398</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>K.</given-names>
            <surname>VanLehn,</surname>
          </string-name>
          <article-title>The relative efectiveness of human tutoring, intelligent tutoring systems, and other tutoring systems</article-title>
          ,
          <source>Educational psychologist 46</source>
          (
          <year>2011</year>
          )
          <fpage>197</fpage>
          -
          <lpage>221</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>A. T.</given-names>
            <surname>Corbett</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. R.</given-names>
            <surname>Anderson</surname>
          </string-name>
          ,
          <article-title>Student modeling and mastery learning in a computer-based programming tutor</article-title>
          , in: Intelligent Tutoring Systems: Second International Conference, ITS'92 Montréal, Canada, June 10-12
          <source>1992 Proceedings 2</source>
          , Springer,
          <year>1992</year>
          , pp.
          <fpage>413</fpage>
          -
          <lpage>420</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>A. C.</given-names>
            <surname>Graesser</surname>
          </string-name>
          , K. VanLehn,
          <string-name>
            <given-names>C. P.</given-names>
            <surname>Rosé</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P. W.</given-names>
            <surname>Jordan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Harter</surname>
          </string-name>
          ,
          <article-title>Intelligent tutoring systems with conversational dialogue</article-title>
          ,
          <source>AI</source>
          magazine
          <volume>22</volume>
          (
          <year>2001</year>
          )
          <fpage>39</fpage>
          -
          <lpage>39</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>J. A.</given-names>
            <surname>Kulik</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Fletcher</surname>
          </string-name>
          ,
          <article-title>Efectiveness of intelligent tutoring systems: a meta-analytic review</article-title>
          ,
          <source>Review of educational research 86</source>
          (
          <year>2016</year>
          )
          <fpage>42</fpage>
          -
          <lpage>78</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>M. S.</given-names>
            <surname>Hussain</surname>
          </string-name>
          ,
          <string-name>
            <surname>O. AlZoubi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R. A.</given-names>
            <surname>Calvo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>K. D'Mello</surname>
          </string-name>
          ,
          <article-title>Afect detection from multichannel physiology during learning sessions with autotutor</article-title>
          ,
          <source>in: Artificial Intelligence in Education: 15th International Conference, AIED</source>
          <year>2011</year>
          , Auckland, New Zealand, June 28-July
          <year>2011</year>
          15, Springer,
          <year>2011</year>
          , pp.
          <fpage>131</fpage>
          -
          <lpage>138</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>A. C.</given-names>
            <surname>Strain</surname>
          </string-name>
          , S. K. D 'Mello,
          <article-title>Emotion regulation during learning</article-title>
          ,
          <source>in: Artificial Intelligence in Education: 15th International Conference, AIED</source>
          <year>2011</year>
          , Auckland, New Zealand, June 28-July
          <year>2011</year>
          15, Springer,
          <year>2011</year>
          , pp.
          <fpage>566</fpage>
          -
          <lpage>568</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>B. D.</given-names>
            <surname>Nye</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Karumbaiah</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S. T.</given-names>
            <surname>Tokel</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. G.</given-names>
            <surname>Core</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Stratou</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Auerbach</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Georgila</surname>
          </string-name>
          ,
          <article-title>Engaging with the scenario: Afect and facial patterns from a scenario-based intelligent tutoring system</article-title>
          ,
          <source>in: Artificial Intelligence in Education: 19th International Conference, AIED</source>
          <year>2018</year>
          , London, UK, June 27-30,
          <year>2018</year>
          , Proceedings,
          <source>Part I 19</source>
          , Springer,
          <year>2018</year>
          , pp.
          <fpage>352</fpage>
          -
          <lpage>366</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>N.</given-names>
            <surname>Ruiz</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Yu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D. A.</given-names>
            <surname>Allessio</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Jalal</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Joshi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Murray</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. J.</given-names>
            <surname>Magee</surname>
          </string-name>
          ,
          <string-name>
            <surname>K. M. Delgado</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          <string-name>
            <surname>Ablavsky</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          <string-name>
            <surname>Sclarof</surname>
          </string-name>
          , et al.,
          <article-title>Atl-bp: a student engagement dataset and model for afect transfer learning for behavior prediction</article-title>
          ,
          <source>IEEE Transactions on Biometrics, Behavior, and Identity Science</source>
          (
          <year>2022</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <given-names>J.</given-names>
            <surname>Chen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Liu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Tu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Aragones</surname>
          </string-name>
          ,
          <article-title>Person-specific expression recognition with transfer learning</article-title>
          ,
          <source>in: 2012 19th IEEE International Conference on Image Processing</source>
          , IEEE,
          <year>2012</year>
          , pp.
          <fpage>2621</fpage>
          -
          <lpage>2624</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <given-names>J.</given-names>
            <surname>Chen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Liu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Tu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Aragones</surname>
          </string-name>
          ,
          <article-title>Learning person-specific models for facial expression and action unit recognition</article-title>
          ,
          <source>Pattern Recognition Letters</source>
          <volume>34</volume>
          (
          <year>2013</year>
          )
          <fpage>1964</fpage>
          -
          <lpage>1970</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <given-names>H.</given-names>
            <surname>Kaya</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Gürpınar</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A. A.</given-names>
            <surname>Salah</surname>
          </string-name>
          ,
          <article-title>Video-based emotion recognition in the wild using deep transfer learning and score fusion</article-title>
          ,
          <source>Image and Vision Computing</source>
          <volume>65</volume>
          (
          <year>2017</year>
          )
          <fpage>66</fpage>
          -
          <lpage>75</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [16]
          <string-name>
            <given-names>M.</given-names>
            <surname>Xu</surname>
          </string-name>
          , W. Cheng,
          <string-name>
            <given-names>Q.</given-names>
            <surname>Zhao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Ma</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Xu</surname>
          </string-name>
          ,
          <article-title>Facial expression recognition based on transfer learning from deep convolutional networks</article-title>
          ,
          <source>in: 2015 11th International Conference on Natural Computation (ICNC)</source>
          , IEEE,
          <year>2015</year>
          , pp.
          <fpage>702</fpage>
          -
          <lpage>708</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [17]
          <string-name>
            <given-names>S. J.</given-names>
            <surname>Pan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Q.</given-names>
            <surname>Yang</surname>
          </string-name>
          ,
          <article-title>A survey on transfer learning</article-title>
          .
          <source>ieee transactions on knowledge and data engineering</source>
          ,
          <volume>22</volume>
          (
          <issue>10</issue>
          )
          <fpage>1345</fpage>
          (????).
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          [18]
          <string-name>
            <given-names>I.</given-names>
            <surname>Dua</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A. U.</given-names>
            <surname>Nambi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Jawahar</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Padmanabhan</surname>
          </string-name>
          , Autorate:
          <article-title>How attentive is the driver?</article-title>
          ,
          <source>in: 2019 14th IEEE International Conference on Automatic Face &amp; Gesture Recognition (FG</source>
          <year>2019</year>
          ), IEEE,
          <year>2019</year>
          , pp.
          <fpage>1</fpage>
          -
          <lpage>8</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          [19]
          <string-name>
            <given-names>B. P.</given-names>
            <surname>Woolf</surname>
          </string-name>
          , I. Arroyo,
          <string-name>
            <given-names>K.</given-names>
            <surname>Muldner</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W.</given-names>
            <surname>Burleson</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D. G.</given-names>
            <surname>Cooper</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Dolan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R. M.</given-names>
            <surname>Christopherson</surname>
          </string-name>
          ,
          <article-title>The efect of motivational learning companions on low achieving students and students with disabilities</article-title>
          ,
          <source>in: Intelligent Tutoring Systems: 10th International Conference, ITS</source>
          <year>2010</year>
          ,
          <article-title>Pittsburgh</article-title>
          , PA, USA, June 14-18,
          <year>2010</year>
          , Proceedings,
          <source>Part I 10</source>
          , Springer,
          <year>2010</year>
          , pp.
          <fpage>327</fpage>
          -
          <lpage>337</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          [20]
          <string-name>
            <given-names>S.</given-names>
            <surname>Karumbaiah</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Lizarralde</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Allessio</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Woolf</surname>
          </string-name>
          , I. Arroyo,
          <string-name>
            <given-names>N.</given-names>
            <surname>Wixon</surname>
          </string-name>
          ,
          <article-title>Addressing student behavior and afect with empathy and growth mindset</article-title>
          .,
          <source>International Educational Data Mining Society</source>
          (
          <year>2017</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          [21]
          <string-name>
            <given-names>Y.</given-names>
            <surname>Kim</surname>
          </string-name>
          ,
          <article-title>Empathetic virtual peers enhanced learner interest and self-eficacy</article-title>
          , in: Workshop on Motivation and
          <article-title>Afect in Educational Software, in conjunction with the 12th</article-title>
          <source>International Conference on Artificial Intelligence in Education</source>
          ,
          <year>2005</year>
          , pp.
          <fpage>9</fpage>
          -
          <lpage>16</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          [22]
          <string-name>
            <given-names>K.</given-names>
            <surname>Delgado</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. M.</given-names>
            <surname>Origgi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Hasanpoor</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Yu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Allessio</surname>
          </string-name>
          , I. Arroyo,
          <string-name>
            <given-names>W.</given-names>
            <surname>Lee</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Betke</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Woolf</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S. A.</given-names>
            <surname>Bargal</surname>
          </string-name>
          ,
          <article-title>Student engagement dataset</article-title>
          ,
          <source>in: Proceedings of the IEEE/CVF International Conference on Computer Vision</source>
          ,
          <year>2021</year>
          , pp.
          <fpage>3628</fpage>
          -
          <lpage>3636</lpage>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>