<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <journal-title-group>
        <journal-title>H. Soyel and H. Demirel, “Localized
discriminative scale invariant feature transform
based facial expression recognition,” Computers
and Electrical Engineering</journal-title>
      </journal-title-group>
    </journal-meta>
    <article-meta>
      <article-id pub-id-type="doi">10.1016/j.compeleceng.2011.10.016</article-id>
      <title-group>
        <article-title>Instructor Presence and Learner Response: Analyzing Learning Gain, Cognitive Load, Visual Attention and Affective States in Video and Metaverse Environments</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Yuli Sutoto Nugroho</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Queen Mary University of London</institution>
          ,
          <addr-line>Mile End Road, London</addr-line>
          ,
          <country country="UK">United Kingdom</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2012</year>
      </pub-date>
      <volume>38</volume>
      <issue>5</issue>
      <fpage>282</fpage>
      <lpage>289</lpage>
      <abstract>
        <p>This research explores the impact of different learning delivery modes, across video and virtual space via the metaverse, on student responses such as learning gain, visual attention, cognitive load, and affective states. Utilizing eye-tracking technology and facial expression analysis, this study examines the effect of instructor presence, both physical and in avatars, in educational videos and virtual environments, including individual and group lectures. The experimental design includes video-based learning and metaverse settings, where the instructor and other learners' presence is manipulated to measure their effects on the learners' cognitive and emotional responses. The mode of instructor presence appears to significantly impact learning outcomes, cognitive load, affective state, and attention, according to preliminary data. These findings underscore the potential of sophisticated educational technology to improve learning outcomes. This study closes essential gaps in the literature and provides valuable information for improving teaching strategies in technologically enhanced learning environments.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;Learning gain</kwd>
        <kwd>Cognitive load</kwd>
        <kwd>Eye-tracking technology</kwd>
        <kwd>Facial Expression</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>1. Introduction1</p>
      <p>Technological developments, increasing student
demands, and new research on successful teaching
approaches all contribute to the evolution of teaching
and learning. As these changes continue, educators must
stay flexible and adaptable, seeking new ideas and
strategies to improve student learning and success.
Suitable learning resources, delivery styles, and
environments assist students to attain their learning
goals [1]. Learning resources can be offered in a variety
of ways. One option is to use video as the primary
medium. Furthermore, the metaverse has become a
model of a learning environment in which both the
lecturer and the students can be virtually present in the
same space. As a result, video is an effective medium, and
the metaverse is a modern learning environment where
instructors and students can virtually interact.</p>
      <p>Video-based learning (VBL) has become more
popular recently, with many students taking online
courses [2][3]. Moreover, in this cutting-edge period,
studying the metaverse is garnering more and more
attention. Metaverse education is becoming an essential
and exciting educational trend. It is considered a platform
for sustainable education free of time and space barriers
[4]. In the metaverse, the lecturer and the learners can be
virtually present in the shared learning environment.
However, the effect of the instructor's physical and
learner presence in the videos and virtual environments
like the metaverse has not yet been extensively
researched, especially the impact of instructor and
learners' presence on learners' responses, such as
learning gain, visual attention and cognitive load (eye
tracking measures), and affective state (valence and
arousal).</p>
      <p>The project examines how the instructor's and other
learners' presence in different learning delivery methods,
such as videos and metaverse, affects learners' learning
gain, visual attention, cognitive load (eye tracking
measures), and affective state (valence and arousal).
Detecting learners' responses will involve selected
instruments and technologies.</p>
    </sec>
    <sec id="sec-2">
      <title>2. Problem Identification</title>
      <p>Many previous studies have discussed the effects of
learning delivery and environment on learners. Other
studies have focused on eye tracking and other body
measurement tools. However, how “presence” in video
materials and the metaverse impacts learners, such as
understanding, cognitive load, visual attention, and
affective state, has yet to be studied. Previous research on
the effect of presence focused on learning gain; in this
research, we also intend to measure cognitive load (using
various measure instruments) and affective state
dynamics.</p>
      <p>Video-based learning has become more familiar,
although its effectiveness in learning and usability needs
to be better understood [5]. Moreover, the merits and
limitations of each video lecture type for learning have
yet to be thoroughly investigated [6]. Regarding the
extent of learning, [7] found that students taught through
videos (e.g., YouTube) obtained considerably more
significant learning gains on the post-test than those who
never thought about using videos. Further studies are
needed on teacher presence (real human and avatar)
versus slides and voice only in learning videos.</p>
      <p>In addition, the metaverse is a breakthrough in
education. The effect of the metaverse shared learning
environment (with and without other students) on
learning needs deeper investigation.</p>
      <p>It is challenging to find prior work comparing the
effect of instructor and learners' presence in lecture
videos and the metaverse on students' affective states,
eye-tracking measures, and learning gain. Hence,
measuring the effect of instructors' and learners'
presence in different learning delivery modes and
learning environments is novel.</p>
      <p>Finally, using body measurement tools for learning
purposes, such as eye-tracking technology and facial
expression recognition, warrants further investigation.
From the problem related to technology-enhanced
learning above, we formulate some research questions
into two different experiments.</p>
    </sec>
    <sec id="sec-3">
      <title>2.1. Research Question</title>
    </sec>
    <sec id="sec-4">
      <title>Experiment 1 (Video Experiment) of</title>
      <p>The research questions for video experiments are as
follows:
1. How does an instructor's physical presence in
a learning video affect learners' learning gains and
cognitive load (measured by eye-tracking) compared
to the instructor's complete absence from the video?
2. How does an instructor's physical presence in
a learning video affect learners' learning gains and
cognitive load (measured by eye-tracking) compared
to the instructor's presence as a virtual avatar?</p>
    </sec>
    <sec id="sec-5">
      <title>2.2. Research</title>
    </sec>
    <sec id="sec-6">
      <title>Experiment</title>
    </sec>
    <sec id="sec-7">
      <title>Experiment) 2</title>
    </sec>
    <sec id="sec-8">
      <title>Question of (Metaverse</title>
      <p>The research questions for metaverse experiments
are as follows:
3. How does the shared virtual presence of
instructors and learners (i.e., co-presence) in the
metaverse, where the student only sees the lecturer
(individual lecture), evoke more pleasant emotions
(higher valence) and more intense emotions (higher
arousal) compared to a scenario where the student
can also see other students (group lecture)?
4. How does a group lecture, where students can
see both the lecturer and other students, affect
learners' learning gains compared to an individual
lecture, where the student only sees the lecturer?</p>
    </sec>
    <sec id="sec-9">
      <title>3. Current</title>
    </sec>
    <sec id="sec-10">
      <title>Hypothesis</title>
    </sec>
    <sec id="sec-11">
      <title>Knowledge and</title>
      <p>A word's representation becomes more complex
when a gesture is added [8]. In contrast, [9]found that
gestures contribute to listening comprehension and
communication effectiveness. Gestures help the speaker
to organize their thoughts [10]. Videos featuring the
presence of a teacher can significantly improve academic
achievements and raise cognitive loads compared to
those that do not include a teacher [11]. However, in our
experiment, we locate the central content/slide and
lecturer position separately. This is possibly a different
result from the previous study since the student can split
their focus and miss critical information on the slide
when focusing on the lecturer's gesture. Moreover, some
answers to the question can only be seen on the slide
without being mentioned by the lecturer.</p>
      <p>On the other hand, as a compensating technique,
speakers and listeners can use gestures during
communication and thinking, interacting with people's
cognitive dispositions [12]. The gesture might have
drawbacks to learning, making the instructor's
explanation seem more complex. Besides, learners need
to switch their attention from the topic to the lecturer
back and forth.From those prior works, we can formulate
some hypotheses (H).</p>
      <p>H1: The lecturer's presence in the video leads to lower
learning gain than in the non-lecturer presence video.</p>
      <p>H2: The lecturer's presence in the video leads to a higher
cognitive load than in the non-lecturer presence video.</p>
      <p>A digital representation of a person in a virtual
setting is called an avatar. An avatar's appearance and
behaviour can be m ore superficial than a real human's.
Avatar-based teachers are more attractive than
face-toface [13], which can be designed to have fewer gestures.
Besides, Avatar-based instructors can enhance the sense
of presence and instructing satisfaction. [14] discovered
that participants using avatars learned substantially
more while memorising pairs of letters.</p>
      <p>Using avatar lecturers or teachers in virtual settings
can improve learning by providing students an
interesting, immersive, and enriching educational
experience [15], [16]. Avatars may be helpful when
students are more comfortable communicating with a
virtual avatar or when the subject contains delicate or
possibly painful issues.</p>
      <p>H3: The instructor's presence as an avatar in the
learning video leads to more significant learning gain than
their physical presence.</p>
      <p>H4: The instructor's virtual presence as an avatar in the
learning video leads to a lower cognitive load than the
physical presence.</p>
      <p>During learning sessions, students naturally
experience a variety of affective states [17] [18]. [19]
discovered that social presence influenced the impact of
agent behaviour on valence and arousal during virtual
presentations. The physical and social environment can
affect students' affective experiences while learning.
Higher valence and arousal are usually fostered in a
positive, supportive, and inclusive learning environment.</p>
      <p>[20] contends that distractions can be minimised
through independent learning. On the other hand, [21]
asserts that group learning dramatically increases
performance. Dividing up complex tasks is the whole
idea of group learning [22].</p>
      <p>Nevertheless, in our research, where there is no
direct interaction among students, individual lectures are
anticipated to provide a more concentrated virtual
setting with fewer distractions. However, group lectures
may create a more pronounced sense of presence in
virtual learning environments.</p>
      <p>H5: There is a statistically significant difference
between the learner's level of valence and arousal during
individual lectures (a student alone) and group lectures (a
student with other students)</p>
      <p>H6: The individual lecture leads to more significant
learning gain than the group lecture.</p>
    </sec>
    <sec id="sec-12">
      <title>4. Research Methodology</title>
      <p>The research was conducted at Queen Mary
University of London and involved 3rd-year Electronic
Engineering and Computer Science students. The
involvement is voluntary. There are two experiments.
Experiment 1 (video) involved 33 participants, and
Experiment 2 (metaverse) involved 20 participants. This
ensures enough data is collected and a robust statistical
analysis can be conducted.
4.1. Video</p>
    </sec>
    <sec id="sec-13">
      <title>Methodology</title>
    </sec>
    <sec id="sec-14">
      <title>Experiment</title>
      <p>Eye-tracking Technology captured and measured
participants' gaze and visual attention when learning.
The data collected is analysed statistically to see the
effects of various experimental conditions on fixations
and pupil dilation, which are indicators of cognitive load
[23], [24], [25]. The eye-tracker that we used is a single
eye camera 120Hz Pupil Core eye-tracker from Pupil
Labs https://pupil-labs.com/products/core (See figure 1).</p>
      <p>Before watching the videos, participants were asked
to answer a knowledge test (pre -test Multiple Choice
Questions) on the topic of the videos. After watching the
videos, they were asked again to answer the same
knowledge test (post-test) to measure learning gain. We
modified a video from Lex Fridman about “Deep
Learning”:
https://www.youtube.com/watch?v=O5xeyoRL95U. We
cut the video into three parts, each about 5 minutes long
(Topic A, B, and C). The videos contain some slide
presentations (Topic A has 4, topic B has 6 , and Topic C
has five slides.</p>
      <p>We randomized the order of the experimental
conditions (no lecturer, lecturer presence, and avatar
lecturer) to avoid any habituation effect. For example, the
first participant starts watching videos using
nonlecturer presence, while the second and third participants
start using lecturer physical presence and avatar lecturer
presence, respectively. Figure 2 shows an example of the
video with the lecturer's presence as an avatar.</p>
      <p>The experiment was about 60 minutes in total. The
visual timeline can be seen in Figure 3.</p>
    </sec>
    <sec id="sec-15">
      <title>4.2. Metaverse</title>
    </sec>
    <sec id="sec-16">
      <title>Methodology</title>
    </sec>
    <sec id="sec-17">
      <title>Experiment</title>
      <p>We recorded the participant's faces when watching
videos using an external camera. Participants' affective
states were measured continuously in the 2D space of
emotional valence (how pleased or unpleased) and
arousal (the intensity of power activation of the
emotion). We use an established system to detect the
level of valence and arousal developed by Daeha Kim and
Byung Cheol Song [26]. We recorded the faces of the
participants and analysed them afterwards. The arousal
scale ranged from 0 (not active) to 1, and the Valence
scale varied from 1 (unpleasant facial expression) to 1
(pleasant) [27]. The Axis of valence and arousal can be
seen in Figure 4.</p>
      <p>We also employed pre- and post-test questions and
questionnaires to ask participants their opinions on their
metaverse learning experience. We modified a video by
Alexander Amini about “AI Bias and Fairness”:
https://www.youtube.com/watch?v=wmyVODy_WD8.
In this experiment, we cut the video into two parts, each
about 4 minutes long (Topic D and E).</p>
      <p>The metaverse or virtual environment used was
ENGAGEVR (https://engagevr.io); see Figure 5. This
platform has many features that support teaching
purposes, and the subscription price is affordable for
everyone in the educational field.</p>
      <p>The metaverse experiment took about 50 minutes.
The visual timeline can be seen in Figure 6.</p>
    </sec>
    <sec id="sec-18">
      <title>5. Potential Contribution</title>
      <p>This research explores the impact of video materials
and the metaverse on students' understanding, cognitive
load, and affective states, enhancing our knowledge of
how these environments affect learning. It addresses
gaps in earlier studies, which mainly looked at learning
improvements, and points out that while video-based
learning is more common, its effectiveness still needs
more exploration. The study also guides creators of
educational videos to help improve the quality of their
content.</p>
      <p>Furthermore, the metaverse emerging as a new
educational platform opens new possibilities and
difficulties. This study aims to evaluate the utility of
advanced techniques such as eye-tracking technology
and facial expression recognition for monitoring
students' mental load and emotional states, which will
provide more in-depth insights into their experiences
during various learning activities. It focuses on the
influence of teacher and student presence in different
educational settings, providing fresh perspectives on
how it affects learning gain, visual attention, cognitive
load, and affective states.</p>
    </sec>
    <sec id="sec-19">
      <title>6. Discussion</title>
      <p>Some researchers study the effect of presence in
learning videos, while others explore how shared
learning space impacts learning. [29] investigates the
effect of lecturer presence in videos using interviews
only as a research method. The interview method has
drawbacks because it is subjective. Meanwhile, several
tools related to human response measurement can be
used, such as eye-tracking technology and facial
expression recognition [30] [31] [32]. Eye-tracking seems
particularly beneficial for studying mental processes
[33]. Thus, it is necessary to investigate the effect of
teacher presence on eye-tracking measures when
studying using videos.</p>
      <p>On the other hand, due to the rapid development of
artificial intelligence, automatic recognition of facial
expressions has been intensively studied in recent years
[34]. Emotion recognition through facial expression
detection is one of the essential fields of study for
humancomputer interaction [35]. However, more researchers
currently focus only on discreet emotion. At the same
time, the dimensional model is more effective than the
discrete model in providing a nuanced and continuous
representation of emotional experiences because it
focuses on measuring intensity along the dimensions of
valence and arousal instead of classifying emotions into
distinct categories. The API system developed by Daeha
Kim and Byung Cheol Song [26], used in this research, is
beneficial because it can detect learners' affective states
during learning in the metaverse.</p>
      <p>The chosen methodology is appropriate for filling the
research gap because combining tests, eye tracking, facial
expression recognition, and questionnaires is a
comprehensive method for collecting data on numerous
aspects of the student experience. The test permits the
measurement of learning gain. Eye tracking allows the
objective measurement of participants' pupil diameter
and visual attention, revealing their cognitive processes.
Facial expression recognition adds another layer of
objective data by documenting participants' affective
states, allowing for a more in-depth comprehension of
their emotional experiences. The questionnaire will
enable participants to give subjective assessments based
on their expertise when studying using a virtual
environment (metaverse).</p>
    </sec>
    <sec id="sec-20">
      <title>7. Current Data Analysis</title>
    </sec>
    <sec id="sec-21">
      <title>7.1. Video Experiment Analysis</title>
      <p>We have successfully gathered data from 33
participants using knowledge tests and eye-tracking
measures.</p>
    </sec>
    <sec id="sec-22">
      <title>7.1.1. Learning</title>
    </sec>
    <sec id="sec-23">
      <title>Experiment</title>
    </sec>
    <sec id="sec-24">
      <title>Gain of</title>
    </sec>
    <sec id="sec-25">
      <title>Video</title>
      <p>From the pre-test and post-test results, we calculate
the learning gain using this formula [36]</p>
      <p>We will compare the learning gain across three
different conditions. In examining the influence of a
lecturer's physical presence on students' learning gains,
we perform statistical analysis to see the difference
between no lecturer presence on the video vs. lecturer
physical presence, lecturer physical presence vs. lecturer
presence as an avatar, and no lecturer presence vs.
presence as an avatar.</p>
      <p>Besides that, it is essential to know the correlation
between learning gain and gaze transition. Transition is
switching gaze activity from 1 Area of Interest (AOI) to
another, back and forth. We aim to examine the
correlation between learning gain and gaze transition
between the two AOIs across all conditions (AOI 1 is the
main content/slide area, and AOI 2 is the
presenter/lecturer area).</p>
    </sec>
    <sec id="sec-26">
      <title>7.1.2. Pupil Diameter</title>
      <p>The output from eye tracking measures gives us
information about the participant’s pupil diameter. By
measuring pupil diameter or pupil dilation, we can link
to the cognitive load of the participant during learning
[23], [24], [25]. The participants watched all videos under
three conditions (no lecturer presence, lecturer physical
presence, and lecturer presence as an avatar), so
comparing which conditions gave them more pupil
dilations was interesting.</p>
      <p>To assess whether significant differences exist in
pupil diameter under three different conditions, we
conducted a series of statistical comparisons among no
lecturer presence on the video vs. lecturer physical
presence, lecturer physical presence vs. lecturer presence
as an avatar, and no lecturer presence vs. presence as an
avatar.</p>
      <p>Knowing the correlation between pupil diameter and
learning gain is also important. We want to explore this
relationship across three conditions. Moreover, we want
to see the correlation between pupil diameter and gaze
transition between the two AOIs across all conditions.</p>
    </sec>
    <sec id="sec-27">
      <title>7.1.3. Comparison of Time Spent</title>
      <p>on Hotspot Looked by Participants
in the Slides and the Number of</p>
    </sec>
    <sec id="sec-28">
      <title>Gaze Transition</title>
      <p>To assess attention allocation, we calculated the
percentage of time participants spent focusing on
designated hotspots in a slide presentation. These
analyses were designed to determine how the physical
presence of a lecturer influenced participant interaction
with the content, as highlighted by the hotspots on the
slides. Figure 7 shows an example of a hotspot on the
slide visualization.</p>
    </sec>
    <sec id="sec-29">
      <title>7.2. Metaverse</title>
    </sec>
    <sec id="sec-30">
      <title>Analysis</title>
    </sec>
    <sec id="sec-31">
      <title>7.2.1. Valence and Arousal</title>
    </sec>
    <sec id="sec-32">
      <title>Experiment</title>
      <p>We have collected data from 20 participants. We
compare valance and arousal levels between conditions
in this experiment (Individual Lecture vs. Group
Lecture). The result will be linked with the learning. For
example, how do participants’ valance and arousal levels
have meaning to their learning?</p>
    </sec>
    <sec id="sec-33">
      <title>7.2.2. Learning Gain of Metaverse</title>
    </sec>
    <sec id="sec-34">
      <title>Experiment</title>
      <p>We calculate the learning gain from the pre and
posttest scores and compare the results between two
conditions (individual vs group lecture) to see if there is
a significant difference.</p>
    </sec>
    <sec id="sec-35">
      <title>7.2.3. Questionnaire</title>
      <p>After learning on metaverse, participants were asked
to give their opinions to compare individual and group
lectures. We analyse their responses, which is valuable in
capturing how virtual environments give a sense of
presence. The questionnaire results can be combined
with the facial expression data to gain comprehensive
analysis.</p>
    </sec>
    <sec id="sec-36">
      <title>Acknowledgements</title>
      <p>The Indonesian Ministry of Education funds this
research. I also want to thank my supervisors, Dr.
MarieLuce Bourguet, Dr. Hamit Soyel, and Prof. Isabelle
Mareschal, for their guidance and support.</p>
      <p>S. Park and S. Kim, “Identifying World Types to
Deliver Gameful Experiences for Sustainable
Learning in the Metaverse,” Sustainability
(Switzerland), vol. 14, no. 3, 2022, doi:
10.3390/su14031361.</p>
      <p>K. Chorianopoulos and M. N. Giannakos,
“Usability design for video lectures,” in
Proceedings of the 11th European Conference on
Interactive TV and Video, EuroITV 2013, 2013. doi:
10.1145/2465958.2465982.</p>
      <p>C. M. Chen and C. H. Wu, “Effects of different
video lecture types on sustained attention,
emotion, cognitive load, and learning
performance,” Comput Educ, vol. 80, 2015, doi:
10.1016/j.compedu.2014.08.015.</p>
      <p>
        K. Ndihokubwayo, J. Uwamahoro, and I.
Ndayambaje, “Effectiveness of PhET
Simulations and YouTube Videos to Improve the
Learning of Optics in Rwandan Secondary
Schools,” African Journal of Research in
Mathematics, Science an
        <xref ref-type="bibr" rid="ref2">d Technology Education,
2020</xref>
        , doi: 10.1080/18117295.2020.1818042.
      </p>
      <p>M. Macedonia and T. R. Knösche, “Body in mind:
How gestures empower foreign language
learning,” Mind, Brain, and Education, vol. 5, no.
4, 2011, doi: 10.1111/j.1751-228X.2011.01129.x.
J. E. Driskell and P. H. Radtke, “The Effect of
Gesture on Speech Production and
Comprehension,” Human Factors: The Journal of
the Human Factors and Ergonomics Society, vol.
45, no. 3, pp. 445–454, Sep. 2003, doi:
10.1518/hfes.45.3.445.27258.</p>
      <p>S. Kita, M. W. Alibali, and M. Chu, “How do
gestures influence thinking and speaking? The
gesture-for-conceptualization hypothesis.,”
Psychol Rev, vol. 124, no. 3, pp. 245 –266, Apr.
2017, doi: 10.1037/rev0000059.</p>
      <p>
        Z. Yu, “The effect of teacher presence in videos
on intrinsic cognitive loads and academic
achievements,” Innovations in Education an
        <xref ref-type="bibr" rid="ref2">d
Teaching International, 2021</xref>
        , doi:
10.1080/14703297.2021.1889394.
      </p>
      <p>D. Özer and T. Göksun, “Gesture Use and
Processing: A Review on Individual Differences
in Cognitive Resources,” Frontiers in Psychology,
vol. 11. 2020. doi: 10.3389/fpsyg.2020.573555.</p>
      <p>X. Guo et al., “The Sense of Presence between
Volumetric-Video and Avatar-Based
Augmented Reality and Physical-Zoom
Teaching Activities,” Presence: Teleoperators and
Virtual Environments, vol. 28, pp. 267–280, May
2022, doi: 10.1162/PRES_a_00351.</p>
      <p>A. Steed, Y. Pan, F. Zisch, and W. Steptoe, “The
impact of a self-avatar on cognitive load in
immersive virtual reality,” in Proceedings - IEEE
Virtual Reality, IEEE Computer Society, Jul.
2016, pp. 67–76. doi: 10.1109/VR.2016.7504689.
K. Oestreicher, J. Kuzma, and D. Yen, “Avatar
Supported Learning in a Virtual University
Abstract : Avatar – Human Interaction,”
Worcester Journal of Learning and Teaching, no.
4, 2010.</p>
      <p>J. W. Woodworth, N. G. Lipari, and C. W. Borst,
“Evaluating teacher avatar appearances in
educational VR,” in 26th IEEE Conference on
Virtual Reality and 3D User Interfaces, VR 2019
Proceedings, 2019. doi: 10.1109/VR.2019.8798318.
[17]
[18]
[19]
[20]
[21]
[22]
[23]
[24]
[25]
[26]
[27]
[28]
[29]
[30]</p>
      <p>S. D’Mello and A. Graesser, “The half -life of
cognitive-affective states during complex
learning,” Cogn Emot, vol. 25, no. 7, 2011, doi:
10.1080/02699931.2011.613668.</p>
      <p>J. Z. Lim, J. Mountstephens, and J. Teo, “Emotion
recognition using eye-tracking: Taxonomy,
review, and current challenges,” Sensors
(Switzerland), vol. 20, no. 8. 2020. doi:
10.3390/s20082384.</p>
      <p>
        M. Pfaller, L. O. H. Kroczek, B. Lange, R. Fülöp,
M. Müller, and A. Mühlberger, “Social Presence
as a Moderator of the Effect of Agent Behavior
on Emotional Experience in Social Interactions
in Virtual Reality,” Front Virtual Real, vol. 2,
        <xref ref-type="bibr" rid="ref2">Dec.
2021</xref>
        , doi: 10.3389/frvir.2021.741138.
      </p>
      <p>J. K. Olsen, N. Rummel, and V. Aleven,
“Learning alone or together? A combination can
be best!,” in Computer-Supported Collaborative
Learning Conference, CSCL, 2017.</p>
      <p>D. W. Johnson and R. T. Johnson, “Learning
Together and Alone: Overview and Meta‐
analysis,” Asia Pacific Journal of Education, vol.
22, no. 1, 2002, doi: 10.1080/0218879020220110.
F. Kirschner, F. Paas, and P. A. Kirschner,
“Individual versus group learning as a function
of task complexity: An exploration into the
measurement of group cognitive load,” in
Beyond Knowledge: The Legacy of Competence:
Meaningful Computer-based Learning
Environments, 2008. doi:
10.1007/978-1-40208827-8_4.</p>
      <p>J. Zagermann, U. Pfeil, and H. Reiterer,
“Measuring cognitive load using eye-tracking
technology in visual computing,” in ACM
International Conference Proceeding Series, 2016.
doi: 10.1145/2993901.2993908.</p>
      <p>Z. Zheng, S. Gao, Y. Su, Y. Chen, and X. Wang,
“Cognitive load-induced pupil dilation reflects
potential flight ability,” Current Psychology,
2022, doi: 10.1007/s12144-022-03430-2.</p>
      <p>V. Peysakhovich, F. Dehais, and M. Causse,
“Pupil Diameter as a Measure of Cognitive Load
during Auditory-visual Interference in a Simple
Piloting Task,” Procedia Manuf, vol. 3, 2015, doi:
10.1016/j.promfg.2015.07.583.</p>
      <p>D. Kim and B. C. Song, “Optimal Transport
based Identity Matching for Identity-invariant
Facial Expression Recognition,” 2022.</p>
      <p>T. T. A. Höfling, A. B. M. Gerdes, U. Föhl, and G.
W. Alpers, “Read My Face: Automatic Facial
Coding Versus Psychophysiological Indicators
of Emotional Valence and Arousal,” Front
Psychol, vol. 11, 2020, doi:
10.3389/fpsyg.2020.01388.</p>
      <p>S. Bianco et al., “A Smart Mirror for Emotion
Monitoring in Home Environments,” Sensors,
vol. 21, no. 22, p. 7453, Nov. 2021, doi:
10.3390/s21227453.</p>
      <p>
        Z. Yu, “The effect of teacher presence in videos
on intrinsic cognitive loads and academic
achievements,” Innovations in Education an
        <xref ref-type="bibr" rid="ref2">d
Teaching International, 2021</xref>
        , doi:
10.1080/14703297.2021.1889394.
      </p>
      <p>J. Zhang, M.-L. Bourguet, and G. Venture, “The
Effects of Video Instructor’s Body Language on
Students’ Distribution of Visual Attention: an
Eye-tracking Study,” 2018. doi:
10.14236/ewic/hci2018.101.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          <string-name>
            <given-names>B. A.</given-names>
            <surname>Pribadi</surname>
          </string-name>
          and
          <string-name>
            <given-names>R.</given-names>
            <surname>Susilana</surname>
          </string-name>
          , “
          <article-title>The use of mind mapping approach to facilitate students' distance learning in writing modular based on printed learning materials</article-title>
          ,”
          <source>European Journal of Educational Research</source>
          , vol.
          <volume>10</volume>
          , no.
          <issue>2</issue>
          , pp.
          <fpage>907</fpage>
          -
          <lpage>917</lpage>
          , Apr.
          <year>2021</year>
          , doi: 10.12973/EU-JER.
          <year>10</year>
          .2.907.
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          <string-name>
            <given-names>D.</given-names>
            <surname>Pal</surname>
          </string-name>
          and
          <string-name>
            <given-names>S.</given-names>
            <surname>Patra</surname>
          </string-name>
          , “University Students'
          <article-title>Perception of Video-Based Learning in Times of COVID-19: A TAM/TTF Perspective,”</article-title>
          <source>Int J Hum Comput Interact</source>
          , vol.
          <volume>37</volume>
          , no.
          <issue>10</issue>
          ,
          <year>2021</year>
          , doi: 10.1080/10447318.
          <year>2020</year>
          .
          <volume>1848164</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          <string-name>
            <given-names>S.</given-names>
            <surname>Lackmann</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P. M.</given-names>
            <surname>Léger</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Charland</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Aubé</surname>
          </string-name>
          , and
          <string-name>
            <given-names>J.</given-names>
            <surname>Talbot</surname>
          </string-name>
          , “
          <article-title>The influence of video format on engagement and performance in online learning</article-title>
          ,
          <source>” Brain Sci</source>
          , vol.
          <volume>11</volume>
          , no.
          <issue>2</issue>
          ,
          <year>2021</year>
          , doi: 10.3390/brainsci11020128.
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>