<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <journal-title-group>
        <journal-title>E. Iacobelli);</journal-title>
      </journal-title-group>
      <issn pub-type="ppub">1613-0073</issn>
    </journal-meta>
    <article-meta>
      <title-group>
        <article-title>Application for Engagement Detection</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Emanuele Iacobelli</string-name>
          <email>iacobelli@diag.uniroma1.it</email>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Samuele Russo</string-name>
          <email>samuele.russo@uniroma1.it</email>
          <xref ref-type="aff" rid="aff2">2</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Christian Napoli</string-name>
          <email>cnapoli@diag.uniroma1.it</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
          <xref ref-type="aff" rid="aff4">4</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Department of Computational Intelligence, Czestochowa University of Technology</institution>
          ,
          <addr-line>42-201 Czestochowa</addr-line>
          ,
          <country country="PL">Poland</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Department of Computer, Control and Management Engineering, Sapienza University of Rome</institution>
          ,
          <addr-line>00185 Roma</addr-line>
          ,
          <country country="IT">Italy</country>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>Department of Psychology, Sapienza University of Rome</institution>
          ,
          <addr-line>00185 Roma</addr-line>
          ,
          <country country="IT">Italy</country>
        </aff>
        <aff id="aff3">
          <label>3</label>
          <institution>Engagement Detection</institution>
          ,
          <addr-line>Eye Tracking, Face Expression Recognition, Machine Learning</addr-line>
          ,
          <country>Residual Neural Networks</country>
        </aff>
        <aff id="aff4">
          <label>4</label>
          <institution>Institute for Systems Analysis and Computer Science, Italian National Research Council</institution>
          ,
          <addr-line>00185 Roma</addr-line>
          ,
          <country country="IT">Italy</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>1846</year>
      </pub-date>
      <volume>000</volume>
      <fpage>0</fpage>
      <lpage>0002</lpage>
      <abstract>
        <p>The study of human engagement has significantly grown in recent years, particularly accelerated by the interaction with a growing number of smart computing machines [1, 2, 3]. Engagement estimation has significant importance across various domains of study, including advertising, marketing, human-computer interaction, and healthcare [4, 5, 6]. In this paper, we propose a real-time application that leverages a single RGB camera to capture user behavior. Our approach implements a novel method for estimating human engagement in real-world scenarios by extracting valuable information from the combination of facial expressions and gaze direction analysis. To acquire this data, we employed fast and accurate machine learning algorithms from the external library dlib, along with custom versions of Residual Neural Networks implemented from scratch. For training our models, we used a modified version of the DAiSEE dataset, a multi-label user afective states classification dataset that collects frontal videos of 112 diferent people recorded in real-world scenarios. In the absence of a baseline for comparing the results obtained by our application, we conducted experiments to assess its robustness in estimating engagement levels, leading to very encouraging results.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>CEUR
ceur-ws.org</p>
    </sec>
    <sec id="sec-2">
      <title>1. Introduction</title>
      <p>In today’s rapidly evolving digital landscape, humanity
interacts with a growing number of smart computing
machines. This situation highlights the increasing trend
of direct interactions with smart devices in various
doand industrial applications. Despite this technological
advancement, many devices lack algorithms capable of
perceiving and responding to users’ attentional states.</p>
      <sec id="sec-2-1">
        <title>Traditional user interfaces still heavily rely on explicit input or predefined triggers, resulting in often ineficient and mechanical interactions.</title>
      </sec>
      <sec id="sec-2-2">
        <title>The potential for automatic acquisition and interpreta</title>
        <p>tion of users’ engagement represents a huge usability
improvement for Human-Computer Interaction (HCI) and
Human-Robot Interaction (HRI) systems. This capability
itive interactions, elevating system responsiveness, and
enhancing overall user experience. In detail, engagement
is a fundamental aspect of the human experience and
nEvelop-O
(C. Napoli)
SYSYEM 2023: 9th Scholar’s Yearly Symposium of Technology,
Engiment, focus, and interaction with their surroundings. For
detecting it, facial expressions and gaze direction are
crucial elements. In particular, the motion of the eyes is
an important element to employ since it highlights the
psychological mechanisms behind the human mind and
naturally gravitates toward objects, people, or specific</p>
        <p>In this paper, we propose a real-time application that
combines gaze direction and face expression analysis to
determine the engagement level of a person while
interacting with intelligent systems. To achieve this, we
defined two machine-learning pipelines leveraging RGB
videos of a person interacting with the system. The first
pipeline focuses on the user’s facial expressions analysis
and employs a residual neural network architecture. The
second pipeline concentrates on the user’s gaze
direction estimation by combining pre-trained face and facial
algorithm that we developed. Predictions of the user’s
engagement level are ultimately calculated by merging
the outputs of these two pipelines using a weighted linear</p>
      </sec>
      <sec id="sec-2-3">
        <title>Addressing the challenge posed by the absence of a</title>
        <p>baseline for reference, our primary hurdle in handling
this task involved creating an appropriate dataset for
training our models. We opted to customize the Afective</p>
      </sec>
      <sec id="sec-2-4">
        <title>States in E-Environment Dataset (DAiSEE) [7], a compre</title>
        <p>hensive collection of multi-label videos designed for
idenholds the promise of ushering in more advanced and intu- landmark detection models with a fast computer vision
captures in depth the quality of an individual’s involve- interpolation formula.
(S. Russo); 0000-0002-3336-5853 (C. Napoli)</p>
        <p>© 2023 Copyright for this paper by its authors. Use permitted under Creative Commons License tifying user afective states. Given that the estimation of
engagement levels necessitates both temporal and spatial ease of implementation and their proven efectiveness
information, videos proved to be an ideal choice. How- in achieving accurate results. The prominence of such
ever, to mitigate the high computational and resource techniques underscores the significance of visual cues,
costs associated with treating videos as opposed to single particularly facial expressions, in gauging user
engageimages, we implemented mandatory preprocessing steps ment levels during online interactions.
to optimize memory and computational eficiency. The work presented in [18] investigates the suitability</p>
        <p>
          Continuing to tackle the absence of a reference base- of three popular models: All-CNN [19], NiN-CNN [20],
line, we conducted experiments to assess the robustness and VD-CNN [21], along with a customized
Convoluand efectiveness of our application. This evaluation was tional Neural Network (CNN) [
          <xref ref-type="bibr" rid="ref39">22</xref>
          ] for detecting
engagecarried out using quantitative metrics. ment level of online learners in educational activities.
All the analyzed models leverage facial expressions for
1.1. Roadmap scalable and accessible engagement detection. Each of
the three base models has its distinct features and the
cusThis paper is organized in the following way: first of all, a tomized CNN combines these advantageous features. For
summary of the state-of-the-art systems and techniques instance, by replacing linear convolutional layers with a
to recognize the human engagement level is presented multilayer perceptron, increasing depth with small
con(see Section 2). Subsequently, a description of the dataset volutional filters, and replacing some max-pooling layers
that we have developed for training our models is illus- with convolutional layers with increased stride. All the
trated (see Section 3). Following this, a detailed overview analyzed models were evaluated on the DAiSEE dataset
of the architectures employed for our application is pro- (extensively explained in Section 3) and the results reveal
vided (see Section 4). Then, the results obtained by test- that the customized CNN outperforms the base models
ing our system considering the quantitative metrics are in detecting the engagement level.
presented (see Section 5). Finally, we summarize the In a similar study proposed in [23], the automatic
article’s content and outline the possible viable improve- recognition of student engagement from facial
expresments that can be made to our application (see Section sions is examined using a three-stage pipeline. The initial
6). step involves face registration, detection, and the
estimation of key facial landmarks (e.g., eyes, nose, and mouth)
2. Related Works by using the approach described in [24]. The second stage
employs four binary classifiers to classify the cropped
The field of engagement level detection has seen signifi- face, distinguishing whether it belongs to one of four
cant growth, particularly fueled by the global pandemic. engagement levels ( ∈ 1, 2, 3, 4 ), where 1 signifies no
With many individuals compelled to participate in re- engagement and 4 represents full focus. The authors
mote meetings, analyzing engagement in online sessions compared three models for the binary classifier:
Suphas become a pivotal focus, leading to the development port Vector Machines with Gabor features (SVM (Gabor))
of numerous systems. Some studies have explored phys- [24], Multinomial Logistic Regression with expression
iological factors like fatigue [8], brain status and data outputs from the Computer Expression Recognition
Tool[
          <xref ref-type="bibr" rid="ref16 ref7">9, 10</xref>
          ], blood flow and heart rate [ 11], and galvanic skin box (MLR(CERT)) [24], and GentleBoost with Box Filter
conductance [12]. However, due to the recent needs and features (Boost(BF)) [25]. This study reveals that SVM
the remote nature of this task, there has been widespread (Gabor) yields the best results. The third stage integrates
exploration of inexpensive and unobtrusive technolo- the outputs of all four binary classifiers, utilizing a
Multigies. Eye trackers [13, 14] and facial expression recogni- nomial Logistic Regressor model to estimate the final
tion models [
          <xref ref-type="bibr" rid="ref56">15, 16</xref>
          ] using simple RGB cameras are now engagement level.
among the most promising options. In [26], the authors introduced a regression model
        </p>
        <p>In a comprehensive review treated in [17], the state- for predicting engagement level as a single scalar value
of-the-art engagement detection techniques within the from RGB video streams captured by two cameras on the
context of online learning are explored. The authors torso and head of an autonomous mobile robot, utilized
classify existing methods into three primary categories: for tours at The Collection museum in Lincoln, UK. The
automatic, semi-automatic, and manual. This classifica- model incorporates CNN and Long Short-Term Memory
tion is based on the methods’ dependencies on learners’ (LSTM) [27] networks for video data analysis.
Trainparticipation. Furthermore, each category is subdivided ing and evaluation of this regressor network were
conbased on the type of input data used (e.g., audio, video, ducted using a dataset built from the recordings of the
text). Among these, video-based methods in the auto- autonomous tour guide robot in the public museum. The
matic category that leverage facial expressions emerge as dataset, manually annotated by three independent
peothe most prevalent. These methods are favored for their ple, assigns scalar values in the range [0,1] to represent
the user’s engagement level. The model demonstrates
optimal engagement level predictions, achieving a Mean driver attention estimation, generating a heat map on the
Squared Error (MSE) prediction loss of up to 0.126 on the images representing the road. The training dataset for
test dataset. this model is constructed using virtual reality and a
driv</p>
        <p>The research conducted in [28] focused on investigat- ing simulator, incorporating images from the DR(eye)VE
ing the Deep Facial Spatiotemporal Network (DFSTN). dataset [32] that depict the frontal view of the road
obComprising two integral modules, namely the pretrained served by the driver. Experimental results showcase the
SE-ResNet-50 (SENet) utilized for extracting facial spatial feasibility and superiority of the proposed method over
features and an LSTM network with Global Attention existing approaches.
for generating an attentional hidden state, the DFSTN
synergistically captures both facial spatial and
temporal information. This combined information is crucial 3. Dataset
for enhancing engagement prediction performance. The
model underwent testing on the DAiSEE dataset, achiev- The baseline dataset utilized for training our networks is
ing an accuracy of 58.84%, showcasing its capability to a customized version of the Dataset for Afective States
outperform numerous existing engagement prediction in E-Environments (DAiSEE), a large collection of
multinetworks trained on the same dataset. label videos designed for identifying user afective states,</p>
        <p>In [29], the estimation of human attention is based on including boredom, confusion, engagement, and
frustrathe direction of the user’s face, considering five diferent tion in real-world scenarios. This dataset comprises 9068
directions: central, lateral to the left, lateral to the right, frontal view videos featuring 112 distinct individuals
extowards up, and towards down. If the user looks in any di- pressing diferent levels of afective states. Each of these
rection other than the central one, they are assumed to be states was manually ranked utilizing the following scale:
distracted, with only the central gaze indicating full focus. very low, low, high, and very high.
The authors created a dataset for training, comprising To create our customized dataset we initially modified
270 videos of approximately 20 seconds each from 18 dif- the task from which DAiSEE was originally built. We
ferent individuals. To enhance data diversity, GAN-based switched from multi-label to multi-class classification,
data augmentation techniques were employed to gener- associating only the level of engagement with each video
ate new samples, diversifying somatic features in the and removing the labels for the other afective states.
Exrecorded videos. Transfer Learning [30] was utilized to ample instances present inside our customized version
construct the classifier. Specifically, a pre-trained VGG16 of the DAiSEE are displayed in Fig. 1. Subsequently, we
[21] architecture was employed, with three additional divided the dataset into Training, Validation, and Test
dense layers attached at the end for attention estimation. sets, with proportions of 60%, 20%, and 20%, respectively.</p>
        <p>The approach presented in [31] ofers a novel method However, the resulting sets were highly unbalanced due
for estimating driver attention. Departing from conven- to a small portion of videos classified as very low and
tional methods that primarily focus on a single frontal low engagement. To address this issue, we downsampled
scene image to analyze driver gaze or head pose, this the dataset in several ways to achieve a more balanced
method introduces a dual-view scene. The additional distribution. First of all, redundancy in subjects was
input data includes the frontal view of the car that the reduced by removing multiple videos of the same
indidriver is observing. Specifically, the gaze direction is viduals. Then, through the use of a normal distribution,
detected and transformed into a probability map of the we sampled the remaining data instances considering
same size as the road view image, while salient features the frequency of labels in the videos with the following
of temporal and spatial dimensions are extracted from the formula:
road view images. These features are then combined and   =   ⋅   −   ⋅  (1)
fed into a multi-resolution neural network tasked with  
where  is the reduction coeficient (that we have set
to 0.25),   represents the total number of samples in a
given set, and   denotes the frequency of label  in that
set. Table 1 displays information both before and after
the preprocessing procedures on the dataset.</p>
        <p>Since DAiSEE contains recordings captured in dynamic
environments, each of these videos may present diferent
and various disturbances, such as changing light
conditions, visual occlusions, or unconstrained user motion.</p>
        <p>To improve video quality, we applied manual color and
intensity adjustments, focusing on enhancing contrast,
brightness, and sharpness for optimal detail resolution.</p>
        <p>Examples include adjustments to the gamma value, which
efectively improves visibility in varying light conditions
or exposure levels by normalizing image histograms,
making videos more suitable for continuous analysis;
Another example is the sharpness adjustments, which
enhance fine details and edges, making facial features
more prominent.</p>
        <p>Despite these modifications, the dataset still demanded
excessive memory requirements. Consequently, we
opted for further adjustments. Considering that the
majority of engagement information is likely derived from
human expressions and gaze attention, with a smaller
contribution from gestures, we decided to crop from each
video only the user’s faces. This step also aimed to
eliminate potential issues and biases arising from background
data. The face cropping was automated using a
pretrained Single Shot Multibox Detector (SSD) model from
the Cafe framework [ 33].</p>
        <p>To prevent the generation of unstable videos, we
applied a stabilization algorithm (see pseudocode in Fig. 2)
that facilitates smooth transitions between subsequently
detected faces by stabilizing the position of their
bounding boxes. At the start of each video, the size of the first
detected face’s bounding box is stored. In all the
following frames, this dimension is used to resize the bounding
box of the subsequently detected faces. Additionally, if
the distance between the centers of two consecutive
detected faces is smaller than a manually adjusted threshold
 , the center of the newest detected face is replaced with
the center of the bounding box of the previously detected
face. Finally, each frame is converted to grayscale, and
histogram equalization is applied to normalize the color
information.
frame is passed to the Gaze Direction Module and the Face Engagement Model. The predictions of these models are then
combined to produce the actual output of our system.
4.</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>Methodology</title>
      <sec id="sec-3-1">
        <title>The complete architecture of the real-time application we</title>
        <p>developed is illustrated in Fig. 3. Specifically, the system
utilizes a single input video stream captured through
a webcam reader module, implemented in the external
library OpenCV [34], to feed two distinct models. The
Face Engagement Model evaluates engagement based
on facial expressions, while the Gaze Direction Model
predicts engagement by analyzing where the user’s focal
point. Lastly, the predictions of these two models are
combined to derive the final engagement level estimated
by our application.</p>
        <sec id="sec-3-1-1">
          <title>4.1. Face Engagement Model</title>
        </sec>
      </sec>
      <sec id="sec-3-2">
        <title>This model is designed to estimate the user engagement</title>
        <p>
          level from frontal recording videos. We designed it as a
customized version of the ResNet architecture [
          <xref ref-type="bibr" rid="ref46">35</xref>
          ] and
we implemented diferent versions to identify the most
efective one. In essence, a residual network employs skip
connections to address the vanishing gradient problem.
These connections allow information to directly
backpropagate, circumventing previous layers. Moreover, a
skip connection facilitates a residual block in learning
the residual, which is the diference between the desired
output and the current input of the layer. This approach
makes it easier for the network to understand what input
modifications are needed to achieve the desired output,
rather than altering the entire input from scratch. This
often translates to a more straightforward learning process
for the network.
        </p>
        <p>To address the human engagement level classification
problem, our models needed to capture both spatial and
temporal information. To enable the network to learn
temporal information by analyzing multiple frames
simultaneously in the same layer, we opted for 3D
convolutional layers instead of the traditional 2D convolutional
layers implemented in the original ResNet architecture.
Learning temporal information is crucial for video
analysis, as it allows the network to recognize complex
patterns such as actions, gestures, or sequences of facial
expressions. Due to this requirement, the model
necessitates an initial period to populate a bufer of 60 frames,
ensuring a suficient amount of data for the correct
utilization of the 3D convolutional layers. Once the bufer
reaches its capacity, the prediction of the engagement
level can begin. Subsequently, with the arrival of each
new frame, the bufer is updated, and the oldest frame is
discarded. We tested three versions of this architecture,
difering mainly in the depth and the internal structure
of the convolutional block used. Specifically, we
implemented the 18-, 34-, and 50-layer versions.</p>
        <p>For all these architectures, we introduced 3D layers
for batch normalization, max pooling, and average
pooling. In detail, each convolutional block includes a batch
normalization layer, and all convolutional layers employ
the ReLU activation function. Only the last fully
connected layer, responsible for the final prediction of the
human’s engagement level, uses the Softmax activation
function. During training, we utilized the He/Kaiming
initialization technique [36], which initializes weights
us
ing a normal distribution with zero mean and a variance
of 2 , where  is the total number of inputs to the neuron.</p>
      </sec>
      <sec id="sec-3-3">
        <title>This initialization is specifically tailored for networks</title>
        <p>employing the ReLU activation function, mitigating the
vanishing or exploding gradient problem.</p>
      </sec>
      <sec id="sec-3-4">
        <title>Additionally, we employed the Focal Loss [37] as the</title>
        <p>training function, opting for it over the conventional
Categorical Cross-Entropy. The principal reason is that the</p>
      </sec>
      <sec id="sec-3-5">
        <title>Focal Loss addresses the issue of unbalanced data by prioritizing examples the model struggles with, rather than those it confidently predicts. This ensures continuous</title>
        <p>improvement on challenging examples, preventing the
model from becoming overly confident with easy ones.
We implemented the following Focal Loss formula:
− (1 −   ) ln(  )
where  represents the focusing parameter (typically a
positive number) to be fine-tuned using cross-validation,
and   denotes the predicted probability of the correct
class. Also, our training process incorporates early
stopping with a learned patience value of 10 epochs, L2
regularization featuring a weight decay set to 1 −3, and an
Adam Optimizer accompanied by a Learning Rate
Scheduler [38] with a maximum learning rate of 1 − 4 and a
Gradient Scaler to reduce the range of magnitudes in the
gradients. All the implementation details of the tested
models are reported in the Table 2. Following the
training phase, we opted for the 50-layer version model, with
a batch size equal to 16, as the engagement network for
our application, as it demonstrated the highest accuracy
among the tested versions.</p>
        <sec id="sec-3-5-1">
          <title>4.2. Gaze Direction Model</title>
          <p>This model is designed to extract attention information
from a person’s gaze in frontal recording videos. The
gaze direction provides valuable insights into a person’s
engagement during a task. The complete workflow of
this model is displayed in Fig. 4. To implement this
model, we combined two pre-trained neural networks
available in the dlib library [39].</p>
          <p>The first network is the Face Cropping Model, a CNN
trained for face detection in general images. It not only
identifies faces but also provides their bounding box
coordinates and converts the input image to grayscale.
Although the use of this network may appear redundant
considering the customized dataset that we have
employed for training the engagement model, it plays a
crucial role in the real-time application. Specifically, it
crops faces from the live stream frames and passes these
images to both the Engagement Model and the Facial
Landmark Detector.</p>
          <p>The Facial Landmark Detector, the second network
x2
x2
x2
x2
Convolution</p>
          <p>Block
k=[3,3,3],f=128
k=[3,3,3],f=128
k=[3,3,3],f=128
k=[3,3,3],f=128
Convolution</p>
          <p>Block
k=[3,3,3],f=256
k=[3,3,3],f=256
k=[3,3,3],f=256
k=[3,3,3],f=256
Convolution</p>
          <p>Block
k=[3,3,3],f=512
k=[3,3,3],f=512
k=[3,3,3],f=512
k=[3,3,3],f=512
Average Pool</p>
          <p>Dropout
Linear</p>
          <p>Output Size = 1x1x1</p>
          <p>Rate = 0.4
Neurons = 1024
x3
x4
x6
x3
k=[1,1,1],f=64
k=[3,3,3],f=64
k=[1,1,1],f=256</p>
          <p>x3
k=[1,1,1],f=128
k=[3,3,3],f=128
k=[1,1,1],f=512</p>
          <p>x4
k=[1,1,1],f=256
k=[3,3,3],f=256
k=[1,1,1],f=1024</p>
          <p>x6
k=[1,1,1],f=512
k=[3,3,3],f=512
k=[1,1,1],f=2048
x3
that we have employed from the dlib library, recognizes
68 2D facial landmarks (e.g., nose tip, corners of the
mouth, and eyes) in a given face image. These facial
landmarks serve two purposes: they are used to crop the
eye regions based on the eye landmarks and to calculate
the face orientation with respect to the vertical axis (yaw
angle). This orientation is determined through the use
of a vector starting from the midpoint between the eyes
and terminating at the nose tip.</p>
          <p>Estimating the focal point of the user is accomplished
through the Gaze Direction Module, a simple computer
vision pipeline. Initially, the eye landmarks outlining
the eye contours are employed to create a mask that
removes extraneous pixels from each cropped eye image.
Subsequently, the Otsu’s method [40] is applied to
automatically threshold the image, distinguishing between
foreground (iris and pupil pixels) and background (sclera
pixels).</p>
          <p>The resulting image is then horizontally and vertically
divided around its center to estimate the gaze direction.
Both vertical and horizontal gaze directions are
quantiifed as values within the range of [-1,1]. Regarding
horizontal gaze direction, a value approaching -1 indicates
the user is looking to the left, while a value approaching
1 suggests a rightward gaze. A value around 0 indicates
the user is looking at the center of the screen. Similarly,
for vertical gaze direction, a value nearing 1 signifies
a downward gaze and a value nearing -1 indicates an
upward gaze.</p>
          <p>To compute these directions, the density of white
pixels representing the sclera is analyzed. For each eye
image, the total number of white pixels is calculated. If this
value is zero, it implies incorrect eye detection, and the
current frame is skipped. Otherwise, for each sub-image
generated, the percentage of white pixels in relation to
the total number of white pixels in the corresponding
original eye image is calculated. Then, the percentages
belonging to the same direction of both eyes are averaged
(e.g., the percentage of white pixels in the left sub-image
of the left eye is averaged with the percentage of white
pixels in the left sub-image of the right eye). Finally, the
diference between these averages produces the value
within the range of [-1,1] described earlier.</p>
          <p>To efectively use the estimated gaze direction, it’s
crucial to consider the limits of the user’s field of view,
which may vary based on the task. In our screen-based
task implementation, we assume that a face orientation
deviation exceeding 30 degrees from the camera-aligned
orientation indicates the user is no longer looking at the
monitor.</p>
          <p>Initially, these limits are set at the task’s beginning and
dynamically adjusted based on the user’s face position
and orientation relative to the camera frame’s center. If
the face orientation exceeds 20 degrees from the frontal
position, the horizontal limits shift proportionally based
on the sine of the face orientation. Updates related to
face position involve calculating the distance between
the face bounding box center and the frame center. If this
distance exceeds one-sixth of the total frame dimension,
the right and left limits are adjusted. The adjustment is
determined by normalizing the distance between the face
and frame centers between 0 and 0.5. If the face shifts
to the right, the distance is subtracted from the limits;
otherwise, it is added.</p>
          <p>The engagement level, derived from the gaze direction,
is within the range [0,1]. It is obtained by subtracting
the sum of horizontal and vertical gaze errors from 1. A
score of 1 indicates complete focus on the screen, with
no gaze exceeding the defined limits. A score of 0 implies
no face detection in the current frame. The closer the
engagement level is to zero, the more the user surpasses
the admissible field of view limits, indicating a lack of
focus on the task. Specifically, when the gaze exceeds the
limits, horizontal and vertical gaze errors are calculated
as the diference in modulo between the estimated gaze
direction and the corresponding limits.</p>
        </sec>
        <sec id="sec-3-5-2">
          <title>4.3. Engagement Level Estimation</title>
          <p>To obtain the final detected engagement level, we
combined the predictions from the Face Engagement Model
and the Gaze Direction Model using a linear interpolation
formula:
 ⋅  
 
+ (1 −  ) ⋅  

(3)</p>
        </sec>
      </sec>
      <sec id="sec-3-6">
        <title>Where  is a learnable parameter used to weigh the</title>
        <p>importance of the models’ predictions. In addition, to
correctly apply this formula, the prediction of the Face
Engagement Model needs to be converted from labels to a
value within the range [0,1]. The conversion is performed
according to the rules displayed in Table 3.</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>5. Results</title>
      <sec id="sec-4-1">
        <title>To evaluate the accuracy of our system, we measured the</title>
        <p>disparity between the predicted engagement scores and
the ground truth values using two regression metrics:
Mean Absolute Error (MAE) and Mean Absolute
Percentage Error (MAPE). To facilitate the application of these
0.5, our system achieved an accuracy of approximately
58% (57.7%), closely aligning with the performance of
state-of-the-art works in engagement level detection
discussed in Section 2 that work with the original version
of the DAiSEE.</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>6. Conclusions</title>
      <sec id="sec-5-1">
        <title>Our work introduces a novel approach to engagement</title>
        <p>level estimation by integrating two distinct machine
learning pipelines focused on analyzing facial
expressions and gaze direction. Noteworthy is our real-time
application’s emphasis on cost-efectiveness and
accessibility, achieved through the utilization of a single RGB
Figure 5: Trend of the Mean Absolute Error (MAE) with vary- camera, fast and lightweight machine learning
algoing values of the parameter  in Eq. (3). rithms, and computationally eficient computer vision
techniques.</p>
        <p>In terms of system training, we customized the DAiSEE
dataset to optimize memory usage, reduce class
imbalance, mitigate bias introduced by repeated instances of
the same individuals, and focus exclusively on facial
cropping to eliminate potential background-related biases.</p>
        <p>The achieved results underscore the potential of our
system as a robust foundation, ofering a secure benchmark
for the development of innovative applications
integrating automatic user engagement recognition, thereby
dynamically adapting to user interactions. This not only
enhances overall usability but also heralds a new era in
application interfaces, promising heightened levels of
user experience and interaction.</p>
        <p>Looking forward, future improvements to our system
can be directed towards enhancing the accuracy,
robustFigure 6: Trend of the Mean Absolute Percentage Error ness, and generalization capabilities by expanding the
(MAPE) with varying values of the parameter  in Eq. (3). dataset’s dimensions. This expansion may involve
incorporating data from a more diverse group, encompassing
individuals with varying demographic characteristics,
metrics, we converted the engagement level labels associ- cultural backgrounds, and engagement patterns.
ated with the samples in our customized dataset using the Also, exploring attention estimation in multi-face
conconversion rules outlined in Table 3. This transformation texts, where multiple individuals are present
simultaneefectively turned the multi-label class problem, designed ously, represents another intriguing avenue for future
refor the DAiSEE dataset, into a regression problem. search. Lastly, a significant refinement to our application</p>
        <p>During training, we experimented with diferent val- involves substituting the CNN layers in the Face
Detecues for the parameter  in Eq. (3) to maximize the sys- tion Model with Visual Transformers [41](ViTs), known
tem’s accuracy. As illustrated in Figs. 5 and 6, the lowest for their excellence in image manipulation and
longerror for both MAE and MAPE occurred when  was set range dependency modeling compared to traditional
conto 0.5. This indicates that both predictions from the Face volutional layers. This substitution could enhance the
Engagement Model and the Gaze Direction Model carry precision of engagement level estimation from facial
exequal importance and are essential for achieving accurate pressions, as diferent facial regions can be efectively
predictions. combined at the same time.</p>
        <p>Analysis of the scenarios where  is 0 (using only the
Face Model) or 1 (using only the Gaze Model) reveals References
significantly higher errors in both performance metrics.</p>
        <p>Independently, these predictions struggle to accurately
gauge the user’s engagement level. With  initialized to
[1] G. Capizzi, G. L. Sciuto, C. Napoli, M. Woźniak,</p>
        <p>G. Susi, A spiking neural network-based
long</p>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          <article-title>term prediction system for biogas production, Neu- tion on driver distraction</article-title>
          ,
          <source>in: 2015 IEEE 12th inter-</source>
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          <source>ral Networks</source>
          <volume>129</volume>
          (
          <year>2020</year>
          )
          <fpage>271</fpage>
          -
          <lpage>279</lpage>
          . doi:
          <volume>10</volume>
          .1016/j. national
          <source>conference on wearable and implantable</source>
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          <string-name>
            <surname>neunet.</surname>
          </string-name>
          <year>2020</year>
          .
          <volume>06</volume>
          .001.
          <article-title>body sensor networks (BSN)</article-title>
          , IEEE,
          <year>2015</year>
          , pp.
          <fpage>1</fpage>
          -
          <lpage>6</lpage>
          . [2]
          <string-name>
            <given-names>N.</given-names>
            <surname>Brandizzi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Russo</surname>
          </string-name>
          , G. Galati,
          <string-name>
            <given-names>C.</given-names>
            <surname>Napoli</surname>
          </string-name>
          , Address- [13]
          <string-name>
            <given-names>E.</given-names>
            <surname>Iacobelli</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Ponzi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Russo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Napoli</surname>
          </string-name>
          , Eye-
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          <article-title>solution to user clustering using recency-frequency- opment and evaluation</article-title>
          ,
          <source>Information</source>
          <volume>14</volume>
          (
          <year>2023</year>
          )
          <fpage>644</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          <source>monetary and vehicle relocation based on neigh-</source>
          [14]
          <string-name>
            <given-names>F.</given-names>
            <surname>Fiani</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Russo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Napoli</surname>
          </string-name>
          , An advanced solu-
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          <string-name>
            <surname>borhood</surname>
            <given-names>splits</given-names>
          </string-name>
          ,
          <source>Information (Switzerland) 13</source>
          (
          <year>2022</year>
          ).
          <article-title>tion based on machine learning for remote emdr</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          <source>doi:10</source>
          .3390/info13110511. therapy,
          <source>Technologies</source>
          <volume>11</volume>
          (
          <year>2023</year>
          ). doi:
          <volume>10</volume>
          .3390/ [3]
          <string-name>
            <given-names>A.</given-names>
            <surname>Alfarano</surname>
          </string-name>
          , G. De Magistris,
          <string-name>
            <given-names>L.</given-names>
            <surname>Mongelli</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Russo</surname>
          </string-name>
          ,
          <year>technologies11060172</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          <string-name>
            <given-names>J.</given-names>
            <surname>Starczewski</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Napoli</surname>
          </string-name>
          , A novel convmixer trans- [15]
          <string-name>
            <given-names>P.</given-names>
            <surname>Kaur</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Krishan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S. K.</given-names>
            <surname>Sharma</surname>
          </string-name>
          , T. Kanchan,
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          <source>tection 14126 LNAI</source>
          (
          <year>2023</year>
          )
          <fpage>3</fpage>
          -
          <lpage>16</lpage>
          . doi:
          <volume>10</volume>
          .1007/ Medicine,
          <source>Science and the Law</source>
          <volume>60</volume>
          (
          <year>2020</year>
          )
          <fpage>131</fpage>
          -
          <lpage>139</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          978- 3-
          <fpage>031</fpage>
          - 42508-
          <issue>0</issue>
          _
          <fpage>1</fpage>
          . [16]
          <string-name>
            <surname>G. De Magistris</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          <string-name>
            <surname>Romano</surname>
            , J. Starczewski, [4]
            <given-names>G.</given-names>
          </string-name>
          <string-name>
            <surname>Capizzi</surname>
            ,
            <given-names>G. L.</given-names>
          </string-name>
          <string-name>
            <surname>Sciuto</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          <string-name>
            <surname>Napoli</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          <string-name>
            <surname>Tramontana</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          <string-name>
            <surname>Napoli</surname>
          </string-name>
          ,
          <article-title>A novel dwt-based encoder for human</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          <article-title>A multithread nested neural network architecture pose estimation</article-title>
          , volume
          <volume>3360</volume>
          ,
          <year>2022</year>
          , pp.
          <fpage>33</fpage>
          -
          <lpage>40</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          <article-title>to model surface plasmon polaritons</article-title>
          propagation, [17]
          <string-name>
            <given-names>M.</given-names>
            <surname>Dewan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Murshed</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Lin</surname>
          </string-name>
          , Engagement detec-
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          <issue>Micromachines 7</issue>
          (
          <year>2016</year>
          ). doi:
          <volume>10</volume>
          .3390/mi7070110.
          <article-title>tion in online learning: a review</article-title>
          ,
          <source>Smart Learning</source>
          [5]
          <string-name>
            <given-names>C.</given-names>
            <surname>Napoli</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Pappalardo</surname>
          </string-name>
          , E. Tramontana, R. K. Now- Environments 6 (
          <year>2019</year>
          )
          <fpage>1</fpage>
          -
          <lpage>20</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          <string-name>
            <surname>icki</surname>
            , J. T. Starczewski,
            <given-names>M.</given-names>
          </string-name>
          <string-name>
            <surname>Woźniak</surname>
            , Toward work [18]
            <given-names>M.</given-names>
          </string-name>
          <string-name>
            <surname>Murshed</surname>
            ,
            <given-names>M. A. A.</given-names>
          </string-name>
          <string-name>
            <surname>Dewan</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          <string-name>
            <surname>Lin</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          <string-name>
            <surname>Wen</surname>
          </string-name>
          , En-
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          <source>network approach</source>
          , volume
          <volume>9119</volume>
          ,
          <year>2015</year>
          , pp.
          <fpage>79</fpage>
          -
          <lpage>89</lpage>
          .
          <article-title>ing convolutional neural networks</article-title>
          , in: 2019 IEEE
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          <source>doi:10.1007/978- 3- 319- 19324- 3_8. Intl Conf on Dependable, Autonomic and Secure</source>
          [6]
          <string-name>
            <given-names>N.</given-names>
            <surname>Brandizzi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Russo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Brociek</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Wajda</surname>
          </string-name>
          , First Computing,
          <source>Intl Conf on Pervasive Intelligence and</source>
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          <string-name>
            <surname>ing</surname>
          </string-name>
          , volume
          <volume>3118</volume>
          ,
          <year>2021</year>
          , pp.
          <fpage>71</fpage>
          -
          <lpage>76</lpage>
          . Congress (DASC/PiCom/CBDCom/CyberSciTech), [7]
          <string-name>
            <given-names>A.</given-names>
            <surname>Gupta</surname>
          </string-name>
          ,
          <string-name>
            <surname>A. D'Cunha</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          <string-name>
            <surname>Awasthi</surname>
          </string-name>
          , V. Balasubrama- IEEE,
          <year>2019</year>
          , pp.
          <fpage>80</fpage>
          -
          <lpage>86</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          <string-name>
            <surname>nian</surname>
            , Daisee: Towards user engagement recogni- [19]
            <given-names>J. T.</given-names>
          </string-name>
          <string-name>
            <surname>Springenberg</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Dosovitskiy</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          <string-name>
            <surname>Brox</surname>
          </string-name>
          , M. Ried-
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          <article-title>tion in the wild</article-title>
          ,
          <source>arXiv preprint arXiv:1609</source>
          .
          <year>01885</year>
          miller,
          <article-title>Striving for simplicity: The all convolutional</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          (
          <year>2016</year>
          ). net,
          <source>arXiv preprint arXiv:1412.6806</source>
          (
          <year>2014</year>
          ). [8]
          <string-name>
            <given-names>Z.</given-names>
            <surname>Wan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>He</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Voisine</surname>
          </string-name>
          , An attention level moni- [20]
          <string-name>
            <given-names>M.</given-names>
            <surname>Lin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Q.</given-names>
            <surname>Chen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Yan</surname>
          </string-name>
          , Network in network, arXiv
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          <article-title>toring and alarming system for the driver fatigue</article-title>
          preprint arXiv:
          <volume>1312</volume>
          .4400 (
          <year>2013</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          <article-title>in the pervasive environment</article-title>
          , in: Brain and Health [21]
          <string-name>
            <given-names>K.</given-names>
            <surname>Simonyan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Zisserman</surname>
          </string-name>
          , Very deep convolu-
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          <string-name>
            <surname>Informatics</surname>
            : International Conference,
            <given-names>BHI</given-names>
          </string-name>
          <year>2013</year>
          ,
          <article-title>tional networks for large-scale image recognition,</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          <string-name>
            <surname>Maebashi</surname>
          </string-name>
          , Japan,
          <source>October 29-31</source>
          ,
          <year>2013</year>
          . Proceedings, arXiv preprint arXiv:
          <volume>1409</volume>
          .1556 (
          <year>2014</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref25">
        <mixed-citation>
          <string-name>
            <surname>Springer</surname>
          </string-name>
          ,
          <year>2013</year>
          , pp.
          <fpage>287</fpage>
          -
          <lpage>296</lpage>
          . [22]
          <string-name>
            <given-names>S.</given-names>
            <surname>Albawi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T. A.</given-names>
            <surname>Mohammed</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Al-Zawi</surname>
          </string-name>
          ,
          <year>Under</year>
          [9]
          <string-name>
            <given-names>V.</given-names>
            <surname>Ponzi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Russo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Wajda</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Brociek</surname>
          </string-name>
          ,
          <string-name>
            <surname>C.</surname>
          </string-name>
          <article-title>Napoli, standing of a convolutional neural network</article-title>
          , in:
        </mixed-citation>
      </ref>
      <ref id="ref26">
        <mixed-citation>
          <article-title>Analysis pre and post covid-19 pandemic rorschach 2017 international conference on engineering and</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref27">
        <mixed-citation>
          <article-title>test data of using em algorithms and gmm models</article-title>
          ,
          <source>technology (ICET)</source>
          , Ieee,
          <year>2017</year>
          , pp.
          <fpage>1</fpage>
          -
          <lpage>6</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref28">
        <mixed-citation>
          volume
          <volume>3360</volume>
          ,
          <year>2022</year>
          , pp.
          <fpage>55</fpage>
          -
          <lpage>63</lpage>
          . [23]
          <string-name>
            <given-names>J.</given-names>
            <surname>Whitehill</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Serpell</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.-C.</given-names>
            <surname>Lin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Foster</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. R.</given-names>
            [10]
            <surname>C.-M. Chen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.-Y.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <surname>C.-M. Yu</surname>
          </string-name>
          ,
          <article-title>Assessing the Movellan, The faces of engagement: Automatic</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref29">
        <mixed-citation>
          <source>British Journal of Educational Technology</source>
          <volume>48</volume>
          (
          <year>2017</year>
          )
          <article-title>ing 5 (</article-title>
          <year>2014</year>
          )
          <fpage>86</fpage>
          -
          <lpage>98</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref30">
        <mixed-citation>
          348-
          <fpage>369</fpage>
          . [24]
          <string-name>
            <given-names>G.</given-names>
            <surname>Littlewort</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Whitehill</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Wu</surname>
          </string-name>
          , I. Fasel, M. Frank, [11]
          <string-name>
            <given-names>S.</given-names>
            <surname>Di Palma</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Tonacci</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Narzisi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Domenici</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Movellan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Bartlett</surname>
          </string-name>
          , Computer expression
        </mixed-citation>
      </ref>
      <ref id="ref31">
        <mixed-citation>
          et al.,
          <article-title>Monitoring of autonomic response to so- ture Recognition (FG'11) 20 (</article-title>
          <year>2011</year>
          )
          <fpage>24</fpage>
          -
          <lpage>25</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref32">
        <mixed-citation>
          <article-title>ciocognitive tasks during treatment in children</article-title>
          with [25]
          <string-name>
            <given-names>P.</given-names>
            <surname>Viola</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Jones</surname>
          </string-name>
          , et al.,
          <string-name>
            <surname>Robust</surname>
          </string-name>
          real-time object
        </mixed-citation>
      </ref>
      <ref id="ref33">
        <mixed-citation>
          <article-title>gies: A feasibility study</article-title>
          ,
          <source>Computers in biology and 4</source>
          (
          <year>2001</year>
          )
          <article-title>4</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref34">
        <mixed-citation>
          <source>medicine 85</source>
          (
          <year>2017</year>
          )
          <fpage>143</fpage>
          -
          <lpage>152</lpage>
          . [26]
          <string-name>
            <given-names>F.</given-names>
            <surname>Del Duchetto</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Baxter</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Hanheide</surname>
          </string-name>
          , Are you [12]
          <string-name>
            <given-names>O.</given-names>
            <surname>Dehzangi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Williams</surname>
          </string-name>
          ,
          <article-title>Towards multi-modal still with me? continuous engagement assessment</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref35">
        <mixed-citation>
          <source>and AI 7</source>
          (
          <year>2020</year>
          )
          <fpage>116</fpage>
          . [27]
          <string-name>
            <given-names>S.</given-names>
            <surname>Hochreiter</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Schmidhuber</surname>
          </string-name>
          ,
          <string-name>
            <surname>Long</surname>
          </string-name>
          short-term
        </mixed-citation>
      </ref>
      <ref id="ref36">
        <mixed-citation>
          <string-name>
            <surname>memory</surname>
          </string-name>
          ,
          <source>Neural computation 9</source>
          (
          <year>1997</year>
          )
          <fpage>1735</fpage>
          -
          <lpage>1780</lpage>
          . [28]
          <string-name>
            <given-names>J.</given-names>
            <surname>Liao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Liang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Pan</surname>
          </string-name>
          , Deep facial spatiotempo-
        </mixed-citation>
      </ref>
      <ref id="ref37">
        <mixed-citation>
          <source>learning, Applied Intelligence</source>
          <volume>51</volume>
          (
          <year>2021</year>
          )
          <fpage>6609</fpage>
          -
          <lpage>6621</lpage>
          . [29]
          <string-name>
            <given-names>S.</given-names>
            <surname>Pepe</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Tedeschi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Brandizzi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Russo</surname>
          </string-name>
          , L. Iocchi,
        </mixed-citation>
      </ref>
      <ref id="ref38">
        <mixed-citation>
          <string-name>
            <surname>dataset</surname>
          </string-name>
          ,
          <source>OBM Neurobiology 6</source>
          (
          <year>2022</year>
          )
          <fpage>1</fpage>
          -
          <lpage>10</lpage>
          . [30]
          <string-name>
            <given-names>S. J.</given-names>
            <surname>Pan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Q.</given-names>
            <surname>Yang</surname>
          </string-name>
          ,
          <article-title>A survey on transfer learning,</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref39">
        <mixed-citation>
          <source>ing 22</source>
          (
          <year>2009</year>
          )
          <fpage>1345</fpage>
          -
          <lpage>1359</lpage>
          . [31]
          <string-name>
            <given-names>Z.</given-names>
            <surname>Hu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Lv</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Hang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Huang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Xing</surname>
          </string-name>
          ,
        </mixed-citation>
      </ref>
      <ref id="ref40">
        <mixed-citation>
          <source>Transactions on Industrial Electronics</source>
          <volume>69</volume>
          (
          <year>2021</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref41">
        <mixed-citation>
          1800-
          <fpage>1808</fpage>
          . [32]
          <string-name>
            <given-names>A.</given-names>
            <surname>Palazzi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Abati</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Solera</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Cucchiara</surname>
          </string-name>
          , et al.,
        </mixed-citation>
      </ref>
      <ref id="ref42">
        <mixed-citation>
          <source>ysis and machine intelligence</source>
          <volume>41</volume>
          (
          <year>2018</year>
          )
          <fpage>1720</fpage>
          -
          <lpage>1733</lpage>
          . [33]
          <string-name>
            <given-names>E.</given-names>
            <surname>Cengil</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Çınar</surname>
          </string-name>
          , E. Özbay, Image classification
        </mixed-citation>
      </ref>
      <ref id="ref43">
        <mixed-citation>
          <article-title>with cafe deep learning framework</article-title>
          , in: 2017 In-
        </mixed-citation>
      </ref>
      <ref id="ref44">
        <mixed-citation>
          <string-name>
            <surname>Engineering</surname>
          </string-name>
          (UBMK), IEEE,
          <year>2017</year>
          , pp.
          <fpage>440</fpage>
          -
          <lpage>444</lpage>
          . [34]
          <string-name>
            <given-names>I.</given-names>
            <surname>Culjak</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Abram</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Pribanic</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Dzapo</surname>
          </string-name>
          , M. Cifrek,
        </mixed-citation>
      </ref>
      <ref id="ref45">
        <mixed-citation>
          <article-title>A brief introduction to opencv</article-title>
          , in: 2012 proceedings
        </mixed-citation>
      </ref>
      <ref id="ref46">
        <mixed-citation>
          <article-title>of the 35th international convention MIPRO</article-title>
          , IEEE,
        </mixed-citation>
      </ref>
      <ref id="ref47">
        <mixed-citation>
          <year>2012</year>
          , pp.
          <fpage>1725</fpage>
          -
          <lpage>1730</lpage>
          . [35]
          <string-name>
            <given-names>K.</given-names>
            <surname>He</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Zhang</surname>
          </string-name>
          , S. Ren,
          <string-name>
            <given-names>J.</given-names>
            <surname>Sun</surname>
          </string-name>
          , Deep residual learn-
        </mixed-citation>
      </ref>
      <ref id="ref48">
        <mixed-citation>
          <string-name>
            <surname>recognition</surname>
          </string-name>
          ,
          <year>2016</year>
          , pp.
          <fpage>770</fpage>
          -
          <lpage>778</lpage>
          . [36]
          <string-name>
            <given-names>K.</given-names>
            <surname>He</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Zhang</surname>
          </string-name>
          , S. Ren,
          <string-name>
            <given-names>J.</given-names>
            <surname>Sun</surname>
          </string-name>
          , Delving deep into
        </mixed-citation>
      </ref>
      <ref id="ref49">
        <mixed-citation>
          <source>international conference on computer vision</source>
          ,
          <year>2015</year>
          ,
        </mixed-citation>
      </ref>
      <ref id="ref50">
        <mixed-citation>
          pp.
          <fpage>1026</fpage>
          -
          <lpage>1034</lpage>
          . [37]
          <string-name>
            <surname>T.-Y. Lin</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          <string-name>
            <surname>Goyal</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          <string-name>
            <surname>Girshick</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          <string-name>
            <surname>He</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          <string-name>
            <surname>Dollár</surname>
          </string-name>
          , Fo-
        </mixed-citation>
      </ref>
      <ref id="ref51">
        <mixed-citation>
          <string-name>
            <surname>vision</surname>
          </string-name>
          ,
          <year>2017</year>
          , pp.
          <fpage>2980</fpage>
          -
          <lpage>2988</lpage>
          . [38]
          <string-name>
            <given-names>L. N.</given-names>
            <surname>Smith</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Topin</surname>
          </string-name>
          , Super-convergence: Very
        </mixed-citation>
      </ref>
      <ref id="ref52">
        <mixed-citation>
          <source>volume 11006, SPIE</source>
          ,
          <year>2019</year>
          , pp.
          <fpage>369</fpage>
          -
          <lpage>386</lpage>
          . [39]
          <string-name>
            <given-names>D. E.</given-names>
            <surname>King</surname>
          </string-name>
          ,
          <article-title>Dlib-ml: A machine learning toolkit,</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref53">
        <mixed-citation>
          <source>The Journal of Machine Learning Research</source>
          <volume>10</volume>
          (
          <year>2009</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref54">
        <mixed-citation>
          1755-
          <fpage>1758</fpage>
          . [40]
          <string-name>
            <given-names>N.</given-names>
            <surname>Otsu</surname>
          </string-name>
          ,
          <article-title>A threshold selection method from gray-</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref55">
        <mixed-citation>
          <source>man, and cybernetics 9</source>
          (
          <year>1979</year>
          )
          <fpage>62</fpage>
          -
          <lpage>66</lpage>
          . [41]
          <string-name>
            <given-names>A.</given-names>
            <surname>Dosovitskiy</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Beyer</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Kolesnikov</surname>
          </string-name>
          , D. Weis-
        </mixed-citation>
      </ref>
      <ref id="ref56">
        <mixed-citation>
          <article-title>worth 16x16 words: Transformers for image recog-</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref57">
        <mixed-citation>
          nition at scale, arXiv preprint arXiv:
          <year>2010</year>
          .11929
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>