<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Keeping Eyes on the Road: Understanding Driver Attention and Its Role in Safe Driving</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Francesca Fiani</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Valerio Ponzi</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Samuele Russo</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Department of Computer, Control and Management Engineering, Sapienza University of Rome</institution>
          ,
          <addr-line>00185 Roma</addr-line>
          ,
          <country country="IT">Italy</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Department of Psychology, Sapienza University of Rome</institution>
          ,
          <addr-line>00185 Roma</addr-line>
          ,
          <country country="IT">Italy</country>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>Institute for Systems Analysis and Computer Science, Italian National Research Council</institution>
          ,
          <addr-line>00185 Roma</addr-line>
          ,
          <country country="IT">Italy</country>
        </aff>
      </contrib-group>
      <fpage>85</fpage>
      <lpage>95</lpage>
      <abstract>
        <p>Monitoring the driver's attention is an important task to maintain the driver's safety. The estimation of the driver's gaze direction can help us to evaluate if the drivers are not focusing their attention on the street. For an evaluation of this type, comparing the inside view and outside scenery of the vehicle is essential, therefore we decided to create a specific dataset for this task. In this work, we realize a machine-learning-oriented approach to driver's attention evaluation using a coupled visual perception system. By analyzing the road and the driver's gaze simultaneously it is possible to understand if the driver is looking at the trafic signs detected. We evaluate if a determined Region Of Interest (ROI) contains a road sign or not through YOLOv8.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;Visual Attention Estimation</kwd>
        <kwd>Machine Learning</kwd>
        <kwd>Artificial Intelligence</kwd>
        <kwd>ADAS (Autonomous Driver Assistance Systems)</kwd>
        <kwd>YOLO</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>
        1. Introduction
the analysis of the vehicle cabin and the driver’s gaze is
conducted independently, without considering the
evaluArtificial Intelligence (AI) employed in assessing driver ation of the surrounding environment, road conditions,
attention within assisted driving scenarios is swiftly ad- and the driver’s reaction to specific events.
vancing, propelled by the evolution of autonomous ve- Several studies focus either on observing the driver’s
hicles and the integration of hybrid systems designed to behavior through internal vehicle cameras or analyzing
assist drivers. These systems encompass a range of func- external road conditions using external cameras and
sentionalities, including cruise control, lane-keeping assis- sors [
        <xref ref-type="bibr" rid="ref10 ref6 ref7 ref8 ref9">6, 7, 8, 9, 10</xref>
        ]. However, a gap exists in
comprehentance, automatic parking, and various other features inte- sive research that integrates both internal and external
grated into modern vehicles. It is well known that driver perspectives without relying on complex and
inaccessiinattention is a major cause of road accidents [
        <xref ref-type="bibr" rid="ref1 ref2 ref3">1, 2, 3</xref>
        ], ble equipment. To address this gap, our research adopts
with violations of the expected driver behavior being a a novel approach. We simultaneously analyze internal
fundamental factor [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ]. Due to its significant contribu- driver information, such as posture and gaze, and
extertion to accidents, monitoring driver attention has become nal data about road conditions and points of interest, like
a critical necessity for automotive safety systems, aiming signs and pedestrians, during driving. This integrated
apto detect potential risks and proactively prevent accidents. proach allows for a more holistic understanding of driver
To achieve comprehensive attention monitoring, it is im- attention and behavior.
perative to conduct precise analyses of various factors, Machine learning is playing a pivotal role in creating
including the driver’s posture, head position, rotation a safer society. In the realm of energy [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ], machine
angles, and gaze direction. These insights into driver learning algorithms are optimizing data systems [
        <xref ref-type="bibr" rid="ref12 ref13">12, 13</xref>
        ],
behavior enable the identification of factors influencing improving supply-demand forecasting, and enhancing
reactions to diferent conditions and scenarios, thereby the eficiency of renewable energy sources. This not only
mitigating distractions and drowsiness-related incidents ensures a stable energy supply but also reduces the risk
in the future [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ]. of blackouts. When it comes to fostering a green
environ
      </p>
      <p>Literature primarily addresses driver attention by di- ment, machine learning is at the forefront of monitoring
viding the internal and external components. Typically, and predicting environmental changes, enabling us to
take timely action against potential threats [14, 15].
SoSYSYEM 2023: 9th Scholar’s Yearly Symposium of Technology, Engi- cial benefits are manifold, including improved healthcare
neering and Mathematics, Rome, December 3-6, 2023 through predictive diagnostics, personalized education,
$ fiani@diag.uniroma1.it (F. Fiani); ponzi@diag.uniroma1.it and efective public services, all contributing to an
im(V. 0P0o0n9z-i0)0;0sa5m-0u39e6le-.7r0u1s9so(@F. uFniairnoim); a010.0i9t-(0S0.0R0u-2ss9o1)0-0273 (V. Ponzi); proved quality of life [16, 17, 18]. In the context of urban
0000-0002-9421-8566 (S. Russo) driving, machine learning is the driving force behind
© 2023 Copyright for this paper by its authors. Use permitted under Creative Commons License autonomous vehicles [19]. These vehicles promise to
sigCPWrEooUrckReshdoinpgs IhStpN:/c1e6u1r3-w-0s.o7r3g ACttEribUutRion W4.0oInrtekrnsahtioonpal (PCCroBYce4.0e).dings (CEUR-WS.org)
nificantly reduce trafic accidents, improve trafic flow, and point of focus. Our novel approach involves the use
and reduce carbon emissions, making our cities safer of a grid of nine cells to predict the Regions Of
Interand more sustainable. Thus, machine learning is a key est (ROIs) of the driver’s gaze, as illustrated in Figure 1.
enabler in our pursuit of a safer society. To achieve this, we employ a VGG16 network to extract</p>
      <p>In this research, we merge various internal and exter- features from facial video frames, augmenting this
innal techniques for gaze recognition and correlate them formation with head-pose data (i.e. roll, pitch, and yaw
with external Regions Of Interest (ROIs) to develop an angles) to enhance gaze-position prediction [20, 21, 22].
easily applicable solution that comprehensively tackles The diference between tracking gaze-position when a
the issue of driver attention. This approach holds sig- person is looking at a monitor and while they are driving
nificant practical implications for everyday scenarios, is, in fact, substantial [21, 22]. When looking at a monitor,
including: head movements are imperceptible, so the only
discriminant is the position of the pupil. During driving, however,
• Autonomous vehicle development: Understand- the driver tends to rotate their head to look at vehicles
ing the driver’s focus during critical driving situ- and pedestrians or tilt it to see street names, signs, or
ations, including the duration of their attention higher trafic lights. They also shift their gaze to look at
to specific elements and their perception of ir- mirrors or to initiate a reverse maneuver. For these
rearelevant factors, plays a pivotal role in the ad- sons, analyzing only pupil movement was insuficient for
vancement of Advanced Driver Assistance Sys- our task and it was necessary to have additional
informatem (ADAS) solutions. tion about head pose (rotation angles) and characteristics
• Car crashes: Having information about driver of eyes or facial images, see Figure 2.</p>
      <p>attention during a road accident could facilitate In addition to methodological research, another
sigthe execution of investigations, checks, and in- nificant challenge we faced was sourcing an
approprisurance procedures. By utilizing an afordable ate dataset for driver attention monitoring. We
encouncamera system, video data on the driver involved tered existing datasets with comprehensive
documentain the accident could be collected and provided tion of driver behavior, but they lacked corresponding
to an application. real-world external observations. Furthermore, datasets
• Emergency services: Emergency response vehi- focused solely on gaze analysis typically consisted of
cles, including ambulances and fire trucks, of- images of individuals looking at points on a computer
ten need to navigate through trafic quickly and screen, which did not align with our real-world driving
safely. Driver attention monitoring systems can scenario. To address this gap, we decided to create our
help emergency service providers ensure their own dataset, encompassing both internal and external
drivers remain vigilant while responding to emer- videos captured during driving sessions. This approach
gencies, minimizing the likelihood of accidents enabled our final application to process and correlate
and delays.
• Public transportation infrastructure: Driver
attention monitoring systems can also be integrated
into public transportation infrastructure, such as
trafic lights and pedestrian crossings. By
detecting instances of driver distraction or inattention,
these systems can improve trafic flow and
pedestrian safety, reducing the risk of accidents and
congestion in urban areas.</p>
      <p>To advance driver attention monitoring, we have
directed our eforts towards computer vision-based
methodologies, which are gaining traction over
physiologybased approaches. Unlike physiological methods, in fact,
vision-based techniques rely solely on cameras to observe
and analyze driver behaviors, eliminating the need for
intrusive devices such as eye-tracking glasses or
brainwave recognition gadgets and consequently reducing the
cost associated with experiments.</p>
      <p>The most in-depth analysis in our work focused on
ifnding the best method and features to extract from
images to accurately determine the driver’s gaze direction
information from multiple perspectives simultaneously. most classical metric being the gaze direction, generally</p>
      <p>
        To train the two components of our application, we uti- assessed by analyzing facial features such as the face
lized two additional datasets. For the internal component, mesh. Other approaches are however available, for
inwhich involves predicting the driver’s gaze position, we stance, the position of the hands and arms, which can be
curated the HEAD-POSE dataset, featuring data from dif- used to assess whether the driver keeps their hands on
ferent subjects. Unlike many existing datasets that often the steering wheel or in other positions, such as holding
focus on a single subject, our dataset ofers a broader a phone [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ].
and more diverse range of observations. For the external On the other hand, other approaches solely focus on
excomponent, which entails predicting the position of road ternal factors by studying the surrounding environment
signs, we leveraged a customized dataset of road signs and collecting information about the vehicle’s movement
sourced from the internet. We carefully selected images (speed, position) to study the driver’s reactivity in
spefrom datasets such as MAPILLARY and GTSDB, ensuring cific circumstances. For example, various sensors such
that they adhered to European trafic regulations gov- as cameras and lidar, applied to the external part of the
erned by the Vienna Convention of 1968. This meticulous vehicle, can allow the observation of the driver’s reaction
curation process ensured the relevance and accuracy of in certain situations [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ]. Another classical study when
the data used in our research and development eforts. analyzing the external environment surrounding the car
is the analysis of road elements present in the scene via
neural networks such as YOLO [
        <xref ref-type="bibr" rid="ref14">25, 26</xref>
        ].
2. Related Works While poor in number compared to decoupled
approaches, some studies simultaneously analyze both
inAs previously discussed, in recent years a growing in- ternal and external images of the vehicle while assessing
terest in analyzing driver attention during driving has driver attention to the road from the driver’s
perspecbeen noticed. This includes understanding whether a tive. In cases where interior cabin images are
associperson is observing the road, being distracted, remaining ated with external frames, the driver’s viewpoint is often
vigilant, or experiencing drowsiness. Most state-of-the- recorded using glasses or equipment that track eye
moveart approaches are based on the unique observation of ments, which directly indicates what is being observed
the driver’s interior cabin to understand their behaviors [
        <xref ref-type="bibr" rid="ref15">27</xref>
        ]. There are also some recent datasets created in a
[
        <xref ref-type="bibr" rid="ref2">23, 24, 2, 20</xref>
        ]. One or more internal cameras are used controlled setting that simulate common driving
situato observe the driver and determine if they are looking tions, such as the DGAZE dataset with its corresponding
at the infotainment system, the road, the mirrors, or, for algorithm I-DGAZE [
        <xref ref-type="bibr" rid="ref16">28</xref>
        ].
example, other passengers. Several methods can be used Regarding specifically the gaze detection task, various
to determine the attention level of the driver, with the approaches are used in simulated or real environments,
both indoors and outdoors [
        <xref ref-type="bibr" rid="ref17">29</xref>
        ]. In most literature works
and datasets, recordings are made using a personal
computer’s webcam while the subject looks at specific points
on the screen for certain moments. With this regression
problem, the aim is to recognize the precise gaze position
on the monitor by studying the direction of the pupils
and gaze triangulation [
        <xref ref-type="bibr" rid="ref18 ref19 ref20 ref21 ref22">30, 22, 31, 32, 33, 34</xref>
        ]. These types
of problems can, however, also be approached through
classification. For example, images can be taken of a
stationary subject in front of a personal computer screen,
ideally divided into a 9-cell grid, and the gaze position
can then be returned not as precise point coordinates,
but instead as the ID number of the observed cell
(classification) [ 21]. In such cases, pupil characteristics are
generally extracted and then classified using classic
machine learning methods such as Support Vector Machines
(SVMs), Convolutional Neural Networks (CNNs), or Deep
Figure 2: Example of internal image with facial features ex- Neural Networks (DNNs). This last approach in
particutraction. The red rectangle is the measured face bounding lar has inspired our choice to implement a classification
box and the blue dots represent the found facial landmarks. algorithm, given the problems and requirements already
The face of the subject has been blurred according to privacy described and specific to the field of driving.
regulations. Other existing algorithms for studying gaze position
start, as previously mentioned, from datasets of tens
• Trafic Objects Dataset: This dataset is a
modiifed version of the Mapillary Dataset, containing
images depicting various trafic signs.
• Trafic Signs Dataset in YOLO format (TSDY):
      </p>
      <p>This dataset comprises images sourced from
the German Trafic Sign Detection Benchmark
(GTSDB), available for download from Kaggle.
of thousands of photos collected using a personal
computer’s webcam, and then extract facial information from
the given images to crop eye images and pass them to
networks such as VGG-16 [22]. In addition, information
related to head position (rotation angles - roll, pitch, yaw)
can also be considered [20]. It is particularly of note that
to improve regression on the viewpoint position it is
fundamental to collect images from multiple subjects, in
multiple vehicles, and under diferent weather conditions.</p>
      <p>
        Finally, additional approaches make use of recordings
in simulated environments using various technologies,
from simulators to simple computer-played videos. For
example, the user’s gaze position can be recorded while
watching driving videos shortly before certain incidents,
in order to understand which objects the driver
(simulated in this case) would have focused on [
        <xref ref-type="bibr" rid="ref17">29</xref>
        ].
      </p>
    </sec>
    <sec id="sec-2">
      <title>To create our dataset, we maintained a consistent</title>
      <p>equipment setup as depicted in Figure 3. For both the
Gaze Directions and Head Posture Dataset and People
Driving Dataset we utilized a city car, while an iPhone
15 camera was employed to capture internal images and
record internal videos. The iPhone was strategically
positioned behind the steering wheel to ensure clear visibility
of the driver while minimizing extraneous details.
Furthermore, we positioned a GoPro Hero10 camera at the
center of the car’s dashboard to capture external footage
3. Methods throughout the drive.</p>
      <p>For the GDHPD, we compiled images from distinct
This research explores an innovative method for recogniz- subjects, consisting of two males and two females. In
ing gaze patterns while driving to evaluate driver atten- some cases, subjects wore glasses, while in the other
tion. Subsequently, we focused on the internal aspect of they did not. The image collection process encompassed
the vehicle, where we trained and tested neural networks various times of the day and diverse lighting conditions,
for gaze classification. Our experimentation involved resulting in a total of 1012 images. Table 1 provides a
various models, including SVM, ClassNET, VGG16-based breakdown of the distribution of these images.
Net, and HEGClass Net. Additionally, we conducted a Subjects were positioned inside a car, with their
seattraining phase for the external aspect using a custom ing adjusted to achieve a standard driving posture.
Subdataset comprising trafic sign objects. Once we obtained sequently, images were captured while subjects varied
results for both components, we merged the two mod- their gaze and head positions. To facilitate
classificaules to conduct a comprehensive analysis of video record- tion, we devised a virtual grid dividing the external view
ings obtained during real-world driving scenarios. This and the driver’s gaze into a 9-cell configuration. This
integrated approach facilitated a more holistic
comprehension of gaze behavior and its correlation with driver
attention in typical driving situations.
3.1. Dataset
A variety of images and videos were gathered and
utilized at diferent stages of development to construct
custom datasets tailored to our research objectives. These
datasets can be categorized into four distinct collections:
• Gaze Directions and Head Posture Dataset
(GDHPD): This dataset comprises images
captured by us, featuring individuals in a driving
environment. The images are utilized to
categorize the gaze position of individuals within a grid
consisting of nine cells, including the exterior of
the grid.
• People Driving Dataset (PDD): This dataset
consists of both external and internal videos,
recorded by our team, showcasing the driving
activities.</p>
      <p>GDHPD dataset specifications
Classes Male Female Total
0 50 64 114
1 50 64 114
2 51 71 122
3 51 71 122
4 48 60 108
5 48 60 108
6 50 50 100
7 51 50 101
8 51 61 112
9 50 61 111
Total 500 612 1012
grid facilitated the association of head and eye positions
with specific regions of the external images, enabling
the identification of Regions Of Interest (ROIs) during
experimentation. 3.2. Gaze Classification</p>
      <p>The dataset comprises ten classes: the nine cells of the
grid, alongside an additional class representing situations Several algorithms, were analyzed to identify the
apwhere the subject’s attention is not directed towards the proach with the best trade-of between accuracy in gaze
road (e.g., face turned sideways, gaze directed upwards direction prediction and generalization capabilities,
alor downwards, etc.). lowing to eficiently recognize images with varied
con</p>
      <p>The PDD dataset comprises videos captured during trast and/or brightness or diferent drivers. The structure
driving sessions, utilizing the same recording setup as takes as input an image (either singular or an extracted
the previous dataset. Subjects were filmed while driv- frame from the recording) and generates a label
predicing under various conditions, capturing both the driver tion through two diferent but subsequent sub-models.
and the road view simultaneously. To ensure synchro- The head-pose estimation algorithm is common for all
nization, the videos underwent pre-processing using a tested approaches, while the classification algorithm has
third-party software, DaVinci Resolve. Synchronization been severely varied in model type, structure, and input
was achieved through voice cues, guaranteeing precise during the search for the optimal solution.
alignment between internal and external footage.</p>
      <p>
        Subsequently, the videos were segmented into sub- 3.2.1. Head-Pose Estimation Part
clips of 30 seconds each to streamline subsequent
processing steps. Each of the extracted sub-clips was anno- The face detection is performed through a pre-trained
tated with labels indicating whether the driver exhibited Multi-task Cascaded Convolutional Network (MTCNN)
a "CAREFUL" or "NOT CAREFUL" driving style, along model, used for both face detection and alignment in
with information regarding the driver’s use of glasses. literature [
        <xref ref-type="bibr" rid="ref23">35</xref>
        ]. MTCNN consists of a cascade of
convoAfter pre-processing, the external images displayed a lutional networks (P-Net, R-Net, and O-Net), for face
resolution of 1440x1080, while the internal images were landmarks identification. The model first identifies the
resized to 1080x1920. bounding box of the face region through candidate
gen
      </p>
      <p>For trafic sign detection and recognition, the primary eration with P-Net and refinement with R-Net and then
dataset utilized was the Trafic Object dataset from the extracts the main 5 landmarks of the face (left and right
Mapillary Trafic Sign Dataset, encompassing tens of eye, nose tip, and mouth corners) with O-Net. Among
thousands of images sourced from roads worldwide. Fo- similar methods, such as Haar Cascade Classifiers [ 36],
cusing solely on Italian/European trafic signs, around MTCNN has shown the best results even in the presence
3,000 images were selected from the dataset after filter- of glasses, partially occluded eyes and beards, and has
ing out images with significantly diferent sign shapes therefore been selected as our chosen method.
or contents. The chosen images ofer a varied range of The identified landmarks are then used to analytically
brightness, positioning within frames, and contextual calculate the roll, pitch and yaw angles of the driver’s
head, while the extracted pupil positions will be used network. This Classifier Network architecture consists of
as input features for the final classifier to determine the two convolutional layers with ReLU activation functions,
observed ROI. This feature is fundamental for our clas- followed by a max pooling layer, and culminates in two
sification task compared to other facial features and is fully connected layers with ReLU and Softmax activation
therefore particularly important to determine accurately. functions. Training of ClassNet spanned 300 epochs,
employing the MSE loss function (Mean Squared Error)
3.2.2. Classification Part and Adam optimizer.</p>
      <p>Finally, the last experimental iteration involved
emOur approach involves classifying observed ROI and road ploying the VGG-16 architecture. Training data consisted
signs through a classification method. The field of view of 1012 samples from the GDHPD dataset, each composed
is divided into nine sections, with an additional label of the face image tensor, alongside head rotation angles
for identifying any gaze position outside these sections (roll, pitch, yaw), and pupil centers. In this setup,
fea(e.g., distracted driving or maneuvering). We’ve explored tures from each image were directly extracted within the
various methods for analysis, ranging from traditional classification network from the RGB image of the entire
SVM to CNNs, all accurately adapted for our application. face, rather than solely from the eyes. Additionally, other
We will first introduce our novel method, followed by an features (roll, pitch, yaw, and pupil centers) obtained
preoverview of other models considered in the analysis. viously through MTCNN were incorporated alongside</p>
      <p>The standout classification model is HEGClass (Head- the 512 features of the image in the rfist fully connected
Eyes-Gaze Classifier), a hybrid approach outlined in this layer. Training was executed over 100 epochs, using
simpaper. It takes cropped face images from head-pose esti- ilarly to the HEGClass model the Cross-Entropy Loss
mation, along with head rotation angles and pupil cen- function and Adam optimizer.
ter coordinates, as inputs. This combined approach has
yielded high precision in classifying the Region of
Interest toward which the gaze is directed. In the HEGClass 3.3. YOLO Training for Trafic Signs
network, as depicted in Figure 4, initial features are ex- For the detection and recognition of trafic signs, we start
tracted from cropped face RGB images using a pre-trained from the pre-trained YOLOv8 model, with
experimentaVGG-16 network. The features are then flattened and tion also conducted using the YOLOv5 model prior to
concatenated with a normalized array containing head transitioning to the v8 version. Fine-tuning of YOLOv8
roll-pitch-yaw and pupil center coordinates. This com- was carried out using two distinct datasets: Trafic
Obbined feature vector of dimension 4096+7 passes through jects and TSDY. The final set of weights chosen for the
two fully connected linear layers, followed by ReLU ac- application was derived from the dual fine-tuning of
tivation functions, and finally through a last fully con- YOLOv8 with both datasets.
nected linear layer with Softmax activation to determine The initial fine-tuning with the Trafic Objects dataset
class membership among the 10 possibilities (9 ROIs for involved 1802 images for training and 919 for validation.
frontal regions and 1 for others). Model training utilized Despite starting with 3000 images, adjustments were
our GDHPD dataset, with 10 epochs to mitigate overfit- made to the training and validation sets due to
imbalting, 32 samples per batch, Cross-Entropy loss function, ance issues within the original dataset, which persisted
and Adam optimizer. even after categorizing labels based on sign categories as</p>
      <p>The first classical model used for comparison is Sup- described in the Dataset section. Subsequently,
proceedport Vector Machine (SVM). In our scenario, where we’re ing from the fine-tuned weights, the model underwent
classifying 10 distinct classes, we employed an SVM with retraining with images from the TSDY dataset, utilizing
a polynomial kernel of degree 4, regularization param- 600 images for training and 141 for validation.
eter set at 100, and coeficient set to 10. Training ex- The dual fine-tuning approach resulted in enhanced
clusively used images from the GDHPD dataset. We ex- performance, as evidenced by improved final accuracy
tracted roll, pitch, yaw, and pupil centers from the images and heightened generalization capabilities in detecting
with MTCNN. Then, using Haar Cascade Classifier [ 36], road signs, even in images sourced from the PDD dataset
we isolated eye patches from each grayscale image and and exhibiting varied lighting conditions. To further
enpassed them through a pre-trained ResNet to obtain 2048 rich the diversity and generalization capabilities of the
features for each eye. These two sets of features were ifne-tuned YOLO model, diverse image augmentation
then averaged to create a unified array of 2048 elements techniques from the Albumentations library were
emcontaining information from both eyes. The resulting ployed during training to simulate real-world conditions,
samples underwent L2 normalization before being fed including Blur, MedianBlur, and CLAHE (Contrast
Liminto the SVM for both training and testing phases. ited Adaptive Histogram Equalization). To streamline the</p>
      <p>Using the same 2048 features extracted ResNet and roll, training process, the Stochastic Gradient Descent (SGD)
pitch, yaw and pupil centers, we also trained the ClassNet optimizer was utilized, with an initial learning rate of
0.01. module. The image and information are then passed</p>
      <p>Post-training, the model returns a text file for each through the network to classify the driver’s gaze
posiimage processed, containing one line per detected sign tion. Upon obtaining the prediction of the observed cell
alongside its position. By processing the pixel coordi- in a given frame, it is compared with the position of the
nates from this file, sign position information was ex- corresponding sign. In frames with multiple detected
tracted to reconstruct the sign’s center and edges within road signs, each corresponding ROI is considered active,
grid cells. In instances where signs spanned multiple thus rendering the driver alert when focusing on any of
cells, multiple coordinates were necessary for accurate them without favoring any type of signal.
identification. Given that a single road sign can span multiple cells, 5
points characterize the object’s position: the four corners
3.4. Application: Merging the Methods and the center. If the driver looks at a cell containing a
partial view of the sign, accounting for the peripheral
The final application, depicted in Figure 5, comprises vision of human eyes we consider them attentive.
Moretwo primary components and generates a CSV report over, in the absence of signals or when the driver’s gaze
detailing the overall behavior of a driver. It requires is directed to cells 4/5/6 (representing the entire road
as input an internal video capturing the driver and an surface), they are still deemed attentive to the street. If
external video recording the street view. For ease of gaze is directed to cells 7 or 8, indicating focus on the
analysis, synchronization of a 30-second video between car’s dashboard or infotainment system, the driver’s
enthe two components is necessary. gagement is noted accordingly.</p>
      <p>After inizialization, external frames undergo analysis Finally, a CSV file is generated to store the analysis
using the YOLOv8 model trained on road signs, produc- results from the 30-second videos. Each row contains
ing a text file detailing the detected signals along with data including the frame count (consistent across internal
their specifications, including the Regions of Interest and external videos), the number of road signs detected
(ROIs) where these signals are located. For frames with in that frame, the cell number(s) housing the detected
at least one detected sign, the corresponding internal signals, the predicted cell value from the driver’s gaze
image is used to extract information via the GDHPD network, the number of observed signals following ROI
ifnal classification network. However, occasional failures
in face detection or slight misplacements of landmarks
introduce minor errors in this initial phase.</p>
      <p>For predicting gaze direction, a zone-based
classification approach was chosen over regression due to the
dificulty in precisely determining the exact point on the
road the individual is looking at, coupled with the
human eye’s ability to perceive a broad area. Despite testing
various methods, the SVM-based approach struggled to
TFhigeuinrete5rn:aPliapnedlineextoefrnthalevaidpepoliscaarteiosnimpureltsaennetoeudsilny tphroecpeaspseedr. exceed a 70% accuracy level, likely due to the similarity
to extract relevant features (facial landmarks and road signs in feature values across the 1012 samples, particularly
bounding boxes respectively). The features are then used those derived from ResNet for eye images.
to determine the active ROIs, which are then compared to Transitioning to neural network-based methods, the
generate the attention level of the driver. ClassNet network yielded lower accuracy than SVM, even
after experimenting with diferent feature combinations.</p>
      <p>Training a network based on VGG-16 architecture from
matching, and an indication of the driver’s attentiveness. scratch yielded better results with an accuracy level of
81%. However, the limited size of our dataset and
computational constraints hindered achieving satisfactory
4. Results performance through this approach. Hence, we adopted
the hybrid HEGClass approach, achieving an impressive
The objective of this research is to develop a comprehen- 96% accuracy and 94.3% F1-score without additional data.
sive system capable of analyzing an individual’s atten- Comprehensive accuracy and F1-score results are shown
tion while driving using only two synchronized videos in Table 2.
as input. Given the scarcity of references on the
simultaneous analysis of internal and external perspectives, all 4.2. YOLOv8 Classification and Detection
subsequent evaluations and comparisons will focus on
the individual components constituting the final system.</p>
      <p>Nonetheless, through extensive testing conducted with
the PDD dataset, comprising approximately 194 videos
each lasting 30 seconds, the final application
demonstrates commendable performance.</p>
      <p>Through dual fine-tuning of the YOLOv8 network by
using the Trafic Object Dataset and the TSDY dataset, an
impressive final F1-score of around 95% was achieved,
with an example of prediction on the PDD dataset shown
in Figure 6. We have observed an interesting
phenomenon when training the network solely with the
4.1. Gaze Classification Trafic Objects dataset, where the F1-score is significantly
lower. Specifically, the training of YOLOv8 with the
For what concerns face detection and landmark extrac- Trafic Objects dataset yielded an overall accuracy of
tion for facial rotation angle calculation, the MTCNN approximately 65%-70%. Performing the same process
model outperformed the Haar Cascade Classifier. This with YOLOv5, instead, showed unexpectedly a higher
superiority stems from MTCNN’s ability to handle vari- accuracy (around 80%), albeit with occasional
misclassifious facial orientations, which is crucial for our GDHPD cations of elements such as empty spaces between tree
dataset as it contains images with rotated or profiled faces. branches. In any case, for the scope of this project this
Additionally, MTCNN’s prediction of landmarks, includ- comparison is not particularly relevant, given the much
ing the center of the pupil, proved vital for training the higher accuracy with dual fine-tuning.
dataset. Following initial evaluations, videos
compromised by low light conditions (such as those filmed in
almost night-time environments) or excessive blurring
of frames, rendering accurate prediction unfeasible, were
eliminated.</p>
      <p>In instances where images lack clarity, are blurry, or
exhibit excessive shaking, the predominant predicted class
is 0, indicating the model’s failure to accurately identify
the correct gaze position. Moreover, during nighttime or
low-light scenarios, accurate gaze evaluation is
significantly impeded by diminished brightness. Additionally,
YOLO struggles with precise detection of relevant signs,
often leading to confusion. Classes 5 and 6 are frequently
identified as the gaze position during driving, aligning
with the fact that these areas correspond to central
regions of the windscreen. In some situations, such as
when the vehicle is stationary at a trafic light or in
trafifc congestion, the system may recognize the same signs
across multiple frames. However, drivers may not
consistently attend to them throughout, as they may have
already observed them and they may not be of immediate
significance at that moment.</p>
      <p>Despite the substantial improvement in generalization
capabilities achieved through dual training, errors in sign
recognition persist. Certain objects along the road may
be mistaken for road signs, such as advertisements
containing elements that, with low resolution, could be
confused. While this issue is present, its impact on the overall
results remains manageable and could potentially be mit- 5. Conclusions
igated with a wider variety of images. Another challenge
arises from grouping signs of diferent shapes and colors This work aimed to develop and assess a comprehensive
into the same class, creating a bias in their classification. system for evaluating attention to trafic signs in driving
Additionally, signs containing other signs within them environments. We accomplished this by creating two
may only become relevant in specific situations, such new datasets (GDHPD, PDD) and modifying two
existas parking signs reserved for disabled individuals. For ing ones (Trafic Objects, TSDY) to better suit our task
this reason, some signs were excluded from the training requirements. The final application was divided into two
phase. parts, utilizing YOLOv8 for sign prediction and MTCNN</p>
      <p>Given the high accuracy values in detection, adjusting + HEGClass for gaze position classification.
the confidence threshold can help alleviate misclassifi- Despite encountering challenges during various
traincation issues. Signs may be recognized even when ro- ing and testing phases, as described in the Results section,
tated, facing the opposite direction of the lane, or located the overall accuracy of the final system remains very
in irrelevant areas. In such cases, they are counted as high, notwithstanding the partial errors accumulated by
points of inattention. Due to class imbalance, accurately its constituent parts.
classifying the type of road sign remains a challenge. These challenges serve as valuable insights for future
Consequently, for our purposes, only information related research endeavors. Opportunities for improvement
into the bounding box defining the sign’s position is ex- clude implementing mechanisms to track seen and
untracted, without specifying the type of sign. Despite seen signals, enhancing prediction accuracy in diverse
attempts to simplify the dataset to recognize only one lighting and atmospheric conditions through dataset
augclass, "TRAF_SIGN", challenges persisted between detect- mentation or pre-processing techniques, and expanding
ing signs and identifying unrelated environmental areas. datasets to ensure greater completeness.
Therefore, the decision was made to revert to using the Overall, this work presents significant potential for
furoriginal labels. ther refinement and advancement, promising avenues for
enhancing the performance and robustness of attention
4.3. Overall Analysis evaluation systems in driving contexts.</p>
    </sec>
    <sec id="sec-3">
      <title>The final results of the application, pertaining to the pre</title>
      <p>diction of the driver’s average attention while viewing a
video, exhibit high performance across most cases, with
a few notable exceptions. Tests were conducted on 191
videos, each lasting 30 seconds, sourced from the PDD</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>G.</given-names>
            <surname>Fitch</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Soccolich</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Guo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>McClaferty</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Fang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Olson</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Pérez-Toledano</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Hanowski</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Hankey</surname>
          </string-name>
          , T. Dingus,
          <article-title>The impact of hand-held and</article-title>
          <string-name>
            <surname>hands-</surname>
          </string-name>
          978-3-
          <fpage>031</fpage>
          -18050-7_
          <fpage>38</fpage>
          .
          <article-title>free cell phone use on driving performance and</article-title>
          [14]
          <string-name>
            <given-names>A.</given-names>
            <surname>Alfarano</surname>
          </string-name>
          , G. De Magistris,
          <string-name>
            <given-names>L.</given-names>
            <surname>Mongelli</surname>
          </string-name>
          , S. Russo, safety-critical
          <source>event risk</source>
          ,
          <year>2013</year>
          . J.
          <string-name>
            <surname>Starczewski</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          <string-name>
            <surname>Napoli</surname>
          </string-name>
          ,
          <article-title>A novel convmixer trans-</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>W.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Lu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Zhang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Xie</surname>
          </string-name>
          , W. Zeng,
          <article-title>Driver former based architecture for violent behavior deaction recognition based on attention mechanism, tection 14126 LNAI (</article-title>
          <year>2023</year>
          )
          <fpage>3</fpage>
          -
          <lpage>16</lpage>
          . doi:
          <volume>10</volume>
          .1007/ in: 2019
          <source>6th International Conference on Sys- 978-3-031-42508-0_1. tems and Informatics (ICSAI)</source>
          ,
          <year>2019</year>
          , pp.
          <fpage>1255</fpage>
          -
          <lpage>1259</lpage>
          . [15]
          <string-name>
            <given-names>V.</given-names>
            <surname>Ponzi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Russo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Wajda</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Brociek</surname>
          </string-name>
          , C. Napoli, doi:10.1109/ICSAI48974.
          <year>2019</year>
          .
          <volume>9010589</volume>
          .
          <article-title>Analysis pre and post covid-19 pandemic rorschach</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>A. J.</given-names>
            <surname>McKnight</surname>
          </string-name>
          ,
          <string-name>
            <surname>A. S. McKnight</surname>
          </string-name>
          ,
          <article-title>The efect of cel- test data of using em algorithms and gmm models, lular phone use upon driver attention</article-title>
          ,
          <source>Accident</source>
          volume
          <volume>3360</volume>
          ,
          <year>2022</year>
          , pp.
          <fpage>55</fpage>
          -
          <lpage>63</lpage>
          .
          <source>Analysis &amp; Prevention</source>
          <volume>25</volume>
          (
          <year>1993</year>
          )
          <fpage>259</fpage>
          -
          <lpage>265</lpage>
          . [16]
          <string-name>
            <given-names>S.</given-names>
            <surname>Russo</surname>
          </string-name>
          ,
          <string-name>
            <surname>S. I. Illari</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Avanzato</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Napoli</surname>
          </string-name>
          , Reducing
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <surname>J. C. de Winter</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          <string-name>
            <surname>Dodou</surname>
          </string-name>
          ,
          <article-title>The driver behaviour the psychological burden of isolated oncological questionnaire as a predictor of accidents: A meta- patients by means of decision trees</article-title>
          , volume
          <volume>2768</volume>
          ,
          <article-title>analysis</article-title>
          ,
          <source>Journal of safety research 41</source>
          (
          <year>2010</year>
          )
          <fpage>463</fpage>
          -
          <lpage>2020</lpage>
          , pp.
          <fpage>46</fpage>
          -
          <lpage>53</lpage>
          .
          <fpage>470</fpage>
          . [17]
          <string-name>
            <given-names>E.</given-names>
            <surname>Iacobelli</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Ponzi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Russo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Napoli</surname>
          </string-name>
          , Eye-
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>K. J.</given-names>
            <surname>Anstey</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Wood</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Lord</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. G.</given-names>
            <surname>Walker</surname>
          </string-name>
          ,
          <article-title>Cogni- tracking system with low-end hardware: Develtive, sensory and physical factors enabling driving opment and evaluation, Information (Switzerland) safety in older adults</article-title>
          ,
          <source>Clinical psychology review 14</source>
          (
          <year>2023</year>
          ).
          <source>doi:10.3390/info14120644. 25</source>
          (
          <year>2005</year>
          )
          <fpage>45</fpage>
          -
          <lpage>65</lpage>
          . [18]
          <string-name>
            <given-names>F.</given-names>
            <surname>Fiani</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Russo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Napoli</surname>
          </string-name>
          , An advanced solu-
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>E.</given-names>
            <surname>Yurtsever</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Lambert</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Carballo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Takeda</surname>
          </string-name>
          ,
          <article-title>A tion based on machine learning for remote emdr survey of autonomous driving: Common practices therapy</article-title>
          ,
          <source>Technologies</source>
          <volume>11</volume>
          (
          <year>2023</year>
          ). doi:
          <volume>10</volume>
          .3390/ and emerging technologies,
          <source>IEEE access 8</source>
          (
          <year>2020</year>
          )
          <year>technologies11060172</year>
          .
          <fpage>58443</fpage>
          -
          <lpage>58469</lpage>
          . [19]
          <string-name>
            <given-names>N.</given-names>
            <surname>Brandizzi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Russo</surname>
          </string-name>
          , G. Galati,
          <string-name>
            <given-names>C.</given-names>
            <surname>Napoli</surname>
          </string-name>
          , Address-
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>Y.</given-names>
            <surname>Hou</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Xue</surname>
          </string-name>
          ,
          <string-name>
            <surname>X. L. Zhang,</surname>
          </string-name>
          <article-title>ing vehicle sharing through behavioral analysis: A J</article-title>
          .
          <string-name>
            <surname>Zhu</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          <string-name>
            <surname>Wang</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          <string-name>
            <surname>Chen</surname>
          </string-name>
          ,
          <article-title>Visual evaluation for au- solution to user clustering using recency-frequencytonomous driving</article-title>
          ,
          <source>IEEE Transactions on Visualiza- monetary and vehicle relocation based on neightion and Computer Graphics</source>
          <volume>28</volume>
          (
          <year>2021</year>
          )
          <fpage>1030</fpage>
          -
          <lpage>1039</lpage>
          . borhood splits,
          <source>Information (Switzerland) 13</source>
          (
          <year>2022</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>Y.</given-names>
            <surname>Xia</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Zhang</surname>
          </string-name>
          , J. Kim,
          <string-name>
            <given-names>K.</given-names>
            <surname>Nakayama</surname>
          </string-name>
          , K. Zipser, doi:10.3390/info13110511. D. Whitney, Predicting driver attention in critical [20]
          <string-name>
            <given-names>D.</given-names>
            <surname>Yang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Dai</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Zhang</surname>
          </string-name>
          , L. Qi, W. Zhang, situations, in: Computer Vision-ACCV
          <year>2018</year>
          : 14th
          <string-name>
            <given-names>Z.</given-names>
            <surname>Jiang</surname>
          </string-name>
          ,
          <article-title>All in one network for driver attenAsian Conference on Computer Vision</article-title>
          , Perth, Aus- tion monitoring,
          <source>in: ICASSP 2020 - 2020 IEEE Intralia, December 2-6</source>
          ,
          <year>2018</year>
          , Revised Selected Papers, ternational Conference on Acoustics,
          <source>Speech and Part V 14</source>
          , Springer,
          <year>2019</year>
          , pp.
          <fpage>658</fpage>
          -
          <lpage>674</lpage>
          .
          <source>Signal Processing (ICASSP)</source>
          ,
          <year>2020</year>
          , pp.
          <fpage>2258</fpage>
          -
          <lpage>2262</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>D.</given-names>
            <surname>Yang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Wei</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Guan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Huang</surname>
          </string-name>
          , W. Cai, doi:10.1109/ICASSP40776.
          <year>2020</year>
          .9053659.
          <string-name>
            <given-names>Z.</given-names>
            <surname>Jiang</surname>
          </string-name>
          ,
          <article-title>An eficient multi-task learning cnn for</article-title>
          [21]
          <string-name>
            <given-names>D.</given-names>
            <surname>Melesse</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Khalil</surname>
          </string-name>
          , E. Kagabo,
          <string-name>
            <given-names>T.</given-names>
            <surname>Ning</surname>
          </string-name>
          ,
          <string-name>
            <surname>K.</surname>
          </string-name>
          <article-title>Huang, driver attention monitoring</article-title>
          ,
          <source>Journal of Systems Appearance-based gaze tracking through superArchitecture</source>
          (
          <year>2024</year>
          )
          <article-title>103085. vised machine learning</article-title>
          ,
          <source>in: 2020 15th IEEE</source>
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>E.</given-names>
            <surname>Yüksel</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Acarman</surname>
          </string-name>
          ,
          <article-title>Experimental study on International Conference on Signal Processing driver's authority and attention monitoring</article-title>
          ,
          <source>in: (ICSP)</source>
          , volume
          <volume>1</volume>
          ,
          <year>2020</year>
          , pp.
          <fpage>467</fpage>
          -
          <lpage>471</lpage>
          . doi:
          <volume>10</volume>
          .1109/ Proceedings of 2011 IEEE International Conference ICSP48669.
          <year>2020</year>
          .
          <volume>9321075</volume>
          .
          <source>on Vehicular Electronics and Safety</source>
          ,
          <year>2011</year>
          , pp.
          <fpage>252</fpage>
          -
          <lpage>[</lpage>
          22]
          <string-name>
            <given-names>X.</given-names>
            <surname>Zhang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Sugano</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Fritz</surname>
          </string-name>
          ,
          <string-name>
            <surname>A</surname>
          </string-name>
          . Bulling, MPIIgaze: 257. doi:
          <volume>10</volume>
          .1109/ICVES.
          <year>2011</year>
          .
          <volume>5983824</volume>
          .
          <article-title>Real-world dataset and deep appearance-based gaze</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>G.</given-names>
            <surname>Capizzi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G. L.</given-names>
            <surname>Sciuto</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Napoli</surname>
          </string-name>
          , M. Woźniak, estimation,
          <year>2017</year>
          . URL: https://arxiv.org/abs/1711. G.
          <article-title>Susi, A spiking neural network-based long-term 09017. prediction system for biogas production</article-title>
          ,
          <source>Neural</source>
          [23]
          <string-name>
            <given-names>S.</given-names>
            <surname>Vora</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Rangesh</surname>
          </string-name>
          ,
          <string-name>
            <surname>M. M. Trivedi</surname>
          </string-name>
          ,
          <source>Driver gaze zone Networks</source>
          <volume>129</volume>
          (
          <year>2020</year>
          )
          <fpage>271</fpage>
          -
          <lpage>279</lpage>
          . doi:
          <volume>10</volume>
          .1016/j. estimation
          <source>using convolutional neural networks: A neunet</source>
          .
          <year>2020</year>
          .
          <volume>06</volume>
          .001.
          <article-title>general framework and ablative analysis</article-title>
          ,
          <year>2018</year>
          . URL:
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>B. A.</given-names>
            <surname>Nowak</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R. K.</given-names>
            <surname>Nowicki</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Woźniak</surname>
          </string-name>
          , C. Napoli, https://arxiv.org/abs/
          <year>1802</year>
          .02690.
          <string-name>
            <surname>Multi</surname>
          </string-name>
          <article-title>-class nearest neighbour classifier for incom-</article-title>
          [24]
          <string-name>
            <given-names>N.</given-names>
            <surname>Mizuno</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Yoshizawa</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Hayashi</surname>
          </string-name>
          , T. Ishikawa,
          <article-title>plete data handling</article-title>
          , volume
          <volume>9119</volume>
          ,
          <year>2015</year>
          , pp.
          <fpage>469</fpage>
          -
          <article-title>Detecting driver's visual attention area by using 480</article-title>
          . doi:
          <volume>10</volume>
          .1007/978-3-
          <fpage>319</fpage>
          -19324-3_
          <fpage>42</fpage>
          .
          <article-title>vehicle-mounted device</article-title>
          ,
          <source>in: 2017 IEEE 16th Inter-</source>
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <given-names>C.</given-names>
            <surname>Ciancarelli</surname>
          </string-name>
          , G. De Magistris, S. Cognetta, national Conference on Cognitive Informatics &amp;
          <string-name>
            <surname>D. Appetito</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          <string-name>
            <surname>Napoli</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          <string-name>
            <surname>Nardi</surname>
          </string-name>
          ,
          <article-title>A gan ap- Cognitive Computing (ICCI*CC)</article-title>
          ,
          <year>2017</year>
          , pp.
          <fpage>346</fpage>
          -
          <lpage>352</lpage>
          .
          <article-title>proach for anomaly detection in spacecraft teleme</article-title>
          - doi:10.1109/ICCI-CC.
          <year>2017</year>
          .
          <volume>8109772</volume>
          . tries 531 LNNS (
          <year>2023</year>
          )
          <fpage>393</fpage>
          -
          <lpage>402</lpage>
          . doi:
          <volume>10</volume>
          .1007/ [25]
          <string-name>
            <given-names>J.</given-names>
            <surname>Terven</surname>
          </string-name>
          ,
          <string-name>
            <surname>D.-M. Córdova-Esparza</surname>
          </string-name>
          , J.
          <article-title>-A. RomeroGonzález, A comprehensive review of yolo archi-</article-title>
          [36]
          <string-name>
            <given-names>P.</given-names>
            <surname>Viola</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Jones</surname>
          </string-name>
          ,
          <article-title>Rapid object detection using a tectures in computer vision: From yolov1 to yolov8 boosted cascade of simple features</article-title>
          ,
          <source>in: Proceedings and yolo-nas, Machine Learning and Knowledge of the 2001 IEEE Computer Society Conference on Extraction 5</source>
          (
          <year>2023</year>
          )
          <fpage>1680</fpage>
          -
          <lpage>1716</lpage>
          . Computer Vision and Pattern Recognition. CVPR
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [26]
          <string-name>
            <given-names>A. A.</given-names>
            <surname>Lima</surname>
          </string-name>
          ,
          <string-name>
            <surname>M. M. Kabir</surname>
          </string-name>
          , S.
          <string-name>
            <surname>C. Das</surname>
            ,
            <given-names>M. N.</given-names>
          </string-name>
          <string-name>
            <surname>Hasan</surname>
          </string-name>
          ,
          <year>2001</year>
          , volume
          <volume>1</volume>
          ,
          <year>2001</year>
          , pp.
          <source>I-I. doi:10</source>
          .1109/
          <string-name>
            <surname>CVPR. M. Mridha</surname>
          </string-name>
          ,
          <article-title>Road sign detection using variants of 2001.990517. yolo and r-cnn: An analysis from the perspective of bangladesh</article-title>
          ,
          <source>in: Proceedings of the International Conference on Big Data, IoT, and Machine Learning: BIM 2021</source>
          , Springer,
          <year>2022</year>
          , pp.
          <fpage>555</fpage>
          -
          <lpage>565</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [27]
          <string-name>
            <given-names>A.</given-names>
            <surname>Palazzi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Abati</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Calderara</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Solera</surname>
          </string-name>
          ,
          <string-name>
            <surname>R.</surname>
          </string-name>
          <article-title>CUcchiara, Predicting the driver's focus of attention: The dr(eye)ve project abs/</article-title>
          <year>1807</year>
          .02588 (
          <year>2018</year>
          ). URL: https://arxiv.org/abs/1705.03854. arXiv:
          <volume>1705</volume>
          .
          <fpage>03854</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [28]
          <string-name>
            <given-names>I.</given-names>
            <surname>Dua</surname>
          </string-name>
          ,
          <string-name>
            T. A. John,
            <given-names>R.</given-names>
            <surname>Gupta</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Jawahar</surname>
          </string-name>
          , Dgaze:
          <article-title>Driver gaze mapping on road</article-title>
          ,
          <source>in: 2020 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS)</source>
          , IEEE,
          <year>2020</year>
          , pp.
          <fpage>5946</fpage>
          -
          <lpage>5953</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [29]
          <string-name>
            <given-names>A.</given-names>
            <surname>Yoshizawa</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Iwasaki</surname>
          </string-name>
          ,
          <article-title>Analysis of driver's visual attention using near-miss incidents</article-title>
          ,
          <source>in: 2017 IEEE 16th International Conference on Cognitive Informatics &amp; Cognitive Computing (ICCI*CC)</source>
          ,
          <year>2017</year>
          , pp.
          <fpage>353</fpage>
          -
          <lpage>360</lpage>
          . doi:
          <volume>10</volume>
          .1109/ICCI-CC.
          <year>2017</year>
          .
          <volume>8109773</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          [30]
          <string-name>
            <surname>H. M. Peixoto</surname>
            ,
            <given-names>A. M. G.</given-names>
          </string-name>
          <string-name>
            <surname>Guerreiro</surname>
          </string-name>
          ,
          <string-name>
            <surname>A. D. D. Neto</surname>
          </string-name>
          ,
          <article-title>Image processing for eye detection and classification of the gaze direction</article-title>
          , in: 2009
          <source>International Joint Conference on Neural Networks</source>
          ,
          <year>2009</year>
          , pp.
          <fpage>2475</fpage>
          -
          <lpage>2480</lpage>
          . doi:
          <volume>10</volume>
          .1109/IJCNN.
          <year>2009</year>
          .
          <volume>5178924</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          [31]
          <string-name>
            <given-names>K.</given-names>
            <surname>Guo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Yu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <article-title>An new algorithm for analyzing driver's attention state</article-title>
          ,
          <source>in: 2009 IEEE Intelligent Vehicles Symposium</source>
          ,
          <year>2009</year>
          , pp.
          <fpage>21</fpage>
          -
          <lpage>23</lpage>
          . doi:
          <volume>10</volume>
          .1109/IVS.
          <year>2009</year>
          .
          <volume>5164246</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          [32]
          <string-name>
            <given-names>H.</given-names>
            <surname>Lee</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Seo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Jo</surname>
          </string-name>
          ,
          <article-title>Gaze tracking system using structure sensor &amp; zoom camera</article-title>
          ,
          <source>in: 2015 International Conference on Information and Communication Technology Convergence (ICTC)</source>
          ,
          <year>2015</year>
          , pp.
          <fpage>830</fpage>
          -
          <lpage>832</lpage>
          . doi:
          <volume>10</volume>
          .1109/ICTC.
          <year>2015</year>
          .
          <volume>7354677</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          [33]
          <string-name>
            <given-names>A. G.</given-names>
            <surname>Mavely</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. E.</given-names>
            <surname>Judith</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P. A.</given-names>
            <surname>Sahal</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S. A.</given-names>
            <surname>Kuruvilla</surname>
          </string-name>
          ,
          <article-title>Eye gaze tracking based driver monitoring system</article-title>
          ,
          <source>in: 2017 IEEE International Conference on Circuits and Systems (ICCS)</source>
          ,
          <year>2017</year>
          , pp.
          <fpage>364</fpage>
          -
          <lpage>367</lpage>
          . doi:
          <volume>10</volume>
          .1109/ICCS1.
          <year>2017</year>
          .
          <volume>8326022</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          [34]
          <string-name>
            <given-names>H.</given-names>
            <surname>Mohsin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S. H.</given-names>
            <surname>Abdullah</surname>
          </string-name>
          ,
          <article-title>Pupil detection algorithm based on feature extraction for eye gaze</article-title>
          ,
          <source>in: 2017 6th International Conference on Information and Communication Technology and Accessibility (ICTA)</source>
          ,
          <year>2017</year>
          , pp.
          <fpage>1</fpage>
          -
          <lpage>4</lpage>
          . doi:
          <volume>10</volume>
          .1109/ICTA.
          <year>2017</year>
          .
          <volume>8336048</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          [35]
          <string-name>
            <given-names>K.</given-names>
            <surname>Zhang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Zhang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Qiao</surname>
          </string-name>
          ,
          <article-title>Joint face detection and alignment using multi-task cascaded convolutional networks</article-title>
          ,
          <year>2022</year>
          . URL: https://arxiv. org/abs/1604.02878. doi:
          <volume>10</volume>
          .48550/ARXIV.2210. 07548.
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>