<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Seeing in the Dark: A Diferent Approach to Night Vision Face Detection with Thermal IR Images</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Kinshuk Gaurav Singh</string-name>
          <email>kgaurav11286@gmail.com</email>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Charulkumar Chodvadiya</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Chintan Bhatt</string-name>
          <email>chintan.bhatt@sot.pdpu.ac.in</email>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Pooja Shah</string-name>
          <email>pooja.shah@sot.pdpu.ac.in</email>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Alessandro Bruno</string-name>
          <email>alessandrobruno10@gmail.com</email>
        </contrib>
      </contrib-group>
      <abstract>
        <p>Identifying faces in low-light situations can be a dificult feat because of the reduced visibility and substandard image quality. Conventional methods of face detection rely on visible light, which is insuficient in environments with low-light conditions. Our paper introduces a fresh approach to identifying faces in night vision. We make use of thermal infrared (IR) images to detect faces. Thermal IR images capture the thermal signatures of objects, which remain unafected by low-light conditions, providing valuable information for accurate face detection. Our method utilizes a deep learning model trained on thermal IR images to detect faces in conditions with low lighting. To evaluate our approach, we tested it on a dataset of thermal IR images captured in diferent lighting scenarios and compared its performance with traditional face detection methods. The results of our experiments indicate that our proposed approach surpasses traditional face detection methods in low-light conditions, achieving high accuracy in detecting faces. Through a qualitative analysis of thermal IR images, we delved into the key factors that lead to our approach's success. Our findings show that the thermal signature of the face ofers valuable insights that aid in precise face detection even in low-light environments. Furthermore, we assessed the efectiveness of our approach in detecting faces under diferent pose and expression variations, and the results indicate that our method is highly eficient in detecting faces in various pose and expression conditions. Our solution presents a novel and eficient method for detecting faces in low-light conditions using thermal infrared images, which has the potential to be utilized in various applications including surveillance, security and law enforcement.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;Face detection</kwd>
        <kwd>Thermal images</kwd>
        <kwd>U-net</kwd>
        <kwd>dlib</kwd>
        <kwd>yolo</kwd>
        <kwd>CNN</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>Face detection is a critical task in computer vision with a wide range of applications such as
surveillance, security, and law enforcement. However, detecting faces in low-light conditions is
a valuable asset and is a challenging problem due to limited visibility and poor image quality.
Traditional methods of face detection rely on visible light, which is inadequate in low-light
environments. Thus, there is a need for approaches to detect faces in low-light conditions.
The ability to see in the dark is a valuable asset in many situations, including security, law
enforcement, and search and rescue. Traditional night vision methods are limited by the amount
of light available, which can make it dificult to see objects in complete darkness, or in low-light
conditions.</p>
      <p>Night vision face detection is an area of research that has gained significant attention in
recent years. Various approaches have been proposed for this problem, such as using visible
light cameras with high-intensity infrared illuminators or active 3D imaging systems. While
these approaches have shown promising results, they have some limitations, such as high power
consumption, limited range, and high cost.</p>
      <p>In recent years, thermal infrared (IR) imaging has emerged as a promising technology for
night vision face detection. Thermal IR images are created by detecting the heat emitted by
objects. This makes them ideal for night vision, as they can be used to see objects in complete
darkness, which are not afected by low-light conditions, and can provide valuable information
for face detection. Thermal IR cameras are also less afected by environmental factors such as
fog, smoke, and dust, making them suitable for outdoor applications. However, using thermal
IR images for face detection requires the development of new methods and models that can
efectively utilize this modality.</p>
      <p>In this paper, we propose a novel approach to night vision face detection using thermal
IR images. Our approach involves the use of a deep learning model trained on thermal IR
images to detect faces in low-light conditions. We evaluated our method using a dataset of
thermal IR images captured in various lighting conditions and compared the performance of our
approach with traditional face detection methods. Our approach has a number of advantages
over traditional night vision methods. First, it is not afected by light conditions. This means
that it can be used to see faces in complete darkness, or in low-light conditions. Second, it can
be used to see through camouflage. This means that it can be used to identify people who are
trying to hide their faces.</p>
      <p>To investigate the factors that contribute to the success of our approach, we conducted a
qualitative analysis of the thermal IR images. Our analysis reveals that the thermal signature of
the face provides valuable information that can be used for accurate face detection in low-light
conditions. We also evaluated the robustness of our approach to diferent pose and expression
variations and found that our method is efective in detecting faces under various pose and
expression conditions. Night vision face detection with thermal IR images has a number of
advantages over traditional night vision methods. For example, it is not afected by light
conditions, and it can be used to see through camouflage. Night vision face detection with
thermal IR images has a number of potential applications in security and privacy. For example,
it can be used to identify people in dark environments, and it can be used to detect people who
are trying to hide their faces.</p>
      <p>There has been a lot of research on night vision face detection in recent years. One of the
most common approaches is to use a thermal IR camera to capture images of faces. The images
are then processed using a computer vision algorithm to detect faces. One of the challenges of
night vision face detection is that thermal IR images can be noisy. This can make it dificult for
computer vision algorithms to detect faces. Another challenge is that faces can be obscured by
objects, such as hats or sunglasses.</p>
    </sec>
    <sec id="sec-2">
      <title>2. Literature Review</title>
      <p>
        The method uses the reliable Hough transform approach to retrieve objects while filtering out
background activity noise. By using LC-Harris to extract 2D features, the depth of each detected
item is calculated. For eficient avoidance, the asynchronous adaptive collision avoidance
(AACA) method is used [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. In [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ], The hotspot and background removal techniques were used
on a dataset of thermal images. When compared to the hot-spot approach, the accuracy of the
background subtraction algorithm was higher at 79%. The background removal approach proved
more resilient to variations in lighting and brightness conditions than the other technique,
which both employed dynamic thresholds.
      </p>
      <p>
        While the tracking phase incorporates mean shift tracking and Kalman filter prediction, the
detection phase uses a support vector machine (SVM) that uses size-normalized pedestrian
candidates. To improve the detection phase, the road-detection module verifies pedestrian
observations. using an onboard forward-looking infrared (FLIR) camera, an autonomous vehicle
detection system for pedestrians in poor light. To distinguish between infrared pedestrians,
low-level Haar-like characteristics are used, while the AdaBoost learning algorithm selects
the most pertinent information. A keypoint-based region of interest (ROI) selection approach
in IR imageries is suggested in order to eficiently scan sub-windows [
        <xref ref-type="bibr" rid="ref4 ref5">4, 5</xref>
        ]. By combining
Haar features and HOG features in Cascade to classify and validate pedestrians, the system
is made robust, reducing the false alarm rate during pedestrian detection and eliminating
non-pedestrians in cluttered background situations [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ]. A face detector based on HOG-SVM,
landmark identification methods based on feature-based active appearance models, deep
alignment networks, a deep shape regression network, and an emotion recognition system are all
included. Unconstrained input can be transformed into frontal images using facial fractalization
algorithms [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ].
      </p>
      <p>
        CNN designs employ hierarchical, discriminative processing for facial recognition and mark
regression. The hierarchical filtering approach uses both global and local filtering and addresses
the hard "face drifting" and "landmark shaking" challenges. The Landmark quality evaluation
system and the Kalman filter both ensure the stability and durability of regional face components
[
        <xref ref-type="bibr" rid="ref7 ref9">7, 9</xref>
        ]. Facial emotion recognition method divides the face into parts and employs active regions
(ARs) like eyes and lips. Employing Convolutional Neural Networks (CNNs) and ten-fold
cross-validation improves recognition accuracy, further aided by parallelism techniques cutting
processing time in half. Decision-level fusion results in a 96.87% recognition accuracy, proving
the scheme’s robustness and usefulness in adverse conditions. The optimized approach increases
average recognition rates by 1%, addressing natural disasters, harsh weather, and low-light
scenarios. The technique efectively tracks ARs and enhances pose prediction, leading to
improved authentication accuracy [
        <xref ref-type="bibr" rid="ref18">18</xref>
        ].
      </p>
      <p>
        The Viola-Jones approach has the shortest detection time and an acceptable accuracy of
90.5±4.34% The pattern-matching method has an accuracy of 91.6±0.03%, while the
activecontour approach has an accuracy of 89.8±0.06%. A Viola-Jones and pattern-matching algorithm
combination might improve system accuracy [
        <xref ref-type="bibr" rid="ref13 ref14">13, 14</xref>
        ]. The stereo-based technology and
longwave infrared-based technology systems [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ] can robustly identify the location of the head
with high reliability, which can then be utilized to select how the airbag should be deployed.
Both algorithms identify and track the head with great precision, with success rates of 96.4%
and 90.1% for varied occupant categories.
      </p>
      <p>
        The camera processes images using OpenCV software and a Raspberry Pi (RPI), a tiny
computer about the size of a credit card. It uses control algorithms to handle alerts and sends
recorded photographs through Wi-Fi to the user’s email. For both human and smoke detection,
the system has an accuracy of 83.56% and 83.33 % respectively. Comparing OpenCV and dlib,
the OpenCV library is more efective, performs better for face identification and detection, and
has greater productivity [
        <xref ref-type="bibr" rid="ref4 ref6">4, 6</xref>
        ]. The night-vision system employs a dedicated ODROID XU4
microprocessor running the Ubuntu MATE operating system to process thermal images.The
deep learning strategy outperformed the Haar+Adaboost algorithm in terms of false detections
[
        <xref ref-type="bibr" rid="ref8">8</xref>
        ].
      </p>
      <p>
        Multiprocess method based on networking middleware that enables real-time face tracking
and emotion identification in thermal infrared photos. Approch can detect facial landmarks and
recognize facial emotions, even when the faces move around. Furthermore, the approach can
recognize face and head postures, which increases its capacity to handle arbitrary head poses
[
        <xref ref-type="bibr" rid="ref15">15</xref>
        ]. The inherent spatial and spectral aspects of several histogram-based feature descriptors
and a set of classifiers to classify thermal emotion. The existence of mixed emotions and
inter-person variability are just two of the most likely causes of low accuracy and precision
[
        <xref ref-type="bibr" rid="ref19">19</xref>
        ].
      </p>
      <p>
        Dynamic group diference coding is a screening method based on examining influencing
factors (D.G.D.C.). D.G.D.C. computes temperature diferences between the target person and the
recently passed crowd (dynamic group) by describing facial temperature with a face temperature
encoder (FTE) and constructing an embedding feature diference matrix. MLPs are used to
capture intrinsic information by describing the diference matrix in vertical and horizontal
directions [
        <xref ref-type="bibr" rid="ref17">17</xref>
        ].
      </p>
    </sec>
    <sec id="sec-3">
      <title>3. Methodology</title>
      <p>
        In this section, we explain the approach to achieving our objectives by applying advanced
computer vision techniques to thermal-visual image data while constantly fine-tuning our
techniques through rigorous testing and experimentation. We will briefly overview the datasets
we use and where they come from to provide some context. Then, we will discuss the primary
methods we use, such as Dlib, Haar-cascade, and Yolo v8, and explain why they are important
for our analysis process.
3.1. Data
The dataset consists of a total of 2556 pairs of thermal-visual images. Each image pair features
participants with meticulously labelled facial areas, including manually defined face bounding
boxes and precisely positioned 54 facial landmarks. The dataset was constructed from our
large-scale (SpeakingFaces dataset) [
        <xref ref-type="bibr" rid="ref20">20</xref>
        ].
      </p>
      <sec id="sec-3-1">
        <title>3.2. Dlib Approach</title>
        <p>A complete and potent open-source software ml toolkit, the dlib library provides a wide range
of features and capabilities for computer vision applications. The face identification capabilities
of the dlib package are efective and precise. Modern algorithms for identifying faces in pictures
or video streams are used. The Histogram of Oriented Gradients (HOG) approach and the
Convolutional Neural Network (CNN)-based method are the two main face identification
techniques ofered by dlib.</p>
        <sec id="sec-3-1-1">
          <title>3.2.1. Histogram of Oriented Gradients (HOG)</title>
          <p>
            In order to capture the distinguishing characteristics of faces, the HOG-based face identification
method implemented in dlib makes use of the idea of local picture gradient orientations. It
entails calculating gradient orientations in tiny picture areas, creating histograms of these
orientations, and representing the local image structure using these histograms. To categorise
areas as either face or non-facial, these histograms are then given into a classifier, such as a
support vector machine (SVM). The HOG-based technique in dlib is renowned for its success
in identifying frontal and near-frontal faces and functions well in a variety of illumination
situations[
            <xref ref-type="bibr" rid="ref23">23</xref>
            ].
          </p>
          <p>The gradient is obtained by combining magnitude and angle from the image. First Gx and Gy
is calculated using the formulae below for each pixel value .
(1)
(2)
(3)
(4)
where ,  refer to rows and columns respectively. After calculating Gx and, the magnitude
and angle of each pixel is calculated using the formulae mentioned below.</p>
          <p>(, ) = (,  + 1) − (,  − 1)
(, ) = ( − 1, ) − ( + 1, )</p>
          <p>Magnitude( ) = √︁2 + 2</p>
          <p>Angle( ) = ⃒⃒ tan− 1 (/)⃒⃒</p>
        </sec>
        <sec id="sec-3-1-2">
          <title>3.2.2. Convolutional Neural Networks (CNN)</title>
          <p>The CNN-based face identification technique in dlib makes use of deep neural networks to
extract distinguishing characteristics from the raw image data directly. A convolutional neural
network is trained using methods like convolutional layers, pooling layers, and fully connected
layers using a sizable dataset of labeled face photos. The next step is to use the trained network
to identify whether a section of an input picture contains a face or not. When detecting faces
with complicated postures, occlusions, and scale fluctuations, the CNN-based method in dlib
performs very well. It can record both subtle features and broad facial structures, producing
extremely accurate face detection outcomes.</p>
        </sec>
      </sec>
      <sec id="sec-3-2">
        <title>3.3. Haar-cascade Approach</title>
        <p>By examining the characteristics of an item, the Haar cascade approach uses machine learning
to find it. It employs a collection of basic rectangular characteristics known as Haar-like features
that are computed at various sizes and locations throughout a picture. The intensity changes
between neighboring rectangular picture sections are the source of these characteristics.</p>
        <p>
          The Haar-like features capture Local image fluctuations, which are then fed into a classifier
built on the AdaBoost algorithm—as input. AdaBoost creates a strong classifier by combining
a number of weak classifiers, successfully teaching the strong classifier how to distinguish
between diferent types of objects. The Haar cascade approach arranges the features in a cascade
structure, with a subset of all the characteristics comprising each step. Because the cascade
structure may swiftly reject portions of the picture that are unlikely to contain the item being
recognized, the processing is eficient[
          <xref ref-type="bibr" rid="ref22">22</xref>
          ].
        </p>
        <p>At each cascade step, the classifier assesses the Haar-like characteristics and decides based on
a threshold. The area moves on to the next step if the features meet the threshold for possible
detection. If not, the algorithm rejects it and continues to the following area of the image.</p>
        <p>At each stage, the Haar-like features are calculated for the image’s region of interest (ROI).
Let’s denote the ROI as (, , , ℎ), where (, ) is the top-left corner coordinates of the ROI,
and (, ℎ) is its width and height.</p>
        <p>The diference between the sums of pixel intensities in two rectangular sections inside the
ROI is used to calculate the Haar-like features. A single Haar-like feature can be described by
the following formula:</p>
        <p>= ∑︀ (pixels in the white rectangle) − ∑︀ (pixels in the black rectangle)
where ∑︀ denotes the sum of pixel intensities.</p>
        <p>A collection of weak classifiers, often based on straightforward decision functions, make up
the classifier at each level. The weak classifier may be modelled as follows:
ℎ( ) =
{︃


if  ≥ 
if  &lt; 
where ℎ( ) is the output of the weak classifier,  and  are the weights assigned to positive and
negative detections, respectively,  is the computed Haar-like feature, and  is the threshold.</p>
        <p>The outputs of the weak classifiers are combined at each stage of the cascade structure.
Applying a combination rule, such as a weighted sum or a voting mechanism, to the outputs of
all steps will yield the final choice.</p>
      </sec>
      <sec id="sec-3-3">
        <title>3.4. YoloV8 Architecture</title>
        <p>
          An important development in real-time object identification, the YOLOv8 architecture provides
outstanding performance. The backbone, head, and neck, which together comprise its design
and are essential to its outstanding skills, are its three main parts. The backbone uses Powerful
convolutional layers to extract high-level information from the input picture, allowing the
model to collect detailed characteristics and semantic data required for accurate object detection.
YOLOv8 creates a strong foundation for following processing stages by using a backbone
network that is resilient[
          <xref ref-type="bibr" rid="ref21">21</xref>
          ].
        </p>
        <p>The head module is essential for the extracted characteristics to be optimised and for reliable
predictions to be produced. It uses cutting-edge methods to predict bounding boxes, class
probabilities, and objectness ratings. The detecting head successfully accumulates object properties at
multiple sizes using a variety of convolutional layers, including 1x1 convolutions, allowing for
accurate localization and classification. The neck module in YOLOv8 also improves the model’s
capacity to recognise objects of various sizes and aspect ratios. This intermediate component
extends the backbone network’s feature collection by including extra context and semantic data.
To further improve the model’s capacity to handle objects of various sizes and looks, feature
pyramid networks (FPN) and spatial pyramid pooling (SPP) are often used for the neck module.
(5)
(6)</p>
      </sec>
      <sec id="sec-3-4">
        <title>3.5. Hybrid Approach</title>
        <p>Our research proposes a hybrid approach that combines the strengths of Dlib and YOLOv8
for object detection, with the aim of optimizing performance in diverse lighting conditions.
Initially trained a Dlib-based detector using a carefully annotated dataset collected specifically
for low-light, night conditions. This model underwent thorough training and validation to
improve its performance in nocturnal scenarios. The Dlib model was then integrated into the
YOLOv8 framework, and additional training was carried out using datasets. In the second phase,
we fine-tuned the model using a separate dataset captured in daylight conditions, allowing
YOLOv8 to adapt and specialize for optimal performance during the day while leveraging the
knowledge acquired during the initial Dlib training. This process leverages the strengths of
each model to create a hybrid architecture that exhibits superior object detection proficiency in
both day and night settings. We took care to maintain consistency in annotation format and
hyperparameter tuning for enhanced accuracy. Comprehensive evaluations on independent
test sets, including both day and night scenarios, validate the efectiveness of this hybrid model
in achieving superior robust detection across diverse lighting environments.</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>4. Result</title>
      <p>Our results comprise using multiple models to perform better at night and day. Face detection
can extend beyond day only or night only; the best way is to integrate multi-model architecture
trained on various data domains. This multi-headed model can perform well in many conditions.</p>
      <p>The U-net results can be observed in fig.3 for shape prediction of the face, which can be
helpful in recognizing persons. The red dots show the predicted points, and the green points
are the original points.</p>
      <p>
        The dlib face detector is trained and tested on [
        <xref ref-type="bibr" rid="ref20">20</xref>
        ] with a shape predictor for night images
and shows the accurate test results.
      </p>
      <p>
        The dlib face detector, on top of its results, can be observed in fig.5 for detecting the persons
or people’s faces at night. It is trained with [
        <xref ref-type="bibr" rid="ref20">20</xref>
        ]. The model predicts faces for the night and
thermal data on random images for testing model performance.
      </p>
      <p>For daylight conditions, this model fails to predict all faces. Thus for the day condition,
YoloV8 is used, fig 6. Which performs good for daylight conditions.</p>
      <p>Training on a common dataset for day and light can lead to limitations like sometimes
detecting noise instead of purely detecting faces, and some background clutter or occlusions can
lead to failures in predicting. Combining both models can give face detection far better results
than training on common ground, fig. 7 shows promising results.Our hybrid model combines
day and night features through multiple models, resulting in a more robust model.</p>
      <sec id="sec-4-1">
        <title>4.1. Model Performance</title>
        <sec id="sec-4-1-1">
          <title>4.1.1. Dlib Approach</title>
          <p>Testing revealed encouraging results for the face identification system built using the Dlib
library, which combines HoG and CNN algorithms. The test error was slightly greater at 3.29
than the train error, which was measured at 2.72. This shows that the model has learned to
generalize to new data quite efectively, even at night.</p>
        </sec>
        <sec id="sec-4-1-2">
          <title>4.1.2. Haar cascade Approach</title>
          <p>As we see in 1, it is clear that the Haar cascade algorithm exhibits a 48.15% accuracy in identifying
faces in test photos that contain both day and nighttime faces, correctly identifying 13 out of
27 faces. However, the algorithm’s accuracy increases to 50%, correctly identifying 9 out of 18
faces, when examining only nighttime photos. These probabilities show the algorithm’s varying
performance with respect to the time of day and indicate the need for more advancements to
increase face detection rates in both situations.</p>
        </sec>
        <sec id="sec-4-1-3">
          <title>4.1.3. Final Result Comparison</title>
          <p>The table 2 contains all the details for the day and night general case, comparing YoloV8, dlib,
and hybrid models.</p>
        </sec>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>5. Conclusions</title>
      <p>This research significantly contributes to the advancement of night vision face detection in
thermal infrared (IR) imaging. A comprehensive examination of various techniques, including
Haar Cascade, HOG+SVM, YOLOv8, and hybrid provides valuable insights for professionals
in computer vision and surveillance technologies. Whereas the Yolo trained on the night data
performs quite well than others, combining the weights of day and night light as a hybrid model
can make a generalized and more accurate ground truth model. The study underscores the
importance of method selection based on parameters like speed, accuracy, and computational
demands while emphasizing the pivotal role of high-quality, diverse training datasets. The
resulting improvements in accurate and eficient face detection systems for low-light
environments have the potential to meaningfully enhance security and safety measures, aligning with
the overarching objectives of this research endeavour.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <surname>Yasin</surname>
            ,
            <given-names>J. N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Mohamed</surname>
            ,
            <given-names>S. A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Haghbayan</surname>
            ,
            <given-names>M. H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Heikkonen</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Tenhunen</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Yasin</surname>
            ,
            <given-names>M. M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Plosila</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          (
          <year>2020</year>
          ,
          <article-title>October)</article-title>
          .
          <article-title>Night vision obstacle detection and avoidance based on Bio-Inspired Vision Sensors</article-title>
          . In 2020 IEEE SENSORS (pp.
          <fpage>1</fpage>
          -
          <lpage>4</lpage>
          ). IEEE.
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <surname>Sharma</surname>
            ,
            <given-names>S. K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Agrawal</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Srivastava</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Singh</surname>
            ,
            <given-names>D. K.</given-names>
          </string-name>
          (
          <year>2017</year>
          , March).
          <article-title>Review of human detection techniques in night vision</article-title>
          .
          <source>In 2017 International Conference on Wireless Communications, Signal Processing and Networking (WiSPNET)</source>
          (pp.
          <fpage>2216</fpage>
          -
          <lpage>2220</lpage>
          ). IEEE.
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <surname>Abaya</surname>
            ,
            <given-names>W. F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Basa</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sy</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Abad</surname>
            ,
            <given-names>A. C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Dadios</surname>
            ,
            <given-names>E. P.</given-names>
          </string-name>
          (
          <year>2014</year>
          , November).
          <article-title>Low cost smart security camera with night vision capability using Raspberry Pi and OpenCV</article-title>
          . In 2014 International conference
          <article-title>on humanoid, nanotechnology, information technology, communication and control, environment and management (HNICEM) (pp</article-title>
          .
          <fpage>1</fpage>
          -
          <lpage>6</lpage>
          ). IEEE.
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <surname>Xu</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Liu</surname>
            ,
            <given-names>X.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Fujimura</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          (
          <year>2005</year>
          ).
          <article-title>Pedestrian detection and tracking with night vision</article-title>
          .
          <source>IEEE Transactions on Intelligent Transportation Systems</source>
          ,
          <volume>6</volume>
          (
          <issue>1</issue>
          ),
          <fpage>63</fpage>
          -
          <lpage>71</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <surname>Sun</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Wang</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Wang</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          (
          <year>2011</year>
          , J'anuary).
          <article-title>Night vision pedestrian detection using a forward-looking infrared camera</article-title>
          .
          <source>In 2011 International Workshop on Multi-Platform/MultiSensor Remote Sensing and Mapping</source>
          (pp.
          <fpage>1</fpage>
          -
          <lpage>4</lpage>
          ). IEEE.
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>N.</given-names>
            <surname>Boyko</surname>
          </string-name>
          ,
          <string-name>
            <given-names>O.</given-names>
            <surname>Basystiuk</surname>
          </string-name>
          and
          <string-name>
            <given-names>N.</given-names>
            <surname>Shakhovska</surname>
          </string-name>
          ,
          <article-title>"Performance Evaluation and Comparison of Software for Face Recognition, Based on Dlib and Opencv Library,"</article-title>
          <source>2018 IEEE Second International Conference on Data Stream Mining Processing (DSMP)</source>
          , Lviv, Ukraine,
          <year>2018</year>
          , pp.
          <fpage>478</fpage>
          -
          <lpage>482</lpage>
          , doi: 10.1109/DSMP.
          <year>2018</year>
          .
          <volume>8478556</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>Yi</given-names>
            <surname>Jin</surname>
          </string-name>
          , Xingyan Guo,
          <string-name>
            <given-names>Yidong</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Junliang</given-names>
            <surname>Xing</surname>
          </string-name>
          , Hui Tian,
          <article-title>"Towards stabilizing facial landmark detection and tracking via hierarchical filtering: A new method,"</article-title>
          <source>Journal of the Franklin Institute</source>
          , Volume
          <volume>357</volume>
          ,
          <string-name>
            <surname>Issue</surname>
            <given-names>5</given-names>
          </string-name>
          ,
          <year>2020</year>
          , Pages
          <fpage>3019</fpage>
          -
          <lpage>3037</lpage>
          , ISSN 0016-0032, https://doi.org/10.1016/j.jfranklin.
          <year>2019</year>
          .
          <volume>12</volume>
          .043.
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <surname>Nowosielski</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Małecki</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Forczmański</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Smoliński</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Krzywicki</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          (
          <year>2020</year>
          ).
          <article-title>Embedded night-vision system for pedestrian detection</article-title>
          .
          <source>IEEE Sensors Journal</source>
          ,
          <volume>20</volume>
          (
          <issue>16</issue>
          ),
          <fpage>9293</fpage>
          -
          <lpage>9304</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>Y.</given-names>
            <surname>Wu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Hassner</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Kim</surname>
          </string-name>
          , G. Medioni and
          <string-name>
            <given-names>P.</given-names>
            <surname>Natarajan</surname>
          </string-name>
          ,
          <article-title>"Facial Landmark Detection with Tweaked Convolutional Neural Networks,"</article-title>
          <source>in IEEE Transactions on Pattern Analysis and Machine Intelligence</source>
          , vol.
          <volume>40</volume>
          , no.
          <issue>12</issue>
          , pp.
          <fpage>3067</fpage>
          -
          <issue>3074</issue>
          , 1 Dec.
          <year>2018</year>
          , doi: 10.1109/TPAMI.
          <year>2017</year>
          .
          <volume>2787130</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <surname>Govardhan</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Pati</surname>
          </string-name>
          , U. C.
          <article-title>(2014, May)</article-title>
          .
          <article-title>NIR image based pedestrian detection in night vision with cascade classification and validation</article-title>
          .
          <source>In 2014 IEEE International Conference on Advanced Communications, Control and Computing Technologies</source>
          (pp.
          <fpage>1435</fpage>
          -
          <lpage>1438</lpage>
          ). IEEE.
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <surname>Kopaczka</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Nestler</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Merhof</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          (
          <year>2017</year>
          ).
          <article-title>Face detection in thermal infrared images: A comparison of algorithm-and machine-learning-based approaches</article-title>
          .
          <source>In Advanced Concepts for Intelligent Vision Systems: 18th International Conference, ACIVS</source>
          <year>2017</year>
          , Antwerp, Belgium,
          <source>September 18-21</source>
          ,
          <year>2017</year>
          , Proceedings
          <volume>18</volume>
          (pp.
          <fpage>518</fpage>
          -
          <lpage>529</lpage>
          ). Springer International Publishing.
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <surname>Krotosky</surname>
            ,
            <given-names>S. J</given-names>
          </string-name>
          ., Cheng, S. Y.,
          <string-name>
            <surname>Trivedi</surname>
            ,
            <given-names>M. M.</given-names>
          </string-name>
          (
          <year>2004</year>
          ,
          <article-title>October)</article-title>
          .
          <article-title>Face detection and head tracking using stereo and thermal infrared cameras for" smart" airbags: a comparative analysis</article-title>
          .
          <source>In Proceedings. The 7th International IEEE Conference on Intelligent Transportation Systems (IEEE Cat. No. 04TH8749)</source>
          (pp.
          <fpage>17</fpage>
          -
          <lpage>22</lpage>
          ). IEEE.
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <surname>Ribeiro</surname>
            ,
            <given-names>R. F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Fernandes</surname>
            ,
            <given-names>J. M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Neves</surname>
            ,
            <given-names>A. J.</given-names>
          </string-name>
          (
          <year>2017</year>
          ).
          <article-title>Face detection on infrared thermal image</article-title>
          .
          <source>Signal</source>
          ,
          <volume>45</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <surname>Kwaśniewska</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rumiński</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          (
          <year>2016</year>
          ,
          <article-title>July)</article-title>
          .
          <article-title>Face detection in image sequences using a portable thermal camera</article-title>
          .
          <source>In Proceedings of the 13th Quantitative Infrared Thermography Conference</source>
          (pp.
          <fpage>4</fpage>
          -
          <lpage>8</lpage>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <surname>Kopaczka</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Schock</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Nestler</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kielholz</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Merhof</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          (
          <year>2018</year>
          ,
          <article-title>October). A combined modular system for face detection, head pose estimation, face tracking and emotion recognition in thermal infrared images</article-title>
          .
          <source>In 2018 IEEE International Conference on Imaging Systems and Techniques (IST)</source>
          (pp.
          <fpage>1</fpage>
          -
          <lpage>6</lpage>
          ). IEEE.
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [16]
          <string-name>
            <surname>Kopaczka</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Breuer</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Schock</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Merhof</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          (
          <year>2019</year>
          ).
          <article-title>A modular system for detection, tracking and analysis of human faces in thermal infrared recordings</article-title>
          .
          <source>Sensors</source>
          ,
          <volume>19</volume>
          (
          <issue>19</issue>
          ),
          <fpage>4135</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [17]
          <string-name>
            <surname>Yan</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Qian</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Wang</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gao</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Yang</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          (
          <year>2023</year>
          ).
          <article-title>Dynamic Group Diference Coding based on Thermal Infrared Face Image for Fever Screening</article-title>
          .
          <source>IEEE Transactions on Instrumentation and Measurement.</source>
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          [18]
          <string-name>
            <surname>Assiri</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hossain</surname>
            ,
            <given-names>M. A.</given-names>
          </string-name>
          (
          <year>2023</year>
          ).
          <article-title>Face emotion recognition based on infrared thermal imagery by applying machine learning and parallelism</article-title>
          .
          <source>Mathematical Biosciences and Engineering</source>
          ,
          <volume>20</volume>
          (
          <issue>1</issue>
          ),
          <fpage>913</fpage>
          -
          <lpage>929</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          [19]
          <string-name>
            <surname>Rooj</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Routray</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Mandal</surname>
            ,
            <given-names>M. K.</given-names>
          </string-name>
          (
          <year>2023</year>
          ).
          <article-title>Feature based analysis of thermal images for emotion recognition</article-title>
          .
          <source>Engineering Applications of Artificial Intelligence</source>
          ,
          <volume>120</volume>
          ,
          <fpage>105809</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          [20]
          <string-name>
            <surname>Kuzdeuov</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Koishigarina</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Aubakirova</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Abushakimova</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Varol</surname>
            ,
            <given-names>H. A.</given-names>
          </string-name>
          (
          <year>2022</year>
          ).
          <article-title>SF-TL54: A Thermal Facial Landmark Dataset with Visual Pairs</article-title>
          .
          <source>In 2022 IEEE/SICE International Symposium on System Integration</source>
          ,
          <string-name>
            <surname>SII</surname>
          </string-name>
          <year>2022</year>
          (pp.
          <fpage>748</fpage>
          -
          <lpage>753</lpage>
          ).
          <article-title>(2022 IEEE/SICE International Symposium on System Integration, SII</article-title>
          <year>2022</year>
          ).
          <article-title>Institute of Electrical and Electronics Engineers Inc</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          [21]
          <string-name>
            <surname>Glenn</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          <string-name>
            <surname>Ultralytics</surname>
          </string-name>
          , YOLOv8.https://github.com/ultralytics/ultralytics(
          <year>2023</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          [22]
          <string-name>
            <given-names>L.</given-names>
            <surname>Cuimei</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Q.</given-names>
            <surname>Zhiliang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Nan</surname>
          </string-name>
          and
          <string-name>
            <given-names>W.</given-names>
            <surname>Jianhua</surname>
          </string-name>
          ,
          <article-title>"Human face detection algorithm via Haar cascade classifier combined with three additional classifiers,"</article-title>
          <source>2017 13th IEEE International Conference on Electronic Measurement Instruments (ICEMI)</source>
          , Yangzhou, China,
          <year>2017</year>
          , pp.
          <fpage>483</fpage>
          -
          <lpage>487</lpage>
          , doi: 10.1109/ICEMI.
          <year>2017</year>
          .
          <volume>8265863</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          [23]
          <string-name>
            <given-names>N.</given-names>
            <surname>Dalal</surname>
          </string-name>
          and
          <string-name>
            <given-names>B.</given-names>
            <surname>Triggs</surname>
          </string-name>
          ,
          <article-title>"Histograms of oriented gradients for human detection,"</article-title>
          <source>2005 IEEE Computer Society Conference on Computer Vision and Pattern Recognition (CVPR'05)</source>
          , San Diego, CA, USA,
          <year>2005</year>
          , pp.
          <fpage>886</fpage>
          -
          <lpage>893</lpage>
          vol.
          <volume>1</volume>
          , doi: 10.1109/CVPR.
          <year>2005</year>
          .
          <volume>177</volume>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>