<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Face Recognition using Naive Bayes Classifier*</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Aleksandra Stachecka</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Tomasz Procek</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Faculty of Applied Mathematics, Silesian University of Technology</institution>
          ,
          <addr-line>Kaszubska 23, 44100 Gliwice</addr-line>
          ,
          <country country="PL">POLAND</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>IVUS2024: Information Society and University Studies 2024</institution>
        </aff>
      </contrib-group>
      <abstract>
        <p>The article is about the implementation and the performance of Face Recognition using Naive Bayes Classifier. Firstly,it is explained how huge impact on today's world has the face recognition system and how it has change over past 60 years. Then the idea of PCA and eigenfaces is brought closer to a reader. Moreover, the methodology of the program is explained. Not only, the most important functions and variables but also the schema of Naive Bayes Classifierare shown on code fragments. The next part is "Experiments" where viewer can find plots and specific information about dataset, examples of generated eigenfaces and the performance and accuracy of the program which is estimated to be nearly 75%. Finally, there is a conclusion. Authors one more time remind the most important information, explain the role of PCA used in the project and look to the future in order to improve their program.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;NaiveBayes</kwd>
        <kwd>FaceRecognition</kwd>
        <kwd>PCA</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
    </sec>
    <sec id="sec-2">
      <title>2. Methodology</title>
      <p>The process begins with acquiring and preprocessing the dataset. The "Labeled Faces in the Wild"
(LFW) dataset is fetched using the ℎ__ function from .. This dataset contains labeled images of
faces, and for this analysis, only inviduals with at least 50 face images are included to ensure
sufficient data per class. The images are resized to 50% of their original dimensions to reduce
computational complexity. The dataset is then split into a feature matrix  and a target vector .
The feature matrix  contains the pixel values of the images, while the target vector  contains
the corresponding class labels for each image. Additionally, the shape parameters of the images,
including the number of samples, height and width are stored for reference throughout the
analysis.</p>
      <p>Listing 1: Fetching and Preprocessing LFW dataset
1 lfw_people = fetch_lfw_people(min_faces_per_person=50, resize=0.5)
2 X = lfw_people.data
3 y = lfw_people.target
4 target_names = lfw_people.target_names
5 n_samples, h, w = lfw_people.images.shape
6
7 original_shape = (h, w)</p>
      <p>The feature matrix X is standarized using StandardScaler from sklearn.preprocessing. This
standarization is crucial for ensuring that the principal component analysis (PCA) operates
effectively. PCA is then applied to the standarized feature matrix  to reduce its dimensionality
while retaining most of the variance. This reduction in dimensionality is achieved by extracting the
most significantfeatures, known as principal components, from the data. In this analysis, 100
principal components are retained, as determined by the parameter _. The transformed data is
represented in a lower-dimensional space, resulting in the matrix  _ . PCA is particularly
wellsuited for this task because it effectively reduces the high dimen- sionality of image data
while preserving essential features that contribute to variance. This reduction is crucial for
computational efficiency and helps in avoiding the curse of dimensional- ity, which can adversely
affect machine learning algorithms. PCA also helps in noise reduction and improves the
performance of subsequent classifiersby focusing on the most significant
features.</p>
      <p>Listing 2: Standarization and PCA on the feature matrix
1 scaler = StandardScaler()
2 X_scaled = scaler.fit_transform(X)
3
4 # Applying PCA
5 n_components = 100 # Number of principal components to keep
6 pca = PCA(n_components=n_components, svd_solver=’randomized’, whiten=True)
7 X_pca = pca.fit_transform(X_scaled)
8
9 print(f"Original shape: {X_scaled.shape}")
10 print(f"Transformed shape: {X_pca.shape}")</p>
      <p>To visualize the principal components, known as eigenfaces, the components are reshaped to
the original image dimensions. The first 10 eigenfaces are displayed to show the main features
captured by PCA. Additionally, a function is definedto visualize the original and reconstructed
faces, allowing for an assessment of how well PCA captures important features. The
PCAtransformed data  _ is inversely transformed to reconstruct the original images, and a
subset of these reconstructed images is displayed alongside their original counterparts.</p>
      <p>Listing 3: Eigenfaces visualization and reconstruction of original faces
13
14
15
16
17
18
19
ax = plt.subplot(2, n_faces, i + 1 + n_faces)
plt.imshow(X_reconstructed[i].reshape((h, w)), cmap=’gray’)
plt.xticks(())
plt.yticks(())</p>
      <p>For classification, a Naive Bayes classifier 6[] is implemented from scratch. This involves
defining methods for fitting the model to training data, calculating likelihoods and posteriors and
making predictions. The dataset is split into training and testing sets using and 80-20 ratio with the
__ function from ._. The Naive Bayes classifieris trained on the training set and used
to predict labels for the test set. The accuracy of the classifier is then calculated using_
from ., providing a quantitative measure of the model’s performance.</p>
      <p>Listing 4: Naive Bayes classifierimplementation and prediction
1 class NaiveBayes:
2 def fit(self, X, y):
3 self.classes = np.unique(y)
4 self.mean = {}</p>
      <p>Finally, a single image from the test set is selected to demonstrate the classifier’s prediction
capability. The selected image is transformed back to its original dimensions, and its true and
predicted labels are displayed. This visual representation helps in understanding the model’s
performance on individual instances, complementing overall accuracy metric.</p>
      <p>Listing 5: Demonstrating classifier’s prediction capability
1 index = 12
2 single_image = X_test[index]
3 single_image_original = pca.inverse_transform(single_image).reshape(original_shape)
4
5 predicted_label = nb.predict(np.array([single_image]))[0]
6 true_label = y_test[index]
7</p>
    </sec>
    <sec id="sec-3">
      <title>3. Experiments</title>
      <p>
        The ___ℎ__ is widely used for facial recognition tasks [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ]. It contains JPEG images of
various famous people, and each image is labaled with the name of the person. Scikit-learn
[8] provides two loaders that will automatically download, parse the metadata files, decode
the JPEG and convert the slices into memmapped numpy arrays.
      </p>
      <p>By performing PCA we can notice a decrease in variance ratio in increasing number of
components. So that the plot helps to decide how many principal components to select to retain as
much information as possible while reducing the number of dimensions. We calculated that the
best option for us is 100 components.</p>
      <p>To check the performance of our Face Recognition using PCA and Bayes Classificator
programe we created the confusion matrix which is fundamental tool in this field.It provides a
detailed breakdown of how well the model’s predictions the actual class labels.</p>
      <p>Finally, in order to check detailed indicators as precision( ratio of true positive predictions to
the total predicted positives), recall(ratio of true positive predictions to the total actual
positives), F1-Score(average of precision and recall), support(the number of occurrences of
each class in dataset) the classification report was made.</p>
      <p>This report clearly shows the accuracy of our application which is estimated to be around
74%.</p>
    </sec>
    <sec id="sec-4">
      <title>4. Conclusion</title>
      <p>This project demonstrates the application of machine learning techniques for facial recognition
using the "Labeled Faces in the Wild" (LFW) dataset. Through meticulous preprocessing and
feature extraction, we prepared the dataset for analysis, ensuring its suitability for subsequent
machine learning algorithms. Principal Component Analysis (PCA) played a pivotal role in
reducing the dimensionality of the dataset while retaining essential variance, effectively
capturing the underlying structure of the facial images. The visualization of eigenfaces provided
valuable insights into the primary features captured by PCA, enhancing our understanding
of the dataset’s characteristics. The implementation of a Naive Bayes classifierfacilitated the
classification of facial images with satisfactory accuracy.By training the classifier on a subsetof
the dataset and evaluating its performance on unseen data, we gained valuable insights into
its generalization capabilities. Furthermore, the visual representation of prediction results provided
a tangible demonstration of the classifier’sability to accurately identify induviduals from facial
images, showcasing the practical utility of the developed model. Overall, this project
exemplifies the effectiveness of machine learning techniques in facial recognition tasks and
underscores their potential for diverse real-world applications, ranging from security and
surveillance to personalized user experience as beyond. As technology continues to evolve,
further advancements in machine learning algorithms and datasets hold promise for even more
accurate and robust facial recognition systems.
[8] Scikit-learn developers, Scikit-learn: Machine learning in python, https://scikit-learn.org/
stable/, 2023. Accessed: 2024-05-20.
[9] AIMonks, Principal component analysis (pca) in machine learning, https://medium.com/
aimonks/principal-component-analysis-pca-in-machine-learning-407224cb4527, 2023.
Accessed: 2024-05-20.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <surname>M.-H. Le</surname>
            ,
            <given-names>N. Carlsson,</given-names>
          </string-name>
          <article-title>Iddecoder: A face embedding inversion tool and its privacy and security implications on facial recognition systems</article-title>
          ,
          <source>in: Proceedings of the Thirteenth ACM Conference on Data and Application Security and Privacy</source>
          ,
          <year>2023</year>
          , pp.
          <fpage>15</fpage>
          -
          <lpage>26</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>A.</given-names>
            <surname>Jaszcz</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Połap</surname>
          </string-name>
          , Aimm:
          <article-title>Artificialintelligence merged methods for flood ddos attacks detection</article-title>
          ,
          <source>Journal of King Saud University-Computer and Information Sciences</source>
          <volume>34</volume>
          (
          <year>2022</year>
          )
          <fpage>8090</fpage>
          -
          <lpage>8101</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>K.</given-names>
            <surname>Prokop</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Połap</surname>
          </string-name>
          , G. Srivastava,
          <string-name>
            <given-names>J. C.-W.</given-names>
            <surname>Lin</surname>
          </string-name>
          ,
          <article-title>Blockchain-based federated learning with checksums to increase security in internet of things solutions</article-title>
          ,
          <source>Journal of Ambient Intelligence and Humanized Computing</source>
          <volume>14</volume>
          (
          <year>2023</year>
          )
          <fpage>4685</fpage>
          -
          <lpage>4694</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>M.</given-names>
            <surname>Wieczorek</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Siłka</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Woźniak</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Garg</surname>
          </string-name>
          ,
          <string-name>
            <surname>M. M. Hassan</surname>
          </string-name>
          ,
          <article-title>Lightweight convolutional neural network model for human face detection in risk situations</article-title>
          ,
          <source>IEEE Transactions on Industrial Informatics</source>
          <volume>18</volume>
          (
          <year>2021</year>
          )
          <fpage>4820</fpage>
          -
          <lpage>4829</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>B. U. H.</given-names>
            <surname>Sheikh</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Zafar</surname>
          </string-name>
          ,
          <article-title>Unlocking adversarial transferability: a security threat towards deep learning-based surveillance systems via black box inference attack-a case study on face mask surveillance</article-title>
          ,
          <source>Multimedia Tools and Applications</source>
          <volume>83</volume>
          (
          <year>2024</year>
          )
          <fpage>24749</fpage>
          -
          <lpage>24775</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>A.</given-names>
            <surname>Vidhya</surname>
          </string-name>
          , Naive bayes explained, https://www.analyticsvidhya.com/blog/2017/09/ naive-bayes-explained/,
          <year>2017</year>
          . Accessed:
          <fpage>2024</fpage>
          -05-20.
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>A. A.</given-names>
            <surname>Jha</surname>
          </string-name>
          , Lfw people dataset, https://www.kaggle.com/datasets/atulanandjha/lfwpeople,
          <year>2023</year>
          . Accessed:
          <fpage>2024</fpage>
          -05-20.
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>