<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <journal-title-group>
        <journal-title>Information Control Systems &amp; Technologies, September</journal-title>
      </journal-title-group>
    </journal-meta>
    <article-meta>
      <title-group>
        <article-title>RF-PSO: An Optimized Approach for Diabetes Prediction</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Houda El Bouhissi</string-name>
          <email>houda.elbouhissi@gmail.com</email>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Amine Ziane</string-name>
          <email>amine.ziane@univ-bejaia.dz</email>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Lamia Rahmani</string-name>
          <email>lamia.rahmani@univ-bejaia.dz</email>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Meriem Medbal</string-name>
          <email>meriem.medbal@univ-bejaia.dz</email>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Mariia Kostiuk</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Supérieure</string-name>
          <xref ref-type="aff" rid="aff2">2</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Random</string-name>
          <xref ref-type="aff" rid="aff2">2</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Khmelnytskyi National University</institution>
          ,
          <addr-line>Institutska str., 11, Khmelnytskyi, 29016</addr-line>
          ,
          <country country="UA">Ukraine</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>LIMED Laboratory, Faculty of Exact Sciences, University of Bejaia</institution>
          ,
          <addr-line>06000, Bejaia</addr-line>
          ,
          <country country="DZ">Algeria</country>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>Machine Learning, Diabetes</institution>
          ,
          <addr-line>Feature extraction, Random</addr-line>
        </aff>
        <aff id="aff3">
          <label>3</label>
          <institution>Numérique</institution>
          ,
          <addr-line>RN 75, Amizour, 06300 Bejaia</addr-line>
          ,
          <country country="DZ">Algeria</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2023</year>
      </pub-date>
      <volume>2</volume>
      <fpage>1</fpage>
      <lpage>23</lpage>
      <abstract>
        <p>Diabetes is a chronic disease due to a malfunction of the pancreas, which leads to high concentration of blood sugar in the blood and can affect the functioning of the body system. High blood sugar levels contribute to complications, over time as it can damage the heart, blood vessels, eyes, kidneys, and nerves...etc. Therefore, we need to develop a system capable of effectively diagnosing diabetic patients using medical details.</p>
      </abstract>
      <kwd-group>
        <kwd>Optimization</kwd>
        <kwd>1</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>for
training</p>
    </sec>
    <sec id="sec-2">
      <title>Random</title>
    </sec>
    <sec id="sec-3">
      <title>Forest</title>
      <p>algorithm.
performance of the
proposal
was
compared.</p>
      <p>The
results
demonstrate
that
combination of the Random</p>
    </sec>
    <sec id="sec-4">
      <title>Forest algorithm with the Particle Swarm Optimization for</title>
      <p>the
The
the
algorithm provides better accuracy.
Prediction,
Optimization.</p>
      <sec id="sec-4-1">
        <title>1. Introduction and Motivation</title>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>Diabetes is a</title>
      <p>widely chronic
malady that occurs
when the pancreas does not produce
enough insulin, or when the body cannot effectively use it has produced insulin. Insulin is a
hormone that regulates blood sugar levels.</p>
      <p>Hyperglycemia, also known as high blood sugar, is a common effect of uncontrolled diabetes and,
over time, leads to considerable damage to many of the body's systems, particularly nerves and blood
vessels.</p>
      <p>There has been drastic increase in rate of people with diabetes since a decade. Current human
lifestyle is the main reason behind growth in diabetes, due to unhealthy diet, lifestyle, stress, insufficient
physical effort, and sport.</p>
      <p>
        According to the World Health Organization [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ] there are 400 million people with diabetes in the
world. Each year, 1.5 million people die of diabetes, due to its complications and diabetes-related death
rates rose 13% in lower-middle income nations.
      </p>
      <p>In general, diabetics, live in low-and middle-income countries. Both the number of cases and the
prevalence of diabetes have been steadily increasing over the past few decades.</p>
      <p>2023 Copyright for this paper by its authors.</p>
      <p>In Algeria, the authorities are aware of the deadly clutch of diabetes and are implementing
mechanisms to support diabetic patients through the media, care and follow-up centers and charity
associations, however, people are not aware of the scale of this disease which is becoming rampant over
time.</p>
      <p>
        We distinguish two types of diabetes [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]: type 1 and type 2.
• Diabetes type 1: also called insulin-dependent, this disease affects 5 to 10% of all diabetics, is
brought on by cellular-mediated autoimmune destruction of the pancreatic B-cells.
• Diabetes type 2: also called non-insulin-dependent diabetes, this kind of diabetes affects people
who have insulin resistance and typically have a relative (as opposed to an absolute) insulin shortage.
It accounts for 90–95% of people with diabetes.
• In addition, due to a variety of clinical, social, and lifestyle reasons, "gestational diabetes" is a
different type of diabetes and one of the most common diseases in pregnant women.
      </p>
      <p>With its growing prevalence in recent years, diabetes has become a disease of the century. Health
professionals recommend that individuals take preventive measures by following certain guidelines to
reduce risk factors and thus avoid the onset and complications of this disease.</p>
      <p>However, as the number of individuals with this disease rises, so does the complexity of
its therapy. To enable early diagnosis of infections and gain a better knowledge of the
variables that significantly influence their occurrence, researchers have developed several
methods.</p>
      <p>
        Several machine-learning based systems have been implemented to support decreasing
the infection risk. Risk prediction is useful in many domains in our daily life [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]. Early
diabetes prediction makes to knowledge the most critical factors to control the disease.
      </p>
      <p>
        The present study is a continuation of our previous contribution [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ], which focused on
gestational diabetes. In [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ], we used three classifiers, and concluded that Random Forest
(RF) performs better.
      </p>
      <p>
        The purpose of this paper is to improve the results achieved in [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ] and propose a novel
approach for diabetes prediction. The proposal, called RF-PSO, involves four main steps
and uses the PSO (Particle Swarm Optimization) algorithm and the RF classifier:
• RF [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ] is belongs to the supervised ML techniques which is an ensemble learning method that
combines multiple decision trees to make predictions.
• PSO [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ] is a metaheuristic optimization algorithm inspired by the social behavior of bird
flocking or fish schooling for feature selection and the for the prediction process.
      </p>
    </sec>
    <sec id="sec-6">
      <title>Thus, the contributions of this paper are:</title>
      <p>• Automatic prediction of diabetes in healthcare systems to early detection of the disease,
therefore, healthcare systems can identify women at risk of diabetes before the disease’s beginning.
In this sense, our work can be considered as a making-decision proposal in healthcare.
• Implementation of an effective diabetes tool for prediction to contribute of the population
health, especially women in pregnancy.
• Highlighting the benefits of bioinspired algorithms for the optimization of machine leaning
algorithms.</p>
      <p>The paper is organized as follows. Main related works on diabetes prediction is
discussed in Section 2. Section 3 describes the proposed approach based on PSO and RF,
while experiments and results are reported in Section 4. Finally, Section 5 concludes the
paper and present future works.</p>
      <sec id="sec-6-1">
        <title>2. Related works</title>
        <p>Various prediction models have been developed and implemented by researchers using
different techniques such as machine-learning algorithms. The analysis of related work
gives an insight of healthcare datasets used, where analysis and predictions have been
realized with several methods and techniques.</p>
        <p>Diabetes prediction is a significant area of research that has captured the attention of
many researchers. While numerous systems have been developed, to the best of our
knowledge, most of the works use RF algorithm, few works are dedicated to gestational
diabetes and the utilization of swarm intelligence remains unexplored.</p>
        <p>In this section, we present the main related works according to our research and attempt
to explore the key features to improve the proposed methodology.</p>
        <p>
          The aim of the study proposed in [
          <xref ref-type="bibr" rid="ref7">7</xref>
          ] is to build a model for early diabetes prediction
with multiple machine-learning models. The authors use four supervised machine-learning
algorithms: RF, Naïve Bayes (NB), Decision Tree (DT) and Support Vector Machine
(SVM) and an unsupervised algorithm, K-means clustering to classify the patients into
categories. The experiments were performed on the PIMA Indian Diabetes and the patient
dataset provided by the hospital of Frankfurt in Germany. The authors start by
implementing the k-means classifier simultaneously on the two databases to clean the
columns where there is missing data and replace it with the cluster, and then they split the
PIMA dataset into 80% for training and 20% for testing and the German dataset into 70%
for training and 30% for testing. They tested each machine-learning algorithm on the
datasets, next they calculated the accuracy, the precision, f1-measure and recall of each
algorithm to evaluate and compare them. The results of the evaluation process showed that
RF performed the best on the Frankfurt dataset with 97.60% of accuracy and that SVM was
the best model on the PIMA dataset with an accuracy of 83.10%, the initialization of cluster
centers and the number of clusters influenced the increase and decrease of the performance
of algorithms.
        </p>
        <p>
          The authors in [
          <xref ref-type="bibr" rid="ref8">8</xref>
          ] realized a system of early diabetes mellitus prediction based on
artificial neural network and machine-learning algorithms. The test of the model has been
done in the National Institute of Diabetes, Digestive, and kidney diseases. The authors
proposed three models: artificial neural network (ANN), RF and k-means clustering. First,
they implemented the ANN, the RF algorithm and finally the k-means clustering. ANN
outperformed the two other models with an accuracy of 75.70%.
        </p>
        <p>
          Another approach proposed in [
          <xref ref-type="bibr" rid="ref9">9</xref>
          ], which combines four algorithms: RF, J48, K-Nearest
Neighbor (KNN) and NB, with WEKA hybrid system. The authors have built a hybrid
model based on different machine-learning algorithms to increase the accuracy of the
diabetes mellitus prediction. The approach uses the PIMA Indian diabetes database and
involves different steps. The first step concerns the data processing, and then they split the
dataset into two sets: 90% for training set and 10% for testing. The evaluation of the model
was based on the accuracy, F-measure, Recall and Precision, the authors tested the four
algorithms individually and the hybrid model. The authors claim that the combination of the
four algorithms provided a better accuracy than each algorithm individually.
        </p>
        <p>
          The authors in [
          <xref ref-type="bibr" rid="ref10">10</xref>
          ] use a variety of machine learning algorithms. The authors indicated
that approaches based on machine learning provide better prediction outcomes than
achieved by building models from patient-collected datasets. The main purpose of the
proposal is to create a system, which, by combining several machine-learning approaches,
can accurately conduct early diabetes prediction for a patient. The used algorithms are
Gradient Boosting (GB), RF, DT, SVM, KNN and logistic Regression (LR). The
emergence of different symptoms has been used to detect the presence of disease. The
methodologies, metrics, and features that were employed affect the outcome of the
prediction. A Disease Influence Measure (DIM) based diabetic prediction has been
provided as a step toward diabetes prediction. The approach performs a preliminary
processing on the input data set, removing the noisy records. The method calculates disease
influence measure (DIM) in the second step using the characteristics of the input data point.
The technique performs diabetic diagnosis depending on the DIM value. Various disease
prediction methods have been taken into consideration, and their effectiveness has been
compared. The analysis's results have been provided in-depth for development. The project
work reveals that the model is capable of accurately predicting diabetes with an accuracy
that exceeds 95%.
        </p>
        <p>
          The authors in [
          <xref ref-type="bibr" rid="ref11">11</xref>
          ] built a model for diabetes prediction based on various
machinelearning algorithms. The study was performed on the PIMA dataset. The authors
implemented k-means clustering algorithm on two attributes of the dataset that are "age"
and "Glucose", to classify each patient into a diabetic or non-diabetic class. The
machinelearning model was built with SVM, RF, DT, Extra Tree Classifier, Ada Boost algorithm,
Perceptron, Linear Discriminant Analysis algorithm, LR, KNN, Gaussian NB, Bagging
algorithm and Gradient Boost Classifier, next created a pipeline for the algorithms that gave
the highest accuracy and then evaluated the pipelines. The proposal evaluation showed that
LR provide the highest accuracy of 96%, but using the pipeline Ada Boost algorithm was
the best model with an accuracy of 98.80%.
        </p>
        <p>
          The authors in [
          <xref ref-type="bibr" rid="ref12">12</xref>
          ] used various machine-learning techniques to predict type-2 diabetic
mellitus disease. The authors' purpose was to create a predictive model that can accurately
predict whether a person has diabetes or not. The chosen classifiers are logistic DT,
ExtraTrees, RF, regression, XGBoost, gradient boosting, and the light gradient boosting
machine (LGBM). Initially gathered and kept in the database are the data required for the
investigation. The authors performed their tasks using the PIMA dataset. After that, the
dataset is pre-processed using various strategies for exploratory data analysis. The dataset is
splited into "training data" and "testing data." The best algorithm that works and has the
highest accuracy is selected as the best model for predicting the disease after various
experimentations and comparison of the obtained results. After evaluating each algorithm,
the LGBM algorithm showed the best performance with an accuracy of 95.20 %.
        </p>
        <p>Analysis of related work yields results on various healthcare datasets, where analysis
and predictions were performed using a variety of methods and techniques. Some of these
research utilized use of unique datasets or a mixture of other databases.</p>
        <p>Various researchers have proposed different prediction models. These models use
machine-learning, deep learning algorithms, data mining techniques or a combination of
these techniques. All these works provide a promising outcome and differ in precision and
accuracy according to used datasets. A review of the relevant papers leads to the conclusion
that researchers have successfully combined several machine learning algorithms with
various data preprocessing approaches for the automatic detection of diabetes.</p>
        <p>Our work is inspired by related works and building on our initial contribution where we
combined machine-learning algorithms with swarm intelligence optimization to achieve
better accuracy.</p>
      </sec>
      <sec id="sec-6-2">
        <title>3. Proposed Methodology</title>
        <p>
          Over the past decade, the number of diabetic people has grown dramatically. The current
human lifestyle is the main reason for the increase in diabetes, as are unhealthy eating
habits, anxiety, use of tools that reduce physical effort [
          <xref ref-type="bibr" rid="ref13">13</xref>
          ], and many other reasons.
        </p>
        <p>Our study focuses on gestational diabetes, which is common among women during
pregnancy who do not already have diabetes and can lead to various complications.</p>
        <p>Every year, gestational diabetes affects 2% to 10% of pregnancies in the United States
[14]. Managing gestational diabetes is crucial to ensuring a healthy pregnancy and the
wellbeing of your baby. In Algeria, the prevalence of gestational diabetes among pregnant
women varies between 2% and 5%.</p>
        <p>
          In our earlier research [
          <xref ref-type="bibr" rid="ref4">4</xref>
          ], we sought to predict early gestational diabetes more
accurately. We chose the most significant machine learning (ML) algorithms that produce
the greatest results in the literature, namely Deep Neural Network (DNN) [15], SVM, and
RF, and we found that RF outperformed the other two algorithms and produced the best
results.
        </p>
        <p>
          The present study is a continuation of our previous contribution [
          <xref ref-type="bibr" rid="ref4">4</xref>
          ] and attempts to
improve the RF algorithm by selecting the optimal features using the PSO algorithm for
gestational diabetes prediction. In this sense, the present paper makes use of PSO and RF
classifier for diabetes prediction with an improved accuracy.
        </p>
        <p>The aim of our research is to predict in the future whether a healthy woman will become
diabetic or not, i.e., affected by gestational diabetes, regarding woman data such as age,
weight, personal information, and her history.</p>
        <p>The Figure 1 depicted the overall architecture of the RF-PSO proposal, which involves
four main steps:
• The first step concerns the data collection.
• The second step concerns data processing. This step involves many phases for treating and
managing data.
• The third step is the main step which includes the selection of the best features for the RF
algorithm; this step is performed using the PSO algorithm.</p>
        <p>• Finally, the last step concerns the prediction process using the RF algorithm.</p>
        <p>The proposal is supported by a software tool deployed on a smartphone.</p>
      </sec>
    </sec>
    <sec id="sec-7">
      <title>These steps will be examined in detail in the following sections.</title>
      <p>3.1.</p>
      <sec id="sec-7-1">
        <title>Data collection</title>
        <p>Data collection is an essential initial step in the diabetes prediction. It enables proper treatment and
accurate assessment of the study population. The purpose of this step is to ensure that the information
gathered is complete and reliable, enabling better evaluation of results and more accurate anticipation
of future probabilities and trends.</p>
        <p>This step includes data gathering and understanding to study the patterns and trends, which helps
in prediction, and evaluating the results.</p>
        <p>The dataset, whose description is provided in table 1, is public and originated from Kaggle and is
made up of several prognostic medical variables obtained from the hospital in Frankfurt, Germany [16].
The dataset includes 2000 person between diabetic and no diabetic and 9 attributes with a size 12.0 kB
and two diabetes classes:
• Class 0 means healthy person.</p>
        <p>• Class 1 means diabetic person.</p>
        <p>It is important to note that this dataset was expanded by 200 records (patient cases), the persons’ data
were collected according to the same attributes. These additional records were gathered, with
confidentiality, from the Internal Medicine Department of Khellil Amrane Hospital in Bejaia City,
Algeria [17] during student internships at the hospital.
This model phase manages inconsistent data to get more accurate and precise results, which
is essential for our prediction process. The dataset contains missing values. Therefore, we
delete missing values for few selected attributes like Glucose level, Blood Pressure, Skin
Thickness, BMI, and Age because these attributes cannot be null. Then we scale the dataset
to normalize all values.</p>
        <p>The processing at all goes through different phases; first, the data was refined to
eliminate errors. Second, a data cleaning and filtering process was conducted to avoid the
formation of inappropriate rules and patterns, which included the removal of noisy data and
missing values, as well as the removal of duplicates and irrelevant data. To resolve
discrepancies, anomalies in noisy data were removed.</p>
        <p>For example, if an attribute such as blood glucose had zero values (this is not possible in
daily life), all such values were substituted by the median value of that attribute.
3.3.</p>
      </sec>
      <sec id="sec-7-2">
        <title>Feature selection</title>
        <p>The purpose of this step is to achieve the first objective, which is to increase prediction
accuracy by minimizing the feature selection subsets and select the best features from the
disease dataset.</p>
        <p>In this phase, feature selection was carried out using the PSO [18] algorithm to select the
best features and reduce space calculation. The reduced subset includes significant features
related to the dataset. Next, the RF algorithm was used for classification for a better
prediction accuracy.</p>
        <p>The fundamental concept behind the fusion of RF with PSO is to leverage PSO for
parameter optimization within the RF algorithm. This optimization procedure aims to
enhance the performance of the RF model by identifying the optimal parameter subset that
minimizes prediction errors and maximizes accuracy.</p>
        <p>In this way, the application of swarm intelligence optimization algorithms on machine
learning has produced good results in other fields, such as image classification [19].</p>
        <p>PSO is a swarm intelligence algorithm, which is advantageous in many aspects. PSO is
very popular and usually used to resolve several optimization problems in different
domains.</p>
        <p>PSO is simple to implement and uses only few parameters. In addition, it is fast with
simple complexity and includes an efficient global search process.</p>
        <p>As depicted in figure 2, the PSO technique goes through various steps, starting from the
initialization of all parameters, constants, values for position and velocity, then evaluate the
fitness of each particle to update position and velocity of each particle and finally determine
the best position (in our case, best features).</p>
        <sec id="sec-7-2-1">
          <title>Initialisation</title>
        </sec>
        <sec id="sec-7-2-2">
          <title>Many iteration s</title>
          <p>We note that choosing the best features improve substantially the model running time. For
our approach, we use just a basic version of PSO.</p>
          <p>The idea is first to see the impact of applying swarm intelligence algorithms on machine
learning algorithms, considering the results obtained from our previous publication.
3.4.</p>
        </sec>
      </sec>
      <sec id="sec-7-3">
        <title>Prediction process</title>
        <p>
          The aim of this step is to achieve the main objective, which is to predict whether a patient is
diabetic or not, using the RF classifier. Previous research [
          <xref ref-type="bibr" rid="ref4">4</xref>
          ] has convincingly confirmed
that RF has a better predictive accuracy. However, the evaluation of data collected from
patients and the decisions of experts is an essential factor in the prediction process at all.
        </p>
        <p>RF algorithm is a strong supervised machine-learning algorithm that is efficient of
performing both Regression and Classification tasks. RF is very popular in medicine
domain, through the application of RF in various healthcare applications, disease trends and
risks of the disease can be identified.</p>
        <p>In simple terms, RF works like this: we develop first different trees to classify a new
object, next, based on the produced attributes, each tree gives a different classification and
the tree votes for that class are saved. As a result, the forest chooses the classification with
the higher number of votes for all trees.</p>
        <p>During this step, we apply the RF classifier to the dataset, incorporating feature selection
using PSO. This process aims to classify each patient as either diabetic or non-diabetic.</p>
        <p>RF proceeds in two phases (see algorithm 1):
• The first phase involves creating the RF by combining N DT.</p>
        <p>• The second makes predictions for each created.</p>
        <p>At the end, we choose the class with the majority voting. The process at all is presented in
the following algorithm:</p>
      </sec>
    </sec>
    <sec id="sec-8">
      <title>Algorithm 1: Prediction process based on RF.</title>
      <p>Input: Dataset, K
Output: predicted class
Begin
1: Select random K data points from the dataset.
2: Build the decision trees associated with the selected data points
3: Choose the number N for DT to build
4: Go to 1 then 2
5: For new data points, find the predictions of each decision tree, and assign the new data points to the
class with the majority votes.</p>
      <p>End</p>
      <p>After implementation of this classification, we obtain class labels (0 or 1) for each of
row of our dataset. Therefore, the proposed system will deal with extra-values according to
the dataset format.</p>
      <sec id="sec-8-1">
        <title>4. Experiment and evaluation</title>
        <p>To evaluate and validate the proposed approach, we implemented a software tool in python
with the environment JupyterLab. JupyterLab [20] is an interactive development
environment for notebooks; it is very popular for data science, computing and artificial
intelligence.</p>
        <p>In this study, the main objective was to improve classification performance and increase
diagnostic accuracy by reducing feature dimension using PSO. This section discusses the
application of the RF-PSO Model on the dataset of the Frankfurt Hospital.</p>
        <p>For this purpose, various scenarios (table 2) were considered in order to evaluate the
proposed method. The distribution selection is arbitrary, and the idea is to see how the
system reacts with several distributions and avoid overfitting and underfitting in machine
learning.</p>
        <p>The distribution of data gives a better insight into the behavior of the built module. The
scenarios are presented as follows:</p>
        <p>Data training (%)</p>
        <p>Data testing (%)
80
70
Machine learning metrics are used to evaluate the performance and effectiveness of
RFPSO model [20]. These metrics provide insights into how well the models are performing
and help in comparing different models or tuning the model's parameters.</p>
        <p>The performance of the proposed method was evaluated using measures of accuracy
(formula 1), and precision (formula 2), based on the following terms: true positive (TP),
true negative (TN), false negative (FN) and false positive (FP). Precision is a well-known
metric in the prediction field, measuring the accuracy with which a model can predict class
membership. It therefore measures the quality of classification results. Accuracy is a metric
for evaluating the performance of classification models with 2 or more classes.</p>
        <p>These measures are calculated as follows:
Every time, the techniques were executed, the scenarios proposed were selected for each
run, thereby generating different results. To account for this variability, each proposal was
executed 2 times for every scenario, once for RF and once for RF-PSO.</p>
        <p>The table below shows the accuracy of the different experiments based on RF classifier
and PSO according to the three scenarios presented before.
The prediction column is the main column of interest and returns the result of the
prediction: positive indicates that the person is likely to be diabetic and negative means that
they are not. The accuracy column provides accuracy rate for each scenario. While the first
column presents the different scenario of splitting dataset.</p>
        <p>For a better analysis, we define a minimum accuracy value to achieve, which presents a
tolerance threshold 75% for results comparison.</p>
        <p>As Table 3 shows, all the accuracy results are higher than the threshold value, the
combination of RF and PSO delivers encouraging results in terms of accuracy according to
the three scenarios.</p>
        <p>After applying RF and PSO on dataset, we obtain precision and accuracies as mentioned
in table 4. The combination of RF and PSO gives highest accuracy of 90% with a 10%
increase regarding RF model. introducing the PSO has contributed to the performance of
the machine learning algorithm, as only attributes of interest to the machine learning
algorithm are selected.</p>
        <p>The We conclude that by using PSO for feature selection, we achieve significantly
improved accuracy compared to our previous approach and the proposed state-of-the-art
systems. However, data
especially with large datasets.</p>
        <p>quality has a significant influence on the obtained results,</p>
        <p>Results of accuracy</p>
        <p>Accuracy (%)
Random Forest</p>
        <p>Random Forest + PSO</p>
        <p>The performance comparison of the curve-based classification algorithms is shown in
Figure 3. The area under a curve grows as a curve's values increase, and the classifier
makes fewer errors as a result. The RF-PSO classifier performs 90% more accurately on
average than RF.</p>
        <p>However, there is a minor increase in precision, and this is due to the size of the data, so
with a larger dataset, the results will be better.</p>
        <p>Finally, we conclude from this study that the combination of the RF classification model
for diabetes prediction with the PSO algorithm performs better for the dataset than the RF
classifier. It is important to consider multiple datasets to have better accuracy for the
diabetes prediction.</p>
      </sec>
      <sec id="sec-8-2">
        <title>5. Conclusion</title>
        <p>In this paper, we have explored and analyzed the main related works for diabetes diagnosis.
In addition, we developed a hybrid model based on a combination of a swarm intelligence
technique and a machine-learning algorithm. This model revealed a highest accuracy.</p>
        <p>In this study, we used PSO and RF for feature selection and data classification,
respectively. The aim was to achieve the highest accuracy for the early prediction of
diabetes. Combining RF with PSO can be a powerful approach for solving various
machine-learning problems.</p>
        <p>To evaluate the model's performance, we used a dataset of 2000 patients gathered from
the hospital of Frankfurt in Germany and improved by 200 patients gathered from the khelil
Amrane hospital of Bejaia, Algeria.</p>
        <p>We developed a software tool in Python. The experiments were performed in two
distinct scenarios, The aim is to assess the contribution of the EHO algorithm to prediction
in terms of speed, quality, and accuracy.</p>
        <p>In the first scenario, we use only the RF classifier for diabetes prediction and in the
second scenario we use the RF-PSO model.</p>
        <p>The best results were obtained using RF-PSO model, which boosted the classification
accuracy. The proposed approach increases accuracy with 10.00%, compared with other
algorithms.</p>
        <p>Regarding the promising results, we claim that this approach could be applied to the
diagnosis of other diseases in different domains.</p>
        <p>The use of machine learning for medical is increasingly needed in the future. Especially
for the disease prediction, machine learning can speed up the process of diagnosing and
triage patients.</p>
        <p>In our future work, we intend to explore other swarm intelligence optimization
algorithms such as Elephant Herding Optimization (EHO) and Grey Wolf Optimization
(GWO) [22] with different dataset to assess the potential impact of the proposed approach.</p>
      </sec>
      <sec id="sec-8-3">
        <title>6. Acknowledgements</title>
        <p>The authors would like to express their gratitude to Dr. Djamel Eddine OUAIL, from the
Internal Medicine Service at Khelil Amrane Hospital in Bejaia city, Algeria, for his
unwavering support and encouragement since the beginning of this work.</p>
      </sec>
      <sec id="sec-8-4">
        <title>7. References</title>
        <p>[14] American Diabetes Association. URL:https://diabetes.org/diabetes/gestational-diabetes (Last
accessed: June 2023).
[15] S. Vivienne, C. Yu-Hsin, Y. Tien-Ju, J.S. Emer, Efficient processing of deep neural networks: A
tutorial and survey, Proceedings of the IEEE, 105 12 (2017) 2295-2329.
[16] Frankfurt Hospital. URL:https://www.kaggle.com/datasets/johndasilva/diabetes.
[17] Diabetics Data, Khelil Amran Hospital, Bejaia, Algeria. Collected July 2023.
[18] M. Jain, V. Saihjpal, N. Singh, S.B. Singh, S. B, An overview of variants and advancements of</p>
        <p>PSO algorithm. Applied Sciences, 12 17 (2022) 83-92.
[19] M. Xu, C. L, D. Lu, Z. Hu, Y. Yue, Application of Swarm Intelligence Optimization Algorithms
in Image Processing: A Comprehensive Review of Analysis, Synthesis, and Optimization,
Biomimetics (Basel), 8 2 (2023) 235. DOI:10.3390/biomimetics8020235.
[20] JupyterLab. URL: https://jupyter.org/.
[21] A. Mahyar, A. Rahmani, Machine learning process evaluating damage classification of
composites, International Journal of Science and Advanced Technology, 9 (2023) 240-250.
[22] A. Chakraborty, A.K. Kar, Swarm Intelligence: A Review of Algorithms, in: Patnaik, S., Yang,
XS., Nakamatsu, K. (eds) Nature-Inspired Computing and Optimization. Modeling and
Optimization in Science and Technologies, Springer, Cham, vol 10, 2017.
DOI:org/10.1007/9783-319-50920-4_19.</p>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <surname>Diabetes</surname>
          </string-name>
          . URL: https://www.who.
          <source>int/health-topics/diabetes (Last accessed: June</source>
          <year>2023</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>S.</given-names>
            <surname>Roshi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S. Sanjay</given-names>
            <surname>Kumar</surname>
          </string-name>
          ,
          <string-name>
            <surname>G.</surname>
          </string-name>
          <article-title>Manali, A Comprehensive review of various diabetic prediction models: a literature survey</article-title>
          .
          <source>Journal of Healthcare Engineering</source>
          ,
          <volume>8100697</volume>
          (
          <year>2022</year>
          ). DOI:
          <volume>10</volume>
          .1155/
          <year>2022</year>
          /8100697
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>R.</given-names>
            <surname>Bekka</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Kherbouche</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>El Bouhissi</surname>
          </string-name>
          ,
          <article-title>Distraction detection to predict vehicle crashes: a deep learning approach</article-title>
          , Computación y Sistemas,
          <volume>26 1</volume>
          (
          <year>2022</year>
          )
          <fpage>373</fpage>
          -
          <lpage>387</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>H.</given-names>
            <surname>El Bouhissi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R. E.</given-names>
            <surname>Al-Qutaish</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Ziane</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Amroun</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Yaya</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Lachi</surname>
          </string-name>
          ,
          <source>Towards Diabetes Mellitus Prediction Based on Machine-Learning, in: 2023 International Conference on Smart Computing and Application (ICSCA)</source>
          ,
          <year>2023</year>
          , pp.
          <fpage>1</fpage>
          -
          <lpage>6</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>E.</given-names>
            <surname>Goel</surname>
          </string-name>
          , Er. Abhilasha, Random Forest: A Review,
          <source>International Journal of Advanced Research in Computer Science and Software Engineering</source>
          ,
          <volume>7 1</volume>
          (
          <issue>2017</issue>
          )
          <fpage>251</fpage>
          -
          <lpage>257</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>E.</given-names>
            <surname>Ali</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>El-Sehiemy</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>El‐Ela</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Kamel</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Baseem</surname>
          </string-name>
          ,
          <article-title>Optimal planning of uncertain renewable energy sources in unbalanced distribution systems by a multi‐objective hybrid PSO-SCO algorithm</article-title>
          ,
          <source>IET Renewable Power Generation</source>
          ,
          <volume>16</volume>
          <fpage>10</fpage>
          (
          <year>2022</year>
          )
          <fpage>2111</fpage>
          -
          <lpage>2124</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>M.O.</given-names>
            <surname>Edeh</surname>
          </string-name>
          ,
          <string-name>
            <surname>O.I. Khalaf</surname>
          </string-name>
          , CA. Tavera,
          <string-name>
            <given-names>S.</given-names>
            <surname>Tayeb</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Ghouali</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.M.</given-names>
            <surname>Abdulsahib</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.E.</given-names>
            <surname>Richard-Nnabu</surname>
          </string-name>
          <string-name>
            <surname>NE</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Louni</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A Classification</given-names>
            <surname>Algorithm-Based Hybrid</surname>
          </string-name>
          Diabetes Prediction Model, Front. Public Health, (
          <year>2022</year>
          ). DOI:
          <volume>10</volume>
          .3389/fpubh.
          <year>2022</year>
          .
          <volume>82951</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>T.</given-names>
            <surname>Alam</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. Atif</given-names>
            <surname>Iqbal</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Ali</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Wahab</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Ijaz</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T. Imtiaz</given-names>
            <surname>Baig</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Hussain</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. Awais</given-names>
            <surname>Malik</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. Mehdi</given-names>
            <surname>Raza</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Ibrar</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Zunish</given-names>
            <surname>Abbas</surname>
          </string-name>
          ,
          <article-title>A model for early prediction of diabetes</article-title>
          ,
          <source>Informatics in Medicine Unlocked</source>
          ,
          <volume>16</volume>
          (
          <year>2019</year>
          ). DOI:
          <volume>10</volume>
          .1016/j.imu.
          <year>2019</year>
          .
          <volume>100204</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>A.</given-names>
            <surname>Minyechil</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Joshi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Alehegn</surname>
          </string-name>
          ,
          <article-title>Analysis and prediction of diabetes diseases using machine learning algorithm: Ensemble approach</article-title>
          ,
          <source>International Research Journal of Engineering and Technology</source>
          ,
          <volume>4</volume>
          <fpage>10</fpage>
          (
          <year>2017</year>
          )
          <fpage>426</fpage>
          -
          <lpage>436</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>P.</given-names>
            <surname>Sonar</surname>
          </string-name>
          ,
          <string-name>
            <surname>K.</surname>
          </string-name>
          <article-title>JayaMalini, Diabetes prediction using different machine learning approaches</article-title>
          ,
          <source>in: 3rd International Conference on Computing Methodologies and Communication (ICCMC)</source>
          ,
          <year>2019</year>
          , pp.
          <fpage>367</fpage>
          -
          <lpage>371</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>M.</given-names>
            <surname>Aishwarya</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Vaidehi</surname>
          </string-name>
          ,
          <article-title>Diabetes prediction using machine learning algorithms</article-title>
          ,
          <source>Procedia Computer Science</source>
          <volume>165</volume>
          (
          <year>2019</year>
          )
          <fpage>292</fpage>
          -
          <lpage>299</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>B. S.</given-names>
            <surname>Ahamed</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.S.</given-names>
            <surname>Meenakshi Sumeet</surname>
          </string-name>
          ,
          <string-name>
            <given-names>O.</given-names>
            <surname>Auxilia</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Nancy</surname>
          </string-name>
          ,
          <article-title>Prediction of type-2 diabetes mellitus disease using machine learning classifiers and techniques</article-title>
          ,
          <source>Frontiers in Computer Science</source>
          ,
          <volume>4</volume>
          (
          <year>2022</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <given-names>O.</given-names>
            <surname>Pavlova</surname>
          </string-name>
          ,
          <string-name>
            <given-names>O.</given-names>
            <surname>Soltyk</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Shvaiko</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Ilchyshyna</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>El Bouhissi</surname>
          </string-name>
          ,
          <article-title>Human Morphofunctional Indicators based Decision Support System for Choosing Kind of Sport</article-title>
          , MoMLeT+DS, (
          <year>2023</year>
          )
          <fpage>322</fpage>
          -
          <lpage>333</lpage>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>