<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <contrib-group>
        <aff id="aff0">
          <label>0</label>
          <institution>Extra Trees</institution>
          ,
          <addr-line>Gradient Boosting, MLP Classifier, model, importance</addr-line>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Lviv Polytechnic National University</institution>
          ,
          <addr-line>Stepan Bandera 12 79013 Lviv</addr-line>
          ,
          <country country="UA">Ukraine</country>
        </aff>
      </contrib-group>
      <fpage>0000</fpage>
      <lpage>0002</lpage>
      <abstract>
        <p>The rapid development of smart cities opens up new opportunities for improving road safety using predictive technologies. This article focuses on predicting road accidents in smart cities using big data, artificial intelligence (AI), and machine learning models. The paper analyzes a dataset of 45 features and about 8 million incidents, including factors such as time of the event, coordinates, distance of the road incident, city, region, zip code, time zone, temperature, airport, wind, humidity, pressure, visibility, precipitation, weather conditions, amenity, bump, junction, crossing, railway, roundabout, station, stop, period of day, and others. Different machine learning models, including Random Forest, Extreme Gradient Boosting, Gradient Boosting, Logistic Regression, Extra Tree, Decision Tree, MLP Classifier, and others, were evaluated for their prediction accuracy. The most effective model was Gradient Boosting, which achieved 85% accuracy while offering better interpretability. The study highlights the potential of AI and machine learning in traffic accident prediction, with Gradient Boosting offering the most effective solution due to its balance of accuracy and clarity. The research helps integrate predictive analytics into smart city infrastructure, improve road safety, and minimize the social and economic costs associated with road accidents. Future research should focus on incorporating real-time data streams from IoT-based systems and extending models that can be adapted to different cities, thereby improving the accuracy of predictions and extending the generalizability of results to the broader urban environment. This work contributes to developing safer and more efficient transportation systems as part of the evolving concept of smart cities.</p>
      </abstract>
      <kwd-group>
        <kwd>Road accident</kwd>
        <kwd>data</kwd>
        <kwd>dataset</kwd>
        <kwd>classification</kwd>
        <kwd>prediction</kwd>
        <kwd>feature</kwd>
        <kwd>Random Forest</kwd>
        <kwd>SMV</kwd>
        <kwd>Logistic Regression</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>With the development of technology, smart cities are becoming a reality, providing new opportunities
to improve road safety. One of the important tasks that can be solved with the help of intelligent
systems is the prediction of road traffic accidents (RTAs). Using big data, artificial intelligence (AI),
and analytical tools, we can create models that predict possible accidents, allowing preventive
measures to be taken in advance.</p>
      <p>Such solutions have the potential to significantly reduce the number of road accidents, minimize
medical costs, and improve the overall efficiency of the urban transportation system. This thesis
discusses modern approaches to traffic accident forecasting, including machine learning methods,
data analytics, and factors that influence the occurrence of accidents.</p>
      <p>The object of research is the intellectual transportation systems of smart cities, particularly their
ability to analyze and predict road traffic accidents (RTAs). In today's context of growing urbanization
and the increasing number of vehicles, road safety is becoming increasingly important, making it
necessary to find new solutions to reduce the number of road accidents.</p>
      <p>The subject of the study is methods and technologies for predicting road accidents in smart cities
using big data, data mining, artificial intelligence (AI), machine learning, and other analytical
approaches. The study of factors that affect the likelihood of accidents, as well as tools for their
prediction, is a key aspect of this research.</p>
      <p>The purpose of the research is to develop an effective intellectual model for predicting accidents
in smart cities, which will reduce the number of accidents by preventing risks in advance. To achieve
this goal, it is planned to apply modern methods of data analysis, use integrated traffic monitoring
systems, and identify key factors affecting road safety.</p>
      <p>
        A dataset consisting of 45 features and about 8 million incidents was used [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ] to predict traffic
accidents. These characteristics include a time of the event, coordinates, distance of the road incident,
city, region, zip code, time zone, temperature, airport, wind, humidity, pressure, visibility,
precipitation, weather conditions, amenity, bump, junction, crossing, railway, roundabout, station,
stop, period of day, and others [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ].
      </p>
      <p>
        In this work, was solved several diverse tasks that covered all stages of working with the dataset.
First, a detailed analysis of the dataset was conducted, including a review of its content, identification
of the main quantitative and qualitative characteristics, and study of possible types of these
characteristics. After that, statistical information about the data was collected, formatted it, and
checked for zero values. In cases where the amount of missing data exceeded 40%, a mechanism for
generating missing data was implemented [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ].
      </p>
      <p>Particular attention was paid to studying the number of different types of incidents depending on
several factors, such as region of the country, city, year, month, day of the week, and weather
conditions. The distribution of incidents by time of day, weather conditions, and duration of events
was analyzed in detail. In addition, a correlation matrix was constructed to examine the relationships
between the quantitative and qualitative characteristics of the dataset. Qualitative characteristics
were converted into quantitative ones by encoding them.</p>
      <p>
        The next step was to split the dataset into training and test samples in the ratio of 70% to 30%,
respectively. Next, various machine learning models were researched and developed that could be
used to predict the probability of traffic incidents. The models considered included Random Forest,
Extreme Gradient Boosting, Gradient Boosting, Logistic Regression, Extra Tree, Decision Tree, MLP
Classifier, and others [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ].
      </p>
    </sec>
    <sec id="sec-2">
      <title>2. Data preparation</title>
      <sec id="sec-2-1">
        <title>2.1. Source dataset</title>
        <p>
          The first step involves examining the original dataset to understand its structure, including the number
of observations, the characteristics it contains, and the target variable for prediction [
          <xref ref-type="bibr" rid="ref4">4</xref>
          ].
        </p>
        <p>
          In our case, the dataset includes a wide range of characteristics, such as time of incident, geographic
coordinates, distance to the incident, city, region, zip code, time zone, temperature, airport proximity,
wind, humidity, atmospheric pressure, visibility, precipitation, and various road and weather
conditions. It also includes attributes such as intersections, roundabouts, stations, and time of day. The
dataset consists of 45 characteristics and approximately 8 million recorded incidents, which are used
to predict traffic accidents [
          <xref ref-type="bibr" rid="ref6">6</xref>
          ].
        </p>
        <p>Now let’s take a closer look at the first 23 independent features of our dataset in more detail Table 1.</p>
        <sec id="sec-2-1-1">
          <title>Feature name Description</title>
          <p>Time The local time of the accident. Interval
Coordinate (Latitude, GPS coordinates of the accident Interval
Longitude) location.
X3</p>
          <p>As for the description of the target class labels, it is given below in Table 2.</p>
          <p>N
Y1</p>
          <p>Missing values check
Now, we need to check for blank values of the dataset features, and for this purpose, we can build bar
chart Figure 1.</p>
          <p>As we can see from the graph, for the End Latitude and End Longitude features, the percentage of
missed values is more than 40. Therefore, these missing values need to be filled with new ones using
one of the techniques, in our case, filling using the mean value.
2.2. Investigation of the Target Class (Severity)
To investigate target class, it is better to draw a graph of the distribution of the number of traffic
events by severity (Figure 2). The graph above shows the general distribution of incidents by severity.
It can be seen that there are 4 types of severity levels in total:</p>
        </sec>
        <sec id="sec-2-1-2">
          <title>1 - Least impact; 2 - Small impact; 3 - Moderate impact; 4 - Significant impact;</title>
          <p>It can also be seen (Figure 5) that the total number of incidents by severity is as follows:



</p>
          <p>Small impact equals 79.4% of all accidents in the dataset, which is 6129159 values.
Moderate impact is 16.8%, which is 1295336 records.</p>
          <p>Significant impact belongs to 2.8% of accidents which is only 203120 of all.</p>
          <p>Only 67066 accidents have the least impact, which is almost 1%.
Let’s build a country figure that will represent the top 10 states with the highest number of traffic
incidents.</p>
          <p>Figure 3 shows that the states with the highest number of incidents are shown in blue, and the
states with the lowest number of incidents are shown in white. So, the top 3 states with the highest
number of incidents are:</p>
          <p>The graph above (Figure 4) shows the top 10 states by incidents in the form of a bar chart. It shows
that incidents occurred most frequently in the following states: California, Florida, Texas, New York
City, South Carolina, New York, Oregon, Virginia, Pennsylvania, and Illinois.</p>
          <p>The next step is to build a graph of incidents in the states depending on their severity. To do this,
let’s divide the total number of incidents in the state into four parts. Namely, incidents with the least
(blue), small (green), moderate (red), and significant (purple) impact on traffic.</p>
          <p>As can be seen from the Figure 5 above, the distribution of incident severity across the states is
uneven, with the number of small-impact incidents being much higher than the other types.</p>
          <p>However, the distribution is the same for all ten states, with the highest number of:



small-impact incidents, followed by
moderate-impact,
significant-impact, and
</p>
          <p>Now let’s look at the mean severity of incidents depending on weather conditions. To do this, we
need to group the data by two characteristics: weather conditions and incident severity. And after that,
let’s draw Figure 6. This graph shows that:


</p>
          <p>The worst severity of an incident, namely an incident with significant impact, is typical for a
weather condition such as Light Blowing Snow.</p>
          <p>For moderate-impact incidents, the weather conditions are usually as follows: Patches of Fog
/ Windy, Light Fog, Partial Fog / Windy, Heavy Freezing Rain / Windy.</p>
          <p>The lightest severity of the incident is typical for such weather conditions as: Low Drifting
Snow, and Heavy Rain Showers.</p>
          <p>The graphic above (Figure 7) shows the total duration of each accident depending on its severity.
It shows that the more complex the incident, the longer it takes to resolve. For example, it takes only
about 0.6 hours to resolve a least impact accident, and about 1 day to resolve a significant accident.
2.4. Correlation Matrix</p>
          <p>In addition, it is important to create a correlation matrix to get a clearer understanding of how
different characteristics are related to each other and influence each other. This type of chart allows
you to observe the relationships between various variables. The resulting correlation matrixes that
visually represent these relationships are shown in Figure 8, and Figure 9 below.</p>
          <p>Temperature (F), and Start Latitude;
Humidity (%), and Temperature (F);</p>
          <p>Visibility (mi), and Humidity (%).</p>
        </sec>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>3. Classification</title>
      <p>3.1.</p>
      <sec id="sec-3-1">
        <title>Classifiers Types</title>
        <p>A variety of machine learning algorithms and models were used to predict the probability of road
accidents. These algorithms were selected based on their unique capabilities and strengths in
performing the classification task. These classifiers are as follows: extreme gradient boosting
(xgboost), light gradient boosting machine (lightgbm), gradient boosting classifier (gbc), random
forest (rf), extra trees (et), logistic regression (lr), ridge classifier, dummy classifier, adaboost (Ida), k
- nearest neighbors (knn), decision tree (dt), and SVM with a linear kernel (Table 3).</p>
        <p>N
xgboost
lightgbm
gbc
rf
et
ada</p>
        <p>Ir
ridge
dummy</p>
        <p>
          Ida
svm
knn
dt
NB
[
          <xref ref-type="bibr" rid="ref7">7</xref>
          ]
        </p>
        <p>Before starting the model building process, it is important to divide the dataset into two separate
parts: one for training and one for testing. This separation ensures that the performance of the model
can be properly evaluated. Given that the dataset contains a large number of records, it was decided
to select only 1% of the total data to prevent the risk of overfitting. Once the dataset was reduced, the
next step was to split the data for modeling into training and test sets, as shown in Figure 10.</p>
        <p>Specifically, 70% of the data was allocated for training the model, allowing it to learn on a
significant portion of the dataset, while the remaining 30% was reserved for testing. This split ensures
that the model can be tested on data it has not seen before, allowing for a more accurate assessment
of its predictive capabilities.
3.3.</p>
      </sec>
      <sec id="sec-3-2">
        <title>Standard Classifier</title>
        <p>Figure 11 below shows a typical model building and training process using standard classification
algorithms. The diagram illustrates the steps involved in building and training a model and
emphasizes the iterative nature of the process, where the classifier was run ten times. After
completing these ten runs, the average accuracy achieved by the model is approximately 85%, which
is a good indicator of its performance.</p>
        <p>The classification report generated as part of the evaluation includes several important metrics
that provide a detailed understanding of the model's performance. These metrics are as follows:







</p>
        <p>Fold - refers to the breakdown of data during cross-validation.</p>
        <p>The main attribute of a classification report is Accuracy. It is measured as the percentage of
correctly predicted cases out of the total number of prognoses.</p>
        <p>The next indicator is the receiver operating characteristic curve, in other words, ROC curve.
It is measured by the area under the curve, which is estimated as the model’s ability to
recognize the existing classes.</p>
        <p>Sensitivity or Recall is a measure of the rate of actual positive cases that were correctly
recognized.</p>
        <p>The relation of correctly predicted positive observations to the number of predicted positive
observations is often referred to as Precision.</p>
        <p>The mean value of accuracy and recall, which provides a balance between these two
indicators, is F1 score.</p>
        <p>Cohen’s kappa – statistic that measures the consistency between annotators, taking into
account the possibility of random agreement [21].</p>
        <p>Matthew’s correlation coefficient (MCC) a balanced metric that considers true and false
positive and negative responses, providing a comprehensive assessment of binary
classifications [22].</p>
        <p>Together, these metrics provide a comprehensive view of the model’s performance in various
aspects, providing a comprehensive evaluation.</p>
        <p>Model Comparison and Result Analysis</p>
        <p>In reference to Figure 12, it can be seen that the Extreme Gradient Boosting classification model
performed best in this analysis, achieving an 85% accuracy rate. It is followed by Light Gradient
Boosting with 84% accuracy and Gradient Boosting, which showed a good 83% of accuracy.
Conversely, the model that performed the worst in this particular task was Support Vector Machine
(SVM), which recorded a relatively low prediction performance of only 51%. In addition, Naїve Bayes
performed even worse, achieving only 20% accuracy, and the Quadratic Discriminant Model
performed terribly, showing only 2% accuracy.</p>
        <p>Moreover, other classification quality metrics confirmed this assessment and produced consistent
results similar to those illustrated by the ROC curve, Precision, Recall, F1, Cohen’s kappa, and MCC.
In the end, it is clear that the best model in this analysis is the Extreme Gradient Boosting classifier
(xgboost), which provided a robust classification accuracy of 85%, as shown in Figure 11.</p>
        <p>The last step of the research is to carefully study the importance of the features in the dataset.
Here, we need to focus on the features and how they affect the performance of different classifiers
and the overall severity of the accidents. To do this, we can plot the importance of the feature
permutation. This graph is a good tool to understand the behavior of the model in machine learning.</p>
        <p>The Feature permutation importance [23] gives a general idea of how the model makes decisions,
namely, it allows to evaluate the contribution of individual features to the model's classification
efficiency. This graph allows you to effectively group the importance of each feature and assess its
impact on the model's classification efficiency. By studying these relationships, we can better
understand which features are the most impactful and how they can be optimized to improve
classification accuracy. This step is important to increase the reliability of the model and ensure that
it accurately reflects the factors that trigger road accidents.</p>
        <p>The Figure 13 above shows a histogram that displays the importance of different features for an
Extreme gradient boosting model (xgboost). The importance of each feature is represented by the
length of the corresponding column, and the features are sorted in descending order of importance.</p>
        <p>The most important features, according to this chart, are:
1. Wind Cloud (highest importance);
2. Traffic Signal;
3. Stop;
4. Crossing;
5. Weather Clear;
6. Duration;
7. Weather Fair;
8. Junction;
9. Wind Speed (mph);
10. Give Way.</p>
        <p>On the other hand, features such as “Weather Tornado”, “Weather Dust”, and “Weather Hail”
have very low importance because they have minimal impact on the model’s prediction.</p>
        <p>Conclusions
This work demonstrated the potential for predicting traffic accidents in smart cities using big data
models and machine learning, with Gradient Boosting proving to be the most effective approach due
to its high accuracy (85%) and interpretability. By analyzing a dataset of 45 characteristics and
approximately 8 million incidents, the study identified key factors that influence road accidents, such
as time, place, and weather conditions, and emphasized the importance of data preprocessing to
ensure reliable results.</p>
        <p>The study showed that predictive models can significantly improve road safety in cities by
enabling proactive measures such as adjusting traffic signals and warning of high-risk conditions.
Although the study showed promising results, future work should focus on improving the models
with real-time data and incorporating additional sources, such as IoT-based traffic monitoring
systems, to improve the accuracy of prediction and generalization across different smart cities.
Overall, the research contributes to the development of safer and more efficient urban transportation
systems.</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>Acknowledgements</title>
      <p>This work was realized within the framework of the Erasmus+ Jean Monnet Chair 2022 «101085612
– DataProEU – Data Protection in the EU».
[11] Yazhi Gao, W. Rong, Y. Shen and Z. Xiong, "Convolutional Neural Network based sentiment
analysis using Adaboost combination," 2016 International Joint Conference on Neural Networks
(IJCNN), Vancouver, BC, Canada, 2016, pp. 1333-1338, doi: 10.1109/IJCNN.2016.7727352.
[12] Thorn J. Logistic Regression Explained. Medium. URL:
https://towardsdatascience.com/logisticregression-explained-9ee73cede081.
[13] H. Luo and Y. Liu, "A prediction method based on improved ridge regression," 2017 8th IEEE
International Conference on Software Engineering and Service Science (ICSESS), Beijing, China,
2017, pp. 596-599, doi: 10.1109/ICSESS.2017.8342986.
[14] Tezcan B. Why Using a Dummy Classifier is a Smart Move. Medium. URL:
https://towardsdatascience.com/why-using-a-dummy-classifier-is-a-smart-move-4a55080e3549.
[15] J. Ghosh and S. B. Shuvo, "Improving Classification Model's Performance Using Linear
Discriminant Analysis on Linear Data," 2019 10th International Conference on Computing,
Communication and Networking Technologies (ICCCNT), Kanpur, India, 2019, pp. 1-5, doi:
10.1109/ICCCNT45670.2019.8944632.
[16] José Luis Rojo-Álvarez; Manel Martínez-Ramón; Jordi Muñoz-Marí; Gustau Camps-Valls,
"Support Vector Machine and Kernel Classification Algorithms," in Digital Signal Processing
with Kernel Methods, IEEE, 2018, pp.433-502, doi: 10.1002/9781118705810.ch10.
[17] Christopher A. K-Nearest Neighbor. Medium. URL:
https://medium.com/swlh/k-nearestneighbor-ca2593d7a3c4.
[18] What is a Decision Tree | IBM. IBM - Deutschland | IBM. URL:
https://www.ibm.com/topics/decision-trees.
[19] Yıldırım S. Naive Bayes Classifier – Explained. Medium. URL:
https://towardsdatascience.com/naive-bayes-classifier-explained-50f9723571ed (date of access:
02.09.2024).
[20] E. Pȩkalska and B. Haasdonk, "Kernel Discriminant Analysis for Positive Definite and Indefinite
Kernels," in IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 31, no. 6, pp.
1017-1032, June 2009, doi: 10.1109/TPAMI.2008.290.
[21] Performance Measures: Cohen's Kappa statistic. The Data Scientist. URL:
https://thedatascientist.com/performance-measures-cohens-kappa-statistic/ (date of access:
02.09.2024).
[22] sklearn.metrics.matthews_corrcoef. scikit-learn. URL:
https://scikitlearn.org/stable/modules/generated/sklearn.metrics.matthews_corrcoef.html (date of access:
02.09.2024).
[23] Feature Permutation Importance Explanations – ADS 1.0.0 documentation. Moved. URL:
https://docs.oracle.com/en-us/iaas/tools/adssdk/latest/user_guide/mlx/permutation_importance.html (date of access: 02.02.2024).</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>US</given-names>
            <surname>Accidents</surname>
          </string-name>
          (
          <year>2016</year>
          -
          <fpage>2023</fpage>
          ).
          <article-title>Kaggle: Your Machine Learning and Data Science Community</article-title>
          . https://www.kaggle.com/datasets/sobhanmoosavi/us-accidents
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <surname>Doroshenko</surname>
            ,
            <given-names>Anastasiya.</given-names>
          </string-name>
          (
          <year>2019</year>
          ).
          <article-title>Application of global optimization methods to increase the accuracy of classification in the data mining tasks</article-title>
          .
          <source>Computer Modeling and Intelligent Systems</source>
          ,
          <volume>2353</volume>
          ,
          <fpage>98</fpage>
          -
          <lpage>109</lpage>
          . https://doi.org/10.32782/cmis/2353-
          <fpage>8</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <surname>Savchuk</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Doroshenko</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          (
          <year>2021</year>
          ).
          <source>Investigation of Machine Learning Classification Methods Effectiveness</source>
          .
          <source>2021 IEEE 16th International Conference on Computer Sciences and Information Technologies (CSIT)</source>
          . https://doi.org/10.1109/csit52700.
          <year>2021</year>
          .9648582
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>A.</given-names>
            <surname>Batyuk</surname>
          </string-name>
          and
          <string-name>
            <given-names>V.</given-names>
            <surname>Voityshyn</surname>
          </string-name>
          ,
          <article-title>"Streaming Process Discovery Method for Semi-Structured Business Processes,"</article-title>
          <source>2020 IEEE Third International Conference on Data Stream Mining &amp; Processing (DSMP)</source>
          , Lviv, Ukraine,
          <year>2020</year>
          , pp.
          <fpage>444</fpage>
          -
          <lpage>448</lpage>
          , doi: 10.1109/DSMP47368.
          <year>2020</year>
          .
          <volume>9204201</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <surname>Alshamrani</surname>
            ,
            <given-names>F. H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Syed</surname>
            ,
            <given-names>H. F.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>Elhussein</surname>
            ,
            <given-names>M. A.</given-names>
          </string-name>
          (
          <year>2019</year>
          ).
          <article-title>Machine learning based model for traffic prediction in Smart Cities</article-title>
          .
          <source>2nd Smart Cities Symposium (SCS</source>
          <year>2019</year>
          ). https://doi.org/10.1049/cp.
          <year>2019</year>
          .0195
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <surname>Nagy</surname>
            ,
            <given-names>A. M.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>Simon</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          (
          <year>2018</year>
          ).
          <article-title>Survey on traffic prediction in Smart Cities</article-title>
          .
          <source>Pervasive and Mobile Computing</source>
          ,
          <volume>50</volume>
          ,
          <fpage>148</fpage>
          -
          <lpage>163</lpage>
          . https://doi.org/10.1016/j.pmcj.
          <year>2018</year>
          .
          <volume>07</volume>
          .004
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <surname>Obelovska</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Snaichuk</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Selecky</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Liskevych</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Valkova</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          <article-title>An Approach Toward Packet Routing in the OSPF-based Network with a Distrustful Router WSEAS Transactions on Information Science</article-title>
          and Applications,
          <year>2023</year>
          ,
          <volume>20</volume>
          , pp.
          <fpage>432</fpage>
          -
          <lpage>443</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>Gradient</given-names>
            <surname>Boosting. WallStreetMojo</surname>
          </string-name>
          . URL: https://www.wallstreetmojo.com/gradient-boosting/.
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          <article-title>[9] Introduction to Random Forest in Machine Learning. Engineering Education (EngEd) Program | Section</article-title>
          . URL: https://www.section.io/engineering-education/
          <article-title>introduction-to-random-forest-inmachine-learning/.</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <article-title>How to Develop an Extra Trees Ensemble with Python - MachineLearningMastery.com</article-title>
          . MachineLearningMastery.com. URL: https://machinelearningmastery.com
          <article-title>/extra-treesensemble-with-python/.</article-title>
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>