<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Information system module for analysis viral infections data based on machine learning</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Nickolay Rudnichenko</string-name>
          <email>nickolay.rud@gmail.com</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Vladimir Vychuzhanin</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Tetiana Otradskya</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Igor Petrov</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>National University “Odessa Maritime Academy”</institution>
          ,
          <addr-line>8 Didrichson Str., Odesa, 65029</addr-line>
          ,
          <country country="UA">Ukraine</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Odesa Polytechnic National University</institution>
          ,
          <addr-line>1 Shevchenko Ave., Odesa, 65001</addr-line>
          ,
          <country country="UA">Ukraine</country>
        </aff>
      </contrib-group>
      <fpage>63</fpage>
      <lpage>74</lpage>
      <abstract>
        <p>The article presents results of the development information system module for analysis viral infections data. The relevance of the problem of automating the process of analyzing large volumes of data based on the use of intelligent technologies and machine learning methods is considered. The structure of the system has been developed and described, the results of design modeling of the key functionality and capabilities of the system based on the use of the UML language are presented, the basic components and technologies for implementing software are described, allowing for modularity and dynamic expandability of the potential for conducting data analysis research. The process of creating, training and testing the created machine learning models is detailed, the results of assessing the significance of the input features of the collected data set on viral diseases and the obtained values of the error matrices are described. The profiling of the operation process of the created models was carried out, the most productive and eficient of them were determined in terms of the consumption of computing resources and overall accuracy, taking into account their generalization ability.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;data analysis</kwd>
        <kwd>data visualization</kwd>
        <kwd>machine learning</kwd>
        <kwd>viral infections</kwd>
        <kwd>information systems development</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>
        predicting future values using machine learning (ML). In the context of data analysis, one of
the current research directions is the issue of viral diseases [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ].
      </p>
      <p>
        Although COVID-19 has become less critical due to regular vaccinations, infection prevention,
and the use of preventive measures such as quarantines, viruses have not disappeared from the
planet and continue to mutate [
        <xref ref-type="bibr" rid="ref5 ref6 ref7">5, 6, 7</xref>
        ].
      </p>
      <p>
        Identifying disease patterns, forecasting, and other analyses can help prevent the further
spread of the virus. Automation of symptom detection and human body reactions to viruses is
essential, as it enables faster modification of efective vaccines and the introduction of various
protective measures into the virus transmission environment [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ].
      </p>
      <p>The relevance of this work lies in the search for existing methods of data analysis, for their
further application to the data visualization and the identification of various patterns based on
obtained graphical representations.</p>
      <p>The goal of this work is to investigate the dependencies of various indicators in virology
using data analysis models, creating a cross-platform information system module for practical
purposes.</p>
      <p>
        Assessing and automating data analysis for various diseases can help identify the source
of the virus, its symptoms, and prevent its further spread. Collecting such information is a
challenging aspect due to the need for access to medical databases or conducting surveys of
infected individuals. This complexity reduces the speed of information updates, and the human
factor may compromise data reliability, especially when dealing with a significant number of
input features [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ].
      </p>
      <p>
        To solve this task efectively, it is important to use the right methods and algorithms during
data collection and preprocessing. Specifically, among contemporary approaches in this
direction, dimensionality reduction methods are relevant to facilitate the transformation of data into
a form suitable for sequential analysis and interpretation of results [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ].
      </p>
      <p>
        The eficiency of data analysis for such data largely depends on the models used and their
hyperparameter values [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ]. Some of the less deterministic ML models may perform data
analysis quickly but produce inaccurate values and lack suficient generalization capability
[
        <xref ref-type="bibr" rid="ref12">12</xref>
        ]. On the other hand, more stochastic models may take longer to perform computational
operations, but their precision will be higher as a result [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ].
      </p>
      <p>
        For both cases, setting the right attributes for the models and their relationships with each
other is a critical factor. This can reduce the likelihood of full-scale quarantines and expedite
the development of vaccine modifications [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ].
      </p>
    </sec>
    <sec id="sec-2">
      <title>2. Results</title>
      <sec id="sec-2-1">
        <title>2.1. Project structure</title>
        <p>During the system development process, UML modeling language was used. The system is built
using a client-server architecture and includes a relational database (DB) for storing data from
experiments. The developed use case diagram for the system’s operation is shown in figure 1.</p>
        <p>Within the scope of this diagram, the user can perform the following actions:
• Authenticate in the system using personal data by verifying requests from the database;</p>
        <p>• Import datasets into the system for further analysis;
• Visualize processed data in the form of diagrams, graphs, tables, and textual information;
• Modify arguments and parameters of ML models during their training;
• Register in the system to gain access to data analysis tracking functions.</p>
        <p>Let’s formalize the main sequence of system operation by applying a sequence diagram
for this purpose. For demonstration, we’ll take a user who uploads a dataset for conducting
exploratory analysis by creating graphical visualizations of static data calculation results (in the
form of a small image with statistics and a report with obtained results in a separate file). The
user uploads data, selects ML models for analysis, and can set parameter values. The sequence
diagram is shown in figure 2.</p>
        <p>The user’s first action is to access the data analysis model selection form for loading features
with parameter editing visualization. After that, they can upload a new dataset in the form of a
CSV file or use one of the previously imported data sets serialized to the database. The system
user confirms their actions, after which the system dynamically changes the parameter values
in the form’s interface, depending on which data the user wants to input.</p>
        <p>Next, the user fills in all the necessary parameters and data processing algorithms, after
which the generated data is sent from the form to the system’s handler. Using a REST request,
the server-side of the system processes and manages the data to perform further control actions.</p>
        <p>In the case of a POST request, the server activates the process of building an ML model based
on the specified parameters. The process of using models, in general terms, is as follows:
• Data processing involves receiving input data that will be further processed. Afterward,
the system selects software methods for their processing;</p>
        <p>• Creating an instance of an ML model, all imported and processed data are then passed on,
and they are processed internally using the fit function;
• Output of a concise report on the evaluation of the performance metrics of the ML model.</p>
        <p>After generating the ML model and processing the data, they need to be sent to the system’s
data visualization module, where, through a GET request, images and other graphical elements
specified by the user for the report are displayed. For more detailed planning and formalization
of the main relationships between the system components and their structure, it is advisable to
construct a diagram of its main components. A generalized component diagram is provided in
ifgure 3.</p>
        <p>The main module of the system is flsite. Given that all system modules are developed in
Python using the Flask framework, flsite is responsible for the operability and integration of
the entire system’s functionality. It utilizes the Flask library, allowing the system to be used in
the form of a web application, which is convenient due to its compactness and speed.</p>
        <p>lfsite also contains a set of REST requests that are sent to or received by the system during
its operation for rendering authorization pages, user profiles, and model management.</p>
        <p>The user profile page contains information about the results of analytical experiments
performed by users and datasets uploaded by them. This is necessary for repeating or correcting
data analysis.</p>
        <p>On the authorization page, the user needs to enter login and password data, as well as
a personal token generated by the system after the registration process through a separate
modular window. Additionally, for introductory actions in working with the system, the user
can choose a limited version in which they don’t need to enter authentication data, but they
will have restrictions in terms of uploading their data; in other words, they will be able to work
with only preinstalled test datasets.</p>
        <p>The scikit-learn library is used for the programmatic implementation of ML models due to its
convenience, good documentation, and integration with the Python language. The advantage
of using this dependency is its support for a significant set of objects for conducting ML model
research, performing statistical modeling, including classification, regression, clustering, and
dimensionality reduction, through a sequential Python interface. Additionally, libraries such as
Pandas, NumPy, SciPy, and Matplotlib were applied for data processing and visualization.</p>
        <p>It’s worth noting the preprocess.py class, responsible for obtaining input model values and
their transformations. This is implemented by integrating functionality from the
imbalancedlearn library, which is based on scikit-learn and provides a range of objects for simplified and
fast classification work in cases of class imbalance detection.</p>
        <p>The utils.py class is responsible for retrieving data from the database, writing them to a
ifle, generating graphical visualizations for reports, transforming data structures and data into
matrices of various dimensions, and a variety of other utility functions.</p>
        <p>We use SQLite to create the database, which comes bundled with Python3. Its convenience is
due to its implementation of support for autonomous and transactional relational mechanisms.
The relationship between the database and the system is depicted in figure 4.</p>
        <p>First, a database with a set of tables is created separately (without launching the web server).
Based on this, request handlers access the construction of database tables, reading and writing
information to them. Afterward, the process of creating a general function for establishing
a connection to the database and an auxiliary function that will initiate the construction
process of a relational database model with the required tables is carried out. At this stage,
the open_resource programmatic method is introduced, which opens the ‘sq_db.sql’ file for
reading, located in the project’s system working directory. Then, for the open database, the
script contained in the ‘sq_db.sql’ file is executed. Finally, the commit method is called to apply
the changes to the current database, and the close method closes the established connection.</p>
      </sec>
      <sec id="sec-2-2">
        <title>2.2. ML models implementation</title>
        <p>As part of using the system’s data analysis module, the following 5 ML models have been
programmatically implemented: Gaussian Naive Bayes (GNB), Decision Tree, Random Forest,
SVM, and Neural Network (NN) based on a perceptron model. In the system’s interface, these
models can be assigned various hyperparameter values, allowing clients to input data manually
from the web page through corresponding fields. The steps of implementing the data analysis
process are as follows:
• Importing input dependencies and libraries.
• Fetching input data into the working environment by downloading them via links.
• Splitting data into training and testing subsets. To perform this step, a cross-validation
splitter has been defined, which divides the data into training/test sets according to a
specified scheme. Each sample can be assigned to no more than one fold of the test set, as
indicated by the user using the test_fold parameter.
• Creating ML model objects.
• Tuning hyperparameter values of ML models using the GridSearchCV method for
optimal selection. This is essentially an object implementing a systematic approach to
hyperparameter optimization, which involves defining a grid of possible values for each
hyperparameter and then exhaustively searching this grid to find the best combination.
This approach is called grid search because it involves creating a grid of hyperparameter
values and evaluating the algorithm’s performance for each combination of values in the
grid. It’s worth noting that there’s no way to know the best hyperparameter values in
advance, so ideally, you should try all possible values to find the optimal ones.
• Doing this manually can be time-consuming and resource-intensive, so we use
GridSearchCV for automating hyperparameter tuning. The following models have similar
code but diferent parameters within.
• Training models and assessing their accuracy metrics based on the use of methods
such as classification_report, taking into account particularities including permutations
(permutation_importance). Permutation importance is defined as the reduction in model
score when the values of a single feature are randomly shufled. This procedure breaks
the relationship between the feature and the target variable, so a drop in the model’s score
indicates how much the model depends on that feature. One advantage of this method is
that it’s model-agnostic and can be calculated many times with diferent feature shufles.</p>
        <p>At this stage, we obtain a finished model for further interaction.</p>
      </sec>
      <sec id="sec-2-3">
        <title>2.3. Data analysis</title>
        <p>In order to conduct research, a dataset was created based on information from the European
Centre for Disease Prevention and Control regarding the most common COVID-19 symptoms in
2020. The dataset contains information about patients who underwent COVID-19 testing, with
a total of 5434 rows and 211 columns. Each row corresponds to a patient’s symptoms, where
each column represents a symptom. To properly preprocess the data, a dimensionality reduction
procedure was performed using PCA. Some features, such as ‘had contact with someone with
the virus” or “recently traveled”, were either removed or combined into more meaningful ones.
This allows us to focus on significant symptoms and identify which symptoms indicate early
stages of COVID-19.</p>
        <p>Out of 5434 patients, 81% later tested positive for the virus. To address the dataset’s imbalance,
the imblearn library was used, which helps minimize the dataset’s class imbalance. We chose
the resampling method, which aggregates repetitions and reduces data imbalance. The dataset’s
target variable is expressed in a binary form, with two values: “yes” and “no”. Essentially, this
allows us to frame the task as binary classification. The data was split into training, validation,
and testing sets, with 60%, 20%, and 20% of the data in each part, respectively. Below is an
example of the data processing results and model creation with the described datasets.</p>
        <p>The results of building the Gaussian Naive Bayes model are shown in figure 5.</p>
        <p>As we can see, the top three symptoms for the Naive Bayes model are Dry cough, Sore Throat,
and Breathing Problems. Fever was close to the third position, but it’s worth noting that its
weight is significantly higher than other symptoms. As observed from the matrix, there are
77 False Negatives (FN) and 80 False Positives (FP). Considering that GNB is a basic model,
its accuracy is quite high. It’s evident that we aim to have a stronger model, and we hope to
achieve this with better models ahead. In terms of metrics, the F1 score for negative cases is
0.824, while the F1 score for positive cases is 0.910.</p>
        <p>Let’s examine the results of the Decision Tree model (figure 6).</p>
        <p>As we can see, the top three symptoms for the Decision Tree model are Breathing Problems,
Sore Throat, and Dry Cough. Once again, Fever was close to the third position, but, unlike GNB,
other symptoms have more weight. In this case, we used the output of the feature_importances
from the Decision Tree. There is an interesting shift in the importance of Breathing Problems
and Sore Throat. After analyzing the data, we observe an increase in the importance of Sore</p>
        <p>Throat, while Breathing Problems decrease. This suggests that Sore Throat is much more
important than Breathing Problems. As we can see, the Decision Tree works much better than
GNB. Previously, there were 77 FN and 80 FP, and now there are 16 FN and 10 FP. In terms of
metrics, the F1 score for negative cases is 0.969, and the F1 score for positive cases is 0.985.</p>
        <p>The most important features for the Decision Tree are Breathing Problems, Sore Throat, and
Dry Cough. Currently, the symptoms are consistent with each other in every model we have
built. Next, we will conduct research and analyze the Random Forest models (figure 7).</p>
        <p>As we can see, the top three symptoms for the scaled Random Forest model are Dry Cough,
Breathing Problems, and Sore Throat.</p>
        <p>Overall, the Random Forest model produces results similar to the Decision Tree. Previously,
there were 16 FN and 10 FP, and we still have 16 FN and 10 FP. In terms of metrics, the F1 score
for negative cases is 0.968, while the F1 score for positive cases is 0.985.</p>
        <p>The most important features for the Random Forest are Dry Cough, Breathing Problems, and
Sore Throat. Currently, the symptoms are consistent in every built and analyzed model.</p>
        <p>Next, let’s analyze the performance of the Support Vector Machine (SVM) model. The
demonstration of SVM is shown in figure 8. In our case, SVM outputs a hyperplane that separates
both classes. The weight coeficients, represented by coef, form a hyperplane orthogonal to the
original division boundary. If the hyperplane finds a feature useful for separating the data, the
plane will be orthogonal to that axis. Therefore, the coeficient shows how important it was in
dividing the two datasets.</p>
        <p>We will move on to using more complex kernels (rbf, poly) and use permutation importance
to compute feature importance (since coef_ is valid only for linear kernels). As we can see, the
top features are Breathing Problems, Dry Cough, and Sore Throat.</p>
        <p>In our case, using the RBF kernel works faster and provides a more balanced analysis of
feature importance.</p>
        <p>Let’s conduct an investigation of the created neural network model (figure 9). For the neural
network, we compare functions using permutation importance.</p>
        <p>From the parameter grid, we used a validation set to find the best combination. For the
activation function, we used ReLU, the number of hidden layers is 80, and we utilized the Adam
optimizer for better performance.</p>
        <p>We wanted to understand whether our network prefers having one larger bank of hidden
layers or several smaller banks of hidden layers. The artificial neural network model achieved
the most significant results with only 6 cases of TN errors. The most important features are
“breathing problems”, “dry cough”, and “sore throat”, depending on the dataset used.</p>
        <p>Based on the results obtained, let’s analyze the performance speed of all the ML models
discussed above. The speed of the two neural networks (ReLU and Adam) is presented through
profiling in figure 10. In this case, we can see that GNB and DT have the minimum operating
speed. RF, due to its complexity, operates more slowly. SVM and NN1 work significantly faster,
even with input value analysis.</p>
        <p>Therefore, the developed module is suficiently functional and allows for the exploration of
datasets to uncover hidden patterns within the data.</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>3. Summary</title>
      <p>As a result of this research, a software module for the system has been developed, capable of
performing data preprocessing and analysis steps related to human virus-related illnesses using
machine learning methods. The following evaluation criteria for the models were determined:
speed, accuracy, predictions, best hyperparameters for the models, error matrix assessments.</p>
      <p>The research findings have established that the most efective models are artificial neural
networks, but decision trees also showed one of the best results, considering that this model
is based on a basic algorithm. The study also revealed the varying speeds of diferent models,
with a significant workload dependence on the chosen hyperparameter values. Additionally,
data inaccuracies can complicate the process of determining the best model.</p>
      <p>Therefore, it can be concluded that the best model for working with the data collected during
the development of the system is the neural network. However, its speed is significantly lower
than that of the decision tree model. Therefore, future research in this area could involve
exploring more eficient models and forming a more extensive dataset for a more balanced
analysis.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>C.</given-names>
            <surname>Peng</surname>
          </string-name>
          ,
          <article-title>Digital Inclusive Finance Data Mining and Model-Driven Analysis of the Impact of Urban-Rural Income Gap</article-title>
          ,
          <source>Wireless Communications and Mobile Computing</source>
          <volume>7</volume>
          (
          <year>2022</year>
          )
          <fpage>1</fpage>
          -
          <lpage>8</lpage>
          . doi:
          <volume>10</volume>
          .1155/
          <year>2022</year>
          /5820145.
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>N.</given-names>
            <surname>Rudnichenko</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Vychuzhanin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>I.</given-names>
            <surname>Petrov</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Shibaev</surname>
          </string-name>
          ,
          <article-title>Decision Support System for the Machine Learning Methods Selection in Big Data Mining</article-title>
          ,
          <source>CEUR Workshop Proceedings</source>
          <volume>2608</volume>
          (
          <year>2020</year>
          )
          <fpage>872</fpage>
          -
          <lpage>885</lpage>
          . URL: https://ceur-ws.
          <source>org/</source>
          Vol-
          <volume>2608</volume>
          /paper65.pdf.
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>I. M.</given-names>
            <surname>Shpinareva</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A. A.</given-names>
            <surname>Yakushina</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L. A.</given-names>
            <surname>Voloshchuk</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N. D.</given-names>
            <surname>Rudnichenko</surname>
          </string-name>
          ,
          <article-title>Detection and classification of network attacks using the deep neural network cascade</article-title>
          ,
          <source>Herald of Advanced Information Technology</source>
          <volume>4</volume>
          (
          <year>2021</year>
          )
          <fpage>244</fpage>
          -
          <lpage>254</lpage>
          . doi:
          <volume>10</volume>
          .15276/hait.03.
          <year>2021</year>
          .
          <volume>4</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>J.</given-names>
            <surname>Andry</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Tannady</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Rembulan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Dinata</surname>
          </string-name>
          ,
          <article-title>Analysis of the Omicron virus cases using data mining methods in rapid miner applications</article-title>
          ,
          <source>Microbes and Infectious Diseases</source>
          <volume>4</volume>
          (
          <year>2023</year>
          ). doi:
          <volume>10</volume>
          .21608/mid.
          <year>2023</year>
          .
          <volume>194619</volume>
          .1469.
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>H.</given-names>
            <surname>Wijaya</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Syaifudin</surname>
          </string-name>
          , T. Siswanto,
          <article-title>Visualization of corona virus disease 2019 deoxyribonucleic acid data analysis</article-title>
          ,
          <source>AIP Conference Proceedings</source>
          <volume>2659</volume>
          (
          <year>2022</year>
          )
          <article-title>090005</article-title>
          . doi:
          <volume>10</volume>
          .1063/5.0118893.
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>A. D.</given-names>
            <surname>Kubegenova</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E. S.</given-names>
            <surname>Kubegenov</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z. M.</given-names>
            <surname>Gumarova</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G. A.</given-names>
            <surname>Kamalova</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G. M.</given-names>
            <surname>Zhazykbaeva</surname>
          </string-name>
          ,
          <article-title>Using Data Mining Technology in Monitoring and Modeling the Epidemiological Situation of the Human Immunodeficiency Virus in Kazakhstan</article-title>
          , in: A.
          <string-name>
            <surname>Gibadullin</surname>
          </string-name>
          (Ed.),
          <source>Information Technologies and Intelligent Decision Making Systems</source>
          , Springer Nature Switzerland, Cham,
          <year>2022</year>
          , pp.
          <fpage>57</fpage>
          -
          <lpage>65</lpage>
          . doi:
          <volume>10</volume>
          .1007/978-3-
          <fpage>031</fpage>
          -21340-
          <issue>3</issue>
          _
          <fpage>6</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>P.</given-names>
            <surname>Srinivas</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Bhattacharyya</surname>
          </string-name>
          ,
          <string-name>
            <surname>D.</surname>
          </string-name>
          <article-title>MidhunChakkaravarthy, Prediction of Swine Flu (H1N1) Virus Using Data Mining and Convolutional Neural Network Techniques</article-title>
          , in: A.
          <string-name>
            <surname>Kumar</surname>
            , G. Ghinea,
            <given-names>S.</given-names>
          </string-name>
          <string-name>
            <surname>Merugu</surname>
          </string-name>
          , T. Hashimoto (Eds.),
          <source>Proceedings of the International Conference on Cognitive and Intelligent Computing</source>
          , Springer Nature Singapore, Singapore,
          <year>2023</year>
          , pp.
          <fpage>557</fpage>
          -
          <lpage>573</lpage>
          . doi:
          <volume>10</volume>
          .1007/
          <fpage>978</fpage>
          -981-19-2358-6_
          <fpage>51</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>H.</given-names>
            <surname>Gunawan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Purwayoga</surname>
          </string-name>
          ,
          <article-title>Data mining menggunakan algoritma k-means clustering untuk mengetahui potensi penyebaran virus corona di kota cirebon, Jurnal Sisfokom (Sistem Informasi dan Komputer</article-title>
          )
          <volume>11</volume>
          (
          <year>2022</year>
          )
          <fpage>1</fpage>
          -
          <lpage>8</lpage>
          . doi:
          <volume>10</volume>
          .32736/sisfokom.v11i1.
          <fpage>1316</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>V.</given-names>
            <surname>Dev</surname>
          </string-name>
          , P. Singh,
          <source>Efect of Corona Virus on Multi-Disease Patients using Association Rule Mining</source>
          ,
          <year>2023</year>
          . URL: https://www.researchgate.net/publication/372312135.
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>E.</given-names>
            <surname>Luna-Ramírez</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Soria-Cruz</surname>
          </string-name>
          ,
          <string-name>
            <given-names>I.</given-names>
            <surname>Castillo-Zúñiga</surname>
          </string-name>
          ,
          <string-name>
            <surname>J. I.</surname>
          </string-name>
          <article-title>López-Veyna, COVID-19 Social Lethality Characterization in some Regions of Mexico through the Pandemic Years Using Data Mining</article-title>
          , in: Y. P. Rybarczyk (Ed.),
          <source>Research Advances in Data Mining Techniques and Applications</source>
          , IntechOpen, Rijeka,
          <year>2023</year>
          . doi:
          <volume>10</volume>
          .5772/intechopen.113261.
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>A. Y. O.</given-names>
            <surname>Allmuttar</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S. K. D.</given-names>
            <surname>Alkhafaji</surname>
          </string-name>
          ,
          <article-title>Using data mining techniques deep analysis and theoretical investigation of COVID-19 pandemic</article-title>
          ,
          <source>Measurement: Sensors</source>
          <volume>27</volume>
          (
          <year>2023</year>
          )
          <article-title>100747</article-title>
          . doi:
          <volume>10</volume>
          .1016/j.measen.
          <year>2023</year>
          .
          <volume>100747</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>M.</given-names>
            <surname>Pandey</surname>
          </string-name>
          , Machine Learning,
          <source>International Journal for Research in Applied Science and Engineering Technology</source>
          <volume>11</volume>
          (
          <year>2023</year>
          )
          <fpage>864</fpage>
          -
          <lpage>869</lpage>
          . doi:
          <volume>10</volume>
          .22214/ijraset.
          <year>2023</year>
          .
          <volume>55224</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <given-names>D. P. F.</given-names>
            <surname>Möller</surname>
          </string-name>
          ,
          <source>Machine Learning and Deep Learning</source>
          , in: Guide to Cybersecurity in Digital Transformation: Trends, Methods, Technologies, Applications and Best Practices, Springer Nature Switzerland, Cham,
          <year>2023</year>
          , pp.
          <fpage>347</fpage>
          -
          <lpage>384</lpage>
          . doi:
          <volume>10</volume>
          .1007/978-3-
          <fpage>031</fpage>
          -26845-
          <issue>8</issue>
          _
          <fpage>8</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <given-names>R.</given-names>
            <surname>Thakur</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Panse</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Bhanarkar</surname>
          </string-name>
          ,
          <article-title>Machine Learning and Deep Learning Techniques</article-title>
          , in: U. N.
          <string-name>
            <surname>Dulhare</surname>
            ,
            <given-names>E. H.</given-names>
          </string-name>
          <string-name>
            <surname>Houssein</surname>
          </string-name>
          (Eds.),
          <source>Machine Learning and Metaheuristics: Methods and Analysis</source>
          , Springer Nature Singapore, Singapore,
          <year>2023</year>
          , pp.
          <fpage>235</fpage>
          -
          <lpage>253</lpage>
          . doi:
          <volume>10</volume>
          .1007/
          <fpage>978</fpage>
          -981-99-6645-5_
          <fpage>11</fpage>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>