<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <journal-title-group>
        <journal-title>Science of The Total Environment</journal-title>
      </journal-title-group>
    </journal-meta>
    <article-meta>
      <article-id pub-id-type="doi">10.1088/1742-6596/801/1</article-id>
      <title-group>
        <article-title>Pattern Recognition and Context Prediction of COVID-19 cases in European Countries</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Arzu Tosayeva</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Ermiyas Birihanu</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Tsegaye Misikir Tashu</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Cases in country Austria Austria Austria</institution>
          ,
          <addr-line>Belgium, Bulgaria, Cyprus Germany</addr-line>
          ,
          <country>Belgium Austria Austria</country>
          <addr-line>Spain, Germany Austria Finland, Estonia</addr-line>
          ,
          <country country="GR">Greece</country>
          ,
          <addr-line>Denmark, Belgium Spain, Germany, Belgium Spain, Germany</addr-line>
          ,
          <country country="GR">Greece</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Department of Artificial Intelligence, Bernoulli Institute of Mathematics</institution>
          ,
          <addr-line>Computer Science and Artificial Intelligence</addr-line>
          ,
          <institution>University of Groningen</institution>
          ,
          <addr-line>Groningen</addr-line>
          ,
          <country country="NL">The Netherlands</country>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>ELTE Eötvös Loránd University</institution>
          ,
          <addr-line>Budapest</addr-line>
          ,
          <country country="HU">Hungary</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2019</year>
      </pub-date>
      <volume>728</volume>
      <issue>012071</issue>
      <fpage>202</fpage>
      <lpage>216</lpage>
      <abstract>
        <p>The global impact of the COVID-19 pandemic has been significant, which requires data analysis to understand trends and patterns. However, this endeavor is challenging due to the complex transmission dynamics and diverse factors that influence the virus's spread. The data associated with COVID-19 is extensive and constantly evolving, and extracting meaningful insights from it is dificult. Therefore, the objective of this study is to analyze the impact of COVID-19 in various European countries, to identify common patterns, and to make predictions within the relevant context. To accomplish this, we used clustering techniques to reveal patterns in COVID-19 cases among European countries. The implementation involved cluster analysis to estimate labels based on cluster size and density while considering relevant background information. Subsequently, a classification model was applied to the labeled dataset. Using the K-Prototypes algorithm and leveraging the Silhouette score for identification, we determined the optimal number of clusters. These clusters were then combined based on density, and the degree of sparsity was assessed. As a result, two clusters emerged: one labeled as "low chance of infection" and the other as "high chance of infection." Using these results, we implemented a classification algorithm, achieving an accuracy rate of 90%. For this study, we gathered data from five diferent sources, consolidating them into a single dataset. Our findings demonstrate that combining COVID-19 datasets with diverse features enables trend analysis, while the use of clustering algorithms facilitates successful label identification in unsupervised learning scenarios involving unlabeled data. The density and size of clusters prove valuable in estimating labels, enhancing our overall understanding of the data. Our code is publicly available here.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;Context prediction</kwd>
        <kwd>COVID-19</kwd>
        <kwd>Label estimation</kwd>
        <kwd>Pattern recognition</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>implemented various measures such as decontamination,
curfews, travel bans, and vaccination campaigns.</p>
      <p>The COVID-19 pandemic has had a profound impact on As governments implement various containment and
the global population, causing significant disruptions in social distancing measures, the demand for healthcare
healthcare systems, industries, and societies around the systems has increased significantly. This poses a
chalworld. To control the spread of the virus and alleviate lenging problem in efectively managing infected
papressure on healthcare systems, numerous countries have tients in hospitals. Having an efective modeling method
implemented strict measures. In Europe, like in other re- that can identify patterns and predict the spread of the
gions, the COVID-19 outbreak emerged in January 2020 virus within the population would be highly valuable for
and quickly escalated, leading to a surge in cases and preparing and formulating health and economic policies
fatalities in hospitals [1]. While some European coun- for governments, administrators, and decision-makers.
tries are currently experiencing new waves of infections, This would aid in slowing down or halting the spread of
others are still dealing with a relatively low number of the virus. With the increase in cases of COVID-19 and
COVID-19 cases. Throughout the epidemic, France, Italy, the availability of more data, several studies have utilized
Spain and the United Kingdom have documented a sig- mathematical models [3], [4], [5] to analyze the spread
nificant number of cases and fatalities[ 2]. To combat the of the virus. In addition, [6], [7] have also used LSTM
virus, the European Union and many member states have models to forecast. However, these models often rely
on outdated data from the same country, which limits
ITAT’23: Information Technologies – Applications and Theory, Septem- their efectiveness. In a cluster analysis study [ 8],
simiber 22–26, 2023, Tatranské Matliare, Slovakia larities were observed in the dynamics of the spread of
* Corresponding author.
the disease between countries such as Italy, France, and
†$Thne0snedanui@thoinrfs.ecloten.thruib(uAte.dToesqauyaellvya.); ermiyasbirihanu@inf.elte.hu Germany, which implemented similar intervention
strate(E. Birihanu); t.m.tashu@rug.nl (T. M. Tashu) gies. In another study [9], supervised machine-learning
0009-0004-0113-9298 (A. Tosayeva); 0000-0001-7081-0365 approaches were used to predict the future of COVID-19.
(E. Birihanu); 0000-0002-4498-2486 (T. M. Tashu) Furthermore, [10] demonstrated the potential of Machine
CPWrEooUrckReshdoinpgs IhStpN:/c1e6u1r3-w-0s.o7r3g ©ACt2tEr0i2bU2utCRioonpWy4r.0igoIhnrttekfornsrahtthiooisnppaalp(PCerCrboByYcite4s.0ea)u.dthionrsg.Usse( CpeErmUittRed-uWndeSr.Correagti)ve Commons License Learning and cloud computing to improve the prediction
of epidemic growth. consider qualitative data, which can influence the spread</p>
      <p>Most existing works are confined to forecasting within of the disease.
specific countries or regions, overlooking the inclusion of In the study conducted by Lai et al. [14], the authors
newly reported cases, recovery rates, and mortality num- analyzed the incidence and mortality rates of 57 countries
bers that vary across diferent countries and over time. in 2020. They used Spearman’s rank-order correlation to
In this study, we utilized COVID-19 data from diverse examine the relationship between cases and deaths.
Howsources. We employed clustering techniques for pattern ever, this research overlooked trends and patterns, which
recognition and applied the concept of estimating class can be crucial for further analysis. Several other research
labels derived from the work of [11] for estimating class works have focused on clustering COVID-19 cases for
labels. Subsequently, we applied classification methods diferent countries, employing the same K-Means
algoto develop, train, and assess a supervised model aimed at rithm.
predicting COVID-19 cases. Labeling unlabeled data has been a significant area of</p>
      <p>The remainder of this study is organized as follows. research, with various ideas and approaches proposed.
Section 2 discusses the existing research conducted in According to Fredriksson et al., [15], approximately 80%
the relevant area, along with its limitations. Section 3 of engineering tasks in a machine learning (ML) project
outlines in detail the proposed research methodology. involve data preparation and labeling. Data preparation
Section 4 presents information on the datasets used in the and labeling often require extensive efort due to
incomresearch, along with the preprocessing and analysis steps plete datasets or the lack of labels for some or all
intaken. Section 5 provides a comprehensive overview of stances. Moreover, even when labels are available, they
the conducted experiments. Finally, Section 6 presents may not be of good quality, leading to incorrect or
parthe conclusions drawn from the research and outlines tially correct labels for data points. High-quality labels
potential future work in the given research area. are crucial for successful supervised machine learning,
as the model’s performance during operations is directly
influenced by the quality of the training data.
2. Related Works Diferent techniques have been suggested by previous
researchers to address the labeling challenge. Cui et al.</p>
      <p>Clustering techniques have been utilized by researchers [16] proposed an approach where samples are divided
since the initial spread of COVID-19 cases. These tech- into clusters, and classification models are applied to
niques have been instrumental in grouping and identi- each cluster based on the dataset’s behaviors. The results
fying distinct patterns to discern diferences or similar- from each cluster are then combined. To improve
clasities among country-specific cases. Cheema et al. [ 12] sification performance, the swarm algorithm was used
proposed the implementation of the K-Means algorithm for clustering, classification, and ensemble learning. The
to establish diverse patterns among countries based on cluster-based ensemble learning method proved efective
various features such as disease prevalence, health sys- with cross-validation practices. However, this research
tems, and environmental indicators. The elbow method only utilized labeled samples and did not consider
unlawas employed to determine the optimal number of clus- beled ones.
ters, considering the sum of squared distances between In another study by Kusumaningrum et al., [17],
Chisamples and their closest cluster centers. Their study Square was used for labeling with the assistance of
Ksuccessfully demonstrated the reliability of Centroid- Means clustering. The homogeneity test was conducted
based Partition clustering in identifying patterns among using the Silhouette coeficient, followed by employing
country-specific cases. However, some limitations of this the Chi-Square Test for automatic cluster labeling in
genresearch include the need to consider both categorical eral.
and numerical features, as K-Means may not be the most In the study conducted by Yogesh [9], the aim was to
suitable choice in such cases. Additionally, as the data develop an LSTM (long-short-term memory) model for
only covers the year 2021, it may not accurately reflect forecasting COVID-19 deaths and cases, specifically for
the current situation for further analysis. Gohari et al. Italy and the United States. The model was subsequently
[13] also employed the K-Means algorithm to analyze lon- evaluated using data from Germany, France, Brazil, India,
gitudinal patterns of change in quantitative COVID-19 and Nepal. On the other hand, Zeroual et al. [18]
conincidence and mortality rates. They utilized dimension- ducted a comparative study of five deep learning methods
reduction techniques to identify correlations between (RNN, LSTM, BILSTM, GRU, and VAR) to forecast the
features, which enhanced the model’s performance. The number of new cases and recovered cases. However, it
research successfully identified three distinct patterns should be noted that both researchers focused on a
limthrough experiments and compared diferent trajectories. ited set of features during their investigation, potentially
Although this research presents an innovative approach limiting the scope and comprehensiveness of their
findin the field, it is limited to quantitative data and does not ings.The work by [19] introduced the concept of creating
a transmission dynamics predictor that exploits tempo- In our research, we aimed to assess the prevalence of
ral variations among diferent countries in relation to specific clusters and determine their commonality. To
the disease’s spread. This is significant because certain achieve this, we utilized the k-prototypes clustering
techcountries encountered outbreaks before others. How- nique, which is suitable for both categorical and
numerever, it’s important to note that the data collected by the ical data. We manipulated the size and density of the
researchers spanned only a duration of three days. clusters by varying the values of parameters  and  . By
applying the k-prototypes clustering algorithm, we were
able to cluster the data efectively. Following the
clus3. Methodology tering process, we assigned a class label to each cluster
based on the features it contained. In COVID-19 cases,
3.1. Pattern Recognition infections are categorized as high or low chance of
infecAssociating a classification with a label is known as recog- tion. We assign labels to clustered dataset points using
nition. Pattern recognition, as the science of drawing two rules: (1) Dense or sparse clusters are designated
conclusions based on data, aims to categorize items or as "high chance," and (2) while others are categorized as
events into groups based on shared characteristics. In "low chance." This involves classifying clusters into "low"
our work, we are utilizing clustering, which falls under or "high" likelihood of infection groups based on their
unsupervised learning, to create patterns and uncover characteristics. As a result, both the cluster and the data
commonalities among the data. Clustering is an efective points contained within it are assigned identical labels,
technique for identifying inherent structures or group- reflecting their infection likelihood.
ings within a dataset without the need for pre-existing In general, if a cluster exhibits high density or sparsity,
labels or categories. it is labeled as having a high chance of infection. On the
other hand, if a cluster is not extremely dense or sparse,
3.2. Label Estimation it can be labeled as having a low chance of infection. This
approach allows us to categorize clusters based on their
Labeling data for COVID-19 manually is a time- characteristics and assign relevant infection likelihood
consuming task that demands significant human re- labels accordingly.
sources. However, our proposed approach tackles this
challenge by estimating relevant labels for data points, 3.3. Context Prediction
enabling supervised learning without the need for
explicit user-provided labels. To separate the instances in the training data into
appro</p>
      <p>When it comes to defining the exact formula for small priate classes, SVM is used. SVM utilizes a hyperplane
clusters, it is important to note that there is no univer- defined as   +  = 0, where  is the weight vector
sally accepted formula. In our approach, we consider the and  is the bias term. The marginal hyperplanes, 1
number of data objects within a cluster to determine its and 2, are given as [20]:
size. Let us assume that we have  clusters and a cluster
set  containing  data points. 1 : (  + ) = 1</p>
      <p>Let the number of data points in each cluster be  and
 be the parameter used to determine whether a cluster 2 : (  + ) = − 1
is considered small or not. a cluster  is considered small Thus, correctly classified points satisfy the inequality:
if
(1)</p>
      <p>(  + ) ≥ 1

|| &lt;  · 
where || is the number of data points of the cluster. For
example,  =0.2 indicates that if a cluster contains less
than 20% then the cluster is considered small. As a next
step, we are exploring whether the cluster is sparse or
dense. For partitioning-based clustering, we used the
sum of squares within the cluster  .</p>
      <p>In order to find the degree of sparsity, we have  , and
we assume that the cluster is sparse when:
  &lt;  · median()</p>
      <p>(2)
where  is the set of   for all . As a result, if  =2.0
means that  of a cluster is greater than 2.0 · ()
and it is sparse.</p>
      <p>In SVM, the margin refers to the distance between
the marginal hyperplanes, also known as the decision
boundary. Specifically, for a linear SVM, the margin is
equal to |2| , where  represents the weight vector of
the hyperplanes. Support vectors are the data points that
lie on either the 1 or 2 hyperplanes, which denfie
the margin. These data points play a crucial role in
determining the position and orientation of the decision
boundary.</p>
    </sec>
    <sec id="sec-2">
      <title>4. Dataset</title>
      <sec id="sec-2-1">
        <title>4.1. Dataset Description</title>
        <sec id="sec-2-1-1">
          <title>The datasets used in this paper were collected from the</title>
        </sec>
        <sec id="sec-2-1-2">
          <title>European Centre for Disease Prevention and Control</title>
          <p>1and [2] Our World in Data 2. In total, five datasets were
used for our investigation.</p>
        </sec>
        <sec id="sec-2-1-3">
          <title>The primary dataset used is titled "Data on the daily</title>
          <p>number of newly reported COVID-19 cases and deaths
by EU/EEA country." This dataset covers the period from
February 2020 to October 2022 from 30 European
countries. The second dataset contains vaccination
information, providing details on the COVID-19 vaccination
progress across diferent countries. The third dataset
includes information on the Gross Domestic Product (GDP)
of countries impacted by COVID-19. The fourth dataset
focuses on travel restrictions implemented by each
country. The fifth dataset concentrates on school restrictions
in European countries. It considers diferent levels and
time periods of school closures, reopening, and other
related restrictions. Finally, we created a comprehensive
compilation of all five datasets, resulting in a final dataset
with a total of 2,246,240 entries. This comprehensive
dataset serves as the basis for our analysis and research
ifndings. It encompasses a diverse range of numerical
and categorical units with varying scales.</p>
        </sec>
      </sec>
      <sec id="sec-2-2">
        <title>4.2. Exploratory Data Analysis</title>
        <p>2022. However, there was a significant decrease in cases
towards the end of 2022, indicating a decline in
COVID19 cases. Figure 2 shows the overall death figures caused
by the spread of COVID-19 from 2021 to 2022. Initially,
during the early stages of this period, the number of
deaths was notably high. This could be attributed to the
limited knowledge and understanding of the pandemic,
and the lack of efective solutions and methods to handle
it. However, as vaccinations and other measures were
implemented, the number of COVID-19 deaths began to
decline. Nevertheless, in 2022, there were still
considerable rates of deaths reported. Towards the end of 2022,
there was a significant decrease in these rates, as several
decisive steps were taken in response to the increased
severity of the pandemic.</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>5. Experimental Setup</title>
      <sec id="sec-3-1">
        <title>5.1. Data Preprocessing</title>
        <p>During the data preprocessing phase, we encountered
missing values in some features of the dataset. To address
these missing values, we applied imputation techniques,
which involve estimating the missing values based on
the available data. In our case, we found that median and
mode imputation were suitable approaches. Therefore,
we performed median or mode imputation wherever
necessary, substituting the missing values with the median
or mode of the corresponding feature. Additionally, we
opted to drop insignificant missing values to ensure the
integrity of the data.</p>
        <sec id="sec-3-1-1">
          <title>Lastly, to ensure data integrity, we conducted a check</title>
          <p>for duplicate columns in the dataset. Any duplicate
columns identified were removed from the dataset to
avoid redundancy. Additionally, we applied
normalization and standardization techniques to scale the data
within specific ranges using Z-score standardization.</p>
        </sec>
        <sec id="sec-3-1-2">
          <title>These pre-processing steps help to enhance the quality</title>
          <p>and comparability of the data for further analysis.</p>
        </sec>
        <sec id="sec-3-1-3">
          <title>Z-score can be defined as follows:</title>
          <p>=
 −</p>
        </sec>
        <sec id="sec-3-1-4">
          <title>Where  represents standardized version of original</title>
          <p>value,  is the dataset we want to normalize  represents
mean of dataset or column and  demonstrates standard
deviation of column or dataset.</p>
        </sec>
      </sec>
      <sec id="sec-3-2">
        <title>5.2. Data Analysis</title>
        <p>Figure 3 shows the direct impact of vaccination on
COVID-19 cases during the entire period of analysis. The
results were obtained by calculating the total number of
vaccinations, which involved summing the administered
doses. We further analyzed the percentage of the
population that received at least one dose of the vaccination.</p>
        <p>In Figure 4, we observed a correlation between the
daily new COVID-19 cases and vaccination rates. When
the daily new cases were high, the percentage of
vaccinated individuals remained low. Conversely, as the
vaccination rates increased, the number of new cases
decreased. As time progressed, with a gradual decline
in the number of COVID-19 cases, the demand for
vaccinations also decreased. These findings demonstrate the
significant role vaccination plays in curbing the spread of
COVID-19 and reducing the number of cases over time.
vice versa. By utilizing the Pearson correlation
coeficient and examining the correlation graph, we can gain
insights into the relationships between diferent features
and identify patterns within the data. This assists in
understanding the underlying structure and dependencies
among the variables, thereby aiding the cluster analysis
process [21].</p>
        <p>In Figure 4, we observe and interpret the correlation
between various features. A positive correlation exists
between Covid-19 cases and deaths (textbf0.62), as well
as between Covid-19 cases and the Population feature
(textbf0.22). On the other hand, we find a negative
correlation between international school closures and Covid-19
cases (textbf-0.24). This indicates that when international
school closures are implemented, there is a reduction in
the number of Covid-19 cases. Similarly, a negative
correlation exists between travel controls and Covid-19 cases
(textbf-0.037), suggesting that as travel controls become
more stringent, the number of Covid-19 cases decreases.</p>
        <p>These correlations provide valuable insights into the
relationships between diferent factors and Covid-19 cases,
Figure 3: Vaccination Vs COVID-19 cases deaths, and control measures. Understanding these
connections can aid in making informed decisions and
formulating efective strategies to manage and mitigate the</p>
        <p>In cluster analysis, various methods can be utilized, impact of the pandemic.
such as Chi-Square, ExtraTreeClassifier, Forward Feature
Selection, and Correlation-based Feature Selection [21].</p>
        <p>For our analysis of the clustering method, we employed a 5.3. Experimentation
correlation graph, specifically utilizing the Pearson Cor- In our experiment, we employed two approaches:
Krelation Coeficient. This coeficient is a measure of the Prototypes [22] and SVM [23]. For both approaches, we
strength of the relationship between diferent features. utilized specific hyperparameters, which are listed in
The Pearson correlation coeficient ranges between -1 Table 1. These hyperparameters were chosen to optimize
and +1. A value close to +1 indicates a strong positive the performance and accuracy of the models.
correlation between features. In such cases, if one feature Rather than using the default parameters for
Kincreases, the other feature is also likely to increase, and Prototypes and SVM, we chose to determine the best
vice versa. Conversely, a value close to -1 indicates a hyperparameters through grid search. The specific
pastrong negative correlation. In this scenario, when one rameter values we chose are presented in Table 1.
feature increases, the other feature tends to decrease, and</p>
      </sec>
      <sec id="sec-3-3">
        <title>5.4. Evaluation Metrics</title>
        <sec id="sec-3-3-1">
          <title>In our clustering evaluation, we utilized the Silhouette Score as a metric to assess the quality and efectiveness of the clustering results. The Silhouette score can be computed as follows:</title>
          <p>=
 − 
max(, )</p>
        </sec>
        <sec id="sec-3-3-2">
          <title>Where  is the mean distance to the points in the</title>
          <p>nearest cluster And,  is the mean intra-cluster distance
to all the points.</p>
          <p>To assess the SVM model’s performance, we employed
metrics such as accuracy, precision, and recall.
 =</p>
          <p>+  
  +   +   +  
  =</p>
          <p>+  
 =</p>
          <p>+</p>
        </sec>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>6. Results</title>
      <sec id="sec-4-1">
        <title>Based on the information provided in Table 2, we ob</title>
        <p>served the distribution of countries across diferent
clusters as follows. Cluster0, Cluster1, Cluster4, Cluster5,
Cluster7, and Cluster9 each consist of one country.
Cluster10 and Cluster11 contain three countries each.
Cluster6 is assigned to two countries and Cluster2 is assigned
to four countries. Cluster8 includes nine countries and
has the highest mean of confirmed cases. During the
Label Estimation phase, all twelve clusters were analyzed
based on their sparsity and density. As a result, it was
determined that Cluster2 and Cluster8 exhibited factors
(4)
(5)
(6)
(7)
associated with a low risk of infection. The specific
characteristics of these clusters contributed to a lower risk of
COVID-19 infection.</p>
        <p>On the other hand, the remaining clusters were
combined to represent the high-risk category. These clusters
likely had higher case densities and lower overall
population vaccinations, indicating a higher risk of COVID-19
spread in those regions. By categorizing the clusters into
high and low-risk categories, we gain valuable insights
into the diferent patterns of COVID-19 prevalence and
can further explore factors contributing to these
variations. As indicated in Table 2, certain countries are
present in multiple clusters. This phenomenon can be
attributed to the fact that certain regions within a
country pose higher risks, while other areas exhibit lower
risks. This pattern emerges due to variations in risk
levels across diferent parts of the same country.</p>
        <p>For context prediction, we employed the Support
Vector Machine (SVM) algorithm. We split the dataset into
a testing set comprising 25 percent of the data and a
training set comprising 75 percent. The table below dis- accuracy and efectiveness of our clustering analysis. By
plays the performance metrics of the model. The SVM refining data scaling, enhancing labeling techniques, and
model achieved an accuracy of 0.84 , indicating that it exploring advanced label estimation approaches, we can
correctly predicted the context in 85 percent of the cases. advance our understanding of COVID-19 patterns and
The precision, which measures the proportion of cor- provide valuable insights for future studies in this field.
rectly predicted positive instances, was 0.90. The recall,
representing the proportion of actual positive instances
correctly identified, was 0.71. These metrics provide in- 8. Conclusion and future work
sights into the performance of the SVM model for context
prediction.</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>7. Discussion</title>
      <p>COVID-19 data has garnered significant attention in
various research domains. However, to the best of our
knowledge, the integration of clustering techniques for
pattern estimation, label exploration, and context prediction
has not been previously investigated. We have
successfully achieved our objective of grouping countries based
on various COVID-19 features, labeling the data using
data characteristics, and implementing context
prediction. However, there is still much more to be done in this
research. One crucial aspect that deserves consideration
is the integration of heterogeneity and coherence of the
clusters to enhance accuracy. Combining clustering
results from multiple algorithms or incorporating ensemble
methods can lead to more robust and reliable clustering
outcomes.</p>
      <p>In our work, we have utilized the K-Prototypes
algorithm for handling both numeric and categorical features.</p>
      <p>However, there are other algorithms, such as the Cluster
Ensemble algorithm (CEBMC), that can handle both types
of features and may provide additional insights.
Exploring and implementing multiple algorithms can lead to a
comprehensive understanding of the data and improve
the analysis. Furthermore, to enhance the performance
of the model, we aim to incorporate multiple clustering
algorithms in a combined manner. Ensemble learning
techniques, where several clustering algorithms work
together, can produce more accurate and stable results.</p>
      <p>Additionally, we seek to enhance the robustness of label
estimation to ensure more accurate results. Exploring
and implementing various label estimation techniques,
considering the characteristics of the dataset, can lead to
more meaningful clustering analysis.</p>
      <p>Overall, our research has made significant progress in
the domain of COVID-19 data analysis using clustering
techniques. However, there are still exciting avenues to
explore and improvements to be made. By addressing
the limitations and considering the integration of
multiple algorithms and ensemble methods, we can advance
our understanding of COVID-19 patterns and contribute
valuable insights to the research community.</p>
      <p>Based on the results obtained from our experiments, it
is important to acknowledge and discuss the potential
challenges and areas for improvement in our work. One
significant challenge we encountered was obtaining
positive results for the Silhouette Score, which is an important
measure to evaluate the quality of cluster grouping.
Initially, we faced negative or low scores due to the lack
of appropriate scaling in our dataset. Scaling the data
correctly is crucial to ensure reliable and meaningful
results. Going forward, it is essential to pay attention to
data scaling techniques and implement them efectively
to enhance the accuracy of our analysis.</p>
      <p>Another challenge we faced was the labeling of the
dataset. Since the initial dataset did not include labels, we
had to undertake the task of labeling the data ourselves.</p>
      <p>Accurate labeling is essential for cluster analysis, as it
provides meaningful interpretation and understanding
of the clusters. It required careful analysis of the features
of each country to assign appropriate labels. In future
research, it would be valuable to explore automated or
semi-automated labeling techniques that can expedite
the process and enhance the accuracy of labeling.
Furthermore, while label estimation techniques were not
the focus of our study, it is an aspect that could
beneift from improvement. Considering the specific features
of the dataset, exploring alternative label estimation
approaches can provide more nuanced and detailed analysis
objectives. Incorporating advanced label estimation
techniques could improve the overall analysis and provide 9. Online Resources
more comprehensive insights into the patterns and
dynamics of COVID-19 spread. Our dataset and code are publicly available here. at https:</p>
      <p>Addressing these challenges and limitations in our //github.com/Tsegaye-misikir/NCC.
research will contribute to further improvements in the</p>
    </sec>
  </body>
  <back>
    <ref-list />
  </back>
</article>