<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <journal-title-group>
        <journal-title>O. Havryliuk);</journal-title>
      </journal-title-group>
    </journal-meta>
    <article-meta>
      <title-group>
        <article-title>Decentralized Segmentation and Prediction of E- Commerce Efficiency Using Machine Learning Methods</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Oleh Havryliuk</string-name>
          <email>o.havryliuk@e-u.edu.ua</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Ihor Ponomarenko</string-name>
          <email>i.ponomarenko@knute.edu.ua</email>
          <xref ref-type="aff" rid="aff4">4</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Maryna Petchenko</string-name>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Oleksandr Yakushev</string-name>
          <email>o.yakushev@chdtu.edu.ua</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Cherkasy State Technological University</institution>
          ,
          <addr-line>460 Shevchenko blvd., 18006 Cherkasy</addr-line>
          ,
          <country country="UA">Ukraine</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>European University</institution>
          ,
          <addr-line>16В Akademika Vernads'koho blvd., 03115 Kyiv</addr-line>
          ,
          <country country="UA">Ukraine</country>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>Kherson National Technical University</institution>
          ,
          <addr-line>11 Instytutska str., 29016 Khmelnytskyi</addr-line>
          ,
          <country country="UA">Ukraine</country>
        </aff>
        <aff id="aff3">
          <label>3</label>
          <institution>State University of Information and Communication Technologies</institution>
          ,
          <addr-line>7 Solomyanska str., 03110 Kyiv</addr-line>
          ,
          <country country="UA">Ukraine</country>
        </aff>
        <aff id="aff4">
          <label>4</label>
          <institution>State University of Trade and Economics</institution>
          ,
          <addr-line>19 Kyoto str., 02156 Kyiv</addr-line>
          ,
          <country country="UA">Ukraine</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2026</year>
      </pub-date>
      <volume>000</volume>
      <fpage>0</fpage>
      <lpage>0003</lpage>
      <abstract>
        <p>The feasibility of using various machine learning algorithms for processing heterogeneous data (structured, semi-structured, and unstructured) is substantiated, and the results obtained can be used to form effective management decisions in the field of marketing. The need to ensure information security is obvious due to the significant risks of unauthorized access to data by third parties for illegal purposes. The need to use decentralization technologies to secure confidential information based on high-performance machine learning algorithms is emphasized. The feasibility of ensuring decentralization by processing an aggregated data set and building a global model with subsequent implementation at the level of an individual company is argued, which minimizes the possibility of loss of commercial data. The study was conducted based on data on the activities of 81 online stores in the consumer electronics market in the Kyiv region for JanuaryMarch 2025 using the specialized web resource Similarweb. The implementation of machine learning algorithms was based on 11 metrics. The importance of cluster analysis and various regression models for data processing and ensuring the efficiency of their integration into decentralized models is proven. The selection of optimal algorithms was based on special metrics and visualization methods. The effectiveness of using the obtained models for effective decentralized implementation at the level of individual online stores is highlighted.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;consumer electronics</kwd>
        <kwd>digitization</kwd>
        <kwd>decentralization</kwd>
        <kwd>cluster analysis</kwd>
        <kwd>machine learning</kwd>
        <kwd>marketing</kwd>
        <kwd>online stores</kwd>
        <kwd>1</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Introduction</title>
      <p>
        Various approaches are used for data processing, among which artificial intelligence has become
particularly widespread. The presented field of knowledge encompasses machine learning, expert
systems, natural language processing, evolutionary computing, and genetic algorithms, among
others. In the fields of marketing and e-commerce, machine learning is widely used due to the
presence of a large number of algorithms that are selected based on the specifics of the data
(numerical expressions, text, photos, videos, audio) and the requirements for the quality of the results
obtained. High-performance mathematical algorithms enable companies to process large datasets
quickly, enhance the quality of models using self-learning principles, and deliver results that
optimize business operations in the digital environment. The digital era is characterized by the
possibility of collecting large amounts of private information that companies need in the process of
forming personalized communications with consumers. However, at the same time, ethical and legal
issues arise regarding the collection of personalized information, which may be perceived negatively
by a large number of users and lead to an increase in the risks of illegal data appropriation by a third
party for criminal activities. The active development of data leakage minimization technologies based
on cybersecurity technologies and the implementation of regulatory acts (for example, the General
Data Protection Regulation (EU) and the Personal Information Protection and Electronic Documents
Act (Canada) stimulate consideration of the principles of decentralization [
        <xref ref-type="bibr" rid="ref2 ref3">2, 3</xref>
        ]. For Web3
ecosystems, decentralization has become the foundation, implemented through technologies such as
edge computing and federated learning, utilizing mobile devices and the Internet of Things. The
current level of development in digital technologies enables the successful integration of machine
learning algorithms into decentralized systems, ensuring a high level of quality in big data processing
[
        <xref ref-type="bibr" rid="ref4">4</xref>
        ].
      </p>
      <p>2.</p>
      <p>The Aim
The research tasks include:
‒ outlining the possibilities of implementing machine learning algorithms while ensuring the
principle of decentralization;</p>
      <p>‒ conducting a cluster analysis on a set of online stores based on the optimization of mathematical
approaches with the assessment of relationships using regression;</p>
      <p>‒ substantiating the possibilities of minimizing the risks of data loss constituting a trade secret, as
well as the consequences of cyberattacks, by adhering to the principle of decentralization.
3.</p>
    </sec>
    <sec id="sec-2">
      <title>Models and Methods</title>
      <p>
        The active development of e-commerce stimulates the formation of modern information systems
that enable the accumulation of relevant information in large sets. Data processing necessitates the
development of effective analytical tools that lay the groundwork for implementing successful
marketing strategies in the digital environment. The presence of a large number of modern machine
learning algorithms allows building high-performance models in accordance with the specifics of
available information and the strategic goals of each company. Ensuring decentralization based on
machine learning algorithms can be achieved thanks to web analytics tools and other approaches to
collecting data on the Internet. This study used information on the activities of 81 online stores in the
consumer electronics market in the Kyiv region for January-March 2025 [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ]; the data was collected
using the Similarweb resource, and their set includes the following metrics: Name of store; Total
visits; Monthly traffic; Mobile web share; Country rank; Visit duration; Pages per visit; Bounce rate;
Organic traffic; Male share; Average age.
      </p>
      <p>The study selected online stores in the consumer electronics market, since modern generations Y,
Z, and Alpha - the largest buyers - regularly purchase new gadgets. The high level of digitalization
has led to the development of e-commerce and the rapid dynamics of purchasing necessary
innovative products online. The choice of the Kyiv region as the study's object was determined by the
concentration of many online stores within the capital of Ukraine, as well as the presence of a large
solvent population.</p>
    </sec>
    <sec id="sec-3">
      <title>Experiment</title>
      <p>Important practical areas studied based on a specific set of observations include classification
carried out according to the existing system of metrics. After testing various algorithms, it was found
that the best grouping results for this set of online stores can be achieved using hierarchical cluster
analysis. To assess the quality of the latter, it is considered advisable to use a silhouette diagram (Fig.
1). The average silhouette value of 0.202 indicates a moderate quality of cluster formation. The best
quality is inherent in Cluster 3, which demonstrates the highest silhouette value (approximately
0.350), indicating the homogeneity and clarity of the represented group. In general, this approach can
also be applied to the process of updating data and increasing the sample size.</p>
      <p>Cluster 2 includes medium-sized online stores that are quite popular among modern users,
although they do not belong to the group of leaders in the consumer electronics market in the Kyiv
region at the beginning of 2025. Cluster 3 includes large online stores in the Kyiv region, which are
characterized by significant popularity among the target audience (for example: FOXTROT, MOYO,
CITRUS).</p>
      <p>Table 1
Average values of online store metrics in January-March 2025 by clusters</p>
      <p>• The greatest activity in e-commerce is inherent in the female audience, due to interest in
specialized online stores with innovative electronics. Additionally, men tend to be less interested in
spending a significant amount of time familiarizing themselves with gadget offers on websites.</p>
      <p>• The given fragment of gradient boost trees demonstrates that the tree-like structure represents
complex nonlinear relationships that cannot be identified based on the usual linear model.</p>
    </sec>
    <sec id="sec-4">
      <title>Further Research</title>
      <p>The results obtained indicate the undeniable potential of using machine learning algorithms for
the effective processing of big data in the field of e-commerce, as well as creating conditions for the
protection of both personal and commercial data. The evolution of digital technologies, cloud
computing, and specialized algorithms has led to the emergence and application of new approaches
in the field of marketing, enabling the identification of hidden connections in information arrays. In
the future, this involves the accumulation of relevant information in the field of e-commerce from
various web resources, primarily from social media. Conducting scientific research and making
effective management decisions in the field of marketing should be based on high-performance
machine learning algorithms and a diverse range of content. The combination of different
information types will contribute to achieving more accurate results, which is especially relevant in
conditions of intensive development in the digital environment.</p>
      <p>6.</p>
    </sec>
    <sec id="sec-5">
      <title>Conclusions</title>
      <p>To optimize business processes in a digital environment, it is advisable to utilize decentralized
segmentation and forecasting of key e-commerce indicators. The implementation of machine
learning algorithms for processing company data in compliance with the principles of
decentralization holds significant prospects, as it enables the achievement of effective results while
also securing personal and commercial data. The Internet is a valuable source of collecting various
information, particularly on the functioning and key performance indicators of 81 online stores,
which are used in the presented study. Securing data in a decentralized system of processing using
machine learning methods requires implementation based on an aggregated dataset with direct
implementation at the local level. Hierarchical cluster analysis and gradient boosting demonstrated
the best results for the obtained data and can be implemented by a specific online store without the
need to display information to third parties.</p>
    </sec>
    <sec id="sec-6">
      <title>Declaration on Generative AI</title>
      <p>During the preparation of this work, the authors used ChatGPT to: translate certain text fragments
into English, perform grammar and spelling checks, and paraphrase or reword content. After using
these tools, the authors carefully reviewed and edited the content as needed and take full
responsibility for the publication’s content.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>F.</given-names>
            <surname>Sabbagh</surname>
          </string-name>
          ,
          <article-title>Digital Economy and Communication Technologies: Methods and Mechanisms of Promotion through E-Commerce and</article-title>
          <string-name>
            <surname>E-Marketing</surname>
          </string-name>
          ,
          <source>Indian Journal of Data Communication and Networking</source>
          , vol.
          <volume>1</volume>
          , no.
          <issue>3</issue>
          ,
          <issue>2021</issue>
          , pp.
          <fpage>10</fpage>
          -
          <lpage>22</lpage>
          . doi:
          <volume>10</volume>
          .54105/ijdcn.B5003.061321
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <surname>Regulation</surname>
          </string-name>
          (EU)
          <year>2016</year>
          /
          <article-title>679 of the European Parliament</article-title>
          and of the Council,
          <year>2016</year>
          . [Online]. URL: https://eur-lex.europa.eu/eli/reg/2016/679/oj/eng
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          <source>[3] The Personal Information Protection and Electronic Documents Act (PIPEDA)</source>
          ,
          <year>2025</year>
          . URL: https://www.priv.gc.ca/en/privacy-topics/
          <article-title>privacy-laws-in-canada/the-personalinformation-protection-and-electronic-documents-act-pipeda/</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>A.</given-names>
            <surname>Brecko</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Kajati</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Koziorek</surname>
          </string-name>
          ,
          <string-name>
            <surname>and I. Zolotova</surname>
          </string-name>
          ,
          <article-title>Federated Learning for Edge Computing: A Survey, Applied Sciences</article-title>
          , vol.
          <volume>12</volume>
          , no.
          <issue>18, article</issue>
          9124,
          <year>2022</year>
          . doi:
          <volume>10</volume>
          .3390/app12189124.
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <surname>Similarweb</surname>
          </string-name>
          ,
          <year>2025</year>
          . URL: https://pro.similarweb.com/
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <surname>Cumulative</surname>
            <given-names>charts</given-names>
          </string-name>
          ,
          <year>2025</year>
          . URL: https://docs.datarobot.com/en/docs/modeling/analyzemodels/evaluate/roc-curve-tab/cumulativecharts.html#:~:text=
          <source>Cumulative%20Gain%20represents%20the%20sensitivity</source>
          ,to%
          <source>2020%25%2 0of%20total%20customers</source>
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>