<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Automated System for Determining a Damage Class for Sections of a Wastewater Disposal Network</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Olena Shapovalova</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Olga Starkova</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Ganna Solodovnyk</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Kharkiv National University of Civil Engineering and Architecture</institution>
          ,
          <addr-line>Sumska street, 40, Kharkiv, 61002</addr-line>
          ,
          <country country="UA">Ukraine</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>The purpose of the study is to provide rationale and describe the process of designing an automated system for determining a damage class for sections of a wastewater disposal network using one of the up-to-date algorithms of data mining. The current stage of development of society requires the introduction of mathematical methods and mining models and the latest developments in the field of IT technology, even in such a well-established industry as public utilities. The proposed automated system provides the user with the opportunity to distribute sections of the network by clusters with the determination of their centers with the subsequent prioritization of repair work. The paper proposes a new approach to determining a damage class for sections of a wastewater disposal network using clustering algorithms with the option of choosing one of two metrics (Manhattan distance and Euclidean metric). The software component implemented in the Java programming language with the use and within the operation of the system allows determining by a number of section's criteria their affiliation with a cer class with the option of further ranking and prioritization of renovation work. The software is a web application which, on condition of good Internet connection, can be useful for professionals engaged in maintaining trouble-free operation of wastewater disposal networks at any time and at any location.</p>
      </abstract>
      <kwd-group>
        <kwd>1 Damage class</kwd>
        <kwd>wastewater disposal network sections</kwd>
        <kwd>clustering</kwd>
        <kwd>web application</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>
        Among the challenges faced by public utilities in cities and localities across Ukrainian, it is the
issues of ensuring the smooth operation of wastewater disposal networks that head the list [
        <xref ref-type="bibr" rid="ref1 ref2 ref3">1-3</xref>
        ].
Sewer canals, the surfaces of which are constantly exposed to aggressive substances, over time suffer
from typical damage, which results in the failure of certain sections and the occurrence of an
emergency situation. A timely diagnostics of the most vulnerable sections and the network as a whole,
followed by preventive measures to maintain the reliable operation of the system may significantly
reduce the cost of repair in the event of an accident and ensure the provision of quality services to
consumers [
        <xref ref-type="bibr" rid="ref4 ref5 ref6">4-6</xref>
        ].
      </p>
      <p>
        Automating the processing of survey data about the condition of canals and ranking sections by
the nature and extent of their damage is a relevant task at the moment, which will minimize the
associated labor costs. Involving the latest developments in mathematical modeling, data mining and
IT technology in solving practical problems, in particular in the construction, renovation and
maintenance of complex systems [
        <xref ref-type="bibr" rid="ref7 ref8">7-8</xref>
        ] is a common practice among modern researchers trying to find
optimum solutions and reduce costs through implementation of high technology [
        <xref ref-type="bibr" rid="ref10 ref8 ref9">8-10</xref>
        ].
      </p>
      <p>
        The authors of a number of papers [
        <xref ref-type="bibr" rid="ref10 ref5 ref6 ref7 ref8 ref9">5-10</xref>
        ] suggest different approaches to deal with the issue. For
instance, the paper [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ] considers the issues of the strategy of selection of potential objects for
rehabilitation of water supply networks based on the analysis of the results of technical diagnostics of
the condition of the sections according to a number of appropriate criteria. In this case, the criteria are
the dynamics of changes in the reliability of particular sections, the presence of external destabilizing
_____________________________________
ITTAP’2021: 1nd InternationaWlorkshop on Information Technologies: Theoretical and Applied Problems, November 16–18, 2021,
Ternopil, Ukraine
EMAIL: shapovalova.olena@kstuca.kharkov.ua (A. 1); starkova@kstuca.kharkov.ua (A. 2); solodovnik@kn-it.info (A. 3)
ORCID: 0000-0003-4566-6634 (A. 1); 0000-0002-9034-8830 (A. 2); 0000-0001-6323-5083 (A. 3)
      </p>
      <p>2021 Copyright for this paper by its authors.</p>
      <p>Use permitted under Creative Commons License Attribution 4.0 International (CC BY 4.0).</p>
      <p>CEUR Workshop Proceedings (CEUR-WS.org)
factors, age, technical condition and repairability of pipelines, likely residual useful life period,
preliminary costs of rehabilitation, the ratio of the actual cost of pipelines to the required costs in case
of their rehabilitation, restrictions on financial costs.</p>
      <p>
        The method and information system proposed in the paper [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ] for determining the classification
of the sewer network sections by condition allows determining the condition class of a sewer section,
rating points, and lists of both individual sections and entire sewers that need priority rehabilitation.
Unfortunately though, the effective use of this method requires special training and a high level of
mastery of special software tools, which often does not allow fully applying it in the practice of
planning repair and rehabilitation work.
      </p>
      <p>
        The information and technical support offered by the authors allows storing the results of the
technical diagnostics of the condition of sections in the data bank and, retrieving, if necessary, the
necessary data on request to evaluate the listed criteria and determine the area of the urban water
supply network or the area with the highest accident rate in pipelines (according to the diameters,
materials and service life selected for analysis) [
        <xref ref-type="bibr" rid="ref5 ref6">5, 6</xref>
        ].
      </p>
      <p>
        Attempts made by domestic researchers to assess the condition and elaborate a priority list for
rehabilitation of wastewater disposal network facilities have been quite successful [
        <xref ref-type="bibr" rid="ref10 ref7 ref8 ref9">7-10</xref>
        ]. The
information system proposed by the authors allows automating the process of elaborating a list of
priority facilities of a wastewater disposal network to be rehabilitated, and visualizing, based on the
data about the condition of the sewer network, the network section which the user is interested in with
all its appropriate parameters [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ].
      </p>
    </sec>
    <sec id="sec-2">
      <title>2. Materials and methods of research</title>
      <p>
        The reviewed studies do not give the unambiguous and general description which would allow
determining the classification of the sewer network sections, being guided by their technical
condition. Therefore, it is of relevance to consider the existing methods for determining the
classification of sewers by condition. Moreover, at the moment there are no automated programs for
determining the damage class using cluster analysis [
        <xref ref-type="bibr" rid="ref11 ref12 ref13">11-13</xref>
        ].
      </p>
      <p>
        The relevance of the study and the development of a software application based on the stud
findings are confirmed by the permanent interest of researchers in the issue [
        <xref ref-type="bibr" rid="ref1 ref10 ref2 ref3 ref4 ref5 ref6 ref7 ref8 ref9">1-10</xref>
        ] and involve the
need for creating an automated system for determining the damage class of the wastewater disposal
network sections. This application will be useful to specialists in the area of maintenance of
troublefree operation of the wastewater disposal network and will provide an opportunity to best develop
repair plans based on the analysis of large amounts of information and classification of sections
according to the level of their damage.
      </p>
      <p>
        The study aims at providing rationale and describing the process of designing an automated
system for determining a damage class for sections of a wastewater disposal network using a
clustering algorithm [
        <xref ref-type="bibr" rid="ref14 ref15 ref16">14-16</xref>
        ]. The automated system is meant to distribute the wastewater disposal
network sections by clusters (with the determination of their centers) with the subsequent
prioritization of repair work [
        <xref ref-type="bibr" rid="ref16 ref17">16-17</xref>
        ]. The object of research is a software application for the
implementation of clustering algorithms to differentiate sections of the wastewater disposal network
according to the extent of their damage. The subject of research is the technology for implementing
algorithms of analysis for the wastewater disposal network sections with the use of the IntelliJ Idea
integrated development environment tools and Java programming language.
      </p>
    </sec>
    <sec id="sec-3">
      <title>3. Research findings</title>
      <p>According to the results of the analysis of the subject area, the input data for the automated
system for determining the damage class of a wastewater disposal network section included the name
of the section and parameters of cross section reduction, cracks, deformation, pipe ruptures,
reinforced concrete corrosion, and the output data included the affiliation of the section with a
particular cluster according to the appropriate parameters of its center. The user is given the
opportunity to initiate, according to the input data, the calculation process, review and download the
data to a file. In the course of loading the input data, the input information is prepared in a certain way
and loaded into the system for further analysis and identification of consistent patterns.</p>
      <p>
        The initiated calculation process activates the data clustering process, during which the user
selects the metrics, the number of clusters and iterations and receives the distribution of sections by
clusters [
        <xref ref-type="bibr" rid="ref18 ref19 ref20">18-20</xref>
        ]. The calculation results in the form of a file with information about each section with
the indication of the damage class, and the coordinates of the centroids of clusters are stored and can
be reviewed at the user’s request and used in the future to make decisions on prioritization of
work[
        <xref ref-type="bibr" rid="ref20 ref21 ref22 ref23">20-23</xref>
        ].
      </p>
      <p>An activity diagram shows the sequence of actions of the automated system and the division of
sections into clusters (Fig. 1).</p>
      <p>
        In the course of operation, the system performs the algorithmic and mathematical calculations as
follows:
 Forming a file from the output data;
 Cluster distribution of sections selected for review using the k-means clustering method
[
        <xref ref-type="bibr" rid="ref24 ref25">24-25</xref>
        ];
 Searching for cluster centroids using Euclidean metric [
        <xref ref-type="bibr" rid="ref26">26</xref>
        ] and Manhattan metric [
        <xref ref-type="bibr" rid="ref27">27</xref>
        ].
      </p>
      <p>To perform data clustering using the k-means clustering method, the number of clusters and the
number of iterations are set automatically. Each object is represented by a sequence of standardized
attributes, i.e. the values of all parameters (attributes) are reduced to the interval [0,100]. In the course
of operation the algorithm of the object is given the cluster to which centroid it appears to be the
nearest. The described operations are performed until the number of operations exceeds the maximum
allowable number. An additional prerequisite to exiting the loop is to reach a state when objects are
no longer assigned to new clusters. At the output of the algorithm, the user is provided with a set of
objects divided into clusters, the centroids of which are equidistant from each other.</p>
      <p>The diagram of the classes is shown in Fig. 2</p>
      <p>The classes of the web application are placed in 5 packages, which are responsible for calculations,
configuration, display and application layers, such as Model, Service and Controller (Fig. 4).
The following tools were used to create the web application:
 Apache Maven framework, designed to automate the collection of projects based on the
description of their structure in files in the POM (Project Object Model) language;.
 Tomcat open source servlet container, which allows running web applications and contains a
number of self-configuring programs;
 Java TestNG programming language testing system;
 Library for modular testing of software in Java JUnit and its extension DBUnit;
 A set of Bootstrap tools for creating websites and web applications;
 Spring Boot Java-based open source environment.</p>
      <p>To start working with the web application, the user should go to the main page of the website
(Fig.5).</p>
      <p>Determining a damage class for sections of wastewater</p>
      <p>a disposal network
K-means clustering method
K-means
Select data
Select file
Number of clusters</p>
      <p>Calculate</p>
      <p>In this form, the user can choose the number of clusters, the number of iterations, calculate or not
the centroids, apply the Euclidean or Manhattan metrics. Also, a .txt file with calculation data must be
uploaded.</p>
      <p>The uploaded file must contain in the lines the numerical data of the section in terms of 5
parameters (cross section reduction, crack, deformation, pipe rupture, and reinforced concrete
corrosion) and its text name, separated by commas. If any parameter is missing, it is acceptable to
replace it with the number 0.</p>
      <p>Clicking the “Calculate”button after uploading the input data will get the result. For ease of
further use and analysis, the data can be presented in tabular form using the Excel program built-in
tools (Fig. 6).</p>
    </sec>
    <sec id="sec-4">
      <title>4. Discussion of the findings</title>
      <p>
        The developed automated system is used to determine the damage class for sections of a
wastewater disposal network. The system is based on the main criteria, principles and methodological
tools for classifying potential objects for renovation, which are covered in the research papers [
        <xref ref-type="bibr" rid="ref2 ref3">2,3</xref>
        ]
and cluster analysis algorithms, the area of application of which is wide enough and discussed in the
papers [
        <xref ref-type="bibr" rid="ref11 ref12 ref13 ref14">11-14</xref>
        ].
      </p>
    </sec>
    <sec id="sec-5">
      <title>1. Conclusions</title>
      <p>According to the findings of the study, using the example of well-known automated systems for
determining the damage class for the wastewater disposal network sections, approaches to dealing
with the issue were considered, the subject environment was analyzed and an activity diagram was
plotted, which was the basis for designing an automated system. Based on the analysis of the subject
area, the input and output data for the system were selected; a cluster analysis was chosen as the
method for determining the priority of repair work.</p>
      <p>A system has been designed and implemented, the software component of which is a
webapplication with client and server parts that automate the determination of the damage class for the
wastewater disposal network sections.</p>
      <p>The user is given the opportunity to store the results of calculations in the database and perform
their further processing using spreadsheets. The system was tested and showed high speed of
statistical data processing and 100% accuracy of distribution of sections by clusters on the test
sample. In the future, it is planned to use the designed system for the analysis of sections of a
wastewater disposal network both in real time and for drawing up plans for renovation work.</p>
      <p>2. References
cluster</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <surname>Honcharenko</surname>
            <given-names>D.F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bondarenko</surname>
            <given-names>A.I.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bulgakov</surname>
            <given-names>V.V.</given-names>
          </string-name>
          and
          <article-title>Garmash A.A. On the issue of ensuring the maintainability of sewer tunnels in Kharkov</article-title>
          .
          <source>Naukovyi visnyk budivnytstva</source>
          ,
          <source>2016. Issue 2</source>
          , pp.
          <fpage>144</fpage>
          -
          <lpage>148</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <surname>Starkova</surname>
            <given-names>O.V.</given-names>
          </string-name>
          <article-title>Models of a reasonable choice of the method of repair and restoration of a section of the sewer network</article-title>
          .
          <source>Naukovyi visnyk budivnytstva</source>
          .
          <source>2016. Issue 3</source>
          , pp.
          <fpage>80</fpage>
          -
          <lpage>85</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <surname>Starkova</surname>
            <given-names>O.V.</given-names>
          </string-name>
          ,
          <string-name>
            <given-names>Shapovalova</given-names>
            <surname>Ye</surname>
          </string-name>
          .A.,
          <string-name>
            <surname>Gnuchikh L</surname>
          </string-name>
          .
          <article-title>A. Modeling the choice of the method of restoration of sewerage networks</article-title>
          .
          <source>Komunalne hospodarstvo mist</source>
          .
          <source>2008. Issue 85</source>
          , pp.
          <fpage>19</fpage>
          -
          <lpage>26</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>Bulgakov</given-names>
            <surname>Yu</surname>
          </string-name>
          .V.
          <article-title>Research of the process of destruction of the construction of the sewer tunnel collector</article-title>
          .
          <source>Naukovyi visnyk budivnytstva</source>
          ,
          <source>2015. Issue 5</source>
          , pp.
          <fpage>79</fpage>
          -
          <lpage>84</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <surname>Orlov</surname>
            <given-names>V.A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kharkin</surname>
            <given-names>V.A.</given-names>
          </string-name>
          <article-title>Development of a strategy for the restoration of wastewater disposal networks</article-title>
          .
          <source>Stroitelstvo i Arkhitektura</source>
          ,
          <year>2001</year>
          , pp.
          <fpage>15</fpage>
          -
          <lpage>32</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <surname>Korynko</surname>
            <given-names>Y.V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Starkova</surname>
            <given-names>O.V.</given-names>
          </string-name>
          ,
          <article-title>and Shevchenko A.A. Methodological foundations of computer modeling of wastewater disposal systems</article-title>
          .
          <source>Naukovyi visnyk budivnytstva</source>
          ,
          <year>2003</year>
          . Issue 23, pp.
          <fpage>223</fpage>
          -
          <lpage>229</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <surname>Honcharenko</surname>
            <given-names>D.F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Starkova</surname>
            <given-names>O.V.</given-names>
          </string-name>
          and
          <string-name>
            <surname>Aleynykova</surname>
            <given-names>A.Y.</given-names>
          </string-name>
          <article-title>Development of an automated system for selecting a method for restoring water pipelines using a fuzzy logic</article-title>
          .
          <source>Systemy obrobky informatsiyi</source>
          ,
          <source>2014. Issue</source>
          <volume>8</volume>
          (
          <issue>124</issue>
          ), pp.
          <fpage>18</fpage>
          -
          <lpage>23</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <surname>Solodovnyk</surname>
            <given-names>H.V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Deyneha</surname>
            <given-names>A.A.</given-names>
          </string-name>
          <article-title>Making multi-stage decisions using information technology</article-title>
          .
          <source>Naukovyi visnyk budivnytstva</source>
          ,
          <source>2019. Issue 2</source>
          , pp.
          <fpage>165</fpage>
          -
          <lpage>166</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <surname>Starkova</surname>
            <given-names>O.V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Shapovalova</surname>
            <given-names>E.A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hnuchykh</surname>
            <given-names>L.A.</given-names>
          </string-name>
          and
          <string-name>
            <surname>Bondarenko</surname>
            <given-names>D.A.</given-names>
          </string-name>
          <article-title>Development of an automated information system for determining priority objects for the renovation of water supply pipelines</article-title>
          .
          <source>Komunalne hospodarstvo mist</source>
          ,
          <year>2011</year>
          . Issue 99, pp.
          <fpage>312</fpage>
          -
          <lpage>316</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <surname>Honcharenko</surname>
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Shumakov</surname>
            <given-names>I.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Starkova</surname>
            <given-names>O.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Aleinikova</surname>
            <given-names>A.</given-names>
          </string-name>
          and
          <string-name>
            <surname>Mikautadze</surname>
            <given-names>R..</given-names>
          </string-name>
          <article-title>Methodological and computer-based support for choosing underground utility networks renovation method //</article-title>
          <source>MATEC Web of Conferences 230</source>
          ,
          <year>02010</year>
          ,
          <year>2018</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <surname>Quackenbush</surname>
            <given-names>J.</given-names>
          </string-name>
          <article-title>Computational analysis of microarray data</article-title>
          .
          <source>Nature reviews genetics</source>
          ,
          <year>2001</year>
          , vol.
          <volume>2</volume>
          , no.
          <issue>6</issue>
          , pp.
          <fpage>418</fpage>
          -
          <lpage>427</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <surname>Sugar</surname>
            <given-names>C. A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>James G. M.</surname>
          </string-name>
          <article-title>Finding the number of clusters in a dataset</article-title>
          .
          <source>Journal of the American Statistical Association</source>
          ,
          <year>2003</year>
          , vol.
          <volume>98</volume>
          , no.
          <issue>463</issue>
          , pp.
          <fpage>750</fpage>
          -
          <lpage>763</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <surname>Tibshirani</surname>
            <given-names>R.</given-names>
          </string-name>
          , Walther G.
          <article-title>Cluster validation by prediction strength</article-title>
          .
          <source>Journal of Computational and Graphical Statistics</source>
          ,
          <year>2005</year>
          , vol.
          <volume>14</volume>
          , no.
          <issue>3</issue>
          , pp.
          <fpage>511</fpage>
          -
          <lpage>528</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <surname>Tibshirani</surname>
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Walther</surname>
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hastie</surname>
            <given-names>T. Estimating</given-names>
          </string-name>
          <article-title>the number of clusters in a data set via the gap statistic</article-title>
          .
          <source>Journal of the Royal Statistical Society. Series B (Statistical Methodology)</source>
          ,
          <year>2001</year>
          , vol.
          <volume>63</volume>
          , no.
          <issue>2</issue>
          , pp.
          <fpage>411</fpage>
          -
          <lpage>423</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <surname>Cuevas</surname>
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Febrero</surname>
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Fraiman</surname>
            <given-names>R</given-names>
          </string-name>
          .
          <article-title>Cluster analysis: a further approach based on density estimation</article-title>
          .
          <source>Computational Statistics &amp; Data Analysis</source>
          ,
          <year>2001</year>
          , vol.
          <volume>36</volume>
          , no.
          <issue>4</issue>
          , pp.
          <fpage>441</fpage>
          -
          <lpage>459</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [16]
          <string-name>
            <surname>Stuetzle</surname>
            <given-names>W. Estimating</given-names>
          </string-name>
          <article-title>the cluster tree of a density by analyzing the minimal spanning tree of a sample</article-title>
          .
          <source>Journal of classification</source>
          ,
          <year>2003</year>
          , vol.
          <volume>20</volume>
          , no.
          <issue>1</issue>
          , pp.
          <fpage>25</fpage>
          -
          <lpage>47</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [17]
          <string-name>
            <surname>Pelleg</surname>
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Moore</surname>
            <given-names>A. W.</given-names>
          </string-name>
          <article-title>X-means: Extending K-means with Efficient Estimation of the Number of Clusters</article-title>
          . ICML,
          <year>2000</year>
          , pp.
          <fpage>727</fpage>
          -
          <lpage>734</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          [18]
          <string-name>
            <surname>Volkovich</surname>
            <given-names>Z.</given-names>
          </string-name>
          , Brazly
          <string-name>
            <given-names>Z.</given-names>
            ,
            <surname>Toledano-Kitai</surname>
          </string-name>
          <string-name>
            <given-names>D.</given-names>
            ,
            <surname>Avros</surname>
          </string-name>
          <string-name>
            <surname>R.</surname>
          </string-name>
          <article-title>The Hotelling's metric as stability measure</article-title>
          .
          <source>Computer modelling and new technologies</source>
          ,
          <year>2010</year>
          , vol.
          <volume>14</volume>
          , pp.
          <fpage>65</fpage>
          -
          <lpage>72</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          [19]
          <string-name>
            <surname>Barzily</surname>
            <given-names>Z.</given-names>
          </string-name>
          , Volkovich
          <string-name>
            <given-names>Z.</given-names>
            ,
            <surname>Akteke-Ozturk</surname>
          </string-name>
          <string-name>
            <surname>B</surname>
          </string-name>
          .
          <article-title>On a minimal spanning tree approach in the cluster validation problem</article-title>
          .
          <source>Informatica, Lith. Acad. Sci.</source>
          ,
          <year>2009</year>
          , vol.
          <volume>20</volume>
          , no.
          <issue>2</issue>
          , pp.
          <fpage>187</fpage>
          -
          <lpage>202</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          [20]
          <string-name>
            <surname>Hamerly</surname>
            <given-names>Y. F. G.</given-names>
          </string-name>
          <article-title>PG-means: learning the number of clusters in data</article-title>
          .
          <source>Advances in neural information processing systems</source>
          ,
          <year>2007</year>
          , vol.
          <volume>19</volume>
          , pp.
          <fpage>393</fpage>
          -
          <lpage>400</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          [21]
          <string-name>
            <surname>Dudoit</surname>
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Fridlyand</surname>
            <given-names>J.</given-names>
          </string-name>
          <article-title>A prediction-based resampling method for estimating the number of clusters in a dataset</article-title>
          .
          <source>Genome biology</source>
          ,
          <year>2002</year>
          , vol.
          <volume>3</volume>
          , no.
          <issue>7</issue>
          , research0036. Available at: http://www.genomebiology.com/ (assessed:
          <fpage>17</fpage>
          .
          <fpage>08</fpage>
          .
          <year>2021</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          [22]
          <string-name>
            <surname>Lange</surname>
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Roth</surname>
            <given-names>V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Braun</surname>
            <given-names>M. L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Buhmann</surname>
            <given-names>J. M.</given-names>
          </string-name>
          <article-title>Stability-based validation of clustering solutions</article-title>
          .
          <source>Neural computation</source>
          ,
          <year>2004</year>
          , vol.
          <volume>16</volume>
          , no.
          <issue>6</issue>
          , pp.
          <fpage>1299</fpage>
          -
          <lpage>1323</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          [23]
          <string-name>
            <surname>Ben-Hur</surname>
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Elisseeff</surname>
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Guyon</surname>
            <given-names>I.</given-names>
          </string-name>
          <article-title>A stability based method for discovering structure in clustered data</article-title>
          .
          <source>Pacific symposium on biocomputing</source>
          ,
          <year>2002</year>
          , vol.
          <volume>7</volume>
          , no.
          <issue>6</issue>
          , pp.
          <fpage>6</fpage>
          -
          <lpage>17</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          [24]
          <string-name>
            <given-names>Adam</given-names>
            <surname>Coates</surname>
          </string-name>
          and
          <string-name>
            <given-names>Andrew Y.</given-names>
            <surname>Ng</surname>
          </string-name>
          .
          <article-title>Learning Feature Representations with K-means</article-title>
          , Stanford University,
          <year>2012</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref25">
        <mixed-citation>
          [25]
          <string-name>
            <surname>Mirkes</surname>
            <given-names>E.M.</given-names>
          </string-name>
          <article-title>K-means and K-medoids applet</article-title>
          . University of Leicester,
          <year>2011</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref26">
        <mixed-citation>
          [26]
          <string-name>
            <surname>Dyuran</surname>
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Odell</surname>
            <given-names>P</given-names>
          </string-name>
          .
          <article-title>Cluster analyze</article-title>
          . Moscow: Statistika,
          <year>1977</year>
          , 128 p.
        </mixed-citation>
      </ref>
      <ref id="ref27">
        <mixed-citation>
          [27]
          <string-name>
            <surname>Kim</surname>
            ,
            <given-names>J. O.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Mueller</surname>
            ,
            <given-names>C. W.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>Klecka</surname>
            ,
            <given-names>W. R.</given-names>
          </string-name>
          <string-name>
            <surname>Factorial</surname>
          </string-name>
          , discriminant and cluster analyzes.
          <source>Moscow: Finansy i statistika</source>
          ,
          <year>1989</year>
          , 215 p.
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>