<?xml version="1.0" encoding="UTF-8"?>
<TEI xml:space="preserve" xmlns="http://www.tei-c.org/ns/1.0" 
xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" 
xsi:schemaLocation="http://www.tei-c.org/ns/1.0 https://raw.githubusercontent.com/kermitt2/grobid/master/grobid-home/schemas/xsd/Grobid.xsd"
 xmlns:xlink="http://www.w3.org/1999/xlink">
	<teiHeader xml:lang="en">
		<fileDesc>
			<titleStmt>
				<title level="a" type="main">Experimental Evaluation of the Effectiveness of ANN-based Numerical Data Augmentation Methods for Diagnostics Tasks</title>
			</titleStmt>
			<publicationStmt>
				<publisher/>
				<availability status="unknown"><licence/></availability>
			</publicationStmt>
			<sourceDesc>
				<biblStruct>
					<analytic>
						<author>
							<persName><forename type="first">Ivan</forename><surname>Izonin</surname></persName>
							<email>ivanizonin@gmail.com</email>
							<affiliation key="aff0">
								<orgName type="institution">Lviv Polytechnic National University</orgName>
								<address>
									<addrLine>S. Bandera str., 12</addrLine>
									<postCode>79013</postCode>
									<settlement>Lviv</settlement>
									<country key="UA">Ukraine</country>
								</address>
							</affiliation>
						</author>
						<author>
							<persName><forename type="first">Roman</forename><surname>Tkachenko</surname></persName>
							<email>roman.tkachenko@gmail.com</email>
							<affiliation key="aff0">
								<orgName type="institution">Lviv Polytechnic National University</orgName>
								<address>
									<addrLine>S. Bandera str., 12</addrLine>
									<postCode>79013</postCode>
									<settlement>Lviv</settlement>
									<country key="UA">Ukraine</country>
								</address>
							</affiliation>
						</author>
						<author>
							<persName><forename type="first">Roman</forename><surname>Pidkostelnyi</surname></persName>
							<email>roman.pidkostelnii@gmail.com</email>
							<affiliation key="aff0">
								<orgName type="institution">Lviv Polytechnic National University</orgName>
								<address>
									<addrLine>S. Bandera str., 12</addrLine>
									<postCode>79013</postCode>
									<settlement>Lviv</settlement>
									<country key="UA">Ukraine</country>
								</address>
							</affiliation>
						</author>
						<author>
							<persName><forename type="first">Olena</forename><surname>Pavliuk</surname></persName>
							<email>olena.m.pavliuk@lpnu.ua</email>
							<affiliation key="aff0">
								<orgName type="institution">Lviv Polytechnic National University</orgName>
								<address>
									<addrLine>S. Bandera str., 12</addrLine>
									<postCode>79013</postCode>
									<settlement>Lviv</settlement>
									<country key="UA">Ukraine</country>
								</address>
							</affiliation>
						</author>
						<author>
							<persName><forename type="first">Viktor</forename><surname>Khavalko</surname></persName>
							<email>khavalkov@gmail.com</email>
							<affiliation key="aff0">
								<orgName type="institution">Lviv Polytechnic National University</orgName>
								<address>
									<addrLine>S. Bandera str., 12</addrLine>
									<postCode>79013</postCode>
									<settlement>Lviv</settlement>
									<country key="UA">Ukraine</country>
								</address>
							</affiliation>
						</author>
						<author>
							<persName><forename type="first">Anatoliy</forename><surname>Batyuk</surname></persName>
							<email>abatyuk@gmail.com</email>
							<affiliation key="aff0">
								<orgName type="institution">Lviv Polytechnic National University</orgName>
								<address>
									<addrLine>S. Bandera str., 12</addrLine>
									<postCode>79013</postCode>
									<settlement>Lviv</settlement>
									<country key="UA">Ukraine</country>
								</address>
							</affiliation>
						</author>
						<title level="a" type="main">Experimental Evaluation of the Effectiveness of ANN-based Numerical Data Augmentation Methods for Diagnostics Tasks</title>
					</analytic>
					<monogr>
						<imprint>
							<date/>
						</imprint>
					</monogr>
					<idno type="MD5">B4D7AD2D5C35E42C4A8460BDFA576D39</idno>
				</biblStruct>
			</sourceDesc>
		</fileDesc>
		<encodingDesc>
			<appInfo>
				<application version="0.7.2" ident="GROBID" when="2023-03-25T01:38+0000">
					<desc>GROBID - A machine learning software for extracting information from scholarly documents</desc>
					<ref target="https://github.com/kermitt2/grobid"/>
				</application>
			</appInfo>
		</encodingDesc>
		<profileDesc>
			<textClass>
				<keywords>
					<term>Tabular data</term>
					<term>classification</term>
					<term>overfitting risk</term>
					<term>underfitting risk</term>
					<term>data augmentation</term>
					<term>ANN</term>
					<term>GAN</term>
					<term>autoencoder</term>
					<term>small data approach</term>
				</keywords>
			</textClass>
			<abstract>
<div xmlns="http://www.tei-c.org/ns/1.0"><p>Improving the accuracy of diagnostics tasks is essential in various medical fields. When there are small data for training, there are high risks of overfitting or underfitting the machine learning model. This makes it impossible to apply it in practice. To solve such a problem, we can use various data augmentation methods. This paper focuses on neural network methods of data augmentation. The authors have investigated a variational autoencoder and approach based on GAN to generate artificial numerical data and then use it by machine-learning-based classifiers. The authors examined the proposed method for diagnosing diabetes mellitus development task. Experiments confirmed that autoencoders generated a dataset similar to an initial one, with a similarity score being 0.93. The authors established a significant accuracy improvement of Random Forest, AdaBoost, and Logistic regression classifiers based on processing an extended dataset. The application of the new dataset obtained using GAN does not ensure satisfactory accuracy. Such an issue may be due to a lack of samples for the training of this neural networks class. Further research is likely to be carried out into ensembles based on a single machine learning method, which will process decorrelated samples acquired by methods investigated in this paper.</p></div>
			</abstract>
		</profileDesc>
	</teiHeader>
	<text xml:lang="en">
		<body>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="1.">Introduction</head><p>The development of modern medicine has been marked by digitizing a wide variety of information and the automation of many processes <ref type="bibr" target="#b0">[1]</ref>. This makes it possible to collect a large amount of data for analysis. It also opens up new opportunities for applying data mining techniques to intellectualizing specific diagnostics or treatment processes.</p><p>However, the scarce data may impede the implementation of machine learning. Alternatively, abnormal data may lead to increased accuracy, which is a critical point in this area.</p><p>One possible solution to this problem lies in adopting data augmentation methods. This approach can allow synthesizing of enough data to train the selected artificial intelligence tool.</p><p>Nowadays, there are quite a few simple methods for manipulating an available sample of data to increase its size. The data are enlarged both by rows and by columns <ref type="bibr" target="#b0">[1]</ref>. However, these methods do not always introduce helpful information into the expanded dataset and, consequently, only increase the learning time of the selected model. The accuracy of the chosen classifier or regressors is not affected here.</p><p>Many neural network methods have been developed to increase a dataset today. A wide variety of artificial neural network topologies are employed here. The augmentation is performed using a variety of information -from time series to images. Generative adversarial network <ref type="bibr" target="#b1">[2]</ref> is among the most used methods for artificial augmentation of datasets, particularly in the field of image processing. This type of neural network is most commonly used to synthesize new images for further use by deep learning neural networks. Another type of neural network is autoencoders, which is often and successfully applied in time series analysis. However, developing and researching a methodology for effective artificial augmentation of numerical datasets remains to be solved. On the one hand, neural network methods are more sophisticated and should reveal patterns in the dataset that are difficult to detect with simple methods <ref type="bibr" target="#b2">[3]</ref>. Such information can serve as a basis for the synthesis of new patterns in the dataset. Alternatively, a neural network toolkit must obtain sufficient data for training and validating the model. Moreover, generalization properties should be especially emphasized. Only by meeting all these requirements will the selected tool operate adequately and synthesize the required amount of synthetic data of the required quality. Thus this paper aims to investigate neural network methods for enlarging tabular datasets to improve the accuracy of classification based on them.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="2.">Materials and methods</head><p>This section includes a description of two neural-network-based approaches for numerical data augmentation used in this paper. The main objective is to improve the classification accuracy in Clinical Medicine based on expanded datasets.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="2.1.1.">ANN-based numerical data augmentation methods</head><p>The first approach selected is a new method for generating an artificial dataset based on a Generative adversarial network (GAN) <ref type="bibr" target="#b3">[4]</ref>. To this end, the author of the technique modified neural networks to deal specifically with numerical datasets. The modification was as follows. The authors proposed to use Conditional GANs as a generator of numerical data. This approach is explained by: -efficient performance in the event of an unbalanced dataset; -independence of the type of variables: discrete and continuous, with the possibility of modeling them both at the same time;</p><p>-a flexible approach to modeling the distribution of probabilities within the dataset; -the possibility of synthesizing high-quality synthetic samples that are very similar to the observations from the initial dataset.</p><p>A peculiarity of this method is that the authors use a special normalization method and a set of stateof-the-art model learning methods, and a post-annotated network. In other respects, the method works like a conventional GAN.</p><p>Another interesting method is data augmentation based on a variational autoencoder <ref type="bibr" target="#b4">[5]</ref>. It is referred to as generative models. Learning methods of this family consist of mapping objects into a given latent space and reproducing them back. The task related to the autoencoder is to find the functions that will allow mapping the latent variable area to another one, an understandable and simple space. A customarily distributed space is a case in point.</p><p>While designing methods based on variational autoencoder, one should define the number of neurons in the first and the second latent layers and set the number of latent factors. It will contain all applicable information and serve as a decoder to recover all initial inputs. After all the necessary settings have been made, the learning procedure can be performed.</p><p>If the hidden dependencies between variables are linear, the variational autoencoder works as a PCA method <ref type="bibr" target="#b5">[6]</ref>. In this case, to each element according to the method <ref type="bibr" target="#b4">[5]</ref> some random noise will be added to get the best autoencoder performance. As investigated by the author of the method, this approach provides the possibility of obtaining an artificial set closer to the real dataset compared to the method without noise. That is why this method was taken for comparison. Details of its implementation are given in <ref type="bibr" target="#b6">[7]</ref> The synthesis of the new data is as follows. Beforehand we know the variance and the mean of our latent variables, which are determined by the autoencoder. The next step is the use of a normal distribution with the variance and the mean for each of the latent variables. It is needed for the selection of the value for all latent variables. It is these that serve as starting points from which all attributes of the initial data set can be reproduced.</p><p>In case the latent dependencies between variables are linear, the variational autoencoder works like the PCA method <ref type="bibr" target="#b5">[6]</ref>, which suggests that some random noise will be added to each item according to the method of <ref type="bibr" target="#b4">[5]</ref> to obtain better results from the autoencoder. As the reported by author of the method investigated in <ref type="bibr" target="#b4">[5]</ref>, this approach allows producing an artificial set more similar to the real dataset compared with the method without using noise. That is why this method was taken for comparison. Details of its implementation are presented in <ref type="bibr" target="#b6">[7]</ref>.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="2.1.2.">Dataset description</head><p>Diagnostics tasks are widespread in the medical industry. The majority of them are reduced to a classification task, and we can apply machine learning techniques. For example, in <ref type="bibr" target="#b7">[8]</ref>, a dataset is submitted, and the task of predicting the development of diabetes task is formulated. The original variable is represented as 0 or 1. Thus, it is a binary classification problem. Fig. <ref type="figure" target="#fig_0">1</ref> shows a few distributions for some variable pairs.  </p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="3.">Modeling and results</head><p>Simulation of the classifiers investigated based on extended datasets using neural network methods. Experimental studies were carried out by dividing the dataset randomly into two parts at a ratio of 80% to 20%. Cross-validation was then applied (5 times). In this way, the reliability of the results was ensured.</p><p>The paper presents two neural network approaches for the artificial expansion of short datasets. Let us consider the outcomes and evaluations of the synthesized data for each of them.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="3.1.1.">Classification using augmented data via autoencoder</head><p>Autoencoder-based modeling was carried out in order to synthesize a new dataset whose size would match the size of the original dataset. A comparison between the synthesized dataset with the original one based on several indicators from <ref type="bibr" target="#b8">[9]</ref>, has revealed the following results: the mean correlations between fake and real columns are 0.97, MAPE estimator results are 0.84, and a similarity score is 0.93.</p><p>In addition, Figure <ref type="figure" target="#fig_1">2</ref> presents a comparison of the feature distributions for the initial and synthesized datasets.</p><p>It should be noted that the autoencoder generated the same number of instances of each class as the number of instances in the initial dataset.</p><p>The application results of different classifiers on the extended dataset are summarised in Table <ref type="table" target="#tab_1">1</ref>.  <ref type="figure" target="#fig_2">3</ref> show the dependence of the accuracy of the machine learning methods on the amout of the artificially generated vectors added to the initial set.  As can be seen from Figure <ref type="figure" target="#fig_2">3</ref> the increase in the number of additionally added, artificially generated vectors to the initial data set led to an increase in the accuracy of all classifiers. The highest accuracy was obtained by added to the initial set the same dimension of the artificial sample (768 additional vectors).</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="3.1.2.">Classification using augmented data via GAN</head><p>The authors adopted the method from <ref type="bibr" target="#b3">[4]</ref> for numerical data augmentation. As in the previous cases, the size of the new dataset is equal to the size of the initial one. A comparison between the simulated dataset and the initial one based on several indicators from <ref type="bibr" target="#b8">[9]</ref> has shown the following results: mean correlations between fake and real columns are 0.92, MAPE estimator results are 0.67, and a similarity score is 0.59.</p><p>Figure <ref type="figure">4</ref> shows a comparison of the feature distributions for the initial and synthesized datasets.</p><p>It is worth remarking that the GAN-based method attempted to balance the dataset. It generated significantly more instances of the smaller class compared to the initial dataset.</p><p>The application results of different classifiers on the extended dataset are summarised in Table <ref type="table" target="#tab_2">2</ref>.  <ref type="table" target="#tab_2">2</ref> suggests that the ensemble-based classifiers exhibit significantly higher accuracy than the other two methods.</p><p>Figure <ref type="figure">5</ref> show the dependence of the accuracy of the machine learning methods on the amout of the artificially generated vectors added to the initial set. As can be seen from Figure <ref type="figure">5</ref> the increase in the number of additionally added, artificially generated vectors to the initial data set did not always lead to an increase in the accuracy of the classifiers. Only extension of the initial set by more than 60% of new, artificially synthesized data vectors helped to reduce the errors of classifiers. The highest accuracy, as in the previous case, was obtained by added to the initial dataset the same dimension of the artificial sample (768 additional vectors).</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="3.2.">Comparison and discusion</head><p>This section compares both the new datasets generated by both methods under investigation and the classification results based on their application.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="3.2.1.">Numerical evaluation of the synthetic datasets</head><p>In this paper, a comparison of the results obtained by every method investigated was performed on the basis of some indicators from <ref type="bibr" target="#b8">[9]</ref>. The results of the comparison between the real dataset and one synthesized by GAN or autoencoder are summarized in Table <ref type="table">3</ref>.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head>Table 3</head><p>An evaluation of the synthetic datasets Indicator obtained by Autoencoder obtained by GAN Mean correlations between fake and real columns 0,97 0,92 Mape estimator results 0,84 0,67 Similarity score 0,93 0,59</p><p>As shown in Table <ref type="table" target="#tab_1">1</ref>, the data augmentation method based on the autoencoder provides significantly higher results in comparison with the dataset obtained by GAN. This can significantly affect the performance of classifiers with these data.</p><p>However, such dataset decorrelation enables the construction of ensemble models based on a single classifier to process different datasets. This approach can significantly improve the accuracy of classification methods in medicine. </p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="3.2.2.">Comparison of the classification accuracy of the different classifiers</head><p>The performance of both neural network approaches was compared by determining the accuracy of a few known classifiers: Random forest classifier; AdaBoost classifier; Logistic regression classifier; SVM classifier. They were employed for classification based on the initial and new datasets. It should be noted that the dimensionality of the new datasets was doubled, thus the synthesized data were added to the original one. The outcomes that are based on Total accuracy, Precision, and Recall are shown in Fig. <ref type="figure" target="#fig_5">6</ref>. Since the problem is not balanced, F-measure was not taken into account. From the graphs in Fig. <ref type="figure" target="#fig_5">6</ref>, it follows that the highest accuracy based on all performance indicators is achieved by using a synthesized dataset with an autoencoder. The application of GAN for data augmentation shows a much lower performance of the known classifiers compared to processing an initial dataset (based on the total accuracy). However, AdaBoost and Random Forest algorithms provide more accurate results in this case. This can be explained by the insufficient amount of training data for effective GAN performance, which affected the one of SVM and Logistic Regression classifiers.</p></div>
<div xmlns="http://www.tei-c.org/ns/1.0"><head n="4.">Conclusion</head><p>This paper deals with the numerical data augmentation task in Clinical Medicine. The authors have experimentally evaluated the performance of modern neural network methods to solve the problem: autoencoders and a GAN. Such an approach helps to reduce risks of overfitting or underfitting when using machine learning models in case of small data processing. The modeling of the performance of these methods has been carried out using the dataset for solving a classification task. In this case, we tried to predict the possibility of diabetes development. The dataset is not balanced. Experiments have shown that autoencoders generate the most similar data according to the Similiarity score. In addition, the accuracy of classifiers based on these data is significantly higher. Compared with the initial dataset, we have improved the target resolution accuracy by about 10%.</p><p>Given the different results of the similarity evaluation of the synthesized datasets concerning the initial one and the different accuracy of the classifiers based on such data, the ensemble learning approach can be used in further research to improve the accuracy of various diagnostics tasks. In particular, the approach of constructing a stacking ensemble of homotypic classifiers that will process different systematically studied datasets seems promising. This very approach can provide a significant increase in the accuracy of classifiers when solving applied tasks of diagnostics in various fields <ref type="bibr">[10]</ref><ref type="bibr">[11]</ref><ref type="bibr" target="#b9">[12]</ref><ref type="bibr" target="#b10">[13]</ref><ref type="bibr" target="#b11">[14]</ref><ref type="bibr" target="#b12">[15]</ref> when processing average datasets.</p><p>The National Research Foundation of Ukraine funds this study from the state budget of Ukraine within the project "Decision support system for modeling the spread of viral infections" (№ 2020.01 / 0025).</p></div><figure xmlns="http://www.tei-c.org/ns/1.0" xml:id="fig_0"><head>Figure 1 :</head><label>1</label><figDesc>Figure 1: Distribution of features</figDesc><graphic coords="3,86.20,495.38,423.98,235.80" type="bitmap" /></figure>
<figure xmlns="http://www.tei-c.org/ns/1.0" xml:id="fig_1"><head>Figure 2 :</head><label>2</label><figDesc>Figure 2: Distribution per features for initial and synthetic datasets</figDesc><graphic coords="5,82.25,543.00,430.30,154.48" type="bitmap" /></figure>
<figure xmlns="http://www.tei-c.org/ns/1.0" xml:id="fig_2"><head>Figure 3 :</head><label>3</label><figDesc>Figure 3: Dependence of the classification accuracy on the number of generated additional vectors by the autoencoder (0 -initial sample, 768 -100% of additional vectors)</figDesc></figure>
<figure xmlns="http://www.tei-c.org/ns/1.0" xml:id="fig_3"><head>Figure 4 :Figure 5 :</head><label>45</label><figDesc>Figure 4: Distribution per features for initial and synthetic datasets</figDesc></figure>
<figure xmlns="http://www.tei-c.org/ns/1.0" xml:id="fig_5"><head>Figure 6 :</head><label>6</label><figDesc>The outcomes of different classifiers based on the initial datasets, and datasets generated using GAN and using autoencoders data augmentation methods: a) Total Accuracy; b) Recall; c) Precision</figDesc></figure>
<figure xmlns="http://www.tei-c.org/ns/1.0" type="table" xml:id="tab_1"><head>Table 1</head><label>1</label><figDesc></figDesc><table><row><cell cols="4">Classification accuracy for investigated ML-based methods using extended dataset obtained by</cell></row><row><cell>autoencoder</cell><cell></cell><cell></cell><cell></cell></row><row><cell>Machine learning algorithm</cell><cell>Total accuracy</cell><cell>Recall</cell><cell>Precision</cell></row><row><cell>Random forest classifier</cell><cell>0.8249</cell><cell>0.7139</cell><cell>0.7846</cell></row><row><cell>AdaBoost classifier</cell><cell>0.8301</cell><cell>0.6990</cell><cell>0.8230</cell></row><row><cell>Logistic regression classifier</cell><cell>0.8295</cell><cell>0.6773</cell><cell>0.8334</cell></row><row><cell>SVM classifier</cell><cell>0.7985</cell><cell>0.6205</cell><cell>0.8016</cell></row><row><cell cols="3">Table 1 reveals that all methods demonstrate high classification accuracy.</cell><cell></cell></row><row><cell>Figure</cell><cell></cell><cell></cell><cell></cell></row></table></figure>
<figure xmlns="http://www.tei-c.org/ns/1.0" type="table" xml:id="tab_2"><head>Table 2</head><label>2</label><figDesc>Classification accuracy for investigated ML-based methods using extended dataset obtained by GAN</figDesc><table><row><cell>Machine learning algorithm</cell><cell>Total accuracy</cell><cell>Recall</cell><cell>Precision</cell></row><row><cell>Random forest classifier</cell><cell>0.7946</cell><cell>0.5853</cell><cell>0.7603</cell></row><row><cell>AdaBoost classifier</cell><cell>0.8015</cell><cell>0.6991</cell><cell>0.7734</cell></row><row><cell>Logistic regression classifier</cell><cell>0.7352</cell><cell>0.5358</cell><cell>0.6907</cell></row><row><cell>SVM classifier</cell><cell>0.7326</cell><cell>0.5101</cell><cell>0,6714</cell></row><row><cell>Table</cell><cell></cell><cell></cell><cell></cell></row></table></figure>
		</body>
		<back>
			<div type="references">

				<listBibl>

<biblStruct xml:id="b0">
	<monogr>
		<title level="m" type="main">DeltaPy: A Framework for Tabular Data Augmentation in Python</title>
		<author>
			<persName><forename type="first">D</forename><surname>Snow</surname></persName>
		</author>
		<idno type="DOI">10.2139/ssrn.3582219</idno>
		<ptr target="https://doi.org/10.2139/ssrn.3582219" />
		<imprint>
			<date type="published" when="2020">2020</date>
			<publisher>Social Science Research Network</publisher>
			<pubPlace>Rochester, NY</pubPlace>
		</imprint>
	</monogr>
</biblStruct>

<biblStruct xml:id="b1">
	<analytic>
		<title level="a" type="main">An intelligent system for cytological and histological image analysis</title>
		<author>
			<persName><forename type="first">O</forename><surname>Berezsky</surname></persName>
		</author>
		<author>
			<persName><forename type="first">G</forename><surname>Melnyk</surname></persName>
		</author>
		<author>
			<persName><forename type="first">T</forename><surname>Datsko</surname></persName>
		</author>
		<author>
			<persName><forename type="first">S</forename><surname>Verbovy</surname></persName>
		</author>
		<idno type="DOI">10.1109/CADSM.2015.7230787</idno>
		<ptr target="https://doi.org/10.1109/CADSM.2015.7230787" />
	</analytic>
	<monogr>
		<title level="m">The Experience of Designing and Application of CAD Systems in Microelectronics</title>
				<meeting><address><addrLine>Lviv -Polyana, Ukraine</addrLine></address></meeting>
		<imprint>
			<publisher>IEEE</publisher>
			<date type="published" when="2015">2015</date>
			<biblScope unit="page" from="28" to="31" />
		</imprint>
	</monogr>
</biblStruct>

<biblStruct xml:id="b2">
	<analytic>
		<title level="a" type="main">Fractal Distribution of Medical Data in Neural Network</title>
		<author>
			<persName><forename type="first">N</forename><surname>Boyko</surname></persName>
		</author>
		<author>
			<persName><forename type="first">M</forename><surname>Kuba</surname></persName>
		</author>
		<author>
			<persName><forename type="first">L</forename><surname>Mochurad</surname></persName>
		</author>
		<author>
			<persName><forename type="first">S</forename><surname>Montenegro</surname></persName>
		</author>
	</analytic>
	<monogr>
		<title level="j">CEUR-WS.Org</title>
		<imprint>
			<biblScope unit="volume">2488</biblScope>
			<biblScope unit="page" from="307" to="318" />
			<date type="published" when="2019">2019</date>
		</imprint>
	</monogr>
</biblStruct>

<biblStruct xml:id="b3">
	<monogr>
		<title level="m" type="main">Modeling Tabular data using Conditional GAN</title>
		<author>
			<persName><forename type="first">L</forename><surname>Xu</surname></persName>
		</author>
		<author>
			<persName><forename type="first">M</forename><surname>Skoularidou</surname></persName>
		</author>
		<author>
			<persName><forename type="first">A</forename><surname>Cuesta-Infante</surname></persName>
		</author>
		<author>
			<persName><forename type="first">K</forename><surname>Veeramachaneni</surname></persName>
		</author>
		<idno>ArXiv:1907.00503</idno>
		<ptr target="http://arxiv.org/abs/1907.00503" />
		<imprint>
			<date type="published" when="2019-12-26">2019. December 26, 2020</date>
		</imprint>
	</monogr>
	<note>Cs, Stat</note>
</biblStruct>

<biblStruct xml:id="b4">
	<monogr>
		<ptr target="https://lschmiddey.github.io/fastpages_/2021/04/10/DeepLearning_TabularDataAugmentation.html" />
		<title level="m">Deep Learning for tabular data augmentation</title>
				<imprint>
			<date type="published" when="2021-05-16">2021. May 16, 2021</date>
		</imprint>
		<respStmt>
			<orgName>Data Science Blog von Lschmiddey</orgName>
		</respStmt>
	</monogr>
</biblStruct>

<biblStruct xml:id="b5">
	<analytic>
		<title level="a" type="main">New Approaches in the Learning of Complex-Valued Neural Networks</title>
		<author>
			<persName><forename type="first">V</forename><surname>Kotsovsky</surname></persName>
		</author>
		<author>
			<persName><forename type="first">A</forename><surname>Batyuk</surname></persName>
		</author>
		<author>
			<persName><forename type="first">M</forename><surname>Yurchenko</surname></persName>
		</author>
		<idno type="DOI">10.1109/DSMP47368.2020.9204332</idno>
	</analytic>
	<monogr>
		<title level="m">IEEE Third International Conference on Data Stream Mining &amp; Processing (DSMP)</title>
				<imprint>
			<date type="published" when="2020">2020. 2020</date>
			<biblScope unit="page" from="50" to="54" />
		</imprint>
	</monogr>
</biblStruct>

<biblStruct xml:id="b6">
	<monogr>
		<author>
			<persName><surname>Lschmiddey</surname></persName>
		</author>
		<ptr target="https://github.com/lschmiddey/deep_tabular_augmentation" />
		<title level="m">lschmiddey/deep_tabular_augmentation</title>
				<imprint>
			<date type="published" when="2021-05-16">2021. May 16, 2021</date>
		</imprint>
	</monogr>
</biblStruct>

<biblStruct xml:id="b7">
	<monogr>
		<ptr target="https://kaggle.com/uciml/pima-indians-diabetes-database" />
		<title level="m">Pima Indians Diabetes Database</title>
				<imprint>
			<date type="published" when="2021-05-16">May 16, 2021</date>
		</imprint>
	</monogr>
</biblStruct>

<biblStruct xml:id="b8">
	<monogr>
		<ptr target="https://baukebrenninkmeijer.github.io/table-evaluator/table_evaluator.html" />
		<title level="m">TableEvaluator table evaluator 15-08-2019 documentation</title>
				<imprint>
			<date type="published" when="2021-05-16">May 16, 2021</date>
		</imprint>
	</monogr>
</biblStruct>

<biblStruct xml:id="b9">
	<analytic>
		<title level="a" type="main">Fuzzy recurrent mappings in multiagent simulation of population dynamics systems</title>
		<author>
			<persName><forename type="first">D</forename><surname>Chumachenko</surname></persName>
		</author>
		<author>
			<persName><forename type="first">O</forename><surname>Sokolov</surname></persName>
		</author>
		<author>
			<persName><forename type="first">S</forename><surname>Yakovlev</surname></persName>
		</author>
		<idno type="DOI">10.47839/ijc.19.2.1773</idno>
		<ptr target="https://doi.org/10.47839/ijc.19.2.1773" />
	</analytic>
	<monogr>
		<title level="j">IJC</title>
		<imprint>
			<biblScope unit="page" from="290" to="297" />
			<date type="published" when="2020">2020</date>
		</imprint>
	</monogr>
</biblStruct>

<biblStruct xml:id="b10">
	<analytic>
		<title level="a" type="main">Situation diagnosis based on the spatially-distributed dynamic disaster risk assessment</title>
		<author>
			<persName><forename type="first">M</forename><surname>Zharikova</surname></persName>
		</author>
		<author>
			<persName><forename type="first">V</forename><surname>Sherstjuk</surname></persName>
		</author>
	</analytic>
	<monogr>
		<title level="m">IEEE 14th Intern. Conf. CSIT</title>
				<imprint>
			<date type="published" when="2019">2019. 2019</date>
			<biblScope unit="page" from="205" to="209" />
		</imprint>
	</monogr>
</biblStruct>

<biblStruct xml:id="b11">
	<analytic>
		<title level="a" type="main">Technology of Gene Expression Profiles Filtering Based on Wavelet Analysis</title>
		<author>
			<persName><forename type="first">Sergii</forename><surname>Babichev</surname></persName>
		</author>
		<author>
			<persName><forename type="first">Jiří</forename><surname>Škvor</surname></persName>
		</author>
		<author>
			<persName><forename type="first">Jiří</forename><surname>Fišer</surname></persName>
		</author>
		<author>
			<persName><forename type="first">Volodymyr</forename><surname>Lytvynenko</surname></persName>
		</author>
	</analytic>
	<monogr>
		<title level="j">International Journal of Intelligent Systems and Applications(IJISA)</title>
		<imprint>
			<biblScope unit="volume">10</biblScope>
			<biblScope unit="issue">4</biblScope>
			<biblScope unit="page" from="1" to="7" />
			<date type="published" when="2018">2018</date>
		</imprint>
	</monogr>
</biblStruct>

<biblStruct xml:id="b12">
	<analytic>
		<title level="a" type="main">Estimation of the inductive model of objects clustering stability based on the k-means algorithm for different levels of data noise</title>
		<author>
			<persName><forename type="first">S</forename><forename type="middle">A</forename><surname>Babichev</surname></persName>
		</author>
		<author>
			<persName><forename type="first">V</forename><forename type="middle">I</forename><surname>Lytvynenko</surname></persName>
		</author>
		<author>
			<persName><forename type="first">M</forename><forename type="middle">A</forename><surname>Taif</surname></persName>
		</author>
	</analytic>
	<monogr>
		<title level="j">Computer Science, Control</title>
		<imprint>
			<biblScope unit="issue">4</biblScope>
			<biblScope unit="page" from="54" to="60" />
			<date type="published" when="2016">2016</date>
		</imprint>
	</monogr>
	<note>Radio Electronics</note>
</biblStruct>

				</listBibl>
			</div>
		</back>
	</text>
</TEI>
