<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Modeling of wheat yield in the steppe region of Ukraine using machine learning techniques</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Petro Hrytsiuk</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Tetiana Babych</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Olena Hladka</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Maryna Nehrey</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Collegium Helveticum, ETH Zurich</institution>
          ,
          <addr-line>Zürich</addr-line>
          ,
          <country country="CH">Switzerland</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>National University of Water and Environmental Engineering</institution>
          ,
          <addr-line>Soborna str., 11, Rivne, 33000</addr-line>
          ,
          <country country="UA">Ukraine</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>The task of this study is to evaluate the climatic factors impact on the detrended values of wheat yield using machine learning techniques. The average decadal temperature values for April, May, June and monthly amounts of precipitation for this period for five regions of the steppe zone of Ukraine were selected for the study. The work uses an innovative approach, according to which the detrended yield values are divided models were used, which were fitted to the available data and demonstrated classification accuracy above 80% on test samples. The support vector method and the random forest method are the most effective classifiers and provide 85% classification accuracy (on test data).</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;Wheat yield</kwd>
        <kwd>climatic factors</kwd>
        <kwd>machine learning</kwd>
        <kwd>classification</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>
        Grain production is one of the most important branches of the economy of Ukraine, ensuring the
food needs of the population and a stable inflow of currency. The average annual production of
cereals in Ukraine for 2019-2021 reached the level of 75 million tons (in 2021, a record crop of 84
million tons was harvested in Ukraine), and the average annual export during this time was 50
million tons [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ].
      </p>
      <p>
        At the same time, it is necessary to note the significant instability of grain production in Ukraine,
associated with the impact of changing climatic factors, which have undergone significant changes
in the last 30 years. This led to a change in the assortment of cultivated grain crops and the geography
of their location [
        <xref ref-type="bibr" rid="ref2 ref3">2, 3</xref>
        ]. There is an increase in the production of heat-loving crops, such as corn,
soybeans, and sunflowers in the chernozem zone of Ukraine and in the Polissia zone. In recent years,
against the backdrop of climate change, the wheat share in the total grain harvest has decreased from
50% to 40%, and the corn share has increased from 15% to 42% [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. Warming, which is accompanied
by a decrease in the amount of precipitation, causes a negative impact on the yield of grain crops.
The steppe region of Ukraine is particularly sensitive to changes in climatic factors, where frequent
droughts lead to a significant drop in grain yields. Therefore, this region is losing its leading position
in the grain production, instead, the share of the central and western regions of Ukraine is increasing.
      </p>
      <p>Domestic consumption of grain in recent years did not exceed 20 million tons. This is
approximately 30% of all grain production, and 70% of grain is exported. Thus, grain production from
the main food resource of the country, which it was in the 20th century, turned into the largest
source of foreign exchange for Ukraine and the key of its economic development. In the last three
years alone, revenues from grain exports amounted to approximately 30 billion US dollars.</p>
      <p>The basis for planning a long-term grain export strategy is the yield forecasting. This is a complex
task, the essence of which is determined by the random nature of many influencing factors.
Therefore, to solve this problem, it is advisable to apply intelligent data analysis techniques with
modern computer technologies using.
__________________
ICST-2024: Information Control Systems &amp; Technologies, September 23-25, 2023, Odesa, Ukraine.</p>
      <p>p.m.hrytsiuk@nuwm.edu.ua (P. Hrytsiuk); t.iu.babych@nuwm.edu.ua (T. Babych); o.m.hladka@nuwm.edu.ua (O. Hladka);
mnehrey@ethz.ch (M. Nehrey)</p>
      <p>0000-0002-3683-4766 (P. Hrytsiuk); 0000-0001-6927-7313 (T. Babych); 0000-0003-4728-0663 (O. Hladka);
0000-00019243-1534 (M. Nehrey)
© 2024 Copyright for this paper by its authors.</p>
      <p>Use permitted under Creative Commons License Attribution 4.0 International (CC BY 4.0).</p>
      <p>At the current stage, the most modern concepts of mathematical modeling are used to build
predictive models, among which the machine learning techniques occupy a leading place. In this
research, there was used such a powerful machine learning tool as classification methods.</p>
      <p>
        The number of works devoted to the research of climatic factors impact on the grain crops yield
in Ukraine is limited [
        <xref ref-type="bibr" rid="ref2 ref4 ref5">2, 4, 5</xref>
        ]. Complicated access to agroclimatic data is one of the reasons for the
insufficient number of publications. From the view point of the grain crops cultivation, the territory
of Ukraine can be divided into several agro-climatic zones: the steppe region, the black soil zone of
the forest-steppe region, the western region. For each of these zones, the nature of yield dependence
on climatic factors will be different. The main purpose of this study is to analyze and model the
impact of climatic factors on wheat yields fluctuations in the steppe region of Ukraine.
      </p>
    </sec>
    <sec id="sec-2">
      <title>2. Literature Review</title>
      <p>
        Wheat production is the basis of Ukrainian agriculture, but climate change threatens it at risk in
some regions of Ukraine. In a comprehensive analytical review conducted within the framework of
the German-Ukrainian Agricultural Policy Dialogue project [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ], the impact of climate changes on
winter wheat yields in the three agroecological zones of Ukraine, as previously mentioned, was
assessed. According to the authors' conclusion, the main concern is the fertile steppe zone, where
the climate is hotter and drier, and frequent droughts are also observed.
      </p>
      <p>
        The increase in the droughts frequency in recent years is seen as a major threat to agriculture.
The author of [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ] investigated the impact of climate change on the level of major agricultural crops
production, as well as on Hungary's GDP. The paper [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ] shows that the machine learning models
shows have stronger predictive power than standard econometric approaches.
      </p>
      <p>Scientific and technical progress contributed to the arrival of large volumes of statistical data from
various branches of agriculture. This greatly expanded the possibilities of using computer
technologies for the analysis and modeling of climatic effects on the agricultural crops yield. In
recent years, there have been publications describing the machine learning methods application to
forecasting the agricultural crops yield.</p>
      <p>
        When developing a crop yield forecasting model in India to determine whether a given climate
factor would affect yield using machine learning, a logistic regression model was found to be the
most accurate [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ]. The paper [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ] provides an overview of some of the existing supervised and
unsupervised machine learning models related to crop yield. Analytical models such as decision
trees, random forests, support vector machines, Bayesian networks, and artificial neural networks
are used to analyze the key factors impact on yield. These methods make it possible to analyze soil,
climate and water regimes that significantly affect crop growth and yield. The review [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ] presents
machine learning (ML) approaches from the point of view of an applied economist.
      </p>
      <p>
        The paper [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ] examines the impact of extreme values of climatic factors on global agricultural
yields. The paper [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ] aims to identify the best yield prediction model that can help farmers decide
which crop to grow based on climate conditions and nutrients present in the soil. In an analysis of
yield prediction by three different supervised machine learning models, the authors concluded that
the best accuracy was achieved with the Random Forest Classifier in both Entropy and Gini Criterion.
      </p>
      <p>
        The study [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ] proposed a machine learning-based forecasting system to forecast the yield of six
agricultural crops at the countries in West Africa. Climatic and weather data and agricultural yields
were combined to predict crop yields and build a decision support system for planning crop
plantings. To build such a system, decision tree, multivariate logistic regression and k-model of
nearest neighbors were used. It was found that the prediction results of the decision tree model and
the K-Nearest Neighbor model are correlated to the expected data.
      </p>
      <p>
        The structure of deep learning for forecasting yield using remote sensing data is presented in the
paper [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ]. An approach to dimensionality reduction based on histograms is proposed and the
structure of a deep Gaussian process is demonstrated, with the help of which spatially correlated
errors are eliminated and the accuracy of soybean yield forecasting (in a US county) is significantly
increased.
      </p>
      <p>
        One of the most powerful tools of machine learning is artificial neural networks. The paper [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ] uses
a semiparametric variant of a deep neural network, which can simultaneously account for complex
nonlinear relationships in high-dimensional datasets. Using data on corn yield from the US Midwest, it
was shown that this approach outperforms both classical statistical methods and fully non-parametric
neural networks in yield prediction.
      </p>
      <p>
        In a previous study [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ] Hrytsyuk et al. demonstrated that in terms of the influence of climate on
wheat yield, all regions of Ukraine are divided into three agro-climatic zones. Annual changes in
yields can be separated into a trend component and a deviation from the trend, explained by the
influence of climate. Application of the binarization method to the yield trend deviation facilitated
the development of machine learning classification models that can predict wheat yields with a
prediction horizon of three months.
      </p>
    </sec>
    <sec id="sec-3">
      <title>3. Methodology</title>
      <p>In our model, the impact of climate on wheat yield is quantified through the cumulative effects of
temperature and precipitation factors, each influencing distinct intervals of the growing season, as
delineated in Table 1. Our research is divided into two main parts. In the first part, we use correlation
and regression analyses to assess the effects of specific climatic factors -  1,  2, ⋯ ,  9,  10,  20,  30 on the
deviations of yields  from their expected trend values. This analysis results in a regression model that
is capable of predicting wheat yields for the current year.</p>
      <p>In the second part, we perform binarization of these yield deviations  . Each value of  is
transformed into a binary factor  1, which can be either 0 or 1. This binary classification enables
us to treat the data for a specific area and year as a sample that belongs to one of two categories:
high yield ( 1 = 0) or low yield ( = 1). This approach enables the application of machine
learning techniques to develop classification-based predictive models for wheat yields.</p>
      <sec id="sec-3-1">
        <title>3.1. Data Collection</title>
        <p>
          The main food crop in Ukraine is wheat. The average annual production of wheat in Ukraine for
2019-2021 reached the level of 26.5 million tons. The weight share of wheat in grain exports during
this time was 38%. This work is devoted to the study of the influence of climatic factors on
fluctuations in wheat yield in the steppe region of Ukraine. Statistical climate data and wheat yield
data for the period 2000-2021 for the Kherson, Mykolaiv, Odesa, Zaporizhzhya, Dnipro and
Kirovohrad regions, which are located in the steppe region of Ukraine, were used for this research.
Climatic characteristics were taken from [
          <xref ref-type="bibr" rid="ref16">16</xref>
          ], yield data were gotten from [
          <xref ref-type="bibr" rid="ref1">1</xref>
          ]. Successful wheat
vegetation in the period from April to June has a decisive impact on the crop yield [
          <xref ref-type="bibr" rid="ref4">4</xref>
          ]. Average
tenday temperature values of April, May, and June and monthly amounts of precipitation for this period
were used to assess the impact of climate on wheat yield (Table 1).
        </p>
      </sec>
      <sec id="sec-3-2">
        <title>3.2. Analysis of wheat yield dynamics</title>
        <p>
          An analysis of wheat yield dynamics in the regions of Ukraine over the past 22 years shows that the
yield is increasing [
          <xref ref-type="bibr" rid="ref1 ref4">1,4</xref>
          ]. The yield increase was the result of investment attractiveness increase of
the grain industry and the significant investment that has flowed into the industry. As a result, the
seed base has improved, agrotechnical culture has increased and the logistics network (elevators,
grain wagons, ports) has developed. In 2021 a record cereals and legumes crop was harvested in
Ukraine
        </p>
        <p>
          84 million tons. However, the tendency to increase grain yield is accompanied by
significant yield fluctuations, the cause of which is mostly the weather and climate factors impact. The
wheat yield dynamic in the Kherson region can serve as an illustration (Figure 1). The magnitude of
deviations from the trend (detrended yield) directly depends on the impact of climatic factors, the main
of which are droughts (2003 and 2012). To modeling of yield dynamics a linear trend model we used
Here a0, a1 - the trend coefficients, determined by statistical data using the least squares method [
          <xref ref-type="bibr" rid="ref17">17</xref>
          ].
An interval forecast is built on the linear trend basis, and for him the forecasting reliability level can
be established. To construct an interval forecast of yield, it is necessary to check the hypothesis about
a normal distribution of detrended yield eps
        </p>
        <p>
          Ten-day temperature values make it possible to more accurately take into account the impact of
external temperature at different stages of plant vegetation. Monthly precipitation amounts are used
because many ten-day precipitation amounts in the steppe zone are close to zero. Statistical
parameters of climatic factors and yield are given in Table 2. The parameter eps represents the
deviation of yield from the trend value. Its magnitude and sign are determined by the impact of
climatic factors on wheat yield in the current year.
lines are high and low yield boundaries. Author's calculations according to [
          <xref ref-type="bibr" rid="ref1">1</xref>
          ]
        </p>
        <p>To test the hypothesis of a normal distribution of detrended yields, a combined sample of
detrended yields for six regions of the steppe zone of Ukraine (132 observations) was used. Statistical
data on climate and wheat yield for the Kherson, Mykolaiv, Odesa, Zaporizhzhia, Dnipro and
Kirovohrad regions were used. Similar weather and climate conditions and soil type allow these
regions to be united into one homogeneous region. The hypothesis of a normal distribution of
detrended yields was confirmed by the Kolmogorov-Smirnov test.</p>
      </sec>
      <sec id="sec-3-3">
        <title>3.3. Binarization of detrended yield</title>
        <p>To solve many problems when planning an agrarian business, it is not necessary to have an accurate
yield forecast. For example, to make a decision about investing in a specific project, it is enough to
T
means a detrended yield value that is significantly lower than the average detrended yield value. All</p>
        <p>This approach enables the use of classification methods
in yield forecasting.</p>
        <p>We used the hypothesis of a normal distribution of detrended yield for the binary classification
of detrended The main task of this study
is to forecast low wheat yield values. To th
a probability of p &lt; 0.33 are located on the integral curve of the normal distribution of detrended
yields, that is, those for which the condition is fulfilled
 (</p>
        <p>) &lt; 0.33.</p>
        <p>Yield
implement a classification approach to yield prediction, a binary variable eps1 is introduced, which
has only two values: 1 ("low yield") and 0 ("high yield"). By the same time, the value of the eps1
variable is determined by the rule</p>
        <p>1,   ( ) &lt; 0.33; (4)
 1 = {</p>
        <p>0,   ( ) ≥ 0.33.</p>
        <p>According to the classification results, it was found that the</p>
        <p>25.5%, the number of cases
distribution of detrended yields is not a necessary condition for their classification. This hypothesis
only simplifies the classification procedure. The number of cases classified as 'low yield' represents
25%.
(3)</p>
      </sec>
      <sec id="sec-3-4">
        <title>3.4. Analysis of climatic factors impact on the wheat yield</title>
        <p>
          Such climatic factors as the average 10-day temperature and monthly precipitation cause fluctuations in
wheat yield relative to the trend. Therefore, assessing the climatic factors impact on grain yield is an
important tool when planning the placement of future crops and when planning future investments in
the agricultural sector [
          <xref ref-type="bibr" rid="ref18">18</xref>
          ].
wheat yield in Kherson region are shown in Table 3 As can be seen, the most noticeable impact on
the wheat yield is caused by the mean ten-day temperature in May and June and the total monthly
precipitation in April.
        </p>
      </sec>
      <sec id="sec-3-5">
        <title>3.5. The multiple linear regression model. Features selection</title>
        <p>
          According to the formulated assumptions, the wheat yield is formed under the impact of 12 climatic
factors (9 temperature and 3 related to precipitation). To build a model of such a relationship, we will
use the methods of multivariate correlation-regression analysis [
          <xref ref-type="bibr" rid="ref17">17</xref>
          ]. At the same time, the response
eps is connected through the multiple regression equation with the factor features t1, t2, t3, t4, t5, t6, t7, t8,
t9, R10, R20, R30. In the study, it is considered that climatic factors affect not the average yield, but the
deviation of the yield from trend value (detrended yield). Therefore, the detrended yield eps will be used
as response
        </p>
        <p>A linear multiple regression equation of the following form will be used to model the dependence:
Here 0 1 2 3 4 5 6 7 8 10 20 30 are model parameters; t1, t2, t3, t4, t5, t6, t7, t8, t9, R10, R20, R30
are model factors; eps is response;
determine model parameters.</p>
        <p>To build regression models with a large number of parameters, it is necessary to have large data
samples. For further research, it will be used the steppe region of Ukraine, which includes the
Kherson, Mykolaiv, Odesa, Zaporizhzhia, Dnipro and Kirovohrad regions. The corresponding data
sample contains 132 observations, each containing detrended yield and 12 climate factors. To process
such large data sets, it is advisable to use specialized software.</p>
        <p>We used the Python software environment and machine learning tools for data processing [19].
When studying statistic dependencies and developing a statistical model of a phenomenon, the
problem lies in choosing the algorithm that is optimal for a specific case. In recent decades, the
introduction of machine learning methods to solve the problems of classification and regression
(quantitative response prediction) has begun. These methods include: multiple regression method,
logistic regression method, linear discriminant analysis, random forest method, support vector
machines, artificial neural networks.</p>
      </sec>
      <sec id="sec-3-6">
        <title>3.6. Machine Learning Algorithms</title>
        <p>
          Recently machine learning-based systems are growing in popularity in research applications. In
particular, the classification is an essential form of data analysis that formulates models while
describing significant data classes [
          <xref ref-type="bibr" rid="ref19">20</xref>
          ]. In this work the several of classification algorithms for
categorical predicting of wheat yield were used.
        </p>
        <p>Logistic regression model. As noted above, to solve many problems when planning an agrarian
This approach enables the use of classification methods in yield forecasting. The detrended yield was
binarized according to the rule (4). As a result, a new data set, which differs from the one described
in section 3.1 by replacing the numerical factor eps with the categorical factor eps1 was gotten. Each
of the 132 observations of the new data set is characterized by a 12-dimensional feature vector.</p>
        <p>The logistic regression model looks like this</p>
        <p>P = F ( X  ') .</p>
        <p>F ( z ) =
1 + ez
.</p>
        <p>
          Here, F is a function whose values fall within the [
          <xref ref-type="bibr" rid="ref1">0, 1</xref>
          ] range and determine the probability P of a low
yield occurrence. To implement the function F, a logistic distribution function is usually used:
ez
Here, the parameter z is calculated from the ratio
 =  0 +  1 1 +  2 2 +  3 3 +  4 4 +  5 5 +  6 6 +  7 7 +  8 8 +  9 9 +  10 10 +  20 20 +
 30 30. (9)
        </p>
        <p>
          To choose the best model, it is necessary to estimate the value of the coefficients
regression model. Usually, the maximum likelihood method [
          <xref ref-type="bibr" rid="ref17">17</xref>
          ] is used for this.
        </p>
        <p>The logistic regression model allows you to classify the samples according to the rule
of the logistic
formula (7).</p>
        <p>In (10)</p>
        <p>1,   &gt; 0.5;
 = {0,   ≤ 0.5.
(7)
(8)
(10)</p>
        <p>
          Evaluation of classifiers. The following indicators are used for evaluating the performance of
the classifiers: matrix of errors (Confusion matrix), overall accuracy of classification (Accuracy),
sensitivity of classification (Sensitivity), specificity of classification (Specificity) and the area under
the ROC curve [
          <xref ref-type="bibr" rid="ref20">21</xref>
          ]. The Confusion matrix is built based on the results of classification by the model
and the actual belonging of observations to classes [19]. Four cases are distinguished in the matrix:
• TP (True Positives) the model correctly detected a low yield value;
• FP (False Positives) the model wrongly recognized a high yield as a low yield;
• FN (False Negatives) the model wrongly recognized a low yield as a high yield;
• TN (True Negatives) the model correctly identified a case of high yield.
        </p>
        <p>In the general case, the Confusion matrix has the following form (table 4):
which passes through the point cloud in such a way that projections onto it provide the best
resolution into two classes.</p>
        <p>Decision tree model. Decision trees used in data mining are of two main types: a classification
tree and a regression tree (the predicted result is a real number) [23]. Decision trees split the space
of objects according to some set of splitting rules. These rules make it possible to implement
sequential dichotomous data segmentation. At each partitioning step, the amount of information
about the variable under study (response) increases. When building a tree, it is important to set the
optimal branching level.</p>
        <p>The disadvantage of the decision tree method is instability: two trees built on the same training
sample can give completely different resulting classes. This shortcoming can be eliminated by
building ensembles of decision trees
bagging [24]. At the same time, several decision trees are built, repeatedly interpolating the data with
replacement (bootstrap), and as a consensus answer, it gives the result of the voting of the trees (their
average forecast). Boosting is another method for constructing a Random Forest [25]. Method of
support vectors. The basic idea of a support vectors classifier is to build a separating surface using
only a small subset of points that lie in the zone critical for separation, while other correctly classified
points of the training sample outside this zone are ignored by the algorithm [26]. Since there can be
many separating hyperplanes, the hyperplane that is the most distant from the training points is
selected from among them. Method of cross-validation. Even with a large data set and random
sampling has been applied to the training sample, the resulting model may be statistically unreliable.
After all, another set of samples can lead to another model, which is significantly different from the
first one. This shortcoming can be eliminated by cross-validation method [27].
the same size.</p>
        <p>2. One of the folds is selected as a data set for testing the model (testing set). The model is built based on
the data of the remaining k-1 folds that form the training set. The MSE test error based on the observations
of the testing set was calculated</p>
        <p>MSE =
1
n</p>
        <p>2
in=1( yi − yi ) .</p>
        <p>(12)
3. The process described above is repeated k times, each time using a different set as a testing set.
4. The total test MSE was calculated as the average of  test MSEs. Similarly, other parameters of
the model were averaged.</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>4. Results</title>
      <sec id="sec-4-1">
        <title>4.1. Linear regression model</title>
        <p>Linear regression model. We will build a linear regression model that will allow us to estimate the
influence of climatic factors on wheat yield fluctuations. The values of the linear regression model
estimates are shown in Table 5. The LM1 model is generally adequate (F-statistic=10.46; Prob
(Fstatistic) = 7.44e-14), but many factors in this model are insignificant (t1, t8, t9, P30). As can be seen
from the table, factors t3, t5, t7, R10 have the greatest influence on yield.
4.2. Logistic regression model
As described above, the excess yield over the trend eps can be translated into the binary form eps1
according to rule (3). This makes it possible to build classification models of yield forecasting. First,
let's build a GLM logistic regression model for a data set that describes 6 regions of the steppe region
of Ukraine. To increase the statistical significance of the model during its construction, the basic
principles of statistical modeling should be followed [28]. All data should be divided into two parts:
the training sample (most of the original data used to build the model) and the control sample (the
rest of the data that did not make it into the training sample). The control sample data are new
(unknown) to the built model, so they are used to evaluate the quality of the built model. In this
work, we used the ratio of the amount of data in the training and control sample as 75% to
25%. Based on the GLM model, a forecast is built on the test sample. The accuracy of the predictive
model presented in the table is 0.727 (24 out of 33 results matched).</p>
      </sec>
      <sec id="sec-4-2">
        <title>4.3. A random forest model</title>
        <p>A fragment of the decision tree of the problem is presented in Figure 2. At the first step, the algorithm
determines the most significant factor and builds a dichotomy rule for it. Such a rule is the logical
.785
the main factor affecting wheat yield. If its value exceeds 19.785°C, the yield is likely to be low. In
the next step, the obtained classes are again divided into subclasses according to another rule. This
makes it possible to clarify the general rule of classification. At the next stage, a group of trees is
combined into a random forest.</p>
        <p>Based on one of the random datasets, a model of a random forest of regression type is built using
all influencing factors. As can be seen from Figure 3, in order to achieve high accuracy in classifying
our data, it is necessary to use between 40 and 180 trees. This configuration of the model provides
its best parameters: classification accuracy 0.88, average classification error MSE = 0.121. The
importance of various traits for classification by the random forest method is illustrated in Figure 4.</p>
      </sec>
      <sec id="sec-4-3">
        <title>4.4. Comparison of the classification models effectiveness</title>
        <p>The study used six methods to build binary classification models: linear discriminant analysis (LDA),
support vector method with linear kernel function (SVML), support vector
method with radial kernel function (SVMR), decision tree method (CART), random forest method
(RF), logistic regression method (GLM). The Python software environment was used to develop all
models. The data were not standardized due to the same scale of indicators. In the SVMR model, the
type of kernel function was taken as default - a Gaussian kernel with a radial basis function (RBF).
The following parameters of the model were used: sigma = 0.4, C = 2. Here C means "the box
constraint level", sigma "kernel scale mode". Let's compare this models using the method described
in [29]. The stages of this technique
are as follows:
• Data Division: The initial dataset is split into two parts: 75% is designated for constructing
the training sample, and 25% is reserved for the control (test) sample.
• Model Training and Testing: The models are trained on the training sample and
subsequently used to classify the control sample. Among the tested machine learning methods,
the random forest method and the support vector method showed the best accuracy (Table 6).
• Cross-Validation Procedure: This process involves partitioning the initial dataset into
several equal groups, with one acting as the control group at a time. Each group serves as the
control group in rotation. During each cycle, the model is trained on the remaining data and
tested on the control group. At the end of the process, the average performance metrics for the
models are compiled. These metrics include accuracy (AC), sensitivity (SE), specificity (SP),
and the area under the ROC curve (AUC).</p>
        <p>A universal method for comparing the classifiers accuracy is ROC analysis [28]. ROC curves were
constructed for the three used models (SVL, SVR, and RF) based on the complete table of initial data,
which includes training and test samples (Figure 5). The area under the ROC curve AUC is a criterion
for evaluating the classifier. For an ideal classifier, the ROC curve has the shape of a right angle.
When evaluating the models by the AUC criterion, the best classifiers are the random forest method.</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>5. Conclusions</title>
      <p>The importance of this research is determined by the fact that today there is an insufficient number
of publications devoted to the impact of climate on the crop yield in Ukraine. There are even fewer
publications that investigate this problem using machine learning techniques. The task of this work
was to evaluate the climatic factors impact on detrended yield fluctuations using machine learning
methods. This approach requires a large amount of data. To solve this problem, the data of six regions
of the steppe zone, the climatic characteristics of which are similar, were combined.</p>
      <p>We have shown that for assessing the weather factors impact on yield, it is sufficient to use the
average ten-day values of temperature and monthly amounts of precipitation for the period from
April to June. Models that reflect the impact of climatic factors on detrended wheat yield were built.
It is shown that the temperature indicators in mid-May and early June and the amount of
precipitation in April commit the greatest influence on the yield.</p>
      <p>In this study, an approach was used, according to which trend deviations were divided into two
s reduced
to a classification problem. This approach, on the one hand, simplifies yield modeling, and on the
other hand, allows for high accuracy in the classification of yield values. Six machine learning models
were used as classifiers: discriminant analysis model, support vector models with linear and radial
kernel functions, decision tree model, random forest model, and logistic regression model. To
increase statistical significance, the cross-validation procedure with subsequent averaging of model
parameters was used. All models were fitted to the available data set and demonstrated classification
accuracy above 80% on test samples. The support vector method and logistic regression model
showed better accuracy in classifying real data than other methods and provide a forecasting
accuracy of 85% (on test data). This predicting accuracy is very good for complex natural processes.</p>
      <p>The classification models built in this work make it possible to estimate in advance (in 3 months)
the future wheat yield in te
investment and marketing decisions. Since the algorithms used in this study are entirely accessible
in terms of implementation, grain producers can use them for short-term yield forecasting. The
proposed method can be used to study the climate impact on the agricultural yield in other regions
and countries.</p>
      <p>Our research is a contribution to solving the problem of ensuring the sustainability of grain
production in Ukraine. The obtained results can be used to stabilize the economic development of
Ukraine and solve the food problem in the world.
[22] A. Afifi, S. Azen, Statistical Analysis, Second Edition: A Computer Oriented Approach.</p>
      <p>Academic Press, New York, 1979.
[23] L. Breiman, J.H. Friedman, R.A. Olshen, C.J. Stone, Classification and regression trees.</p>
      <p>Brooks/Cole Publishing, Monterey, 1984.
[24] L. Breiman, Bagging Predictors. Mach. Learn, 24, 1996, pp. 123 140.
[25] J. Friedman, Stochastic Gradient Boosting. Computational Statistics and Data analysis, 38 4
(2002) 367-378.
[26] T. Hastie, R. Tibshirani, J. Friedman, Model Assessment and Selection. The Elements of</p>
      <p>Statistical Learning, Springer Series in Statistics, 2009, pp. 219-259.
[27] D. Berrar, Cross-Validation. The Encyclopedia of Bioinformatics and Computational Biology,</p>
      <p>Academic Press, 2019.
[28] J. Gareth, D. Witten, T. Hastie, R. Tibshirani, An Introduction to Statistical Learning. Springer,</p>
      <p>New York, 2013.
[29] M. Kuhn, K. Johnson, Applied Predictive Modeling. Springer, New York, 2013.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>[1] State Statistics Service of Ukraine. URL: http://www.ukrstat.gov.ua.</mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>T.</given-names>
            <surname>Adamenko</surname>
          </string-name>
          ,
          <article-title>Climate change and agriculture in Ukraine: what farmers should know. German-Ukrainian agropolitical dialogue</article-title>
          , Zapovit, Kyiv,
          <year>2019</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3] Mitigation of Climate Change, Contribution of Working Group III to the
          <source>Sixth Assessment Report of the Intergovernmental Panel on Climate Change</source>
          . Cambridge University Press, Cambridge-New York,
          <year>2022</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>P.</given-names>
            <surname>Hrytsiuk</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Babych</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Baranovsky</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Havryliuk</surname>
          </string-name>
          ,
          <article-title>Assessing of Climate Impact on Wheat Yield using Machine Learning Techniques</article-title>
          .
          <source>CEUR Workshop Proceedings</source>
          ,
          <volume>3513</volume>
          ,
          <year>2023</year>
          , pp.
          <fpage>314</fpage>
          <lpage>329</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>D.</given-names>
            <surname>Muller</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Jungandreas</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Koch</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Shirhorn</surname>
          </string-name>
          ,
          <article-title>The impact of climate change on wheat production in Ukraine</article-title>
          .
          <source>Report on agricultural policy (APD)</source>
          ,
          <year>2016</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>H.M.</given-names>
            <surname>Zemankovics</surname>
          </string-name>
          ,
          <article-title>Mitigation and adaptation to Climate Change in Hungary</article-title>
          .
          <source>In: J. Central Eur. Agric</source>
          . Vol.
          <volume>13</volume>
          (
          <issue>1</issue>
          ),
          <year>2012</year>
          , pp.
          <fpage>58</fpage>
          -
          <lpage>72</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>J.</given-names>
            <surname>Ifft</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Kuhns</surname>
          </string-name>
          ,
          <string-name>
            <surname>K. Patrick,</surname>
          </string-name>
          <article-title>Can machine learning improve prediction an application with farm survey data</article-title>
          .
          <source>Int. Food and Agribus. Manag. Rev 21 8</source>
          (
          <year>2018</year>
          ) 1083
          <fpage>1098</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>B.</given-names>
            <surname>Sharma</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Yadav</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Yadav</surname>
          </string-name>
          ,
          <article-title>Predict Crop Production in India Using Machine Learning Technique: A Survey</article-title>
          , in: 8th International Conference on Reliability,
          <article-title>Infocom Technologies and Optimization (Trends and Future Directions)</article-title>
          , Noida, India,
          <year>2020</year>
          , pp.
          <fpage>993</fpage>
          <lpage>997</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>D.</given-names>
            <surname>Elavarasan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.R.</given-names>
            <surname>Vincent</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Sharma</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.Y.</given-names>
            <surname>Zomaya</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Srinivasan</surname>
          </string-name>
          ,
          <article-title>Forecasting yield by integrating agrarian factors and machine learning models: a survey</article-title>
          .
          <source>Comput. Electron. Agric</source>
          ,
          <volume>155</volume>
          ,
          <year>2018</year>
          , pp.
          <fpage>257</fpage>
          <lpage>282</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>H.</given-names>
            <surname>Storm</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Baylis</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Heckelei</surname>
          </string-name>
          ,
          <article-title>Machine learning in agricultural and applied economics</article-title>
          .
          <source>Eur. Rev. Agric. Econ</source>
          ,
          <volume>47 3</volume>
          (
          <year>2020</year>
          ) 849
          <fpage>892</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>E.</given-names>
            <surname>Vogel</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.G.</given-names>
            <surname>Donat</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.V.</given-names>
            <surname>Alexander</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Meinshausen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.K.</given-names>
            <surname>Ray</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Karoly</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Meinshausen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Frieler</surname>
          </string-name>
          ,
          <article-title>The effects of climate extremes on global agricultural yields</article-title>
          .
          <source>Environ. Res. Let., 14</source>
          <volume>5</volume>
          (
          <year>2019</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>M.</given-names>
            <surname>Kalimuthu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Vaishnavi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Kishore</surname>
          </string-name>
          ,
          <article-title>Crop Prediction using Machine Learning</article-title>
          ,
          <source>in: 2020 Third International Conference on Smart Systems and Inventive Technology (ICSSIT)</source>
          . Tirunelveli, India,
          <year>2020</year>
          , pp.
          <fpage>926</fpage>
          <lpage>932</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <given-names>L.S.</given-names>
            <surname>Cedric</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W.Y.</given-names>
            <surname>Hamilton</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Aworkaa</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.T.</given-names>
            <surname>Zoueucd</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.K.</given-names>
            <surname>Mutomboa</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Krichen</surname>
          </string-name>
          ,
          <article-title>Crops yield prediction based on machine learning models: Case of West African countries</article-title>
          .
          <source>Sm. Agri. Tech</source>
          ,
          <volume>2</volume>
          (
          <year>2022</year>
          ) 1
          <fpage>14</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <given-names>J.</given-names>
            <surname>You</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Low</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Lobell</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Ermon</surname>
          </string-name>
          ,
          <article-title>Deep gaussian process for crop yield prediction based on remote sensing data</article-title>
          ,
          <source>in: the Proceedings of the AAAI Conference on Artificial Intelligence</source>
          ,
          <fpage>31</fpage>
          <lpage>1</lpage>
          (
          <year>2017</year>
          )
          <fpage>4559</fpage>
          -
          <lpage>4565</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <given-names>A.</given-names>
            <surname>Crane-Droesch</surname>
          </string-name>
          ,
          <article-title>Machine learning methods for crop yield prediction and climate change impact assessment in agriculture</article-title>
          .
          <source>Environ. Res. Lett., 13</source>
          <volume>11</volume>
          (
          <year>2018</year>
          ) 1
          <fpage>12</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [16]
          <article-title>Meteorological data archive</article-title>
          . URL: https://meteopost.com/weather/archive/
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [17]
          <string-name>
            <given-names>N.R.</given-names>
            <surname>Draper</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Smith</surname>
          </string-name>
          ,
          <source>Applied Regression Analysis. 3th Edition</source>
          , Wiley, New York,
          <year>1998</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          [18]
          <string-name>
            <given-names>P.</given-names>
            <surname>Hrytsiuk</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Babych</surname>
          </string-name>
          ,
          <string-name>
            <given-names>O.</given-names>
            <surname>Mandziuk</surname>
          </string-name>
          ,
          <article-title>Region sown areas portfolio optimization taking into account crop production economic risk</article-title>
          .
          <source>Global Journal Environemental Science Management</source>
          ,
          <volume>5</volume>
          (
          <year>2019</year>
          )
          <fpage>140</fpage>
          -
          <lpage>150</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          [20]
          <string-name>
            <given-names>S.</given-names>
            <surname>Ramaswamy</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Rastogi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Shim</surname>
          </string-name>
          ,
          <article-title>Efficient algorithms for mining outliers from large data sets</article-title>
          ,
          <source>in: Proceedings of the 2000 ACM SIGMOD international conference on Management of data</source>
          ,
          <year>2000</year>
          , pp.
          <fpage>427</fpage>
          <lpage>438</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          [21]
          <string-name>
            <given-names>T.</given-names>
            <surname>Fawcett</surname>
          </string-name>
          ,
          <article-title>An Introduction to ROC Analysis. Pattern Recognit</article-title>
          .
          <source>Lett., 27</source>
          <volume>8</volume>
          (
          <year>2006</year>
          ) 861
          <fpage>874</fpage>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>