<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Kirill Smelyakov 1, Oleksandr Bizkrovnyi 1, Natalia Sharonova 2, Serhii Smelyakov 1, Anastasiya Chupryna 1</article-title>
      </title-group>
      <contrib-group>
        <aff id="aff0">
          <label>0</label>
          <institution>Kharkiv National University of Radio Electronics</institution>
          ,
          <addr-line>14 Nauky Ave., Kharkiv, 61166</addr-line>
          ,
          <country country="UA">Ukraine</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>National Technical University "KhPI"</institution>
          ,
          <addr-line>Kyrpychova str. 2, Kharkiv, 61002</addr-line>
          ,
          <country country="UA">Ukraine</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>This article is an investigation of factors that can affect cryptocurrency price and their usage in regression models to determine which model type and the algorithm itself is best-suited for predicting crypto price. The determination of the best algorithm is based on an experiment that includes training and validation of models. Comparative analysis of models validation results defines the best-suited algorithm; the type of cryptocurrency that is being analyzed is Defi, namely Ethereum; the study is based on a one-year time frame; the paper does not consider political factors and factors of infrastructure destruction that may affect cryptocurrency prices. The factor types, which are used to create regression models, consist of fundamental factors. The technical factors were omitted and can be investigated in other works. Factors include: network statistics, exchange statistics, mining statistics, social statistics, transactions data, etc. The models performance is calculated by regression metrics. The JMH is used to calculate models time to train.</p>
      </abstract>
      <kwd-group>
        <kwd>1 Cryptocurrency</kwd>
        <kwd>machine learning</kwd>
        <kwd>price forecasting</kwd>
        <kwd>prediction model</kwd>
        <kwd>impacting factors for cryptocurrency price</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>Cryptocurrency and bitcoin in particular, has demonstrated its value in recent years, and there are
now 14 million bitcoins in circulation. Investors speculating on the future possibilities of this new
technology have provided much of the current market capitalization, and this will likely continue until
a certain degree of price stability and market acceptance is achieved. Beyond the announced price of a
cryptocurrency, those who invest in it rely on the perceived "intrinsic value" of the cryptocurrency. This
includes the technology itself and the network, the integrity of the cryptographic code and the
decentralized network. Blockchain public ledger technology (the underlying cryptocurrency) is capable
of disrupting a range of transactions beyond the traditional payment system. These include stocks, bonds
and other financial assets whose records are stored digitally and for which there is currently a need for
a trusted third party to validate the transaction. At present, a huge number of models, algorithms and
technologies have been developed to improve the speed of fraud detection, mining efficiency,
cybersecurity and privacy, as well as to improve the efficiency of price forecasting, volatility, portfolio
volume and structure, etc. At the same time, algorithms to solve these problems are often unsustainable
because they do not take into account a number of important influencing factors. In this regard, it is
now relevant to make a deeper analysis of the factors that have an impact on the price formation of
cryptocurrency in order to build regression models to predict prices.</p>
    </sec>
    <sec id="sec-2">
      <title>2. Related Works</title>
      <p>
        A cryptocurrency is a software with a specific way to use it which allows including them into the
currencies market and do trading. The papers [
        <xref ref-type="bibr" rid="ref1 ref2 ref3">1-3</xref>
        ] present a modern review of Cryptocurrency systems,
Models and algorithms. In particular, the work [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ] has comparative analysis of mining algorithms, in
[
        <xref ref-type="bibr" rid="ref2 ref3">2, 3</xref>
        ] the main Challenges and Opportunities are formulated, as well as an analysis of the basic methods
of artificial intelligence, which are used to solve the most important problems of cryptocurrency, related
to price forecasting, risks, cybersecurity threats and a number of others. The main idea of the “coin”
type of cryptocurrency is the ability to prepare anonymous transactions [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ]; work [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ] shows the features
of the model of decentralized confidential payment system. Furthermore (major restriction regarding
the type of cryptocurrency), other types of cryptocurrencies exist, but this research orients just to the
investigation of the factors that make an impact on the “coin" type of cryptocurrency [
        <xref ref-type="bibr" rid="ref6 ref7 ref8">6-8</xref>
        ].
      </p>
      <p>
        All of the cryptocurrencies are software and their price is a difficult analysis of many factors [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ].
Globally, there are two parts which contain their own factors: product itself (cryptocurrency as product
and their mechanism) and trading market. Nothing lives outside the environment and crypto is not an
exclusion. The relationship between the choice of factors, models and algorithms of blockchain
cryptocurrency ecosystem functioning are described in papers [
        <xref ref-type="bibr" rid="ref10 ref11 ref12">10-12</xref>
        ]. Crypto is a common software,
so general rules that affect any product in this area can affect cryptocurrency too. Examples of such
functions are described in [
        <xref ref-type="bibr" rid="ref13 ref14 ref15">13-15</xref>
        ]: each product has competitors; each product should give some
specific features to survive; each product depends on the buying ability of potential customers, etc.
      </p>
      <p>
        A cryptocurrency is a specific software with a peculiar mechanism of their work. In general, all of
the cryptocurrencies have users and transactions validators, the role that validators play miners, which
mine each of the next blocks of the blockchain. As a result, there are two role needs of which need to
be addressed. If users will not have the ability to use crypto coins, then crypto will die [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ]. Another
part of that, If the miners will not have required profit to cover all of the costs, then transactions will be
approved with huge delay, which leads to decreasing popularity of crypto and as a result, this reason
might be as root cause of cryptocurrency to go from the market [
        <xref ref-type="bibr" rid="ref17">17-19</xref>
        ].
      </p>
      <p>The solution of these problems is decisive for the effective application of private models of artificial
intelligence, machine learning and computer vision [20, 21], including using these models to improve
efficiency algorithms for the formation and processing of network information [22-24].</p>
      <p>There are many approaches that use different factors to predict the crypto price. There is an attempt
to predict price using GRU, LSTM and bi-LSTM Machine Learning Algorithms [25]. This approach
uses the following factors for training data set:
 Open price;
 High price;
 Low price;
 Close price;
 Date.</p>
      <p>Factors selection is not enough confident, because the price factors do not give the root cause of
values for given factors for particular date, in other words, the factors are not descriptive. Despite on
this assumption the model validation process shows the following results (Figure 1, Figure2).</p>
      <p>There is no description in the article of how the models were trained and validated, but historic
price values never can be used for future price prediction.</p>
      <p>Another work [26] uses technical metrics of cryptocurrency for price prediction. All of the article
frameworks attempt to predict the Bitcoin prices starting from five technical indicators:
 Simple Moving Average (SMA);
 Exponential Moving Average (EMA);
 Momentum (MOM);
 Moving Average Convergence Divergence (MACD);
 Relative Strength Index (RSI).</p>
      <p>One of the ML frameworks is described below (Figure 3).</p>
      <p>Technical indicators also cannot be a comprehensive data source for ML model training, because
technical analysis does not live without fundamental analysis which includes: network statistics,
exchanges statistics, worldwide economic state, etc. The main goal of this investigation is to determine
the informative factors that can be used for the price forecasting, determine the effectiveness of the
usage of regression machine learning models [27] and figure out the best suited algorithm for the
mentioned problem.</p>
    </sec>
    <sec id="sec-3">
      <title>3. Methods and Materials</title>
    </sec>
    <sec id="sec-4">
      <title>3.1. Data Description</title>
      <p>Consider initial data for the methods and experiments, metrics, factors and methods proposed to
solve the problem under consideration.</p>
      <p>The problem is characterized by time series data because crypto metrics and worldwide economic
indices are being updated each day. This dataset includes only worldwide economic data and metrics
data for a particular cryptocurrency. The exclusion here is data that can be spread between countries.
The crypto price is worldwide, as a result, no data per country can be used in a data set.</p>
      <p>The factors used in the dataset are not the same dates. While regular exchanges work only 5 days a
week during business hours, crypto exchanges work every day around the clock. This fact leads to a
need for data preparation. Values that are found at the end of the week are used during the weekends
for crypto exchanges.</p>
      <p>Another data issue that was found is data missing in some places. This issue is resolved by deleting
the full row to avoid the creation of incorrect relationships between factors. In case the dataset is large,
the gaps recovering is possible, but when the dataset is small, each data row is important to create valid
relationships. The dataset cannot be found on the Internet in public access. It consists of different parts
that are being retrieved from different sources in CSV format and combined using "Spark" after.</p>
      <p>Here are the examples of the data that formed the dataset (Table 1, Table 2).</p>
      <p>The “Price” column is not used during forming the dataset, because a separate CSV file with crypto
price and date exists. The dataset is available by the following link [28].</p>
      <p>About data source. There are a huge amount of different ways for retrieving particular info regarding
particular cryptocurrencies. For example, to retrieve network and mining data, users can run nodes for
a particular cryptocurrency and aggregate required information. This way is time and resource
consumptions.</p>
      <p>The social media data can be retrieved from required portals such as GitHub, Telegram, Twitter, in
a direct way, but these are raw data that requires preparing for extracting data sentiment.</p>
      <p>In general, many services avoid such resources and time costs. The “IntoTheBlock” was selected as
the source data provider because this system allows a 7-day trial and already has all precomputed and
aggregated data that was mentioned above.</p>
      <p>Also, worldwide economic information for the dataset can be retrieved from the following resource
[29].</p>
    </sec>
    <sec id="sec-5">
      <title>3.2. Informative Factors’ Selection</title>
      <p>The mentioned above global factors give understanding which of them may signalize the price
direction. There is the following list of the factors from different areas that were selected to build ML
models.</p>
    </sec>
    <sec id="sec-6">
      <title>3.2.1. Worldwide Data</title>
      <p>There are factors that are not directly related to cryptocurrencies, but allegedly affect them. These
are global economic factors that demonstrate global economic behavior that may affect the demand for
cryptocurrency.</p>
      <p>S&amp;P 1200 – factor means the global economic situation in the world based on indexes of 1200
biggest companies. This factor was included as an index of investor buying ability. If the world
economic situation will be better, then more investors may spend funds to buy such volatile investments
as cryptocurrency.</p>
      <p>Dow Jones Global – factor means the economic state of industrial companies. This factor also can
be used as mentioned before.</p>
    </sec>
    <sec id="sec-7">
      <title>3.2.2. Ethereum Data</title>
      <p>The following factors are closely related to cryptocurrencies and their work principles. In general,
these metrics show different aspects inside cryptocurrencies during their lifecycle.
Inflow volume – total amount (in $ or tokens) entering exchange(s) deposit wallets. All exchanges refer
to all supported exchanges. The sharp jumps of inflows tend to coincide with and sometimes precede
periods of high volatility. This can potentially be interpreted as a sign of holders looking to sell in
centralized exchanges.</p>
      <p>Outflow volume – total amount (in $ or tokens) leaving exchange(s) withdrawal wallets. All
exchanges refer to all supported exchanges. Outflow Volume often spikes following either a crash or a
significant break-out. This could potentially be interpreted as users going long and opting to hold their
crypto outside centralized exchanges.</p>
      <p>ETH Price – a dependent variable in the regression model. Means Ethereum price.</p>
      <p>ETH – BTC correlation – factor is used to display correlation between prices of largest
cryptocurrencies. If some product loses the buyer's confidence, then other products in that sphere may
lose it too.</p>
      <p>Large transaction – indicator shows transactions where an amount greater than $100,000 USD was
transferred.</p>
      <p>Transaction volume is USD – large transactions are those where an amount greater than $100,000
USD was transferred. In this case, the Large Transactions Volume in USD indicator measures the
aggregate dollar amount transferred in such transactions. Large transaction Volume metric shows the
total amount transacted by whales players in a given day. This indicator may give an idea of changes in
the cryptocurrency market if huge amounts of crypto volume transfers between addresses.
cryptocurrency quickly grows.
may indicate upcoming market moves.</p>
      <p>Transaction count – indicator displays activity in blockchain networks which can show the general
market behavior. If the transaction count increases, then popularity of the industry or particular
cryptocurrency rises.</p>
      <p>Miners’ inflows – indicators may point to general miner activity and how much they earn. Huge
amount of miner inflow can mean an increased need for them.</p>
      <p>Miners’ outflows – indicator may point to miner behavior, when they sell their crypto holdings into
exchanges.
confirmation.</p>
      <p>Miners’ reward – metric describes the miner’s reward. In case if reward is low, then crypto currency
may be stuck with a long transaction confirmation delay. Also, if the huge reward consists of fees that
users pay, then popularity of crypto may decrease.</p>
      <p>Average transaction fees – metric can point to increasing cryptocurrency demand. In case, when a
huge amount of transactions exists in the queue, customers start to pay extra fee to up their transaction</p>
      <p>Average transaction volume – transaction volume can indicate both trading and non-speculative
activity. Similar to trading volume observed in exchanges, transaction volume can be useful for
identifying reversals and breakouts.</p>
      <p>GitHub activity is a couple of indicators that refer to this: opened issues, closed issues, watchers
count, forks count, opened and closed pull requests count. This may point to an idea of how the
Search trends – indicates how often cryptocurrency rises to the spotlight. The increased attention
Telegram sentiment – indicator helps measure traders' emotions. In the case of bitcoin, positive
sentiment on Telegram has on several occasions preceded a price movement, as seen in December 2019,
April, and June 2020. At the same time, the percentage of messages perceived as negative tends to
increase during market crashes.</p>
      <p>Finally, the total number of messages is indicative of the level of activity in these group chats. It is
not necessarily reflected in crypto-activity, but it is worth noting the fluctuations in it as a rough
indicator of community engagement.</p>
      <p>Twitter sentiment – indicate a measurement of the emotions of market participants. Sometimes
sentiment can be a leading indicator, as was the case with Ethereum in June and July. In most cases,
however, sentiment tends to be a reactive indicator. In other words, there is more positive sentiment
when prices are rising and negative sentiment when prices are falling.
3.3.</p>
    </sec>
    <sec id="sec-8">
      <title>ML Model Validation and Metrics</title>
      <p>The correctness of the created regression model is a relative value. Despite on how the model
validness is determined for classification problems, the regression validation does not include
determining count of the “false negative” values. The regression model validity is defined by the
following metrics.
where  ̂– predicted value,  – real value.</p>
      <p>The R2 metric decided to not include into model validation metric set. Despite the same
Rsquared statistic produced, the predictive validity would be rather different depending on what the true</p>
      <p>Mean absolute error</p>
      <sec id="sec-8-1">
        <title>Mean square error</title>
      </sec>
      <sec id="sec-8-2">
        <title>Root mean square error Explained Variance</title>
        <p>( −̂),
( )
where N – count of record in the test dataset,  ̂– predicted value,  – real value.
where N – count of record in the test dataset,  ̂– predicted value,  – real value.
(1)
(2)
(3)
(4)
dependency is. If it is truly linear, then the predictive accuracy would be quite good. Otherwise, it will
be much poorer. In this sense, R-Squared is not a good measure of predictive error.</p>
      </sec>
    </sec>
    <sec id="sec-9">
      <title>3.4. ML Models and Methods</title>
      <p>There are bunches of different machine learning regression algorithms that can be selected. First of
all, regression model selection should be based on requirements and data specifications based on which
the model will be trained and validated.</p>
      <p>This research focuses on time series data, because cryptocurrency price changes continuously,
depending on the selected time frame. There are many types of regression models, but this article
focuses on a nonlinear model.</p>
      <p>There are a few reasons for these algorithms types selection:
 Logistic regression is not suitable because that algorithm has only two values (1, 0) for
dependent variable. Investigated problem has time series data;
 The relationships between data variables that was explained above is not linear, because
increasing one variable in couple of decreasing another one variable may affect the dependent
variable in an unpredictable way. That means that usage of linear regression algorithms can lead to
incorrect data fitting;
 Polynomial regression models a non-linear dataset using a linear model. It works in a similar
way to multiple linear regression (which is just linear regression but with multiple independent
variables) but uses a non-linear curve. It is used when data points are present in a non-linear fashion
[30]. This algorithm is not fit for the current purpose, because there is a need to make rules or
decisions instead of calculating an average between all points to “draw line”.</p>
      <p>The best way here is usage of non-linear regression algorithms. There are three representers of
nonlinear algorithms will be used:
 Decision trees;
 Random forest;
 Gradient boosted trees.</p>
      <p>The decision tree regression has as a main function is to split the dataset into smaller sets. The subsets
of the dataset are created to plot the value of any data point that connects to the problem statement. The
splitting of the data set by this algorithm results in a decision tree that has decision and leaf nodes. ML
experts prefer this model in cases where there is not enough change in the data set [31].</p>
      <p>The decision tree algorithm has hyperparameters for model tuning. One of them is tree depth. If the
maximum depth of the tree is set too high, the decision trees learn too fine details of the training data
and learn from the noise, i.e. they overfit (Figure 4).</p>
      <p>As a result, to make a good fitted model, the optimal count of tree depth is required. The optimal
count of depth can be found experimentally. If the regression model validity metrics have a good
performance for the training dataset, but on the test dataset is low, then the model overfitting happens.
There is a need to decrease the tree's depth.</p>
      <p>The random forest is also a widely-used algorithm for non-linear regression in Machine Learning.
Unlike decision tree regression (single tree), a random forest uses multiple decision trees for predicting
the output. Random data points are selected from the given dataset (say k data points are selected), and
a decision tree is built with them via this algorithm. Several decision trees are then modeled that predict
the value of any new data point. There is an example on Figure 5 of how a random forest algorithm
works.</p>
      <p>Since there are multiple decision trees, multiple output values will be predicted via a random forest
algorithm. You have to find the average of all the predicted values for a new data point to compute the
final output. The only drawback of using a random forest algorithm is that it requires more input in
terms of training. This happens due to the large number of decision trees mapped under this algorithm,
as it requires more computational power [32].</p>
      <p>The Gradient Boosted Regression Trees (GBRT) model (also called Gradient Boosted Machine or
GBM) is one of the most effective machine learning models for predictive analytics, making it an
industrial workhorse for machine learning. The Boosted Trees Model is a type of additive model that
makes predictions by combining decisions from a sequence of base models. For boosted trees model,
each base classifier is a simple decision tree. This broad technique of using multiple models to obtain
better predictive performance is called model ensembling. Unlike Random Forest which constructs all
the base classifier independently, each using a subsample of data, GBRT uses a particular model
assembling technique called gradient boosting [34].</p>
    </sec>
    <sec id="sec-10">
      <title>4. Experiment</title>
      <p>The main goal of this experiment is training of selected regression models, and determining which
of the models is the best for small amounts of data. The experiment consists of two steps:
 Experiment planning;
 Results overview.</p>
    </sec>
    <sec id="sec-11">
      <title>4.1. Experiment Planning</title>
      <p>First of all, experiment requires creation of the dataset. This point is achieved by manually
downloading the sources and making hierarchical folder structure for convenient files accessing. All of
the source files have a column which describe a day when event is happened, and this column is used
to merge each of the source files into full dataset.</p>
      <p>The dataset is being created or combined from different sources using Apache Spark Framework
abilities.</p>
      <p>Apache Spark is a unified analytics engine for large-scale data processing. It provides high-level
APIs in Java, Scala, Python and R, and an optimized engine that supports general execution graphs. It
also supports a rich set of higher-level tools including Spark SQL for SQL and structured data
processing, MLlib for machine learning, GraphX for graph processing, and Structured Streaming for
incremental computation and stream processing.</p>
      <p>The next step is splitting the dataset into two parts in the ratio of 70% and 30%. This action is
required to train and test our model.</p>
      <p>The spark.mllib supports decision trees for binary and multiclass classification, as well as for
regression, using both continuous and categorical features. The implementation splits the data by rows,
which allows distributed training with millions of instances.</p>
      <p>Random forests are ensembles of decision trees. Random forests are one of the most successful
machine learning models for classification and regression. They combine multiple decision trees to
reduce the risk of overshoot. Like decision trees, random forests work with categorical features, extend
to multi-class classification, do not require feature scaling, and can account for non-linearity and feature
interaction. The spark.mllib supports random forests for binary and multi-class classification as well as
regression, using both continuous and categorical features. The spark.mllib implements random forests
using an existing decision tree implementation.</p>
      <p>Gradient-Boosted Trees (GBTs) are ensembles of decision trees. GBTs iteratively train decision
trees to minimize the loss function. Like decision trees, GBTs work with categorical features, extend to
multi-class classification, do not require feature scaling, and are able to account for non-linearities and
feature interactions.</p>
      <p>The spark.mllib supports GBT for binary classification and regression using both continuous and
categorical features. The spark.mllib implements GBT using an existing decision tree implementation.
For more information about decision trees, see the decision tree guide.</p>
      <p>Note that GBTs do not yet support multi-class classification. Use decision trees or Random Forests
to solve multi-class problems.</p>
      <p>After the dataset is created, ML models will be trained and validated using metrics that were
mentioned above. An important note is a calculation of training time for each of the models during the
training process to determine the fastest model for the proposed dataset.</p>
      <p>The next step of the experiment is the validation of trained models, providing info about that and
determining the correctness of assumption that selected factors have relationships with ETH price.
4.2.</p>
    </sec>
    <sec id="sec-12">
      <title>ML Models Training</title>
      <p>The data that is selected to train the regression models is in a one-year time frame, because that time
period has robust rules of market behavior for particular cryptocurrencies. The models training process
that will be recorded below will be prepared with the best suited hyperparameters for this data. The
Decision Tree max depth: 30, The Random Forest max depth: 5, Gradient Boosted Tree: max iterations:
10, loss type: absolute, max depth: 5, subsampling rate: 0.4.</p>
      <p>In general, data is grained by day, so there are 365 examples of how the indicators affect the
Ethereum price. This is a small amount of data to make robust forecasting, but the experiment will give
more valuable facts of that assumption. The Java 8, Spark framework and MacOS Monterey were
selected to build and validate the regression models. Also, hardware consists of 2,6 GHz Quad-Core
Intel Core i7, 16 GB 2133 MHz LPDDR3.</p>
      <p>The microbenchmark for measuring how long the models are being trained was prepared by Java
JMH Benchmark. JMH is a Java library for writing benchmarks on the JVM, developed as part of the
OpenJDK project. JMH provides a very solid foundation for writing and executing benchmarks whose
results will not be corrupted by unwanted virtual machine optimizations. There are the following
benchmark modes, check Table 3.</p>
      <p>Description
Measures the number of operations per second, meaning the
number of times per second your benchmark method could be</p>
      <p>executed.</p>
      <p>Measures the average time it takes for the benchmark method to</p>
      <p>execute (a single execution).</p>
      <p>Measures how long time it takes for the benchmark method to</p>
      <p>execute, including max, min time etc.</p>
      <p>Measures how long time a single benchmark method execution
takes to run. This is good to test how it performs under a cold</p>
      <p>start (no JVM warm up).</p>
      <p>Measures all of the above.
The experiment shows that the quickest algorithm for that amount of data is a Random Forest regressor.</p>
    </sec>
    <sec id="sec-13">
      <title>5.1. Decision Tree Model</title>
    </sec>
    <sec id="sec-14">
      <title>5. Results</title>
      <p>This investigation part contains aggregated obtained results during testing of created regression
models and provides it in a table view.</p>
      <p>The next table (Table 5) represents data that is retrieved during Decision Tree model validation.</p>
    </sec>
    <sec id="sec-15">
      <title>5.2. Random Forest Regression Model</title>
      <p>The next table (Table 7) represents data that is retrieved during Random forest regression model
validation.</p>
    </sec>
    <sec id="sec-16">
      <title>5.3. Gradient Boosted Trees</title>
      <p>The next table (Table 9) represents data that is retrieved during Gradient Boosting Trees regression
model validation.</p>
      <p>As we can show, the Table 10 gives an example of output that Gradient Boosted Trees regression
model produces.</p>
      <p>This algorithm is the most accurate from each other, but training time is an issue here.</p>
    </sec>
    <sec id="sec-17">
      <title>6. Discussions</title>
      <p>There are the following aggregated results in charts that display which of the algorithms is the best
suited for a particular problem.</p>
      <p>All charts represent a comparison of algorithms by particular regression accuracy metric. These
charts include the following regression validation metrics:
 Root Mean Squared Error (RMSE);
 Mean absolute error (MAE);
 Mean square error (MSE);
 Explained Variance.</p>
      <p>In the end of this section, the general conclusions regarding usage results of mentioned regression
algorithms are extracted.</p>
      <p>The Figure 6 shows algorithms comparison by RMSE metric, which means differences between
values (sample or population values) predicted by a model or an estimator and the values observed. The
RMSD represents the square root of the second sample moment of the differences between predicted
values and observed values or the quadratic mean of these differences.</p>
      <p>The smallest error is observed for Random Forest algorithm, along with the smallest training time.</p>
      <p>The next chart (Figure 7) displays the comparison between algorithms by MAE metric which refers
to the magnitude of difference between the prediction of an observation and the true value of that
observation. MAE takes the average of absolute errors for a group of predictions and observations as a
measurement of the magnitude of errors for the entire group. MAE can also be referred as L1 loss
function. The results are the same as previous: Random Forest algorithm has highest accuracy, Gradient
Boosted Trees is on the second place and Decision Tree is the last.</p>
      <p>The next chart (Figure 8) displays the comparison between algorithms by MSE metric which defines
as Mean or Average of the square of the difference between actual and estimated values.</p>
      <p>The Random forest algorithm has the smallest value by that metric. The next chart (Figure 9)
displays results for Explained variance (also called explained variation) is used to measure the
discrepancy between a model and actual data. In other words, it’s the part of the model’s total variance
that is explained by factors that are actually present and aren’t due to error variance.</p>
    </sec>
    <sec id="sec-18">
      <title>7. Conclusions</title>
      <p>The experiment for determining the factors which affect crypto price was setup in this work. Related
works give understanding what was already done in scope of this theme. The statement regarding not
exhaustive of reviewed works was extracted. The own opinion for investigation was proposed. This
assumption is based on relationships between the following metrics: crypto data and worldwide metrics.
The assumption of relationships existence here is based on dependencies in the world. It means
dependencies between software and how this software is used in real life. Which affects software usage.</p>
      <p>The next step was determining the best suited family of regression algorithms, and the non-linear
family was selected. Also, the metrics which are required for validation were selected too.</p>
      <p>The experiment execution was the next step. And the Apache Spark was used to create data set from
source files and for creating and training the regression models. The JMH tool determined that the
random forest algorithm is fastest between others in time elapsed for training perspective.</p>
      <p>The result of validation of created models gives information that Random Forest algorithm is the
most accurate between each other, also as a training time is smallest in comparison to other algorithms.
The Gradient Boosted Tree algorithm stays in the middle of performance and Decision Tree algorithm
does not suite for prepared data and problem.</p>
      <p>Further investigations may be focused on including into model additional cryptocurrency factors:
spreading crypto between exchanges, more financial factors for the worldwide economic and political
situation like country financial institute openness which describes its ready to economic development
and infrastructure failures.</p>
    </sec>
    <sec id="sec-19">
      <title>8. References</title>
      <p>[18] X. Li and C. A. Wang, "The technology and economic determinants of cryptocurrency exchange
rates: The case of bitcoin", Decision Support Systems, vol. 95, pp. 49-60, 2017.
[19] A. Park, J. Kietzmann, L. Pitt and A. Dabirian, "The Evolution of Nonfungible Tokens:
Complexity and Novelty of NFT Use-Cases," in IT Professional, vol. 24, no. 1, pp. 9-14, 1
Jan.Feb. 2022, doi: 10.1109/MITP.2021.3136055.
[20] K. Smelyakov, A. Chupryna, M. Hvozdiev and D. Sandrkin, "Gradational Correction Models
Efficiency Analysis of Low-Light Digital Image," 2019 Open Conference of Electrical, Electronic
and Information Sciences (eStream), 2019, pp. 1-6, doi: 10.1109/eStream.2019.8732174.
[21] K. Smelyakov, M. Shupyliuk, V. Martovytskyi, D. Tovchyrechko and O. Ponomarenko,
"Efficiency of image convolution," 2019 IEEE 8th International Conference on Advanced
Optoelectronics and Lasers (CAOL), 2019, pp. 578-583, doi:
10.1109/CAOL46282.2019.9019450.
[22] O. Lemeshko, M. Yevdokymenko, O. Yeremenko, A. M. Hailan, P. Segeč and J. Papán, "Design
of the Fast ReRoute QoS Protection Scheme for Bandwidth and Probability of Packet Loss in
Software-Defined WAN," 2019 IEEE 15th International Conference on the Experience of
Designing and Application of CAD Systems (CADSM), 2019, pp. 1-5, doi:
10.1109/CADSM.2019.8779321.
[23] Ageyev D., Radivilova T. Traffic monitoring and abnormality detection methods for decentralized
distributed networks // CEUR Workshop Proceedings. 2021. Vol. 2923. P. 283–288.
[24] K. Smelyakov, A. Datsenko, V. Skrypka and A. Akhundov, "The Efficiency of Images Reduction
Algorithms with Small-Sized and Linear Details," 2019 IEEE International Scientific-Practical
Conference Problems of Infocommunications, Science and Technology (PIC S&amp;T), 2019, pp.
745750, doi: 10.1109/PICST47496.2019.9061250.
[25] A Novel Cryptocurrency Price Prediction Model Using GRU, LSTM and bi-LSTM Machine</p>
      <p>Learning Algorithms. URL: https://www.mdpi.com/2673-2688/2/4/30/pdf.
[26] Predictions of bitcoin prices through machine learning based frameworks. URL:
https://www.ncbi.nlm.nih.gov/pmc/articles/PMC8022579.
[27] Y. Xin et al., "Machine Learning and Deep Learning Methods for Cybersecurity," in IEEE Access,
vol. 6, pp. 35365-35381, 2018, doi: 10.1109/ACCESS.2018.2836950.
[28] Data Repository. URL: https://drive.google.com/drive/folders/17wcLX2VVw1cCo_6RsCEas2H</p>
      <p>SiDnZhfj4?usp=sharing.
[29] Data for Analysis. URL:
https://www.marketwatch.com/investing/index/spg1200/downloaddata?countrycode=xx.
[30] Five Types of Regression Analysis And When To Use Them. URL:
https://www.appier.com/blog/5-types-of-regression-analysis-and-when-to-use-them.
[31] Eight popular regression algorithms in machine learning of 2021. URL: https://www.jigsawacade
my.com/popular-regression-algorithms-ml.
[32] DTR. URL: https://scikit-learn.org/stable/auto_examples/tree/plot_tree_regression.html.
[33] The flowchart of random forest (RF) for regression. URL:
https://www.researchgate.net/figure/The-flowchart-of-random-forest-RF-for-regression-adaptedfrom-Rodriguez-Galiano-et_fig3_303835073.
[34] The Gradient Boosted Regression Trees (GBRT) model. URL: https://apple.github.io/turicreate/d
ocs/userguide/supervised-learning/boosted_trees_regression.html.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>U.</given-names>
            <surname>Mukhopadhyay</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Skjellum</surname>
          </string-name>
          ,
          <string-name>
            <given-names>O.</given-names>
            <surname>Hambolu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Oakley</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Yu</surname>
          </string-name>
          and
          <string-name>
            <given-names>R.</given-names>
            <surname>Brooks</surname>
          </string-name>
          ,
          <article-title>"A brief survey of Cryptocurrency systems,"</article-title>
          <source>2016 14th Annual Conference on Privacy, Security and Trust (PST)</source>
          ,
          <year>2016</year>
          , pp.
          <fpage>745</fpage>
          -
          <lpage>752</lpage>
          , doi: 10.1109/PST.
          <year>2016</year>
          .
          <volume>7906988</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>F.</given-names>
            <surname>Sabry</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W.</given-names>
            <surname>Labda</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Erbad</surname>
          </string-name>
          and
          <string-name>
            <given-names>Q.</given-names>
            <surname>Malluhi</surname>
          </string-name>
          ,
          <article-title>"</article-title>
          <source>Cryptocurrencies and Artificial Intelligence: Challenges and Opportunities," in IEEE Access</source>
          , vol.
          <volume>8</volume>
          , pp.
          <fpage>175840</fpage>
          -
          <lpage>175858</lpage>
          ,
          <year>2020</year>
          , doi: 10.1109/ACCESS.
          <year>2020</year>
          .
          <volume>3025211</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>J.</given-names>
            <surname>Bonneau</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Miller</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Clark</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Narayanan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. A.</given-names>
            <surname>Kroll</surname>
          </string-name>
          and
          <string-name>
            <given-names>E. W.</given-names>
            <surname>Felten</surname>
          </string-name>
          ,
          <article-title>"SoK: Research Perspectives and Challenges for Bitcoin and Cryptocurrencies,"</article-title>
          <source>2015 IEEE Symposium on Security and Privacy</source>
          ,
          <year>2015</year>
          , pp.
          <fpage>104</fpage>
          -
          <lpage>121</lpage>
          , doi: 10.1109/SP.
          <year>2015</year>
          .
          <volume>14</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>F.</given-names>
            <surname>Béres</surname>
          </string-name>
          ,
          <string-name>
            <given-names>I. A.</given-names>
            <surname>Seres</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A. A.</given-names>
            <surname>Benczúr and M.</surname>
          </string-name>
          Quintyne-Collins,
          <article-title>"Blockchain is Watching You: Profiling and Deanonymizing Ethereum Users,"</article-title>
          <source>2021 IEEE International Conference on Decentralized Applications and Infrastructures (DAPPS)</source>
          ,
          <year>2021</year>
          , pp.
          <fpage>69</fpage>
          -
          <lpage>78</lpage>
          , doi: 10.1109/DAPPS52256.
          <year>2021</year>
          .
          <volume>00013</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>Yu</given-names>
            <surname>Chen</surname>
          </string-name>
          , Xuecheng Ma,
          <source>Cong Tang and Man Ho Au</source>
          ,
          <article-title>"Pgc: Pretty good decentralized confidential payment system with auditability"</article-title>
          ,
          <source>Cryptology ePrint Archive Report</source>
          <year>2019</year>
          /319,
          <year>2019</year>
          , [online] Available. URL: https://eprint.iacr.org/
          <year>2019</year>
          /319.
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          <article-title>[6] Understanding The Different Types of Cryptocurrency</article-title>
          . URL: https://www.sofi.com/learn/conten t/understanding
          <article-title>-the-different-types-of-cryptocurrency.</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          <article-title>[7] The 10 Most Popular Cryptocurrencies, and What You Should Know About Each Before You Invest</article-title>
          . URL: https://time.com/nextadvisor/investing/cryptocurrency/types-of-cryptocurrency.
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>P.</given-names>
            <surname>Tasatanattakool</surname>
          </string-name>
          and
          <string-name>
            <given-names>C.</given-names>
            <surname>Techapanupreeda</surname>
          </string-name>
          ,
          <article-title>"Blockchain: Challenges and applications,"</article-title>
          <source>2018 International Conference on Information Networking (ICOIN)</source>
          ,
          <year>2018</year>
          , pp.
          <fpage>473</fpage>
          -
          <lpage>475</lpage>
          , doi: 10.1109/ICOIN.
          <year>2018</year>
          .
          <volume>8343163</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>A</given-names>
            <surname>Guide to Cryptocurrency Fundamental</surname>
          </string-name>
          <article-title>Analysis</article-title>
          . URL: https://academy.binance.com/en/articles/a
          <article-title>-guide-to-cryptocurrency-fundamental-analysis</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <article-title>The 7 Key Factors Influencing Cryptocurrency Value</article-title>
          . URL: https://www.makeuseof.
          <article-title>com/factorsinfluencing-the-cryptocurrency-value/</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>S.</given-names>
            <surname>Boshuis</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Braam</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A. Pedroza</given-names>
            <surname>Marchena</surname>
          </string-name>
          and
          <string-name>
            <given-names>S.</given-names>
            <surname>Jansen</surname>
          </string-name>
          ,
          <article-title>"</article-title>
          <source>The Effect of Generic Strategies on Software Ecosystem Health: The Case of Cryptocurrency Ecosystems," 2018 IEEE/ACM 1st International Workshop on Software Health (SoHeal)</source>
          ,
          <year>2018</year>
          , pp.
          <fpage>10</fpage>
          -
          <lpage>17</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <surname>Jiangtao</surname>
            <given-names>Ma</given-names>
          </string-name>
          , Yaqiong Qiao, Guangwu Hu, Yongzhong Huang, Arun Kumar Sangaiah,
          <string-name>
            <given-names>Chaoqin</given-names>
            <surname>Zhang</surname>
          </string-name>
          , et al.,
          <article-title>"De-anonymizing social networks with random forest classifier"</article-title>
          ,
          <source>IEEE Access</source>
          , vol.
          <volume>6</volume>
          , pp.
          <fpage>10139</fpage>
          -
          <lpage>10150</lpage>
          ,
          <year>2017</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <surname>Vynokurova</surname>
            <given-names>O.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Peleshko</surname>
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zhernova</surname>
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Perova</surname>
            <given-names>I.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kovalenko</surname>
            <given-names>A.</given-names>
          </string-name>
          (
          <year>2021</year>
          )
          <article-title>Solving Fraud Detection Tasks Based on Wavelet-Neuro Autoencoder</article-title>
          . In: Babichev S.,
          <string-name>
            <surname>Lytvynenko</surname>
            <given-names>V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Wójcik</surname>
            <given-names>W.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Vyshemyrskaya</surname>
            <given-names>S</given-names>
          </string-name>
          .
          <source>(eds) Lecture Notes in Computational Intelligence and Decision Making. ISDMCI 2020. Advances in Intelligent Systems and Computing</source>
          , vol
          <volume>1246</volume>
          . Springer, Cham. https://doi.org/10.1007/978-3-
          <fpage>030</fpage>
          -54215-3_
          <fpage>34</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <given-names>T.</given-names>
            <surname>Radivilova</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Kirichenko</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Ageiev</surname>
          </string-name>
          and
          <string-name>
            <given-names>V.</given-names>
            <surname>Bulakh</surname>
          </string-name>
          ,
          <article-title>"Classification Methods of Machine Learning to Detect DDoS Attacks,"</article-title>
          <source>2019 10th IEEE International Conference on Intelligent Data Acquisition and Advanced Computing Systems: Technology and Applications (IDAACS)</source>
          ,
          <year>2019</year>
          , pp.
          <fpage>207</fpage>
          -
          <lpage>210</lpage>
          , doi: 10.1109/IDAACS.
          <year>2019</year>
          .
          <volume>8924406</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <given-names>F. A.</given-names>
            <surname>Cahyadi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A. I.</given-names>
            <surname>Owen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Ricardo</surname>
          </string-name>
          and
          <string-name>
            <given-names>A. A. S.</given-names>
            <surname>Gunawan</surname>
          </string-name>
          ,
          <article-title>"Blockchain Technology behind Cryptocurrency and Bitcoin for Commercial Transactions,"</article-title>
          <source>2021 1st International Conference on Computer Science and Artificial Intelligence (ICCSAI)</source>
          ,
          <year>2021</year>
          , pp.
          <fpage>115</fpage>
          -
          <lpage>119</lpage>
          , doi: 10.1109/ICCSAI53272.
          <year>2021</year>
          .
          <volume>9609790</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [16]
          <string-name>
            <given-names>How</given-names>
            <surname>Bitcoin</surname>
          </string-name>
          <article-title>Works</article-title>
          . URL: https://www.investopedia.com/news/how-bitcoin-works.
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [17]
          <string-name>
            <given-names>S.</given-names>
            <surname>Pillai</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Biyani</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Motghare</surname>
          </string-name>
          and
          <string-name>
            <given-names>D.</given-names>
            <surname>Karia</surname>
          </string-name>
          ,
          <article-title>"Price Prediction and Notification System for cryptocurrency Share Market Trading,"</article-title>
          <source>2021 International Conference on Communication information and Computing Technology (ICCICT)</source>
          ,
          <year>2021</year>
          , pp.
          <fpage>1</fpage>
          -
          <lpage>7</lpage>
          , doi: 10.1109/ICCICT50803.
          <year>2021</year>
          .
          <volume>9510122</volume>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>