<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <journal-title-group>
        <journal-title>Information Control Systems &amp; Technologies, September</journal-title>
      </journal-title-group>
    </journal-meta>
    <article-meta>
      <title-group>
        <article-title>Computer Modelling and Investigation of Investment Portfolios</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Nataliia Kuznietsova</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Oleksii Shevchuk</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Kyiv</institution>
          ,
          <addr-line>03056</addr-line>
          ,
          <country country="UA">Ukraine</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>National Technical University of Ukraine “Igor Sikorsky Kyiv Polytechnic Institute”</institution>
          ,
          <addr-line>ave. Peremohy 37</addr-line>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>University Claude Bernard Lyon 1</institution>
          ,
          <addr-line>43 boulevard du 11 Novembre 1918, 69622 Villeurbanne cedex</addr-line>
          ,
          <country country="FR">France</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2023</year>
      </pub-date>
      <volume>2</volume>
      <fpage>1</fpage>
      <lpage>23</lpage>
      <abstract>
        <p>In this paper the main methods and models for investment portfolio forming are studied. The novelty of the paper is development of the method for forming an investment portfolio based on artificial intelligence. It differs from classical models that study the economic metrics not separately for each company but in aggregate effect. Thus, it gives to the model an opportunity to consider the market as a whole and determine which stocks are overvalued and undervalued. The method solves the problem of forming an investment portfolio in two stages. At the first stage, the method finds the true price of the company's shares based on the company's information, its quarterly reports and daily financial metrics that describe the company's economic situation dynamics and its shares. At the second stage the found price is compared with the current market price and based on this information a decision is made to add shares to the portfolio, and with what percentage or discard them. It was also developed an information technology implemented the packages in Python and realized the data preparation and investigation modules, modules for formation an investment portfolio using various types of classical methods and based on them models as well as developed method based on artificial intelligence. An experimental study was performed on a sufficient number of real datasets and where the developed approach showed high enough prediction accuracy scores. Investment analysis, computer modelling, investment portfolio, artificial intelligence-based models, Markowitz model, Bayesian networks, gradient boosting</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>Nowadays, humanity moreover understands the proper of good knowledge of financial management
importance. The main function of the management is to set financial assets in motion with the aim to
them. Moreover, any country significantly depends on investment processes in the modern world.</p>
      <p>
        The concept of financial investment dates to the time of ancient Babylon. The first written references
to investing and competent financial management date back to those times [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. Since that times,
scientists have been looking for ways to improve investment efficiency by increasing returns and
reducing the risk. One of the found solutions is the portfolio investment method.
      </p>
      <p>The essence of portfolio investing is to improve investment opportunities by providing a set of
investment objects with those investment qualities that cannot be achieved in the case of one single
investment object, but are quite possible in the case of these objects combination. Thus, a new challenge
for mathematicians and economists has emerged: finding the most optimal combination of investment
portfolio object’s.</p>
      <p>An investment portfolio is a purposefully formed set of investments in investment objects that
corresponds to a certain investment strategy of the investor by himself. It follows from this definition
that the main goal of forming an investment portfolio is to ensure the developed investment policy</p>
      <p>2023 Copyright for this paper by its authors.
implementation by selecting the most effective and reliable investments. Depending on the direction of
the chosen investment policy and the specifics of conducting investment activities, a system of specific
goals is determined, the main and most common of which are: the capital growth maximization, the
profit growth maximization, the investment risks minimization, ensuring the liquidity of the investment
portfolio that meets the requirements.</p>
      <p>The variety of types of investment objects, investment goals, their priority, as well as other
conditions led to the creation of a huge list of types of investment portfolios characterized by a certain
ratio of profitability and risk. The classification of investment portfolios by types of investment objects
is primarily related to the direction and volumes of investment activity.</p>
    </sec>
    <sec id="sec-2">
      <title>2. Problem statement</title>
      <p>In this article the analysis of existing methods for an investment portfolio formation and
investigation was carried out, their main advantages and disadvantages were identified. Some classical
methods and models for investment portfolio formation such as the decision-making method with a
fuzzy preference relation on a set of alternatives, Bayesian networks, Markowitz model and artificial
intelligence-based method with gradient boosting models were discussed. The theoretical foundations
of these methods were analyzed in accordance to their requirements, limitations and possibility for their
implementation. It was also conducted the simulations based on these methods and comparative
analysis on real data. The main aim in this article was to propose a new proprietary method for
investment portfolio formation based on artificial intelligence methods. This AI method should offset
main drawbacks of mentioned methods at all or at least partial to avoid limitations and give better results
on the same dataset. It was also decided to develop an information technology that forms an investment
portfolio using classical methods and a developed method based on artificial intelligence.</p>
    </sec>
    <sec id="sec-3">
      <title>3. Methods and materials</title>
      <p>Throughout the history of the investment portfolio forming task, a huge number of approaches and
methods have been developed. A special contribution was made by such well-known economists as
Sharpe and Markowitz. Also, a popular method for solving such problems is the Bayesian network
method, which is based on Bayes’ probabilistic theorem. With the advent of neural networks, many
scientists are trying to solve some specific problems and tasks. The formation of an investment portfolio
task was also the area of application and testing of different methods and approaches. Below will be
discussed the most appropriate methods and models for solving the problem of investment strategy
development with the application of classical and artificial intelligence methods and models.
3.1.</p>
    </sec>
    <sec id="sec-4">
      <title>Decision-making method</title>
      <p>
        A forming investment portfolio process includes a key decision-making stage. It is very important
that this stage is based on certain consequence of principles. Of course, the process of decision-making
can be influenced by personal preferences or gut feelings, but the final decision should still be based on
mathematical and statistical principles. This will help make the decision-making process somewhat
idempotent [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]. This theory helps to formally describe the fuzzy concepts that people use to describe
their desires, goals and perceptions of the system.
      </p>
      <p>So now it is possible to determine the main advantage of the decision-making theory such as its
simplicity and intuitiveness. Nevertheless, the factor of the existence of personal preferences rates
usually cause inaccuracy in results and make the most significant disadvantage.</p>
      <p>The decision-making process of choosing the most optimal alternative among a set of all alternatives
can take place when different amounts of information about these alternatives are available. A universal
way to describe input information is to write it in the form of a preference relation on a set of
alternatives.</p>
      <p>When modeling real systems, cases with clear relations of non-strict preference are rare. The case
of fuzzy relations is more typical. Such relations are described by the membership function   ( ,  ),
 
= 1⁡ − ⁡</p>
      <p>∈ [  ( ,  )],  ∈ 
  ( 0) ≥   ( ), ∀ = 1, . . ,  , ∀ ∈ 
  2( ,  ) = ∑     ( ,  )

 =1
following property must be satisfied:</p>
      <sec id="sec-4-1">
        <title>We must not forget that the main goal is to find the most profitable alternative  0. To do this, the</title>
        <p>If the condition is met, then the alternative  0 is called Pareto-optimal and to solve the problem, we
need to use another tool ‒ convolution of multiple criteria into a scalar one. One of the most common
ways of convolution is to use the intersection, which will give us the following:
  1
( ,  ) =</p>
        <p>{ 1 1( ,  ), … ,     ( ,  )}
where  1,…,</p>
        <p>are the weights of each criteria.
relations {  }, which is written as the sum:</p>
        <p>
          Another convolution important for finding optimal alternatives is the convolution of the original
which has the property of reflexivity. Given a fuzzy relation of non-strict preference R on X, we can
unambiguously define three corresponding fuzzy relations [
          <xref ref-type="bibr" rid="ref3">3</xref>
          ]:
(1)
(2)
(3)
(4)
(5)
(6)
(7)
(8)
(9)
where ∑ =1   = 1,   ⁡ ≥ 0.
sets  1
        </p>
        <p>and  2
3.2.</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>Markowitz model</title>
      <p>Since the optimal values belong to the space of non-dominated alternatives, they will belong to the
obtained as a result of convolution. To obtain the optimal alternatives, it is necessary
value will be maximal for optimal alternatives.
to find such elements from the joint set of non-dominated alternatives  1 and  2 . The non-dominance</p>
      <p>
        The investment portfolio formation theory according to H. Markowitz [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ] is based on the behavioral
specifics of an investor who wants to invest its resources in companies with the lowest risk and receive
certain dividends for this risk. Markowitz’s approach assumes that an investor takes into account only
two parameters: risk and return [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ].
      </p>
      <p>The Markowitz optimal portfolio model is based on the following principles:</p>
      <p>An investor wants to maximize the income for a given risk level.</p>
      <p>Investors always try to avoid risk. Between two assets with the same income, the one with the
1.
2.
3.</p>
      <p />
      <p>Fuzzy relation of indifference (equation 1)
Fuzzy equivalence relation (equation 2)</p>
      <p>Fuzzy strict preference relation (equation 3)
 ( ,  ) = max⁡[min{1 −   ( ,  ); 1 −   ( ,  )} ; min{  ( ,  );   ( ,  )}]</p>
      <p>( ,  ) = min{  ( ,  );   ( ,  )}
  ( ,  ) = {  ( ,  ) − ⁡   ( ,  ),  ⁡  ( ,  ) &gt;   ( ,  )</p>
      <p>0,</p>
      <p>According to the definition of the intersection operation of fuzzy sets, the expression for the
membership function of the set of non-dominated alternatives will be as follows:
lower risk is chosen.</p>
      <p>Risk is the uncertainty of a future outcome.</p>
      <p>An investor’s portfolio consists of all his assets and liabilities.</p>
      <p>Investors make investment decisions based on expected returns and investment risk. The
“usefulness of investments” is calculated by the formula:</p>
      <p>⁡ = ⁡   ⁡ − ⁡ (  )/2
where  ‒ the degree of risk aversion by the investor,   ‒ expected income,   ‒ expected risk.</p>
      <p>According to the Markowitz model, the portfolio income is the weighted average income on its
components which is determined by the formula:

where  is the number of stocks,   ‒ income of a particular stock,   ‒ the stock percentage.</p>
      <p>The following equation is used to determine the riskiness of a portfolio:</p>
      <p>where   is the stock percentage,  
and  (standard deviation).</p>
      <p>For Markowitz’s model, the efficient set theorem is important, which states that an investor will
choose his optimal portfolio among a set of portfolios, each of which will provide:
is a linear correlation coefficient,   ,   are the risks of stocks 
•
•
maximum profitability for a given risk level;
minimum risk for a given value of expected income.</p>
      <p>The set of portfolios satisfying the theorem is called the efficient set or efficient frontier. Within this
set or on the frontier are all portfolios that can be formed from a certain number of stocks. The effective
set is the area in which the points are located, and the effective frontier is the line that graphically
delineates this set (Figure 1).</p>
      <p>⁡ = ⁡
  ⁡ − ⁡</p>
      <p>set. On figure 1 such portfolio located at the point  ∗.</p>
      <p>
        The choice of investment portfolio objects depends not only on the objects themselves, but also on
the investor’s strategy and goals. Thus, the most obvious strategies are those of minimum risk and
maximum return. Another way to form an investment portfolio is to maximize the Sharpe ratio [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ].
      </p>
      <p>This ratio helps to solve the problem of choosing between two investment objects by comparing
their income with a risk-free object. The Sharpe ratio is calculated using the following equation:
where   is a portfolio income,   is a risk-free object income,   are the portfolio riskiness.</p>
      <p>Thus, the Sharpe ratio will show the relative income on an asset per unit of risk. Maximizing this
value will form a portfolio from the assets with the most “profitable” risk.</p>
      <p>The main Markowitz model’s advantage is that it is based on the economy theory and it was really
efficient for solving economy tasks. But such model requires a big dataset of historical economy data
which is not always possible to retrieve. Also for some tasks as the historical data could be not always
in the same way as the relevant current financial processes. So the Markowitz model requires additional
learning of actual data, otherwise it will lose accuracy.
(10)
(11)</p>
    </sec>
    <sec id="sec-6">
      <title>Bayesian network</title>
      <p>
        Bayesian networks are widely used for processing statistical data represented by time series. These
networks play a special role in risk management. Bayesian networks help to establish cause-and-effect
relationships between certain attributes and the conclusion obtained under such conditions [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ].
      </p>
      <p>
        A Bayesian network can be viewed as a model for representing the relationships between the vertices
of an acyclic graph, which are represented as causal probability dependencies. Formally, a Bayesian
network is a triple  ⁡ =⁡&lt; ⁡ ,  ,  &gt;, where  is a set of variables,  is a directed acyclic graph, and 
is the joint probability distribution of the variables  ⁡ = ⁡ {⁡ 1, . . . ,   }. It should be mentioned, that the
Markov condition is satisfied for the set of variables  , which means that each variable in the network
is independent of all other variables except for the parental predecessors of this variable [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ].
      </p>
      <p>Bayesian networks are based on probability theory and, in particular, Bayes’ theorem. In the general
case where an event depends on several other events, Bayes’ theorem is as follows:
(12)
(13)
(14)
 ( | 1, . . . ,   ) ⁡ = ⁡
 ( ) ( 1, . . . ,   | )</p>
      <p>( 1, . . . ,   )
by analyzing historical data).
where  ( | 1, . . . ,   ) is the probability of an event  in the presence of a disturbances  1, . . . ,   , ⁡ ( )
is the probability of the target event  (this value can be estimated using historical data),  ( | ) is the
probability of the reverse course of the event, the appearance of  1, . . . ,   in the presence of  ,
 ( 1, . . . ,   ) is the probability of occurrence of the events  1, . . . ,   (this value can also be estimated</p>
      <p>But this equation needs to be transformed to be used in a Bayesian network. Assuming that all events
 1, . . . ,   are independent and  is known, the following equation will hold:</p>
      <sec id="sec-6-1">
        <title>With further normalization of equation 12, we can get rid of the denominator  ( 1, . . . ,   ), which</title>
        <p>will simplify the task of forming a conclusion. Thus, we obtain a generalized equation for drawing a
conclusion by Bayes’ theorem:</p>
        <p>( 1, . . . ,   | ) ⁡ = ⁡ ( 1| ). . .  (  | )
 ( | 1, . . . ,   ) ⁡ = ⁡ ( ) ( 1| ). . .  (  | )</p>
        <p>The Bayessian model is really good for classification and it can give a probability of each event, but
it requires also a big dataset, and the model based on the Bayesian network grows with incresing number
of features. Another drowback is that Bayesian model is not able to predict an accurate value but just
classify the sample to a group or gives the probability of the appearance of some fact. So for financial
tasks you need carefully to make a problem statement as well the main “probability” evidence which
you are going to forecast.
3.4.</p>
      </sec>
    </sec>
    <sec id="sec-7">
      <title>Artificial intelligence-based models</title>
      <p>
        There are numerous of different artificial intelligence models, ranging from various variations of
gradient boosting to full-fledged neural networks. Often, a single model is not enough to solve a certain
task. To solve this problem, an ensemble of models is used, i.e., several machine learning algorithms
assembled into a single one. This approach is often used to enhance the positive qualities of individual
algorithms, which may be weak on their own, but show excellent results in a group. When using
ensemble methods, algorithms are trained simultaneously and can correct each other’s mistakes. Model
ensembles are usually built on the basis of a decision tree model. Trees are added one at a time to the
ensemble and trained to mutually correct prediction errors made by previous models. This type of
ensemble is called boosting [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ].
      </p>
      <p>
        Gradient boosting is one of the classic models of artificial intelligence (AI). It also belongs to the
ensemble models of artificial intelligence, i.e., this AI model will consist of several models [
        <xref ref-type="bibr" rid="ref10 ref9">9, 10</xref>
        ].
      </p>
      <p>Consider the task of recognizing objects in a multidimensional space  with a label space Y. Suppose


we are given a training set {  } =1 , where   ⁡ ∈ ⁡ , and the true labels {  } =1,   ⁡ ∈ ⁡ , of each object
from the set. The learning task is to use the training set to find the approximating function ̂( ) to the
function  ( ) which minimizes the expected value of some given loss function according to the
formula:
̂( ) =</p>
      <p>Ε , [ ( ,  ( ))]

 ̂( ) = ∑</p>
      <p>=1   ℎ ( ) + 
where  ( ,  ( )) is a loss function.</p>
      <p>The gradient boosting method searches for an approximating function ̂( ) as a weighted sum of
functions ℎ( ) and some class  , which are considered as weak models. With this in mind, the equation
(15) can be represented:
where  is a number of weak models,   is a weighted coefficient and ℎ ( ) is a weak model function.</p>
      <p>Thus, the output of the learning algorithm is a set of 
decision trees, and to make a prediction,
i.e., to determine the output  for a new object  , we should calculate the sum with the equation:

 =  0 +  ∗ ∑=1   ( )</p>
      <p>where  0 is the first decision tree,  is the scaling factor,  is the total number of constructed decision
trees,   the m-th decision tree.</p>
      <p>
        The final classifier is represented as a linear combination of classifiers. Finding the optimal values
of the coefficients of this linear combination is a rather time-consuming task, so gradient boosting uses
a greedy algorithm for gradually adding classifiers [
        <xref ref-type="bibr" rid="ref10 ref11">10, 11</xref>
        ].
      </p>
      <p>The most significant advantages of AI-based models are their flexibility and possibility to be
specialized for the certain task. It makes results more accurate then with other models. Another
advantages goes from using ensembles of the models. It gives opportunity to resolve tasks with a lack
of data in the dataset. So, the approach based on using of AI models seems to us more perspective,
modern and relevant for using to practical financial tasks. In this paper we will focus on development
of the AI models for forming an investment portfolio. Here we propose the new method based on AI
which differs from classical that study the economic metrics not separately for each company but their
aggregate effect. Consequently, it gives to the model an opportunity to consider the whole financial
market and to classify which companies’ stocks are overvalued and undervalued and to choose the most
attractive and valuable companies for investment.
(15)
(16)
(17)</p>
    </sec>
    <sec id="sec-8">
      <title>4. Modeling and simulation of the practical task</title>
      <p>effectiveness in different ways, such as:</p>
      <p>It was decided to build the following models for testing the forming investment portfolios
•
•
•
•
decision-making model with fuzzy preference relation on a set of alternatives;
Bayesian network model;
Markowitz model;
artificial intelligence model.
4.1.</p>
    </sec>
    <sec id="sec-9">
      <title>Data sample description</title>
      <p>The key data for forming an investment portfolio task is the general information about the company,
as well as its financial position and the dynamics of its development. Thus, three types of data were
used: general information about companies, quarterly companies’ reports, daily information about
companies and their stocks. Basic data refers to all information about a company that generally does
not change over time. This includes information about the industry in which the company specializes,
the sector in which the company operates, etc. (Table 1).</p>
      <sec id="sec-9-1">
        <title>Sector</title>
      </sec>
      <sec id="sec-9-2">
        <title>Consumer Cyclical</title>
      </sec>
      <sec id="sec-9-3">
        <title>Technology</title>
      </sec>
      <sec id="sec-9-4">
        <title>Technology</title>
      </sec>
      <sec id="sec-9-5">
        <title>Consumer Cyclical</title>
      </sec>
      <sec id="sec-9-6">
        <title>Consumer Cyclical</title>
      </sec>
      <sec id="sec-9-7">
        <title>Industry</title>
      </sec>
      <sec id="sec-9-8">
        <title>Entertainment</title>
      </sec>
      <sec id="sec-9-9">
        <title>Semiconductors</title>
      </sec>
      <sec id="sec-9-10">
        <title>Consumer Electronics</title>
      </sec>
      <sec id="sec-9-11">
        <title>Auto Manufactures</title>
      </sec>
      <sec id="sec-9-12">
        <title>Auto Manufactures</title>
      </sec>
      <sec id="sec-9-13">
        <title>Currency USD USD JPY</title>
        <p>JPY
USD</p>
        <p>Quarterly reports include a variety of information on the company’s development dynamics, as well
as typical metrics used to assess the company’s financial performance. Such data include the amount of
debt, the amount of the company’s profit, information on annual stockholder dividends, profit before
tax, etc. (Table 2).
capitalization and the closing price of the companies’ stocks (Table 3).</p>
      </sec>
    </sec>
    <sec id="sec-10">
      <title>Development of the Artificial Intelligence-based model</title>
      <p>The decision-making model, that was built to select the optimal investment objects for an investment
portfolio, received input features that describe the company’s financial position and growth dynamics.
The most important features were the industry in which the company operates, changes in the stock
price, company capitalization and the debt amount. The data was taken for the last quarter and the
difference data was calculated as the difference between the last and the penultimate quarter.</p>
      <p>The model output was a metric of each alternative importance, which ranges from 0 to 1. The higher
metric value corresponds to the higher priority of choosing this alternative. The number of companies
to be added to the portfolio was set by a user. The stocks’ percentage in each company was calculated
by the formula:

the portfolio and in experiments this number was chosen as 7.
where</p>
      <p>is a metric returned by the model for the  -th company,  is a number of companies in
The Bayesian network model receives various data about the companies from which the portfolio is
planned to be formed, as well as their financial indicators’ dynamics. As an output, the model returns
the probability that the company’s stock will grow in value. Only those companies will be selected for
the portfolio whose growth probability exceeds a certain cut-off point set by the investor. The
percentage of each investment object is determined by the following formula:

companies whose stock price growth probability exceeded the cut-off point.
where</p>
      <p>is the probability of increase in the stock price of the  -th company,  is a number of
The Markowitz method was implemented by searching through one hundred million potential
portfolios. The portfolios from the portfolio effective set that were optimal according to the risk
minimization strategy and the Sharpe ratio maximization strategy were selected.</p>
      <p>Specific data pre-processing for the artificial intelligence-based model was carried out. The model
received all types of data: basic, daily, and quarterly. Basic text data is converted into numeric data. For
daily and quarterly data, statistics were obtained for each metric for different specified intervals, such
as standard deviation, an average value for the period, and others.</p>
      <p>It was developed the method based on artificial intelligence for building investor’s portfolio which
works in two stages. In the first stage, it uses an ensemble of gradient boosting models to determine the
“fair/true” share price. In the second stage it uses the equation (20) to determine whether the company
is undervalued or overvalued.
where "Fair"⁡
is the price determined by the gradient boosting
method and
is the price from real dataset. If the 
&lt; 1 it means that the company is
undervalued and in next period its shares will increase, so it is recommended to add this company in
investor’s portfolio. If the</p>
      <p>&gt; 1 it means that the company is overvalued and in next period the
shares will fall and it is not recommended to investors to focus on this company and to include it in
portfolio.</p>
      <p>Companies for which this ratio exceeds a certain cut-off point determined by analysts beforehand
will also be rejected as erroneous. The share of each investee is calculated as the ratio of the investee’s
coefficient to the sum of all selected investees coefficients.
(19)
(20)</p>
    </sec>
    <sec id="sec-11">
      <title>5. Results 5.1.</title>
    </sec>
    <sec id="sec-12">
      <title>Decision-making model</title>
      <p>The first built model was a decision-making model with a fuzzy preference relation on a set of
alternatives. An investment portfolio with seven investment objects was formed (Table 4).
desirable. On the positive side, we can point out the choice of WWE as a target. The stocks of this
company grew the most out of the entire set, but it is not enough to compensate the losses of this
portfolio.</p>
      <p>Thus, the decision-making model did not produce a satisfactory result. It selected only those
investment targets that the investor liked, with little regard for financial feasibility. The assignment of
weights to each criterion has a negative impact on the model’s efficiency, as they are based on certain
subjective judgments rather than analytical decisions.
5.2.</p>
    </sec>
    <sec id="sec-13">
      <title>Bayesian network model</title>
      <p>The Bayesian network model receives various companies input data from which the portfolio is
planned to be formed, as well as their financial indicators dynamics. At the output model returns the
probability that the company’s shares will increase in price. Only those companies whose growth
probability will exceed a certain cut-off limit set by the investor will be selected for the portfolio. Daily
data were converted to quarterly and combined with preliminary quarterly data for Bayesian network
model. Basic information about each company’s sector and industry was also added to them. The first
difference for numerical data was taken in order to form the task of determining the share price’s growth
or fall based on the input data. The value 0.55 was chosen as the cut-off, i.e. all companies whose stock
growth probability exceeds 55%. Received results for Bayesian modeling are illustrated in the Table 5.
The total portfolio income is -740.65. The resulting portfolio created by the Bayesian network model
was also found to be unprofitable. Compared to the previous portfolio, created by decision-making
model, it has some advantages, in addition to a smaller loss. This includes numerous of profitable
stocks. Only three of the eight investment objects fell in value, and unfortunately, the fall of one of
them was significant. All the other objects were characterized by price growth. In addition, if the
cutoff value was increased, the portfolio risk would decrease, and, in this particular case, there would be
only one SONY investment target that had an increase in price.
5.3.</p>
    </sec>
    <sec id="sec-14">
      <title>Markowitz model</title>
      <p>The Markowitz model was implemented by sifting through one hundred million portfolios. As a
result, the cloud of portfolios in the Figure 2 was obtained. Two investment strategies for the Markowitz
model were considered: the minimum risk strategy and the maximum Sharpe ratio strategy. On the
Figure 2, red point represents a portfolio for minimum risk strategy, and green – for maximum Sharp
ratio strategy. The results of each strategy are presented in Tables 6 and 7.</p>
      <p>Table 6 shows the most significant elements of the Markowitz model of minimal risk portfolio. The
total minimum risk Markowitz model portfolio income is -105.67. The portfolio showed much better
result than the previous ones, but it is still unprofitable. The benefits of the portfolio are that the most
unprofitable investment objects such as ADBE, TSLA, NVDA are taken with a very low percentage.
In addition, the most profitable stocks, such as WWE and ORCL, account for almost a third part of the
portfolio. The biggest portfolio loss was caused by buying TM shares with a significant percentage.
Usually, these stocks are relatively stable, but this time they fell sharply during the quarter.</p>
      <p>It is also significant to mention that the market tended to fall during 2022 due to geopolitical factors.
If in the sample were investment objects that had risen approximately as much as TSLA or ADBE fell,
the portfolio could have made zero losses or even a profit.</p>
      <p>The total maximum Sharp ratio Markowitz model portfolio income is 96.75. The portfolio was
profitable, but this profitability was achieved through the adventurous purchase of WWE shares, which
make up almost half of the total portfolio. In addition, the shares of TSLA, ADBE, and NVDA
increased, which reduced the portfolio potential returns.
5.4.</p>
    </sec>
    <sec id="sec-15">
      <title>Artificial intelligence-based model</title>
      <p>The portfolio formed by the artificial intelligence-based model showed the best results (Table 8).
The total income is 195.93. In addition, all the most unprofitable stocks, such as TSLA and ADBE,
were rejected, while the most profitable ones, such as ORCL or WWE, were added to the portfolio. It
is also worth noting that the portfolio is quite balanced. The share of each investment objects is fairly
equal, which makes the portfolio relatively stable, unlike the portfolio of the Markowitz maximum
Sharpe ratio model. The main drawback of this model was training time. It took much more time and
processing resources than any previous model.</p>
    </sec>
    <sec id="sec-16">
      <title>6. Discussion</title>
      <p>After obtaining the simulation results, it is possible to discuss the effectiveness of the classical
approaches as well as the developed method for investment portfolio formation. It is obvious that the
results of decision-making methods with a fuzzy preference ratio on a set of alternatives and Bayesian
networks showed too bad results, so it was decided not to add them to the final comparative Table 9.</p>
      <p>The comparison table clearly shows that the developed artificial intelligence-based method for
investment portfolio formation rejected the most unprofitable objects that were included in the
portfolios created by the Markowitz models of minimum risk and maximum Sharpe ratio. The following
advantage of the developed method is also clearly visible ‒ the objects of its portfolio are maximally
equivalent, which is the most weighted. Even in a minimal risks Markowitz portfolio there is an asset
like TM, that clearly outperforms all other assets in the portfolio in percentage terms. The main
disadvantage of the developed method is its demand for computational capabilities and input data
volumes. The method uses a huge amount of data about different companies in order to determine the
true shares price. Accordingly, it is needed to spend a lot of time on this. The same Markowitz method
will run much faster because it doesn’t need to process as much data. On the other hand, a trained model
can form a portfolio from a new set of companies that could be included into the model, whereas a
Markowitz model needs to start forming a portfolio from beginning with new data. In this case, the
developed artificial intelligence-based method will work faster.
intelligence-based decision-making method. This problem could be solved by optimizing the program
as well as by uploading the IT to cloud services. It is also advisable to use parallel calculations.</p>
    </sec>
    <sec id="sec-17">
      <title>7. Conclusions</title>
      <p>The main existed methods for forming an investment portfolio, from the simplest classical to the
most modern methods of AI were considered. Based on the most relevant to this task methods, the
models for investment portfolio formation were built and compared. Obtained result showed that the
portfolios formed by the decision-making method and the Bayesian network method showed
unsatisfactory results. For the decision-making method, the explanation for such unsatisfactory results
is that it does not analyze the financial market, but simply helps in decision-making. The results on
investment modelling by Markowitz model were much better. This explains its popularity among
analysts and investors. In addition, Markowitz model gives the option to calculate the portfolio risk.</p>
      <p>As a result, a method of forming an investment portfolio based on artificial intelligence was
developed and compared with classical methods. Then it was tested on the real data and received results
of comparison showed that the developed method has a higher efficiency than classical methods. The
resulting investment portfolio produced the highest return among the other portfolios and was also
balanced, which shows its reliability. Results could also be improved by using even more numbers of
different metrics and increasing the size of the dataset. The advantage of the artificial intelligence-based
method is that it analyzes the market as a whole and can compare companies with each other. In
comparison the Markowitz model only analyzes the time series of share prices.</p>
      <p>
        As it is shown by experiments, the proposed artificial intelligence-based method should be also
improved in future research. The first way to improve modeling results is to expand the dataset. This
can be done by extending the period for which the data was collected, for example, for 10 years. It
should be also increased the number of metrics themselves, add other financial indicators [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ] that
describe the company’s position, in addition to those mentioned earlier. In addition it could be also
proposed to use other methods for investment portfolio formation. For example, neural networks, in
particular those that work with fuzzy data. The approach of detecting shares and companies for investing
could also be adapted by the investor’s behavior and its attitude to risks.
8. References
      </p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>F.</given-names>
            <surname>Ponziani</surname>
          </string-name>
          ,
          <article-title>The evolution of investing from Babylon to</article-title>
          MELD,
          <year>2022</year>
          . URL: https://www.meld.com/blog/the-evolution
          <article-title>-of-investing-from-babylon-to-meld.</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>M.</given-names>
            <surname>Voskoglou</surname>
          </string-name>
          . Fuzzy sets,
          <source>Fuzzy Logic and Their Applications</source>
          ,
          <year>2021</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>Yu.P.</given-names>
            <surname>Zaichenko</surname>
          </string-name>
          , Operations research, Slovo Publishing House,
          <year>2006</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>H.</given-names>
            <surname>Markowitz</surname>
          </string-name>
          , Portfolio selection:
          <source>Efficient Diversification of Investments</source>
          ,
          <year>1991</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>P.</given-names>
            <surname>Bidyuk</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Polozhaenko</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Kuznietsova</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Levenchuk</surname>
          </string-name>
          ,
          <article-title>Probabilistic Data Analysis in NonStationary Processes Forecasting</article-title>
          ,
          <source>in: Proc. 2020 IEEE 2nd International Conference on System Analysis &amp; Intelligent Computing (SAIC)</source>
          , Kyiv, Ukraine,
          <year>2020</year>
          , pp.
          <fpage>297</fpage>
          -
          <lpage>302</lpage>
          , doi: 10.1109/SAIC51296.
          <year>2020</year>
          .
          <volume>9239200</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>S. E.</given-names>
            <surname>Paw</surname>
          </string-name>
          . The Sharpe Ratio: Statistics and Application, ed. 1,
          <year>2021</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>J.</given-names>
            <surname>Pearl</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Russell</surname>
          </string-name>
          , Bayesian Networks, in: M.A.
          <string-name>
            <surname>Arbib</surname>
          </string-name>
          (Ed.),
          <source>Handbook of Brain Theory and Neural Networks</source>
          , Cambridge, MA: MIT Press,
          <year>2003</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <surname>Freund</surname>
            <given-names>Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Schapire</surname>
            <given-names>R.</given-names>
          </string-name>
          ,
          <source>Boosting: Foundations and Algorithms (Adaptive Computation and Machine Learning Series)</source>
          ,
          <year>2014</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>C.</given-names>
            <surname>Zhang</surname>
          </string-name>
          , Y.
          <source>Ma Ensemble Machine Learning: Methods and Applications</source>
          ,
          <year>2012</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <surname>Mason</surname>
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Baxter</surname>
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bartlett</surname>
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Frean</surname>
            <given-names>M.</given-names>
          </string-name>
          , Boosting Algorithms as Gradient Descent,
          <source>Neural Information Processing Systems</source>
          ,
          <volume>12</volume>
          (
          <year>2000</year>
          )
          <fpage>512</fpage>
          -
          <lpage>518</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <surname>Galkina</surname>
            <given-names>S.</given-names>
          </string-name>
          ,
          <source>Optimal Portfolio Selection Using Machine Learning Techniques</source>
          ,
          <source>International Journal of Open Information Technologies</source>
          ,
          <volume>2 6</volume>
          (
          <issue>2014</issue>
          )
          <fpage>14</fpage>
          -
          <lpage>19</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>P.I.</given-names>
            <surname>Bidyuk</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N. V.</given-names>
            <surname>Kuznietsova</surname>
          </string-name>
          ,
          <article-title>Forecasting the volatility of financial processes with conditional variance models</article-title>
          ,
          <source>Journal of Automation and Information Sciences, 46</source>
          <volume>10</volume>
          (
          <year>2014</year>
          )
          <fpage>11</fpage>
          -
          <lpage>19</lpage>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>