<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <journal-title-group>
        <journal-title>A. Sachenko);</journal-title>
      </journal-title-group>
    </journal-meta>
    <article-meta>
      <title-group>
        <article-title>Evaluation of ensemble machine learning models for movie recommendation systems</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Anatoliy Sachenko</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Taras Lendiuk</string-name>
          <email>as@wunu.edu.ua</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Khrystyna Lipianina-Honcharenko</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Vasyl Koval</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Grygoriy Hladiy</string-name>
          <email>ghladiy@wunu.edu.ua</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Yurii Halias</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Kazimierz Pulaski University of Technology and Humanities in Radom</institution>
          ,
          <addr-line>Radom, 26 600</addr-line>
          ,
          <country country="PL">Poland</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>West Ukrainian National University</institution>
          ,
          <addr-line>Lvivska str., 11, Ternopil, 46000</addr-line>
          ,
          <country country="UA">Ukraine</country>
        </aff>
      </contrib-group>
      <volume>000</volume>
      <fpage>0</fpage>
      <lpage>0001</lpage>
      <abstract>
        <p>This article is dedicated to evaluating the effectiveness of ensemble machine learning models in the context of movie recommendation systems. It explores various ensemble methods, including Random Forest, AdaBoost, XGBoost, LightGBM, CatBoost, and Gradient Boosting Machine, to enhance the accuracy of predicting user preferences. The study is based on the MovieLens 100K dataset, which contains 100,000 ratings from 943 users across 1,682 movies. The application of feature engineering, data normalization methods, and iterative feature selection has improved the model's ability to accurately predict user interests. The analysis showed that the XGBoost model exhibits the best results with the lowest RMSE value of 0.902, indicating higher prediction accuracy compared to other models considered. LightGBM and CatBoost also showed competitive results with RMSE values of 0.910 and 0.919, respectively. The study highlights the importance of an integrated approach to developing recommendation systems that adapt to the diverse preferences and contexts of users, opening wide perspectives for further research in this area.</p>
      </abstract>
      <kwd-group>
        <kwd>ensemble models</kwd>
        <kwd>machine learning</kwd>
        <kwd>movie recommendation systems 1</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>In today's world, where the amount of digital content is growing every day, recommender
systems play a key role in helping users find information, products or services that best suit
their interests and needs. Among the various applications of recommender systems, film
recommendation systems are particularly important, helping users navigate the vast world of
cinema by suggesting films based on their preferences. The development of machine learning
technologies and ensemble methods opens up new opportunities to improve the accuracy and
adaptability of such systems.</p>
      <p>In recent years, significant advances in machine learning and ensemble methods have
greatly expanded the capabilities of recommender systems. Algorithms such as Random Forest,</p>
      <p>AdaBoost, XGBoost, LightGBM, CatBoost, and Gradient Boosting Machine have demonstrated
high performance in classification and regression tasks, making them ideal tools for developing
advanced recommender systems. These methods not only improve the accuracy of predicting
users' interests, but also ensure high adaptability of the system to changing preferences and
contexts.</p>
      <p>Despite the significant progress in this area, there are certain challenges, in particular,
related to the processing of large amounts of data, effective consideration of socio-demographic
information and browsing context, as well as optimization of the choice of hyperparameters.
This study aims to address these challenges by proposing a comprehensive approach to building
a film recommendation system that integrates advanced machine learning techniques.</p>
      <p>The main objective of this research is to evaluate the effectiveness of ensemble machine
learning models in improving the accuracy of film recommendations. We aim to investigate
how different ensemble methods can affect the model's ability to accurately predict user
preferences using both traditional and innovative approaches to data processing and analysis.
Through a comparative analysis of different models, we plan to identify the most effective
strategies to implement in film recommendation systems, paving the way for further research
and development in this exciting area.</p>
      <p>This paper focuses on the evaluation of ensemble machine learning models for film
recommendation systems and is structured as follows. Section 2 describes the analysis of related
work, highlighting key algorithms and their applications in the field of recommender systems.
Section 3 presents an integrated approach to building an intelligent film recommendation
system, including a description of the research methodology and data analysis. Section 4
implements the proposed method, demonstrating the process of data preparation, model
training, and evaluation of their effectiveness. Section 5 presents the results of the study,
analyzing the performance of different ensemble models and their ability to accurately predict
user preferences.</p>
    </sec>
    <sec id="sec-2">
      <title>2. Related Work</title>
      <p>
        In the field of machine learning and ensemble methods, a number of studies highlight key
algorithms and their application to improve forecasting accuracy. The Random Forest algorithm
discussed in [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ] is the foundation of ensemble learning, while AdaBoost [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ] and XGBoost [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]
optimize the boosting process to improve weak classifiers and demonstrate high performance
in prediction. LightGBM [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ] uses innovative decision tree algorithms to efficiently process big
data, while CatBoost [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ] provides accuracy for categorical data without complex
hyperparameter selection. Gradient Boosting Machine [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ] improves models through stochastic
gradient descent. The foundations of statistical learning and deep learning are presented in [
        <xref ref-type="bibr" rid="ref7">7,
8, 13-15, 19, 20</xref>
        ], respectively, providing a theoretical basis and practical directions for
development in the field of data mining.
      </p>
      <p>In the area of film recommendation systems, research has used a variety of machine learning
approaches to improve the accuracy and relevance of suggestions to users. For example, a study
focusing on multimodal trusted recommendations uses machine learning algorithms such as
backpropagation, SVD, and deep learning to identify trusted users whose recommendations are
then offered to active users [9]. Another study introduces context-aware approaches, using
signal processing and machine learning to recommend films that take into account the user's
specific context [10]. Approaches such as deep learning are used to predict user ratings based
on the MovieLens dataset, demonstrating the use of collaborative filtering based on deep
learning strategies [11].</p>
      <p>This study is distinguished by the use of a comprehensive approach that integrates advanced
machine learning techniques to create a more accurate and adaptive film recommendation
system. By using feature engineering, data normalisation, and iterative feature selection, our
method improves the model's ability to accurately predict users' interests, taking into account
not only their prior ratings, but also socio-demographic information and viewing context. A
special feature of our approach is the use of ensemble methods, such as XGBoost, to optimise
the prediction accuracy, which demonstrates a significant reduction in RMSE error compared
to other models mentioned in [9-11]. This makes our study particularly valuable for the
development of effective recommender systems that can adapt to a wide range of user
preferences and contexts.</p>
      <p>Thus, the main goal of this study is to determine the optimal strategy for improving the
accuracy of intelligent film selection using ensemble machine learning methods. In particular,
a comparative analysis of various ensemble approaches such as Random Forest, AdaBoost,
XGBoost, Stacking Ensemble, and SVR (Support Vector Regression) is planned to evaluate their
ability to minimise the prediction error measured by the RMSE (Root Mean Square Error)
metric.</p>
    </sec>
    <sec id="sec-3">
      <title>3. An integrated approach to creating an intelligent film selection system</title>
      <sec id="sec-3-1">
        <title>3.1. Method description</title>
        <p>Creating and evaluating machine learning models for intelligent film selection can be
represented as a sequential step-by-step process that includes the following steps:</p>
        <p>Step 1. Data collection. Collection of datasets with user reviews, film metadata (genres,
directors, actors, ratings) and socio-demographic information of users.</p>
        <p>Step 2. Initial data analysis.</p>
        <p>Step 2.1. Descriptive statistics to calculate means, medians, and standard deviations. The
mean (μ) is a fundamental indicator in statistics, which is the arithmetic mean of a set of values
calculated by dividing the sum of all values by their number, which allows you to get the overall
central tendency of the data. The median, on the other hand, is defined as the value that divides
an ordered set of data into two equal parts, serving as a reliable indicator of central tendency,
especially in the presence of outliers. The standard deviation (σ) describes the dispersion or
variability of the data relative to the mean, indicating how far apart the values in a set are from
their average.</p>
        <p>Step 2.2. Visualize the distributions of scores and ratings. Histograms and Boxplots are used
as key visualization tools to clearly represent the distributions of scores and ratings in a dataset.
Histograms make it easy to identify underlying trends in the distribution by showing the
frequency of different values, which helps determine how the data is distributed along the rating
or rating scale.</p>
        <p>Step 2.3. Detect anomalies and outliers to identify errors or special cases. Outliers can be
identified using, for example, the Z-score or interquartile range (IQR).</p>
        <p>The Z-score determines how far a value is from the mean, expressed in standard deviations.
Values with ∣Z∣&gt;3 are usually considered outliers.</p>
        <p>=
(
−  )</p>
        <p>Interquartile range (IQR). The difference between the third (Q3) and first (Q1) quartiles.
Outliers are defined as values that fall outside of 1.5×IQR from Q1 and Q3.</p>
        <p>Step 2.4. Processing of missing values through imputation or deletion. Deletion involves
simply eliminating the rows or columns containing the missing data, which can be effective but
potentially leads to the loss of valuable information. Imputation fills in gaps in the data using a
variety of methods, such as replacing the missing values with the mean, median, or mode for
numeric data and most commonly, values or applying algorithms such as K-nearest neighbours
for categorical data.</p>
        <p>Step 2.5. Correlation analysis to identify relationships between features. Correlation analysis
allows you to determine the strength and direction of the relationship between pairs of
variables. The most commonly used correlation is Pearson's correlation for continuous
variables:
 =
∑( −  )( −  )
∑( −  ) ∑( −  )
where  and  are the values of variables X and Y, respectively, and 
and 
are their
average values.</p>
        <p>Step 3. Feature engineering.</p>
        <p>Step 3.1. Coding of categorical variables using the methods of one-hot encoding, label
encoding, or binary encoding.</p>
        <p>One-hot encoding. Each categorical variable is divided into as many binary variables
(columns) as there are unique categories. If a variable belongs to a certain category, the
corresponding column will have a value of 1, and the other columns will have a value of 0.</p>
        <p>Label Encoding. Each unique category is assigned a unique integer. For example, if we have
the categories {Red, Green, Blue}, they can be encoded as {0, 1, 2}, respectively.</p>
        <p>Binary Encoding. First, categories are converted to integers using label encoding. Then, these
integers are converted into a binary code, and each bit of the binary representation becomes a
separate feature.</p>
        <p>Step 3.2. Normalisation of numeric variables to ensure the same scale.</p>
        <p>
          Min-Max normalisation. The feature is scaled to the specified range, usually [
          <xref ref-type="bibr" rid="ref1">0, 1</xref>
          ].
        </p>
        <p>Z-score normalisation (standardisation). The data is scaled so that its mean value is 0 and
the standard deviation is 1.
or iterative methods.</p>
        <p>where μ is the mean value of the feature, σ is the standard deviation of the feature.
Step 3.3. Selection of features based on mutual information, importance of features in models,
A measure of the relationship between two variables that helps to identify how useful the
information from one variable is in predicting the other.</p>
        <p>( ;  ) = ∑ ∈ ∑ ∈  ( ,  )
(</p>
        <p>( ,  )
 ( ) ( )
)
where p(x,y) is the joint probability of two random variables X and Y, and p(x) and p(y) are
the marginal probabilities of X and Y, respectively.</p>
        <p>Importance of features in models. Some machine learning algorithms, such as random
forests, can provide estimates of the importance of features based on how much the feature
improves the partitioning criterion (e.g., reducing uncertainty).</p>
        <p>Iterative methods. Include the use of algorithms that sequentially add or remove features to
determine the optimal set of features. For example, recursive feature extraction (RFE) works by
training the model, evaluating the importance of the features, and removing the least important
features, repeating the process until a given number of features is reached.</p>
        <p>Step 4. Preparing data for modelling.</p>
        <p>Step 4.1. Cleaning the data from duplicates and errors.</p>
        <p>Identification of duplicates is carried out by identifying and deleting records that are exact
copies of other records.
logical or statistical checks.
reduces the impact of outliers:</p>
        <p>= { ∈  ∣ ∃ ′ ∈  :  =  ′ ∧  =  ′}
Error correction. Correction of inconsistencies or input errors that may be detected through
Step 4.2. Handling missing values through median imputation for numerical data, which

= 
∪ {
( ) ∣  ∈ 
}</p>
        <p>Step 4.3. Split the data into training and test samples. The partitioning can be done using a
percentage or a fixed number of records. Let D be a complete dataset, then:</p>
        <p>Training sample</p>
        <p>is a proportion p of D, where 0 &lt;  &lt; 1.</p>
        <sec id="sec-3-1-1">
          <title>The test sample  is the rest of the records.</title>
        </sec>
        <sec id="sec-3-1-2">
          <title>Step 5. Evaluation of machine learning models.</title>
          <p>=  ∣  ∣

= (1 −  ) ⋅∣  ∣</p>
        </sec>
        <sec id="sec-3-1-3">
          <title>Section 3.2)</title>
          <p>defined as:</p>
          <p>Step 5.1. Training ensemble models using machine learning algorithms, namely Random
Forest, AdaBoost, XGBoost, LightGBM, CatBoost, Gradient Boosting Machine (GBM) (see</p>
          <p>Step 5.2. Validate the models on test data and display the scores and select the best model
using the RMSE (Root Mean Square Error) metric. RMSE is the root mean square error and is

1

=
( −  )</p>
          <p>Where n is the number of observations in the test dataset,  is the actual value of the i-th
observation, and  is the predicted value for the i-th observation.</p>
          <p>The best model is selected by comparing the RMSE values for each model. The model with
the lowest RMSE value is considered to be the best, as it indicates a lower average prediction
error on the test data.</p>
          <p>Step 5.3. Optimise the best model using RandomisedSearchCV. This approach can be more
time efficient, especially when working with a large hyperparameter space. Let's say we have a
hyperparameter space H that defines all possible combinations of parameters that can be used
by a machine learning model. RandomizedSearchCV selects n random combinations of
hyperparameters from H to train and evaluate the model. The selection process can be described
as follows:</p>
          <p>Step 5.3.1. Define the hyperparameter space  = {ℎ , ℎ , . . . , ℎ }, where each ℎ can have
a different range or set of values.</p>
          <p>Step 5.3.2. Randomly select n combinations of parameters from  .</p>
          <p>Step 5.3.3. For each random combination ℎ</p>
          <p>of  :
Training the model with hyperparameters ℎ .</p>
          <p>Evaluate the model by cross-validation on the training data set.</p>
          <p>Measuring the quality of the model using a given evaluation metric, such as the mean
square error (MSE) for regression tasks or accuracy for classification tasks.</p>
          <p>Step 5.3.4: Select the combination of hyperparameters that shows the best result according
to the given evaluation metric.</p>
          <p>Mathematically, the model evaluation for each combination of hyperparameters can be
represented as follows:  (ℎ ) =
∑
 (
, 
,  )
where:  (ℎ ) is the performance estimate for the  th combination of hyperparameters, 
is the number of folds in cross-validation, L is a loss function (e.g. MSE), 
is the model
trained with the  th combination of hyperparameters, 
,  is the validation dataset for
the jth fold.</p>
          <p>Step 5.4. Diagnose the model using training and validation curves.</p>
          <p>This approach provides a systematic and comprehensive approach to analysing and
evaluating machine learning data and models, contributing to the development of accurate and
reliable intelligent film selection systems.</p>
        </sec>
      </sec>
      <sec id="sec-3-2">
        <title>3.2. Study of ensemble models</title>
        <p>In the field of machine learning, the use of ensemble methods and specialised algorithms to
improve prediction accuracy is critical for solving complex problems. Ensemble methods, such
as Random Forest, AdaBoost, XGBoost, LightGBM, CatBoost, Gradient Boosting Machine
(GBM), provide powerful tools for building more reliable and accurate models.</p>
        <p>Random Forest (Fig. 1) is an ensemble machine learning method that uses multiple decision
trees to achieve higher prediction accuracy than is possible with a single decision tree. The basic
idea is to build a large number of decision trees, each of which contributes to the final solution,
which provides a high level of accuracy and control over retraining.</p>
        <p>Random Forest models the answer as the aggregate result of the predictions of a set of
decision trees. Mathematically, the answer Y for a classification task can be defined as the most
frequently predicted class among all N trees, while for a regression task, the answer is the
average of the predictions of all trees. For classification:
 =</p>
        <p>{ ,  , … ,  }
where  is the prediction of the i-th tree, and for regression:
 =

1

where  is the prediction of the i-th tree, and N is the total number of trees in the forest.
This approach reduces the variability and errors inherent in single decision trees and improves
the accuracy and generalisability of the model, providing effective management of overfitting
through tree diversification.</p>
        <p>AdaBoost (Adaptive Boosting) (Fig. 2) is a boosting technique that creates a strong classifier
by combining many weak classifiers. It works by sequentially improving weak classifiers by
focusing on cases that were misclassified by previous classifiers.</p>
        <p>Mathematically, for each iteration t, each training example i is assigned a weight  , which
is adaptively updated depending on whether the observation was correctly classified. The final
prediction of the model is the weighted sum of the predictions of all weak classifiers:
where ℎ () is the prediction of the t-th weak classifier on the input example  ,  is the
weight assigned to this classifier, which depends on its accuracy, and T is the total number of
weak classifiers. The weights αt are determined based on the classification error  , and the
smaller the error, the higher the weight assigned to the classifier. The weights of the training
examples are updated so that examples that were misclassified receive higher weights, forcing
the next classifier to focus on these harder cases.</p>
        <p>XGBoost (Extreme Gradient Boosting) (Fig. 3) is an efficient and scalable implementation of
gradient boosting. It includes a number of optimisations for speed and performance, and has
built-in tools to prevent overfitting.</p>
        <p>Mathematically, XGBoost seeks to minimise the following objective function in the t-th step,
which includes both a loss function L and a regularisation ΩΩ to control the model complexity:

( ) =
( ,  (
) +  ( )) + (
)
where  are the actual values,  ( ) are the predictions at step  ( ) is the prediction
made by the t-th tree on the input example  , and ( ) is the regularisation term for the t-th
tree, which typically includes both the number of leaves in the tree and the sum of the squares
of the leaf weights to avoid overfitting. XGBoost uses this formula to improve the predictions
at each step, effectively finding the direction in which to go to reduce errors while keeping the
model simple enough to avoid overfitting.</p>
        <p>LightGBM is an efficient implementation of the gradient boosting algorithm that is
optimised for speed and performance. LightGBM uses histogram-based methods to reduce
computational and memory consumption, making it particularly useful for processing large
datasets.</p>
        <p>CatBoost is an algorithm that specialises in working with categorical data, using special
coding techniques to process this type of data without preprocessing. It also includes
mechanisms to combat overfitting, which allows for stable results.</p>
        <p>Formally, the model seeks to minimise the objective function:</p>
        <p>= (,  ) + ()
where (,  , ) defines the loss between the actual values of y and the model predictions of
 ̂, and () expresses the regularisation term that controls the model complexity. A key
feature of CatBoost is its ability to automatically handle categorical variables, efficiently
encoding them and using them to improve model accuracy, making it particularly powerful in
situations where other gradient boosting algorithms may require complex data preprocessing.</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>4. Implementation</title>
      <p>To implement the proposed approach, we chose the MovieLens 100K dataset [12], developed by
GroupLens Research, which is a classic dataset for recommender systems that contains 100,000
ratings from 943 users for 1,682 films. The ratings are on a scale from 1 to 5, where each rating
is associated with a specific user and film, including a timestamp when the rating was made. In
addition to ratings, the dataset includes demographic information about users (age, gender,
profession, postal code) and metadata about films (genre, title).</p>
      <p>To implement the proposed approach, the following parameters were selected from the
dataset (see Table 1) and used to store information about films and their ratings by users. Each
row of the table will display a unique user and the film they have rated, along with the rating
they have given, as well as additional information about the film such as director, title, release
date and genre. The genre can be represented as a single line that includes all genres of the film,
or as multiple binary columns, each indicating the presence or absence of a particular genre.</p>
      <p>The graph (Fig. 7) shows the distribution of film ratings that users have left in the dataset.
The X-axis represents the possible ratings from 1 to 5, where each rating is displayed as a
separate column. The Y-axis shows the frequency of the rating distribution as a proportion of
the total number of ratings. The graph shows that the least popular ratings are "1" and "2",
which make up a smaller proportion of all ratings given. The rating "3" has a slightly higher
frequency, but a much larger number of users preferred the higher ratings of "4" and "5", with
"4" being the most frequently given rating. This may indicate a positive slope in the distribution
of ratings, indicating a tendency for users to leave higher ratings.</p>
      <p>Next, we compare the RMSE values for different machine learning models: Random Forest,
AdaBoost, XGBoost, LightGBM, CatBoost, and Gradient Boosting Machine (GBM). The graph
shows (Fig. 8) that XGBoost has the lowest RMSE (0.902), which indicates higher prediction
accuracy compared to other models. LightGBM and CatBoost also perform competitively with
RMSEs of 0.910 and 0.919, respectively, indicating their effectiveness in the prediction task.
GBM has a slightly higher RMSE of 0.942, which is better than Random Forest and AdaBoost
with RMSEs of 1.074 and 1.037, respectively. The higher RMSE values for Random Forest and
AdaBoost may indicate a lower ability of these models to accurately predict the data compared
to the other techniques considered.</p>
      <p>Next, based on the best model, namely XGBoost, we will test the output of the results
displayed on the user interface (Fig. 9) for the web application for film rating. The UI allows
users to get a predicted film score based on the data they enter. Users enter a film title
("Interstellar"), a release date ("November 5, 2014"), a genre (in this case, "Adventure, Drama,
Sci-Fi", separated by commas), and a director's name ("Christopher Nolan"). Once the data is
entered, the XGBoost model processes this information and produces a predicted rating for the
film (in this case, "Rating: 4.07"), which is displayed in the interface window.</p>
      <p>Model RMSE
Random Forest 1.074
AdaBoost 1.037
XGBoost 0.902
LightGBM 0.910
CatBoost 0.919
GBM 0.942</p>
      <p>In summary, this study is distinguished by a unique integrated approach that combines
advanced machine learning techniques to create a highly accurate and flexible film
recommendation system. Through the use of feature engineering, data normalisation
techniques, and iterative feature selection, our model effectively predicts user preferences based
on their rating history, socio-demographic data, and viewing context. The use of ensemble
methods, namely XGBoost, significantly reduced the RMSE error compared to alternative
models described in [9-11], highlighting the importance of this study in the development of
effective recommender systems that adapt to diverse user preferences and contexts.</p>
    </sec>
    <sec id="sec-5">
      <title>5. Conclusions</title>
      <p>In this study, we examined the use of ensemble machine learning methods to create an
intelligent film selection system that is highly accurate and adaptive to individual user
preferences. By analysing different algorithms such as Random Forest, AdaBoost, XGBoost,
LightGBM, CatBoost, and Gradient Boosting Machine, we found that XGBoost demonstrates
the best results with the lowest RMSE value of 0.902. This indicates that XGBoost is highly
effective in predicting film scores compared to the other models considered.</p>
      <p>The use of feature engineering, data normalisation, and iterative feature selection allowed
us to improve the model's ability to accurately predict users' interests, taking into account not
only their previous ratings, but also socio-demographic information and viewing context. This
approach provided a significant reduction in the RMSE error compared to other models
mentioned in studies [9-11], emphasising the importance of an integrated approach to the
development of recommender systems.</p>
      <p>The implementation of the proposed approach on the MovieLens 100K dataset has shown
its practical applicability and effectiveness. The use of detailed data analysis, including
descriptive statistics, visualisation, anomaly detection, missing value processing, and
correlation analysis, allowed us to better understand the features of the dataset and prepare it
for effective modelling.</p>
      <p>As a result, this study demonstrates that the use of ensemble machine learning methods,
such as XGBoost, combined with careful data preparation and feature engineering, can
significantly improve the accuracy of film recommendation systems. This opens up wide
prospects for further research in this area, including the development of new methods to
improve prediction accuracy, as well as the adaptation of the system to different conditions and
user needs.
[8] Y. Bengio, A. Courville, I. Goodfellow, Deep Learning. MIT Press, 2016.</p>
      <p>https://www.deeplearningbook.org/
[9] S. S. Choudhury, S. N. Mohanty, A. K. Jagadev, Multimodal trust based recommender
system with machine learning approaches for movie recommendation. International
Journal of Information Technology 13 (2021) 475-482.
https://doi.org/10.1007/s41870-02000553-2
[10] C. Biancalana, F. Gasparetti, A. Micarelli, A. Miola, G. Sansonetti, Context-aware movie
recommendation based on signal processing and machine learning. Proceedings of the 2nd
Challenge on Context-Aware Movie Recommendation CAMRa '11, October 2011, pp. 5-10.
https://doi.org/10.1145/2096112.2096114
[11] J. Lund, Y.-K. Ng, Movie Recommendations Using the Deep Learning Approach.</p>
      <p>Proceedings of the 2018 IEEE International Conference on Information Reuse and
Integration for Data Science (IRI), 2018, pp. 47-54. IEEE.
https://doi.org/10.1109/iri.2018.00015
[12] MovieLens 100K Dataset. (n.d.). GroupLens. URL:
https://grouplens.org/datasets/movielens/100k/
[13] V. Golovko, Y. Savitsky, T. Laopoulos, A. Sachenko and L. Grandinetti, Technique of
learning rate estimation for efficient training of MLP, Proceedings of the IEEE-INNS-ENNS
International Joint Conference on Neural Networks, IJCNN 2000, Neural Computing: New
Challenges and Perspectives for the New Millennium, Como, Italy, 2000, vol. 1, pp.
323328. https://doi.org/10.1109/IJCNN.2000.857856
[14] S. Anfilets, S. Bezobrazov, V. Golovko, A. Sachenko, M. Komar, R. Dolny, V. Kasyanik, P.</p>
      <p>Bykovyy, E. Mikhno, &amp; O. Osolinskyi, Deep multilayer neural network for predicting the
winner of football matches. International Journal of Computing 19 (2020) 70-77.
https://doi.org/10.47839/ijc.19.1.1695
[15] I. Paliy, A. Sachenko, V. Koval and Y. Kurylyak, "Approach to Face Recognition Using
Neural Networks," 2005 IEEE Intelligent Data Acquisition and Advanced Computing
Systems: Technology and Applications, Sofia, Bulgaria, 2005, pp. 112-115,
https://doi.org/10.1109/IDAACS.2005.282951
[16] R. Gramyak, H. Lipyanina-Goncharenko, A. Sachenko, T. Lendyuk, D. Zahorodnia,
Intelligent Method of a Competitive Product Choosing based on the Emotional Feedbacks
Coloring. In IntelITSIS, 2021, pp. 246-257. https://ceur-ws.org/Vol-2853/paper31.pdf
[17] H. Lipyanina, S. Sachenko, T. Lendyuk, V. Brych, V. Yatskiv, O. Osolinskiy, (2021). Method
of detecting a fictitious company on the machine learning base. In International
Conference on Computer Science, Engineering and Education Applications (pp. 138-146).
Cham: Springer International Publishing (2021).
https://doi.org/10.1007/978-3-030-804725_12
[18] H. Lipyanina, V. Maksymovych, A. Sachenko, T. Lendyuk, A. Fomenko, I. Kit, Assessing
the investment risk of virtual IT company based on machine learning. In International
Conference on Data Stream Mining and Processing (pp. 167-187). Cham: Springer
International Publishing (2020). https://doi.org/10.1007/978-3-030-61656-4_11
[19] V. Turchenko, E. Chalmers, &amp; A. Luczak, A deep convolutional auto-encoder with pooling
– unpooling layers in caffe. International Journal of Computing 18 (2019) 8-31.
https://doi.org/10.47839/ijc.18.1.1270
[20] A. R. Marakhimov, &amp; K. K. Khudaybergenov, Approach to the synthesis of neural network
structure during classification. International Journal of Computing 19 (2020)
2026. https://doi.org/10.47839/ijc.19.1.1689</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>L.</given-names>
            <surname>Breiman</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Random</given-names>
            <surname>Forests</surname>
          </string-name>
          .
          <source>Machine Learning</source>
          <volume>45</volume>
          (
          <year>2001</year>
          )
          <fpage>5</fpage>
          -
          <lpage>32</lpage>
          . https://doi.org/10.1023/A:1010933404324
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>Y.</given-names>
            <surname>Freund</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R. E.</given-names>
            <surname>Schapire</surname>
          </string-name>
          ,
          <article-title>A decision-theoretic generalization of on-line learning and an application to boosting</article-title>
          .
          <source>Journal of Computer and System Sciences</source>
          <volume>55</volume>
          (
          <year>1997</year>
          )
          <fpage>119</fpage>
          -
          <lpage>139</lpage>
          . https://doi.org/10.1006/jcss.
          <year>1997</year>
          .1504
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>I. M.</given-names>
            <surname>Sukarsa</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N. N. Pandika</given-names>
            <surname>Pinata</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N. K.</given-names>
            <surname>Dwi Rusjayanthi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N. W.</given-names>
            <surname>Wisswani</surname>
          </string-name>
          ,
          <article-title>Estimation of Gourami supplies using gradient boosting decision tree method of XGBoost</article-title>
          .
          <source>TEM Journal</source>
          <volume>10</volume>
          (
          <year>2021</year>
          )
          <fpage>144</fpage>
          -
          <lpage>151</lpage>
          . https://doi.org/10.18421/tem101-
          <fpage>17</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>J.</given-names>
            <surname>Lyu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Zheng</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Qi</surname>
          </string-name>
          , G. Huang,
          <string-name>
            <surname>LightGBM-LncLoc</surname>
          </string-name>
          :
          <article-title>A LightGBM-Based Computational Predictor for Recognizing Long Non-Coding RNA Subcellular Localization</article-title>
          .
          <source>Mathematics</source>
          <volume>11</volume>
          (
          <year>2023</year>
          )
          <article-title>602</article-title>
          . https://doi.org/10.3390/math11030602
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>L.</given-names>
            <surname>Qian</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Chen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Huang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R. J.</given-names>
            <surname>Stanford</surname>
          </string-name>
          ,
          <article-title>Employing categorical boosting (CatBoost) and meta-heuristic algorithms for predicting the urban gas consumption</article-title>
          .
          <source>Urban Climate</source>
          <volume>51</volume>
          (
          <year>2023</year>
          )
          <article-title>101647</article-title>
          . https://doi.org/10.1016/j.uclim.
          <year>2023</year>
          .101647
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>Z.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S. Ameenuddin</given-names>
            <surname>Irfan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Teoh</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Hriday</surname>
          </string-name>
          <string-name>
            <surname>Bhoyar</surname>
          </string-name>
          ,
          <article-title>Gradient Boosting</article-title>
          .
          <source>In Numerical Machine Learning</source>
          (pp.
          <fpage>116</fpage>
          -
          <lpage>159</lpage>
          ).
          <source>BENTHAM SCIENCE PUBLISHERS</source>
          ,
          <year>2023</year>
          . https://doi.org/10.2174/9789815136982123010007
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>Support</given-names>
            <surname>Vector</surname>
          </string-name>
          <article-title>Machines for Regression. (n.d.)</article-title>
          .
          <source>In Support Vector Machines</source>
          (pp.
          <fpage>330</fpage>
          -
          <lpage>351</lpage>
          ). Springer New York,
          <year>2008</year>
          . https://doi.org/10.1007/978-0-
          <fpage>387</fpage>
          -77242-
          <issue>4</issue>
          _
          <fpage>9</fpage>
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>