<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Utilizing Metadata to Select a Recommendation Algorithm for a User or an Item</article-title>
      </title-group>
      <contrib-group>
        <aff id="aff0">
          <label>0</label>
          <institution>ITMO University</institution>
          ,
          <addr-line>St. Petersburg</addr-line>
          ,
          <country country="RU">Russia</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Vyatka State University</institution>
          ,
          <addr-line>Kirov</addr-line>
          ,
          <country country="RU">Russia</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>In general, recommender systems solve the problem of information overload by helping users of services to nd items in which they are interested. There are plenty of algorithms and models that can be used to build such a subset of items, and their performance may vary not only for di erent datasets but for separate parts of a single one. This issue leads to the "algorithm selection problem". Many state-of-art solutions to this problem are based on meta-learning. It associates the task features with the performance of the base-level algorithms and models. This paper presents the method of recommendation algorithm selection for particular users (or items) that uses binary representations of both explicit metadata of them and computable statistical meta-features. There are two di erent techniques within the proposed method, which are based on classi cation or clustering of such binary data, respectively. The metalearning process is almost automated. The ndings of the experiments prove that the usage of the method for recommendation algorithm selection is reasonable and e ective. In most cases, a recommender system that uses the metamodel shows lower rating prediction errors compared to any other one utilizing a single model or algorithm for all the users (or items), while in a small number of tests their performance is just the same. The detailed analysis of the evaluation results allows for a rming that the described metamodels can be used in real-world systems to improve the experience of particular users.</p>
      </abstract>
      <kwd-group>
        <kwd>Recommender System Algorithm Selection Problem MetaLearning Collaborative Filtering</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>Numerous modern companies provide their clients with a great variety of
services, products, and content. Such diversity leads to the problem of information
overload that has an adverse e ect on both user experience and service growth.
Recommender systems (RSs) are used to solve this problem e ectively. They
automatically create suggestions of items that most likely lie in a user's area of
interest.</p>
      <p>
        Modern recommendation algorithms are actively used in real-world systems [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ].
Even a moderate improvement in the personalization of the collections of items
is bene cial for both services' clients and its owners. Moreover, the existence of
instruments predicting a user's interests makes it possible to use recommender
systems in such non-standard areas as medicine and law. Nonetheless, the
accuracy of these algorithms is relatively low. The need for functioning in the poorly
formalizable domain of people's interests leads to the vast complexity of the
development of e ective algorithms as well as to the inability to create universal
ones. Usually, RS designers implement and evaluate di erent models and
several versions of their compositions. This approach needs a signi cant number of
computational resources, but the result can be suboptimal [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ].
      </p>
      <p>The crucial problem that arises is that the performance of a model may
vary for di erent users and items within a single dataset. For example, a matrix
factorization model can predict the interests of a particular service user precisely,
while the k-Nearest-Neighbors algorithm is much more applicable for making
recommendations for another individual. A state-of-art solution to this problem
is the use of meta-learning.</p>
      <p>
        Meta-learning lies in the creation of metamodels that associate task
features with algorithms' performance. Such metamodels can be separated into
the global-level (of a dataset), middle-level (of particular users or items), and
micro-level (of particular rows) ones [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]. Micro-level metamodels have recently
been proposed and not been studied enough, since they need massive and dense
datasets to be e ective. There are plenty of researches within the rst two levels,
including the meta-recommender system methods. Within this eld, the
contribution of Cuhna et al. [
        <xref ref-type="bibr" rid="ref4 ref5">4,5</xref>
        ] and their meta-research [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ] is particularly notable.
The authors' studies provide an in-depth review of the existing meta-learning
approaches to the problem of algorithm selection in the domain of recommender
systems.
      </p>
      <p>In most cases, a metamodel utilizes only the data about the base models'
performance. The real-world systems are generally aware of users' and items'
features, even if they are not used in the recommendations generation process.
The core idea behind this research is that a metamodel can take a cue from
both such metadata and standard meta-features. Moreover, when a metamodel
training dataset is represented in a binary format, it allows for applying the same
algorithms in di erent business cases.</p>
      <p>This paper proposes the middle-level automated meta-learning method that
uses the recommendation algorithms' evaluation results as well as the binary
representation of both users' (or items') metadata and their computable
metafeatures.</p>
    </sec>
    <sec id="sec-2">
      <title>Method</title>
      <p>The formal statement of the research problem is as follows. There is a dataset
consisting of (u; i; rui) rows that show what rating rui has the user u given to
the item i. Also, there are users' and items' metadata stored as binary vectors.
Several recommendations models (base models) are trained on the part of a
dataset, and evaluated on another part, which has given the predictions (u; i; rbui).
The task is to create a metamodel that selects the best base model for a particular
user (or item), i.e. makes a personalized selection.</p>
      <p>It should be stated that below is the solution that utilizes users' metadata.
The solution that uses items' metadata is the same.</p>
      <p>
        At this stage, two di erent solutions to the problem are proposed. On the
one hand, one can take the evaluation results as the starting point and make a
dataset showing the correspondence between metadata of the user and the best
model for them. In this case, the task is to solve a classi cation problem. On the
other hand, one can start from the nature of users, group them using particular
features, and select the best model for each built group. In this case, the task is
to solve a clustering problem [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ].
      </p>
      <p>While a metamodel is being trained, its input is built from two parts. The
rst one is a set of explicit users' metadata, taken from a dataset, e.g. age, sex,
or profession. The second one contains meta-features computed on the basis of
users' rating information, i.e., some standard values used in meta-learning.</p>
      <p>Fig. 1 shows the proposed recommender system structure.</p>
      <p>Both parts are represented as binary vectors. It is assumed that users'
metadata are already binary, while the computable features are converted after a
calculation.</p>
      <p>Both proposed meta-learning techniques suggest the users are grouped. The
best base model must be chosen for each users' group U from the whole set of
models M . The selection is performed by the equation
mbest = argmax( X f (m; u));</p>
      <p>m2M u2U
where f (m; u) is the utility function for the model m and the user u that grows
together with the model's performance. The de nition of such a function depends
on the error metric in use. Since RMSE is highly applicable to the task of rating
prediction, the utility function is de ned as follows.</p>
      <p>Let r 2 RN be the vector of actual ratings rj , rm 2 RN be the vectors of
b
predictions rbmj which model m has given for the row j. The vector emin 2 RN
contains values
ejmin = mm2iMn((rbmj
rj )2)
equal to minimal squares of error for each row. Thus, the utility function for the
model m and the user u is computed by the formula
f (m; u) =</p>
      <p>X
i.e. as the cumulative di erence between the minimal squared errors and the
models' squared errors.</p>
      <p>When one chooses an algorithm for a users' group, they should consider the
fact that several base-level models can show roughly equal performance, but one
of them is the best on a whole dataset, i.e., ts the task and its domain more
than others. Another model may show a higher level of utility not only because
of its accordance with this users' group but due to the existence of outlying
values in a training dataset, causing errors in modeling of interests. It may lead
to the situation when a wrong base-level model that has learned several random
records, but not the actual patterns of users' behavior, is selected.</p>
      <p>Consequently, there should be a strong reason to prefer a particular algorithm
for a users' group over the overall best one. In this research, the stated problem
is solved as follows. After the model mU best that has the highest utility for a
group is chosen, it is calculated how much higher its utility is compared to the
one of the overall best model moverall best. If this value is lower than the constant
, the model moverall best is used for a users' group.</p>
      <p>Since a utility function may take strictly non-negative values as well as
strictly non-positive ones, the nal selection of a model mU for a users' group U
is performed by equations:
utilityU =</p>
      <p>X f (mU best; u);
utilityoverall =
u2U
X f (moverall best; u);
u2U
kimp = max (</p>
      <p>utilityU
utilityoverall
; utilityoverall );</p>
      <p>utilityU
mU =
(mU best
moverall best
if kimp &gt;=
if kimp &lt;
:
(4)
(5)
(6)
(7)</p>
      <p>The meta-learning techniques used within the method are described below.
2.1</p>
      <sec id="sec-2-1">
        <title>Computation of Meta-Features</title>
        <p>
          Meta-features, which are commonly used in metamodel training, can be
separated into three groups [
          <xref ref-type="bibr" rid="ref7">7</xref>
          ].
1. Statistical and/or information-theoretical meta-features describe a dataset
or its parts using particular metrics from corresponding sciences, which are
applied to data analysis. Such metrics provide certain knowledge about users'
behavior and can be used for distinguishing and grouping them. Examples of
these features are the record count, the expected value for separate columns,
their entropy, skewness, and kurtosis.
2. Base-level model features. This group includes hyperparameters and the
features of trained base-level models. The examples of the latter include the leaf
count in a decision tree or weights of input neurons in the feed-forward
network. These metrics can be used to discover implicit dependencies between
a model and separate parts of a dataset.
3. The last kind of metrics that are commonly used in meta-features
computation is "landmarkers", i.e., fast and rough estimations of di erent models'
performance. To calculate them, one can use simpli ed models or small parts
of a whole dataset.
        </p>
        <p>In terms of the current research, the meta-features from the rst group are of
particular interest. Base-level model features are not studied, since it is assumed
that all the models are of di erent kinds, and not the ones of a single kind that
di er only by hyperparameters. The usage of "landmarkers" is also not studied,
as the stated task implies the hybridization of several models that are trained
already.</p>
        <p>Several statistical features have been chosen; they describe rating dataset
parts that correspond to separate users. For each of these features, the algorithm
of conversion to 0-or-1 columns is described. The meta-features count has been
restricted to three to avoid the increase in the width of the metamodel training
dataset. The meta-features are shown in Table 1.
Meta-Feature Conversion to the Binary Format
Relative count of user's ratings. Three columns for ranges
Equals to count of user's rating divided by [0; 0:75), [0:75; 1:25), [1:25; +1).
the average rating count per user
in a whole dataset.</p>
        <p>Standard deviation of user's ratings.</p>
        <p>Two columns for ranges:
[0; 0:1 of maximal possible rating),
[0:1 of maximal possible rating; +1).</p>
        <p>Two columns for positive and negative
values respectively.</p>
        <p>User's ratings skewness.
2.2</p>
      </sec>
      <sec id="sec-2-2">
        <title>User Classi cation</title>
        <p>Since the input consists of binary vectors, a classi er can distinguish any two
users having non-equal metadata vectors. Thus, the set U is de ned as the group
of users having equal metadata vectors. The best model is selected for each set
U using formulas (1) { (7). These actions build the dataset for a classi er, which
associates the metadata vector with the base model as a class label. Obviously,
if the dataset includes all the possible varieties of binary metadata vectors, the
classi er is not needed. However, most likely, the dataset does not cover all of
them.</p>
        <p>It is considered that the model count is no less than three, so it is better to
train an ensemble of binary classi ers using the "One-vs-Rest" strategy. Since
the metadata consists of binary vectors, the ensemble can be e ectively founded
on any decision-tree based classi ers. In this research, the random forest classi er
is used.
2.3</p>
      </sec>
      <sec id="sec-2-3">
        <title>User Clustering</title>
        <p>The clustering-based technique brings a speci c problem. A clusterer can
consider some users "noise" and not attach them to any group. In this case, the
meta-learning process is organized as follows.</p>
        <p>At rst, the overall set of users' metadata is clustered. As a result, several
clusters are obtained, each of which can be considered as a target group U , along
with the separate part of users not belonging to any cluster. The best models for
clustered users are chosen using the formulas (1) { (7). For the "noise" group,
a metamodel uses the overall best base model, since it is incorrect to consider
these users as a neighborhood of any kind.</p>
        <p>
          The task of clustering of binary vectors can be considered as the one of point
clustering in the multi-dimensional Euclidean space. It is possible if at least one
of the two following conditions is met [
          <xref ref-type="bibr" rid="ref8">8</xref>
          ]. The rst one is that a dataset must
not be very sparse. The second one states that a dataset must not be very wide.
It is challenging to provide more precise restrictions, but when both conditions
are not met, a clustering model trains on a wide, sparse dataset. In this case,
distance metrics applied in Euclidean spaces become almost useless and cannot
be utilized to distinguish points properly. Nonetheless, it is assumed in this
research that the count of binary columns is less than 100, while the acceptable
density can at least be provided by the computable meta-features.
        </p>
        <p>Thus, the OPTICS model with cosine similarity as a distance metric has
been chosen initially to perform clustering. It has been stated heuristically that
the farthest two users can di er in no more than two columns and coincide in
at least one another column. So, according to the distance metric, the maximal
cluster size equals 1 1=(p2 p2) = 0:5.</p>
        <p>
          However, the OPTICS model has a signi cant disadvantage. Its accuracy is
high, but the training process on the datasets used in the experimental study
can take an unacceptably long amount of time. This problem can be solved
by using the HDBSCAN model [
          <xref ref-type="bibr" rid="ref9">9</xref>
          ]. Like OPTICS, it is an improved version of
DBCSAN that performs hierarchical clustering without the xed restriction of
maximal cluster radius. It has been discovered empirically that the HDBSCAN
model with L2 distance metric shows almost the same level of performance on
the used datasets as the OPTICS one with the stated above con guration. It
shows highly similar results, but learns much faster. So, this is the HDBSCAN
model that is used further.
3
        </p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>Results and Discussion</title>
      <p>
        The initial pool of algorithms in this research included ve collaborative ltering
models, implemented in the Surpriselib library [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ]: SVD, SVD++, Baseline only,
KNN Baseline, Co-Clustering.
      </p>
      <p>
        Five datasets [
        <xref ref-type="bibr" rid="ref11 ref12 ref13">11,12,13</xref>
        ] were used. They are described in 2. The presented
features do not include computable ones. Metadata rows were converted to
binary vectors manually in such a way that most of the information was saved,
but the width of the dataset did not increase drastically.
      </p>
      <p>Each dataset was used in the following way. It was shu ed randomly and
split into three parts A, B, and C with 60, 20%, and 20% of records from the
initial dataset, respectively. Each model from the initial pool was trained on part
A and evaluated on part B. Based on the evaluation results, the pool of the best
base models was built. Next, four metamodels were trained, each of which was
classifying/clustering users/items. The constant was de ned as 1.25. During
the process of classi er training, only 70% of users/items from part B were used.
Finally, all the base models and created metamodels were evaluated on the part
C. The RMSE was used as the evaluation metric.</p>
      <p>The evaluation results are presented in Table 3. Additionally, the table
includes the results for a perfect abstract metamodel that chooses the truly best
algorithm for each user or item.</p>
      <p>In four out of ve cases, the best metamodel had better performance than the
best base-level one, since its RMSE on a whole dataset was lower. The greatest
bene t from the meta-learning was taken for the Modcloth dataset. In the case
of Movielens 1M dataset, the best metamodel had shown exactly the same result
as the best base-level one.</p>
      <p>The improvement achieved by applying the proposed method was di erent
for di erent datasets. It is possible to suggest that it depended both on the best
base-level model accuracy and on the level of maximal theoretical improvement
that could be obtained by per-user (or per-item) algorithm selection. Figure 2
illustrates the relation between these values.</p>
      <p>According to the graphs, the actual improvement that the metamodel usage
made could be truly dependent on the maximal possible one. At the same time,
the possibility of such dependency on the relative best base-level model error
was questionable due to the results for the Rent the Runway dataset. It was
important that previous experiments that had been performed without statistical
meta-features and the usage of the constant showed that this dependency
could exist, and the results had been generally worse. It allows for a rming
that combining explicit metadata together with computable meta-features and
performing a restricted selection of the best model for a users' (or items') group
is reasonable and increases the accuracy of metamodels.</p>
      <p>Both the numerical results and the graphs showed that the actual
improvement was far from theoretically possible. Along with that, training accuracy was
reasonably good. The graphs in Figure 3 demonstrate how often metamodels
have selected the truly best algorithm for a user or an item.</p>
      <p>30%
20%
10%
0%</p>
      <p>MM (user class.)</p>
      <p>MM (item class.)</p>
      <p>MM (user clust.)</p>
      <p>MM (item clust.)
According to these measurements, classi cation and clustering models had a
rather similar performance that led to the best selection rate that equaled
2428%. Worse results showed by item-based models for three out of ve datasets
could be explained by relatively small item counts in them. It allows for
concluding that the bottleneck lies in both dataset size and the complexity of
determining the truly best model for users' or items' groups before training.</p>
      <p>A pending question within this research is how to choose an optimal value for
the constant . Figure 4 shows the ratios between the lowest metamodel RMSE
and the lowest base-level model RMSE with di erent values. When is less
or equal to 1, it is equivalent to to be ignored (following formulas (4) { (7)
from Section 2).</p>
      <p>In three out of ve cases, the usage of the constant was e ective. In the case
of the Modcloth dataset, there was almost no di erence between the evaluated
values. For the Amazon Phones dataset, the restrictions in the algorithm
selection the constant brought had a ected metamodels' performance negatively.
Consequently, the optimal value depended on the used dataset and
performance of the base-level models. Since the evaluation of several metamodels with
di erent values has a high computational cost, it is highly preferable to
deter</p>
      <p>Modcloth
mine its value before the training. Possibly, this task can be done on the basis
of the base-level models' errors. This problem requires additional research.</p>
      <p>The impact of the usage of the metamodels on the time needed for
recommendation generation has also been considered. Figure 5 shows the average times the
base-level and meta- models need to make a single rating prediction on the used
datasets. The experiments have been performed using the Intel Core i7-7700HQ
CPU and DDR-4 RAM on 2133 MHz frequency. It is signi cant that the
Surpriselib framework evaluates test requests individually, and not in batches. This
peculiarity allows the measured time values to be as realistic as possible.
500
400
300
200
100
0</p>
      <p>SVD</p>
      <p>SVD++</p>
      <p>KNN
Baseline</p>
      <p>Baseline Co-Clustering MM (user
Only class.)</p>
      <p>MM (item
class.)</p>
      <p>MM (user
clust.)</p>
      <p>MM (item
clust.)</p>
      <p>Depending on a dataset, when a metamodel was used, the average prediction
time was ranged between 143 and 413 microseconds. This interval allowed an RS
to process thousands of requests per second. The speed of base-level models, in
some cases, was comparable to the metamodels' one, but, in some other cases, it
was noticeably higher. However, the base-level models have particularly native
implementation, while the metamodels are implemented in pure Python without
any optimizations. Besides, it is obvious and seen from the graphs that the speed
of a hybrid model depends on the speed of all the models it combines directly.
Thus, the prediction time of a recommender system that uses the metamodel
is enough for the production but it can be decreased by optimizations and the
native implementation.</p>
      <p>The results show that all the metamodels are equally performant. To choose
one of them in real-world tasks, one should consider the number of users and
items in the available data. The classi cation-based metamodels are preferred
in this case, since they can be trained faster than the clustering-based ones. To
employ the proposed method in an existing system that includes several
recommenders, one can convert users' metadata to the binary vectors format, train
the classi er using the existing recommenders' evaluation results, and redirect
recommendation requests to the metamodel.</p>
      <p>Although the improvement is moderate, the usage of the metamodel that
utilizes the peculiarities of a particular user can noticeably increase the quality of
their experience. To examine this assumption, one should conduct additional
research that aims at the evaluation of ranking measures and uses the appropriate
datasets.
4</p>
    </sec>
    <sec id="sec-4">
      <title>Conclusions</title>
      <p>The presented work describes a solution to the problem of recommendation
algorithm selection for particular users or items within a single service. The paper
proposes the middle-level automated meta-learning method that uses the rating
prediction algorithms' evaluation results as well as the binary representation of
both users' (or items') metadata together with their computable meta-features.
Two techniques of creating the metamodel are suggested; they are based on
classi cation and clustering, respectively. Both techniques are applicable for users'
features as well as for an items' features.</p>
      <p>The metamodels have been evaluated on ve open datasets, with the usage
of collaborating ltering models on a base level. The experiments prove that the
usage of the proposed method is reasonable and e ective. The performance in
comparison to the base-level models improved up to 3.81% of RMSE value in
a best-case scenario and did not decrease in any case. The measures obtained
from the evaluation have been analyzed in detail, and the possible directions of
future work been described, including the ones aimed at performance
enhancement. Current results can be used in real-world systems; they can improve the
experience of their users. The additional advantage of the method is that the
process of creation of a metamodel is almost automated, and the only manual
task is to convert explicit metadata to the binary format.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Ricci</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rokach</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Shapira</surname>
          </string-name>
          , B. eds: Recommender Systems Handbook. Springer US (
          <year>2015</year>
          ). https://doi.org/10.1007/978-1-
          <fpage>4899</fpage>
          -7637-6
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Cunha</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Soares</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>de Carvalho</surname>
            ,
            <given-names>A.C.P.L.F.</given-names>
          </string-name>
          :
          <article-title>Metalearning and Recommender Systems: A literature review and empirical study on the algorithm selection problem for Collaborative Filtering</article-title>
          .
          <source>Information Sciences</source>
          .
          <volume>423</volume>
          ,
          <issue>128</issue>
          {
          <fpage>144</fpage>
          (
          <year>2018</year>
          ). https://doi.org/10.1016/j.ins.
          <year>2017</year>
          .
          <volume>09</volume>
          .050
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Collins</surname>
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Beel</surname>
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Tkaczyk</surname>
            <given-names>D</given-names>
          </string-name>
          .
          <article-title>One-at-a-time: A Meta-Learning RecommenderSystem for Recommendation-Algorithm Selection on Micro Level</article-title>
          . ArXiv abs/
          <year>1805</year>
          .12118 (
          <year>2018</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Cunha</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Soares</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>de Carvalho</surname>
            ,
            <given-names>A.C.P.L.F.</given-names>
          </string-name>
          :
          <article-title>Selecting Collaborative Filtering Algorithms Using Metalearning</article-title>
          .
          <source>In: Machine Learning and Knowledge Discovery in Databases</source>
          . pp.
          <volume>393</volume>
          {
          <fpage>409</fpage>
          . Springer International Publishing (
          <year>2016</year>
          ). https://doi.org/10.1007/978-3-
          <fpage>319</fpage>
          -46227-1 25
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Cunha</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Soares</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>de Carvalho</surname>
          </string-name>
          , A.
          <string-name>
            <surname>C.P.L.F.:</surname>
          </string-name>
          <article-title>CF4CF</article-title>
          .
          <source>In: Proceedings of the 12th ACM Conference on Recommender Systems - RecSys '18</source>
          . ACM Press (
          <year>2018</year>
          ). https://doi.org/10.1145/3240323.3240378
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Meltsov</surname>
            <given-names>V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Novokshonov</surname>
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Repkin</surname>
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Nechaev</surname>
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zhukova</surname>
            <given-names>N.</given-names>
          </string-name>
          :
          <article-title>Development of an intelligent module for monitoring and analysis of client's bank transactions</article-title>
          . In: Conference of Open Innovations Association,
          <source>FRUCT. N 24</source>
          .
          <fpage>255</fpage>
          -
          <lpage>262</lpage>
          (
          <year>2019</year>
          ). https://doi.org/10.23919/fruct.
          <year>2019</year>
          .8711931
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>Brazdil</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Giraud-Carrier</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Soares</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Vilalta</surname>
          </string-name>
          , R.:
          <source>Metalearning</source>
          . Springer Berlin Heidelberg (
          <year>2009</year>
          ). https://doi.org/10.1007/978-3-
          <fpage>540</fpage>
          -73263-1
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>Smieja</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hajto</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Tabor</surname>
            ,
            <given-names>J.:</given-names>
          </string-name>
          <article-title>E cient mixture model for clustering of sparse high dimensional binary data</article-title>
          .
          <source>Data Mining and Knowledge Discovery</source>
          .
          <volume>33</volume>
          ,
          <issue>1583</issue>
          {
          <fpage>1624</fpage>
          (
          <year>2019</year>
          ). https://doi.org/10.1007/s10618-019-00635-1
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <surname>Campello</surname>
            ,
            <given-names>R.J.G.B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Moulavi</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sander</surname>
          </string-name>
          , J.:
          <source>Density-Based Clustering Based on Hierarchical Density Estimates. In: Advances in Knowledge Discovery and Data Mining</source>
          . pp.
          <volume>160</volume>
          {
          <fpage>172</fpage>
          . Springer Berlin Heidelberg (
          <year>2013</year>
          ). https://doi.org/10.1007/978-3-
          <fpage>642</fpage>
          -37456-2 14
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <surname>Surprise</surname>
          </string-name>
          {
          <article-title>A Python skicit for recommender systems</article-title>
          . http://surpriselib.com/
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <string-name>
            <surname>Harper</surname>
            ,
            <given-names>F.M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Konstan</surname>
            ,
            <given-names>J.A.</given-names>
          </string-name>
          :
          <article-title>The MovieLens Datasets</article-title>
          .
          <source>ACM Transactions on Interactive Intelligent Systems. 5</source>
          ,
          <issue>1</issue>
          {
          <fpage>19</fpage>
          (
          <year>2015</year>
          ). https://doi.org/10.1145/2827872
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          12.
          <string-name>
            <surname>Amazon Cell Phones Reviews</surname>
          </string-name>
          | Kaggle - https://www.kaggle.com/grikomsn/amazoncell-phones-reviews
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          13.
          <string-name>
            <surname>Misra</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Wan</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>McAuley</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          :
          <article-title>Decomposing t semantics for product size recommendation in metric spaces</article-title>
          .
          <source>In: Proceedings of the 12th ACM Conference on Recommender Systems - RecSys '18</source>
          . ACM Press (
          <year>2018</year>
          ). https://doi.org/10.1145/3240323.3240398
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>