<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <journal-title-group>
        <journal-title>June</journal-title>
      </journal-title-group>
    </journal-meta>
    <article-meta>
      <title-group>
        <article-title>Stream Recommendation using Individual Hyper-Parameters</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Bruno Veloso∗</string-name>
          <email>bruno.m.veloso@inesctec.pt</email>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Benedita Malheiro</string-name>
          <email>mbm@isep.ipp.pt</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Jeremy D. Foss</string-name>
          <email>Jeremy.Foss@bcu.ac.uk</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Stream Content Personalisation, Collaborative Filtering, Hyper-</string-name>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>DMT Lab - Birmingham City, University</institution>
          ,
          <addr-line>Birmingham</addr-line>
          ,
          <country country="UK">United Kingdom</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>ISEP/IPP - School of Engineering, Polytechnic of Porto</institution>
          ,
          <addr-line>Porto, Portugal, INESC TEC, Porto</addr-line>
          ,
          <country country="PT">Portugal</country>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>Parameter Optimisation</institution>
        </aff>
        <aff id="aff3">
          <label>3</label>
          <institution>REMIT - Research on Economics, Management and Information, Technologies, Portucalense University</institution>
          ,
          <addr-line>Porto, Portugal, INESC TEC, Porto</addr-line>
          ,
          <country country="PT">Portugal</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2019</year>
      </pub-date>
      <volume>05</volume>
      <issue>2019</issue>
      <abstract>
        <p>Nowadays, with the widely usage of on-line stream video platforms, the number of media resources available and the volume of crowd-sourced feedback volunteered by viewers is increasing exponentially. In this scenario, the adoption of recommendation systems allows platforms to match viewers with resources. However, due to the sheer size of the data and the pace of the arriving data, there is the need to adopt stream mining algorithms to build and maintain models of the viewer preferences as well as to make timely personalised recommendations. In this paper, we propose the adoption of optimal individual hyper-parameters to build more accurate dynamic viewer models. First, we use a grid search algorithm to identify the optimal individual hyper-parameters (IHP) and, then, use these hyper-parameters to update incrementally the user model. This technique is based on an incremental learning algorithm designed for stream data. The results show that our approach outperforms previous approaches, reducing substantially the prediction errors and, thus, increasing the accuracy of the recommendations.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>CCS CONCEPTS</title>
      <p>• Information systems → Data streams; Collaborative
filtering; • Computing methodologies → Online learning settings.</p>
    </sec>
    <sec id="sec-2">
      <title>INTRODUCTION</title>
      <p>
        The number of multimedia sources, resources and volume of
feedback data – ratings, likes, shares and posts/reviews – available
on-line are a challenge to standard recommendation algorithms
and makes real time processing almost impossible. This problem
requires the development of dedicated tools, involving profiling and
recommendation, to provide viewers with suggestions matching
their preferences in near real time. The user generated data, which
corresponds to explicit preferences and intrinsic behaviours, can
be used for extracting patterns and define viewer profiles. Dynamic
∗All authors contributed equally to this research.
viewer profiling, i.e., the ability to build and maintain profiles based
on the continuous stream of viewer interactions (likes, posts,
ratings, watched items, etc.), can be addressed as stream mining Gama
[
        <xref ref-type="bibr" rid="ref6">6</xref>
        ].
      </p>
      <p>From the several data sets available on-line, we chose
MovieLens 100k (ML 100k), MovieLens 1M (ML 1M), CiaoDVD, Jester,
FilmTrust and EachMovie. The data sets were prepared and
partitioned for the application of stream mining techniques. The initial
model and the updating process is achieved using the Stochastic
Gradient Descent (SGD) with the individual hyper-parameters. The
model is updated every time a viewer rates a resource, i.e., updates
the viewer latent matrix using the SGD. This incremental
methodology is validated by calculating, for the active viewer, the Root Mean
Square Error (RMSE), the Recall@10 and TRecall@10 between the
predictions generated by the model and the actual viewer ratings.</p>
      <p>
        In this paper, we propose the adoption of incremental matrix
factorisation algorithm together with individual hyper-parameters. We
determine the individual optimal learning rate and regularisation
parameters using: (i) of-line holdout training and validation to find
and evaluate the hyper-parameters; and (ii) on-line predictive
sequential data stream evaluation protocol to verify the performance
of the proposed data stream processing approach. This
methodology, when compared with the pre-existing approaches referred in
Section 2, leads to predictions with increased accuracy. This work
extends the work of [
        <xref ref-type="bibr" rid="ref24">24</xref>
        ] with an enriched related work discussion
and evaluation, including new data sets, statistical analyses and
prequential evaluation.
      </p>
      <p>In terms of organisation, this document contains five sections.
Section 2 is dedicated to the optimisation algorithms and on-line
recommendation. Section 3 describes our approach, including the
grid search and the recommendation algorithm. Section 4 describes
the experiments and discusses the results obtained. Finally, Section
5 draws the conclusions and suggests future developments.
2</p>
    </sec>
    <sec id="sec-3">
      <title>RELATED WORK</title>
      <p>
        Nowadays, recommendation algorithms are widely used in diferent
domains to suggest items or products, taking in consideration the
behaviour and preferences of the active user. There are several
diferent techniques to generate personalised recommendations, namely,
content-based and collaborative filtering algorithms. We focus on
collaborative filters, more specifically, on incremental model-based
algorithms. These algorithms rely on hyper-parameters, which
must be tuned according to the data used.
Regarding hyper-parameter optimisation algorithms, there are
multiple solutions available in the literature. These solutions can
be grouped in five categories: ( i) the gradient descent algorithms,
which detect the down-hills of the objective function to find local
optima Vorontsov et al. [
        <xref ref-type="bibr" rid="ref27">27</xref>
        ] and Shamir and Zhang [
        <xref ref-type="bibr" rid="ref19">19</xref>
        ]; (ii) the
particle swarm algorithms, which identify a set of candidate solutions
called particles Poli et al. [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ] and Kennedy [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ]; (iii) the grid search
algorithms, which perform a sequential search with progressive
increments of the target hyper-parameters Coope and Price [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ]
and LaValle et al. [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ]; (iv) random search algorithms, which try
with random values until they satisfy a convergence criteria Price
[
        <xref ref-type="bibr" rid="ref16">16</xref>
        ] and Bergstra and Bengio [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]; and (v) direct search algorithms,
which apply heuristics to converge to an optimal solution Storn
and Price [
        <xref ref-type="bibr" rid="ref21">21</xref>
        ] and Audet and Dennis Jr [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ].
      </p>
      <p>
        Vorontsov et al. [
        <xref ref-type="bibr" rid="ref27">27</xref>
        ] proposed and applied the gradient descent
algorithm, an approach identical to the stochastic parallel
gradient descent, to minimise image perturbations. The algorithm was
designed to work in parallel to be applied to the multiple control
channels provided by camera. Shamir and Zhang [
        <xref ref-type="bibr" rid="ref19">19</xref>
        ] presented an
extensive study on the SGD convergence properties by applying
averaging schemes on the SGD iterations to obtain optimal
performances. The authors found that the polynomial-decay averaging
improves the results with the same complexity in computational
terms as the standard average.
      </p>
      <p>
        Poli et al. [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ] and Kennedy [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ] performed an extended overview
of particle swarm optimisation algorithms, including the adoption
of nearest neighbours’ velocity and craziness, cornfield vectors, the
elimination of ancillary variables and distance acceleration.
      </p>
      <p>
        Coope and Price [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ] used grid search to follow a given direction
of the gradient descent to reduce the search space. This suggestion
tries to minimise the number of search iterations required to
minimise the objective function, focusing the grid on locations where
the gradient descent is higher. LaValle et al. [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ] adopted
probabilistic road maps which share similarities with grid search algorithms.
The authors proposed low discrepancy and low dispersion samples
to evaluate the objective function.
      </p>
      <p>
        Price [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ] proposed a controlled random search method to
optimise parameters. The constrains are applied at variable level,
specifying for each case the upper and lower bounds. This solution
is more efective and less computational expensive than the typical
random search. Bergstra and Bengio [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ] presented random search
as a substitute of common optimisation algorithms such as manual
or grid search. Grid search algorithms perform an exponential
number of irrelevant search trials if they have insuficient resolution to
identify a local minimum. The manual search requires experts with
confidence and experience to tune the parameters.
      </p>
      <p>
        Storn and Price [
        <xref ref-type="bibr" rid="ref21">21</xref>
        ] described a diferential evolution direct
search method. This method is composed of three operations: (i)
mutation, where a vector is generated with diferent indexes; ( ii)
crossover, which is designed to increase the diversity; and (iii)
selection of the generated and target vectors. Audet and Dennis Jr
[
        <xref ref-type="bibr" rid="ref1">1</xref>
        ] proposed a mesh adaptive direct search, which is an extension
of the pattern search algorithm. The main diference is that the
proposed algorithm is not constrained to a finite number of directions.
The proposed algorithm outperforms the pattern algorithm when
the number of function evaluations is limited to a fixed number.
      </p>
      <p>
        In terms of incremental collaborative filtering, we found several
proposals in the literature, including contributions designed to
minimise problems aflicting collaborative recommendation. It is
the case of: (i) the addition of new entities, e.g., using folding-in
techniques Sarwar et al. [
        <xref ref-type="bibr" rid="ref17">17</xref>
        ]; (ii) the minimisation of the cold-start
efect, e.g., using linear regression Sedhain et al. [
        <xref ref-type="bibr" rid="ref18">18</xref>
        ]; (iii) the slow
decrease of importance of items with time, e.g., using rating or latent
factor fading Vinagre et al. [
        <xref ref-type="bibr" rid="ref26">26</xref>
        ]; (iv) the identification of outliers
Gouvert et al. [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ]; (v) the exploration of additional information, e.g.,
using extension and sketching techniques of the latent matrices
Song et al. [
        <xref ref-type="bibr" rid="ref20">20</xref>
        ]; and (vi) the imputation of missing data, e.g., based
on the popularity of the items He et al. [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ] or on explicit user
comments Manotumruksa et al. [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ] and Zhang et al. [
        <xref ref-type="bibr" rid="ref30">30</xref>
        ].
      </p>
      <p>
        Sarwar et al. [
        <xref ref-type="bibr" rid="ref17">17</xref>
        ] adopted the folding-in technique to add new
entities to the singular-value decomposition (SVD) model. This
technique is less computational expensive than to recompute the
SVD every time a new entity arrives. However, the authors do not
present a mechanism to update the latent matrices.
      </p>
      <p>
        Xiang and Yang [
        <xref ref-type="bibr" rid="ref29">29</xref>
        ] proposed a factorised model with four
stages: time bias and user bias shifting, item bias shifting and user
preference shifting. First, they build the model, using SVD
matrix factorisation, and apply the four bias efects (user, item, time
and user preference) to optimise the factorised model. Then, they
use SGD to update the model and learn over time, determining a
global learning rate and over-fitting parameter for all users. Finally,
they evaluate the results with RMSE. Takács et al. [
        <xref ref-type="bibr" rid="ref22">22</xref>
        ] proposed
several matrix factorisation approaches as well as neighbour
selection methodologies for matrix factorisation models. The authors
use bias efects and weights to optimise the factorised model and
adopt a global learning rate and over-fitting parameter for all users.
Gower [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ] explored diferent on-line recommender models using
SVD. The author applied learning algorithms to enhance the model
with time-based biases. This approach adopts two bias efects (user
and item) to optimise the factorised model and calculates a global
learning rate and over-fitting parameter (for all users). Vinagre et al .
[
        <xref ref-type="bibr" rid="ref25">25</xref>
        ] introduced an incremental matrix factorisation algorithm for
positive-only feedback and proposed a new evaluation
methodology called prequential protocol. The matrix factorisation and the
learning techniques are SVD and SGD, respectively. The
prequential protocol verifies, every time a new rating event occurs, if the
rated item would have been recommended to that viewer and, if
afirmative, counts as a hit. In 2015, the same authors included a
rating-and-recency-based scheme to perform negative preference
imputation [
        <xref ref-type="bibr" rid="ref26">26</xref>
        ]. This scheme creates and maintains a global item
queue of size n where the top-rated items, i.e., items rated with the
maximum rating, are kept at the head and all other items slide to
the tail till they are, eventually, removed from the queue. Every
time a new item is rated, it is inserted in the queue depending on
its rating: at the top, if it was rated with the maximum rating, or
immediately after the top-rated items, otherwise.
      </p>
      <p>
        Song et al. [
        <xref ref-type="bibr" rid="ref20">20</xref>
        ] presented an alternative updating method to
be applied to factorised matrices. This method has three steps: (i)
extension of the matrix with auxiliary features; (ii) application of a
sketching technique to compress the feature matrix; and (iii) update
of the latent elements based on the diference between the new
and old models or between the new and old prediction errors. He
et al. [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ] exploited popularity to reduce the matrix sparsity and
initialise missing item data. The authors presented the element-wise
Alternating Least Squares technique to build a matrix factorisation
model with this variable-weighted missing item data. Additionally,
the authors presented a refined version of the objective function
which classifies as true negative all the initialised items.
      </p>
      <p>
        Manotumruksa et al. [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ] present a solution combining matrix
factorisation with word embedding to enhance the final predictions.
The user textual reviews are processed to extract semantic
information and are put into latent matrices. These semantic latent matrices
are combined with the matrix factorisation model to improve the
recommendations. Sedhain et al. [
        <xref ref-type="bibr" rid="ref18">18</xref>
        ] propose a model which
minimises the cold-start efect. The system called LoCo implements:
(i) a multivariate linear regression to learn the weights that the
social components have on the user preferences; (ii) a low-rank
parameterisation of the linear model weights, to reduce the
multitude of social parameters; and (iii) a randomised SVD to project
the metadata and make predictions. Gouvert et al. [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ] describe a
solution to deal with outliers. The proposal extends the Poisson
matrix factorisation, taking into consideration the model
exposure, by multiplying the model by a hyper-parameter. The results
show improved performance with dispersed data sets. Zhang et al.
[
        <xref ref-type="bibr" rid="ref30">30</xref>
        ] describe a hybrid matrix factorisation recommender system.
It extracts the user implicit emotional feedback from user reviews,
using sentiment analysis, and, then, uses this information to revise
the previously rated service. Additionally, the authors extract and
incorporate user and service features into a conventional matrix
factorisation algorithm to increase the accuracy of the predictions.
      </p>
      <p>
        Our approach, alternatively, determines optimal individual
learning rates and regularisation parameters with the help of grid search.
The grid search optimisation algorithm is: (i) easy to implement
and parallelise; and (ii) very reliable in low-dimension search space
according to Bergstra and Bengio [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]. Regarding the
recommendation algorithm, we extend the work proposed by Takács et al.
[
        <xref ref-type="bibr" rid="ref22">22</xref>
        ] with the use of individual hyper parameters to maximise the
accuracy of the model. The original algorithm implements matrix
factorisation to generate the user and item representative latent
matrices together with SGD to update the latent factors. This
technique has two model parameters: the learning rate to control the
speed of the learning process, and the regulator parameter to avoid
over-fitting. Finally, the predictions are computed through the dot
product between the corresponding latent vectors of each latent
matrix. The specific version adopted in this work corresponds to the
biased regularised incremental simultaneous matrix factorisation
[
        <xref ref-type="bibr" rid="ref22">22</xref>
        ]. Our extension explores the usage of individual parameters for
each user based on the assumption that each user has a diferent
learning behaviour.
3
      </p>
    </sec>
    <sec id="sec-4">
      <title>PROPOSED METHOD</title>
      <p>Our proposal performs a grid search algorithm to select the best
individual hyper-parameters. We implemented the biased regularised
incremental simultaneous matrix factorisation to update
sequentially the model. The search fulfils two conditions: ( i) the global
output metrics with the found hyper-parameters improve the
baseline results; and (ii) the optimal individual hyper-parameters are
selected in terms of RMSE. Algorithm 1 implements the individual
hyper-parameter (IHP) optimisation process, using the initial 50 %
of the data (train partition). It determines the optimal learning rate
(η) and regularisation parameter (λ) for each user in terms of RMSE.
These individual parameters are then used to update the user latent
matrix. For each new stream rating (line 3), it adds the new rating
to the rating matrix (line 4), where ru,i represents the rating given
by a user u to item i, and calculates the individual and global RMSE
between the prediction (rˆu,i ) and the real rating (lines 5-10). Next,
it updates the viewer latent matrix (the viewer row) (line 11), using
the hyper-parameters found by the grid search and the calculated
rating error (line 6), where p®u is the latent user vector and q®i is the
latent item vector. Finally, if both global and individual RMSE have
decreased (lines 12 and 13), it updates the user’s hyper-parameters
(lines 14-15). Line 16 increments the total number of events.
Algorithm 1 Optimisation of the Individual Hyper-Parameters
10:</p>
      <p>Algorithm 2 illustrates the on-line model update with stream
learning, using the remaining 50 % of the data. For each stream
rating (line 1), it adds the rating to the rating matrix R (line 2),
calculates the RMSE between the predicted and the real rating
(lines 3-6). Finally, it updates the user latent matrix (the user row),
using the optimal individual hyper-parameters and the calculated
rating error (line 7). The variables represent the real rating (ru,i ), the
prediction (rˆu,i ), the latent user vector (p®u ), the latent item vector
(q®i ), the user learning rate (ηu ), the user regularisation parameter
(λu ), the accumulated error (e), the prediction error for the user u
and item i (eu,i ), the root mean square error (rmse) and the total
number of events (n).</p>
      <p>Algorithm 2 Stream learning algorithm
1: for ru,i ← Stream do
2: R ← addN ewRatinд(ru,i )
3: rˆu,i = q®i · p®u
4: eu,i = (ru,i − rˆu,i )2
5: e = e + eu,i
6: rmse = q ne
7: p®u ← p®u + ηu (eu,i q®i − λup®u )
4</p>
    </sec>
    <sec id="sec-5">
      <title>EXPERIMENTS AND RESULTS</title>
      <p>The following subsections present the data sets, the evaluation
metrics together with the protocol, the experiments, the results
obtained and a final discussion. The experiments were performed
with an Intel Xeon CPU E5-2680 2.40 GHz Central Processing Unit
(CPU), 32 GiB DDR3 Random Access Memory (RAM) and 1 TiB of
hard drive platform running the Ubuntu 16.04.
4.1</p>
    </sec>
    <sec id="sec-6">
      <title>Data Sets</title>
      <p>Our proposal was evaluated with six diferent data sets: ( i)
MovieLens 100k (ML 100k) contains information about 943 users and 1682
movies, including 100 000 user ratings together with time stamps
and has a data sparsity of 93.70 %; (ii) MovieLens 1M (ML 1M) has
higher data sparsity 95.5 % and holds information about 6040 users
and 3952 movies, including 1 000 000 user ratings together with
time stamps; (iii) CiaoDVD has a data sparsity of 99.97 % and
contains information about 17 615 users and 16 121 movies, including
72 665 user ratings together with time stamps; (iv) FilmTrust has
a data sparsity of 98.86 % and holds information about 1508 users
and 2071 movies, including 35 497 user ratings together with time
stamps; (v) Jester has a data sparsity of 80.83 % and contains
information about 59 132 users and 150 items, including 1 700 000 user
ratings; (vi) EachMovie data set has a data sparsity of 97.63 % and
holds information about 72 916 users and 1628 movies, including
2 811 983 user ratings together with time stamps.</p>
      <p>Our selection of the data sets was based, first, on the type of the
data, i.e., we chose media-related data sets, and, next on the data
sizes, i.e., with diverse sizes. Diferent data sizes allow us to verify
if the grid search algorithm finds hyper-parameter values which,
in fact, improve the recommendation results.
4.2</p>
    </sec>
    <sec id="sec-7">
      <title>Evaluation Metrics</title>
      <p>
        Regarding the evaluation metrics, we calculate the incremental
error prediction measure (RMSE), which is calculated after each
new rating event Takács et al. [
        <xref ref-type="bibr" rid="ref22">22</xref>
        ]. In terms of accuracy metrics,
we adopt the standard recall metric used by Cremonesi et al. [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ],
which considers a subset of the top-rated items and additionally we
adopt another recall-based metric called – Target Recall (TRecall)
– used by Veloso et al. [
        <xref ref-type="bibr" rid="ref23">23</xref>
        ] which contemplates a subset of items
centred around the target rating. For each new rating event, we
determine the Recall@N and the TRecall@N. First, we predict the
ratings of all items unseen by the viewer, including the newly rated
item, and then we select 1000 unrated items plus the new rated
item and sort them in descending order. Finally, if the newly rated
item belongs to the list of the top N viewer predicted items centred
on the actual viewer rating, we count a hit for the TRecall and if
the item belongs to the top N items we count a hit for the Recall.
4.3
      </p>
    </sec>
    <sec id="sec-8">
      <title>Evaluation Protocol</title>
      <p>The evaluation protocol defines the data ordering, partitions and
distribution. In this work we adopt the holdout and the prequential
protocol. The data was ordered temporally and, then, partitioned
according to the selected protocol.</p>
      <p>
        In the holdout protocol the “Train” subset is used to build the
initial model and to optimise the hyper-parameters, whereas the
remaining data or “Test” subset is used for the validation. Since
in a real streaming environment it is impossible to anticipate the
number events of each user, the training process is initiated with
a predefined number of events and generic hyper-parameters. To
guarantee that there are suficient events to train each individual
hyper-parameter, we ordered the data temporally and ensure that
the train phase holds 50 % of the ratings, leaving the remaining 50 %
of the ratings for validation. We adopt the holdout protocol, which
is a simple version of the cross validation, to determine the general
performance of the model. The main drawback of holdout when
compared with cross-validation is that it presents higher variance
Blum et al. [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ].
      </p>
      <p>
        In the case of the prequential protocol proposed by Gama et al.
[
        <xref ref-type="bibr" rid="ref7">7</xref>
        ], all the data it is used to build, maintain and assess the algorithm
performance. The hyper-parameters used in this step are the same
used with holdout. Each one of data ratings triggers the generation
and immediate evaluation of the predictions as well as the model
updating in both evaluation protocols. We choose the
prequential protocol because it is the unique approach that measures the
learning process. This method uses the predictive sequential error
estimated over a sliding window to verify the model reaction to
non-stationary environments.
4.4
      </p>
    </sec>
    <sec id="sec-9">
      <title>Significance Tests</title>
      <p>
        To detect the statistical diferences between the proposed and the
baseline approaches we applied two diferent significance tests: ( i)
the Wilcoxon test Wilcoxon [
        <xref ref-type="bibr" rid="ref28">28</xref>
        ] to verify if the mean ranks of two
samples difer; and ( ii) the McNemar test McNemar [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ] to assess if a
statistically significant change occurs on a dichotomous trait at two
time points on the same population. We define a 5 % of significance
level for all tests. The goal of the Wilcoxon and McNemar tests is
to reject the null-hypothesis, i.e., that both approaches have the
same performance. For a significance level ( p) of 0.05, the value of
McNemar test (M) is 3.84 and the value of the Wilcoxon test (W ) is
137.
4.5
      </p>
    </sec>
    <sec id="sec-10">
      <title>Experiments</title>
      <p>
        To assess the proposed solution, the algorithm executed thirty times
with each data set to compute the mean and the standard deviation
of the performance metrics. The baseline algorithm corresponds
to the biased regularised incremental simultaneous matrix
factorisation algorithm proposed by Takács et al. [
        <xref ref-type="bibr" rid="ref22">22</xref>
        ] with global
hyperparameters. Table 2 presents the results of the holdout evaluation
protocol with the best results of each data set highlighted in bold.
In the case of ML 100k, the RMSE decreases 2.0 % and the Recall and
TRecall increases 79.8 % and 30.0 %, respectively. The results for the
ML 1M show that the RMSE decreases 49.2 %, whereas the Recall
and TRecall increase 128.4 % and 126.8 %, respectively. For the Jester
data set, the RMSE increases 3.5 % and Recall and TRecall improve
3.0 % and 2.5 %. In the case of FilmTrust, the RMSE increases 22.8 %
and the Recall and TRecall decreases 73.6 % and 62.7 %, respectively.
For the EachMovie data set, the RMSE decreases 34.3 % and the
Recall and TRecall improve 214.5 % and 109.6 %. Finally, for the
CiaoDVD, the RMSE inflates 3.7 % and Recall and TRecall improve
71.3 % and 35.4 %.
      </p>
      <p>The holdout evaluation shows that IHP improves the
Recallbased results of five data sets (the exception is FilmTrust) and
decreases the prediction errors of three data sets (EachMovie, ML 1M
and ML 100k), including two of the largest data sets. The statistical
results for the Wilcoxon and McNemar tests, displayed in Table 3,
reject the null hypothesis for all data sets.</p>
      <p>Figure 1 presents the prequential results for all data sets with a
sliding window of 1000 events. The plots represent the logarithmic
ratio between the predictive sequential error of the baseline (with
global hyper-parameters) and the IHP algorithms. Positive values
mean that IHP is more accurate than the baseline algorithm. We can
observe with all data sets that IHP decreases significantly the
prediction errors in the streaming scenario. Additionally, the prequential
evaluation plots display the behaviour of the learning process over
time with non-monotonic data streams. EachMovie, the larger data
set with more than 2.8 million events and a data sparsity of 97.63 %,
shows a continuous decrease in performance, indicating unstable
user behaviour over time. This behaviour exposes a concept drift
problem, which is dificult to detect in recommendation systems.
Jester, the second larger data set data set with 1.7 million events
and a sparsity of 80.83 %, shows no concept drift signs, suggesting
that its lower sparsity contributes to a stabler performance.
The proposed algorithm is designed to optimise the individual
hyper-parameters of model-based collaborative filters by
minimising the RMSE. However, to find the best individual hyper-parameters,
it requires a significant number of rating events per user. This
limitation is observable in the holdout evaluation (Table 3) where the
three data sets with lower number of average events per user (Jester,
FilmTrust and CiaoDVD in Table 1 perform worst with individual
hyper-parameters than with the baseline approach. Specifically, the
Jester, FilmTrust and CiaoDVD data sets have less than 30
average events per user and the minimal number of events per user
is 1. Since we use 50 % of the user data to select the individual
hyper-parameters, this means that, in the case of these data sets,
the average number of events available to perform this task is less
than 15. In terms of Recall, our algorithm surpasses the baseline
with all data sets except with the smallest data set (FilmTrust),
showing considerable immunity to the problem of low number of
training events. The holdout experiments were repeated 30 times
to compute the average and standard deviation of the metrics used.</p>
      <p>The prequential evaluation protocol assesses the behaviour of
the algorithm in streaming conditions. Figure 1 presents the
diference between the baseline and IHP algorithms. The results show
that IHP reduces the predictive errors with all data sets. This is
due to the implemented search mechanism which, firstly, tries to
optimise the global predictive error and, then, minimise the
individual prediction errors. Finally, we made a set of statistical tests
to verify if our algorithm is statistically diferent from the baseline
algorithm. According to Table 3, we can reject the null hypothesis.</p>
      <p>We can state, based on the obtained results, that the proposed
algorithm represents the user behaviour more accurately, modelling
users with distinct behaviours with diferent individual learning
rate and regularisation parameters. Moreover, it requires an
average of 23 events per user for the convergence of the two individual
hyper-parameters and to outperform the baseline results with both
evaluation protocols (holdout and prequential). Concept drift is
expected to occur with large (size), long (time) and sparse data sets
and, naturally, with data streams. To overcome this problem, it is
necessary to readjust the individual hyper-parameters. However,
concept drift detection is not a trivial task for recommendation
systems since it is very dificult to detect significant data variations
when user feedback can be highly variable, imitating the behaviour
of outliers. A possible solution is to recalculate regularly the
individual hyper-parameters. Regarding the state of the art algorithms,
our algorithm is, as far as we know, the first to adopt independent
user hyper-parameters to improve the final recommendations.</p>
    </sec>
    <sec id="sec-11">
      <title>5 CONCLUSIONS</title>
      <p>This paper describes an individual learning algorithm for updating
user models. This algorithm refines existing stream mining
techniques by building and updating individual user models based on
the user stream of events. The goal of this research is to improve
real time user profiling by considering the continuous stream of
on-line user events. The algorithm uses these events to learn the
preferences of each user, which are subject to external influences
as well as to personal evolution of interests, as they occur in time.</p>
      <p>
        Our approach applies the individual learning rate and
regularisation parameters to build and update the individual user profiles,
while keeping the prediction model isolated from the user event
streams. Regarding the on-line incremental matrix factorisation
algorithm of Vinagre et al. [
        <xref ref-type="bibr" rid="ref25">25</xref>
        ], our algorithm displays for all data
sets improved recall results; in terms of RMSE, our approach has
better results with the data sets with an average of at least 45 events
per user.
      </p>
      <p>As future work, concerning the incremental learning algorithm,
we plan to explore forgetting strategies so that old and less relevant
events slowly fade into oblivion, making the user profile more
realistic and accurate. Further exploration it is necessary to adjust
dynamically the individual hyper-parameters throughout time. The
decision to recompute the individual hyper-parameters can be based
on the detection of concept drifts since our results indicate that
larger data sets display concept drift efects. Finally, to minimise
the number of computational resources necessary, older inactive
users can be deactivated.</p>
    </sec>
    <sec id="sec-12">
      <title>ACKNOWLEDGMENTS</title>
      <p>This work was partially financed by National Funds through the
Portuguese funding agency, FCT – Fundação para a Ciência e
Tecnologia, within project UID/EEA/50014/2019.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>Charles</given-names>
            <surname>Audet and John E Dennis Jr</surname>
          </string-name>
          .
          <year>2006</year>
          .
          <article-title>Mesh adaptive direct search algorithms for constrained optimization</article-title>
          .
          <source>SIAM Journal on optimization 17</source>
          ,
          <issue>1</issue>
          (
          <year>2006</year>
          ),
          <fpage>188</fpage>
          -
          <lpage>217</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>James</given-names>
            <surname>Bergstra</surname>
          </string-name>
          and
          <string-name>
            <given-names>Yoshua</given-names>
            <surname>Bengio</surname>
          </string-name>
          .
          <year>2012</year>
          .
          <article-title>Random search for hyper-parameter optimization</article-title>
          .
          <source>Journal of Machine Learning Research</source>
          <volume>13</volume>
          ,
          <string-name>
            <surname>Feb</surname>
          </string-name>
          (
          <year>2012</year>
          ),
          <fpage>281</fpage>
          -
          <lpage>305</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>Avrim</given-names>
            <surname>Blum</surname>
          </string-name>
          , Adam Kalai,
          <string-name>
            <given-names>and John</given-names>
            <surname>Langford</surname>
          </string-name>
          .
          <year>1999</year>
          .
          <article-title>Beating the hold-out: Bounds for k-fold and progressive cross-validation</article-title>
          .
          <source>In COLT</source>
          , Vol.
          <volume>99</volume>
          . ACM, New York, NY, USA,
          <fpage>203</fpage>
          -
          <lpage>208</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <surname>Ian</surname>
            <given-names>D</given-names>
          </string-name>
          <string-name>
            <surname>Coope</surname>
            and
            <given-names>Christopher John Price.</given-names>
          </string-name>
          <year>2001</year>
          .
          <article-title>On the convergence of grid-based methods for unconstrained optimization</article-title>
          .
          <source>SIAM Journal on Optimization 11</source>
          ,
          <issue>4</issue>
          (
          <year>2001</year>
          ),
          <fpage>859</fpage>
          -
          <lpage>869</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>Paolo</given-names>
            <surname>Cremonesi</surname>
          </string-name>
          , Yehuda Koren, and
          <string-name>
            <given-names>Roberto</given-names>
            <surname>Turrin</surname>
          </string-name>
          .
          <year>2010</year>
          .
          <article-title>Performance of recommender algorithms on top-n recommendation tasks</article-title>
          .
          <source>In Proceedings of the fourth ACM conference on Recommender systems. ACM</source>
          , New York, NY, USA,
          <fpage>39</fpage>
          -
          <lpage>46</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>Joao</given-names>
            <surname>Gama</surname>
          </string-name>
          .
          <year>2010</year>
          .
          <article-title>Knowledge discovery from data streams</article-title>
          .
          <source>Chapman</source>
          and Hall/CRC, Boca Raton, FL, USA.
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>João</given-names>
            <surname>Gama</surname>
          </string-name>
          , Raquel Sebastião, and Pedro Pereira Rodrigues.
          <year>2009</year>
          .
          <article-title>Issues in evaluation of stream learning algorithms</article-title>
          .
          <source>In Proceedings of the 15th ACM SIGKDD international conference on Knowledge discovery and data mining. ACM</source>
          , New York, NY, USA,
          <fpage>329</fpage>
          -
          <lpage>338</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>Olivier</given-names>
            <surname>Gouvert</surname>
          </string-name>
          , Thomas Oberlin, and
          <string-name>
            <given-names>Cédric</given-names>
            <surname>Févotte</surname>
          </string-name>
          .
          <year>2018</year>
          .
          <article-title>Negative Binomial Matrix Factorization for Recommender Systems</article-title>
          . arXiv preprint arXiv:
          <year>1801</year>
          .
          <volume>01708</volume>
          (
          <year>2018</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>Stephen</given-names>
            <surname>Gower</surname>
          </string-name>
          .
          <year>2014</year>
          .
          <article-title>Netflix prize</article-title>
          and SVD.
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <surname>Xiangnan</surname>
            <given-names>He</given-names>
          </string-name>
          , Hanwang Zhang,
          <string-name>
            <surname>Min-Yen Kan</surname>
          </string-name>
          , and
          <string-name>
            <surname>Tat-Seng Chua</surname>
          </string-name>
          .
          <year>2016</year>
          .
          <article-title>Fast matrix factorization for online recommendation with implicit feedback</article-title>
          .
          <source>In Proceedings of the 39th International ACM SIGIR conference on Research and Development in Information Retrieval. ACM</source>
          , New York, NY, USA,
          <fpage>549</fpage>
          -
          <lpage>558</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>James</given-names>
            <surname>Kennedy</surname>
          </string-name>
          .
          <year>2010</year>
          .
          <article-title>Particle swarm optimization</article-title>
          .
          <source>Encyclopedia of machine learning</source>
          (
          <year>2010</year>
          ),
          <fpage>760</fpage>
          -
          <lpage>766</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <surname>Steven</surname>
            <given-names>M LaValle</given-names>
          </string-name>
          , Michael S Branicky, and
          <string-name>
            <surname>Stephen R Lindemann</surname>
          </string-name>
          .
          <year>2004</year>
          .
          <article-title>On the relationship between classical grid search and probabilistic roadmaps</article-title>
          .
          <source>The International Journal of Robotics Research</source>
          <volume>23</volume>
          ,
          <fpage>7</fpage>
          -
          <lpage>8</lpage>
          (
          <year>2004</year>
          ),
          <fpage>673</fpage>
          -
          <lpage>692</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <surname>Jarana</surname>
            <given-names>Manotumruksa</given-names>
          </string-name>
          , Craig Macdonald, and
          <string-name>
            <given-names>Iadh</given-names>
            <surname>Ounis</surname>
          </string-name>
          .
          <year>2017</year>
          .
          <article-title>Matrix factorisation with word embeddings for rating prediction on location-based social networks</article-title>
          .
          <source>In European Conference on Information Retrieval</source>
          . Springer, Cham, Switzerland,
          <fpage>647</fpage>
          -
          <lpage>654</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <given-names>Quinn</given-names>
            <surname>McNemar</surname>
          </string-name>
          .
          <year>1947</year>
          .
          <article-title>Note on the sampling error of the diference between correlated proportions or percentages</article-title>
          .
          <source>Psychometrika</source>
          <volume>12</volume>
          ,
          <issue>2</issue>
          (
          <year>1947</year>
          ),
          <fpage>153</fpage>
          -
          <lpage>157</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <surname>Riccardo</surname>
            <given-names>Poli</given-names>
          </string-name>
          , James Kennedy, and
          <string-name>
            <given-names>Tim</given-names>
            <surname>Blackwell</surname>
          </string-name>
          .
          <year>2007</year>
          .
          <article-title>Particle swarm optimization</article-title>
          .
          <source>Swarm intelligence 1</source>
          ,
          <issue>1</issue>
          (
          <year>2007</year>
          ),
          <fpage>33</fpage>
          -
          <lpage>57</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [16]
          <string-name>
            <given-names>WL</given-names>
            <surname>Price</surname>
          </string-name>
          .
          <year>1983</year>
          .
          <article-title>Global optimization by controlled random search</article-title>
          .
          <source>Journal of Optimization Theory and Applications</source>
          <volume>40</volume>
          ,
          <issue>3</issue>
          (
          <year>1983</year>
          ),
          <fpage>333</fpage>
          -
          <lpage>348</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [17]
          <string-name>
            <surname>Badrul</surname>
            <given-names>Sarwar</given-names>
          </string-name>
          , George Karypis, Joseph Konstan,
          <string-name>
            <given-names>and John</given-names>
            <surname>Riedl</surname>
          </string-name>
          .
          <year>2002</year>
          .
          <article-title>Incremental singular value decomposition algorithms for highly scalable recommender systems</article-title>
          .
          <source>In Fifth international conference on computer and information science</source>
          , Vol.
          <volume>27</volume>
          .
          <string-name>
            <surname>Citeseer</surname>
          </string-name>
          ,
          <volume>28</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          [18]
          <string-name>
            <surname>Suvash</surname>
            <given-names>Sedhain</given-names>
          </string-name>
          ,
          <string-name>
            <given-names>Aditya</given-names>
            <surname>Menon</surname>
          </string-name>
          , Scott Sanner, Lexing Xie, and
          <string-name>
            <given-names>Darius</given-names>
            <surname>Braziunas</surname>
          </string-name>
          .
          <year>2017</year>
          .
          <article-title>Low-Rank Linear Cold-Start Recommendation from Social Data</article-title>
          . AAAI, Palo Alto, CA, USA.
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          [19]
          <string-name>
            <given-names>Ohad</given-names>
            <surname>Shamir</surname>
          </string-name>
          and
          <string-name>
            <given-names>Tong</given-names>
            <surname>Zhang</surname>
          </string-name>
          .
          <year>2013</year>
          .
          <article-title>Stochastic gradient descent for non-smooth optimization: Convergence results and optimal averaging schemes</article-title>
          .
          <source>In International Conference on Machine Learning</source>
          .
          <fpage>71</fpage>
          -
          <lpage>79</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          [20]
          <string-name>
            <surname>Qiang</surname>
            <given-names>Song</given-names>
          </string-name>
          , Jian Cheng, and Hanqing Lu.
          <year>2015</year>
          .
          <article-title>Incremental matrix factorization via feature space re-learning for recommender system</article-title>
          .
          <source>In Proceedings of the 9th ACM Conference on Recommender Systems. ACM</source>
          , New York, NY, USA,
          <fpage>277</fpage>
          -
          <lpage>280</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          [21]
          <string-name>
            <given-names>Rainer</given-names>
            <surname>Storn</surname>
          </string-name>
          and
          <string-name>
            <given-names>Kenneth</given-names>
            <surname>Price</surname>
          </string-name>
          .
          <year>1997</year>
          .
          <article-title>Diferential evolution-a simple and eficient heuristic for global optimization over continuous spaces</article-title>
          .
          <source>Journal of global optimization 11</source>
          ,
          <issue>4</issue>
          (
          <year>1997</year>
          ),
          <fpage>341</fpage>
          -
          <lpage>359</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          [22]
          <string-name>
            <surname>Gábor</surname>
            <given-names>Takács</given-names>
          </string-name>
          , István Pilászy, Bottyán Németh, and
          <string-name>
            <given-names>Domonkos</given-names>
            <surname>Tikk</surname>
          </string-name>
          .
          <year>2009</year>
          .
          <article-title>Scalable collaborative filtering approaches for large recommender systems</article-title>
          .
          <source>Journal of machine learning research 10</source>
          ,
          <string-name>
            <surname>Mar</surname>
          </string-name>
          (
          <year>2009</year>
          ),
          <fpage>623</fpage>
          -
          <lpage>656</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          [23]
          <string-name>
            <surname>Bruno</surname>
            <given-names>Veloso</given-names>
          </string-name>
          , Benedita Malheiro, Juan Carlos Burguillo, and
          <string-name>
            <given-names>Jeremy</given-names>
            <surname>Foss</surname>
          </string-name>
          .
          <year>2017</year>
          .
          <article-title>Personalised fading for stream data</article-title>
          .
          <source>In Proceedings of the Symposium on Applied Computing. ACM</source>
          , New York, NY, USA,
          <fpage>870</fpage>
          -
          <lpage>872</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          [24]
          <string-name>
            <surname>Bruno</surname>
            <given-names>Veloso</given-names>
          </string-name>
          , Benedita Malheiro, Juan Carlos Burguillo, Jeremy Foss, and
          <string-name>
            <given-names>João</given-names>
            <surname>Gama</surname>
          </string-name>
          .
          <year>2018</year>
          .
          <article-title>Personalised Dynamic Viewer Profiling for Streamed Data</article-title>
          .
          <source>In World Conference on Information Systems and Technologies</source>
          . Springer, Cham, Switzerland,
          <fpage>501</fpage>
          -
          <lpage>510</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref25">
        <mixed-citation>
          [25]
          <string-name>
            <surname>João</surname>
            <given-names>Vinagre</given-names>
          </string-name>
          , Alípio Mário Jorge, and
          <string-name>
            <given-names>João</given-names>
            <surname>Gama</surname>
          </string-name>
          .
          <year>2014</year>
          .
          <article-title>Fast incremental matrix factorization for recommendation with positive-only feedback</article-title>
          .
          <source>In International Conference on User Modeling, Adaptation, and Personalization</source>
          . Springer, Cham, Switzerland,
          <fpage>459</fpage>
          -
          <lpage>470</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref26">
        <mixed-citation>
          [26]
          <string-name>
            <surname>João</surname>
            <given-names>Vinagre</given-names>
          </string-name>
          , Alípio Mário Jorge, and
          <string-name>
            <given-names>João</given-names>
            <surname>Gama</surname>
          </string-name>
          .
          <year>2015</year>
          .
          <article-title>Collaborative filtering with recency-based negative feedback</article-title>
          .
          <source>In Proceedings of the 30th Annual ACM Symposium on Applied Computing. ACM</source>
          , New York, NY, USA,
          <fpage>963</fpage>
          -
          <lpage>965</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref27">
        <mixed-citation>
          [27] MA Vorontsov,
          <source>GW Carhart, and JC Ricklin</source>
          .
          <year>1997</year>
          .
          <article-title>Adaptive phase-distortion correction based on parallel gradient-descent optimization</article-title>
          .
          <source>Optics letters 22</source>
          ,
          <issue>12</issue>
          (
          <year>1997</year>
          ),
          <fpage>907</fpage>
          -
          <lpage>909</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref28">
        <mixed-citation>
          [28]
          <string-name>
            <given-names>Frank</given-names>
            <surname>Wilcoxon</surname>
          </string-name>
          .
          <year>1945</year>
          .
          <article-title>Individual comparisons by ranking methods</article-title>
          .
          <source>Biometrics bulletin 1</source>
          ,
          <issue>6</issue>
          (
          <year>1945</year>
          ),
          <fpage>80</fpage>
          -
          <lpage>83</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref29">
        <mixed-citation>
          [29]
          <string-name>
            <given-names>Liang</given-names>
            <surname>Xiang</surname>
          </string-name>
          and
          <string-name>
            <given-names>Qing</given-names>
            <surname>Yang</surname>
          </string-name>
          .
          <year>2009</year>
          .
          <article-title>Time-dependent models in collaborative ifltering based recommender system</article-title>
          .
          <source>In Proceedings of the 2009 IEEE/WIC/ACM International Joint Conference on Web Intelligence and Intelligent Agent TechnologyVolume 01</source>
          , Vol.
          <volume>1</volume>
          . IEEE Computer Society,
          <fpage>450</fpage>
          -
          <lpage>457</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref30">
        <mixed-citation>
          [30]
          <string-name>
            <surname>Yin</surname>
            <given-names>Zhang</given-names>
          </string-name>
          , Min Chen, Dijiang Huang,
          <string-name>
            <surname>Di Wu</surname>
            , and
            <given-names>Yong</given-names>
          </string-name>
          <string-name>
            <surname>Li</surname>
          </string-name>
          .
          <year>2017</year>
          .
          <article-title>iDoctor: Personalized and professionalized medical recommendations based on hybrid matrix factorization</article-title>
          .
          <source>Future Generation Computer Systems</source>
          <volume>66</volume>
          (
          <year>2017</year>
          ),
          <fpage>30</fpage>
          -
          <lpage>35</lpage>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>