<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Cold start Using Stereotypes with Dynamic User and Item Features</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Nourah A. AlRossais</string-name>
          <email>nalrossais@ksu.edu.sa</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>King Saud University, College of Computer and Information Sciences, Information Technology Department</institution>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Workshop Proce dings</institution>
        </aff>
      </contrib-group>
      <abstract>
        <p>A demanding operation regime of Reccomender Systems (RS) is that of extreme cold start followed by a 'warming up' phase, in which a new user begins to interact with the items in the catalogue, or a new item receives the first interactions by existing users. In the RS literature the majority of approaches and techniques have been proposed for a fully 'warm' RS; only a small subset of the research addresses the cold start and within that subset, the 'warming up' phase is an area that has received little attention. During warm up new user (new item) interactions begin to appear but they are too few for a collaborative filtering technique to work, while a smart content based filtering approach (those of extreme cold start) may not be able to capture the emerging personalization traits. In this paper, starting from a stereotype driven approach developed for pure cold starts, we formulate and discuss a dynamic model that uses the few arising user/item interactions to adjust simple personalization features that are embedded in the proposed model. We demonstrate how the little personalization introduced by the dynamic model improves substantially a range of performance metrics during warm up.</p>
      </abstract>
      <kwd-group>
        <kwd>Recommender system</kwd>
        <kwd>cold start</kwd>
        <kwd>warming up</kwd>
        <kwd>new item problem</kwd>
        <kwd>new user problem</kwd>
        <kwd>stereotypes</kwd>
        <kwd>dynamic bias</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>Dynamic User and Item
1. Introduction
ommender System (RS) is that of cold start, namely when
the RS is required to produce content recommendations
to new unknown users, or when new content is added
to the catalogue and the task is to recommend the new
unknown content to its existing users. Diferent
solutions have been proposed to address cold start, some data
driven and some technique driven, as reviewed in [1].
The author in [2] demonstrated how creating rating
agnostic stereotypes for both users and items lead to better
zero interactions available from the new user or
concernrecommendation metrics during extreme cold start (i.e. 10].
ing the new item); only deep learning architectures can
achieve comparable levels of accuracy, serendipity and
fairness during extreme cold start to that of stereotyped
metadata features [3].</p>
      <p>When a new user begins interacting with the items
catalogue, or when new content begins to be rated by
some users, the RS can make use of such information
to update its extreme cold start algorithm in light of the
new evidence. During such ”warm up” phase, too few
interactions are available to switch to a collaborative
ifltering approach, but such little interactions should not
be omitted as they may give important personalization
0009-0008-1062-9001 (N. A. AlRossais)</p>
    </sec>
    <sec id="sec-2">
      <title>2. Approach</title>
      <p>In a generic RS framework, user  consumes content/item
 and attributes an explicit or implicit liking via a rating
tentially non linear or stochastic, e.g. a deep learning
architecture) that models   via the user’s describing
feature vector   , the item’s describing feature vector   , via
all previous ratings  
provided to item  by all other</p>
      <p>
        users ( ≠  ) as well as such other users’ features, and
ifnally all previous ratings 
given by the user  to other
items  and all such other item’s features. In our previous
works [2, 3] we modeled further the functional form of
our RS as in equation (
        <xref ref-type="bibr" rid="ref1">1</xref>
        ):

 =  (  ,   ,   ,   )
(
        <xref ref-type="bibr" rid="ref1">1</xref>
        )
      </p>
      <p>
        Equation (
        <xref ref-type="bibr" rid="ref1">1</xref>
        ) simplifies further the problem by
condensing the explanatory efects provided by all other users
having consumed item  prior to user  in a simplified
term,   , that we call characteristic item bias. In a
similar fashion, the explanatory efects provided by all other
items rated by user  prior to encountering item  , are
incorporated in a simplified term,   , the characteristic user
bias. All previous user to item interactions, not directly
involving user  and item  , together with their metadata
features, are used in the generation of the functional form
 , i.e. training the model (
        <xref ref-type="bibr" rid="ref1">1</xref>
        ).
      </p>
      <p>
        Sterotypization of the features was introduced via
standard algorithms that can take any metadata feature type
and transform it into a stereotyped version of the original
metadata with substantially smaller dimensions. Doing
so for both user and item features aids the sparsity of the
problem, and simultaneously creates a basis,  ̃ and  ̃, that
improves the training process and predictive power of
model (
        <xref ref-type="bibr" rid="ref1">1</xref>
        ), as illustrated in [2].
      </p>
      <sec id="sec-2-1">
        <title>During the new item and new user extreme cold starts, using stereotyped features, the RS can be rewritten for the new user extreme cold start problem as: And for the new item extreme cold start problem as:</title>
        <p>= (  ̃,   ,  ̃ , ( ̄  ̃ ))

 =  (  ̃, ( ̄   ̃),  ̃ ,   )</p>
      </sec>
      <sec id="sec-2-2">
        <title>In the new user cold start, given that no prior rating</title>
        <p>information is available about the new user, ( ̄  ̃ )
represents the typical user bias of all previously observed
users belonging to the same stereotypes. Likewise in the
new item problem, ( ̄   ̃) represents the typical item bias
across all items belonging to the same stereotype.</p>
        <p>
          In the present research we extend the models (
          <xref ref-type="bibr" rid="ref2 ref3">2,3</xref>
          ) to
adapt dynamically during the warming up phase using
the concept of dynamic adaptive features introduced by
[11]. In particular, we can write for the new user warming
up phase, when there are exactly  ratings available for
user  and we are modeling the  + 1 consumption of
item  :

(+1) = (  ̃,   ,  ̃ ,
        </p>
        <p />
        <p>(+1) )
 
(+1) =  ⋅ ( ̄  ̃ ) + (1 −  )⋅ &lt;   &gt;1,..,
 = (/</p>
        <p>
          )
In (
          <xref ref-type="bibr" rid="ref4">4</xref>
          )
        </p>
        <p>(+1) represents the dynamic user bias. &lt;</p>
        <p>&gt;1,.., is the average observed bias of the
particular user  , over its first  reviews. When  is small the
observed bias is not trustworthy, and it receives a low
weight compared to</p>
        <p>(+1) . As the user advances in his
interactions with the items a stronger personalization
model’s parameters optimized during training.
arises, and when  grows to the value of   the user bias
is fully personalized. In this model   and  are also</p>
      </sec>
      <sec id="sec-2-3">
        <title>The new item warming up phase can be modeled similarly: (4) (5)</title>
        <sec id="sec-2-3-1">
          <title>3.1. Cold start warming experiments</title>
          <p>Beginning from pure cold start experiments, as described
in [2, 3], when a user (item) is left out to represent a new
user (item), all its ratings and interactions are blanked
out and temporal consistency is enforced during the
experiments. Only the users/items/ratings that had been
expressed at a time prior to the new user (item) joining


(+1) =  (  ̃,</p>
          <p />
          <p>
            (+1) ,  ̃ ,   )
(+1) =  ⋅ ( ̄   ̃) + (1 −  )⋅ &lt;   &gt;1,..,

 = (/

)

In (
            <xref ref-type="bibr" rid="ref5">5</xref>
            )
          </p>
          <p>(+1) is the dynamic item bias, expressed as
a dynamic weighted sum of (  ̃) and the actual item

 average bias for its first  reviews, &lt;   &gt;1,.., . The
model improves the item characterisation as its number
of reviews  increases, obtaining full characterization
after   interactions.</p>
        </sec>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>3. Experimental settings and</title>
      <p>
        (
        <xref ref-type="bibr" rid="ref2">2</xref>
        )
      </p>
      <p>
        results
The ideal dataset to demonstrate the application of
models (
        <xref ref-type="bibr" rid="ref4 ref5">4, 5</xref>
        ), must exhibit several characteristics; firstly the
(
        <xref ref-type="bibr" rid="ref3">3</xref>
        ) ratings (interactions) should be timestamped, so that the
full history of how new users, new items and user/item
interactions occurred is available. Secondly it should be
rich in metadata features for both users and items, and
ifnally it should be publicly available to make findings
reproducible. Based on these requirements the dataset
used is the Movielens/IMDb integrated dataset described
in [12].
the online platform are retained to train the RS
functional shapes, i.e. the  and the  of (
        <xref ref-type="bibr" rid="ref2 ref3 ref4 ref5">2,3,4,5</xref>
        ). In addition,
the temporal order of the interactions provided by the
new user (or to the new item) must be maintained when
training and evaluating the dynamic biases of (
        <xref ref-type="bibr" rid="ref4">4</xref>
        ), (
        <xref ref-type="bibr" rid="ref5">5</xref>
        ).
      </p>
      <p>
        Following the works and findings of [ 2, 3] we train
two algorithms for both  and  of problems (
        <xref ref-type="bibr" rid="ref2 ref3 ref4 ref5">2,3,4,5</xref>
        ),
the first one is Extreme Gradient Boosting, [ 13] (XGB).
      </p>
      <p>This is chosen as our benchmark, because in the previous
work [2] it represented the best extreme cold start model
among the various machine learning / statistical driven
models discussed in [2]. The second is a deep neural
network algorithm, (DNN) whose deep layers, coupled
by dropout layers, and performance during extreme cold
starts have been the subject of the research [3]. The main
focus of the paper is not to detail which of the functional
algorithms exhibits better performance metrics during
cold start and during the warm up transition, instead the
objective is to investigate the efect of the dynamic bias
model, how that improves the warming up phase and
whether the findings can be ascribed to any particular
RS functional form, or if they are independent of the RS
chosen.</p>
      <p>
        In the results that follow we discuss the behaviour of
the extreme cold start model and of the dynamic bias
warm up model as the new user begins rating existing
items (or existing users begin rating the new item). In
particular we examine the performance metrics at several
increasing review counts,  , i.e. how the metrics change
when we take away the first  reviews of the user for the
new user case (or reviews to the item for the new item)
and use these to train the dynamic bias models (
        <xref ref-type="bibr" rid="ref4">4</xref>
        ), (
        <xref ref-type="bibr" rid="ref5">5</xref>
        ).
      </p>
      <p>When assessing performance only the interactions that
follow, from  + 1 , are utilized to compute the metric
for both warming model as well as for the extreme cold
start model.</p>
      <p>In the experiments presented in this paper we restrict
our attention to users that have more than 100
interactions with the catalogue (approximately 4,000 users), and
items that have been interacted with at least 100 times
(approximately 2,500 items).</p>
      <p>In addition to well known accuracy metrics of single
predictions and ranked lists we will also discuss how the
warming up phase afects serendipity (SER) and fairness
(FRN) of ranked lists. SER is the property of generating
ranked lists that contain useful and unexpected items,
therefore exposing users to larger parts of the catalogue. to the new item.
Fairness (FRN), which is related to SER, focuses on the
tail of the distribution of the items that are recommended
the fewest, it measures how often the least recommended
items are injected in some lists. The latter metric is a
very actual research issue in RS, after ethical concerns
raised by some industry RS [14].</p>
      <sec id="sec-3-1">
        <title>3.2. Results</title>
        <sec id="sec-3-1-1">
          <title>In figures ( 1) and (2) we show the efect on the Rmse of the predicted ratings when warming the RS’s personalization via the dynamic bias models based on stereotypes, versus continuing the recommendations using the stereotyped</title>
          <p>
            extreme cold start model. Three distinct interesting facts be adopted as discussed in [2]. One of the most relevant
can be inferred from these results; firstly the dynamic and popular measure is the Normalized Discounted
Cupersonalization improves the Rmse recommendation per- mulative Gain (Ndcg) whose definition can be found in
formance by 3 to 4% for both the new user and new item [2] and references within. In this context it is suficient
experiments. Secondly, the improvement obtained is in- to remind that the higher the Ndcg, the more valuable
dependent of the functional form of the RS chosen in the content of the ranked list to the user (both in terms
our experiments XGB vs DNN. Each RS has its own cold of items present and how they are ranked).
start base prediction ability and the warming up of the Figures (
            <xref ref-type="bibr" rid="ref3">3</xref>
            ) and (
            <xref ref-type="bibr" rid="ref4">4</xref>
            ) show the standardized Ndcg of both
model bias is beneficial in a similar manner to both func- the top-10 and top-20 ranked lists for the new user and
tional forms and across experiments. Thirdly, there is new item experiments. Given that the Ndcg tends to
a diferent behaviour in the new user vs the new item decay as the top-N list grows, we standardize all the data
experiments. As the user begins interacting with the by the   (=100) (the Ndcg of the top 10 list when there
catalogue it becomes more dificult for the pure stereo- are no user to item interactions, i.e. the extreme cold start
type based cold start model to predict the ratings of the case). In all cases when there are  reviews available the
items encountered further down the catalogue interac- assessment of the   takes the items reviewed into
tions path. This is potentially explained by the fact that account, excluding them from the lists, therefore also the
items that are well known and widely interacted with ranked lists of the extreme cold start base model changes
are easier to predict and usually come earlier in the re- with  . In both experiments we notice the tendency of
view histories. With new items the situation is slightly the   to decay as the number of reviews  increases.
diferent, the personalization efect of a dynamic bias In both experiments and for both functional forms tested
improves performance versus the cold start base model, for our warming up RS, the dynamic bias substantially
but also the base algorithm improves (at a slower rate) improves the ranked list quality as measured by the   ,
as the reviews increase, hence indicating that the first albeit in a slightly diferent manner. In the case of the new
temporal reviews to a new item are the most dificult to user experiment the improvement due to personalization
predict. efects to the dynamic bias happens relatively quickly in
          </p>
          <p>When moving from accuracy of single rating predic- this dataset, where an approximately 5% improvement
tions to a description of how well a RS proposes ranked over cold start base model is obtained within the first
lists of items to users, there are several metrics that can 5, 10 interactions lifting the Ndcg curve. In the new
item case the improvement over pure cold start is more mulation of the recommendation algorithms, and that is
distributed over the entire number of interactions. These indicating that they can be a source of improved accuracy
two distinct behaviours may be related to the diferent independently of the RS functional shape chosen.
statistical nature of users and items. As a future work the author plans to enrich the models</p>
          <p>
            Our previous works [2, 3] introduced operative metrics presented with a collaborative filtering element, to be
for serendipity (SER) and fairness (FRN) of ranked lists, in modeled via a Markov chain. The dynamic bias presented
this context we refer to such definitions. Table (
            <xref ref-type="bibr" rid="ref1">1</xref>
            ) shows in this work, coupled with a Bayesian efect (to be
forthe efect of warming up the dynamic bias on SER and mulated) should provide a construct capable of smoothly
FRN, for brevity we only report a single top-N for the new transitioning from a content based cold start RS, based
user case with the DNN functional shape, as a function on stereotyped features, to a collaborative filtering model
of the new user growing interactions  (for brevity we when there are enough user to items interactions.
do not show XGB results which are very similar). We can Finally, this work presents results on an integrated
see that dynamic bias approach improves substantially dataset rich in user and item metadata, which also fulfils
the SER of the top N list as the personalization increases the necessary requirement of having interactions
timesthe lists cover more and more catalogue. However, the tamped. Part of our future research will be focused on
specialization also reduces the FRN for the new user finding data sets that have the required characteristics
case. As specialization occurs the probability of being and extending our findings to diferent domains.
recommended for the least recommended items decreases
initially and only at larger  the FRN of the top-10 ranked
list surpasses that of the base cold start model. References
          </p>
        </sec>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>4. Conclusions and future work</title>
      <p>The model proposed in this paper addresses the
warming up of RS from cold start, it presents a unified
construct for both the new user and new item problems in a
stereotype driven framework. The model allows for a
degree of personalization via dynamic user/item bias terms.</p>
      <p>The findings discussed in the paper demonstrate how for
the new user problem the adoption of the bias allows
for a rapid improvement in accuracy metrics, while for
the new item problem the improvement is more gently
spread over the initial incoming interactions. The only
metric that using dynamic biases is observed to under
perform the pure content based stereotype model is the
fairness (FRN) metric for the new user problem. Such
small under-performance in fairness may be the price to
pay for a rapid user specialization. The new user and new
item dynamic bias are efective with two alternative
for[6] J. Li, Y. Wang, J. McAuley, Time interval aware start advertisements: Improving ctr predictions via
self-attention for sequential recommendation, in: learning to learn id embeddings, in: Proceedings
Proceedings of the 13th international conference of the 42nd International ACM SIGIR Conference
on web search and data mining, 2020, pp. 322–330. on Research and Development in Information
Re[7] M. Vartak, A. Thiagarajan, C. Miranda, J. Bratman, trieval, 2019, pp. 695–704.</p>
      <p>H. Larochelle, A meta-learning perspective on cold- [11] G. Behera, N. Nain, Collaborative filtering with
temstart recommendations for items, Advances in neu- poral features for movie recommendation system,
ral information processing systems 30 (2017). Procedia Computer Science 218 (2023) 1366–1373.
[8] C. Finn, P. Abbeel, S. Levine, Model-agnostic meta- [12] N. A. ALRossais, D. Kudenko, isynchronizer: A tool
learning for fast adaptation of deep networks, in: In- for extracting, integration and analysis of
movieternational conference on machine learning, PMLR, lens and imdb datasets, in: Adjunct Publication of
2017, pp. 1126–1135. the 26th Conference on User Modeling, Adaptation
[9] Y. Zhu, R. Xie, F. Zhuang, K. Ge, Y. Sun, X. Zhang, and Personalization, 2018, pp. 103–107.</p>
      <p>L. Lin, J. Cao, Learning to warm up cold item embed- [13] T. Chen, T. He, M. Benesty, V. Khotilovich, Y. Tang,
dings for cold-start recommendation with meta scal- H. Cho, K. Chen, R. Mitchell, I. Cano, T. Zhou, et al.,
ing and shifting networks, in: Proceedings of the Xgboost: extreme gradient boosting, R package
44th International ACM SIGIR Conference on Re- version 0.4-2 1 (2015) 1–4.
search and Development in Information Retrieval, [14] A. A. Kodiyan, An overview of ethical issues in
2021, pp. 1167–1176. using ai systems in hiring with a case study of
ama[10] F. Pan, S. Li, X. Ao, P. Tang, Q. He, Warm up cold- zon’s ai based hiring tool (2019).</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>D. K.</given-names>
            <surname>Panda</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Ray</surname>
          </string-name>
          ,
          <article-title>Approaches and algorithms to mitigate cold start problems in recommender systems: a systematic literature review</article-title>
          ,
          <source>Journal of Intelligent Information Systems</source>
          <volume>59</volume>
          (
          <year>2022</year>
          )
          <fpage>341</fpage>
          -
          <lpage>366</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>N.</given-names>
            <surname>AlRossais</surname>
          </string-name>
          , D. Kudenko, T. Yuan,
          <article-title>Improving coldstart recommendations using item-based stereotypes, User Model User-Adap Inter 31 (</article-title>
          <year>2021</year>
          )
          <fpage>867</fpage>
          -
          <lpage>905</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>N.</given-names>
            <surname>AlRossais</surname>
          </string-name>
          ,
          <article-title>Improving cold start stereotype-based recommendation using deep learning</article-title>
          ,
          <year>2023</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>S.</given-names>
            <surname>Rendle</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Freudenthaler</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Schmidt-Thieme</surname>
          </string-name>
          ,
          <article-title>Factorizing personalized markov chains for nextbasket recommendation</article-title>
          ,
          <source>in: Proceedings of the 19th international conference on World wide web</source>
          ,
          <year>2010</year>
          , pp.
          <fpage>811</fpage>
          -
          <lpage>820</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>C.</given-names>
            <surname>Cai</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>He</surname>
          </string-name>
          ,
          <string-name>
            <surname>J. McAuley</surname>
          </string-name>
          ,
          <article-title>Spmc: Socially-aware personalized markov chains for sparse sequential recommendation</article-title>
          ,
          <source>arXiv preprint arXiv:1708.04497</source>
          (
          <year>2017</year>
          ).
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>