<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>into Recom mendation Datasets</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Weizhe Lin</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff5">5</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Linjun Shou</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
          <xref ref-type="aff" rid="aff5">5</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Ming Gong</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
          <xref ref-type="aff" rid="aff5">5</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Pei Jian</string-name>
          <email>jpei@cs.sfu.ca</email>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
          <xref ref-type="aff" rid="aff5">5</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Zhilin Wang</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff4">4</xref>
          <xref ref-type="aff" rid="aff5">5</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Bill Byrne</string-name>
          <email>bill.byrne@eng.cam.ac.uk</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff5">5</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Daxin Jiang</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
          <xref ref-type="aff" rid="aff5">5</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Department of Engineering, University of Cambridge</institution>
          ,
          <addr-line>Cambridge</addr-line>
          ,
          <country country="UK">United Kingdom</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>In summary</institution>
          ,
          <addr-line>our contributions are:</addr-line>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>Microsoft STCA</institution>
          ,
          <addr-line>Beijing</addr-line>
          ,
          <country country="CN">China</country>
        </aff>
        <aff id="aff3">
          <label>3</label>
          <institution>Simon Fraser University</institution>
          ,
          <addr-line>British Columbia</addr-line>
          ,
          <country country="CA">Canada</country>
        </aff>
        <aff id="aff4">
          <label>4</label>
          <institution>University of Washington</institution>
          ,
          <addr-line>Seattle</addr-line>
          ,
          <country country="US">United States</country>
        </aff>
        <aff id="aff5">
          <label>5</label>
          <institution>strong Collaborative Filtering (CF) and Content-based Fil-</institution>
        </aff>
      </contrib-group>
      <abstract>
        <p>Popular book and movie recommendation datasets can be associated with Knowledge Graphs (KG) that enable the development of KG-based recommender systems. However, most of these approaches are based on Collaborative Filtering, leaving Contentbased Filtering approaches unexploited. This is partially due to the lack of items' content-based information (e.g. summary texts of movies and books) in datasets. To facilitate the research in achieving both KG-aware and content-aware recommender systems, we contribute to public domain resources through the creation of a large-scale Movie-KG dataset and an extension of the already public Amazon-Book dataset through incorporation of text descriptions crawled from external sources. Both datasets provide items' descriptive texts that enable recommendations based on unstructured content. We provide benchmark results as well as showing the value of the content-based information in making recommendations.</p>
      </abstract>
      <kwd-group>
        <kwd>Knowledge graph</kwd>
        <kwd>recommender systems</kwd>
        <kwd>recommendation dataset</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>item interactions along with a KG that provides external
knowledge. This makes reasoning and searching over
such as Amazon-Book and MovieLens-20M provide user- such information.</p>
    </sec>
    <sec id="sec-2">
      <title>1. Introduction</title>
      <p>In recent years, modern recommender systems (RS) based
cessful solutions that recommend items (such as movies,
books, and news) to users. There has been a growing
interest in incorporating Knowledge Graphs (KG) into
recommendation since recent study showed that RS can
benefit from the external information provided by KGs
to enrich the user/item representations [1]. KGs link
items to be recommended to other related KG entities,
and these connections serve as “item properties” (e.g.
movie genre, movie actors, and book editors). Datasets
WA, USA.
nEvelop-O
STCA.</p>
      <sec id="sec-2-1">
        <title>KGs possible.</title>
        <p>While KGs can readily incorporate structured content
information and external knowledge, unstructured
content such as item descriptions, is unexploited in these
popular KG-based recommendation datasets. We note
†This work was done during Weizhe Lin’s internship at Microsoft
mary texts from web resources. movie recommendation dataset that was newly collected
2. We have created a brand-new large-scale dataset from real user behaviors (Sec. 3.1); (2) the
Amazon-Bookfor movie recommendation. The dataset was col- Extended dataset based on the Amazon-Book dataset [4]
lected from real users of Microsoft Edge, and KGs and newly augmented with textual book descriptions
and movie descriptions were extracted from in- (Sec. 3.2).</p>
        <p>ternal knowledge bases.
3. We finally provide benchmark results for both 3.1. Movie KG Dataset
datasets with strong RS models. The results show
that the incorporation of unstructured content 3.1.1. Dataset Construction
leads to system improvements.</p>
      </sec>
      <sec id="sec-2-2">
        <title>This dataset was formed by real user browsing behav</title>
        <p>iors that were logged by a popular commercial browser
2. Related Work between 13/05/2021 and 20/06/2021 (39 days in total).
Sensitivity of data and privacy was firstly removed by
Recommendation Datasets. Some popular existing rec- decoupling users’ real identity (e.g. IP, account
identiommendation datasets are shown in the first section of Ta- ifer) from the data and assigning each user an insensitive
ble 1. MovieLens-20M1 is a popular benchmark that has unique virtual identifier. To link users’ browsing
hisbeen widely used, while Flixster2 is large but less popu- tory with movies, we exploited a commercial knowledge
lar in recent work. They are both movie datasets but they graph that consists of a large number of entities intended
do not provide associated KGs to enable the development to cover the domain and a rich set of relationships among
of KG-aware systems. Amazon-Book [4], Last-FM [4], them. Entities with type “movie” were extracted together
Book-Crossing [5]3 and Alibaba-iFashion [6] are ded- with their one-hop neighbors and edges. In this way a
icated for KG-based CF model evaluation. They pro- movie-specific sub-graph was extracted from the original
vide rich user-item interactions but unstructured content- knowledge graph.
based information (e.g. item summary texts in natural Drawing from billions of user browsing logs, KG entity
language) is not included. In news recommendations, in linking was performed to match web page titles with
contrast to the small Yahoo! dataset [7], MIND [8] from the movie titles in the movie-specific KG. Movie entities
Microsoft is a more recent and larger dataset collected were extracted using an internal service that provides
from real user behaviors. However, these datasets only Named Entity Recognition and entity linking. To make
support the development of content-aware CBF models the dataset more compact: (1) consecutive and duplicated
since no KG is available. user interactions with the same movie were merged to</p>
        <p>
          KG-aware Recommender Systems. Beyond tradi- one single record by aggregating their browsing times;
tional CF models that model users and items through only (2) less active users with fewer than 10 interactions in 39
interaction data [9, 10, 11, 12, 13, 14, 15, 13, 16, 17, 18, 19, days were dropped; (3) the most frequent 50,000 movies
20, 21], KG-based CF models fuse external knowledge with the most user interactions were selected, since tail
from auxiliary KGs to improve both the accuracy and ex- records were considered unreliable.
plainability of recommendation [1, 22]. The CF methods 125,218 active users and 50,000 popular movies were
to exploit the KGs can be categorized into Embedding- chosen. KG entities related to these movies were
exbased Methods [23, 24, 25], Path-based Methods [26, 27, 28, tracted to form a sub-KG with 250,327 entities
(includ29], and GNN-based Methods [30, 31, 27, 4, 32]. CBF mod- ing movie items). Each user is represented by [    ,
els match items to a user by considering the metadata      ], where     is a unique identifier that has
(content-based information) of items with which the user been delinked from real user identity, and      is
has interacted [
          <xref ref-type="bibr" rid="ref7">33, 34, 35, 36, 37</xref>
          ], while most research an ID list of movie items that the user has browsed.
in KG-based CBF, a recently popular topic, focuses on For training and evaluating recommender systems,
enhancing the item representations with KG embeddings      was split into train/validation/test sets: the
by mapping relevant KG entities to the content of items, first 80% of users’ historical interactions (ordered by click
e.g., by entity linking [38, 39]. date-time) are in the train set; 80% to 90% serves as the
validation set; and the remaining is reserved for testing.
        </p>
        <p>During validation and test, interactions in the train set
3. Datasets serve as users’ previous clicked items, and systems are
evaluated on their ratings for test set items. In addition
We introduce the collection process of the two new to the traditional data split strategy, a cold-start user
datasets in this section: (1) a large-scale high-quality set was created to evaluate model performance on users
outside the training dataset. 3% of users were moved
to the cold-start set and they are not available in model</p>
      </sec>
      <sec id="sec-2-3">
        <title>1https://grouplens.org/datasets/movielens/ 2https://sites.google.com/view/mohsenjamali/flixter-data-set 3http://www2.informatik.uni-freiburg.de/ cziegler/BX/</title>
        <p>training. To challenge systems’ real abilities under
extreme cold-start scenarios, we chose users that have very
few interactions (lower than 20 interactions per user on
average). This portion of interactions was split into a
cold-start history set with the first 80% interactions of
each user, and a test set with the remaining 20%.</p>
        <p>A comparison with some popular existing
recommendation datasets is provided in Table 1. Our new movie
dataset is based on large user populations and movie
inventory and draws from an extensive movie-specific
KG with millions of triplets. This is a rich resource not
yet provided by existing popular movie recommendation
datasets. Our dataset provides not only titles and
genres (as in MovieLens-20M), but also other content-based
information, such as rich summary texts of movies that
enable content-based recommendation with large
language understanding models such as BERT and GPT-2.
3.1.2. Statistical Analysis
14000
ssu12000
re10000
fo 8000
r
eb 6000
m
uN 4000
2000</p>
        <p>0
2000
ts1500
xe
ftoh1000
t
g
n
Le 500
Fig. 1 presents the key statistics of the dataset. This
dataset contains 125,218 users, 50,000 movie items, and aFnigduArme1az:oKne-yBostoakt-iEstxitcesnodfeMdo(Fviige.-K1(Gd)-)D. ataset (Fig. 1(a)-1(c))
4M+ interactions. The KG contains 250,327 entities with
12 relation types and 12,055,581 triplets. The attributes
of movies include content-based properties (title, descrip- dustrial products 4.
tion, movie length) and properties already tied to
knowledge graph entities (production company, country, lan- 3.2. Extended Amazon Book Dataset
guage, producer, director, genre, rating, editor, writer, The Amazon-Book dataset was originally released by [4]
honors, actor). Fig. 1(a)-1(c) shows the distribution of and has been used in the development of many advanced
length of interactions with movies per user, number of in- recommendation systems. However, this dataset
proteractions with users per movie, and the length of movie vides only CF interaction data without content-based
descriptions. The average length of movie descriptions features needed for CBF evaluation. To fill this gap, we
is 425 words, which is long enough for models to encode extend the dataset with descriptions extracted from
mulmovies from the summary of stories. The average num- tiple data sources, noting that the dataset was originally
ber of user interactions is 32.70, providing rich resources collected in 2014 and many items have since expired on
for modeling long-term user interest. The first 19 days of the Amazon website. The procedure is summarized as
data (with similar data distribution but fewer users and follows:
interactions) is separated as “Standard” for most research (1) We matched product descriptions in Amazon with
purposes, while the full 39 days of data are released as entries in Amazon-Book using their unique item
iden“Extended” for more extended use, such as training
in4We provide results of this extended set in our repository.</p>
        <p>20Length of interactions of users</p>
        <p>40 60 80 100
(a) Length of user history ( =
32.70,  = 45.76 )
0 0 Count of interactions with movies 200</p>
        <p>50 100 150
(b) Users’ interactions with
each movie ( = 83.33 ,  =
582.50)
0 0 500Number of m1ov5i0e0s 2000</p>
        <p>1000
(c) Length of movie
description in Movie-KG-Dataset ( =
425.22,  = 329.93 )
20N0umber of books 600</p>
        <p>400
(d) Length of book description
in Amazon-Book-Extended
( = 151.75 ,  = 206.98 )
#Rel.</p>
        <p>N/A
N/A
39
18
9
51
N/A
N/A
tifiers (asin). 24,555 items were successfully matched KGAT [4]: Knowledge Graph Attention Network
(98.56%); (2) we then matched each of the remaining (KGAT) which explicitly models high-order KG
connecitems with the most relevant entry in a huge commer- tivities in KG. The models’ user/item embeddings were
cial KG which enables rich descriptions to be extracted. initialized from the pre-trained BPRMF weights.
325 items were matched at this step; (3) the remaining KGIN [6]: a state-of-the-art KG-based CF model that
28 items were matched manually to the most relevant models users’ latent intents (preferences) as a
combinaproducts in Amazon. tion of KG relations.</p>
        <p>The distribution of description lengths is given in KMPN [42]: a KG-based CF model that models users
Fig. 1(d). through learning a set of preference embeddings.</p>
        <p>
          NRMS-BERT [42]: a strong CBF model that
incorporates a pre-trained BERT for extracting content-based
4. Benchmarking features from natural language. It was inspired by
4.1. Evaluation Metrics NRMS [
          <xref ref-type="bibr" rid="ref7">34</xref>
          ], a strong news recommender system with
a bi-encoder architecture.
        </p>
        <p>Following common practice [19, 4, 6, 40], we report met- Mixture of Expert: a hybrid system where the output
rics for evaluating model performance: (1) Recall@K : scores of two systems, KMPN and NRMS-BERT, passed
within top- recommendations, how well the system through 3 layers of a Multi-Layer Perception (MLP) to
recalls the test-set browsed items for each user; (2) obtain final item ratings.
ndcg@K (Normalized Discounted Cumulative Gain) [40]: CKMPN [42]: a contrastive learning approach that
increases when relevant items appear earlier in the rec- fuses the features of NRMS-BERT with those of KMPN
ommended list; (3) HitRatio@K : how likely a user finds in training. It is also a hybrid system that leverages both
at least one interesting item in the recommended top-K CBF and CF features.
items.</p>
        <sec id="sec-2-3-1">
          <title>4.2. Baselines</title>
        </sec>
      </sec>
      <sec id="sec-2-4">
        <title>We take the performance of several recently published</title>
        <p>recommender systems as points for comparison5. We
carefully reproduced all these baseline systems from their
repositories6.</p>
        <p>BPRMF [10]: a powerful Matrix Factorization (MF)
method that applies a generic optimization criterion
BPROpt for personalized ranking. Limited by space, other
MF models (e.g. FM [41], NFM [13]) are not presented
since BPRMF outperformed them.</p>
        <p>CKE [24]: a CF model that leverages heterogeneous
information in a knowledge base for recommendation.</p>
      </sec>
      <sec id="sec-2-5">
        <title>5They are also baseline systems being compared in a recent paper</title>
        <p>[6] (WWW’21).
6As a result, the results reported here may difer from those of the
original papers.</p>
        <sec id="sec-2-5-1">
          <title>4.3. Discussion</title>
        </sec>
      </sec>
      <sec id="sec-2-6">
        <title>On the Amazon-Book-Extended dataset, as shown in Ta</title>
        <p>ble 2, KG-based systems (e.g. KGAT, KMPN) leverage
structured information embeded in KGs to achieve ∼0.17
Recall@20. The CBF-system (NRMS-BERT) achieves
0.1142 Recall@20 with only summary texts of books.</p>
        <p>The performance is not far from that of KG-based
models. It shows that our extension to the original dataset
is successful and this content-based information can be
used for making content-aware recommendations. The
best scores are obtained by hybrid methods (MoE and
CKMPN) that combine KMPN and NRMS-BERT to make
recommendations. In particular, the CKMPN model
improves results of @60/@100, showing that even though
NRMS-BERT’s performance is lower, KG-based systems
can still benefit from incorporating its features. This
0.1980
0.1950
0.1997
0.2194
0.2193
0.1728
0.2266
0.1783
0.0807
0.1812
0.3236
0.3155
0.3196
0.3643
0.3602
0.2773
0.3668
0.3414
0.2071
0.3380
also shows that content-based features are useful, but making recommendations. It is challenging to exploit
simple aggregation schemes can not improve the per- them at the same time while overcoming the limitations
formance significantly. One research challenge left for of each item property.
future research is how to better exploit the content-based
features. Table 4</p>
        <p>We observe a similar trend on Movie-KG-Dataset. The Case study for a user who have browsed the movie Tenet
hybrid method, CKMPN, achieves the best performance. (2020). Source Code (2011) has a similar genre, while Dunkirk
We note that the performance of NRMS-BERT is closer (2017) has the same director. Y/N: whether or not the movie
to that of KG-based models (e.g. KMPN). The improve- appears in the top-100 recommendation list of the models.
ment brought by content-based features is also more MoE: Mixture of Expert.
obvious on Movie-KG-Dataset. This is because Movie- Item KMPN NRMS-BERT MoE
KG-Dataset has larger average text lengths compared
to Amazon-Book-Extended (425.22 v.s. 151.75), ofering SoDurucnekCirokd(e20(21071)1) NY NY NN YY
richer information that result in more discriminative item
embeddings.</p>
        <p>An example output of systems is presented in Ta- On the cold-start test set, CF, CBF, and hybrid methods
ble 4. Y/N indicates whether or not the movie appears have achieved much lower results. The performance
rein the top-100 recommendation list of the four models duction is more obvious on the CBF model (NRMS-BERT
(KMPN/NRMS-BERT/Mixture of Expert (MoE)/CKMPN). drops from 0.1241 to 0.0437 Recall@20), since the users
This user has browsed Tenet (2020) directed by Christo- in this set have very few (&lt;20) historical interactions on
pher Nolan. The movie Source Code (2011) and Tenet are average, making a content-based system fail to capture
both about time travel, but they have quite diferent film their characterstics and make suitable recommendations.
crews. As a result, Source Code was considered positive Therefore, another research challenge is to improve the
by NRMS-BERT which evaluates on the movie descrip- cold-start performance of all these models, making
modtion, but was considered negative by KG-based KMPN. els more robust to extreme cases.</p>
        <p>Combining the scores of both systems, MoE did not
recommend the movie. However, CKMPN complemented 5. Conclusion
the failure of KMPN and gave a high score for this movie,
by learning a content-aware item representation based on To facilitate research in developing RS models that are
the representation of NRMS-BERT through contrastive both KG-aware and content-aware, we introduced two
learning. In contrast, Dunkirk (2017) is about war and datasets: (1) Amazon-Book-Extended which inherits a
history which is not in the same topic as Tenet. However, popular book recommendation dataset and has been
since they were directed by the same director, KMPN newly extended with summary descriptions; (2) a new
and CKMPN both recommended this movie, while MoE’s large-scale movie recommendation dataset,
Movie-KGprediction was negatively afected by NRMS-BERT. This Dataset, based on recently collected user interactions on
case study suggests that both KGs and unstructured con- a widely-used commercial web browser, accompanied by
tent (summary texts of movies in this case) are useful for a knowledge graph and summary texts for the movies.
CKMPN
We provided benchmark results that demonstrated the techniques for recommender systems, Computer
value of incorporating content-based item descriptions. 42 (2009) 30–37.</p>
        <p>Both datasets are shared in our github repository: [10] S. Rendle, C. Freudenthaler, Z. Gantner, L.
Schmidthttps://github.com/LinWeizheDragon/Content-Aware- Thieme, Bpr: Bayesian personalized ranking from
Knowledge-Enhanced-Meta-Preference-Networks-for- implicit feedback, in: Proceedings of the
TwentyRecommendation. Fifth Conference on Uncertainty in Artificial
Intelligence, UAI ’09, AUAI Press, Arlington, Virginia,
USA, 2009, p. 452–461.</p>
        <p>References [11] Y. Koren, Factorization meets the neighborhood:
A multifaceted collaborative filtering model, in:
[1] Q. Guo, F. Zhuang, C. Qin, H. Zhu, X. Xie, H. Xiong, Proceedings of the 14th ACM SIGKDD
InternaQ. He, A survey on knowledge graph-based recom- tional Conference on Knowledge Discovery and
mender systems, IEEE Transactions on Knowledge Data Mining, KDD ’08, Association for
Comand Data Engineering (2020). puting Machinery, New York, NY, USA, 2008,
[2] J. Devlin, M.-W. Chang, K. Lee, K. Toutanova, BERT: p. 426–434. URL: https://doi.org/10.1145/1401890.</p>
        <p>Pre-training of deep bidirectional transformers for
language understanding, in: Proceedings of the [12] 1S4.0R1e9n4d4le.,dFoai:c1t0o.ri1z1a4ti5o/n1m40a1c8h9in0e.s1w40it1h94li4b.fm, ACM
2019 Conference of the North American Chap- Transactions on Intelligent Systems and
Technolter of the Association for Computational Linguis- ogy (TIST) 3 (2012) 1–22.
tics: Human Language Technologies, Volume 1 [13] X. He, T.-S. Chua, Neural factorization
ma(Long and Short Papers), Association for Com- chines for sparse predictive analytics, in:
Proputational Linguistics, Minneapolis, Minnesota, ceedings of the 40th International ACM SIGIR
2019, pp. 4171–4186. URL: https://aclanthology.org/ Conference on Research and Development in
InN19-1423. doi:10.18653/v1/N19- 1423. formation Retrieval, SIGIR ’17, Association for
[3] A. Radford, J. Wu, R. Child, D. Luan, D. Amodei, Computing Machinery, New York, NY, USA, 2017,
I. Sutskever, et al., Language models are unsuper- p. 355–364. URL: https://doi.org/10.1145/3077136.
vised multitask learners, OpenAI blog 1 (2019) 9.
[4] X. Wang, X. He, Y. Cao, M. Liu, T.-S. Chua, Kgat: [14] 3R0.8J.0O77e7n.tdaoryi:o1,0E..1-1P4.L5/im3 0, 7J.7-W13.6L.o3w0,8D0 7.7L7o., M.
FineKnowledge graph attention network for recommen- gold, Predicting response in mobile advertising
dation, in: Proceedings of the 25th ACM SIGKDD with hierarchical importance-aware factorization
International Conference on Knowledge Discovery machine, in: Proceedings of the 7th ACM
interna&amp; Data Mining, 2019, pp. 950–958. tional conference on Web search and data mining,
[5] C.-N. Ziegler, S. M. McNee, J. A. Konstan, G. Lausen, 2014, pp. 123–132.</p>
        <p>Improving recommendation lists through topic di- [15] H. Guo, R. TANG, Y. Ye, Z. Li, X. He, Deepfm:
versification, in: Proceedings of the 14th interna- A factorization-machine based neural network for
tional conference on World Wide Web, 2005, pp. ctr prediction, in: Proceedings of the
Twenty22–32. Sixth International Joint Conference on
Artifi[6] X. Wang, T. Huang, D. Wang, Y. Yuan, Z. Liu, X. He, cial Intelligence, IJCAI-17, 2017, pp. 1725–1731.</p>
        <p>T.-S. Chua, Learning intents behind interactions URL: https://doi.org/10.24963/ijcai.2017/239. doi:10.
with knowledge graph for recommendation, in:
Proceedings of the Web Conference 2021, 2021, pp. [16] 2W4.9Z6h3a/nigj,cTa.iD.u2,0J1. 7W/a2n3g9,. Deep learning over
multi878–887. ifeld categorical data, in: European conference on
[7] Z. Yang, C. Xu, W. Wu, Z. Li, Read, attend and information retrieval, Springer, 2016, pp. 45–57.
comment: A deep architecture for automatic news [17] H.-T. Cheng, L. Koc, J. Harmsen, T. Shaked, T.
Chancomment generation, in: 2019 Conference on Em- dra, H. Aradhye, G. Anderson, G. Corrado, W. Chai,
pirical Methods in Natural Language Processing M. Ispir, R. Anil, Z. Haque, L. Hong, V. Jain,
(EMNLP), 2019. X. Liu, H. Shah, Wide deep learning for
rec[8] F. Wu, Y. Qiao, J.-H. Chen, C. Wu, T. Qi, J. Lian, ommender systems, in: Proceedings of the 1st
D. Liu, X. Xie, J. Gao, W. Wu, et al., Mind: A Workshop on Deep Learning for Recommender
large-scale dataset for news recommendation, in: Systems, DLRS 2016, Association for
ComputProceedings of the 58th Annual Meeting of the As- ing Machinery, New York, NY, USA, 2016, p.
sociation for Computational Linguistics, 2020, pp. 7–10. URL: https://doi.org/10.1145/2988450.2988454.
3597–3606.
[9] Y. Koren, R. Bell, C. Volinsky, Matrix factorization [18] dYo.iQ:1u0,.H11.4C5a/i2, 9K8.84R5e0n., 2W98. 8Z4h5a4n. g, Y. Yu, Y. Wen,
J. Wang, Product-based neural networks for user
response prediction, in: 2016 IEEE 16th Interna- [29] H. Zhao, Q. Yao, J. Li, Y. Song, D. L. Lee, Meta-graph
tional Conference on Data Mining (ICDM), IEEE, based recommendation fusion over heterogeneous
2016, pp. 1149–1154. information networks, in: Proceedings of the 23rd
[19] X. He, L. Liao, H. Zhang, L. Nie, X. Hu, T.-S. Chua, ACM SIGKDD international conference on
knowlNeural collaborative filtering, in: Proceedings of edge discovery and data mining, 2017, pp. 635–644.
the 26th international conference on world wide [30] H. Wang, F. Zhang, M. Zhang, J. Leskovec, M. Zhao,
web, 2017, pp. 173–182. W. Li, Z. Wang, Knowledge-aware graph neural
[20] R. Ying, R. He, K. Chen, P. Eksombatchai, W. L. networks with label smoothness regularization for
Hamilton, J. Leskovec, Graph convolutional neu- recommender systems, in: Proceedings of the 25th
ral networks for web-scale recommender systems, ACM SIGKDD international conference on
knowlin: Proceedings of the 24th ACM SIGKDD Interna- edge discovery &amp; data mining, 2019, pp. 968–977.
tional Conference on Knowledge Discovery &amp; Data [31] H. Wang, M. Zhao, X. Xie, W. Li, M. Guo,
KnowlMining, 2018, pp. 974–983. edge graph convolutional networks for
recom[21] X. He, K. Deng, X. Wang, Y. Li, Y. Zhang, M. Wang, mender systems, in: The World Wide Web
Lightgcn: Simplifying and powering graph convolu- Conference, WWW ’19, Association for
Comtion network for recommendation, in: Proceedings puting Machinery, New York, NY, USA, 2019, p.
of the 43rd International ACM SIGIR conference on 3307–3313. URL: https://doi.org/10.1145/3308558.
research and development in Information Retrieval, 3313417. doi:10.1145/3308558.3313417.
2020, pp. 639–648. [32] Z. Wang, G. Lin, H. Tan, Q. Chen, X. Liu, Ckan:
[22] J. Chicaiza, P. Valdiviezo-Diaz, A comprehensive Collaborative knowledge-aware attentive network
survey of knowledge graph-based recommender for recommender systems, in: Proceedings of the
systems: Technologies, development, and contribu- 43rd International ACM SIGIR Conference on
Retions, Information 12 (2021) 232. search and Development in Information Retrieval,
[23] H. Wang, F. Zhang, M. Zhao, W. Li, X. Xie, M. Guo, 2020, pp. 219–228.</p>
        <p>
          Multi-task feature learning for knowledge graph [33] J. Liu, P. Dolan, E. R. Pedersen, Personalized news
enhanced recommendation, in: The World Wide recommendation based on click behavior, in:
ProWeb Conference, 2019, pp. 2000–2010. ceedings of the 15th international conference on
[24] F. Zhang, N. J. Yuan, D. Lian, X. Xie, W.-Y. Ma, Col- Intelligent user interfaces, 2010, pp. 31–40.
laborative knowledge base embedding for recom- [
          <xref ref-type="bibr" rid="ref7">34</xref>
          ] C. Wu, F. Wu, S. Ge, T. Qi, Y. Huang, X. Xie,
Neumender systems, in: Proceedings of the 22nd ACM ral news recommendation with multi-head
selfSIGKDD international conference on knowledge attention, in: Proceedings of the 2019
Conferdiscovery and data mining, 2016, pp. 353–362. ence on Empirical Methods in Natural Language
[25] Y. Cao, X. Wang, X. He, Z. Hu, T.-S. Chua, Uni- Processing and the 9th International Joint
Confying knowledge graph learning and recommen- ference on Natural Language Processing
(EMNLPdation: Towards a better understanding of user IJCNLP), Association for Computational
Linguispreferences, in: The world wide web conference, tics, Hong Kong, China, 2019, pp. 6389–6394. URL:
2019, pp. 151–161. https://aclanthology.org/D19-1671. doi:10.18653/
[26] B. Hu, C. Shi, W. X. Zhao, P. S. Yu, Leveraging meta- v1/D19-1671.
        </p>
        <p>path based context for top-n recommendation with [35] S. Okura, Y. Tagami, S. Ono, A. Tajima,
Embeddinga neural co-attention model, in: Proceedings of based news recommendation for millions of users,
the 24th ACM SIGKDD International Conference in: Proceedings of the 23rd ACM SIGKDD
internaon Knowledge Discovery &amp; Data Mining, 2018, pp. tional conference on knowledge discovery and data
1531–1540. mining, 2017, pp. 1933–1942.
[27] J. Jin, J. Qin, Y. Fang, K. Du, W. Zhang, Y. Yu, [36] J. Lian, F. Zhang, X. Xie, G. Sun, Towards better
repZ. Zhang, A. J. Smola, An eficient neighborhood- resentation learning for personalized news
recombased interaction model for recommendation on mendation: a multi-channel deep fusion approach.,
heterogeneous graph, in: Proceedings of the 26th in: IJCAI, 2018, pp. 3805–3811.</p>
        <p>ACM SIGKDD International Conference on Knowl- [37] C. Wu, F. Wu, M. An, J. Huang, Y. Huang, X. Xie,
edge Discovery &amp; Data Mining, 2020, pp. 75–84. Npa: neural news recommendation with
personal[28] X. Yu, X. Ren, Y. Sun, Q. Gu, B. Sturt, U. Khandel- ized attention, in: Proceedings of the 25th ACM
wal, B. Norick, J. Han, Personalized entity recom- SIGKDD international conference on knowledge
mendation: A heterogeneous information network discovery &amp; data mining, 2019, pp. 2576–2584.
approach, in: Proceedings of the 7th ACM interna- [38] D. Liu, J. Lian, S. Wang, Y. Qiao, J.-H. Chen, G. Sun,
tional conference on Web search and data mining, X. Xie, Kred: Knowledge-aware document
represen2014, pp. 283–292. tation for news recommendations, in: Fourteenth</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>A. Training Details</title>
      <sec id="sec-3-1">
        <title>All experiments were run on 8 NVIDIA A100 GPUs</title>
        <p>with batch size 8192 × 8 for KMPN/CKMPN and 4 × 8
for NRMS-BERT. Adam [43] is used to optimize models.
KMPN/CKMPN is trained for 2000 epochs with linearly
decayed learning rates from 10−3 to 0 for
Amazon-BookExtended and 5 × 10−4 to 0 for Movie-KG-Dataset.
NRMSBERT is trained for 10 epochs at a constant learning rate
of 10−4.</p>
        <p>We follow the oficial repository of KGAT 7 for BPRMF,
CKE, and KGAT training.</p>
        <p>Codes and pre-trained models are released in
https://github.com/LinWeizheDragon/Content-AwareKnowledge-Enhanced-Meta-Preference-Networks-forRecommendation.</p>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          <source>ACM Conference on Recommender Systems</source>
          ,
          <year>2020</year>
          ,
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          pp.
          <fpage>200</fpage>
          -
          <lpage>209</lpage>
          . [39]
          <string-name>
            <given-names>H.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Zhang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Xie</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Guo</surname>
          </string-name>
          , Dkn: Deep
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          tion,
          <source>in: Proceedings of the 2018 world wide web</source>
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          <string-name>
            <surname>conference</surname>
          </string-name>
          ,
          <year>2018</year>
          , pp.
          <fpage>1835</fpage>
          -
          <lpage>1844</lpage>
          . [40]
          <string-name>
            <given-names>W.</given-names>
            <surname>Krichene</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Rendle</surname>
          </string-name>
          ,
          <article-title>On sampled metrics for item</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          recommendation,
          <source>in: Proceedings of the 26th ACM</source>
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          <string-name>
            <surname>Discovery</surname>
          </string-name>
          &amp;
          <string-name>
            <surname>Data Mining</surname>
          </string-name>
          ,
          <year>2020</year>
          , pp.
          <fpage>1748</fpage>
          -
          <lpage>1757</lpage>
          . [41]
          <string-name>
            <given-names>S.</given-names>
            <surname>Rendle</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Gantner</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Freudenthaler</surname>
          </string-name>
          , L. Schmidt-
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          <source>of the 34th International ACM SIGIR Confer-</source>
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          <string-name>
            <surname>tion</surname>
            <given-names>Retrieval</given-names>
          </string-name>
          , SIGIR '11,
          <string-name>
            <surname>Association for</surname>
          </string-name>
          Com-
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          <string-name>
            <surname>puting Machinery</surname>
          </string-name>
          , New York, NY, USA,
          <year>2011</year>
          ,
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          p.
          <fpage>635</fpage>
          -
          <lpage>644</lpage>
          . URL: https://doi.org/10.1145/2009916.
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          2010002. doi:
          <volume>10</volume>
          .1145/2009916.2010002. [42]
          <string-name>
            <given-names>W.</given-names>
            <surname>Lin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Shou</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Gong</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Jian</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          <article-title>content-aware collaborative filtering</article-title>
          , in: 4th Edi-
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          <year>2022</year>
          ,
          <year>2022</year>
          . [43]
          <string-name>
            <given-names>D. P.</given-names>
            <surname>Kingma</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Ba</surname>
          </string-name>
          ,
          <article-title>Adam: A method for stochastic</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          optimization, in: Y. Bengio, Y. LeCun (Eds.),
          <fpage>3rd</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          <string-name>
            <surname>tions</surname>
          </string-name>
          ,
          <source>ICLR</source>
          <year>2015</year>
          , San Diego, CA, USA, May 7-9,
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          2015, Conference Track Proceedings,
          <year>2015</year>
          . URL:
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>http://arxiv.org/abs/1412.6980.</mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>