<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <issn pub-type="ppub">1613-0073</issn>
    </journal-meta>
    <article-meta>
      <article-id pub-id-type="doi">10.1145/3292500.3330961</article-id>
      <title-group>
        <article-title>mender System: A Hierarchical Graph Attention Approach</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Dong Li</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Divya Bhargavi</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Vidya Sagar Ravipati</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="editor">
          <string-name>Recommender System, Graph Neural Network, Multi-behavior</string-name>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Amazon.com</institution>
          ,
          <country country="US">USA</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Kent State University</institution>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>Workshop Proce dings</institution>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2019</year>
      </pub-date>
      <volume>30</volume>
      <fpage>793</fpage>
      <lpage>803</lpage>
      <abstract>
        <p>While recommender systems have significantly benefited from implicit feedback, they have often missed the nuances of multi-behavior interactions between users and items. Historically, these systems either amalgamated all behaviors, such as impression (formerly view), add-to-cart, and buy, under a singular 'interaction' label, or prioritized only the target behavior, often the buy action, discarding valuable auxiliary signals. Although recent advancements tried addressing this simplification, they primarily gravitated towards optimizing the target behavior alone, battling with data scarcity. Additionally, they tended to bypass the nuanced hierarchy intrinsic to behaviors. To bridge these gaps, we introduce the Hierarchical Multi-behavior Graph Attention Network (HMGN). This pioneering framework leverages attention mechanisms to discern information from both inter and intra-behaviors while employing a multi-task Hierarchical Bayesian Personalized Ranking (HBPR) for optimization. Recognizing the need for scalability, our approach integrates a specialized multi-behavior sub-graph sampling technique. Moreover, the adaptability of HMGN allows for the seamless inclusion of knowledge metadata and time-series data. Empirical results attest to our model's prowess, registering a notable performance boost of up to 64% in NDCG@100 metrics over conventional graph neural network methods.</p>
      </abstract>
      <kwd-group>
        <kwd>Approach</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>CEUR
ceur-ws.org
A</p>
    </sec>
    <sec id="sec-2">
      <title>1. Introduction</title>
      <p>Recommender systems are widely employed in online
platforms to provide accurate and relevant content to
implicit feedback has become widely adopted [1, 2, 3, 4,
5, 6, 7], where the user-item relationship is classified as
either interacted or unknown. However, modern online
shopping platforms involve various types of interactions,
such as clicks, add-to-cart actions, and buy. Relying solely
on binary implicit feedback information (interacted or
unknown) while overlooking this rich multi-behavior
information can worsen the cold start problem and
exacerbate data sparsity issues [8]. Furthermore, considering
the naturally existing multi-behavior information in the
modern e-commerce ecosystem, it becomes beneficial to
diferentiate between various user behaviors and
optimize for target behaviors, such as buy, which align with
the ultimate goal of maximizing revenue for businesses.</p>
      <p>Recent years have witnessed the rapid development of
multi-behavior recommender systems [9, 10, 11, 12, 13, 8,
14, 15], due to their ability to supplement and motivate
sparse target behavior signals (buy) with auxiliary
behaviors (view or what’s often termed ”impression”, favorite,
add-to-cart, etc). Despite these eforts, there exist
drawbacks that prohibit these works from fully exploiting the</p>
      <p>Failure to exhaustively utilize and predict
auxiliary behaviors. [8] constructed a heterogeneous graph
convolutional neural network for the multi-behavior
recommendation and apply Bayesian Personalized Ranking
(BPR), where positive and negative pairs are sampled
according to the target behavior. [12] utilized the graph
attention layer together with the item knowledge graph
and temporal encoding to empower the learned
embeddings to be expressive. All these multi-behavior
models [8, 14, 12] merely optimize for single-behavior
(target behavior) despite utilizing multi-behavior
information during the model learning stage. Besides, these
approaches are unable to predict other auxiliary behaviors
thus still sufering from sub-optimal performance and
label-sparsity issues. Overlooking the hierarchical
pattern of multi-behaviors. [16, 13] took the backbone
of heterogeneous graph convolutional neural network
and applied binary Mean Square Error to all behaviors
CEUR
htp:/ceur-ws.org
ISN1613-073</p>
      <p>CEUR
Workshop on Learning and Evaluating Recommendations with Impres- predictions. Despite its eficiency, we argue this method
†This work was completed during the author’s internship at Amazon.
two diferent levels of behaviors. These methods did
© 2023 Copyright for this paper by its authors. Use permitted under Creative Commons License not consider the case where an item would innocently
Attribution 4.0 International (CC BY 4.0).
behavior scenarios. This multi-task optimization
framework is crucial for the behavior prediction
task.
3. To scale our model, we extend sub-graph sampling
strategy to multi-behavior scenario, avoiding bias
in behavior sampling. The corresponding
pairwise negative sampling for HBPR loss is
refurbished in the sub-graph stage.
4. Extensive empirical experiments are conducted on
two practically processed e-commerce datasets
(Taobao and RetailRocket ). The results show that
our framework achieves significant improvement
over state-of-the-art models (SOTA) (35% in
RetailRocket and 64% in Taobao) in terms of ofline
metric ( @100 ) on target-behavior
prediction task.
dation problem; in section 3, we detailed clarify the
proposed frameworks including model architecture and
optimization; in section 4.1, we introduce the sub-graph
serve as a negative sample while belonging to a higher sampling as well as incorporation of temporal and
knowlimportance behavior group. More specifically, a user edge information for multi-behavior recommendation; in
views and buys an item directly (without add-to-cart ), section 5, we conduct extensive empirical experiments;
would possibly serve as a negative sample when opti- in section 6, we briefly introduce the related works and
mizing for add-to-cart behavior, which is contrary to the discuss the resemblance and discrepancy from ours and
hierarchical pattern between multi-behaviors. in section 7, we summarize and conclude the paper.</p>
      <sec id="sec-2-1">
        <title>To tackle these issues, we aim to leverage the graph</title>
        <p>attention neural network which enables us to learn and
represent individual interaction and optimize a multi-task 2. Experiment Formulation
objective. The challenge is direct utilization of Bayesian
Personalized Ranking (BPR) for multi-behavior scenarios
can lead to some conflicts. We use an example in fig. 1 (d)
to clarify the situation.  3 add-to-cart the item  3 and buy
item  2. To optimize the behavior add-to-cart, the BPR
criterion would clarify  3 is positive and  2 is negative for
 3. We argue that this is not practical since the buy
behavior over  2 already demonstrates the user’s preference.</p>
      </sec>
      <sec id="sec-2-2">
        <title>Highly inspired by [9, 17], we propose a Hierarchical</title>
      </sec>
      <sec id="sec-2-3">
        <title>Bayesian Personalized Ranking (HBPR) optimization cri</title>
        <p>terion to deal with the multi-behavior case by taking the
hierarchical relations between behaviors into
consideration. We propose a multi-behavior ad-hoc graph
attention network, aka Hierarchical Multi-behavior Graph</p>
      </sec>
      <sec id="sec-2-4">
        <title>Attention Network (HMGN). We list the main contribu</title>
        <p>tions of this paper as below:
Ad Tech and Media customers always encounter label
sparsity for single-behavior recommendation systems
despite having rich auxiliary information. Incorportating
multiple behaviors can also help in the development of
a model that can predict a customer’s propensity across
diferent stages of purchase funnel. Additionally, our
customers express the need to create a robust
decisionmaking engine that can leverage side information, such
as item metadata, to enhance the recommendation
process. The work we present here represents a
preliminary step towards their long-term goal of building a
resilient pipeline for multi-behavior recommendation
systems based on Graph Neural Networks (GNNs).</p>
      </sec>
      <sec id="sec-2-5">
        <title>We first introduce the concept of a multi-behavior rec</title>
        <p>ommender system in graph terminology. User-Item
Multi-Behavior (Temporal) Bipartite Graph: the
user-item interactions can be treated as a heterogeneous
bipartite graph  = (ℰ ,  ) = {(,   , )| ∈  ,  ∈  ,   ∈
} where  is the set of all user nodes,  is the set of all
item nodes,   indicates user  interacted item  with
behavior   .  is the set of all multi-behavior
interaction edges between a user and an item (view, add-to-cart,</p>
      </sec>
      <sec id="sec-2-6">
        <title>1. We explore and benchmark two light yet efective</title>
        <p>graph attention neural network paradigms
targeting on multi-behavior recommendation. Besides
target behavior, our model is also able to predict
auxiliary behaviors.</p>
      </sec>
      <sec id="sec-2-7">
        <title>2. We propose HBPR loss criterion dedicated to multi</title>
        <p>favor, buy, etc for e-commerce datasets). If temporal in- 3.1. HMGN-intra
view (analogous to ”impression” in the context), add-to- (message propagating) phase is conducted on these them.
formation is to be considered, the dynamic graph can be
exist at most one edge per behavior type.
represented as:   = {(,   , ,   , )| ∈  ,  ∈  , 
where   is the timestamp when  interacted with  under
  . For each user-item node pair (, ) , there can only</p>
        <p>∈ }</p>
      </sec>
      <sec id="sec-2-8">
        <title>The goal of the multi-behavior recommender system</title>
        <p>is to utilize all the interaction information (including
cart, favor, buy, etc) to predict the possible target behavior,
typically buy behavior.</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>3. HMGN: Hierarchical</title>
      <p>Multi-Behavior Graph Attenion</p>
    </sec>
    <sec id="sec-4">
      <title>Network</title>
      <sec id="sec-4-1">
        <title>In this section, we elaborate on the details of our pro</title>
        <p>posed framework - Hierarchical Multi-Behavior Graph
Attention Network (HMGN) , the model illustration of
which is presented in Figure 2. As we formulated in
section 2, the multi-behavior recommendation data is
organized as a heterogeneous bipartite graph. This leads to
a diferent order of information propagation and
aggregation process in GNNs based on each behavior which we
explore in HMGN-intra and HMGN-inter.
Mathematically, we denote e
(0), e</p>
        <p>(0) as the initialized embedding for
user  and item  . After propagating through  layers of
the graph attention network, the output of the last layer
() would be treated as the final representation

(( () e
) ⋅ ( () e() ) ⋅ √1/ )</p>
        <p>target-oriented representation (e(,) ) learning. Since all
the procedure is operated and conducted in each isolated
() ) are
transsingle-behavior graph, the e
the perspective of each behavior  . The next step is to
aggregate each behavior representation.
(,) is representative from</p>
        <sec id="sec-4-1-1">
          <title>3.1.2. Aggregation Phase</title>
        </sec>
      </sec>
      <sec id="sec-4-2">
        <title>All behavior-specific information (for the target node) would be put together to form a final personalized representation. We are trying to keep the structure as light as possible to maintain the most efective and eficient com</title>
      </sec>
      <sec id="sec-4-3">
        <title>To this end, we take a weighted aggregation approach:</title>
        <p>ponents for the GNN as LightGCN [19] demonstrated. to set   ).</p>
        <p>e

(+1) = ∑  , ⋅ e(,)</p>
        <p>∈
In [13],  , is set to be manually determined hyper- nism to aggregate information from neighboring sources
(2)
(3)
 () e</p>
        <p>() is the personalized query in behavior  space
and { () e() } , { () e() } are the corresponding keys and</p>
        <p>values in each behavior space for item  . 
()
scale the weight that how much information the target
node ( ) obtained from diferent behavior spaces of the
neighborhood  node. And the obtained representation
e← contains all the behavior information that passed
from node  to node  . The next step is aggregating all
the information from its neighborhood nodes (belonging</p>
        <p>←(, ) tends to</p>
        <sec id="sec-4-3-1">
          <title>3.2.2. Aggregation Phase</title>
          <p>Unlike GCN [13] architecture, we use an attention
mechanodes to target nodes [11].</p>
          <p>3.3. Behavior Preference Modeling
After information propagating through  layers of GNN,
we obtain final representation
item  , respectively. In real-world multi-behavior
recommendation, a user would have a generalization impression
e

() and e
() for user  and
ization preference  , (view/add-to-cart-buy, etc). Inspired
by this joint relation:
ln   (, ) =
ln   (|) +
ln   ()
We propose the following formula that can capture both
user item generalization preference as well as behavior
specialization preference:
 (, , ) = (
e
() ) ((1 −  ) +  

)e()
(6)
where  is identity matrix and  is diagonal matrix
representing a behavior type  .  is a scalar hyperparameter
that balances the trade-of between generalization and
specialization. Noting that, if  = 1 , eq. (6) would be the
user-item inner product [3] - the most common strategy
to mimic user-item preference; if  = 0 , it would collapse
to the behavior specialization preference, similar to [13].
parameters and is shared between each user and item
(as   ). A drawback for this setting occurs when a node
doesn’t exhibit a specific behavior. It leads to an unstable
amplitude in e
(+1) with e</p>
          <p>(,) being a zero vector. With
our personalized average aggregation, when a user does
not exhibit the specific behavior  , their weight  , shall
be set as 0, and the aggregation operation is performed
over the remaining behaviors:
 , =
⎧0,
⎨ 
⎩</p>
          <p>1
∑ ( e
(,) ≠0)
e

e

(,)
(,)
= 0
≠ 0
where  is the indicator function, equal to 1 when input
is True and 0 for False.
3.2. HMGN-inter
In HMGN-inter framework, the message would first
pass and aggregate across all behaviors within each
userwhere diferent user/item can exhibit varying levels of
multiple behaviors. For example, a particular user can
perform more add-to-cart actions compared to buy.
Similarly, a particular item could experience more page-view
than it has been marked as favorite. Thus learning
useritem behavior involves understanding these complex
dynamics that are user/item centric.</p>
        </sec>
        <sec id="sec-4-3-2">
          <title>3.2.1. Inter-behavior Phase</title>
          <p>For each user-item pair, the cross attention attempts to
maximize the optimal behavior representation for the
target node (user  3 from fig. 2(a)).</p>
          <p>∑
∈
∈
item node pair. We motivate this from real-world patterns  over an item (like/dislike) and also the behavior
special3.4. HBPR: Hierarchical Bayesian</p>
          <p>Personalized Ranking
While recent works [8, 12, 11] utilize multi-behavior
information in modeling phase, they solely optimize for
target behavior (treating only target behavior as positive
label) in their objective. We argue that this single-task
optimization framework won’t be suficient to exploit
the power of auxiliary behavior information. And some
multi-task optimization [13] ignores the relationship
between the behaviors. Thus, we propose a multi-task
optimization criterion - Hierarchical Bayesian Personalized</p>
        </sec>
      </sec>
      <sec id="sec-4-4">
        <title>Ranking (HBPR).</title>
        <p>The behavior-specific personalized formalization is
an extension of [7] and is defined as
 &gt; ,  states as user  prefer item  than item  under the
behavior  . In multi-behavior scenario, we maintain &gt;,
with the properties of totality, antisymmetry, transitivity
(see [7]) , extended it with extra property - hierarchy:
&gt;, ⊂  2, where
∀,  ∈  ,</p>
        <p>1,  2 ∈  ∶
 &gt; , 1  ∧  &gt; , 2  ∧  1 &gt;  2 ⇒  &gt; , 1
,  = , 2 
behaviors.
where &gt; is the heuristically defined priority rank across</p>
      </sec>
      <sec id="sec-4-5">
        <title>Intuition of hierarchy: Using the same example we</title>
        <p>talked in fig. 1 in section 1:  3 add-to-cart the item  3 and
buy item  2. We have  2 &gt; 3,
 3 and  3 &gt; 3, −−
Without defining of hierarchy, the BPR criterion would
classify  3 is positive and  2 is negative for  3 under
be 2.
havior add-to-cart,  3 &gt; 3, −−
set as  buy &gt;  add-to-cart &gt;  view.
practical since the typical case is a person click/view an
item, favor it or add-to-cart, then buy it. Under such
a hierarchical assumption, as long as the user buy an
item, the preference over other behaviors is supposed
to be automatically claimed. Following this logic, the
e-commerce shopping behaviors priority rank &gt; can be
 2. This is not quite
 + = { 0,  1, ⋯ ,  −1 }.</p>
        <p>Definition 1 (Higher priority rank set  +). Given a
behaviors priority rank:  0 &gt; &gt;  1 &gt;
 ⋯ &gt;
  . The

higher priority rank set of   (0 ≤  ≤ ) is denoted as
,
For convenience, we define  + ⊂  as the set of all items
that user  interacted with behavior  and  ,−
as non-interacted or unknown in contrast.</p>
        <p>=  \ ,
Obviously,
+
 +
,
∩  ,− = ∅ and  ,+</p>
        <p>∪  ,− =  .</p>
        <p>+ can be treat as a positive sample for
user  with the behavior  , while the negative one is not</p>
        <p>An item from  ,
of hierarchy.
directly from  ,− =  \ ,</p>
        <p>+ according to the requirements</p>
        <p>Equipped with this definition of behavior higher
priority rank set, we can determine the negative item sets
compatible with hierarchy principle:

,
,− = { ∈ 
users and items. Therefore, the scalability of
recommendation models is crucial for production deployments.</p>
      </sec>
      <sec id="sec-4-6">
        <title>While recent works have exhibited impressive success</title>
        <p>in homogeneous sub-graph sampling strategies
(singlebehavior) [20, 21, 22, 23], there is a dearth of literature
in the heterogeneous (multi-behavior) sampling realm
graph kernel (the inner circle in Figure 3). Then, we
sub-graph sampling; (d)”add-to-cart” sub-graph sampling (e)final sub-graph. For sub-graph BPR optimization, positive
useritem pairs are selected from the kernel of sub-graph and negative pairs come from the entire sub-graph.
sample up to  -hop neighborhood nodes for these kernel
nodes. In the HBPR optimization stage, where a pair of
nodes (positive and negative item samples for a user) are
required, we constrain the positive sample to be derived
from our core sub-graph kernel and the negative sample
to be derived from rest of the sub-graph. (Algorithm 1)</p>
        <p>The justification of why sub-graph kernel is crucial
for optimization is as follows. In the original full-size
graph, nodes and edges are self-contained (connected),
and every node can receive information from all of its
neighborhoods during the graph propagation stage. In
the case of sub-graph, since only few nodes are kept,
there is an unavoidable loss of knowledge during the
information propagation. Especially for those nodes in the
”border” of the sub-graph in Figure 3, they are left with
spare neighboring. In contrast, for those nodes in the
kernel, almost all the neighbors (dense neighborhoods)
would be kept due to our layer-wise sampling strategy.</p>
      </sec>
      <sec id="sec-4-7">
        <title>This leads to a more accurate and informative embedding representation for the nodes in the kernel. Thus it would be beneficial to only optimize the kernel edges as positive relations.</title>
        <p>4.2. Temporal Encoding
latent dimension of the model):
Temporal information can be seamlessly incorporated
into our framework in a way similar to positional
encoding in transformer[18, 12]. Formally, we define the
temporal representation as a vector    ∈ ℝ ( is the
  ,2 = sin(
2 )   ,2+1 = cos(</p>
        <p>2+1 )
and aggregation is in-batch operated by adjacency
(Lapla(,) is used to
subcian) matrix.
4.3. Knowledge Graph (KG) Enhancing</p>
      </sec>
      <sec id="sec-4-8">
        <title>We are also able to enhance and boost our framework with metadata. Similar to [11, 26], we leverage the itemmeta data and train a separate KG loss.</title>
        <p>Collaborative Knowledge Graph (CKG) can be
seamlessly extended from single-behavior recommendation to
multi-behavior recommendation. User-item interactions
can be largely divided into multiple triples, (, , )
for
example, (,  , )
data, we can also define
, (,   , )</p>
        <p>, etc. As to the item-meta
(,  , )
where  ∈ ℛ is the
relation and  ∈ ℰ is an entity (item feature), for example,
(,  ,   )
as {(ℎ,  , )|ℎ,  ∈</p>
        <p>. To this end, the CKG is defined
⋃  ⋃ ℰ ,  ∈ 
⋃ ℛ} The translation
principle which optimizes KG by projecting relations
and entities into a common semantic space, [26] is given
by:   ℎ + e ≈     , where eℎ and     are the
representation of the head and tail nodes in relation  space. The
scoring function  is thus given by
(ℎ,  , ) = ||
  ℎ +   −     ||
(12)</p>
      </sec>
      <sec id="sec-4-9">
        <title>We optimize a separate BPR criterion powered KG scoring objective:</title>
        <p>ℒ 
= ∑ − log  ((ℎ,  ,  ′) − (ℎ,  , ) )
(13)
Combine Equation (8):
ℒ ′ = ℒ + ℒ 
 +
,

,−
,</p>
        <p>← { ∈ 
for all  ∈ 
 ← 0
repeat
← {|(, , ) ∈ 

}
 ,− ← { ∈   | ∈  \ ,+ }
training set   = {(, , , )}
INPUT: kernel users set   ⊂ 
OUTPUT: sub-graph   = {(,   , )| ∈  ,  ∈  }
,
1: for all  ∈ ℬ</p>
        <p>do
construct kernel items set 
construct single-behavior graph  
  from one-hop
neighsampling and obtain single-behavior sub-graph
5: end for
7: sub-graph kernel  
6: sub-graph   ← {(,   , )| ∈
⋃  

 ,  ∈ ⋃    }</p>
        <p>← {(,   , )| ∈   ,  ∈  ∈
8: Empty dataset   ← {}
9: for all  ∈   do
for all  ∈</p>
        <p>do
end for
end for
23: end for</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>5. Experiment</title>
      <sec id="sec-5-1">
        <title>In this section, we experimentally investigate the performance of the proposed framework on two real world multi-behavior datasets. Specifically, we aim to answer the following research questions:</title>
        <p>(RQ1): How do our proposed models: HMGN-intra and
HMGN-inter perform compared to the SOTA baselines
[12, 13] in terms of target-behavior prediction?
(RQ2): How does the proposed hierarchical multi-task
learning perform compared to single-task learning in
terms of target-behavior prediction? And how does it
perform compared to other multi-task learning in terms
of multi-behavior predictions?
(RQ3): How do our models perform when trained on
sampled sub-graph vs non-sampling full graph?
(RQ4): How would diferent gadgets (temporal encoding,
knowledge-graph) contribute to the performance?
2:
3:
4:
10:
11:
12:
13:
14:
15:
16:
17:
18:
19:
20:
21:
22:
tral). However, in practical circumstances, a user cannot
exhibit these behaviors simultaneously. Therefore, we
refrain from using them for our experiments.
and  
1 (with view, add-to-cart, buy
2 (with view,</p>
        <sec id="sec-5-1-1">
          <title>5.1.1. Dataset Processing</title>
          <p>is treated as training set, data in Dec-02-2017 is the
validation set and Dec-03-2017 is the test set. Similarly in
dataset, we split data according to the
following timelines: May-03-2015 to Aug-15-2015 as a
training set, Aug-16-2015 to Sep-01-2015 as the
validation set and Sep-01-2015 to Sep-15-2015 as a test set. For
reproducibility, we would share the details of the data
processing as well as the models in the code.</p>
        </sec>
        <sec id="sec-5-1-2">
          <title>5.1.2. Dataset Statistics</title>
        </sec>
      </sec>
      <sec id="sec-5-2">
        <title>We present the processed dataset statistic in Table 1</title>
        <p>5.2. Experimental Settings</p>
        <sec id="sec-5-2-1">
          <title>5.2.1. Evaluation Metrics</title>
          <p>for a individual user  :</p>
          <p>and
( = 10, 50, 100 ) as our metrics, specifically
∑  (() ≤  )
( , |
∈  
∑
∈ 
 |)

 (() ≤  )
 log(() + 1)</p>
          <p>(14)
1https://tianchi.aliyun.com/dataset/649
2https://www.kaggle.com/datasets/retailrocket/ecommerce-dataset
Performance of the HMGN models w.r.t baselines on RetailRocket dataset. Mean results are present by repeating 5 times.</p>
          <p>NDCG@10</p>
          <p>NDCG@50</p>
          <p>NDCG@100</p>
          <p>RECALL@10</p>
          <p>RECALL@50</p>
          <p>RECALL@100
. Here
• KHGT [12] is a SOTA model that targets on
multi [29, 30, 31, 32].</p>
          <p>cial cases when |
5.2.2. Baselines
 is an indicator function and ()</p>
          <p>is the rank of item
user  in test set. It is worth noting that in some
spe</p>
          <p>is the set of positive items for
 | &gt;  (</p>
          <p>1,  2) and  1 ≥  2.,
2 could happen. This is because
while the denominator could dominate the value.
that numerator wouldn’t change much (from  1 to  2),</p>
        </sec>
      </sec>
      <sec id="sec-5-3">
        <title>We compare the proposed model with several influential recommendation models including both single-behavior and multi-behavior ones.</title>
        <p>Single-behavior Models:
• itemKNN [33] A classical and robust neighborhood
method.</p>
        <p>objectives.
• BPR [7] is one of the most popular methods in
recommendation which learns to optimize pair-wise
• LightGCN [19] is a graph-based model which
simplifies the framework of GCN for
recommendation by removing feature transformation and
nonlinear activation.</p>
        <p>Multi-behavior Models:
• LightGCN-M We extend LightGCN for multi-task
optimization by integrating multi-behaviors
during the modeling stage.
behavior recommendation task enhanced with
knowledge and temporal information.
• KGAT [11] attention-based graph neural network
that incorporates the item meta-data. Here, we
extend and optimize it with multi-task learning.
• GHCF [13] a SOTA model particularly for
multibehavior collaborative filtering without sampling
any negative items by optimizing a mean square
error (MSE) loss.
5.3. Performance Comparison (RQ1)
indicating that behavior-specific learning is more
efective than cross-behavior learning for
multibehavior recommendations. In the HMGN-intra
model, all information propagates first in the
isolated single-behavior graph and then gathers
together. Considering the multi-task objectives,
these would better help explain and contribute to
each individual behavior prediction, leading to a
superior performance.
• most all the multi-behavior models exhibit better
performance compared to single-behavior ones,
this emphasizes the importance of multi-behavior
utilization.
5.4. HBPR Multi-task Optimization (Q2)</p>
      </sec>
      <sec id="sec-5-4">
        <title>We empirically testify the HBPR based multi-task opti</title>
        <p>mization framework from two perspectives:</p>
      </sec>
      <sec id="sec-5-5">
        <title>1. Speciality. Comparing the performance of Light</title>
        <p>GCN with its counterpart LightGCN-M, and base model
(HMGN-intra optimized by HBPR) with its counterpart
 . −  in Figure 4, we would see that the
multitask optimization is crucial for multi-behavior
recommender system. In other words, even if the end goal of
the model is to predict one target behavior, learning to
explicitly optimize auxiliary behaviors in model objective
significantly improves the model metrics.</p>
      </sec>
      <sec id="sec-5-6">
        <title>2. Expressiveness. We also compare our proposed</title>
        <p>model with GHCF on the ability to predict other auxiliary
behaviors. Figure 5 indicate our method consistently
outperforms GHCF on all behaviors prediction in terms
of   metric.
5.5. Sub-graph Sampling (Q3)
To test the scalability of our graph, we utilized the
algorithms that are proposed in Section 4.1. We assign
diferent sampling size and obtain the performance for
each setting in Table 4. (Since the sub-graph size is
dependent on the sampling size, we repeat this experiment 100
times to compute an average sub-graph size.) Note that
for both the datasets, a sub-graph size of around 20 , can
already get pretty good results compared to the full-size
5.6. Study of Graph Enhancement (Q4)
In this section, we investigate how diferent graph
enhancing gadgets afect our best performing HMGN-intra
model. From Figure 6, we have two key takeaways:
• In   dataset (Figure 6(a)), HMGN-intra
with temporal encoding performs worse than
that without temporal information whereas in
 dataset (Figure 6(b)) HMGN-intra
with temporal encoding included gets better
performance than that without temporal encoding.
We think this contradictory efect is caused by
the duration of the data (Figure 6(c)). In   ,
all data is collected within 9 days which all can
be considered as ”recent” events. In this case, the
fusion of temporal encoding would behave just
like a noise injection that decreases the
performance. In  , a long period dataset, the
data lasts 3.5 months, which would enable the
temporal encoding to take efect.
to exploit robust and powerful GNN. We reason such
techniques as graph enrichment and augmentation. KGAT
[11] integrates the KG and attention mechanism for
single behavior recommendations systems where the
objectives are consist of with both KG [26] and collaborative
ifltering optimization. SGL [ 44] generates multiple views
of graphs by node and edge drop and applies InfoNCE
[45] onto that. SimGCL [46]
6.3. Resemblance and Discrepancy</p>
        <p>In this paper, we devised two new graph attention-based
frameworks called HMGN-intra and HMGN-inter for
multi-behavior recommender systems. We discover that
it is crucial for multi-behavior systems to learn via
multitask objectives, aka, optimizing for all behaviors instead
of target behaviors. To this end, we propose a
hierarchical Bayesian Personalized Ranking optimization criterion.</p>
      </sec>
      <sec id="sec-5-7">
        <title>We enable our model with the ability to capture prefer</title>
        <p>ence generalization as well as behavior specialization and
6.2. Graph Enrichment and Augmentation to predict all types of behaviors. Further, we provide a
With the ability to learn high-order topological informa- unified and comprehensive strategy for multi-behavior
tion, Graph Neural Network (GNN) [38, 39] has achieved methods. Specifically, we enhance our model with
mulsignificant success in Recommender System [ 19, 40, 41]. tiple techniques like temporal encoding and metadata
Various techniques like optimizing Contrastive Learning information. To scale up the model, we also extend the
(CL) [42], leveraging metadata or Knowledge Graph (KG) sub-graph sampling to the multi-behavior scenarios.
Ex[11], incorporating temporal information [39], sampling tensive empirical analyses indicate our proposed model
sub-graph [43] for scalability, are exhibiting advantages outperforms the baselines significantly.
963770.963776. doi:10.1145/963770.963776. ’22, Association for Computing Machinery, New
[34] F. Xiao, L. Li, W. Xu, J. Zhao, X. Yang, J. Lang, York, NY, USA, 2022, p. 1273–1282. URL: https:
H. Wang, Dmbgn: Deep multi-behavior graph //doi.org/10.1145/3477495.3532014. doi:10.1145/
networks for voucher redemption rate prediction, 3477495.3532014.
in: Proceedings of the 27th ACM SIGKDD Confer- [42] C. Huang, X. Wang, X. He, D. Yin, Self-supervised
ence on Knowledge Discovery &amp;amp; Data Mining, learning for recommender system, in:
ProceedKDD ’21, Association for Computing Machinery, ings of the 45th International ACM SIGIR
ConNew York, NY, USA, 2021, p. 3786–3794. URL: https: ference on Research and Development in
Infor//doi.org/10.1145/3447548.3467191. doi:10.1145/ mation Retrieval, SIGIR ’22, Association for
Com3447548.3467191. puting Machinery, New York, NY, USA, 2022, p.
[35] C. Gao, X. He, D. Gan, X. Chen, F. Feng, 3440–3443. URL: https://doi.org/10.1145/3477495.</p>
        <p>Y. Li, T. Chua, D. Jin, Neural multi-task rec- 3532684. doi:10.1145/3477495.3532684.
ommendation from multi-behavior data, in: [43] R. Ying, R. He, K. Chen, P. Eksombatchai, W. L.
2019 IEEE 35th International Conference on Hamilton, J. Leskovec, Graph convolutional
neuData Engineering (ICDE), IEEE Computer Soci- ral networks for web-scale recommender systems,
ety, Los Alamitos, CA, USA, 2019, pp. 1554–1557. in: Proceedings of the 24th ACM SIGKDD
InURL: https://doi.ieeecomputersociety.org/10.1109/ ternational Conference on Knowledge Discovery
ICDE.2019.00140. doi:10.1109/ICDE.2019.00140. &amp; Data Mining, KDD18, Association for
Com[36] Y. Wu, R. Xie, Y. Zhu, X. Ao, X. Chen, puting Machinery, New York, NY, USA, 2018,
X. Zhang, F. Zhuang, L. Lin, Q. He, Multi- p. 974–983. URL: https://doi.org/10.1145/3219819.
view multi-behavior contrastive learning in rec- 3219890. doi:10.1145/3219819.3219890.
ommendation, in: Database Systems for Ad- [44] J. Wu, X. Wang, F. Feng, X. He, L. Chen, J. Lian,
vanced Applications: 27th International Confer- X. Xie, Self-supervised graph learning for
recomence, DASFAA 2022, Virtual Event, April 11–14, mendation, in: F. Diaz, C. Shah, T. Suel, P. Castells,
2022, Proceedings, Part II, Springer-Verlag, Berlin, R. Jones, T. Sakai (Eds.), SIGIR ’21: The 44th
InHeidelberg, 2022, p. 166–182. URL: https://doi. ternational ACM SIGIR Conference on Research
org/10.1007/978-3-031-00126-0_11. doi:10.1007/ and Development in Information Retrieval,
Vir978-3-031-00126-0_11. tual Event, Canada, July 11-15, 2021, ACM, 2021,
[37] L. Xia, C. Huang, Y. Xu, J. Pei, Multi-behavior pp. 726–735. URL: https://doi.org/10.1145/3404835.
sequential recommendation with temporal graph 3462862. doi:10.1145/3404835.3462862.
transformer, IEEE Transactions on Knowledge [45] M. Gutmann, A. Hyvärinen, Noise-contrastive
esand Data Engineering 35 (2023) 6099–6112. doi:10. timation: A new estimation principle for
unnor1109/TKDE.2022.3175094. malized statistical models, in: Y. W. Teh, M.
Titter[38] X. Li, D. Li, R. Jin, R. Ramnath, G. Agrawal, Deep ington (Eds.), Proceedings of the Thirteenth
Intergraph clustering with random-walk based scalable national Conference on Artificial Intelligence and
learning, in: 2022 IEEE/ACM International Confer- Statistics, volume 9 of Proceedings of Machine
Learnence on Advances in Social Networks Analysis and ing Research, PMLR, Chia Laguna Resort, Sardinia,
Mining (ASONAM), 2022, pp. 88–95. doi:10.1109/ Italy, 2010, pp. 297–304. URL: https://proceedings.</p>
        <p>ASONAM55673.2022.10068646. mlr.press/v9/gutmann10a.html.
[39] da Xu, chuanwei ruan, evren korpeoglu, sushant [46] J. Yu, H. Yin, X. Xia, T. Chen, L. Cui, Q. V. H.
kumar, kannan achan, Inductive representation Nguyen, Are graph augmentations necessary?
learning on temporal graphs, in: International Con- simple graph contrastive learning for
recommenference on Learning Representations, 2020. URL: dation, SIGIR ’22, Association for
Computhttps://openreview.net/forum?id=rJeW1yHYwH. ing Machinery, New York, NY, USA, 2022, p.
[40] X. Wang, X. He, M. Wang, F. Feng, T. Chua, Neural 1294–1303. URL: https://doi.org/10.1145/3477495.
graph collaborative filtering, in: Proceedings of the 3531937. doi:10.1145/3477495.3531937.
42nd International ACM SIGIR Conference on
Research and Development in Information Retrieval,
SIGIR 2019, Paris, France, July 21-25, 2019., 2019,
pp. 165–174.
[41] S. Peng, K. Sugiyama, T. Mine, Less is more:</p>
      </sec>
      <sec id="sec-5-8">
        <title>Reweighting important spectral graph features for</title>
        <p>recommendation, in: Proceedings of the 45th
International ACM SIGIR Conference on Research
and Development in Information Retrieval, SIGIR</p>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list />
  </back>
</article>