<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Hierarchical Multi-Task Learning Framework for Session-based Recommendations</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Sejoon Oh</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Walid Shalaby</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Amir Afsharinejad</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Xiquan Cui</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Georgia Institute of Technology</institution>
          ,
          <addr-line>Atlanta, Georgia</addr-line>
          ,
          <country country="US">USA</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>The Home Depot</institution>
          ,
          <addr-line>Atlanta, Georgia</addr-line>
          ,
          <country country="US">USA</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>While session-based recommender systems (SBRSs) have shown superior recommendation performance, multi-task learning (MTL) has been adopted by SBRSs to enhance their prediction accuracy and generalizability further. Hierarchical MTL (H-MTL) sets a hierarchical structure between prediction tasks and feeds outputs from auxiliary tasks to main tasks. This hierarchy leads to richer input features for main tasks and higher interpretability of predictions, compared to existing MTL frameworks. However, the H-MTL framework has not been investigated in SBRSs yet. In this paper, we propose HierSRec which incorporates the H-MTL architecture into SBRSs. HierSRec encodes a given session with a metadata-aware Transformer and performs next-category prediction (i.e., auxiliary task) with the session encoding. Next, HierSRec conducts next-item prediction (i.e., main task) with the category prediction result and session encoding. For scalable inference, HierSRec creates a compact set of candidate items (e.g., 4% of total items) per test example using the category prediction. Experiments show that HierSRec outperforms existing SBRSs as per next-item prediction accuracy on two session-based recommendation datasets. The accuracy of HierSRec measured with the carefully-curated candidate items aligns with the accuracy of HierSRec calculated with all items, which validates the usefulness of our candidate generation scheme via H-MTL.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;Session-based Recommendation</kwd>
        <kwd>Hierarchical Multi-task Learning</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>
        Problem Description and Motivation. Multi-task learning (MTL) [
        <xref ref-type="bibr" rid="ref1 ref2 ref3 ref4 ref5">1, 2, 3, 4, 5</xref>
        ] has been
employed to enhance the accuracy of existing recommender systems. MTL prevents the
overiftting of a model via sharing parameters between multiple prediction tasks [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ]. Hierarchical
MTL (H-MTL) [
        <xref ref-type="bibr" rid="ref10 ref7 ref8 ref9">7, 8, 9, 10</xref>
        ] further improves the MTL by exploiting predictions from other tasks
as another task’s input in a hierarchical order. For example, the main task (e.g., next-item
prediction) can use outputs from auxiliary tasks (e.g., next-category prediction) as input features
to its prediction model. Those additional features can serve as rich external knowledge and
enhance the performance of the main task.
      </p>
      <p>
        While such an H-MTL framework gives an implicit data augmentation efect and higher
generalization capability to a machine learning model [
        <xref ref-type="bibr" rid="ref10 ref8">8, 10</xref>
        ], its application to session-based
recommender systems (SBRSs) [
        <xref ref-type="bibr" rid="ref11 ref12 ref13 ref14 ref15 ref16 ref17 ref18 ref19">11, 12, 13, 14, 15, 16, 17, 18, 19</xref>
        ] has not been investigated yet.
SBRSs have gained attention as they capture the latest and evolving interests of a user in a
session, where a session consists of a sequence of user-item interactions occurring within a
short period [
        <xref ref-type="bibr" rid="ref20">20</xref>
        ]. The H-MTL architecture will be beneficial and important to SBRSs as per not
only prediction accuracy [
        <xref ref-type="bibr" rid="ref21">21</xref>
        ] but also interpretability [
        <xref ref-type="bibr" rid="ref22">22</xref>
        ]. Let us assume a user is searching
for a product to add to the current session in an e-commerce platform. With the H-MTL,
we not only enhance the recommendation quality to users by leveraging prior knowledge
obtained from auxiliary tasks but also provide meaningful and reasonable explanations of the
recommendations to users by interpreting outputs from auxiliary tasks (e.g., top-K predicted
categories or interaction types).
      </p>
      <p>
        Challenges. Devising the H-MTL framework optimized for SBRSs is challenging for three
major reasons. First, we need to define appropriate auxiliary tasks (e.g., category or interaction
type predictions) related to the main task; in addition, we ought to set up proper hierarchical
relationships between prediction tasks (e.g., a bottom-up approach from next-category prediction
to next-item prediction). Second, rich and accurate session representations are required to ensure
the high accuracy of the prediction tasks. The session features should contain item metadata
information (e.g., categories) so that they can be used for auxiliary prediction tasks. Finally, it is
computationally prohibitive to test the performance of H-MTL for SBRSs with millions of items
available on online platforms (e.g., Amazon website). While existing methods [
        <xref ref-type="bibr" rid="ref23 ref24 ref25">23, 24, 25</xref>
        ] use
randomly-sampled candidate items for the test, the sampled metrics can be inconsistent with
the original performance measured with all items [
        <xref ref-type="bibr" rid="ref26">26</xref>
        ]. Thus, how can we generate high-quality
candidate items to accurately evaluate the performance of H-MTL for SBRSs?
Proposed Method. To address the above challenges, we propose a novel recommendation
model called HierSRec which incorporates the H-MTL framework to SBRSs. HierSRec is trained
with multiple objectives of predicting the next item (main task) and next category (auxiliary
task) in a session. We choose next-category predictions as our auxiliary task since category
labels are mostly available in recommendation datasets (e.g., Amazon product review [
        <xref ref-type="bibr" rid="ref30">30</xref>
        ],
Diginetica [
        <xref ref-type="bibr" rid="ref31">31</xref>
        ]). If categories labels are partially available or completely unavailable, we can
cluster items to obtain implicit category information [
        <xref ref-type="bibr" rid="ref32">32</xref>
        ] or predict other metadata in a dataset
such as interaction types (e.g., purchase, click, etc.) as auxiliary tasks. The first step of HierSRec
is generating a session representation using a metadata-aware Transformer [
        <xref ref-type="bibr" rid="ref33">33</xref>
        ] encoder. The
Transformer encoder uses item IDs, item categories, and additional item metadata such as titles
and descriptions to produce a precise and rich summary of the session. After that, HierSRec
predicts the next category using the session encoding vector and transforms those category
prediction results into embeddings. Next, we predict the next item in a session using the session
representation and category prediction embeddings. Finally, the next-item and next-category
prediction losses are combined together to train HierSRec. After training, we first perform the
category prediction for all test instances, and we can generate high-quality candidate items for
each test example by aggregating items belonging to top-K predicted categories.
Experiments. Thorough experiments on two large-scale E-commerce datasets show that
HierSRec has superior next-item prediction performance to existing SBRSs by leveraging the
hierarchical prediction framework. HierSRec shows at least 6.7% performance improvements
as per three accuracy metrics compared to baselines. Moreover, HierSRec achieves comparable
accuracy to that of HierSRec tested with full items by deliberately selecting a few items (e.g.,
4% of total items) as ranking candidates using the category prediction. Ablation studies of
HierSRec verify the efectiveness of each component of HierSRec.
      </p>
      <p>Contributions. The main contributions of our paper are summarized as follows.
• To the best of our knowledge, this is the first work to leverage the H-MTL framework for</p>
      <p>SBRSs.
• We propose HierSRec that accurately predicts the next item in a session via employing
the output of next-category prediction. HierSRec ofers a compact set of candidate items
of each test example for scalable ranking.
• Experiments on two recommendation datasets show that HierSRec outperforms existing
SBRSs as per next-item prediction accuracy. We also confirm the efectiveness of our
candidate generation method.</p>
    </sec>
    <sec id="sec-2">
      <title>2. Related Work</title>
      <p>
        Multi-Task Learning (MTL) &amp; Hierarchical MTL. Recent research has shown that the
generalizability and prediction performance of recommendation models can gain substantial
improvements by using multi-task learning (MTL) [
        <xref ref-type="bibr" rid="ref34 ref4 ref5">4, 5, 34</xref>
        ]. In particular, MTL shares the
knowledge learned from other related tasks with the main one, which has been shown to not
only enhance the overall performance of the model, but also decrease the chance of overfitting
and improve the quality of learned representations [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ]. Hierarchical learning can improve
the generalizability and interpretability of MTL by using the predictions of related tasks for
another task. This architecture is called Hierarchical MTL (H-MTL). H-MTL has been utilized
in Natural Language Processing [
        <xref ref-type="bibr" rid="ref35 ref36 ref37 ref38 ref7 ref8 ref9">7, 8, 9, 35, 36, 37, 38</xref>
        ] and Computer Vision [
        <xref ref-type="bibr" rid="ref10 ref37 ref39 ref40">10, 37, 39, 40</xref>
        ]
domains to boost the performance of a model by sharing knowledge from lower-level tasks
for more complex ones. In the context of recommender systems, Chen et al. [
        <xref ref-type="bibr" rid="ref22">22</xref>
        ] use H-MTL
to improve the prediction accuracy and also provide a linguistic explanation of why a user
likes/dislikes an item. Lim et al. [
        <xref ref-type="bibr" rid="ref21">21</xref>
        ] also utilize H-MTL to predict the Point-of-Interest (POI) a
user will visit next.
      </p>
      <p>
        Session-based Recommendation. Neural networks have served as the key component of
the state-of-the-art SBRSs. Recurrent neural networks (RNNs) have been used to capture item
dependencies within sessions [
        <xref ref-type="bibr" rid="ref11 ref12 ref13 ref18 ref19">11, 12, 13, 18, 19</xref>
        ]. However, RNNs are limited in capturing
longer dependencies across items. Thus, graph neural network-based approaches [
        <xref ref-type="bibr" rid="ref15 ref27 ref28 ref41 ref42">15, 27, 28,
41, 42</xref>
        ] and attention-based methods [
        <xref ref-type="bibr" rid="ref11 ref12 ref13 ref18 ref19">11, 12, 13, 18, 19</xref>
        ] have been proposed to incorporate
such dependencies precisely into SBRSs. Transformer [
        <xref ref-type="bibr" rid="ref33">33</xref>
        ]-based approaches provide superior
performance in predictions [
        <xref ref-type="bibr" rid="ref43 ref44 ref45 ref46">43, 44, 45, 46</xref>
        ]. MTL has been utilized to improve the accuracy of
next-item prediction in SBRSs [
        <xref ref-type="bibr" rid="ref1 ref2 ref29 ref3">47, 1, 2, 3, 29</xref>
        ]. User intent prediction [
        <xref ref-type="bibr" rid="ref32">48, 32, 49</xref>
        ] is a well-known
example of applying MTL to SBRSs. However, existing SBRSs are either designed only for
next-item prediction or incompatible with the H-MTL architecture.
      </p>
    </sec>
    <sec id="sec-3">
      <title>3. Proposed Approach: HierSRec</title>
      <p>
        Overview. As shown in Figure 1, our proposed recommendation model HierSRec employs a
Transformer [
        <xref ref-type="bibr" rid="ref33">33</xref>
        ] architecture to encode items in a session accurately and utilizes the obtained
session representation for the hierarchical MTL [
        <xref ref-type="bibr" rid="ref10 ref39 ref7 ref8 ref9">7, 8, 9, 10, 39</xref>
        ]. We define tasks of HierSRec
as the next item prediction (main) and multi-level category predictions (auxiliary), where the
category information is available on many recommendation datasets (e.g., Amazon product
review [
        <xref ref-type="bibr" rid="ref30">30</xref>
        ], Diginetica [
        <xref ref-type="bibr" rid="ref31">31</xref>
        ]). Notice that our framework can be easily extended to other types of
auxiliary tasks such as predicting user actions (e.g., click, add-to-cart, or purchase). Finally, we
ofer a candidate item generation scheme that uses the output from auxiliary tasks for scalable
inference or evaluation.
      </p>
      <p>
        Session Encoding with Metadata-aware Transformer. We use the Transformer encoder [
        <xref ref-type="bibr" rid="ref33">33</xref>
        ]
to create an accurate representation of a user’s interest within a session. Formally, given a
session  with a sequence of  observed items {1, . . . , }, the first step is transforming the
item sequence to the item representation sequence {1 , . . . ,  } using the item ID, the item
category information, and the other item metadata such as titles and descriptions. Given an
observed item , ∀, 1 ≤  ≤  in the current session , its embedding is constructed as
follows.
      </p>
      <p>
        =  ((ℐ ,  , ℳ )),
(1)
where   and  indicate the Transformer encoder and embedding concatenation
operation, and ℐ , , and ℳ denote item ID embeddings, item category embeddings,
and item metadata embeddings, respectively. Diferent encoders such as STAMP [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ] can
be used as   in HierSRec, but our metadata-aware Trasnformer shows the best
empirical performance. Given a set of items ℐ and item ID embedding dimension ℐ , item
ID embeddings ℐ ∈ R|ℐ|× ℐ map each item ID to the ℐ -dimensional feature. Assuming
there are  -level item categories (e.g., 3-level categories for a “dinner plate” item are 1
Kitchen; 2 - Tableware &amp; Bar; 3 - Dinnerware), we create -dimensional embeddings for all
category levels (i.e., 1 , . . . ,  ) and sum up all category embeddings of an item to encode
its category information, i.e.,  = 1 + . . . +  . Finally, item metadata embeddings
ℳ ∈ R|ℐ|× ℳ are concatenations of trainable or pre-trained embeddings (e.g., from BERT [50]
for texts and ResNet [51] for images) for each metadata of an item. For instance, the metadata
description, image).
embedding of an item  can be derived as follows: ℳ = (title ,
      </p>
      <p>Finally, the item representation sequence {1 , . . . ,  } generated by the Transformer is
fed to the average pooling layer to generate an accurate session representation final of a given
session  (i.e., final ∈ Rℐ++ℳ = pooling (1 , . . . ,  )). We choose the average pooling
since it empirically shows the best next-item prediction performance compared to the trainable
pooling or max pooling.</p>
      <p>
        Hierarchical Multi-Task Learning (H-MTL). Given the final session representation ifnal
of a session  = {1, . . . , }, a simple yet efective way to predict the next item +1 in  is
employing a fully-connected layer to transform final to a next-item prediction score vector.
However, this approach can easily make a model overfit the training data compared to multi-task
learning (MTL) with sharing representations [
        <xref ref-type="bibr" rid="ref6">6, 52, 53, 54</xref>
        ].
      </p>
      <p>
        While the MTL method can avoid the overfitting problem, it can be further enhanced by
introducing a “hierarchy” or order between the prediction tasks, which is called hierarchical MTL
(H-MTL) [
        <xref ref-type="bibr" rid="ref10 ref39 ref7 ref8 ref9">7, 8, 9, 10, 39</xref>
        ]. H-MTL models fully exploit the outputs from other tasks as “additional
knowledge” via performing predictions of multiple tasks in a specific order or hierarchically
(i.e., tree-structure), which ofers the implicit data augmentation efect and the generalization
capability to the main model.
      </p>
      <p>To adapt H-MTL to SBRSs, we first predict categories of the next item (e.g., +1) in a session
 = {1, . . . , } and employ the category prediction results to enhance the item prediction.
Specifically, we obtain  -level category prediction score vectors (i.e., {1, . . . ,  }) of the next
item in a session using a session representation final , as shown below.</p>
      <p>next =  next (final +   ∑︁  ),</p>
      <p>=1
 = Proj () ∈ R(ℐ++ℳ), ∀, 1 ≤  ≤ .
(3)
Finally, we sum the session representation final and all category prediction embeddings
{1 , . . . ,  } and feed it to the fully-connected layer to generate the next-item prediction
vector next ∈ R|ℐ|, as shown below.</p>
      <p>=   (final ) ∈ R||, ∀, 1 ≤  ≤ ,
where   indicates a fully-connected layer. Next, we transform category prediction vectors
{1, . . . ,  } to category prediction embeddings {1 , . . . ,  } via projection layers: Proj  ∈
R||× (ℐ++ℳ), ∀, 1 ≤  ≤  , as shown below.
(2)
(4)
where   (a hyperparameter) controls the impact of category predictions on the next-item
prediction. Category prediction results {1, . . . ,  } can be used as implicit explanations of the
next-item prediction next as the category prediction embeddings are used as a part of input
features to the next-item predictor.</p>
      <p>Loss Function and Optimization. Since HierSRec is a multi-task learning method, loss
functions of multi-level next-category prediction and next-item prediction are combined and
jointly optimized together. We use the Cross-Entropy loss for both category and item predictions.
Assuming the ground-truth  -level categories of a next-item +1 we want to predict are
1, . . . ,  , then a loss function of a level- category prediction task is given as follows.
ℒnext (next ) = −</p>
      <p>Softmax (next ) ,
ℒ () = −
() log</p>
      <p>Softmax () ,

where  is a one-hot vector whose ℎ value is 1,  is a level- category prediction score vector,
and Softmax () = ∑︀  . Similarly, the next-item prediction loss is given as follows.
where  is a one-hot vector whose ℎ</p>
      <p>+1 value is 1, and next is a next-item prediction score
vector. The combined loss function for our proposed hierarchical multi-task learning is given as
follows.</p>
      <p>ℒfinal = ℒnext +  ∑︁
ℒ ,
where  is the importance weight of category prediction tasks. We tested several weighting
strategies for  and   such as trainable weights or randomized weights per epoch, but setting
 =   = 1.0 shows the best prediction performance empirically. We optimize the above
loss function (7) with Adam [55] optimizer for all training data.</p>
      <p>
        Candidate Item Generation for Scalable Evaluation. During the test (or inference) stage,
we use the next-item prediction vector next ∈ R|ℐ| generated from the trained model.
Computing next can be computationally expensive on large-scale recommendation datasets (e.g.,
e-commerce domain) with millions of items. Thus, for practicality, existing algorithms [
        <xref ref-type="bibr" rid="ref23 ref24 ref25">23, 24, 25</xref>
        ]
uses a small set of candidate items instead of all items during the inference. However, those
methods use randomly-sampled candidates, and the accuracy measured with such candidates
can be significantly diferent from the accuracy calculated with full items [
        <xref ref-type="bibr" rid="ref26">26</xref>
        ]. Thus, we propose
a more accurate candidate generation method that leverages the category prediction result. First,
each level- category prediction  and construct a candidate item set ℐ′ ⊂ ℐ
given a test session with observed items, we conduct the category prediction and obtain the
score vectors {1, . . . ,  }. We sample top-K (K: hyperparameter) categories {1 , . . . , 
 } from
by the following.
ℐ′ = { |  ∈ , ∀, 1 ≤  ≤ } ∪ · · · ∪ {  |  ∈  , ∀, 1 ≤  ≤ }
      </p>
      <p>
        1
We empirically verify that our candidate selection policy can achieve nearly equivalent accuracy
to that of an original policy that uses all items during the inference (refer to Figure 2 later).
||
∑︁
=1
|ℐ|
∑︁ () log
=1
︂(
︂(

=1
︂)
︂)
(5)
(6)
(7)
What If Category Information Is Unavailable? While item category information is available
in most recommendation datasets, it might be partially available or completely unavailable
in a few cases. To address it, we can find implicit categories of items by utilizing graph
neural networks and clustering [
        <xref ref-type="bibr" rid="ref32">32</xref>
        ]. For instance, we apply an of-the-shelf node embedding
algorithm on a session-item bipartite graph to obtain item representations. Applying a clustering
method (e.g., K-means) on obtained item embeddings will generate clusters of items, which
will approximate the category information. Another solution is employing other metadata (e.g.,
interaction type) for auxiliary tasks of the H-MTL.
      </p>
    </sec>
    <sec id="sec-4">
      <title>4. Experimental Evaluations of HierSRec</title>
      <p>
        Datasets. Table 2 lists the statistics of the datasets. The Home Depot (THD) is an E-commerce
dataset obtained from a large online retailer THD. The dataset is composed of Add-to-Cart (ATC)
events within millions of online sessions. The dataset has rich product metadata including 7
attributes: product title, 3-level categories, brand, manufacturer, color, department name, and
class name. We exclude certain information from the THD dataset according to the company’s
policy. For the THD dataset, we filter users and items with less than 10 interactions. Diginetica 1
1https://competitions.codalab.org/competitions/11161
is a public E-commerce dataset that was a part of CIKM Cup 2016 challenge. We did
preprocess of the Diginetica similar to [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ].
      </p>
      <p>
        Baselines. We use the following state-of-the-art session-based recommenders:2 (1) NARM [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ]:
An attention-based model that employs a hybrid encoder to reflect a user’s global and local
interests with an attention mechanism, (2) STAMP [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ]: An attention/memory-based model that
incorporates a user’s short-term and long-term interests via short-term attention and long-term
memory modules, respectively, (3) CSRM [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ]: a session-based recommendation model that
contextualizes the current and neighborhood sessions with inner and outer memory encoders,
respectively, (4) TAGNN [
        <xref ref-type="bibr" rid="ref27">27</xref>
        ]: a graph neural network (GNN)-based session-based recommender
that utilizes a target-aware attention module for predictions, and (5) COTREC [
        <xref ref-type="bibr" rid="ref28">28</xref>
        ]: a
state-ofthe-art GNN-based recommendation model that combines self-supervised learning with graph
co-training. We exclude several models including nearest-neighbor algorithms if they show
similar or worse performance compared to our existing baselines.
      </p>
      <p>Hyperparameters. Hyperparameters of HierSRec and baseline methods are found by
extensive grid search using a validation set (randomly sampled 10% from training). Specifically,
ℐ =  = ℳ = 128,  =   = 1.0, and batch size and learning rates are set to 1024 and
0.0001, respectively. We use all items as candidates during the test by default. HierSRec also
uses 2 layers of Transformer Encoder with 8 attention heads.</p>
      <p>Reproducibility. While the code of HierSRec and the THD dataset cannot be released due to
the company policy, we release the public dataset (Diginetica) and baseline implementations
used in the paper.</p>
      <p>Next-item Prediction Accuracy of HierSRec. To verify the efectiveness of HierSRec,
we measure the next-item prediction accuracy of HierSRec and baselines on diverse datasets,
with respect to three accuracy metrics: Mean Reciprocal Rank@20 (MRR) [56], HITS@20, and
Recall@20. HITS@20 counts only ground-truth next-items in top-20 lists, while Recall@20
counts all future items (including the next-item) in a session in top-20 lists.</p>
      <p>Table 3 shows the next-item prediction accuracy of HierSRec and baselines on the Diginetica
and THD datasets. HierSRec shows the best performance as per all metrics among all methods
across all datasets, with statistical significance (P-values from one-tailed t-test are ≤ 0.05).
Relative performance improvements of HierSRec compared to the best baseline are 6.7% −
25.8%. The high performance of HierSRec is due to the hierarchical learning architecture,
not additional item metadata (compare the first and last row in Table 4b). In other words,
the key reason for these performance improvements is incorporating prior knowledge from the
next-category prediction into the next-item prediction, so that we can filter out items associated
with irrelevant categories easily while predicting the next item.</p>
      <p>Next-category Prediction Accuracy of HierSRec. We test how accurately HierSRec can
predict the next category in a session on the Diginetica dataset. As shown in Table 4a, HierSRec
outperforms the heuristic and shows almost the same accuracy as an MTL variant of HierSRec
without hierarchical learning. It is expected since the category prediction does not take any
additional feature from the next-item prediction task due to its lower hierarchy, and enhancing
the next-category prediction accuracy is not the main goal of HierSRec.</p>
      <p>Ablation Study of HierSRec. We conduct the ablation study of HierSRec to show how
2We used open-source implementations of baseline algorithms (https://github.com/rn5l/session-rec).</p>
      <sec id="sec-4-1">
        <title>Models / Metrics</title>
        <p>b Ablation study of HierSRec.</p>
      </sec>
      <sec id="sec-4-2">
        <title>MRR HITS @20 @20</title>
        <p>efective the hierarchical learning of HierSRec is for session-based recommendations. We
create two variants of HierSRec, where the first one only performs next-item predictions
only with metadata-aware Transformer (no MTL), and the second one employs normal
MTL architecture without hierarchical predictions. Table 4b shows the ablation study result
of HierSRec on the THD dataset. As we can notice, HierSRec exhibits the highest accuracy
compared to the two variants, with at least 7.3% relative performance improvements and
statistical significance. This result shows that our proposed hierarchical MTL architecture
induces a higher generalization capability of a model and an implicit data augmentation efect.
Verification of Candidate Items Generated by HierSRec. We confirm the quality of
candidate items generated by HierSRec by comparing the accuracy measured with our
candidates, random candidates, and full items. Figure 2a shows the MRR@20 metric of HierSRec
calculated with three diferent candidate generation policies on the Diginetica dataset. Using
the category prediction knowledge, HierSRec can create a small candidate set (e.g., 4% of total
items) consisting of key items that are highly related to the ground-truth next item in a session.
However, the random candidate policy exhibits poor performance as it cannot selectively choose
important items as well as the ground-truth next item.</p>
        <p>Case Study: Hierarchical Predictions of HierSRec. Figure 2b is a case study result on the
a Next-item prediction accuracy of HierSRec as per the b Case-study result of HierSRec on the The Home Depot
number of candidates used for test. dataset.
THD dataset of how HierSRec utilizes category predictions to improve the next-item predictions.
Top-K category prediction results of HierSRec include diverse and evolving preferences of a
user in a session (e.g., Door →− Bathroom) and identify the most important categories (e.g.,
Bath and Bathroom Faucets) for next-item predictions. Providing these meaningful knowledge
from category predictions will make next-item recommendations personalized and accurate.</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>5. Conclusion</title>
      <p>In this paper, we proposed a novel session-based recommendation model HierSRec that employs
a metadata-aware Transformer encoder and a hierarchical multi-task learning framework to
obtain higher model generalizability. Future works of HierSRec include that (1) more complex
hierarchical structures (e.g., tree-shape) between various tasks (e.g., next-item, next-action, and
next-category predictions) can be explored, and (2) extending HierSRec to predict next items
accurately in sessions with cold-start items or only a few items by employing their metadata.
features and post-fusion context for e-commerce session-based recommendation, arXiv
preprint arXiv:2107.05124 (2021).
[47] M. Tavakol, U. Brefeld, Factored mdps for detecting topics of user sessions, in: Proceedings
of the 8th ACM Conference on Recommender Systems, 2014, pp. 33–40.
[48] J. Guo, Y. Yang, X. Song, Y. Zhang, Y. Wang, J. Bai, Y. Zhang, Learning multi-granularity
consecutive user intent unit for session-based recommendation, in: Proceedings of the
iffteenth ACM International conference on web search and data mining, 2022, pp. 343–352.
[49] Z. Liu, H. Chen, F. Sun, X. Xie, J. Gao, B. Ding, Y. Shen, Intent preference decoupling for
user representation on online recommender system, in: Proceedings of the Twenty-Ninth
International Conference on International Joint Conferences on Artificial Intelligence,
2021, pp. 2575–2582.
[50] J. Devlin, M.-W. Chang, K. Lee, K. Toutanova, BERT: Pre-training of deep bidirectional
transformers for language understanding, Association for Computational Linguistics, 2019,
pp. 4171–4186.
[51] K. He, X. Zhang, S. Ren, J. Sun, Deep residual learning for image recognition, in: 2016
IEEE Conference on Computer Vision and Pattern Recognition (CVPR), IEEE, 2016, pp.
770–778.
[52] S. Chowdhuri, T. Pankaj, K. Zipser, Multinet: Multi-modal multi-task learning for
autonomous driving, in: 2019 IEEE Winter Conference on Applications of Computer Vision
(WACV), IEEE, 2019, pp. 1496–1504.
[53] X. Wang, C. Zhang, Z. Zhang, Boosted multi-task learning for face verification with
applications to web image and video search, in: 2009 IEEE Conference on Computer Vision
and Pattern Recognition, IEEE, 2009, pp. 142–149.
[54] S. Zhang, L. Yao, A. Sun, Y. Tay, Deep learning based recommender system: A survey and
new perspectives, ACM Computing Surveys (CSUR) 52 (2019) 1–38.
[55] D. P. Kingma, J. Ba, Adam: A method for stochastic optimization, in: International</p>
      <p>Conference on Learning Representations (ICLR), 2015.
[56] E. M. Voorhees, et al., The trec-8 question answering track report., in: Text Retrieval
Conference, volume 99, 1999, pp. 77–82.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>N.</given-names>
            <surname>Qiu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Gao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Huang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Tu</surname>
          </string-name>
          , W. Luo,
          <article-title>Incorporating global context into multi-task learning for session-based recommendation</article-title>
          ,
          <source>in: International Conference on Knowledge Science, Engineering and Management</source>
          , Springer,
          <year>2021</year>
          , pp.
          <fpage>627</fpage>
          -
          <lpage>638</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>C.</given-names>
            <surname>Huang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Chen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Xia</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Xu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Dai</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Chen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Bo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Zhao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. X.</given-names>
            <surname>Huang</surname>
          </string-name>
          ,
          <article-title>Graph-enhanced multi-task learning of multi-level transition dynamics for session-based recommendation</article-title>
          ,
          <source>in: AAAI Conference on Artificial Intelligence (AAAI)</source>
          ,
          <year>2021</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>W.</given-names>
            <surname>Meng</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Yang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Xiao</surname>
          </string-name>
          ,
          <article-title>Incorporating user micro-behaviors and item knowledge into multi-task learning for session-based recommendation</article-title>
          ,
          <source>in: Proceedings of the 43rd International ACM SIGIR Conference on Research and Development in Information Retrieval</source>
          ,
          <year>2020</year>
          , pp.
          <fpage>1091</fpage>
          -
          <lpage>1100</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>C.</given-names>
            <surname>Gao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>He</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Gan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Chen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Feng</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.-S.</given-names>
            <surname>Chua</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Jin</surname>
          </string-name>
          ,
          <article-title>Neural multi-task recommendation from multi-behavior data, in: 2019 IEEE 35th international conference on data engineering (ICDE)</article-title>
          , IEEE,
          <year>2019</year>
          , pp.
          <fpage>1554</fpage>
          -
          <lpage>1557</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>G.</given-names>
            <surname>Hadash</surname>
          </string-name>
          ,
          <string-name>
            <given-names>O. S.</given-names>
            <surname>Shalom</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Osadchy</surname>
          </string-name>
          ,
          <article-title>Rank and rate: multi-task learning for recommender systems</article-title>
          ,
          <source>in: Proceedings of the 12th ACM Conference on Recommender Systems</source>
          ,
          <year>2018</year>
          , pp.
          <fpage>451</fpage>
          -
          <lpage>454</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>S.</given-names>
            <surname>Ruder</surname>
          </string-name>
          ,
          <article-title>An overview of multi-task learning in deep neural networks</article-title>
          ,
          <source>arXiv preprint arXiv:1706.05098</source>
          (
          <year>2017</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>V.</given-names>
            <surname>Sanh</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Wolf</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Ruder</surname>
          </string-name>
          ,
          <article-title>A hierarchical multi-task approach for learning embeddings from semantic tasks</article-title>
          ,
          <source>in: Proceedings of the AAAI Conference on Artificial Intelligence</source>
          , volume
          <volume>33</volume>
          ,
          <year>2019</year>
          , pp.
          <fpage>6949</fpage>
          -
          <lpage>6956</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>X.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Xu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Zhang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Jin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Sun</surname>
          </string-name>
          ,
          <article-title>End-to-end aspect-based sentiment analysis with hierarchical multi-task learning</article-title>
          ,
          <source>Neurocomputing</source>
          <volume>455</volume>
          (
          <year>2021</year>
          )
          <fpage>178</fpage>
          -
          <lpage>188</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>B.</given-names>
            <surname>Tian</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Zhang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Xing</surname>
          </string-name>
          ,
          <article-title>Hierarchical inter-attention network for document classification with multi-task learning</article-title>
          .,
          <source>in: IJCAI</source>
          ,
          <year>2019</year>
          , pp.
          <fpage>3569</fpage>
          -
          <lpage>3575</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>H.</given-names>
            <surname>Park</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Bharadhwaj</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B. Y.</given-names>
            <surname>Lim</surname>
          </string-name>
          ,
          <article-title>Hierarchical multi-task learning for healthy drink classification</article-title>
          , in: 2019
          <source>International Joint Conference on Neural Networks (IJCNN)</source>
          , IEEE,
          <year>2019</year>
          , pp.
          <fpage>1</fpage>
          -
          <lpage>8</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>J.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Ren</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Chen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Ren</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Lian</surname>
          </string-name>
          , J. Ma,
          <article-title>Neural attentive session-based recommendation</article-title>
          ,
          <source>in: Proceedings of the 2017 ACM on Conference on Information and Knowledge Management</source>
          ,
          <year>2017</year>
          , pp.
          <fpage>1419</fpage>
          -
          <lpage>1428</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>Q.</given-names>
            <surname>Liu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Zeng</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Mokhosi</surname>
          </string-name>
          ,
          <string-name>
            <surname>H. Zhang,</surname>
          </string-name>
          <article-title>Stamp: short-term attention/memory priority model for session-based recommendation</article-title>
          ,
          <source>in: Proceedings of the 24th ACM SIGKDD International Conference on Knowledge Discovery &amp; Data Mining</source>
          ,
          <year>2018</year>
          , pp.
          <fpage>1831</fpage>
          -
          <lpage>1839</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <given-names>M.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Ren</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Mei</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Chen</surname>
          </string-name>
          , J. Ma, M. de Rijke,
          <article-title>A collaborative session-based recommendation approach with parallel memory modules</article-title>
          ,
          <source>in: Proceedings of the 42nd International ACM SIGIR Conference on Research and Development in Information Retrieval</source>
          ,
          <year>2019</year>
          , pp.
          <fpage>345</fpage>
          -
          <lpage>354</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <given-names>C.</given-names>
            <surname>Hansen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Hansen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Maystre</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Mehrotra</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Brost</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Tomasi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Lalmas</surname>
          </string-name>
          ,
          <article-title>Contextual and sequential user embeddings for large-scale music recommendation</article-title>
          ,
          <source>in: ACM RecSys</source>
          ,
          <year>2020</year>
          , pp.
          <fpage>53</fpage>
          -
          <lpage>62</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <given-names>S.</given-names>
            <surname>Wu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Tang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Zhu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Xie</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Tan</surname>
          </string-name>
          ,
          <article-title>Session-based recommendation with graph neural networks</article-title>
          ,
          <source>in: Proceedings of the AAAI Conference on Artificial Intelligence</source>
          , volume
          <volume>33</volume>
          ,
          <year>2019</year>
          , pp.
          <fpage>346</fpage>
          -
          <lpage>353</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [16]
          <string-name>
            <given-names>B.</given-names>
            <surname>Hidasi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Karatzoglou</surname>
          </string-name>
          ,
          <article-title>Recurrent neural networks with top-k gains for session-based recommendations</article-title>
          ,
          <source>in: Proceedings of the 27th ACM international conference on information and knowledge management</source>
          ,
          <year>2018</year>
          , pp.
          <fpage>843</fpage>
          -
          <lpage>852</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [17]
          <string-name>
            <given-names>J.</given-names>
            <surname>You</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Pal</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Eksombatchai</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Rosenburg</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Leskovec</surname>
          </string-name>
          ,
          <article-title>Hierarchical temporal convolutional networks for dynamic recommender systems</article-title>
          ,
          <source>in: The world wide web conference</source>
          ,
          <year>2019</year>
          , pp.
          <fpage>2236</fpage>
          -
          <lpage>2246</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          [18]
          <string-name>
            <given-names>Z.</given-names>
            <surname>Pan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Cai</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Ling</surname>
          </string-name>
          ,
          <string-name>
            <surname>M. de Rijke</surname>
          </string-name>
          ,
          <article-title>An intent-guided collaborative machine for sessionbased recommendation</article-title>
          ,
          <source>in: Proceedings of the 43rd International ACM SIGIR Conference on Research and Development in Information Retrieval</source>
          ,
          <year>2020</year>
          , pp.
          <fpage>1833</fpage>
          -
          <lpage>1836</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          [19]
          <string-name>
            <given-names>P.</given-names>
            <surname>Ren</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Chen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Ren</surname>
          </string-name>
          , J. Ma, M. De Rijke,
          <article-title>Repeatnet: A repeat aware neural recommendation machine for session-based recommendation</article-title>
          ,
          <source>in: Proceedings of the AAAI Conference on Artificial Intelligence</source>
          , volume
          <volume>33</volume>
          ,
          <year>2019</year>
          , pp.
          <fpage>4806</fpage>
          -
          <lpage>4813</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          [20]
          <string-name>
            <given-names>S.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Cao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Q. Z.</given-names>
            <surname>Sheng</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Orgun</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Lian</surname>
          </string-name>
          ,
          <article-title>A survey on session-based recommender systems</article-title>
          , arXiv preprint arXiv:
          <year>1902</year>
          .
          <volume>04864</volume>
          (
          <year>2019</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          [21]
          <string-name>
            <given-names>N.</given-names>
            <surname>Lim</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Hooi</surname>
          </string-name>
          , S.
          <article-title>-</article-title>
          <string-name>
            <surname>K. Ng</surname>
            ,
            <given-names>Y. L.</given-names>
          </string-name>
          <string-name>
            <surname>Goh</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          <string-name>
            <surname>Weng</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          <string-name>
            <surname>Tan</surname>
          </string-name>
          ,
          <article-title>Hierarchical multi-task graph recurrent network for next poi recommendation</article-title>
          ,
          <source>in: Proceedings of the 44th International ACM SIGIR Conference on Research and Development in Information Retrieval</source>
          ,
          <year>2022</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          [22]
          <string-name>
            <given-names>Z.</given-names>
            <surname>Chen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Xie</surname>
          </string-name>
          , T. Wu, G. Bu,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <surname>E. Chen,</surname>
          </string-name>
          <article-title>Co-attentive multi-task learning for explainable recommendation</article-title>
          .,
          <source>in: IJCAI</source>
          ,
          <year>2019</year>
          , pp.
          <fpage>2137</fpage>
          -
          <lpage>2143</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          [23]
          <string-name>
            <given-names>J.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <surname>J. McAuley</surname>
          </string-name>
          ,
          <article-title>Time interval aware self-attention for sequential recommendation</article-title>
          ,
          <source>in: Proceedings of the 13th international conference on web search and data mining</source>
          ,
          <year>2020</year>
          , pp.
          <fpage>322</fpage>
          -
          <lpage>330</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          [24]
          <string-name>
            <surname>W.-C. Kang</surname>
            ,
            <given-names>J. McAuley</given-names>
          </string-name>
          ,
          <article-title>Self-attentive sequential recommendation, in: 2018 IEEE international conference on data mining (ICDM)</article-title>
          , IEEE,
          <year>2018</year>
          , pp.
          <fpage>197</fpage>
          -
          <lpage>206</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref25">
        <mixed-citation>
          [25]
          <string-name>
            <given-names>F.</given-names>
            <surname>Sun</surname>
          </string-name>
          , J. Liu,
          <string-name>
            <given-names>J.</given-names>
            <surname>Wu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Pei</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Lin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W.</given-names>
            <surname>Ou</surname>
          </string-name>
          , P. Jiang,
          <article-title>Bert4rec: Sequential recommendation with bidirectional encoder representations from transformer</article-title>
          ,
          <source>in: Proceedings of the 28th ACM international conference on information and knowledge management</source>
          ,
          <year>2019</year>
          , pp.
          <fpage>1441</fpage>
          -
          <lpage>1450</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref26">
        <mixed-citation>
          [26]
          <string-name>
            <given-names>W.</given-names>
            <surname>Krichene</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Rendle</surname>
          </string-name>
          ,
          <article-title>On sampled metrics for item recommendation</article-title>
          ,
          <source>in: Proceedings of the 26th ACM SIGKDD international conference on knowledge discovery &amp; data mining</source>
          ,
          <year>2020</year>
          , pp.
          <fpage>1748</fpage>
          -
          <lpage>1757</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref27">
        <mixed-citation>
          [27]
          <string-name>
            <given-names>F.</given-names>
            <surname>Yu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Zhu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Q.</given-names>
            <surname>Liu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Wu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Tan</surname>
          </string-name>
          ,
          <article-title>Tagnn: Target attentive graph neural networks for session-based recommendation</article-title>
          ,
          <source>in: Proceedings of the 43rd International ACM SIGIR Conference on Research and Development in Information Retrieval</source>
          ,
          <year>2020</year>
          , pp.
          <fpage>1921</fpage>
          -
          <lpage>1924</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref28">
        <mixed-citation>
          [28]
          <string-name>
            <given-names>X.</given-names>
            <surname>Xia</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Yin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Yu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Shao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Cui</surname>
          </string-name>
          ,
          <string-name>
            <surname>Self-Supervised Graph</surname>
          </string-name>
          Co-
          <article-title>Training for Session-based Recommendation</article-title>
          ,
          <source>in: 30th ACM International Conference on Information and Knowledge Management (CIKM</source>
          <year>2021</year>
          ), ACM,
          <year>2021</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref29">
        <mixed-citation>
          [29]
          <string-name>
            <given-names>W.</given-names>
            <surname>Shalaby</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Oh</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Afsharinejad</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Kumar</surname>
          </string-name>
          ,
          <string-name>
            <surname>X. Cui,</surname>
          </string-name>
          <article-title>M2trec: Metadata-aware multitask transformer for large-scale and cold-start free session-based recommendations</article-title>
          ,
          <source>in: Proceedings of the 16th ACM Conference on Recommender Systems</source>
          ,
          <year>2022</year>
          , pp.
          <fpage>573</fpage>
          -
          <lpage>578</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref30">
        <mixed-citation>
          [30]
          <string-name>
            <given-names>J.</given-names>
            <surname>Ni</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <surname>J. McAuley</surname>
          </string-name>
          ,
          <article-title>Justifying recommendations using distantly-labeled reviews and ifne-grained aspects</article-title>
          ,
          <source>in: Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP)</source>
          ,
          <year>2019</year>
          , pp.
          <fpage>188</fpage>
          -
          <lpage>197</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref31">
        <mixed-citation>
          [31]
          <article-title>Diginetica dataset for cikm cup 2016 challenge</article-title>
          , https://competitions.codalab.org/ competitions/11161,
          <year>2016</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref32">
        <mixed-citation>
          [32]
          <string-name>
            <given-names>S.</given-names>
            <surname>Oh</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Bhardwaj</surname>
          </string-name>
          , J. Han,
          <string-name>
            <surname>S</surname>
          </string-name>
          . Kim,
          <string-name>
            <given-names>R. A.</given-names>
            <surname>Rossi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Kumar</surname>
          </string-name>
          ,
          <article-title>Implicit session contexts for next-item recommendations</article-title>
          ,
          <source>in: Proceedings of the 31st ACM International Conference on Information &amp; Knowledge Management</source>
          ,
          <year>2022</year>
          , pp.
          <fpage>4364</fpage>
          -
          <lpage>4368</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref33">
        <mixed-citation>
          [33]
          <string-name>
            <given-names>A.</given-names>
            <surname>Vaswani</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Shazeer</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Parmar</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Uszkoreit</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Jones</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A. N.</given-names>
            <surname>Gomez</surname>
          </string-name>
          , Ł. Kaiser,
          <string-name>
            <surname>I. Polosukhin</surname>
          </string-name>
          ,
          <article-title>Attention is all you need</article-title>
          ,
          <source>Advances in neural information processing systems</source>
          <volume>30</volume>
          (
          <year>2017</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref34">
        <mixed-citation>
          [34]
          <string-name>
            <given-names>R.</given-names>
            <surname>Caruana</surname>
          </string-name>
          ,
          <article-title>Multitask learning</article-title>
          ,
          <source>Machine Learning</source>
          <volume>28</volume>
          (
          <year>1997</year>
          ). doi:
          <volume>10</volume>
          .1023/A:
          <fpage>1007379606734</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref35">
        <mixed-citation>
          [35]
          <string-name>
            <given-names>Y.</given-names>
            <surname>Farag</surname>
          </string-name>
          ,
          <string-name>
            <surname>H.</surname>
          </string-name>
          <article-title>Yannakoudakis, Multi-task learning for coherence modeling, in: Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, Association for Computational Linguistics (ACL</article-title>
          ),
          <year>2019</year>
          , pp.
          <fpage>629</fpage>
          -
          <lpage>639</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref36">
        <mixed-citation>
          [36]
          <string-name>
            <given-names>M.</given-names>
            <surname>Song</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Zhao</surname>
          </string-name>
          ,
          <article-title>Enhance rnnlms with hierarchical multi-task learning for asr</article-title>
          ,
          <source>in: ICASSP 2022 - 2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)</source>
          ,
          <year>2022</year>
          , pp.
          <fpage>6102</fpage>
          -
          <lpage>6106</lpage>
          . doi:
          <volume>10</volume>
          .1109/ICASSP43922.
          <year>2022</year>
          .
          <volume>9747525</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref37">
        <mixed-citation>
          [37]
          <string-name>
            <surname>D.-K. Nguyen</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          <article-title>Okatani, Multi-task learning of hierarchical vision-language representation</article-title>
          ,
          <source>2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)</source>
          (
          <year>2019</year>
          )
          <fpage>10484</fpage>
          -
          <lpage>10493</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref38">
        <mixed-citation>
          [38]
          <string-name>
            <given-names>W.</given-names>
            <surname>Song</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Song</surname>
          </string-name>
          , L. Liu,
          <string-name>
            <given-names>R.</given-names>
            <surname>Fu</surname>
          </string-name>
          ,
          <article-title>Hierarchical multi-task learning for organization evaluation of argumentative student essays</article-title>
          , in: C.
          <string-name>
            <surname>Bessiere</surname>
          </string-name>
          (Ed.),
          <source>Proceedings of the Twenty-Ninth International Joint Conference on Artificial Intelligence, IJCAI-20, International Joint Conferences on Artificial Intelligence Organization</source>
          ,
          <year>2020</year>
          , pp.
          <fpage>3875</fpage>
          -
          <lpage>3881</lpage>
          . URL: https: //doi.org/10.24963/ijcai.
          <year>2020</year>
          /536. doi:
          <volume>10</volume>
          .24963/ijcai.
          <year>2020</year>
          /536.
        </mixed-citation>
      </ref>
      <ref id="ref39">
        <mixed-citation>
          [39]
          <string-name>
            <given-names>J.</given-names>
            <surname>Fan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Zhao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Kuang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Zheng</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Zhang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Yu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Peng</surname>
          </string-name>
          , Hd-mtl:
          <article-title>Hierarchical deep multitask learning for large-scale visual recognition</article-title>
          ,
          <source>IEEE transactions on image processing 26</source>
          (
          <year>2017</year>
          )
          <fpage>1923</fpage>
          -
          <lpage>1938</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref40">
        <mixed-citation>
          [40]
          <string-name>
            <given-names>Z.</given-names>
            <surname>Yan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Zhang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Piramuthu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Jagadeesh</surname>
          </string-name>
          , D. DeCoste, W. Di,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Yu</surname>
          </string-name>
          ,
          <article-title>Hd-cnn: hierarchical deep convolutional neural networks for large scale visual recognition</article-title>
          ,
          <source>in: Proceedings of the IEEE international conference on computer vision</source>
          ,
          <year>2015</year>
          , pp.
          <fpage>2740</fpage>
          -
          <lpage>2748</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref41">
        <mixed-citation>
          [41]
          <string-name>
            <given-names>C.</given-names>
            <surname>Xu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Zhao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Liu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V. S.</given-names>
            <surname>Sheng</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Xu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Zhuang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Fang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Zhou</surname>
          </string-name>
          ,
          <article-title>Graph contextualized self-attention network for session-based recommendation</article-title>
          .,
          <source>in: IJCAI</source>
          , volume
          <volume>19</volume>
          ,
          <year>2019</year>
          , pp.
          <fpage>3940</fpage>
          -
          <lpage>3946</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref42">
        <mixed-citation>
          [42]
          <string-name>
            <given-names>Z.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W.</given-names>
            <surname>Wei</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Cong</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.-L.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.-L.</given-names>
            <surname>Mao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Qiu</surname>
          </string-name>
          ,
          <article-title>Global context enhanced graph neural networks for session-based recommendation</article-title>
          ,
          <source>in: Proceedings of the 43rd International ACM SIGIR Conference on Research and Development in Information Retrieval</source>
          ,
          <year>2020</year>
          , pp.
          <fpage>169</fpage>
          -
          <lpage>178</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref43">
        <mixed-citation>
          [43]
          <string-name>
            <surname>G. de Souza Pereira Moreira</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          <string-name>
            <surname>Rabhi</surname>
            ,
            <given-names>J. M.</given-names>
          </string-name>
          <string-name>
            <surname>Lee</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          <string-name>
            <surname>Ak</surname>
            ,
            <given-names>E. Oldridge,</given-names>
          </string-name>
          <article-title>Transformers4rec: Bridging the gap between nlp and sequential/session-based recommendation</article-title>
          ,
          <source>in: Fifteenth ACM Conference on Recommender Systems</source>
          ,
          <year>2021</year>
          , pp.
          <fpage>143</fpage>
          -
          <lpage>153</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref44">
        <mixed-citation>
          [44]
          <string-name>
            <surname>G. de Souza Pereira Moreira</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          <string-name>
            <surname>Rabhi</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          <string-name>
            <surname>Ak</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          <string-name>
            <surname>Schiferer</surname>
          </string-name>
          ,
          <article-title>End-to-end session-based recommendation on gpu</article-title>
          ,
          <source>in: Fifteenth ACM Conference on Recommender Systems</source>
          ,
          <year>2021</year>
          , pp.
          <fpage>831</fpage>
          -
          <lpage>833</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref45">
        <mixed-citation>
          [45]
          <string-name>
            <given-names>X.</given-names>
            <surname>Chen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Liu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Lei</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.-J.</given-names>
            <surname>Zha</surname>
          </string-name>
          ,
          <string-name>
            <surname>Z. Xiong,</surname>
          </string-name>
          <article-title>Bert4sessrec: Content-based video relevance prediction with bidirectional encoder representations from transformer</article-title>
          ,
          <source>in: Proceedings of the 27th ACM International Conference on Multimedia</source>
          ,
          <year>2019</year>
          , pp.
          <fpage>2597</fpage>
          -
          <lpage>2601</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref46">
        <mixed-citation>
          [46]
          <string-name>
            <given-names>G. d. S. P.</given-names>
            <surname>Moreira</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Rabhi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Ak</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. Y.</given-names>
            <surname>Kabir</surname>
          </string-name>
          , E. Oldridge,
          <article-title>Transformers with multi-modal</article-title>
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>