=Paper= {{Paper |id=Vol-3837/paper3 |storemode=property |title=RankTower: A Synergistic Framework for Enhancing Two-Tower Pre-Ranking Model |pdfUrl=https://ceur-ws.org/Vol-3837/paper_07_ceur_paper.pdf |volume=Vol-3837 |authors=Yachen Yan,Liubo Li |dblpUrl=https://dblp.org/rec/conf/adkdd/YanL24 }} ==RankTower: A Synergistic Framework for Enhancing Two-Tower Pre-Ranking Model== https://ceur-ws.org/Vol-3837/paper_07_ceur_paper.pdf
                         RankTower: A Synergistic Framework for Enhancing Two-Tower
                         Pre-Ranking Model
                         Yachen Yan1,∗ , Liubo Li1
                         1
                             Credit Karma, 760 Market Street, San Francisco, California, USA, 94012


                                             Abstract
                                             In large-scale ranking systems, cascading architectures have been widely adopted to achieve a balance between efficiency and effec-
                                             tiveness. The pre-ranking module selects candidates for the subsequent ranking module, while maintaining efficiency and accuracy
                                             under online latency constraints. In this paper, we propose a novel neural network architecture called RankTower, which is designed
                                             to efficiently capture user-item interactions while following the user-item decoupling paradigm to ensure online inference efficiency.
                                             The proposed approach employs a hybrid training objective that learns from samples obtained from the full stage of the cascade
                                             ranking system, optimizing different objectives for varying sample spaces. This strategy enhances the pre-ranking model’s ranking
                                             capability and alignment with the existing cascade ranking system. Experimental results conducted on public datasets demonstrate that
                                             RankTower significantly outperforms state-of-the-art pre-ranking models.

                                             Keywords
                                             Recommender Systems, Pre-Ranking, Learning to Rank, Differentiable Sorting



                                                                                                                              Item 1


                                 Item              Recall           Pre-Ranking          Ranking           Re-Ranking         Item 2

                                Corpus                                                                                                                                                                          Online Serving
                                                                                                                              Item 3
                                                                                                                                                                      Maximum Similarity Layer

                                                                                                                                                                                                                   Vector
                                                                                                                                                                                                                  Database
                         Figure 1: The Architecture of Cascade Ranking System                                                                                              Gated Cross-Attention
                                                                                                                                                                                 Network




                         1. Introduction                                                                                                          User
                                                                                                                                            Multi-Embedding
                                                                                                                                                                                                                       Item
                                                                                                                                                                                                                Multi-Embedding




                         In industrial information services, such as recommender                                                                         Multi-Head Gated                   Multi-Head Gated
                                                                                                                                                             Network                            Network
                         systems, search engines, and advertisement systems, the
                         cascading architecture ranking system has been widely used                                                                      Embedding Layer                    Embedding Layer
                         to achieve a balance between efficiency and effectiveness.
                         A typical cascade ranking system, as illustrated in Figure 1,
                         consists of multiple sequential stages, including recall, pre-                                                                       User Tower                           Item Tower

                         ranking, ranking, and re-ranking stages. Pre-ranking is
                         commonly regarded as a lightweight ranking module char-                                                        Figure 2: The Architecture of RankTower
                         acterized by a simpler network architecture and a reduced
                         set of features. Compared to ranking models, pre-ranking
                         models are required to score a larger number of candidate
                         items for each user and demonstrate higher inference ef-                                                               the cascade ranking system. Tightly coupled with
                         ficiency. Given the emphasis on efficiency, pre-ranking                                                                this sampling approach, we strategically integrate a
                         typically employs a straightforward vector-product-based                                                               hybrid loss function that combines distillation and
                         model.                                                                                                                 learning-to-rank losses. This synergistic approach
                            We propose a novel pre-ranking framework called Rank-                                                               facilitates comprehensive learning of the ordering
                         Tower to address these challenges. The primary contribu-                                                               dynamics underlying user interactions while align-
                         tions are as follows:                                                                                                  ing with the inherent patterns of the cascade ranking
                                                                                                                                                system.
                                  • We introduce the RankTower architecture, which
                                    comprises three key components: Multi-Head Gated                                                         • Experiments on public datasets demonstrate that
                                    Network, Gated Cross-Attention Network, and Max-                                                           RankTower significantly outperforms state-of-the-
                                    imum Similarity Layer. This architecture efficiently                                                       art pre-ranking models in terms of prediction accu-
                                    captures user-item interactions while following the                                                        racy and inference efficiency.
                                    user-item decoupling paradigm to ensure online in-
                                    ference efficiency.
                                                                                                                                        2. Model Architecture
                                  • We employ a full-stage sampling strategy by draw-
                                    ing the training samples from different stages of                                                   The RankTower architecture, as shown in Figure 2, intro-
                                                                                                                                        duces three main modules: Multi-Head Gated Network for
                         Woodstock’22: Symposium on the irreproducible science, June 07–11, 2022,
                         Woodstock, NY                                                                                                  computing diversified user and item representations, Gated
                         ∗
                              Corresponding author.                                                                                     Cross-Attention Network for modeling bi-directional user-
                         Envelope-Open yachen.yan@creditkarma.com (Y. Yan); liubo.li@creditkarma.com                                    item interactions, and Maximum Similarity Layer for effi-
                         (L. Li)                                                                                                        ciently capturing user-item interactions to compute the final
                         Orcid 0000-0002-1213-4343 (Y. Yan); 0009-0006-9933-2436 (L. Li)
                                         © 2022 Copyright for this paper by its authors. Use permitted under Creative Commons License
                                                                                                                                        prediction.
                                         Attribution 4.0 International (CC BY 4.0).



CEUR
                  ceur-ws.org
Workshop      ISSN 1613-0073
Proceedings
  RankTower follows the user-item decoupling paradigm,
enabling efficient online serving by pre-computing and
caching user and item embeddings. During online serv-
ing, only the gated cross-attention layers require forward                  Add & LayerNorm             Add & LayerNorm

propagation, while other operations remain parameter-free,
optimizing computational efficiency.                                         Gated Attention             Gated Attention
                                                                                  Unit                        Unit

                                                                              Q         KV                   Q      KV
2.1. Preliminary
The dataset for building the pre-ranking model consists
of instances (𝑥𝑢 , 𝑥𝑖 , 𝑦, 𝑝), where 𝑥𝑢 and 𝑥𝑖 are the user fea-
ture and item feature respectively, 𝑋𝑈 and 𝑋𝐼 are the user
and item input embeddings obtained by concatenating re-                         User                          Item
spective feature embedding vectors, 𝑦 ∈ {0, 1} indicates the              Multi-Embedding              Multi-Embedding

user-item binary feedback label, 𝑝 is the logged ranking
model prediction that for knowledge distillation. 𝑧 and 𝑦̂         Figure 3: The Architecture of Gated Cross-Attention Network
denote the pre-ranking model’s logit and prediction.

2.2. Multi-Head Gated Network
The Multi-Head Gated Network is an enhanced MLP aug-
mented with a gating mechanism for extracting diverse user           ℰ𝑢 = LN(𝐸𝑢 + GAU(𝑄 = 𝐸𝑢 , 𝐾 = 𝐸𝑖 , 𝑉 = 𝐸𝑖 )) ∈ ℝ𝐵×𝐻𝑢 ×𝑘
and item representations. The MLP output is multiplied by                                                                   (3)
an instance-aware gating vector, modeled by a two-layer
MLP. The input embedding does not receive gradients from             ℰ𝑖 = LN(𝐸𝑖 + GAU(𝑄 = 𝐸𝑖 , 𝐾 = 𝐸𝑢 , 𝑉 = 𝐸𝑢 )) ∈ ℝ𝐵×𝐻𝑖 ×𝑘
the gating network during training for stability. For exam-                                                                    (4)
ple, given an user input embedding 𝑋𝑈 , the ℎ-th sub-space
                                                                     The cross-attention mechanism with two parallel
𝑒𝑢ℎ of the user multi-embedding is:
                                                                   branches is designed to simultaneously attend to both user
             𝑒𝑢ℎ = 𝑀𝐿𝑃𝑢 (𝑋𝑈 )ℎ ∘ 𝜎(𝑔𝑀𝐿𝑃𝑢 (𝑋𝑈 ))ℎ                   preferences and item characteristics. This bidirectional ap-
                                                            (1)    proach allows the model to capture user-item interactions
                ∈ ℝ𝐵×𝑘 ,     ℎ = 1, ⋯ , 𝐻𝑢                         more accurately. The overall structure of the Gated Cross-
    where ∘ denotes the Hadamard product, 𝜎 denotes the ac-        Attention Network is illustrated in Figure 3.
tivation function of the gating network: Sigmoid(𝑥), 𝑀𝐿𝑃𝑢
denotes the MLP layer for modeling the user input embed-           2.3.2. Gated Attention Unit
ding, 𝑔𝑀𝐿𝑃𝑢 denotes the gating MLP for facilitating selec-
                                                                   The Gated Attention Unit introduces a gating mechanism
tive attention, 𝐵 is the batch size and 𝑘 is the embedding
                                                                   to facilitate selective attention for better learning the de-
size of each sub-space.
                                                                   pendency between user embedding and item embedding.
    Similarly, for item input embedding 𝑋𝐼 , the ℎ-th sub-space
                                                                   Specifically, the Gated Attention Unit effectively enables an
𝑒𝑖ℎ of the item multi-embedding is:
                                                                   attentive gating mechanism as follows:
              𝑒𝑖ℎ = 𝑀𝐿𝑃𝑖 (𝑋𝐼 )ℎ ∘ 𝜎(𝑔𝑀𝐿𝑃𝑖 (𝑋𝐼 ))ℎ                                  𝑄 = 𝜙(𝑋𝑄 𝑊𝑄 ), 𝐾 = 𝜙(𝑋𝐾 𝑊𝐾 )
                                                            (2)
                 ∈ ℝ𝐵×𝑘 ,     ℎ = 1, ⋯ , 𝐻𝑖                                                                                    (5)
                                                                                   𝑉 = 𝜙(𝑋𝑉 𝑊𝑉 ), 𝑈 = 𝜎(𝑋𝑄 𝑊𝑈 )
    In the offline processing stage, we will periodically batch
                                                                      where 𝑋𝑄 , 𝑋𝐾 , 𝑋𝑉 are the query, key, and value input, 𝜙
inference and store all the user/item’s embeddings 𝑒𝑢ℎ and
                                                                   is the non-linear activation function for projection layer, 𝜎
𝑒𝑖ℎ into the vector database for online serving usage.
                                                                   is the sigmoid function for computing gating value. With
                                                                   the learned projection 𝑄, 𝐾, 𝑉, and the gating value 𝑈, we
2.3. Gated Cross-Attention Network                                 compute the attention weights, followed by gating and a
The Gated Cross-Attention Network employs the cross-               post-attention projection.
attention mechanism to effectively model the interaction
between user embedding and item embedding. It utilizes the                              𝑂 = (𝑈 ⊙ 𝐴𝑉 )𝑊𝑜                        (6)
Gated Attention Unit as the main building block, along with
residual connections and layer normalization for training                                              𝑄𝐾 𝑇
                                                                                        𝐴 = softmax(          )                (7)
stability.
                                                                                                       √𝑑𝑘

2.3.1. Cross Attention Mechanism                                     where 𝐴 ∈ ℝ𝐻𝑢 ×𝐻𝑖 contains user to item attention weights.
                                                                   This example assumes that we use user embedding as the
The Bi-Directional Gated Cross-Attention Network inter-
                                                                   query, and item embedding as key and value.
changeably utilizes user and item embedding as queries
and keys-values for bi-directional attention. Specifically,
                                                      𝐻
with the user multi-embedding 𝐸𝑢 = Concat(e1u , ..., 𝑒𝑢 𝑢 ) and    2.4. Maximum Similarity Layer
                                             𝐻
item multi-embedding 𝐸𝑖 = Concat(e1i , ..., 𝑒𝑖 𝑖 ), the cross-     The Maximum Similarity Layer computes the final proba-
attention compute the user attended embedding ℰ𝑢 and               bility prediction based on the user and item attended em-
item attended embedding ℰ𝑖 as follows:                             beddings. Specifically, each user sub-space computes the
                                                                                   Fine-Grained                      Coarse-Grained
                                           Distillation Loss
                                                                                   Ranking Loss                       Ranking Loss




                                                   ...                                 ...                                  ...




                                                                          Pre-Ranking Model



                      User        Item 1           ...         Item N   Item N+1       ...        Item N+M   Item N+M+1     ...     Item N+M+L




                                             Impression                             Candidate                             Random
                                              Samples                               Samples                               Samples




Figure 4: The Synergistic Framework for Learning User Behavior Ordering and Full-Stage Sample Ordering




maximum cosine similarity with all item sub-spaces, and                               positive user feedback, while negative samples are items
the scalar outputs are summed across user sub-spaces:                                 that have been exposed to the user without receiving user
                                                                                      feedback.
                 𝐻𝑢
                                               𝑝         𝑞
           𝑠 = ( ∑ Max 𝐶𝑂𝑆𝐼 𝑁 𝐸(ℰ𝑢 , ℰ𝑖 ))/𝜏                            (8)
                𝑝=1 𝑞∈{1,⋯,𝐻𝑖 }                                                       3.1.2. Candidate Samples
    where 𝑝 and 𝑞 are the sub-space indexes of user-attended                          The item candidates in the ranking or pre-ranking stages
embedding and item-attended embedding, respectively, and                              that are not viewed by the user are categorized based on their
𝜏 is the learnable temperature scalar for re-scaling the cosine                       progression through the cascade ranking pipeline. Ranking
similarity. Note that the Maximum Similarity Layer does not                           candidates, which have advanced to the ranking stage, are
have any parameters which is suitable for online serving.                             generally considered as hard negative samples due to their
                                                                                      higher relevance and quality compared to the pre-ranking
                                                                                      candidates. Pre-ranking candidates are regarded as rela-
3. Pre-Ranking Model Optimization                                                     tively easy negative samples because they were filtered out
                                                                                      before reaching the ranking stage.
The pre-ranking models trained exclusively on impression
samples, same as ranking models, suffer from sample selec-
tion bias. The pre-ranking model, which operates on the                               3.1.3. Random Samples
outputs of recall models, aims to identify the most relevant                          Items that are randomly sampled from the item corpus to
candidates set for the ranking model. Consequently, align-                            serve as negative samples. These random samples are con-
ing the item distribution between the training and serving                            sidered the easiest negative samples but are included to
phases is essential to mitigate this sample selection bias and                        further enhance the generalization capability of the pre-
improve model effectiveness.                                                          ranking model. The incorporation of random samples en-
   As illustrated in Figure 4, we implemented full-stage sam-                         sures that the model remains effective and adaptable when
pling to draw training data from impression samples, can-                             encountering previously unseen items during the serving
didate samples, and random samples to mitigate sample                                 phase, thereby improving its robustness and ability to han-
selection bias. Moreover, we strategically applied various                            dle diverse item distributions.
distillation and learning-to-rank losses to different sample
scopes to effectively learn the ordering of user behaviors
and the sequencing of the sample stages.
                                                                                      3.2. Label Aggregation
                                                                                      Our framework incorporates two types of labels: hard la-
3.1. Full-Stage Sampling                                                              bels and soft labels. Hard labels represent various types of
                                                                                      positive user feedback on impression samples, while soft
The RankTower model is trained using user-level listwise                              labels are predictions made by the ranking models, used
samples containing multiple positive items and multiple                               knowledge distillation. Both categories of labels require an
objectives. The training samples for each user are sourced                            aggregation function to consolidate the different user behav-
from various stages of the cascade ranking system, as shown                           iors into a single scalar value for the pre-ranking model’s
in Figure 1. Detailed definitions and relationships among                             learning.
these components are provided below:
                                                                                      3.2.1. Hard Labels
3.1.1. Impression Samples
                                                                                      The aggregation of hard labels is highly dependent on the
The items output by the ranking model and viewed by the                               specific business problem, requiring that labels be aggre-
user consist of both positive and negative samples. Posi-                             gated according to their orders of importance.
tive samples are items that have received various types of
   For instance, in online advertising, eCPM can be uti-          model to the pre-ranking model is conducted exclusively
lized based on the pricing model of the platform. In an           on impression samples. As the ranking model is trained
e-commerce context, one might establish a relative prefer-        solely on these impression samples, its ability to generalize
ence order based on the depth of user feedback, such as           to candidate samples and random samples is inherently
Purchase > Add to Cart > Click. For scenarios like feed rank-     limited.
ing or video recommendations, user feedback signals can
be aggregated using a weighted sum approach. Addition-            3.3.2. Fine-Grained Ranking Loss
ally, we incorporate a general impression label applicable
across business scenarios, for learning the pattern of the        The fine-grained ranking loss is applied to both impression
cascade ranking system. The label assigned a value of 1 for       and candidate samples, which directly correspond to the
impression samples and 0 otherwise.                               sample scope used in serving. We employ the SoftSort, a
   The user feedback labels help the pre-ranking model in         differentiable sorting loss, to learn user behavior and the pat-
learning the revenue or engagement level associated with          terns of the cascade ranking system. This loss function aims
different user behaviors. The exposure label facilitates the      to precisely rank items according to the varying degrees
pre-ranking model’s ability to learn and replicate the rank-      of positive feedback they receive and effectively differenti-
ing patterns in the downstream cascade ranking system.            ate positives from impression samples and negatives from
                                                                  candidate samples.
                                                                     Consider the SoftSort operator defined by metric func-
3.2.2. Soft Labels
                                                                  tion d = | ⋅ |𝑝 and temperature parameter 𝜏 for sorting 𝑛-
For soft labels, we use the ranking objective function as         dimensional real vectors 𝑠 ∈ ℝ𝑛 :
aggregation function. This approach ensures that the soft
labels are seamlessly integrated into the training process,
maintaining the consistency between the pre-ranking model                                             −d(𝑠𝑜𝑟𝑡(𝑠)1𝑇 , 1𝑠 𝑇 )
                                                                          SoftSort𝑑𝜏 (𝑠) = softmax(                         )   (10)
and the ranking model.                                                                                        𝜏
                                                                     The output of SoftSort operator is a permutation matrix
3.3. Hybrid Loss Functions                                        of dimension 𝑛. The softmax operator is applied row-wise,
                                                                  thereby relaxing the permutation matrices into a set of uni-
The pre-ranking model focuses on achieving both consis-
                                                                  modal row-stochastic matrices. In simple words: the 𝑟-th
tency and ranking accuracy through the following tech-
                                                                  row of the SoftSort operator is the softmax of the negative
niques:
                                                                  distances to the 𝑟-th largest element [1]. We then employ the
     • Knowledge Distillation: The ranking model’s predic-        softmax cross entropy between the permutation matrices of
       tions are used as soft labels to transfer knowledge        label 𝑦 and the permutation matrices of logit 𝑧. The SoftSort
       from the ranking model (teacher) to the pre-ranking        loss function is hereby defined as:
       model (student).
     • Ranking Capability: Fine-grained and coarse-                ℒSorting (𝑧, 𝑦) = −tr(J𝑛 (SoftSort𝑑𝜏 (𝑦) ∘ log SoftSort𝑑𝜏 (𝑧)))
       grained ranking losses are applied to improve rank-                                                                     (11)
       ing performance and retrieval capability, respec-            where J𝑛 is a 𝑛 × 𝑛 matrix of ones, y = (𝑦𝑖 )𝑖∈𝒟ℐ ∪𝒟𝒞 is the
       tively.                                                    hard label and z = (𝑧𝑖 )𝑖∈𝒟ℐ ∪𝒟𝒞 is the logit of the pre-ranking
                                                                  model. We use the tr to compute the element sum of the
     • Diverse Training Samples: The model is trained on          matrix SoftSort𝑑𝜏 (𝑦) ∘ log (SoftSort𝑑𝜏 (𝑧)).
       samples from different stages and varying easy/hard
       sample levels to achieve robust generalization and
                                                                  3.3.3. Coarse-Grained Ranking Loss
       optimize hierarchical objectives.
                                                                  The coarse-grained ranking loss, applied to all samples (im-
   Our synergistic framework is designed to learn both the        pression, candidate, and random), aims to separate positive
hierarchy of user behaviors and the pattern of the cascade        and negative samples while supporting ranking among pos-
ranking system. For instance, in the context of online adver-     itives by distinguishing varying degrees of relevance. We
tising, the model is expected to understand the following         propose the Adaptive Margin Rankmax (AM-Rankmax) loss,
order of importance: converted items > clicked items > ex-        an extension of the Rankmax loss [2] that introduces an
posed items > candidate items and randomly sampled items.         adaptive margin based on the pair’s nature and label dis-
                                                                  tance, thereby extending the Rankmax loss to address rank-
3.3.1. Distillation Loss                                          ing problems with ordered or continuous positive labels.
                                                                     Consider the Rankmax loss for ranking problems with
As the main goal for the pre-ranking model is to output a
                                                                  binary labels only:
high-quality item set for the ranking model, hence we used
a listwise loss for distilling the knowledge from the ranking
model as follows:                                                                                    𝑛
                                                                       ℒ𝑅𝑎𝑛𝑘𝑚𝑎𝑥 (𝑧, 𝑦) = ∑ log ∑ (𝑧𝑖 − 𝑧𝑗 + 1)+                 (12)
                                                                                          𝑗∶𝑦𝑗 >0   𝑖=1
                                             exp(𝑧𝑖 )
       ℒDistillation (𝑧, 𝑝) = − ∑ 𝑝𝑖 log                    (9)     The Rankmax loss is reminiscent of pairwise losses. To
                              𝑖∈𝒟𝐼
                                           ∑𝑗∈𝒟𝐼 exp(𝑧𝑗 )
                                                                  extend the Rankmax loss to more general ranking prob-
   where 𝑝 is the prediction of the ranking model (soft label),   lems involving multi-level positive labels, we introduce the
𝑧 is the logit of the pre-ranking model, 𝒟ℐ is the impression     adaptive margin with following enhancements:
samples set. Note the distillation process from the ranking
     • The loss is applied only when 𝑦𝑖 < 𝑦𝑗 , which is more                  • Q2: How do different settings on dataset sampling
       suitable for multi-level positive label scenario.                        and training losses influence the performance of
                                                                                RankTower?
     • The margin adjusts based on whether 𝑦𝑖 is positive
       or negative, to further enhancing the differentiality
       between positive and negative samples.                         4.1. Experiment Setup
                                                                      4.1.1. Datasets
     • The margin scales with the label distances between
       samples, reflecting varying degree of positive sam-            We evaluate our model using real-world datasets: Ali-
       ples.                                                          mama1 , Taobao2 , and KuaiRand3 . For each dataset, we
                                                                      keep users with at least 100 impressions and 20 instances
  The adaptive margin function is:                                    of positive feedback. The data is split into 70% for training,
                                                                      10% for validation, and 20% for testing. As all labels in the
                𝑚(𝑖, 𝑗) = 𝛼 ⋅ I(𝑦𝑖 = 0) + 𝛿(𝑦𝑖 , 𝑦𝑗 )         (13)    datasets are binary, we aggregate them by summing the
  where 𝛼 is a constant for adding additional margin be-              labels to form the hard label.
tween negative and positive items, I is the indicator function.
The metric function 𝛿 can take various forms, for example             4.1.2. Evaluation Metrics
𝛿(𝑦𝑖 , 𝑦𝑗 ) = 1 or 𝛿(𝑦𝑖 , 𝑦𝑗 ) = 𝛽|𝑦𝑖 − 𝑦𝑗 |𝑝 . The adaptive margin
                                                                      We consider Recall@K and NDCG@K for evaluating the
Rankmax loss is then given by:
                                                                      performance of the models, and we set 𝑘 to 100 for all ex-
                                                                      periment metrics.
 𝐿AM−Rankmax (𝑧, 𝑦) = ∑ log ∑ (𝑧𝑖 − 𝑧𝑗 + 𝑚(𝑖, 𝑗))+                       Recall@K is the fraction of relevant retrieved within
                         𝑗∶𝑦𝑗 >0     𝑖∶𝑦𝑖 <𝑦𝑗                         the top 𝐾 recommendations. It’s mainly used for measuring
                                                        (14)          ranking system’s capability on retrieving relevant items.
  where y = (𝑦𝑖 )𝑖∈𝒟ℐ ∪𝒟𝒞 ∪𝒟ℛ is the hard label from all the             NDCG@K measures the quality of the ranking by con-
samples and z = (𝑧𝑖 )𝑖∈𝒟ℐ ∪𝒟𝒞 ∪𝒟ℛ is the logit of the pre-            sidering both the relevance and the position of items within
ranking model.                                                        the top 𝐾 recommendations. Items with higher relevance
  The AM-Rankmax loss function can effectively adapt to               ranked at higher position contribute more to the metric.
scenarios with multiple positive labels of varying levels.
This enhancement allows the model to handle different de-             4.1.3. Competing Models
grees of positive feedback, thereby improving its ability to
generalize and accurately rank items in complex settings.             We compare RankTower with the following pre-ranking
                                                                      models: LR [3], Two-Tower [4], DAT [5], COLD [6], Int-
                                                                      Tower [7] and ARF[8].
3.3.4. The Hybrid Ranking Loss
We design a hybrid ranking loss that integrates both distil-          4.2. Model Performance Comparison (Q1)
lation and ranking objectives. The hybrid ranking loss is
the weighted sum of three losses:
                                                                      Table 1
            ℒHybrid (𝑧, 𝑦) = 𝜆1 ℒDistillation (𝑧, 𝑝)                  Performance Comparison of Different Algorithms on Alimama ,
                            + 𝜆2 ℒSorting (𝑧, 𝑦)              (15)    Taobao and KuaiRand Dataset.
                                                                                             Alimama                Taobao               KuaiRand
                            + 𝜆3 ℒAM−Rankmax (𝑧, 𝑦)                         Model     Recall@K NDCG@K       Recall@K NDCG@K        Recall@K NDCG@K
                                                                              LR       0.4802      0.3237    0.4792       0.2685    0.6713      0.5027
                                                                          Two-Tower    0.5123      0.3428    0.5019       0.2921    0.6902      0.5258
   where 𝜆1 , 𝜆2 and 𝜆3 are weights for each sub-objective.                  DAT       0.5161      0.3472    0.5089       0.3013    0.6955      0.5312
Balancing distillation and ranking losses is crucial for the                COLD
                                                                           IntTower
                                                                                       0.5210
                                                                                       0.5215
                                                                                                   0.3518
                                                                                                   0.3519
                                                                                                             0.5123
                                                                                                             0.5101
                                                                                                                          0.3070
                                                                                                                          0.3051
                                                                                                                                    0.7011
                                                                                                                                    0.6960
                                                                                                                                                0.5349
                                                                                                                                                0.5309
pre-ranking model to inherit the ranking model’s capabili-                   ARF       0.5318      0.3655    0.5215       0.3117    0.7096      0.5497
                                                                          RankTower    0.5462      0.3794    0.5301       0.3223    0.7182      0.5551
ties while generalizing to broader sample spaces. Weighting
fine-grained and coarse-grained ranking losses ensures a
                                                                         The overall performance of different model architectures
balance between precise ranking and overall retrieval ro-
                                                                      is listed in Table 1. We have the following observations for
bustness.
                                                                      model effectiveness:

4. Experiments                                                        • LR exhibits the lowest performance compared to the other
                                                                        neural network-based models.
We conduct experiments on three large-scale public datasets
                                                                      • Two-Tower brings the most significant relative improve-
from online advertising, e-commerce, and short video rec-
                                                                        ment in performance, highlighting the importance of
ommendation domains to evaluate the effectiveness of Rank-
                                                                        learning deep feature interactions.
Tower. The experiments provide a comprehensive descrip-
tion of the evaluation metrics, and comparisons with state-           • COLD achieves strong performance among the compet-
of-the-art pre-ranking models. We aim to answer the fol-                ing models, indicating the significance of learning user-
lowing questions through our experiments:                               item feature interactions.
     • Q1: How does our proposed RankTower perform
       for pre-ranking task? Is it effective and efficient
       under extremely high-dimensional and sparse data               1
                                                                        https://tianchi.aliyun.com/dataset/408
                                                                      2
       settings?                                                        https://tianchi.aliyun.com/dataset/649
                                                                      3
                                                                        https://kuairand.com/
• ARF outperform other models without utilizing listwise         4.3.3. Effect of Distillation from Ranking Model
  ranking losses, highlighting the importance of using list-
                                                                 We conduct an ablation study on the distillation component
  wise ranking losses.
                                                                 and further compare Softmax loss with other alternatives.
• RankTower achieves the best prediction performance,              The Table 4 demonstrate the efficacy of transferring
  attributed to its effective modeling of bi-directional user-   knowledge distillation. Among various loss function ex-
  item feature interactions and the design of full-stage sam-    perimented for distillation, the Softmax loss outperforms
  pling and hybrid loss functions.                               the other alternative losses. The Softmax loss, being a list-
                                                                 wise ranking loss, proved more adept at distilling the rank-
                                                                 ing model’s capabilities compared to the weighted logloss,
4.3. Model Study (Q2)
                                                                 which essentially is a pointwise approach and exhibited sub-
To gain deeper insights into the proposed model, we conduct      optimal performance in learning the relative ranking distri-
experiments on the KuaiRand dataset and compare model            bution. In contrast, the pairwise logloss, focusing solely on
performance on different settings, including: 1) the effect      pairwise ordering of ranking model’s predictions without
of full-stage data sampling; 2) the effect of listwise ranking   considering the relative proximity of predictions, exhibited
losses; and 3) the effect of distillation from the ranking       overfitting to the ranking model’s outputs.
model.
                                                                 Table 4
4.3.1. Effect of Full-Stage Sampling                             Experiment Results for Different Distillation Losses.

We conduct an ablation study to evaluate the impact of                                                 Recall@K     NDCG@K
each sampling component on the model’s performance. As             Distillation (Softmax)               0.7182       0.5551
shown in Table 2, the full-stage sampling strategy achieves        Distillation (Weighted Logloss)      0.7130       0.5519
                                                                   Distillation (Pairwise Logloss)      0.7071       0.5432
the best overall performance. Training the pre-ranking
                                                                   No Distillation                      0.7108       0.5495
model solely with impression samples hinders its ability to
generalize to unexposed items, negatively affecting retrieval
performance. We also observe that candidate samples are
more important than random samples, as they significantly        5. Conclusion
enhance the model’s ability to discriminate between relevant
and non-relevant items.                                          This paper introduces the RankTower model, designed to
                                                                 enhance the performance of the two-tower model by effec-
Table 2                                                          tively capturing bi-directional latent interactions between
Experiment Results for Different Sampling Strategies.            user and item. To ensure consistency with existing casecade
                                       Recall@K    NDCG@K        ranking system, a hybrid loss function and full-stage sam-
 Full-Stage Sampling                    0.7182      0.5551       pling approach are integrated into the model’s optimization
 w/o random samples                     0.7125      0.5437       framework. Comprehensive experiments demonstrate that
 w/o candidate samples                  0.7040      0.5401       RankTower significantly outperforms state-of-the-art pre-
 w/o candidate & random samples         0.6981      0.5323       ranking models. In future work, we aim to study how to
                                                                 effectively and jointly optimize the cascade ranking system
                                                                 in an end-to-end fashion.
4.3.2. Effect of Listwise Ranking Losses
To better understand the properties of the proposed hybrid       References
loss, we compare it with several widely used ranking losses
in the industry. The experiment results, as shown in Table 3,    [1] S. Prillo, J. Eisenschlos, Softsort: A continuous relax-
indicate that the hybrid loss consistently outperforms other         ation for the argsort operator, in: International Confer-
alternatives, surpassing both its individual components: the         ence on Machine Learning, PMLR, 2020, pp. 7793–7802.
Sorting loss and the AM-Rankmax loss. Moreover, our pro-         [2] W. Kong, W. Krichene, N. Mayoraz, S. Rendle, L. Zhang,
posed AM-Rankmax demonstrates superior performance                   Rankmax: An adaptive projection alternative to the
compared to the original Rankmax loss and the Softmax                softmax function, Advances in Neural Information
loss.                                                                Processing Systems 33 (2020) 633–643.
                                                                 [3] H. B. McMahan, G. Holt, D. Sculley, M. Young, D. Ebner,
                                                                     J. Grady, L. Nie, T. Phillips, E. Davydov, D. Golovin,
Table 3
Experiment Results for Different Ranking Losses.
                                                                     et al., Ad click prediction: a view from the trenches, in:
                                                                     Proceedings of the 19th ACM SIGKDD international
                            Recall@K     NDCG@K                      conference on Knowledge discovery and data mining,
          Hybrid Loss        0.7182       0.5551                     2013, pp. 1222–1230.
          Sorting            0.7128       0.5516
                                                                 [4] P.-S. Huang, X. He, J. Gao, L. Deng, A. Acero, L. Heck,
          AM-Rankmax         0.7132       0.5507
          Rankmax            0.7105       0.5492
                                                                     Learning deep structured semantic models for web
          Softmax            0.7109       0.5498                     search using clickthrough data, in: Proceedings of
          ApproxNDCG         0.7006       0.5436                     the 22nd ACM international conference on Information
          RankNet            0.7072       0.5452                     & Knowledge Management, 2013, pp. 2333–2338.
                                                                 [5] Y. Yu, W. Wang, Z. Feng, D. Xue, A dual augmented two-
                                                                     tower model for online large-scale recommendation
                                                                     (2021).
[6] Z. Wang, L. Zhao, B. Jiang, G. Zhou, X. Zhu, K. Gai, Cold:
    Towards the next generation of pre-ranking system,
    arXiv preprint arXiv:2007.16122 (2020).
[7] X. Li, B. Chen, H. Guo, J. Li, C. Zhu, X. Long, S. Li,
    Y. Wang, W. Guo, L. Mao, et al., Inttower: the next gen-
    eration of two-tower model for pre-ranking system, in:
    Proceedings of the 31st ACM International Conference
    on Information & Knowledge Management, 2022, pp.
    3292–3301.
[8] Y. Wang, Z. Wang, J. Yang, S. Wen, D. Kong, H. Li,
    K. Gai, Adaptive neural ranking framework: Toward
    maximized business goal for cascade ranking systems,
    arXiv preprint arXiv:2310.10462 (2023).