<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Robust Training Objectives Improve Embedding-Based Retrieval in Industrial Recommendation Systems</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Matthew Kolodner</string-name>
          <email>mkolodner@snap.com</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Mingxuan Ju</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Zihao Fan</string-name>
          <email>zfan3@snap.com</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Tong Zhao</string-name>
          <email>tong@snap.com</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Elham Ghazizadeh</string-name>
          <email>eghazizadeh@snap.com</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Yan Wu</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Neil Shah</string-name>
          <email>nshah@snap.com</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Yozen Liu</string-name>
          <email>yliu2@snap.com</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Snap, Inc.</institution>
          ,
          <addr-line>2772 Donald Douglas Loop N, Santa Monica, CA 90405</addr-line>
          ,
          <country country="US">USA</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>Improving recommendation systems (RS) can greatly enhance the user experience across many domains, such as social media. Many RS utilize embedding-based retrieval (EBR) approaches to retrieve candidates for recommendation. In an EBR system, the embedding quality is key. According to recent literature, self-supervised multitask learning (SSMTL) has showed strong performance on academic benchmarks in embedding learning and resulted in an overall improvement in multiple downstream tasks, demonstrating a larger resilience to the adverse conditions between each downstream task and thereby increased robustness and task generalization ability through the training objective. However, whether or not the success of SSMTL in academia as a robust training objectives translates to large-scale (i.e., over hundreds of million users and interactions in-between) industrial RS still requires verification. Simply adopting academic setups in industrial RS might entail two issues. Firstly, many self-supervised objectives require data augmentations (e.g., embedding masking/corruption) over a large portion of users and items, which is prohibitively expensive in industrial RS. Furthermore, some self-supervised objectives might not align with the recommendation task, which might lead to redundant computational overheads or negative transfer. In light of these two challenges, we evaluate using a robust training objective, specifically SSMTL, through a large-scale friend recommendation system on a social media platform in the tech sector, identifying whether this increase in robustness can work at scale in enhancing retrieval in the production setting. Through online A/B testing with SSMTL-based EBR, we observe statistically significant increases in key metrics in the friend recommendations, with up to 5.45% improvements in new friends made and 1.91% improvements in new friends made with cold-start users. Besides, with a dedicated case study, the benefits of robust training objectives are demonstrated through SSMTL on large-scale graphs with gains in both retrieval and end-to-end friend recommendation.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>
        Recommendation systems (RS) have become a crucial
component for user experience [
        <xref ref-type="bibr" rid="ref1">1, 2</xref>
        ]. Most industrial RS
explore a two-stage process [3]. During the first stage (i.e., the
retrieval phase), among hundreds of millions of candidate
users/items, the RS usually utilizes several models optimized
for recall to select a small set of candidate users/items (e.g.,
1,000 candidates). Whereas during the second stage (i.e., the
ranking phase), within the candidate subset, the RS can
explore complicated expensive models that are optimized for
precision to select top  candidates for the final
recommendation. Such two-stage process enables recommendation
over large quantities of possible users/items and allows for
greater flexibility towards key recommendation metrics.
      </p>
      <p>
        In this two-stage scheme, the retrieval stage is especially
important, as it acts as the bottleneck for possible candidates
provided to the ranker in the second stage. One common
approach [
        <xref ref-type="bibr" rid="ref2 ref3">4, 5</xref>
        ] for the retrieval step is to leverage
embeddingbased retrieval (EBR). Specifically, EBR learns embeddings
for all users and items as vectors in a low-dimensional latent
space. These embeddings are learned in a way such that
the distance between them is reflective of their similarity,
with more similar items being closer together in the latent
space. As a result, candidates can be retrieved through a
nearest-neighbor search across the latent space. In practice,
this is done using an approximate nearest neighbor methods
optimized for large-scale retrieval, such as FAISS [
        <xref ref-type="bibr" rid="ref4">6</xref>
        ] and
HNSW [7].
      </p>
      <p>
        Many methods [
        <xref ref-type="bibr" rid="ref5">8, 9, 10, 11</xref>
        ] have been proposed for
generating high-quality embeddings for EBR, which lead to
more relevant candidates and improved metrics after the
end-to-end recommendation. In this work, we specifically
focus on the friend recommendation EBR setting, where
vast amounts of topological information relating users are
readily available. Recent works [
        <xref ref-type="bibr" rid="ref6 ref7 ref8">12, 13, 14</xref>
        ] have shown
that including this relational information can improve the
embedding quality. The relational information is commonly
modeled with graph neural networks (GNNs), producing
embeddings that leverage neighbor information in graphs,
such as co-friend relationships. For graph-aware EBR in
particular, link prediction has seen success for generating
high-quality embeddings [
        <xref ref-type="bibr" rid="ref9">15</xref>
        ], where we look to predict
the presence of an edge between a query node and set of
candidate nodes.
      </p>
      <p>
        While link prediction is efective in learning nuanced
similarities and distinctions between candidates, there are
several other self-supervised graph learning philosophies that
can provide high-quality embeddings, such as mutual
information maximization [
        <xref ref-type="bibr" rid="ref10">16</xref>
        ], generative reconstruction [
        <xref ref-type="bibr" rid="ref11">17</xref>
        ],
or whitening decorrelation [
        <xref ref-type="bibr" rid="ref12">18</xref>
        ]. Based on these general
philosophies, many graph-based approaches have been
proposed and used to learning embeddings directly, achieving
desirable properties of embeddings without requiring
explicit labels. Recently, Ju et al. [
        <xref ref-type="bibr" rid="ref13">19</xref>
        ] evaluated combining
these self-supervised learning approaches with link
prediction in a multitask (MTL) setting, demonstrating a larger
resilience to the adverse conditions between each downstream
task and thereby increased robustness and generalization
ability through the training objective
      </p>
      <p>
        However, whether or not using SSMTL in academia as
a robust training objective translates to large-scale (i.e.,
over hundreds of millions of users and interactions
inbetween) industrial RSs still requires verification. Simply
adopting academic setups in industrial RSs might result
in several issues. Firstly, many self-supervised objectives
require data augmentations (e.g., embedding
masking/corruption) over a large portion of users and items, which
is prohibitively expensive in industrial RSs. Furthermore,
some self-supervised objectives might not align with the
recommendation task, which might lead to redundant
computational overheads or negative transfer [
        <xref ref-type="bibr" rid="ref14">20</xref>
        ], a phenomenon
where performance can worsen as a result of the complexity
and potentially opposing nature of the various tasks.
      </p>
      <p>
        In this work, we investigate whether robust SSMTL
training objectives are able to improve the link prediction
retrieval performance on large-scale graphs with over
hundreds of millions of nodes and edges. Specifically, we look to
ifnd what combination of SSL approaches can improve
overall robustness and thereby augment retrieval through
complementary yet disjoint information. In our experiments,
we find two SSL approaches, based on philosophies from
whitening decorrelation (e.g., Canonical Correlation
Analysis [
        <xref ref-type="bibr" rid="ref15">21</xref>
        ]) and generative reconstruction (e.g., Masked
Autoencoders [
        <xref ref-type="bibr" rid="ref16">22</xref>
        ]), that are able to augment the performance
of link prediction without negative transfer. We deploy
the proposed framework on an industrial large-scale friend
recommendation system to a community of hundreds of
millions of users. In the online A/B testing, we observe
significant improvements in key metrics like new friends
made, especially with cold-start users on the platform. Our
contributions are summarized as follows:
• We demonstrate the efectiveness of robust training
objectives such as SSMTL in a large-scale industrial
recommendation system.
• We conduct an online study of SSMTL on a massive
real-world recommendation system, and observe a
statistically significant increase in key metrics, with
up to 5.45% improvements in new friends made
and 1.91% improvements in new friends made with
cold-start users.
      </p>
    </sec>
    <sec id="sec-2">
      <title>2. Background</title>
      <sec id="sec-2-1">
        <title>2.1. Graph-Aware Embedding-based</title>
      </sec>
      <sec id="sec-2-2">
        <title>Retrieval</title>
        <p>
          In a two-stage recommendation system with a retrieval then
ranking phase, the retrieval phase plays an important role
filtering out the most relevant candidates to lighten the load of
the ranker. Since the ranking result is largely dependent on
items retrieved in the retrieval phase, a good quality retrieval
model can drastically improve the final ranking. Embedding
based retrieval (EBR) is a method that’s recently adopted
and deployed in many content, product, and friend
recommendation systems[
          <xref ref-type="bibr" rid="ref17 ref18 ref2 ref6">4, 23, 24, 12</xref>
          ], and proved to achieve
superior results. EBR transform users and items into
embeddings, changing the retrieval problem into a
nearestneighbor search problem in a low-dimensional latent space.
These embeddings can be determined in advance and
indexed using an approximate nearest neighbor search such
as FAISS [
          <xref ref-type="bibr" rid="ref4">6</xref>
          ] and HNSW [7] in order to retrieve the top-
most relevant items eficiently at serving.
        </p>
        <p>
          When applying EBR to RS problems, the quality of
embeddings is of upmost importance. In this paper, we use a
friend recommendation system as our subject. In scenarios
like friend recommendation where vast amounts of
topological information relating users and items is readily available,
these embeddings can be augmented with GNNs. Previous
work showed that EBR for friend recommendation systems
see benefits leveraging graph-aware embeddings[
          <xref ref-type="bibr" rid="ref6">12</xref>
          ]. In
this setting, nodes would contain individual user features
while edges map to user-user interactions. This approach
compliments commonly used graph traversal approaches
(eg. friend-of-friend (FoF) [
          <xref ref-type="bibr" rid="ref19">25</xref>
          ]), allowing for retrieval of
candidates from any number of hops away from the target.
        </p>
        <p>
          Here we describe GNNs for generating graph-aware
embeddings for EBR. GNNs have demonstrated state-of-the-art
performance in many problems containing rich topological
information within the graph data [
          <xref ref-type="bibr" rid="ref20">26</xref>
          ], such as
recommendation and forecasting. Formally, we define  = (, ℰ , ),
 where  ∈
        </p>
        <p>
          R× 
where  is the set of  nodes (|| = ), ℰ is the set of
edges (ℰ ∈  ⊆  ), and  is a feature matrix of dimension
. Many modern GNNs also employ
a message-passing structure, consisting of an aggregation
(AGG) and update (UPD) function. The goal of this paradigm
is for nodes to receive information from their neighbors,
collecting messages using its AGG function before updating
their own messages with the UPD function, both of which
are learnable and permutation-invariant. For some node 
at layer , the next message-passing layer can be written as
h(+1) = UPD() (︁
h(), AGG() (︁
{h(), ∀ ∈  ()}
where  () is the neighborhood nodes of node .
Different message-passing GNN models use diferent
combinations of AGG and UPD functions. An example of a more
complex GNN, Graph-attention networks (GATs) [
          <xref ref-type="bibr" rid="ref21">27</xref>
          ], use
an attention mechanism for each pair of nodes  and 
︁)
(1)
  = softmax (att (Wℎ, Wℎ ))
(2)
where W is a linear transformation applied to every node
and att is the attention function parameterized by a weight
vector and a non-linearity function. The AGG function is
then a attention-weighted sum of its neighbors features
while the UPD function is implicitly defined in
W and the
non-linearity function. Typically, to generate graph-aware
embeddings from GNNs, a margin based ranking loss[
          <xref ref-type="bibr" rid="ref6 ref7">13, 12</xref>
          ]
or contrastive[
          <xref ref-type="bibr" rid="ref22">28</xref>
          ] loss can be used, to encourage items that
are closer in the graph to be closer in the embedding space.
        </p>
      </sec>
      <sec id="sec-2-3">
        <title>2.2. Multitask Learning</title>
        <p>
          Multitask learning (MTL) is an approach in machine
learning where a model is trained simultaneously on several tasks.
MTL has been extensively explored in recommendation as a
way to improve key metrics [
          <xref ref-type="bibr" rid="ref23 ref24 ref25 ref26">29, 30, 31, 32</xref>
          ]. Thus, the core
idea behind multitask learning is to improve the robustness
of the model by leveraging the domain-specific information
contained in the training signals of related tasks [
          <xref ref-type="bibr" rid="ref27 ref28">33, 34</xref>
          ].
Hard parameter sharing, one of the most fundamental forms
of MTL, uses a shared representation which then branches
into multiple heads capable of learning task-specific
information [
          <xref ref-type="bibr" rid="ref29 ref30 ref31">35, 36, 37</xref>
          ].
        </p>
        <p>
          For graph-aware EBR in particular, self-supervised
multitask learning (SSMTL) has been proposed as a new approach
to MTL, optimizing the embeddings directly to achieve
desirable embedding properties without the use of positive
or negative labels. In this setting, we combine several
selfsupervised learning (SSL) methods with a downstream
retrieval task to learn both direct and indirect embedding
features. Recent work [
          <xref ref-type="bibr" rid="ref13">19</xref>
          ] has shown that SSMTL can lead to
improved task generalization and embedding quality on
several academic benchmarks through the increasingly robust
training objective. However, many of the SSL approaches
used are constrained to the assumption that global graph
information can be inferred within the graph structure. This is
not valid in the large-scale recommendation setting, where
graphs are constrained to some -hop around a query user
in order to fit in memory. As a result, many of these SSL
methods may lead to negative transfer due to SSL task
conlfict with the target link prediction task, and there remains
work to be done to investigate which methods perform best
in this large-scale setting.
        </p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>3. Self-Supervised Multitask</title>
    </sec>
    <sec id="sec-4">
      <title>Learning for EBR</title>
      <p>In the following sections, we describe details of the SSL
methods used in our SSMTL approach, our experiment set
up and results, highlighting the benefits and impact of
including SSMTL based embedding in EBR for large-scale
industrial recommendation systems.</p>
      <sec id="sec-4-1">
        <title>3.1. Self-Supervised Learning Methods</title>
        <p>We identify two self-supervised learning approaches that
are scalable and lead to improvements in the large-scale
recommendation setting through a more robust training
objective.</p>
        <p>
          Canonical Correlation Analysis. Based on work from
[
          <xref ref-type="bibr" rid="ref15">21</xref>
          ], Canonical Correlation Analysis (CCA) deploys a
noncontrastive, non-discriminitive SSL method to train the
GNN. The self-supervised training objective is described in
Equation 3. First, given a subgraph with  nodes, two
augmented views of the subgraph are created and fed through
the GNN, producing ZA and ZB where ZA, ZB ∈ R× .
Each of these embeddings are fed through a task-specific
head, and then are normalized so that each feature has 0
mean and √1 standard deviation, resulting in Z˜A and Z˜B.
The loss is then computed from Equation 3. The first term
in the equation seeks to minimize the distance of the same
nodes between the two views. The second term enforces
that the feature-wise covariance of all nodes is equal to the
identity matrix.
ℒCCA = ⃦⃦ Z˜ −
⃦
        </p>
        <p>Z˜ ⃦⃦ 2
⃦ 
+</p>
        <p>⃦
︂( ⃦⃦ Z˜ Z˜ − I⃦
⃦ 2
⃦</p>
        <p>
          ⃦
+ ⃦⃦ Z˜ Z˜ − I⃦
⃦ 2 )︂
⃦ 
(3)
Masked Autoencoders. Based on work from [
          <xref ref-type="bibr" rid="ref16">22</xref>
          ], this
approach leverages a graph masked autoencoder (MAE)
that focuses on feature reconstruction. First, an augmented
view of the subgraph is created and the features of the query
users are masked out. This augmented graph is then fed
through the GNN and a task-specific head. The features of
the query users are then re-masked and passed through a
graph convolution layer. As described in Equation 4, for
all masked nodes , the final loss is equal to the average
of the scaled cosine error between the original features X
and generated features Z. This approach only relies on the
local neighborhood surrounding the query node, making it
a good option for large-scale SSMTL.
        </p>
        <p>ℒMAE =
1
∑︁</p>
        <p>︂(
|| ∈
1
−</p>
        <p>x z
‖x‖ · ‖ z‖
︂) 
,  ≥ 1
(4)
We note that these two approaches both utilize
noncontrastive methods. While experimenting with diferent
SSL tasks, we find that contrastive SSL approaches do not
perform very well in the production setting due to their
assumption that global information is readily available in
the original and augmented graphs. This is not necessarily
true for large-scale recommendation, where subgraphs are
constrained to the K-hop neighborhood surrounding each
query node.</p>
      </sec>
      <sec id="sec-4-2">
        <title>3.2. Experimental Setup</title>
        <sec id="sec-4-2-1">
          <title>3.2.1. Problem Breakdown</title>
          <p>We evaluate the SSMTL as a robust training objective on an
industrial friend recommendation system with hundreds of
millions of users and connections. To handle this scale of
training, we sample subgraphs containing the -hop
neighborhood around each query user. Following training, the
embeddings for EBR can be via propagation through the
encoder backbone.</p>
        </sec>
        <sec id="sec-4-2-2">
          <title>3.2.2. Retrieval Baseline</title>
          <p>The baseline model uses a supervised single-task setup for
embedding-based retrieval. We use a GAT as the GNN
encoder backbone to obtain embeddings for the query user and
each candidate, producing a candidate embedding matrix z.
We can then compute the dot product between the query
user and each candidate and apply Softmax to generate the
logits. We then calculate the Categorical Cross Entropy
Loss with the true labels y across the  = 2 classes and 
candidates, outlined in Equation 5.</p>
          <p>ℒretrieval = −
 
∑︁ ∑︁  log</p>
          <p>︃(</p>
        </sec>
        <sec id="sec-4-2-3">
          <title>3.2.3. SSMTL Implementation Details</title>
          <p>In our SSMTL approach, we use both CCA and MAE in
combination with the retrieval baseline as the training
objectives. All three methods share the same GAT GNN backbone.
The augmented views for CCA and MAE occur separately,
with CCA performing edge and feature drop augmentations
while MAE performs edge drop and query node masking.
The task-specific head for CCA is a Linear-ReLU-Linear
block while the task-specific head for MAE is one linear
layer. The final loss with SSMTL is a weighted sum of the
losses.</p>
          <p>ℒcombined =  ℒretrieval +  ℒCCA +  ℒMAE
(6)
where  is the weight for the retrieval loss,  is the weight
of the CCA loss, and  is the weight of the MAE loss. In
practice, we observed best performance when the retrieval
weight was several orders of magnitude larger than the
other loss weights.</p>
        </sec>
      </sec>
      <sec id="sec-4-3">
        <title>3.3. Results</title>
        <p>We evaluated the efectiveness of SSMTL for end-to-end
friend recommendation with online A/B testing. The
control group used candidates retrieved from the production
model trained with retrieval baseline, while the treatment
group instead used candidates retrieved with the new robust
training objective in the SSMTL setting, specifically
combining the previous retrieval loss with whitening decorrelation
and generative reconstruction objectives.</p>
        <p>In the A/B experimental results, we saw statistically
significant improvements across several friend recommendation
metrics. Specifically, we observed up to
5.45%
improvements in new friends made and +1.91% new friends made
with low-degree users in various markets. Overall, from
these results, we see that SSMTL is able to provide improved
recommendation compared with the single-task setting, in
particular helping with candidate generation for low-degree
users.</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>4. Conclusion</title>
      <p>In this paper, we evaluate the efectiveness of a robust
selfsupervised multitask learning objective in embedding-based
retrieval. Through online evaluation, we demonstrate that
self-supervised methods used in a multi task setting are able
to augment the performance of the underlying retrieval task
on the scale of over 800 million nodes and edges, providing
complementary yet disjoint information to enhance the
embedding quality. We observe statistically significant gains
in the number of friendships made for both high and low
degree users.
2023. arXiv:2306.12680.
[2] A. Sun, Y. Peng, A survey on modern recommendation
system based on big data, 2024. arXiv:2206.02631.
[3] P. Covington, J. Adams, E. Sargin,
Deep neural
networks for youtube recommendations, in:
Proceedings of the 10th ACM Conference on
Recommender Systems, RecSys ’16, Association for
Computing Machinery, New York, NY, USA, 2016, p. 191–198.
URL: https://doi.org/10.1145/2959100.2959190. doi:10.
1145/2959100.2959190.
P. Pronin, J. Padmanabhan, G. Ottaviano, L. Yang,
Embedding-based retrieval in facebook search, CoRR
abs/2006.11632 (2020). URL: https://arxiv.org/abs/2006.
11632. arXiv:2006.11632.
Billion-scale
similarity search with gpus,
CoRR abs/1702.08734
(2017).</p>
      <p>URL:</p>
      <p>http://arxiv.org/abs/1702.08734.</p>
      <p>arXiv:1702.08734.
[7] Y. A. Malkov, D. A. Yashunin,
Eficient and
robust approximate nearest neighbor search using
hierarchical navigable small world graphs,
CoRR
abs/1603.09320 (2016). URL: http://arxiv.org/abs/1603.
09320. arXiv:1603.09320.
Divide and conquer: Towards better embedding-based
retrieval for recommender systems from a multi-task
perspective, 2023. arXiv:2302.02657.
[9] G. Linden, B. Smith, J. York, Amazon.com
recommendations: item-to-item collaborative filtering, IEEE
Internet Computing 7 (2003) 76–80. doi:10.1109/MIC.
based retrieval with llm for efective agriculture
information extracting from unstructured data, 2023.
arXiv:2308.03107.
N. Shah, P. Yu, N. Srivastava, L. Shi, G.
Venkataraman, J. Yu, Embedding based retrieval in friend
recommendation, in: Proceedings of the 46th
International ACM SIGIR Conference on Research and
De</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>Y.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Liu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Satapathy</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Wang</surname>
          </string-name>
          , E. Cambria,
          <article-title>Recent developments in recommender systems: A survey,</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>J.</given-names>
            <surname>Huang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Sharma</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Sun</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Xia</surname>
          </string-name>
          , D. Zhang,
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>Y.</given-names>
            <surname>Gan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Ge</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Zhou</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Su</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Xu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Xu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Q.</given-names>
            <surname>Hui</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Chen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Shan</surname>
          </string-name>
          ,
          <article-title>Binary embedding-based retrieval at tencent</article-title>
          ,
          <year>2023</year>
          . arXiv:
          <volume>2302</volume>
          .
          <fpage>08714</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>J.</given-names>
            <surname>Johnson</surname>
          </string-name>
          , M. Douze, H. Jégou, [8]
          <string-name>
            <given-names>Y.</given-names>
            <surname>Zhang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Dong</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W.</given-names>
            <surname>Ding</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Jiang</surname>
          </string-name>
          , K. Gai,
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>R.</given-names>
            <surname>Peng</surname>
          </string-name>
          , K. Liu,
          <string-name>
            <given-names>P.</given-names>
            <surname>Yang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Yuan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Li</surname>
          </string-name>
          , Embedding-
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>J.</given-names>
            <surname>Shi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Chaurasiya</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Liu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Vij</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Wu</surname>
          </string-name>
          , S. Kanduri, velopment in Information Retrieval, SIGIR '23,
          <string-name>
            <surname>Association</surname>
          </string-name>
          for Computing Machinery, New York, NY, USA,
          <year>2023</year>
          , p.
          <fpage>3330</fpage>
          -
          <lpage>3334</lpage>
          . URL: https://doi.org/10.1145/ 3539618.3591848. doi:
          <volume>10</volume>
          .1145/3539618.3591848.
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [13]
          <string-name>
            <given-names>R.</given-names>
            <surname>Ying</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>He</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Chen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Eksombatchai</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W. L.</given-names>
            <surname>Hamilton</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Leskovec</surname>
          </string-name>
          ,
          <article-title>Graph convolutional neural networks for web-scale recommender systems</article-title>
          , CoRR abs/
          <year>1806</year>
          .
          <year>01973</year>
          (
          <year>2018</year>
          ). URL: http://arxiv.org/abs/
          <year>1806</year>
          .
          <year>01973</year>
          . arXiv:
          <year>1806</year>
          .
          <year>01973</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [14]
          <string-name>
            <surname>P. P.-H. Kung</surname>
            ,
            <given-names>Z.</given-names>
          </string-name>
          <string-name>
            <surname>Fan</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          <string-name>
            <surname>Zhao</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          <string-name>
            <surname>Liu</surname>
            ,
            <given-names>Z.</given-names>
          </string-name>
          <string-name>
            <surname>Lai</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          <string-name>
            <surname>Shi</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          <string-name>
            <surname>Wu</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          <string-name>
            <surname>Yu</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          <string-name>
            <surname>Shah</surname>
          </string-name>
          , G. Venkataraman,
          <article-title>Improving embedding-based retrieval in friend recommendation with ann query expansion</article-title>
          ,
          <source>in: Proceedings of the 47th International ACM SIGIR Conference on Research and Development in Information Retrieval</source>
          ,
          <year>2024</year>
          , pp.
          <fpage>2930</fpage>
          -
          <lpage>2934</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [15]
          <string-name>
            <given-names>C.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Peng</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Niu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Zhang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Peng</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Zhou</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <article-title>Learning graph attention-aware knowledge graph embedding</article-title>
          ,
          <source>Neurocomputing</source>
          <volume>461</volume>
          (
          <year>2021</year>
          )
          <fpage>516</fpage>
          -
          <lpage>529</lpage>
          . URL: https://www.sciencedirect.com/science/article/ pii/S0925231221010961. doi:https://doi.org/10. 1016/j.neucom.
          <year>2021</year>
          .
          <volume>01</volume>
          .139.
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [16]
          <string-name>
            <surname>A.</surname>
          </string-name>
          v. d. Oord,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>O.</given-names>
            <surname>Vinyals</surname>
          </string-name>
          ,
          <article-title>Representation learning with contrastive predictive coding</article-title>
          , arXiv preprint arXiv:
          <year>1807</year>
          .
          <volume>03748</volume>
          (
          <year>2018</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [17]
          <string-name>
            <given-names>K.</given-names>
            <surname>He</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Chen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Xie</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Dollár</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Girshick</surname>
          </string-name>
          ,
          <article-title>Masked autoencoders are scalable vision learners</article-title>
          ,
          <source>in: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition</source>
          ,
          <year>2022</year>
          , pp.
          <fpage>16000</fpage>
          -
          <lpage>16009</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [18]
          <string-name>
            <given-names>A.</given-names>
            <surname>Ermolov</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Siarohin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Sangineto</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Sebe</surname>
          </string-name>
          ,
          <article-title>Whitening for self-supervised representation learning</article-title>
          ,
          <source>in: International Conference on Machine Learning, PMLR</source>
          ,
          <year>2021</year>
          , pp.
          <fpage>3015</fpage>
          -
          <lpage>3024</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [19]
          <string-name>
            <given-names>M.</given-names>
            <surname>Ju</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Zhao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Q.</given-names>
            <surname>Wen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W.</given-names>
            <surname>Yu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Shah</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Ye</surname>
          </string-name>
          ,
          <string-name>
            <surname>C.</surname>
          </string-name>
          <article-title>Zhang, Multi-task self-supervised graph neural networks enable stronger task generalization</article-title>
          ,
          <source>in: The Eleventh International Conference on Learning Representations</source>
          ,
          <year>2023</year>
          . URL: https://openreview.net/forum?id= 1tHAZRqftM.
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [20]
          <string-name>
            <given-names>L.</given-names>
            <surname>Torrey</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Shavlik</surname>
          </string-name>
          , Transfer Learning,
          <source>IGI Global</source>
          ,
          <year>2010</year>
          , pp.
          <fpage>242</fpage>
          -
          <lpage>264</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [21]
          <string-name>
            <given-names>H.</given-names>
            <surname>Zhang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Q.</given-names>
            <surname>Wu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Yan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Wipf</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P. S.</given-names>
            <surname>Yu</surname>
          </string-name>
          ,
          <article-title>From canonical correlation analysis to self-supervised graph neural networks</article-title>
          ,
          <source>CoRR abs/2106</source>
          .12484 (
          <year>2021</year>
          ). URL: https: //arxiv.org/abs/2106.12484. arXiv:
          <volume>2106</volume>
          .
          <fpage>12484</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [22]
          <string-name>
            <given-names>Z.</given-names>
            <surname>Hou</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Liu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Cen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Dong</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Yang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Tang</surname>
          </string-name>
          , Graphmae:
          <article-title>Self-supervised masked graph autoencoders</article-title>
          ,
          <year>2022</year>
          . arXiv:
          <volume>2205</volume>
          .
          <fpage>10803</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [23]
          <string-name>
            <given-names>P.</given-names>
            <surname>Covington</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Adams</surname>
          </string-name>
          , E. Sargin,
          <article-title>Deep neural networks for youtube recommendations</article-title>
          ,
          <source>in: Proceedings of the 10th ACM Conference on Recommender Systems</source>
          , New York, NY, USA,
          <year>2016</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          [24]
          <string-name>
            <given-names>T.</given-names>
            <surname>Koh</surname>
          </string-name>
          , G. Wu, , M. Mi,
          <article-title>Manas hnsw realtime: Powering realtime embedding-based retrieval</article-title>
          ,
          <year>2021</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          [25]
          <string-name>
            <given-names>M. E. J.</given-names>
            <surname>Newman</surname>
          </string-name>
          ,
          <article-title>Clustering and preferential attachment in growing networks</article-title>
          ,
          <source>Physical Review E</source>
          <volume>64</volume>
          (
          <year>2001</year>
          ). URL: http://dx.doi.org/10.1103/PhysRevE.64. 025102. doi:
          <volume>10</volume>
          .1103/physreve.64.025102.
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          [26]
          <string-name>
            <given-names>J.</given-names>
            <surname>Zhou</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Cui</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Zhang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Yang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Liu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Sun</surname>
          </string-name>
          ,
          <article-title>Graph neural networks: A review of methods and applications</article-title>
          , CoRR abs/
          <year>1812</year>
          .08434 (
          <year>2018</year>
          ). URL: http: //arxiv.org/abs/
          <year>1812</year>
          .08434. arXiv:
          <year>1812</year>
          .08434.
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          [27]
          <string-name>
            <given-names>P.</given-names>
            <surname>Veličković</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Cucurull</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Casanova</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Romero</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Liò</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Bengio</surname>
          </string-name>
          , Graph attention networks,
          <year>2018</year>
          . arXiv:
          <volume>1710</volume>
          .
          <fpage>10903</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          [28]
          <string-name>
            <given-names>Z.</given-names>
            <surname>Liu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Ma</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Ouyang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Xiong</surname>
          </string-name>
          ,
          <article-title>Contrastive learning for recommender system</article-title>
          ,
          <source>CoRR abs/2101</source>
          .01317 (
          <year>2021</year>
          ). URL: https://arxiv.org/abs/2101. 01317. arXiv:
          <volume>2101</volume>
          .
          <fpage>01317</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          [29]
          <string-name>
            <given-names>Y.</given-names>
            <surname>Lu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Dong</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Smyth</surname>
          </string-name>
          ,
          <article-title>Why i like it: multi-task learning for recommendation and explanation,</article-title>
          <year>2018</year>
          , pp.
          <fpage>4</fpage>
          -
          <lpage>12</lpage>
          . doi:
          <volume>10</volume>
          .1145/3240323.3240365.
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          [30]
          <string-name>
            <given-names>J.</given-names>
            <surname>Ma</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Zhao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Yi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Chen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Hong</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E. H.</given-names>
            <surname>Chi</surname>
          </string-name>
          ,
          <article-title>Modeling task relationships in multi-task learning with multi-gate mixture-of-experts</article-title>
          ,
          <source>in: Proceedings of the 24th ACM SIGKDD International Conference on Knowledge Discovery &amp; Data Mining, KDD '18</source>
          ,
          <string-name>
            <surname>Association</surname>
          </string-name>
          for Computing Machinery, New York, NY, USA,
          <year>2018</year>
          , p.
          <fpage>1930</fpage>
          -
          <lpage>1939</lpage>
          . URL: https://doi.org/10.1145/ 3219819.3220007. doi:
          <volume>10</volume>
          .1145/3219819.3220007.
        </mixed-citation>
      </ref>
      <ref id="ref25">
        <mixed-citation>
          [31]
          <string-name>
            <given-names>X.</given-names>
            <surname>Ma</surname>
          </string-name>
          , L.
          <string-name>
            <surname>Zhao</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          <string-name>
            <surname>Huang</surname>
            ,
            <given-names>Z.</given-names>
          </string-name>
          <string-name>
            <surname>Wang</surname>
            ,
            <given-names>Z.</given-names>
          </string-name>
          <string-name>
            <surname>Hu</surname>
            ,
            <given-names>X.</given-names>
          </string-name>
          <string-name>
            <surname>Zhu</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          <string-name>
            <surname>Gai</surname>
          </string-name>
          ,
          <article-title>Entire space multi-task model: An efective approach for estimating post-click conversion rate</article-title>
          ,
          <year>2018</year>
          . arXiv:
          <year>1804</year>
          .07931.
        </mixed-citation>
      </ref>
      <ref id="ref26">
        <mixed-citation>
          [32]
          <string-name>
            <given-names>H.</given-names>
            <surname>Tang</surname>
          </string-name>
          , J. Liu,
          <string-name>
            <given-names>M.</given-names>
            <surname>Zhao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Gong</surname>
          </string-name>
          ,
          <article-title>Progressive layered extraction (ple): A novel multi-task learning (mtl) model for personalized recommendations</article-title>
          ,
          <source>in: Proceedings of the 14th ACM Conference on Recommender Systems</source>
          , RecSys '20,
          <string-name>
            <surname>Association</surname>
          </string-name>
          for Computing Machinery, New York, NY, USA,
          <year>2020</year>
          , p.
          <fpage>269</fpage>
          -
          <lpage>278</lpage>
          . URL: https://doi.org/10.1145/3383313.3412236. doi:
          <volume>10</volume>
          . 1145/3383313.3412236.
        </mixed-citation>
      </ref>
      <ref id="ref27">
        <mixed-citation>
          [33]
          <string-name>
            <given-names>A.</given-names>
            <surname>Argyriou</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Evgeniou</surname>
          </string-name>
          ,
          <string-name>
            <surname>M.</surname>
          </string-name>
          <article-title>Pontil, Multi-task feature learning</article-title>
          , in: B.
          <string-name>
            <surname>Schölkopf</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          <string-name>
            <surname>Platt</surname>
          </string-name>
          , T. Hofman (Eds.),
          <source>Advances in Neural Information Processing Systems</source>
          , volume
          <volume>19</volume>
          , MIT Press,
          <year>2006</year>
          . URL: https: //proceedings.neurips.cc/paper_files/paper/2006/file/ 0afa92fc0f8a9cf051bf2961b06ac56b-Paper.pdf.
        </mixed-citation>
      </ref>
      <ref id="ref28">
        <mixed-citation>
          [34]
          <string-name>
            <given-names>R.</given-names>
            <surname>Caruana</surname>
          </string-name>
          ,
          <article-title>Multitask learning</article-title>
          ,
          <source>Machine Learning</source>
          <volume>28</volume>
          (
          <year>1997</year>
          )
          <fpage>41</fpage>
          -
          <lpage>75</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref29">
        <mixed-citation>
          [35]
          <string-name>
            <given-names>P.</given-names>
            <surname>Guo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.-Y.</given-names>
            <surname>Lee</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Ulbricht</surname>
          </string-name>
          ,
          <article-title>Learning to branch for multi-task learning</article-title>
          ,
          <year>2020</year>
          . arXiv:
          <year>2006</year>
          .
          <year>01895</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref30">
        <mixed-citation>
          [36]
          <string-name>
            <given-names>X.</given-names>
            <surname>Sun</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Panda</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R. S.</given-names>
            <surname>Feris</surname>
          </string-name>
          ,
          <article-title>Adashare: Learning what to share for eficient deep multi-task learning</article-title>
          , CoRR abs/
          <year>1911</year>
          .12423 (
          <year>2019</year>
          ). URL: http://arxiv.org/abs/
          <year>1911</year>
          . 12423. arXiv:
          <year>1911</year>
          .12423.
        </mixed-citation>
      </ref>
      <ref id="ref31">
        <mixed-citation>
          [37]
          <string-name>
            <given-names>S.</given-names>
            <surname>Vandenhende</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Georgoulis</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B. D.</given-names>
            <surname>Brabandere</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L. V.</given-names>
            <surname>Gool</surname>
          </string-name>
          ,
          <article-title>Branched multi-task networks: Deciding what layers to share</article-title>
          ,
          <year>2020</year>
          . arXiv:
          <year>1904</year>
          .02920.
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>