<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Hybrid Model with Time Modeling for Sequential Recommender Systems</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Marlesson R. O. Santana</string-name>
          <email>marlessonsa@gmail.com</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Recommender systems, Session-based recommendations, Recurrent</string-name>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Anderson Soares</string-name>
          <email>anderson@inf.ufg.br</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Deep Learning Brazil - Federal University of Goiás</institution>
          ,
          <addr-line>Goiânia, Goiás</addr-line>
          ,
          <country country="BR">Brazil</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Deep Learning Brazil - Federal University of Goiás</institution>
          ,
          <addr-line>Goiânia, Goiás</addr-line>
          ,
          <country country="BR">Brazil</country>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>neural networks</institution>
          ,
          <addr-line>Hybrid model, Time modeling</addr-line>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2021</year>
      </pub-date>
      <fpage>53</fpage>
      <lpage>57</lpage>
      <abstract>
        <p>Deep learning based methods have been used successfully in recommender system problems. Approaches using recurrent neural networks, transformers, and attention mechanisms are useful to model users' long- and short-term preferences in sequential interactions. To explore diferent session-based recommendation solutions, Booking.com recently organized the WSDM WebTour 2021 Challenge, which aims to benchmark models to recommend the final city in a trip. This study presents our approach to this challenge. We conducted several experiments to test diferent state-of-the-art deep learning architectures for recommender systems. Further, we proposed some changes to Neural Attentive Recommendation Machine (NARM), adapted its architecture for the challenge objective, and implemented training approaches that can be used in any sessionbased model to improve accuracy. Our experimental result shows that the improved NARM outperforms all other state-of-the-art benchmark methods.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>CCS CONCEPTS</title>
      <p>• Computer systems organization → Embedded systems; •
Information systems → Recommender systems; • Computing
methodologies → Sequential decision making.</p>
    </sec>
    <sec id="sec-2">
      <title>INTRODUCTION</title>
      <p>
        Deep learning approaches for recommender systems have garnered
significant attention owing to their potential for modeling
longand short-term user preferences. Models using recurrent neural
networks (RNNs), transformers, and attention mechanisms are handy
for text, presenting encouraging results when used in session-based
recommendation problems [
        <xref ref-type="bibr" rid="ref15 ref16 ref17 ref2 ref6 ref8">2, 6, 8, 15–17</xref>
        ], where sequential
information and short-term context are fundamental in recommending
the next item of the session.
      </p>
      <p>
        In general, session-based approaches use only the information
from the item’s interaction sequence in the session as a predictor of
the next item. However, in some domains, contextual information
within and between sessions and global information are essential
to modeling user behavior. Recommending travel destinations has
several particularities, such as the sequence of cities on a trip can
contain noise, financial constraints may lead to diversions, and
destinations are highly correlated to the moment in time and duration
of the trip [
        <xref ref-type="bibr" rid="ref1 ref11 ref7">1, 7, 11</xref>
        ] .
      </p>
      <p>
        The WSDM WebTour 2021 Challenge [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ] organized by
Booking.com, focuses on recommending travel destinations in a session.
This study describes our approach for the WSDM WebTour 2021
Challenge and proposes some approaches that can improve
sessionbased models for recommender systems, especially with the highly
time-dependent recommendation domains, noise information, and
imbalanced classes.
2
      </p>
    </sec>
    <sec id="sec-3">
      <title>THE CHALLENGE AND DATASET</title>
      <p>
        Booking.com is the world’s largest online travel agency. It is a
platform where millions of travelers find accommodations for their
trips, and millions of accommodation providers list their hotels,
apartments, guest houses, and other lodgings. [
        <xref ref-type="bibr" rid="ref1 ref7">1, 7</xref>
        ].
      </p>
      <p>
        Booking.com recently organized the WSDM WebTour 2021
Challenge [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ]. The training dataset consists of over a million anonymized
hotel reservations based on real data. Each reservation is a part of
a customer’s trip, which includes at least four consecutive
reservations. The challenge’s goal is to recommend the final city of each
trip, and we evaluated models using an accuracy metric for the
ifrst four items suggested. For Accuracy@4, the metric value is one
when the real city is one of the four main suggestions and zero
otherwise.
3
      </p>
    </sec>
    <sec id="sec-4">
      <title>PROPOSED APPROACHES</title>
      <p>This section describes each approach used in the present study to
improve our final model. First, we chose a state-of-the-art
sessionbased model to improve (described in 3.1). Further, we created new
features from statistical information about users and cities and
added time modeling to focus on the travel problem particularity
(described in 3.2 and 3.3). Also, we have reduced the efect of the
imbalance dataset using specific loss functions and multitask
modeling (described in 3.5 and 3.4). Finally, we improve the generalization
of noise information through data augmentation (described in 3.6).</p>
      <p>Each approach can be applied separately in other session-based
models.
3.1</p>
    </sec>
    <sec id="sec-5">
      <title>Model Architecture</title>
      <p>
        We use the architecture of the neural attentive recommendation
machine (NARM) [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ] and adapt it for the traveling problem
presented in this study. NARM uses an encoder-decoder architecture
to address the session-based recommendations problems.
According to the authors of the paper, the idea of NARM is to build a
hidden representation of the current session using an RNN module.
It converts the input click sequence  = [1, 2, ...,  ] into a set of
high-dimensional hidden representation with an attention signal
that can be used to produce a ranking list of all items that can occur
in the next step of the current session.
      </p>
      <p>Our approach uses categorical and dense features of the user,
city, and trip combined with the trip history. The core of the NARM
module is the same as the original paper, and we just changed the
size of the inputs, the bottleneck with hidden representation, and
the outputs.</p>
      <p>The features are concatenated and follow two paths. The first
group of features passes through an attention layer before being
fed into the NARM module. It is an important step for relating
diferent positions of the same input sequence. In the second path,
the features bypass the feature bottleneck generated by the NARM
module, which improves the decoder by providing it with more
contextual information on the session.</p>
      <p>Figure 1 shows the final architecture proposed with
modifications described in this study.
3.2</p>
    </sec>
    <sec id="sec-6">
      <title>Feature Engineer</title>
      <p>Our approach uses statistical features from users, cities, and trips
combined with the trip’s history information. For every trip, we
extract more than 30 statistical features. Below we provide examples
of the important features we use in the present study:
User Statistics
• Number of trips
• Number of cities visited
• Average number of unique cities visited per trip
• Average trip size
• Average trip duration
• Most frequently traveled month
• Number of unique cities the user visited
• Number of unique countries the user visited
• Device used for booking
• Booker country
City Statistics
• City interactions that have preceded the current trip
• City interactions that have preceded the current trip by user
• City interactions in the current trip
• Country interactions that have preceded the current trip
• Country interactions that have preceded the current trip by
user</p>
      <p>Furthermore, we employ user statistical features to create user
embedding from an autoencoder trained on the same training split
data. This approach proved helpful because only 0.7% of the users
have made more than one trip; the vast majority are new users. In
addition to providing a dense representation of the user features,
the autoencoder ensures that similar users are close to each other
in the created vector space. Moreover, the autoencoder was trained
separately from the recommendation model. Figure 2 shows user
embeddings colored by the most frequently traveled month.
3.3</p>
    </sec>
    <sec id="sec-7">
      <title>Time Modeling</title>
      <p>In terms of tourism, a trip’s location is merely a piece of information
that can tell how good the place is for a trip. The date (moment
in time) and duration of the trip are as crucial as, for example,
sometimes, the traveler can only visit certain places only in specific
seasons or specific time windows. Sometimes, it is not feasible to
stay for a few days more.</p>
      <p>
        Embeddings are highly useful for modeling categorical features
such as the location, and RNNs are remarkably good for capturing
sequential data’s time-dependence. However, they are limited when
the events have irregular time intervals, such as duration of stay on
a specific city trip. Therefore, we need to incorporate time interval
explicitly into the model as a time-space embedding. In this case,
we use the approach presented in [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ] concatenated with a start trip
month embedding for time-space embedding.
      </p>
      <p>Therefore, each city embedding presented in a trip follows through
a time embedding layer. We incorporate time (start trip month) and
duration (how long users stayed in a city) information. In the end,
we obtain a space-time embedding.
Hybrid Model with Time Modeling for Sequential Recommender Systems
3.4</p>
    </sec>
    <sec id="sec-8">
      <title>Multi-Task Learning</title>
      <p>
        The data about the cities in our dataset are imbalanced and have
high dimensionality. In contrast, the data about countries are less
imbalanced and have lesser dimensionality. Thus, we adapt our
model to predict both to try to improve models’ gradient signal.
According to [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ], multi-task learning acts as a regularizer by
introducing an inductive bias into the model. As such, it reduces the
risk of overfitting and the Rademacher complexity of the model,
and thus, it has the ability to fit random noise.
      </p>
      <p>We used two targets in the present study. The city of the next
step interaction was the main target that we used to evaluate the
model, and the city’s country was used as the second target. The
second target was used only for regularization and as inductive
bias in the primary target. Both targets use the same loss function,
and the final loss is a combination of the two.
3.5</p>
    </sec>
    <sec id="sec-9">
      <title>Loss Function</title>
      <p>
        Focal loss [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ] is very useful for training imbalanced datasets. It
adds a weighted term in front of the cross-entropy loss to balance
the gradient from positive and negative samples. Easily classified
negatives comprise most of the loss and dominate the gradient; this
efect can be reduced by using a focal loss strategy. We use focal
loss for both targets, and the final loss is the average of each loss.
      </p>
      <p>Formally, the focal loss is expressed as follows:</p>
      <p>() = − (1 − ) ()
where  is adjusts the rate at which easy examples are
downweighted and  is a prefixed value between 0 and 1 to balance
the positive-labeled samples and negative-labeled samples. We use
Equation 2 for calculation loss.</p>
      <p>( ,  ) =   ( ) + (1 − )  ( )
where  is a balance parameter, and  and  are obtained from
the target prediction.
(1)
(2)
3.6</p>
    </sec>
    <sec id="sec-10">
      <title>Data Augmentation</title>
      <p>
        Data augmentation techniques have been widely used to enhance
image [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ], text [
        <xref ref-type="bibr" rid="ref18">18</xref>
        ], or recommendations based models [
        <xref ref-type="bibr" rid="ref15 ref4">4, 15</xref>
        ].
The principal inputs of the model in the challenge are cities in the
current session’s history. These inputs can have noise, some users
may take a long or short trip, and users can jump one or more cities
on a popular route. We applied three diferent data augmentations
steps to improve the accuracy and generalize the model.
      </p>
      <p>
        The first is a sequence preprocessing step proposed in [
        <xref ref-type="bibr" rid="ref15 ref4">4, 15</xref>
        ],
where we generate new samples using each trip’s time-step. Given
an input training trip [1, 2, 3, ...,  ], we generate the sequences
and corresponding labels ( [1], 2), ( [1, 2], 3), ( [1, 2, ..., −1],  )
for training. We filter only trips with more than four cities.
      </p>
      <p>For each sample in training, we randomly choose a step in the trip
to change. Given an input training trip [1, 2, 3, 4], we change it in
three ways: remove a step as a dropout layer and generate sequences
such as [1, 2, 4]; replace with a mask token (unknown) and
generate sequences with the same size but with mask [1, 2, 3,    ]
; and replace with a similar city, such as [1, 2, 3, 4] where 2 and
2 are similar cities. The mask token will be the same as that used
in production/inference mode when the model does not embed a
specific city, and the similarity definition is the same as that we
used for the Item-KNN model.</p>
      <p>We apply both methods to make the model less susceptible to
overfitting or noise information.
4</p>
    </sec>
    <sec id="sec-11">
      <title>EXPERIMENTS</title>
      <p>In this section, we present the experimental settings and results.
4.1</p>
    </sec>
    <sec id="sec-12">
      <title>Dataset</title>
      <p>
        We randomly partition the training dataset by trip, using 90% of
data for training and 10% for model validation. We filtered trips
with less than four cities or more than 10 and trips with a duration
of more than 22 days in the training split. These values came from
the initial analysis and are outliers. Table 1 shows the statistics of
the dataset used in our experiments.
• Popularity: Popular predictor recommends the most
popular city in last city’s country on the current trip.
• Item-KNN: This baseline recommends the most similar city
to the last city on current trip. The similarity is defined as
the cosine distance between the trip vector of the city. It is
similar to the co-occurrence approach. [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]
• Caser: This baseline is a convolutional neural network (CNN)
approach for a sequential recommendation. The
Convolutional Sequence Embedding Recommendation Model (Caser)
embeds a sequence of recent items into an "image" in time
and learns sequential patterns as local features. [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ]
• SASRec: SASRec is a self-attentive approach that captures
users’ sequential behaviors and achieves state-of-the-art
performance on sequential recommendation. [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ]
• NARM: Neural Attentive Session-based Recommendation
is an encoder-decoder architecture with an attention
mechanism to model the user’s sequential behavior, which is then
combined as a unified session representation later. [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ]
      </p>
      <p>We choose the NARM model, which shows the best performance
among the models listed above, such as our baseline model to
improve during the competition. Therefore, our approach is an
adaptation of the NARM model. Some approaches were modeled
for the click prediction problem. For our experiments, we adapted
the last layer of all deep learning models to produce a ranking list
of all items,  = [1, 2, 3, ...,  ], that can occur on the current trip.
Furthermore, we trained our model using the same loss function.
4.3</p>
    </sec>
    <sec id="sec-13">
      <title>Experimental Settings</title>
      <p>All our models were trained using similar parameters. We used
50-dimensional embeddings for all category features, and we
normalized dense features using standard deviation. Optimization was
carried out using Rectified Adam (RAdam), with a 0.001 learning
rating and 0.01 weight decay, with mini-batch size fixed at 64. We
truncated the trip history size using a fixed window of 10 time-steps
with padding. The number of epochs was defined by early stopping
using 10 steps, and we used cross-entropy as a loss function for
most models or focal loss with 1  and 3  .</p>
      <p>
        Finally, we use the MARS-Gym framework [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ] to model, train,
and evaluate all experiments described in this paper. All
experiments are present and can be reviewed at https://github.com/
marlesson/booking_challenge.
4.4
      </p>
    </sec>
    <sec id="sec-14">
      <title>Experimental Results</title>
      <p>Each user’s embeddings are presented in the 2-D plane in Figure 2.
We can see that the month is a good predictor of user behavior, and,
in general, users who travel in the same month also share
similarities. We can also notice that there is a mixed cluster in the center,
maybe belonging to users who are likely to travel on diferent dates
with greater frequency. Furthermore, we can use this representation
for new users to improve a cold-start recommendation.</p>
      <p>Each model’s performance on the test dataset is summarized in
Table 2. We can see that the Item-KNN method performs similarly
to the Popularity method, which is a cold start method. Therefore,
there is no benefit in adopting the Item-KNN method given the
complexity posed by cold recommendations.</p>
      <p>SASRec and Caser are very diferent approaches, the former
based on attention mechanism and later on horizontal and vertical
convolutions. Both showed similar results, with Caser performing
slightly better. However, the best base baseline model was NARM.
We can see that the original NARM outperformed state-of-the-art
baselines using the same input information. Therefore, we chose
NARM in this study to improve the architecture and training phase.</p>
      <p>We evaluate two versions of NARM. In NARM V1, we use only
improvements that were applied in the training and regularization
steps, with the same input data were used for other baseline models,
but applying approaches 3.3, 3.4, 3.5, 3.6. NARM V1 outperforms
the original NARM with approximately +4.62% accuracy. This is
a promising finding, as this indicates that these techniques can
now be used for any other session-based models that are in need of
improvement.</p>
      <p>Finally, we apply all approaches present in this study to NARM
V2, and it represents our final approach for the challenge. Figure 1
shows our final architecture. We found that NARM V2 outperforms
the original NARM with approximately +9.66% accuracy. This
supports the idea that when it is not possible to obtain an item or user
metadata, session statistics can also be used as features.
5</p>
    </sec>
    <sec id="sec-15">
      <title>CONCLUSION</title>
      <p>In this study, we present our approach for the WSDM WebTour
2021 Challenge. We conducted several experiments using diferent
session-based models to recommend the next destination on a trip.
We modified the existing NARM model to add contextual
information to the session, space-time modeling, and approaches to reduce
the negative efect of class imbalance in the training phase. Our
results show that it is possible to enhance the performance of
stateof-the-art models through simple changes. The improved NARM
outperforms all baseline models, and that the implemented training
approaches can be used in any session-based recommendations
model.
Hybrid Model with Time Modeling for Sequential Recommender Systems</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>Lucas</given-names>
            <surname>Bernardi</surname>
          </string-name>
          , Themistoklis Mavridis, and
          <string-name>
            <given-names>Pablo</given-names>
            <surname>Estevez</surname>
          </string-name>
          .
          <year>2019</year>
          .
          <article-title>150 successful machine learning models: 6 lessons learned at booking. com</article-title>
          .
          <source>In Proceedings of the 25th ACM SIGKDD International Conference on Knowledge Discovery &amp; Data Mining</source>
          .
          <fpage>1743</fpage>
          -
          <lpage>1751</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>Qiwei</given-names>
            <surname>Chen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Huan</given-names>
            <surname>Zhao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Wei</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Pipei</given-names>
            <surname>Huang</surname>
          </string-name>
          , and
          <string-name>
            <given-names>Wenwu</given-names>
            <surname>Ou</surname>
          </string-name>
          .
          <year>2019</year>
          .
          <article-title>Behavior sequence transformer for e-commerce recommendation in alibaba</article-title>
          .
          <source>In Proceedings of the 1st International Workshop on Deep Learning Practice for High-Dimensional Sparse Data. 1-4.</source>
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>James</given-names>
            <surname>Davidson</surname>
          </string-name>
          , Benjamin Liebald, Junning Liu, Palash Nandy, Taylor Van Vleet,
          <string-name>
            <surname>Ullas Gargi</surname>
          </string-name>
          , Sujoy Gupta,
          <string-name>
            <surname>Yu</surname>
            <given-names>He</given-names>
          </string-name>
          , Mike Lambert,
          <string-name>
            <given-names>Blake</given-names>
            <surname>Livingston</surname>
          </string-name>
          , et al.
          <year>2010</year>
          .
          <article-title>The YouTube video recommendation system</article-title>
          .
          <source>In Proceedings of the fourth ACM conference on Recommender systems. 293-296.</source>
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>Alexandre</given-names>
            <surname>De Brébisson</surname>
          </string-name>
          , Étienne Simon, Alex Auvolat, Pascal Vincent, and
          <string-name>
            <given-names>Yoshua</given-names>
            <surname>Bengio</surname>
          </string-name>
          .
          <year>2015</year>
          .
          <article-title>Artificial neural networks applied to taxi destination prediction</article-title>
          .
          <source>arXiv preprint arXiv:1508.00021</source>
          (
          <year>2015</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>Dmitri</given-names>
            <surname>Goldenberg</surname>
          </string-name>
          , Kostia Kofman, Pavel Levin, Sarai Mizrachi, Maayan Kafry, and
          <string-name>
            <given-names>Guy</given-names>
            <surname>Nadav</surname>
          </string-name>
          .
          <year>2021</year>
          .
          <article-title>Booking.com WSDM WebTour 2021 Challenge</article-title>
          .
          <source>In ACM WSDM Workshop on Web Tourism (WSDM WebTour'21).</source>
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <surname>Wang-Cheng Kang</surname>
          </string-name>
          and
          <string-name>
            <surname>Julian McAuley</surname>
          </string-name>
          .
          <year>2018</year>
          .
          <article-title>Self-attentive sequential recommendation</article-title>
          .
          <source>In 2018 IEEE International Conference on Data Mining (ICDM)</source>
          . IEEE,
          <fpage>197</fpage>
          -
          <lpage>206</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>Julia</given-names>
            <surname>Kiseleva</surname>
          </string-name>
          , Melanie JI Mueller, Lucas Bernardi, Chad Davis, Ivan Kovacek, Mats Stafseng Einarsen, Jaap Kamps, Alexander Tuzhilin, and
          <string-name>
            <given-names>Djoerd</given-names>
            <surname>Hiemstra</surname>
          </string-name>
          .
          <year>2015</year>
          .
          <article-title>Where to go on your next trip? Optimizing travel destinations based on user preferences</article-title>
          .
          <source>In Proceedings of the 38th International ACM SIGIR Conference on Research and Development in Information Retrieval</source>
          .
          <fpage>1097</fpage>
          -
          <lpage>1100</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>Jing</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Pengjie</given-names>
            <surname>Ren</surname>
          </string-name>
          , Zhumin Chen, Zhaochun Ren, Tao Lian, and Jun Ma.
          <year>2017</year>
          .
          <article-title>Neural attentive session-based recommendation</article-title>
          .
          <source>In Proceedings of the 2017 ACM on Conference on Information and Knowledge Management</source>
          .
          <fpage>1419</fpage>
          -
          <lpage>1428</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <surname>Yang</surname>
            <given-names>Li</given-names>
          </string-name>
          ,
          <string-name>
            <given-names>Nan</given-names>
            <surname>Du</surname>
          </string-name>
          , and
          <string-name>
            <given-names>Samy</given-names>
            <surname>Bengio</surname>
          </string-name>
          .
          <year>2017</year>
          .
          <article-title>Time-dependent representation for neural event sequence prediction</article-title>
          .
          <source>arXiv preprint arXiv:1708.00065</source>
          (
          <year>2017</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <surname>Tsung-Yi</surname>
            <given-names>Lin</given-names>
          </string-name>
          , Priya Goyal, Ross Girshick, Kaiming He, and
          <string-name>
            <given-names>Piotr</given-names>
            <surname>Dollár</surname>
          </string-name>
          .
          <year>2017</year>
          .
          <article-title>Focal loss for dense object detection</article-title>
          .
          <source>In Proceedings of the IEEE international conference on computer vision</source>
          . 2980-
          <fpage>2988</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>Sarai</given-names>
            <surname>Mizrachi</surname>
          </string-name>
          and
          <string-name>
            <given-names>Pavel</given-names>
            <surname>Levin</surname>
          </string-name>
          .
          <year>2019</year>
          .
          <article-title>Combining Context Features in SequenceAware Recommender Systems</article-title>
          .. In
          <string-name>
            <surname>RecSys (Late-Breaking</surname>
            <given-names>Results</given-names>
          </string-name>
          ).
          <fpage>11</fpage>
          -
          <lpage>15</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>Sebastian</given-names>
            <surname>Ruder</surname>
          </string-name>
          .
          <year>2017</year>
          .
          <article-title>An overview of multi-task learning in deep neural networks</article-title>
          .
          <source>arXiv preprint arXiv:1706.05098</source>
          (
          <year>2017</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <surname>Marlesson</surname>
            <given-names>R. O.</given-names>
          </string-name>
          <string-name>
            <surname>Santana</surname>
          </string-name>
          , Luckeciano C. Melo,
          <string-name>
            <surname>Fernando H. F. Camargo</surname>
          </string-name>
          , Bruno Brandão, Anderson Soares,
          <string-name>
            <surname>Renan M. Oliveira</surname>
            , and
            <given-names>Sandor</given-names>
          </string-name>
          <string-name>
            <surname>Caetano</surname>
          </string-name>
          .
          <year>2020</year>
          .
          <article-title>MARSGym: A Gym framework to model, train, and evaluate Recommender Systems for Marketplaces</article-title>
          . arXiv:
          <year>2010</year>
          .
          <article-title>07035 [cs</article-title>
          .IR]
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <given-names>Connor</given-names>
            <surname>Shorten and Taghi M Khoshgoftaar</surname>
          </string-name>
          .
          <year>2019</year>
          .
          <article-title>A survey on image data augmentation for deep learning</article-title>
          .
          <source>Journal of Big Data</source>
          <volume>6</volume>
          ,
          <issue>1</issue>
          (
          <year>2019</year>
          ),
          <fpage>1</fpage>
          -
          <lpage>48</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <given-names>Yong</given-names>
            <surname>Kiam</surname>
          </string-name>
          <string-name>
            <given-names>Tan</given-names>
            , Xinxing
            <surname>Xu</surname>
          </string-name>
          , and Yong Liu.
          <year>2016</year>
          .
          <article-title>Improved recurrent neural networks for session-based recommendations</article-title>
          .
          <source>In Proceedings of the 1st workshop on deep learning for recommender systems</source>
          .
          <volume>17</volume>
          -
          <fpage>22</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [16]
          <string-name>
            <given-names>Jiaxi</given-names>
            <surname>Tang</surname>
          </string-name>
          and
          <string-name>
            <given-names>Ke</given-names>
            <surname>Wang</surname>
          </string-name>
          .
          <year>2018</year>
          .
          <article-title>Personalized top-n sequential recommendation via convolutional sequence embedding</article-title>
          .
          <source>In Proceedings of the Eleventh ACM International Conference on Web Search and Data Mining</source>
          .
          <fpage>565</fpage>
          -
          <lpage>573</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [17]
          <string-name>
            <surname>Maksims</surname>
            <given-names>Volkovs</given-names>
          </string-name>
          , Anson Wong, Zhaoyue Cheng, Felipe Pérez, Ilya Stanevich, and
          <string-name>
            <given-names>Yichao</given-names>
            <surname>Lu</surname>
          </string-name>
          .
          <year>2019</year>
          .
          <article-title>Robust contextual models for in-session personalization</article-title>
          .
          <source>In Proceedings of the Workshop on ACM Recommender Systems Challenge. 1-5.</source>
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          [18]
          <string-name>
            <given-names>Jason</given-names>
            <surname>Wei</surname>
          </string-name>
          and
          <string-name>
            <given-names>Kai</given-names>
            <surname>Zou</surname>
          </string-name>
          .
          <year>2019</year>
          .
          <article-title>Eda: Easy data augmentation techniques for boosting performance on text classification tasks</article-title>
          . arXiv preprint arXiv:
          <year>1901</year>
          .
          <volume>11196</volume>
          (
          <year>2019</year>
          ).
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>