<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Cascaded Machine Learning Model for Eficient Hotel Recommendations from Air Travel Bookings</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Benoit Lardeux Amadeus SAS Sophia Antipolis</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>France</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Mourad Boudia Amadeus SAS Sophia Antipolis</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>France</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>France</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Rodrigo Acuna Agost Amadeus SAS Sophia Antipolis</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>France</string-name>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Amadeus SAS Sophia Antipolis</institution>
          ,
          <country country="FR">France</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2019</year>
      </pub-date>
      <fpage>9</fpage>
      <lpage>16</lpage>
      <abstract>
        <p>Recommending a hotel for vacations or a business trip can be a challenging task due to the large number of alternatives and considerations to take into account. In this study, a recommendation engine is designed to identify relevant hotels based on features of the facilities and the context of the trip via flight information. The system was designed as a cascaded machine learning pipeline, with a model to predict the conversion probability of each hotel and another to predict the conversion of a set of hotels as presented to the traveller. By analysing the feature importance of the model based on sets of hotels, we are able to construct optimal lists of hotels by selecting individual hotels that will maximise the probability of conversion.</p>
      </abstract>
      <kwd-group>
        <kwd>Recommender systems</kwd>
        <kwd>machine learning</kwd>
        <kwd>hotels</kwd>
        <kwd>conversion</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>CCS CONCEPTS</title>
      <p>• Computing methodologies →
Machine learning;</p>
    </sec>
    <sec id="sec-2">
      <title>INTRODUCTION</title>
      <p>
        In the United States, the travel industry is estimated to be the third
largest industry after the automotive and food sectors and
contributes to approximately 5% of the gross domestic product. Travel
has experienced rapid growth as users are willing to pay for new
experiences, unexpected situations, and moments of meditation
[
        <xref ref-type="bibr" rid="ref28 ref9">9, 28</xref>
        ], while the cost of travel has decreased over time in part due
to low cost carriers and the sharing economy. At the same time,
traditional travel players such as airlines, hotels, and travel
agencies, aim to increase revenue from these activities. The supply side
must identify its market segments, create the respective products
with the right features and prices, and it has to find a distribution
channel. The traveller has to find the right product, its conditions,
its price and how and where to buy it. In fact, the vast quantity
of information available to the users makes this selection more
challenging.
      </p>
      <p>
        Finding the best alternative can become a complicated and
timeconsuming process. Consumers used to rely mostly on
recommendations from other people by word of mouth, known products from
∗Both authors contributed equally to this research.
advertisements [
        <xref ref-type="bibr" rid="ref20">20</xref>
        ] or inform themselves by reading reviews [
        <xref ref-type="bibr" rid="ref18 ref6">6, 18</xref>
        ].
However, the Internet has overtaken word of mouth as the primary
medium for choosing destinations [
        <xref ref-type="bibr" rid="ref23">23</xref>
        ] by guiding the user in a
personalized way to interesting or useful products from a large
space of possible options.
      </p>
      <p>Many players have emerged in the past decades mediating the
communication between the consumers and the suppliers. One type
of player is the Global Distribution System (GDS), which allows
customer-facing travel agencies (online or physical) to search and
book content from most airlines and hotels. Increased conversion
is a benefical goal for the supplier and broker as it implies more
revenue for a lower cost of operation, and for the traveller, as it
implies quicker decision making and thus less time spent on search
and shopping activities.</p>
      <p>In this study, we aim to increase the conversion rate for
hospitality recommendations after users book air travel. In Section 2,
the problem is formulated in order to highlight the
considerations which separate this work from many recommender system
paradigms. Section 3 presents the main techniques and concepts
used in this study. In Section 4, a brief overview is given of the
industry data used in this study. Section 5 discusses the results obtained
for diferent machine learning models including feature analysis.
A discussion of the main outcomes of this study is provided in
Section 6.
2
2.1</p>
    </sec>
    <sec id="sec-3">
      <title>PROBLEM FORMULATION</title>
    </sec>
    <sec id="sec-4">
      <title>Industry background</title>
      <p>
        Booking a major holiday is typically a yearly or bi-yearly activity for
travellers, requiring research for destinations, activities and pricing.
According to a study from Expedia [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ], on average, travellers
visit 38 sites up to 45 days prior to booking. The travel sector is
characterized by Burke and Ramezani [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ] as a domain with the
following factors:
• Low heterogeneity: the needs that the items can satisfy are
not so diverse.
• High risk: the price of items is comparatively high.
• Low churn: the relevance of items do not change rapidly.
• Explicit interaction style: the user needs to explicitly interact
with the system in order to add personal data. Although some
implicit preferences can be tracked from web activity and
past history, mainly the information obtained is gathered in
an explicit way (e.g. when/where do you want to travel?).
• Unstable preferences: information collected from the past
about the user might be no longer trustworthy today.
      </p>
      <p>
        Researchers have tried to relate touristic behavioural patterns
to psychological needs and expectations by 1) defining a
characterization of travel personalities and 2) building a computational
model based on a proper description of these profiles [
        <xref ref-type="bibr" rid="ref27">27</xref>
        ].
Recommender systems are a particular form of information filtering that
exploit past behaviours and user similarities. They have become
fundamental in e-commerce applications, providing suggestions
that adequately reduce large search spaces so that users are directed
toward items that best meet their preferences. There are several
core techniques that are applied to predict whether an item is in
fact useful to the user [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ]. With a content-based approach, items
are recommended based on attributes of the items chosen by the
user in the past [
        <xref ref-type="bibr" rid="ref26 ref3">3, 26</xref>
        ]. In collaborative filtering techniques,
recommendations to each user are based on information provided by
similar users, typically without any characterization of the
content [
        <xref ref-type="bibr" rid="ref19 ref24 ref25">19, 24, 25</xref>
        ]. More recentely, session-based recommenders have
been proposed, where content is selected based on previous activity
made by the user on a website or application [
        <xref ref-type="bibr" rid="ref17">17</xref>
        ].
2.2
      </p>
    </sec>
    <sec id="sec-5">
      <title>Terminology</title>
      <p>In order to clearly define our goal, let us first define some
terminology:
• Hotel Conversion: a hotel recommendation leads to a
conversion when the user books a specific hotel.
• Hotel Model: machine learning model trained to predict
the conversion probability of individual hotels.
• Passenger Name Record (PNR): digital record that
contains information about the passenger data and flight details.
• Session: after a traveller completes a flight booking through
a reservation system, a session is defined by the context of
the flight, the context of the reservation, and a set of five
recommended hotels proposed by the recommender system.
• Session Conversion: a session leads to a conversion when
the user books any of the hotels suggested during the session.
• Session Model: machine learning model trained using
features related with the session context and hotels, its output
is the conversion probability of the session.</p>
      <p>The end goal of the recommender system is to increase session
conversion. We can estimate the probability of booking of a list of
hotels using the session model, and thus we can compare diferent
lists using the session model to determine the one which will
maximise the probability of conversion of the session. Note that in this
case conversion is defined as a selection or "click" of a hotel on the
interface, rather than a booking.
2.3</p>
    </sec>
    <sec id="sec-6">
      <title>Hotel recommendations</title>
      <p>The content sold through a GDS is diverse, including flight
segments, hotel stays, cruises, car rental, and airport-hotel transfers.
The core GDS business concerns the delivery of appropriate travel
solutions to travel retailers. Therefore, state-of-the-art
recommendation engines capable of analysing historical bookings and
automatically recommending the appropriate travel solutions need
to be designed. Figure 1 shows an outline of the rule-based
recommendation system currently in use. After a user books a flight,
information related to the trip is sent to the recommender engine.</p>
      <p>However, this system does not take into account valuable
information such as the context of the request (e.g. where did the
booking originate from?), details about the associated flight (e.g.
how many days is the user staying in the city?) nor historical
recommendations (e.g. are similar users likely to book similar hotels?),
which are key assets to fine tune the recommendations.</p>
      <p>
        The problem is novel due to the richness of available data sources
(bookings, ratings, passenger information) and the variety of
distribution channels: indirect through travel agencies or direct
(website, mobile, mailbox). However, it is important to consider that
by design, no personally identifiable information (PII) or traveller
specific history is used as part of the model, which therefore
excludes collaborative-filtering or content-based approaches. The
contributions of this work are:
• The combination of data feeds to generate the context of
travel, including flights booked by traveller, historical hotels
proposed and booked at destination by other travellers, and
hotel content information.
• The definition of a 2-stage machine learning recommender
tailored for travel context. Two machine learning models are
required to build the new recommendation set. The output
of the first machine learning algorithm (prediction of the
probability of hotel booking) is a key input for the second
algorithm, based on the idea of [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ].
• The comparison of several machine learning algorithms for
modelling the hospitality conversion in the travel industry.
• The design and implementation of a recommendation builder
engine which generates the hotel recommendations that
maximize the conversion rate of the session. This engine is
built based on the analysis of the feature importance of the
session model at individual level [
        <xref ref-type="bibr" rid="ref29">29</xref>
        ].
3
3.1
      </p>
    </sec>
    <sec id="sec-7">
      <title>METHODOLOGY</title>
    </sec>
    <sec id="sec-8">
      <title>Pipeline</title>
      <p>Using machine learning and the historical dataset of
recommendations, we can train a model which is capable of predicting with
high confidence whether a proposed set of recommended hotels
leads to a booking.</p>
      <p>Once we have fit the model, we can evaluate other combinations
of hotels and recommend a list of hotels to the user that maximizes
the conversion. Instead of proposing a completely new set of hotels,
we decide to modify the existing suggestions given by the existing
rule-based system. Our approach, shown in Figure 2, removes one
of the initial hotels and introduces an additional one that increases
the conversion probability:</p>
      <p>We have identified two diferent ways to select the hotel that is
going to be introduced within the set of recommendations:
• We can create and evaluate all possible combinations and
choose the one with the highest conversion probability. This
means, each time one out of the five hotels from the initial
list is removed, and a new one from the pool of hotels is
inserted. However, this brute force solution is computationally
ineficient and time-consuming (e.g., in Paris this results in
5*1,653 diferent combinations for a single swap, the length
of the list multiplied by the number of available hotels).
• Alternatively, a hotel from the list of selected hotels can
be replaced with an available hotel, based on some criteria.
Typically, the criteria might be the price of the hotel room,
or the average review score, or a combination of multiple
indicators. In this work, the criteria used to optimise the
overall list of hotels is determined via feature analysis.</p>
      <p>Nevertheless, the last solution presents some challenges that
need to be discussed and solved:
(1) How to study the feature importance of complex non-linear
models?
(2) How to best interpret the feature importance in an
unbalanced dataset?
(3) How many features should be used during the selection
process of building an optimal list? Initially, we are facing a
multi-objective optimization problem since the choice of a
hotel for enhancing the conversion probability might depend
on diferent features. Furthermore, the existence of
categorical features makes this optimization even harder. Can we
convert it into a univariate optimization problem?</p>
      <p>
        The novelty of this study comes from the use of two related works
to address the above points. First, we design a two-stage cascaded
machine learning model [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ] where the output probabilities of the
ifrst model are a new feature of the second one. Second, we interpret
the feature importance of the positive instances (i.e. conversions)
with a local interpretable model-agnostic (LIME) technique [
        <xref ref-type="bibr" rid="ref29">29</xref>
        ].
Thus, we can study the feature importance of particular instances
in complex models, allowing the switch from a multi-objective to a
univariate optimization problem when one feature is dominant.
3.2
      </p>
    </sec>
    <sec id="sec-9">
      <title>Cascade Generalization</title>
      <p>
        Ensembling techniques consist in combining the decisions of
multiple classifiers in order to reduce the test error on unseen data. After
studying the bias-variance decomposition of the error in bagging
and boosting, Kohavi observed that the reduction of the error is
mainly due to reduction in the variance [
        <xref ref-type="bibr" rid="ref21">21</xref>
        ]. An issue with boosting
is robustness to noise since noisy examples tend to be misclassified
and therefore the weight will increase for these examples [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]. A
new direction in ensemble methods was proposed by Gama and
Brazdil [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ] called Cascade Generalization. The basic idea is to use
sequentially the set of classifiers (similarly to boosting), where at
each step, new attributes are added to the original data. The new
attributes are derived from the probability class distribution given
by the base classifiers.
      </p>
      <p>There are several advantages of using cascade generalization
over other ensemble algorithms:
• The new attributes are continuous since they are probability
class distributions.
• Each classifier has access to the original attributes and any
new attribute included at lower levels is considered exactly
in the same way as any of the original attributes.
• It does not use internal cross validation which afects the
computational eficiency of the method.
• The new probabilities can act as a dimensionality
reduction technique. The relationship between the independent
features and the target variable are captured by these new
attributes.</p>
      <p>As will be shown in further sections, this last point is a key
aspect of the proposed system, as the probabilities generated by the
hotel model can be used to directly select new hotels to include in
the recommendation. However, the session model uses aggregated
features from the hotel model, and as such an interpretable feature
analysis is required to determine how best to select hotels based
on their features.
3.3</p>
    </sec>
    <sec id="sec-10">
      <title>Interpretability in Machine Learning</title>
      <p>
        Machine learning has grown in popularity in the last decade by
producing more reliable, more accurate, and faster results in areas
such as speech recognition [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ], natural language understanding
[
        <xref ref-type="bibr" rid="ref8">8</xref>
        ], and image processing [
        <xref ref-type="bibr" rid="ref22">22</xref>
        ]. Nevertheless, machine learning
models act mostly as black boxes. That is, given an input the system
produces an output with little interpretable knowledge on how it
achieved that result. This necessity for interpretability comes from
an incompleteness in the problem formalisation meaning that, for
certain problems, it is not enough to get the solution, but also how it
came to that answer [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ]. Several studies on the interpretability for
machine learning models can be found on the literature [
        <xref ref-type="bibr" rid="ref1 ref15 ref32">1, 15, 32</xref>
        ].
3.4
      </p>
    </sec>
    <sec id="sec-11">
      <title>Local Interpretable Model-Agnostic</title>
    </sec>
    <sec id="sec-12">
      <title>Explanations (LIME)</title>
      <p>
        In this section, we focus on the work from Ribeiro et al. [
        <xref ref-type="bibr" rid="ref29">29</xref>
        ] called
Local Interpretable Model-Agnostic Explanations. The Local
Interpretable Model-Agnostic Explanations model explains the
predictions of any classifier (model-agnostic) in a interpretable and
faithful manner by learning an interpretable model locally around
the prediction:
• Interpretable. In the context of machine learning systems,
we define interpretability as the ability to explain or to
present in understandable terms to a human [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ].
• Local fidelity . Global interpretability implies describing
the patterns present in the overall model, while local
interpretability describes the reasons for a specific decision on a
unique sample. For interpreting a specific observation, we
assume it is suficient to understand how it behaves locally.
• Model-agnostic. The goal is to provide a set of techniques
that can be applied to any classifier or regressor in contrast
to other domain-specific techniques [
        <xref ref-type="bibr" rid="ref33">33</xref>
        ].
      </p>
      <p>In practice, LIME creates interpretable explanations for an
individual sample by fitting a linear model to a set of perturbed
variations of the sample and the resulting predictions as output
from a complex-model.
3.5</p>
    </sec>
    <sec id="sec-13">
      <title>Predictive Models</title>
      <p>The selection of which machine learning model to use highly
depends on the problem nature, constraints and limitations that are
trying to be solved. In this work, algorithms from diferent families
of machine learning were investigated. Specifically, the Naive Bayes
Classifier (NB) and Generalised linear Model (GLM) were
investigated as linear models, Random Forests (RF), Gradient Boosting
Machines (GBMs) were used to evaluate Decision Tree based
ensembles and fully connected Neural Networks (NN) were also assessed.
Furthermore, the model ensembling technique of Stacking (STK)
was also assessed. Stacking comprises of learning a linear model
to predict the target variable based on the output probabilities of
multiple machine learning algorithms as features.
3.6</p>
    </sec>
    <sec id="sec-14">
      <title>Hotel Model</title>
      <p>
        The first step is to train a machine learning model on individual
hotels, as shown is Figure 3. The features used for training this
model are not exclusively related to hotels, but also with the session
and flight context. Evaluating this model, we get the probability
that a certain hotel will be booked for a given location. The model
is learned by framing the problem as a supervised classification
problem, using the conversion (i.e. click) as a label. Note that for the
hotel model, the probabilities of conversion are independent of other
hotels presented in the session. This leads to several advantages:
• Cold start problem: the model does not penalise items or
users that have not been recommended yet, since no hotel
identifier or personally identifiable information is used. [
        <xref ref-type="bibr" rid="ref31">31</xref>
        ].
• Dimensionality reduction: the output probabilities of the
hotel model can be interpreted as a feature that comprises
the relationship between the independent variables and the
target variable. This is a key concept of the Cascade
Generalization technique, thus the output of the hotel model is
combined with the features to create the feature vector for
the session model, as shown in 4.
      </p>
      <p>Note that the features used as input to the hotel model are
discussed in Section 4.
The second machine learning model predicts whether a session
leads to a conversion or not, see Figure 4. A session is composed
of five diferent hotels and the aim of the recommender system
is to propose a set of hotels that results in the user booking any
one of them. Aggregates of the features from the Hotel Model
(contextual, passenger, and hotel features) are used, as well as the
hotel probabilities obtained from the hotel model. The numerical
features related with the hotels are aggregated in diferent ways
(max, min, std and avg of price and probability for example). The
features related with the context do not change (e.g. attributes about
the session or the flight) as these are identical for each element in
the session.
The Session Model estimates the conversion probability of the
session using contextual and content information. Thus, part of the
session builder is to create and evaluate new lists of hotels to
determine whether these lists will result in higher conversion probability
than the original list. Figure 5 shows how this process is performed.
First, a reference session with the recommendations, given by an
existing rule based system, is scored. For each of the proposed
hotels, we estimate the booking probability of each hotel using
the Hotel Model. Next, we can calculate the booking probability
at session level, using the probabilities of the Hotel Model as an
input feature of the Session Model. Then, we aim to improve the
conversion probability of the session by removing one of the hotels
from the list and introducing a new one. After including the new
hotel, if the booking probability of the current session is greater
than the probability of the previous session, then this new hotel
list is the one that will be proposed to the user.</p>
      <p>
        A rule must be defined to select the hotel to remove and which
new hotel to introduce in the recommendation list. Once we have
trained the Session Model, we can analyse the feature importance of
the variables for the positive cases that were correctly classified (i.e.
true positive cases). With the Local Interpretable Model-Agnostic
Explanations model [
        <xref ref-type="bibr" rid="ref29">29</xref>
        ], we can understand the behaviour of the
model for these particular instances. Based on the importance of
features from LIME, a heuristic can be defined to replace a hotel
from the list in order to improve the session conversion probability.
      </p>
      <p>Note that the LIME analysis is performed only on true positive
cases from the training set. In this dataset, the classes are highly
imbalanced due to a low conversion rate, as such standard feature
analysis techniques may be overly influenced by negative samples,
i.e., sessions which did not result in clicks. As LIME is designed to
be used on individual decisions, a linear model is fitted and analysed
for each true positive. The feature weights for each linear model are
then averaged, given a feature importance ranking for all correctly
classified converted sessions.
3.9</p>
    </sec>
    <sec id="sec-15">
      <title>Evaluation Metrics</title>
      <p>As with many conversion problems, the classes are highly
imbalanced, and as such the metrics used to assess performance must be
carefully chosen.</p>
      <p>
        F-measure (Fβ ). The generalization of the F1 metric is given by
[
        <xref ref-type="bibr" rid="ref7">7</xref>
        ]:
Fβ = (1 + β 2)PR
      </p>
      <p>β 2P + R
β is a parameter that controls a balance between precision P and
recall R. When β = 1, F1 comes to be equivalent to the harmonic
mean of P and R. If β &gt; 1, F becomes more recall-oriented (by
placing more emphasis on false negatives) and if β &lt; 1, it becomes more
precision oriented (by attenuating the influence of false negatives).
Common used metrics are the F2 and F0.5 scores.</p>
      <p>Area Under the ROC curve. The receiver operating
characteristic (ROC) curve is created by plotting the true positive rate (TPR)
against the false positive rate (FPR) at various threshold levels.
However, this can present an optimistic view of a classifier performance
if there is a large skew in the class distribution because the metric
takes into account true negatives.</p>
      <p>Average Precision (AP). The precision-recall curve is a similar
evaluation measure that is based on recall and precision at diferent
threshold levels. An equivalent metric is the Average Precision
(AP) which is the weighted mean of precisions achieved at each
threshold, with the increase in recall from the previous threshold
as the weight:</p>
      <p>AP = Õ</p>
      <p>(Rn − Rn−1)Pn
n</p>
      <p>
        Precision-recall curves are better for highlighting diferences
between models for unbalanced datasets due to the fact that they
evaluate the fraction of true positives among positive instances. In
highly imbalanced settings, the AP curve will likely exhibit larger
diferences and will be more informative than the area under the
ROC curve. Note that the relative ranking of the algorithms does
not change since a curve dominates in ROC space if and only if it
dominates in PR space [
        <xref ref-type="bibr" rid="ref10 ref30">10, 30</xref>
        ].
The dataset in this study consists of 715,952 elements. Out of these
recommendations, there are a total of 3,588 clicks, which are
considered conversions. Therefore, the dataset is unbalanced since only
0.5% of the instances are session conversions.
      </p>
      <p>Each row contains information regarding the context of the
session, the recommended hotel, and whether the recommendation
led to a conversion. In particular, the features are the number of
recommendations (from 1 to 5), date of the recommendation,
country where the booking was made, country where the passenger is
traveling, hotel identifier, hotel provider identifier, price of the hotel
at time of the recommendation, price currency and whether the
recommendation led to a conversion. Additionally, the logs were
enriched with supplementary information regarding each hotel
including a hotel numerical rating (from 0 to 5), hotel categorical
rating and the hotel chain.
information (e.g., flights number, dates) and the passenger
information (e.g., name, gender, and somethime passport details). A PNR
may also include many other data elements such as payment
information (currency, total price, etc), additional ancillary services sold
with the ticket (such as extra baggage and hotel reservation) and
other airline related information (cabin code, special meal request,
etc).</p>
      <p>For the purpose of this study, we retrieve and extract features
related with the air travel of the traveller. These include the date
of PNR creation, airline code, origin city, destination city, date of
departure, time of departure, date of arrival, time of arrival, days
between the departure and booking date, travel class, number of
stops (if any), duration of the flight in minutes (including stops)
and the number of days the passenger is staying at the destination.
5</p>
    </sec>
    <sec id="sec-16">
      <title>RESULTS</title>
      <p>Table 1 shows the results of the experiment comparing diferent
algorithms for the hotel model in terms of AUC, AP, F1 and F0.5
scores. In Figure 6, the ROC and AP curves can be seen in detail.
The low AUC value for the GLM model and Naive Bayes Classifier
suggest that linear classification techniques do not lead to the best
results and more complex models are needed to correctly represent
the data. The non-linear techniques have closer results, with the
Random Forest obtaining the best values for AP, F1 and F0.5. A
Stacked Ensemble using all the previous models is created but it
does not improve the previous outcome.
5.1</p>
    </sec>
    <sec id="sec-17">
      <title>Contribution of PNR data</title>
      <p>The PNR data is an important attribute since it contains rich
attributes related to the trip and passenger. However, is this case
personally identifiable information is not used in the recommender
system, thus the PNR features help to provide context about the
trip rather than the traveller. Incorporating this data to the models
substantially enhanced their performance, as can be observed in
Figure 6. Features of the PNR including the number of travellers
in the booking and trip duration, among others, contributed to an
increase in area under the PR curve from 0.183 to 0.249.
5.2</p>
    </sec>
    <sec id="sec-18">
      <title>Session Model</title>
      <p>After we have trained the hotel model, we predict individually the
probability of conversion of a hotel. Then, we create the sessions
based on 5 recommended hotels.</p>
      <p>In Table 2 the results are shown. In this case, the best model for
both AUC and AP is the Stacked Ensemble composed of a Random
Forest, a Generalized Linear Model and a Naïve Bayes Classifier.
Although the F0.5 score of the GBM model is slightly better than the
STK model, the latter clearly outperforms the rest of the metrics.
5.3</p>
    </sec>
    <sec id="sec-19">
      <title>Feature Importance</title>
      <p>After the Session model has been trained, we analyse its feature
importance to study which variables contribute the most to the model
using LIME. Concretely, we evaluate the model on the true positive
instances from the training dataset, since we want to optimise the
conversion.</p>
      <p>As can be seen in Figure 7, the most important features according
to LIME are all derived from the hotel model: the standard deviation,
maximum, and average individual hotel conversion probabilities.
Some features which are important to the model such as "market"
(country where the booking is made from), the flight class of service,
the destination city, and arrival and departure times of the flight
can not be used to manipulate the results of the session builder,
as these are all part of context of the recommendation. Features
extracted from prices (the diference between the average price and
the minimum, and the ratio of the lowest price to the average price)
are also considered important by the LIME model, but rank lower
than many hotel conversion probability features.</p>
      <p>As the standard deviation of the individual hotel conversions
is the most important criteria, the following rule for the session
builder is defined: from the original hotel list remove one hotel
with the closest conversion probability to the mean conversion
probability of the list, and replace it with the hotel with the
highest conversion probability from the set of available hotels for a
particular city.
5.4</p>
    </sec>
    <sec id="sec-20">
      <title>Simulated conversion using Hotel List</title>
    </sec>
    <sec id="sec-21">
      <title>Builder</title>
      <p>Results from the hotel list builder are shown in Table 3 for the two
largest cities in the dataset and for the complete dataset. For both
cities, we observe a large increase in conversion when using the
LIME based session builder. However, a brute force approach to
evaluating all possible lists does lead to higher conversion rates, at
the cost of a significant increase in processing time. When we
consider the complete dataset, we once again observe a large increase
in conversion from the baseline for the LIME model. With respect
to brute force, we observe that the LIME session builder performs
much closer to the brute force builder in terms of conversion. This
is attributed to the impact of smaller cities in the complete dataset,
and thus less choice in hotels for the builders, resulting in the LIME
session builder finding the optimal list. Additionally, on the
complete dataset, the processing time of the brute force builder is 2.8
times the duration of the LIME builder, whereas larger gains were
observed on the individual cities, where more options for hotels
were available.
6</p>
    </sec>
    <sec id="sec-22">
      <title>DISCUSSION</title>
      <p>In this study, an algorithm was created to improve hotel
recommendations based on historical hotel bookings and flight booking
attributes. Diferent machine learning models are used in a cascaded
fashion. First, a model estimates the conversion probability of the
individual hotels independently. Note that adding trip context, via
PNR based features, resulted in better PR AUC. The output of the
ifrst model is then combined with aggregates of the hotels in the
list in order to create a feature vector for the session model to
estimate the conversion probability that any hotel in the list will be
converted. LIME analysis revealed that the hotel model conversion
probabilities are the most important features, specifically the
standard deviation, mean and maximum individual hotel conversion
probabilities in the list. This allows for a simple heuristic to be
defined to increase the session conversion probability. In this study,
a single change is performed in the list of hotels, however this could
be expanded to allow multiple changes.</p>
      <p>
        Variations on this pipeline could also be considered, for instance
LIME is used in this study for feature importance ranking in the
session builder, however recently a similar methodology was proposed
using a mixture regression model referred to as LEMNA [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ].
      </p>
      <p>Here, the session builder relies on insights gained from analysis
of the feature importance ranking of the session model using LIME
over all sessions which lead to a conversion. Thus, the same
heuristic is applied to all datapoints in the session builder. However, a key
aspect of LIME is that it provides an interpretation of a model for a
single datapoint. As such, an evolution of the approach would be
to compute the most important features for each recommendation
in real time, and to use the information to build an optimal hotel
list based on the attributes most likely to lead to conversion.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>David</given-names>
            <surname>Baehrens</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Timon</given-names>
            <surname>Schroeter</surname>
          </string-name>
          , Stefan Harmeling, Motoaki Kawanabe, Katja Hansen, and
          <string-name>
            <surname>Klaus-Robert Müller</surname>
          </string-name>
          .
          <year>2010</year>
          .
          <article-title>How to explain individual classification decisions</article-title>
          .
          <source>Journal of Machine Learning Research</source>
          <volume>11</volume>
          ,
          <string-name>
            <surname>Jun</surname>
          </string-name>
          (
          <year>2010</year>
          ),
          <fpage>1803</fpage>
          -
          <lpage>1831</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>Eric</given-names>
            <surname>Bauer</surname>
          </string-name>
          and
          <string-name>
            <given-names>Ron</given-names>
            <surname>Kohavi</surname>
          </string-name>
          .
          <year>1998</year>
          .
          <article-title>An empirical comparison of voting classification algorithms: Bagging, boosting, and variants</article-title>
          .
          <source>Machine learning 36, 1</source>
          (
          <year>1998</year>
          ),
          <fpage>2</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>Yolanda</given-names>
            <surname>Blanco-Fernandez</surname>
          </string-name>
          , Jose J Pazos-Arias,
          <article-title>Alberto Gil-Solla, Manuel RamosCabrer</article-title>
          , and
          <string-name>
            <surname>Martin</surname>
          </string-name>
          Lopez-Nores.
          <year>2008</year>
          .
          <article-title>Providing entertainment by contentbased filtering and semantic reasoning in intelligent recommender systems</article-title>
          .
          <source>IEEE Transactions on Consumer Electronics</source>
          <volume>54</volume>
          ,
          <issue>2</issue>
          (
          <year>2008</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>J.</given-names>
            <surname>Bobadilla</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Ortega</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Hernando</surname>
          </string-name>
          ,
          <article-title>and</article-title>
          <string-name>
            <given-names>A.</given-names>
            <surname>Gutiérrez</surname>
          </string-name>
          .
          <year>2013</year>
          .
          <article-title>Recommender systems survey</article-title>
          .
          <source>Knowledge-Based Systems 46 (July</source>
          <year>2013</year>
          ),
          <fpage>109</fpage>
          -
          <lpage>132</lpage>
          . https://doi. org/10.1016/j.knosys.
          <year>2013</year>
          .
          <volume>03</volume>
          .012
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>Robin</given-names>
            <surname>Burke</surname>
          </string-name>
          and
          <string-name>
            <given-names>Maryam</given-names>
            <surname>Ramezani</surname>
          </string-name>
          .
          <year>2011</year>
          .
          <article-title>Matching recommendation technologies and domains</article-title>
          .
          <source>In Recommender systems handbook</source>
          . Springer,
          <fpage>367</fpage>
          -
          <lpage>386</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>Marcirio</given-names>
            <surname>Silveira</surname>
          </string-name>
          <string-name>
            <surname>Chaves</surname>
          </string-name>
          , Rodrigo Gomes, and
          <string-name>
            <given-names>Cristiane</given-names>
            <surname>Pedron</surname>
          </string-name>
          .
          <year>2012</year>
          .
          <article-title>Analysing reviews in the Web 2.0: Small and medium hotels in Portugal</article-title>
          .
          <source>Tourism Management</source>
          <volume>33</volume>
          ,
          <issue>5</issue>
          (
          <year>2012</year>
          ),
          <fpage>1286</fpage>
          -
          <lpage>1287</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>Nancy</given-names>
            <surname>Chinchor</surname>
          </string-name>
          .
          <year>1992</year>
          .
          <article-title>MUC-4 Evaluation Metrics</article-title>
          .
          <source>In Proceedings of the 4th Conference on Message Understanding (MUC4 '92)</source>
          .
          <article-title>Association for Computational Linguistics</article-title>
          , Stroudsburg, PA, USA,
          <fpage>22</fpage>
          -
          <lpage>29</lpage>
          . https://doi.org/10.3115/1072064.1072067
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>Ronan</given-names>
            <surname>Collobert</surname>
          </string-name>
          , Jason Weston, Léon Bottou, Michael Karlen, Koray Kavukcuoglu, and
          <string-name>
            <given-names>Pavel</given-names>
            <surname>Kuksa</surname>
          </string-name>
          .
          <year>2011</year>
          .
          <article-title>Natural language processing (almost) from scratch</article-title>
          .
          <source>Journal of Machine Learning Research</source>
          <volume>12</volume>
          ,
          <string-name>
            <surname>Aug</surname>
          </string-name>
          (
          <year>2011</year>
          ),
          <fpage>2493</fpage>
          -
          <lpage>2537</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>Antónia</given-names>
            <surname>Correia</surname>
          </string-name>
          , Patricia Oom do Valle, and
          <string-name>
            <given-names>Cláudia</given-names>
            <surname>Moço</surname>
          </string-name>
          .
          <year>2007</year>
          .
          <article-title>Why people travel to exotic places</article-title>
          .
          <source>International Journal of Culture, Tourism and Hospitality Research</source>
          <volume>1</volume>
          ,
          <issue>1</issue>
          (
          <year>2007</year>
          ),
          <fpage>45</fpage>
          -
          <lpage>61</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>Jesse</given-names>
            <surname>Davis</surname>
          </string-name>
          and
          <string-name>
            <given-names>Mark</given-names>
            <surname>Goadrich</surname>
          </string-name>
          .
          <year>2006</year>
          .
          <article-title>The relationship between Precision-Recall and ROC curves</article-title>
          .
          <source>In Proceedings of the 23rd international conference on Machine learning. ACM</source>
          ,
          <volume>233</volume>
          -
          <fpage>240</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>Finale</given-names>
            <surname>Doshi-Velez</surname>
          </string-name>
          and
          <string-name>
            <given-names>Been</given-names>
            <surname>Kim</surname>
          </string-name>
          .
          <year>2017</year>
          .
          <article-title>Towards a rigorous science of interpretable machine learning</article-title>
          . (
          <year>2017</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <surname>Expedia</surname>
          </string-name>
          .
          <year>2013</year>
          .
          <article-title>Retail and Travel Site Visitation Aligns As Consumers Plan and Book Vacation Packages</article-title>
          . https://advertising.expedia.com/about/press-releases/ retail-and
          <article-title>-travel-site-visitation-aligns-consumers-plan-and-book-vacation-packages</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <given-names>João</given-names>
            <surname>Gama</surname>
          </string-name>
          and
          <string-name>
            <given-names>Pavel</given-names>
            <surname>Brazdil</surname>
          </string-name>
          .
          <year>2000</year>
          .
          <article-title>Cascade generalization</article-title>
          .
          <source>Machine Learning</source>
          <volume>41</volume>
          ,
          <issue>3</issue>
          (
          <year>2000</year>
          ),
          <fpage>315</fpage>
          -
          <lpage>343</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <surname>Wenbo</surname>
            <given-names>Guo</given-names>
          </string-name>
          , Dongliang Mu, Jun Xu, Purui Su,
          <string-name>
            <given-names>Gang</given-names>
            <surname>Wang</surname>
          </string-name>
          , and
          <string-name>
            <given-names>Xinyu</given-names>
            <surname>Xing</surname>
          </string-name>
          .
          <year>2018</year>
          .
          <article-title>Lemna: Explaining deep learning based security applications</article-title>
          .
          <source>In Proceedings of the 2018 ACM SIGSAC Conference on Computer and Communications Security. ACM</source>
          ,
          <volume>364</volume>
          -
          <fpage>379</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <surname>Jonathan</surname>
            <given-names>L Herlocker</given-names>
          </string-name>
          ,
          <article-title>Joseph A Konstan,</article-title>
          and John Riedl.
          <year>2000</year>
          .
          <article-title>Explaining collaborative filtering recommendations</article-title>
          .
          <source>In Proceedings of the 2000 ACM conference on Computer supported cooperative work. ACM</source>
          ,
          <volume>241</volume>
          -
          <fpage>250</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [16]
          <string-name>
            <surname>Geofrey</surname>
            <given-names>Hinton</given-names>
          </string-name>
          ,
          <source>Li Deng</source>
          ,
          <string-name>
            <given-names>Dong</given-names>
            <surname>Yu</surname>
          </string-name>
          , George E Dahl, Abdel-rahman
          <string-name>
            <surname>Mohamed</surname>
          </string-name>
          , Navdeep Jaitly, Andrew Senior, Vincent Vanhoucke, Patrick Nguyen,
          <string-name>
            <surname>Tara N Sainath</surname>
          </string-name>
          , et al.
          <year>2012</year>
          .
          <article-title>Deep neural networks for acoustic modeling in speech recognition: The shared views of four research groups</article-title>
          .
          <source>IEEE Signal Processing Magazine</source>
          <volume>29</volume>
          ,
          <issue>6</issue>
          (
          <year>2012</year>
          ),
          <fpage>82</fpage>
          -
          <lpage>97</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [17]
          <string-name>
            <surname>Dietmar</surname>
            <given-names>Jannach</given-names>
          </string-name>
          , Malte Ludewig, and
          <string-name>
            <given-names>Lukas</given-names>
            <surname>Lerche</surname>
          </string-name>
          .
          <year>2017</year>
          .
          <article-title>Session-based item recommendation in e-commerce: on short-term intents, reminders, trends and discounts</article-title>
          .
          <source>User Modeling and User-Adapted Interaction 27</source>
          ,
          <fpage>3</fpage>
          -
          <lpage>5</lpage>
          (
          <year>2017</year>
          ),
          <fpage>351</fpage>
          -
          <lpage>392</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          [18]
          <string-name>
            <given-names>Ingrid</given-names>
            <surname>Jeacle</surname>
          </string-name>
          and
          <string-name>
            <given-names>Chris</given-names>
            <surname>Carter</surname>
          </string-name>
          .
          <year>2011</year>
          .
          <article-title>In TripAdvisor we trust: Rankings, calculative regimes and abstract systems</article-title>
          .
          <source>Accounting, Organizations and Society</source>
          <volume>36</volume>
          ,
          <issue>4</issue>
          (
          <year>2011</year>
          ),
          <fpage>293</fpage>
          -
          <lpage>309</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          [19]
          <string-name>
            <given-names>Michael</given-names>
            <surname>Kenteris</surname>
          </string-name>
          , Damianos Gavalas, and
          <string-name>
            <given-names>Aristides</given-names>
            <surname>Mpitziopoulos</surname>
          </string-name>
          .
          <year>2010</year>
          .
          <article-title>A mobile tourism recommender system</article-title>
          .
          <source>In Computers and Communications (ISCC)</source>
          ,
          <source>2010 IEEE Symposium on. IEEE</source>
          ,
          <fpage>840</fpage>
          -
          <lpage>845</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          [20]
          <string-name>
            <surname>Dae-Young</surname>
            <given-names>Kim</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Yeong-Hyeon Hwang</surname>
          </string-name>
          , and Daniel R Fesenmaier.
          <year>2005</year>
          .
          <article-title>Modeling tourism advertising efectiveness</article-title>
          .
          <source>Journal of Travel Research</source>
          <volume>44</volume>
          ,
          <issue>1</issue>
          (
          <year>2005</year>
          ),
          <fpage>42</fpage>
          -
          <lpage>49</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          [21]
          <string-name>
            <given-names>Ron</given-names>
            <surname>Kohavi</surname>
          </string-name>
          ,
          <string-name>
            <surname>David H Wolpert</surname>
          </string-name>
          , et al.
          <year>1996</year>
          .
          <article-title>Bias plus variance decomposition for zero-one loss functions</article-title>
          .
          <source>In ICML</source>
          , Vol.
          <volume>96</volume>
          .
          <fpage>275</fpage>
          -
          <lpage>83</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          [22]
          <string-name>
            <given-names>Yann</given-names>
            <surname>Le</surname>
          </string-name>
          <string-name>
            <surname>Cun</surname>
          </string-name>
          ,
          <string-name>
            <given-names>LD</given-names>
            <surname>Jackel</surname>
          </string-name>
          ,
          <string-name>
            <surname>B Boser</surname>
          </string-name>
          , JS Denker, HP Graf, Isabelle Guyon, Don Henderson, RE Howard, and
          <string-name>
            <given-names>W</given-names>
            <surname>Hubbard</surname>
          </string-name>
          .
          <year>1989</year>
          .
          <article-title>Handwritten digit recognition: Applications of neural network chips and automatic learning</article-title>
          .
          <source>IEEE Communications Magazine</source>
          <volume>27</volume>
          ,
          <issue>11</issue>
          (
          <year>1989</year>
          ),
          <fpage>41</fpage>
          -
          <lpage>46</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          [23]
          <string-name>
            <surname>Asher</surname>
            <given-names>Levi</given-names>
          </string-name>
          , Osnat Mokryn, Christophe Diot, and
          <string-name>
            <given-names>Nina</given-names>
            <surname>Taft</surname>
          </string-name>
          .
          <year>2012</year>
          .
          <article-title>Finding a needle in a haystack of reviews: cold start context-based hotel recommender system</article-title>
          .
          <source>In Proceedings of the sixth ACM conference on Recommender systems. ACM</source>
          ,
          <volume>115</volume>
          -
          <fpage>122</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          [24]
          <string-name>
            <surname>Greg</surname>
            <given-names>Linden</given-names>
          </string-name>
          ,
          <string-name>
            <given-names>Brent</given-names>
            <surname>Smith</surname>
          </string-name>
          , and Jeremy York.
          <year>2003</year>
          .
          <article-title>Amazon. com recommendations: Item-to-item collaborative filtering</article-title>
          .
          <source>IEEE Internet computing 7</source>
          ,
          <issue>1</issue>
          (
          <year>2003</year>
          ),
          <fpage>76</fpage>
          -
          <lpage>80</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref25">
        <mixed-citation>
          [25]
          <string-name>
            <surname>Stanley</surname>
            <given-names>Loh</given-names>
          </string-name>
          , Fabiana Lorenzi, Ramiro Saldaña, and
          <string-name>
            <given-names>Daniel</given-names>
            <surname>Licthnow</surname>
          </string-name>
          .
          <year>2003</year>
          .
          <article-title>A tourism recommender system based on collaboration and text analysis</article-title>
          .
          <source>Information Technology &amp; Tourism 6</source>
          ,
          <issue>3</issue>
          (
          <year>2003</year>
          ),
          <fpage>157</fpage>
          -
          <lpage>165</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref26">
        <mixed-citation>
          [26]
          <string-name>
            <surname>Raymond J Mooney and Loriene Roy</surname>
          </string-name>
          .
          <year>2000</year>
          .
          <article-title>Content-based book recommending using learning for text categorization</article-title>
          .
          <source>In Proceedings of the fifth ACM conference on Digital libraries. ACM</source>
          ,
          <volume>195</volume>
          -
          <fpage>204</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref27">
        <mixed-citation>
          [27]
          <string-name>
            <surname>Julia</surname>
            <given-names>Neidhardt</given-names>
          </string-name>
          , Leonhard Seyfang, Rainer Schuster, and
          <string-name>
            <given-names>Hannes</given-names>
            <surname>Werthner</surname>
          </string-name>
          .
          <year>2014</year>
          .
          <article-title>A picture-based approach to recommender systems</article-title>
          .
          <source>Information Technology &amp; Tourism</source>
          <volume>15</volume>
          ,
          <issue>1</issue>
          (sep
          <year>2014</year>
          ),
          <fpage>49</fpage>
          -
          <lpage>69</lpage>
          . https://doi.org/10.1007/s40558-014-0017-5
        </mixed-citation>
      </ref>
      <ref id="ref28">
        <mixed-citation>
          [28]
          <string-name>
            <given-names>Andreas</given-names>
            <surname>Papatheodorou</surname>
          </string-name>
          .
          <year>2001</year>
          .
          <article-title>Why people travel to diferent places</article-title>
          .
          <source>Annals of tourism research 28</source>
          ,
          <issue>1</issue>
          (
          <year>2001</year>
          ),
          <fpage>164</fpage>
          -
          <lpage>179</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref29">
        <mixed-citation>
          [29]
          <string-name>
            <given-names>Marco</given-names>
            <surname>Tulio</surname>
          </string-name>
          <string-name>
            <surname>Ribeiro</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Sameer</given-names>
            <surname>Singh</surname>
          </string-name>
          ,
          <string-name>
            <given-names>and Carlos</given-names>
            <surname>Guestrin</surname>
          </string-name>
          .
          <year>2016</year>
          .
          <article-title>Why should i trust you?: Explaining the predictions of any classifier</article-title>
          .
          <source>In Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining. ACM</source>
          ,
          <volume>1135</volume>
          -
          <fpage>1144</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref30">
        <mixed-citation>
          [30]
          <string-name>
            <given-names>Takaya</given-names>
            <surname>Saito</surname>
          </string-name>
          and
          <string-name>
            <given-names>Marc</given-names>
            <surname>Rehmsmeier</surname>
          </string-name>
          .
          <year>2015</year>
          .
          <article-title>The precision-recall plot is more informative than the ROC plot when evaluating binary classifiers on imbalanced datasets</article-title>
          .
          <source>PloS one 10</source>
          ,
          <issue>3</issue>
          (
          <year>2015</year>
          ),
          <year>e0118432</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref31">
        <mixed-citation>
          [31]
          <string-name>
            <surname>Andrew</surname>
            <given-names>I Schein</given-names>
          </string-name>
          , Alexandrin Popescul,
          <string-name>
            <surname>Lyle H Ungar</surname>
          </string-name>
          , and David M Pennock.
          <year>2002</year>
          .
          <article-title>Methods and metrics for cold-start recommendations</article-title>
          .
          <source>In Proceedings of the 25th annual international ACM SIGIR conference on Research and development in information retrieval. ACM</source>
          ,
          <volume>253</volume>
          -
          <fpage>260</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref32">
        <mixed-citation>
          [32]
          <string-name>
            <surname>Alfredo</surname>
            <given-names>Vellido</given-names>
          </string-name>
          ,
          <string-name>
            <surname>José David</surname>
          </string-name>
          Martín-Guerrero, and
          <source>Paulo JG Lisboa</source>
          .
          <year>2012</year>
          .
          <article-title>Making machine learning models interpretable</article-title>
          ..
          <source>In ESANN</source>
          , Vol.
          <volume>12</volume>
          .
          <string-name>
            <surname>Citeseer</surname>
          </string-name>
          ,
          <volume>163</volume>
          -
          <fpage>172</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref33">
        <mixed-citation>
          [33]
          <string-name>
            <surname>Peng</surname>
            <given-names>Zhang</given-names>
          </string-name>
          , Jiuling Wang, Ali Farhadi, Martial Hebert, and
          <string-name>
            <given-names>Devi</given-names>
            <surname>Parikh</surname>
          </string-name>
          .
          <year>2014</year>
          .
          <article-title>Predicting failures of vision systems</article-title>
          .
          <source>In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition</source>
          .
          <fpage>3566</fpage>
          -
          <lpage>3573</lpage>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>