<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Visualizable and explicable recommendations obtained from price estimation functions</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Claudia Becerra and</string-name>
          <email>[cjbecerrac,fagonzalezo]@unal.edu.co</email>
          <email>cjbecerrac@unal.edu.co</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Alexander Gelbukh</string-name>
          <email>gelbukh@gelbukh.com</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Center for Computing Research (CIC), National Polytechnic Institute (IPN)</institution>
          ,
          <addr-line>Mexico DF, 07738</addr-line>
          ,
          <country country="MX">Mexico</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Fabio Gonzalez, Intelligent Systems Research Laboratory (LISI), Universidad Nacional de Colombia</institution>
          ,
          <addr-line>Bogota</addr-line>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2011</year>
      </pub-date>
      <fpage>23</fpage>
      <lpage>27</lpage>
      <abstract>
        <p>Collaborative filtering is one of the most common approaches in many current recommender systems. However, historical data and customer profiles, necessary for this approach, are not always available. Similarly, new products are constantly launched to the market lacking historical information. We propose a new method to deal with these “cold start” scenarios, designing price-estimation functions used for making recommendations based on cost-benefit analysis. Experimental results, using a data set of 836 laptop descriptions, showed that such price-estimation functions can be learned from data. Besides, they can also be used to formulate interpretable recommendations that explain to users how product features determine its price. Finally a 2D visualization of the proposed recommender system was provided.</p>
      </abstract>
      <kwd-group>
        <kwd>Apriori recommendation</kwd>
        <kwd>Cold-start recommendation</kwd>
        <kwd>Price estimation functions</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>H.1.2 [User/Machine Systems]: Human information
processing; H.4.2 [Types of Systems]: Decision support</p>
    </sec>
    <sec id="sec-2">
      <title>General Terms</title>
      <p>Experimentation</p>
    </sec>
    <sec id="sec-3">
      <title>1. INTRODUCTION</title>
      <p>The internet and e-commerce grow exponentially. As a
result, decision-making process about products and services
is becoming increasingly complex. These processes involve
hundreds and even thousands of choices and a growing
number of heterogeneous features for each product. This is
mainly due to the introduction and constant evolution of
new markets, technologies and products.</p>
      <p>
        Unfortunately, human capacity for decision-making is too
limited to address the complexity of this scenario. Studies in
psychology field have shown that human cognitive capacities
are limited from five to nine alternatives for simultaneous
comparison [
        <xref ref-type="bibr" rid="ref14 ref17">17, 14</xref>
        ]. Consequently, making a purchasing
decision at an e-commerce store that does not provide tools to
assist decision-making, is a task that largely overwhelms
human capacities. Moreover, several studies have shown that
this problem generates adverse e↵ ects on people such as:
regret due to the selected option, dissatisfaction due to poor
justification for the decision, uncertainty about the idea of
“best option”, and overload of time, attention and memory
(see [
        <xref ref-type="bibr" rid="ref20">20</xref>
        ]).
      </p>
      <p>
        Many recommender systems approaches have addressed
this problem through collaborative filtering [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ] based on
product content (i.e. descriptions) and on customer
information [
        <xref ref-type="bibr" rid="ref2 ref9">2, 9</xref>
        ]. This approach recommends products similar to
those chosen by similar users. On the other hand, latent
semantics approaches [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ] have been successfully used to build
a nity measures between products and users. Most of the
aforementioned approaches have been applied in domains
with products such as books and movies that remain
available long enough to collect enough historical data to build
a model [
        <xref ref-type="bibr" rid="ref12 ref4">4, 12</xref>
        ].
      </p>
      <p>
        While impressive progresses have been made in the field
using collaborative filtering, the relevance of current
approaches in domains with frequent changes in products is
still an open question [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ]. For example, customer-electronics
domain is characterized by products with a very short life
cycle in the market and a constant renewal of technologies
and paradigms. Collaborative approaches face two major
problems in this scenario [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ]. First, product features are
constantly redefined, making di cult for users to identify
relevant attributes. Second, historical sales product data
become obsolete very quickly due to the frequent product
substitution. This problem of making automatic
recommendations without historical data is known as cold-start
recommendation [
        <xref ref-type="bibr" rid="ref18">18</xref>
        ].
      </p>
      <p>In this paper, we propose a new cold-start method based
on an estimate of the benefit to the user when purchasing
a product. This function is formulated as the di↵ erence
between estimated and real prices. Therefore, our approach
recommends products with high benefit-cost ratio to find
“best-deals” on a data set of products. Figure 1 shows an
example of such recommendations based on utility functions
displaying 900 laptop computers. In this figure, the features
of laptops below the line in bold, indicating fair prices, do
not justify prices of laptops.</p>
      <p>The rest of the paper is organized as follows. In Section
2, the necesary background and proposed method are
presented. In Section 3, an evaluation methodology and some
data refinements are proposed and applied to the model.
Finally, in Section 4, some concluding remarks are briefly
discussed.</p>
      <p>APRIORI RECOMMENDATIONS USING
UTILITY FUNCTIONS</p>
      <p>
        The general intuition of method is led by the
lexicographical criterion [
        <xref ref-type="bibr" rid="ref22">22</xref>
        ]. That is, users prefer products that o↵ er
more value for their money. Clearly, this approach is not
applicable to all circumstances, but it is general enough when
customer profiles are not available in cold-start scenarios.
      </p>
      <p>When a user purchases a product xi, a utility function
utility(xi) provides an estimation of the di↵ erence between
the estimated price f (xi) and the market price yi, that is
utility(xi) = f (xi) yi. Thus, the products in the market
are represented as a set X = {x1, x2,..., xn, ..., xN }, where
each product xi is a vector characterized in a feature space
RM. With these data, a regression model, learned from X
and the vector of prices y, generates price estimations f (xi)
required for calculation of the utility. Finally, the utility
function is computed on all products thus providing an
ordered list with the top-n apriori recommendations.</p>
      <p>Estimates of price f (xi) can be obtained by a
linearregression model as:</p>
      <p>X
m2 {1,...,M}
f (xi) = o +
mxim.</p>
      <p>(1)</p>
      <p>
        This model is equivalent to an additive value function
used in the decision-making model SAW (simple additive
weighting) [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ], but with coe cients m learned
automatically. Clearly, the recommendations obtained from these
estimates can be explained to users, since each term mxim
represents the money contribution to the final price estimate
provided by the m-th feature of the i -th product.
      </p>
      <p>
        The quality of the apriori recommendations obtained with
the proposed method depends primarily on three factors:
the amount of training data, the accuracy of price estimates
f (xi), and the ability to extract user-understandable
explanations from the regression model. Certainly, linear models,
such as that of eq. 1, o↵ er good interpretability, but in many
cases, these models generate high rates of error in their
predictions when the interactions among features are complex.
These models are known as weak regression models. On the
other hand, discriminative approaches, such as support
vector regression [
        <xref ref-type="bibr" rid="ref19">19</xref>
        ], provide better models with lower error
rates but also with lower interpretability.
      </p>
      <p>
        This trade-o↵ can be overcome with a hybrid regression
model as 3FML (three-function meta-learner) [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]. This
metaregressor combines two di↵ erent regression methods in a new
improved combined model in a way similar to other
metaalgorithms such as voting, bagging and AdaBoost (see [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]).
Unlike these methods, 3FML uses one regression method to
make price predictions and another to predict the error. As
long as the former regression method is weak, stable and
interpretable, the latter can be any other regression method
regardless its interpretability. As a result, the combined
regression preserves the same interpretability level of the first
regressor but with lower error rate.
      </p>
      <p>
        A linear regression model can be trained to learn
parameters m by minimizing the least squared error from data
[
        <xref ref-type="bibr" rid="ref15">15</xref>
        ]. This first model can be used by 3FML to build a base
regression model f0(x) with the full dataset. Then, this
model is used to divide the data into two additional groups
depending on whether the obtained price predictions were
below or above the training price, given a di↵ erence
threshold ✓ . Next, using the same base-regression method, two
additional models f+1(x) and f 1(x) are trained with the pair
of subsets called respectively, upper model and lower model.
Figure 3 illustrates upper, base and lower models compared
to the target function, which is the price in a data set of
laptop computers. The three resulting models are combined
using an aggregation mechanism – called mixture of experts
[
        <xref ref-type="bibr" rid="ref11">11</xref>
        ] – with the following expression:
having
      </p>
      <p>P
l✏H
fˆ(xi) =</p>
      <p>wl(xi)fl(xi),
P
l✏H
wl(xi) = 1, i 2 {1 . . . n}
(2)
(3)</p>
      <p>H is a dictionary of experts consisting on the base model
and two additional specialized models, H = {f 1(x), f0(x),
f+1(x)}. The gating coe cient wli establishes the level of
relevance of the l model into the final price prediction for
the i -th product.</p>
      <p>
        In 3FML model, coe cients wli = wl(xi) are obtained
by chaining a membership function wl for each regression
model to a function ↵ that depends on the errors of the
three models, wl(xi) = wl(↵ (f 1(xi), f0(xi), f+1(xi), yi)).
These membership functions wl(↵ ) are similar to those used
in fuzzy sets [
        <xref ref-type="bibr" rid="ref23">23</xref>
        ] but these satisfy the constraint given by
eq. 3. Three examples of those functions are shown in
Figure 2; one triangular and two Gaussian. Clearly, the range
of the error function ↵ must agree with the domain of the
membership functions. For instance, if the domain of the
membership functions is [
        <xref ref-type="bibr" rid="ref1">0, 1</xref>
        ], an appropriate function ↵ i
must return a value close to 0.5 when yi is better modeled
by f0(xi). Similarly, reasonable values for ↵ i, if yi is better
modeled by f 1(xi) or f+1(xi), are respectively 0.0 and 1.0.
      </p>
      <p>
        Such function ↵ can be arithmetically constructed (see
[
        <xref ref-type="bibr" rid="ref3">3</xref>
        ] for triangular and Gaussian cases) and ↵ i can be
obtained for every xi. 3FML makes use of a second
regression method to learn a function for ↵ i. This function is
called ↵ -learner, which seeks to predict the same target yi
but indirectly through the errors obtained by f 1, f0 and
f+1. The estimates obtained with ↵ -learner are used in
combination with the membership functions to get coe cients
wl(xi). Therefore, final predictions are obtained with a
different linear model for each target price yi. The resulting
model is also linear, but di↵ erent for each product instance
in function to xi:
      </p>
      <p>P
m✏ {1, ..., M }
fˆ(xi) = ˆ0(xi) +
ˆm(xi)xim,
(4)
where
ˆo(xi) = X lowl(xi) ; ˆm(xi) = X lmwl(xi).</p>
      <p>l2 H
l2 H</p>
      <p>Clearly, the model in eq. 4 is as user-explainable as that
of eq. 1.</p>
      <p>The e↵ ect of ↵ -learner in eq. 4 is that the entire data set
is clustered into three latent classes. These classes can be
considered as market segments namely: high-end, mid-range
and low-end products. Many commercial markets exhibit
this segmentation, e.g. computers, mobile phones, cars, etc.</p>
      <p>EXPERIMENTAL VALIDATION</p>
      <p>The aim of experiments is to build a model that provides a
cost-benefit ranking of a set of products where each product
is represented as a vector of features. To assess the quality
of this ranking, two factors are observed. First, the error
of the price-estimation regression should be low to make
sure that this function provides a reasonable explanation
of the data set. Second, the model must be interpretable
and discovered knowledge must be consistent with market
data. For example, if a proposed model discovers a ranking
of how much money each operating system contributes to
laptop prices, this ranking should be in agreement the prices
of retail versions of the same operating systems.</p>
      <p>In addition, the full features set of the top-10
recommended products is provided along with a 2D visualization
of the entire data set. These resources allow the reader –
guided by a brief discussion – to qualitatively evaluate the
recommendations obtained with the proposed method.</p>
      <p>
        The data is a set of 836 laptop computers each represented
by a vector of 69 attributes including price, which is the
attribute to be estimated. Data were collected by Becerra1
from several U.S. e-commerce sites (e.g. Pricegrabber, Cnet,
Yahoo, etc.), during the second half of 2007 within a month.
A subset of 17 features was selected using the
correlationbased selection method proposed by Hall [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ]. We call this
dataset Laptops 17 836 ; all its features and percentage of
missing values are shown in Table 1.
      </p>
      <p>
        For the construction of the price-estimation function,
several regression methods were used, namely: least mean squares
linear regression (LMS) [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ], M5P regression tree [
        <xref ref-type="bibr" rid="ref16 ref21">21, 16</xref>
        ],
support vector regression (SVR) [
        <xref ref-type="bibr" rid="ref19">19</xref>
        ] and three-function
metalearner (3FML, described in previous section). 3FML
provides three interpretable linear models: upper, base and
lower models, which can be associated with product classes.
Finally, estimated price for each laptop was obtained with
the combination of these three models using eq. 4 with the
weights obtained from ↵ -Learner and Gaussian membership
functions.
      </p>
      <p>The performance of each method was measured using
rootmean-square error (RMSE) defined as:
v
uu Pi ⇣fˆ(xi)
RM SE = t
yi
⌘2</p>
      <p>.
|X|</p>
      <p>The data set was randomly divided into 75% for training
and 25% for testing. Ten di↵ erent runs of this partition ratio
were used for each method. These ten RMSE results were
averaged and reported. Table 2 shows the results, their
standard deviation (in parentheses) and some model parameters.
1http://unal.academia.edu/claudiabecerra/teaching
The method with lowest RMSE was SVR with a
complexity parameter C = 100 using radial basis functions (RBF)
as kernel. However, interpretability of this model is quite
limited, given the embedded feature space induced by the
kernel. On the other hand, LMS and 3FML provide
straightforward interpretation of coe cients, which represent the
amount of the contribution of each feature to the product
estimated price. Clearly, 3FML was the method that better
coped with this interpretability-accuracy trade-o↵ .</p>
      <p>In this section the price estimation function obtained
using 3FML is manually analyzed checking coherence of
coe cients with real facts of the market. Particularly,
coefficients for attributes operating system, processor and
numerical features are reviewed, and – when necessary – some
refinements are proposed to the data sets to deal with
discussed issues.
3.3.1</p>
      <sec id="sec-3-1">
        <title>Operating System attribute analysis</title>
        <p>Table 3 shows the distribution of the di↵ erent operating
systems into the entire data set of laptops and the
abbreviations that we use to refer them at Table 5 and Table 4.</p>
        <p>In order to evaluate the portion of the price estimation
model related to operating system (OS) attribute, coe
cients of this feature are compared with related Microsoft’s
retail prices. Table 4 shows public retail prices for Windows
Vista published at 2007-3Q. In spite that at that date,
Windows Vista operating system had already six months
of launched, many brand new laptops still had pre-installed
previous Windows XP . Thus, we consider for analysis
Windows XP Pro equivalent to Windows Vista Business , as
well as, Windows XP equivalent to Windows Vista Home
Premium . This assumption is also coherent with the
observed behavior in Microsoft’s price policy that keeps prices
of previous product releases invariable during version
transition periods.</p>
        <p>It is interesting to highlight the behavior of 3FML model
with Windows Vista Ultimate . Although this OS version
occurs only at 1.32% of instances (see Table 3), it is
correctly recognized as the most expensive OS (see Table 4) by
the upper model. This fact corrects an erroneous tendency
recognized by base and lower models. In general terms, for
other OS versions, 3FML managed to predict similar
ordering as that of retail prices.
3.3.2</p>
      </sec>
      <sec id="sec-3-2">
        <title>Processor attribute coefficients</title>
        <p>As shown in Table 1, Laptops 17 836 data set has two
features to describe the main processors of laptops , they are:
Processor Speed (numeric) and Processor (nominal). The
former is the processor clock rate and the latter is a text
string that contains — in most of cases – the manufacturer,
the product family and the model (e.g. “Intel Core 2 Duo
Mobile T7200”). Unlike OS attribute, which has only seven
possible alternatives, Processor attribute has 133 possible
processor models. Moreover, the frequencies of occurrence
of each processor model exhibit a Zipf-type distribution (see
Figure 4). Thus, approximately half of the 836 laptops have
only 8 di↵ erent processors and more than 80 processors
occur only in one laptop. Part of this sparseness is due to
missing information, abbreviations and formatting.</p>
        <p>The Processor attribute, as found in the data set, can
generate a detrimental e↵ ect on the price-estimation
function. Besides, coe cients could hardly be explained and
their evaluation against market facts could lead to
misleading results. Thus, the model was withdrawn from Processor
attribute and it was renamed as Proc. Family. In addition,
the data set was enriched manually adding the following four
processor related attributes:
• L2-Cache: processor cache in Kibibytes (210 bytes).
• Hyper Transport: frontal bus clock rate in Mhz.
• Thermal Design: maximum dissipated power in watt.
• Process Technology: CMOS technology in nanometre.
This new data set is referred as Laptops 21 836 data set.
Performance results of new price-estimation functions are
shown in Table 6. Clearly, SVR and 3FML obtained
substantial improvements using this new data set.</p>
        <p>Similarly to the analysis made for OS attribute, processors
families also have a consumer-value ranking given by their
technology, which can be compared to a ranking taken from
an interpretable price-estimation function. The technology
ranking of Intel processors is: (1st) Core 2 Duo , Core
Duo , Core Solo , Pentium Dual Core and Celeron .
Same for AMD’s processors: (1st) Turion , Athlon and
Sempron 2. We extracted a ordering for processor
fami2see
http://www.notebookcheck.net/NotebookProcessors.129.0.html for a short description of mobile
processor families (site consulted in June 2011)
lies by their corresponding coe cients from 3FML models.
Results for this ranking – means and standard deviation –
making 10 runs with di↵ erent samples of 75% training and
25% test are shown in Table 7.</p>
        <p>Results in Table 7 show how upper model better ordered
processor families with high technological ranking.
Similarly, lower model does a similar work recognizing Sempron
family at the lowest rank.
3.3.3</p>
        <sec id="sec-3-2-1">
          <title>Numerical attributes coefficients</title>
          <p>This subsection present a brief discussion on the
interpretation of coe cients extracted from the price-estimation
function for some numeric attributes (shown in Table 8).
Although this interpretation is clearly subjective, it reveals
some laptop-market facts, which were extracted in an
unsupervised way from the data.</p>
          <p>For instance, consider Thermal Design attribute.
Negative values in the coe cients reveal a fact: the lesser
power the CPU dissipates, the higher the laptop’s price.
Besides, these coe cients also shows that this e↵ ect a↵ ects
prices more at mid-range and low-end laptop-market
segments. Similarly, Max. Horizontal Resolution attribute
reveals that this feature has greater impact on the mid-range
laptop market prices.</p>
          <p>Interestingly, there is a phenomenon revealed by the
features that are easy perceived by users, such as Installed
Memory, Max. Horizontal Res. (number of horizontal pixels
on screen), L2-Cache and Processor Speed. That is: those
features have considerably less e↵ ect on prices in high-end
than in mid-range and low-end market segments. This
phenomenon can be explained by the fact that “luxury” goods
justify their price more by attributes such as brand-label,
exclusive features and physical appearance rather than for
their configuration.
3.4</p>
          <p>Recommendations for users
3.4.1</p>
        </sec>
        <sec id="sec-3-2-2">
          <title>Top-10 recommendations</title>
          <p>After the quantitative evaluation (i.e. regression error)
and qualitative assessment (i.e. agreement with market facts)
using 3FML model, the resulting functions provided
reasonable estimates of price and support elements to explainable
recommendations such as rankings and weights of attributes.
After obtaining the estimates of prices, the profit for each
laptop is calculated from the di↵ erence between this
estimate and real price. Table 9 shows the top-10
recommendations with the highest profit among all 836 laptops.</p>
          <p>The first and second top-ranked laptops have similar
configurations, but even small di↵ erences make comparison
difficult at first sight. The second laptop has better price, more
memory, docking station and ports replicator slots. Unlike,
the former has higher screen resolution and a fingerprint
sensor. These di↵ erences can be compared quantitatively with
the help of coe cients provided by the model. However, a
better explanation of the #1 recommended choice is a
market fact extracted from the obtained manufacturer ranking
showed in Table 10. The three regression models identify
the Lenovo brand better ranked than HP . Therefore,
the first recommended laptop becomes a “best deal” given
the standard prices of Lenovo at the time. Similarly,
recommendations #7, #9 and #10 seem to get their high user
profit not because of their configuration features, but
because of their label Sony , which do better positions on the
ranking of manufacturers than its counterparts.</p>
          <p>Second and third recommendations only di↵ er in
Processor Speed attribute. Clearly, the estimated cost of that
differentia is the numerical di↵ erence between their estimated
prices, which is $42. Nevertheless, their real price di↵ erence
is $50. This explains the order of position in the ranking
assigned by the recommender system to the #2 and #3
recommendations. More pair-wise comparisons and evaluations
could be made but are omitted due to space limitations.</p>
          <p>These paired comparisons become cognitively more di
cult when the number of features, di↵ erences and instances
increases. However, the proposed recommender method
provides reasonable explanations no matter how much data is
involved, and these can be provided by user request. This
is important because cold-start recommender systems need
to establish trust in users due of the lack of collaborative
support.
3.4.2</p>
          <p>2D Visualization</p>
          <p>Ordered lists are the most common format to present
recommendations to users. However, despite having such an
ordination, establish the most appropriated choice for a
particular user is a di cult task. Therefore, we propose a novel
visualization method for our recommender system. The
purpose of this is to enable users to build a mental model of the
market. When users do not have a clear aim or a defined
budget, this tool provides a rough idea of the number of
options and prices. In addition, visualization can help the
short-term memory decreasing cognitive load and highlight
the recommended options.</p>
          <p>The proposed 2D visualization is shown in Figure 5. The
horizontal axe represents actual price and the vertical axe
represents the profit, which is the di↵ erence between the
estimated and actual price. Each laptop is represented as a
bubble, where larger radius and warmer colors (darker gray
in the grayscale version) means higher profit-price di↵
erences. Besides, the number of ranking was included in the
top-99 recommendations.</p>
          <p>This visualization highlights other “best deals” that are
hidden in the ranking list. For instance, consider
recommendation #53 (see Figure 5 in the coordinates $1550 price
and $260 profit). Perhaps this is an interesting option to
consider if user’s budget is over $1500. Similarly,
recommendation #26 can be quickly identified as the best option
for buyer on a low budget.</p>
          <p>The proposed visualization also allows a qualitative
assessment of the price-estimation function. For instance,
consider the laptops above $1300, this function has di culties
to predict prices using the current set of features, which in
turn appears to be very e↵ ective for mid-range prices. This
problem could be solved indentifying and adding to the set
of attributes those distinctive features of high-end laptops,
namely: shockproof devices, special designs, colors, housing
materials, exclusive options, etc.</p>
          <p>CONCLUSIONS</p>
          <p>We presented a novel product recommender system based
on an interpretable price-estimation function, which
estimates the economic benefit for the customer to buy a
product in a particular market. Accurate and interpretable price
estimations were obtained using the 3FML (three-function
meta-learner) method. This regression method allows the
combination of an interpretable regressor (e.g. LMS) to
estimate prices and an uninterpretable regressor (e.g. SVR)
to identify the latent class of each product. The combined
model obtained better price estimates than LMS, SVR and
M5P regression tree, while it kept a high level of
interpretation.</p>
          <p>The proposed method was tested with real-market data
from a data set of laptops. The obtained price-estimation
model was interpretable, allowing evaluation and refinement
by domain experts and ensuring that price estimates are
a coherent consequence of the product features. In
addition, the obtained recommendations are easy to understand
by users. For instance, feature rankings (e.g. ranking of
CPU) and feature price contributions (e.g. cost per GB of
main memory) are provided. Importantly, while the price
estimates are obtained in a supervised way, other domain
knowledge is extracted in a non-supervised way. Although
the proposed method was tested in a particular domain (i.e.
laptops), this same process can be applied to other domains
that exhibit similar number of options and features.</p>
          <p>Moreover, a user-friendly visualization method for
recommendations was proposed using a 2D Cartesian metaphor
and concrete variables such as cost and profit. This
visualization allows users to make a quick mental map of a large
market to explore and identify recommendations in di↵ erent
price ranges.</p>
          <p>In conclusion, the proposed method is flexible and can be
useful in e-commerce scenarios with products that allow the
construction of price-estimation functions, such as
customerelectronics products and others. Finally, our method fills a
gap where recommender systems based on historical
information fail because of the lack of such information.</p>
          <p>ACKNOWLEDGEMENTS</p>
          <p>This research is funded in part by the Bogota Research
Division (DIB) at the National University of Colombia, and
throught a grant from Colciencias, project 110152128465.</p>
        </sec>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>E.</given-names>
            <surname>Alpaydin</surname>
          </string-name>
          .
          <article-title>Introduction to Machine Learning</article-title>
          . The MIT Press,
          <year>October 2004</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>M.</given-names>
            <surname>Balabanovic</surname>
          </string-name>
          and
          <string-name>
            <given-names>Y.</given-names>
            <surname>Shoham</surname>
          </string-name>
          . Fab:
          <article-title>Content-based collaborative recommendation</article-title>
          .
          <source>Communications of the Association for Computing Machinery</source>
          ,
          <volume>40</volume>
          (
          <issue>3</issue>
          ):
          <fpage>66</fpage>
          -
          <lpage>72</lpage>
          ,
          <year>1997</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>C.</given-names>
            <surname>Becerra</surname>
          </string-name>
          and
          <string-name>
            <given-names>F.</given-names>
            <surname>Gonzalez</surname>
          </string-name>
          . 3
          <article-title>-functions meta-learner algorithm: a mixture of experts technique to improve regression models</article-title>
          .
          <source>In DMIN08: Proceedings of the 4th international conference on data mining</source>
          , Las Vegas,
          <string-name>
            <surname>NV</surname>
          </string-name>
          , USA.,
          <year>2008</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>J.</given-names>
            <surname>Bennet</surname>
          </string-name>
          and
          <string-name>
            <given-names>S.</given-names>
            <surname>Lanning</surname>
          </string-name>
          .
          <article-title>The netflix prize</article-title>
          .
          <source>In KDD Cup and Workshop</source>
          ,
          <year>2007</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>R. T.</given-names>
            <surname>Eckenrode</surname>
          </string-name>
          .
          <article-title>Weighting multiple criteria</article-title>
          .
          <source>Management Sciences</source>
          ,
          <volume>12</volume>
          :
          <fpage>180</fpage>
          -
          <lpage>192</lpage>
          ,
          <year>1965</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>D.</given-names>
            <surname>Goldberg</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Nichols</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B. M.</given-names>
            <surname>Oki</surname>
          </string-name>
          , and
          <string-name>
            <given-names>D.</given-names>
            <surname>Terry</surname>
          </string-name>
          .
          <article-title>Using collaborative filtering to weave an information tapestry</article-title>
          .
          <source>Commun. ACM</source>
          ,
          <volume>35</volume>
          (
          <issue>12</issue>
          ):
          <fpage>61</fpage>
          -
          <lpage>70</lpage>
          ,
          <year>1992</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>M. A.</given-names>
            <surname>Hall</surname>
          </string-name>
          .
          <article-title>Correlation-based feature selection for discrete and numeric class machine learning</article-title>
          .
          <source>In ICML '00: Proceedings of the 17th International Conference on Machine Learning</source>
          , pages
          <fpage>359</fpage>
          -
          <lpage>366</lpage>
          , San Francisco, CA, USA,
          <year>2000</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>E.</given-names>
            <surname>Han</surname>
          </string-name>
          and
          <string-name>
            <surname>G</surname>
          </string-name>
          . Karypis.
          <article-title>Feature-based recommendation system</article-title>
          .
          <source>In Proceedings of the 14th ACM International Conference on Information and Knowledge Management (CIKM)</source>
          , pages
          <fpage>446</fpage>
          -
          <lpage>452</lpage>
          ,
          <year>2005</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <surname>Borchers</surname>
            <given-names>A</given-names>
          </string-name>
          .
          <string-name>
            <surname>Riedl J. Herlocker</surname>
            <given-names>J. L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Konstan</surname>
            <given-names>J. A.</given-names>
          </string-name>
          <article-title>An algorithmic framework for performing collaborative filtering</article-title>
          .
          <source>In Proc. 22nd ACM SIGIR Conference on Information Retrieval</source>
          , pages
          <fpage>230</fpage>
          -
          <lpage>237</lpage>
          ,
          <year>1999</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>T.</given-names>
            <surname>Hofmann</surname>
          </string-name>
          .
          <article-title>Latent semantic models for collaborative filtering</article-title>
          .
          <source>ACM Trans. on Information Systems</source>
          , pages
          <fpage>89</fpage>
          -
          <lpage>115</lpage>
          ,
          <year>2004</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>R. A.</given-names>
            <surname>Jacobs</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. I.</given-names>
            <surname>Jordan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S. J.</given-names>
            <surname>Nowlan</surname>
          </string-name>
          , and
          <string-name>
            <given-names>G. E.</given-names>
            <surname>Hinton</surname>
          </string-name>
          .
          <article-title>Adaptive mixtures of local experts</article-title>
          .
          <source>Neural Comput.</source>
          ,
          <volume>3</volume>
          (
          <issue>1</issue>
          ):
          <fpage>79</fpage>
          -
          <lpage>87</lpage>
          ,
          <year>1991</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>G.</given-names>
            <surname>Linden</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Smith</surname>
          </string-name>
          ,
          <string-name>
            <given-names>and J.</given-names>
            <surname>York</surname>
          </string-name>
          . Amazon.
          <article-title>com recommendations: Item-to-item collaborative filtering</article-title>
          .
          <source>IEEE Internet Computing</source>
          ,
          <volume>7</volume>
          (
          <issue>1</issue>
          ):
          <fpage>76</fpage>
          -
          <lpage>80</lpage>
          ,
          <year>2003</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <given-names>M.</given-names>
            <surname>Mandl</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Felfernig</surname>
          </string-name>
          , E. Teppan, and
          <string-name>
            <given-names>M.</given-names>
            <surname>Schubert</surname>
          </string-name>
          .
          <article-title>Consumer decision making in knowledge-based recommendation</article-title>
          .
          <source>Journal of Intelligent Information Systems</source>
          ,
          <year>2010</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <given-names>G. A.</given-names>
            <surname>Miller</surname>
          </string-name>
          .
          <article-title>The magical number seven: Plus or minus two: Some limits on our capability for processing information</article-title>
          .
          <source>Psychological Review</source>
          ,
          <volume>63</volume>
          (
          <issue>2</issue>
          ):
          <fpage>81</fpage>
          -
          <lpage>97</lpage>
          ,
          <year>1956</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <given-names>D. C.</given-names>
            <surname>Montgomery</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E. A.</given-names>
            <surname>Peck</surname>
          </string-name>
          , and
          <string-name>
            <given-names>G. G.</given-names>
            <surname>Vining</surname>
          </string-name>
          .
          <article-title>Introduction to linear regression analysis</article-title>
          . Wiley Interscience,
          <year>2006</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [16]
          <string-name>
            <given-names>R. J.</given-names>
            <surname>Quinlan</surname>
          </string-name>
          .
          <article-title>Learning with continuous classes</article-title>
          .
          <source>In 5th Australian Joint Conference on Artificial Intelligence</source>
          , pages
          <fpage>343</fpage>
          -
          <lpage>348</lpage>
          ,
          <year>1992</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [17]
          <string-name>
            <given-names>T. L.</given-names>
            <surname>Saaty</surname>
          </string-name>
          and
          <string-name>
            <given-names>M. S.</given-names>
            <surname>Ozdemir</surname>
          </string-name>
          .
          <article-title>Why the magic number seven plus or minus two</article-title>
          .
          <source>Mathematical and Computer Modelling</source>
          ,
          <volume>38</volume>
          (
          <issue>3</issue>
          ):
          <fpage>233</fpage>
          -
          <lpage>244</lpage>
          ,
          <year>2003</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          [18]
          <string-name>
            <given-names>A. I.</given-names>
            <surname>Schein</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Popescul</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L. H.</given-names>
            <surname>Ungar</surname>
          </string-name>
          , and
          <string-name>
            <given-names>D. M.</given-names>
            <surname>Pennock</surname>
          </string-name>
          .
          <article-title>Methods and metrics for cold-start recommendations</article-title>
          .
          <source>In SIGIR</source>
          ,
          <year>2002</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          [19]
          <string-name>
            <given-names>B.</given-names>
            <surname>Scholkopf</surname>
          </string-name>
          and
          <string-name>
            <given-names>A. J.</given-names>
            <surname>Smola</surname>
          </string-name>
          .
          <article-title>Learning with Kernels: Support Vector Machines, Regularization, Optimization, and Beyond</article-title>
          . MIT Press,
          <year>2001</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          [20]
          <string-name>
            <given-names>B.</given-names>
            <surname>Schwartz</surname>
          </string-name>
          .
          <article-title>The tyranny of choice</article-title>
          .
          <source>Scientific American</source>
          , pages
          <fpage>71</fpage>
          -
          <lpage>75</lpage>
          ,
          <year>April 2004</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          [21]
          <string-name>
            <given-names>Y.</given-names>
            <surname>Wang</surname>
          </string-name>
          and
          <string-name>
            <given-names>I. H.</given-names>
            <surname>Witten</surname>
          </string-name>
          .
          <article-title>Induction of model trees for predicting continuous classes</article-title>
          .
          <source>In Poster papers of the 9th European Conference on Machine Learning</source>
          . Springer,
          <year>1997</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          [22]
          <string-name>
            <given-names>K.</given-names>
            <surname>Yoon</surname>
          </string-name>
          and
          <string-name>
            <given-names>C.</given-names>
            <surname>Hwang</surname>
          </string-name>
          .
          <article-title>Multiple attribute decision making: An introduction</article-title>
          . Sage University papers,
          <volume>7</volume>
          (
          <issue>104</issue>
          ),
          <year>1995</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          [23]
          <string-name>
            <given-names>L.</given-names>
            <surname>Zadeh</surname>
          </string-name>
          . Fuzzy Sets,
          <source>Fuzzy Logic and Fuzzy Systems. World Scientific</source>
          ,
          <year>1996</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>