<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Collaborative Filtering With Adaptive Information Sources</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Neal Lathia</string-name>
          <email>n.lathia@cs.ucl.ac.uk</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Xavier Amatriain, Josep M. Pujol</string-name>
          <email>xar, jmps@tid.es</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Department of Computer Science, University College London</institution>
          ,
          <addr-line>Gower Street, WC1 E6BT</addr-line>
          ,
          <country country="UK">UK</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Telefonica Research</institution>
          ,
          <addr-line>Via Augusta 177, Barcelona 081290</addr-line>
          ,
          <country country="ES">Spain</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>Collaborative ltering (CF) algorithms, which generate recommendations for web users by predicting user-item ratings, are often evaluated according to their predictions; in this context the problem of generating recommendations can be formulated as one of tting a community of users to the best set of predictors. However, the data used to perform CF is sparse, and accuracy is limited by both the quantity and quality of information available. Mining the web has the potential to address these issues: the quality and quantity of ratings can be incremented by collecting external sources of rating information. In this work we introduce a method to perform CF with external data sources; furthermore, we show that a community of users can be partitioned according to what external source acts as a better predictor of each user's preferences. In particular, we nd that a single kNN predictor can achieve remarkably high prediction accuracy if the data sources are selected optimally: designing a recommender system can thus be approached with the focus on data quality rather than algorithmic method.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>Recommender systems, based on collaborative ltering (CF),
are displaying an evermore important and pervasive presence
on the web. The problem of generating recommendations has
been described as a prediction problem: based on a pro le of
user ratings, the system needs to predict future user ratings
for other content in the future. The approaches adopted to
perform CF can be broadly divided into two categories. The
rst are statistical approaches; these draw on the assumption
of like-mindedness between users and therefore focus on a
variety of classi ers that operate on the user-rating data; the
most prominent candidates being based on matrix
factorisation and neighbourhood methods [Koren, 2008][Herlocker et
al., 2004]. The second approach is based on user modeling;
these methods augment statistical approaches by reasoning
on the context and behaviours that emerge when people use
recommender systems; recent examples include the rising
interest in trust modeling for collaborative contexts, including
[O'Donovan and Smyth, 2005].</p>
      <p>Traditional CF suffers from the problem of data sparsity;
the ability that a system has to make predictions for a user
or item is limited by the lack of rating information. The data
also has very high dimensionality; for example, the Net ix
dataset1 includes about half a million users and about twenty
thousand movies. The mere size of the data implies that
generating recommendations is a very expensive process that is
dif cult to scale to large communities. The current focus of
much CF research is on improving the accuracy of the
algorithms applied to generate recommendations. In particular, a
number of successful statistical methods [Koren, 2008]
combine an ensemble of predictors to produce higher accuracy.
However, improving the classi cation method does nothing
to improve the data that is being used when predicting user's
preferences, and a fundamental limitating factor of any
learning algorithm applied to the CF domain is the sparsity and
potential inaccuracy of the data being used.</p>
      <p>Similarly, algorithm-centric research also deters from fully
modeling the implicit ways in which people form their
opinions. While sociologists often model preference formation
according to the principles of homophily (like-mindedness)
and social in uence (adopting the same preferences as
inuential members) [Axelrod, 1997][McPherson et al., 2001],
CF research has mainly centred its assumptions on the former
theory. Although the task of identifying the source of in
uence in a set of user ratings seems daunting, this theory
carries with it the assumption that there are a range of sources
where users may form their opinions; in particular, not all
users form their opinions by eliciting information from
similar neighbours. The problem is thus how to model the way
people form their opinions.</p>
      <p>Mining the web for publicly available ratings has the
potential to address the sparsity problem by drawing on the
assumptions of social in uence: there are a great number of
online resources that contain a vast amount of ratings that
may be accessed by users as they form their opinions. In this
work we therefore propose to explore four different source
datasets and evaluate the predictive power they have on a
test set of user-movie ratings. Two of these source datasets
were collected from the web, while the second two are
de1http://www.net ixprize.com</p>
    </sec>
    <sec id="sec-2">
      <title>Number of ratings per user 10 100</title>
    </sec>
    <sec id="sec-3">
      <title>Number of ratings</title>
    </sec>
    <sec id="sec-4">
      <title>Netflix</title>
    </sec>
    <sec id="sec-5">
      <title>Flixster</title>
      <p>Rot en T.
1000
10000
0.1 Standard0.D2eviation0.3</p>
      <sec id="sec-5-1">
        <title>Flixster</title>
      </sec>
      <sec id="sec-5-2">
        <title>Netflix</title>
      </sec>
      <sec id="sec-5-3">
        <title>Rotten T. 0.4</title>
        <p>rived from the a set of training data, based on neighbours and
power users; Section 2 describes these datasets, and Section
3 highlights the statistical features that emerge between the
sets. In Section 4 we introduce the method we implement to
perform cross-dataset predictions, and Section 5 reports and
analyses the results when each source dataset is used to make
predictions on a common test set.</p>
        <p>Our main result is that the accuracy of a CF prediction
algorithms heavily depends on the quality of the information
used to generate predictions, and the most appropriate source
is user-dependent. In particular, matching users to the
correct source of rating information has the potential to produce
highly accurate recommendations when using a simple
userbased NN algorithm. We evaluate a number of benchmark
methods that attempt to achieve this goal in Section 6; we thus
introduce a novel perspective to CF, where the focus should
not be so much on the method applied, but on the data that is
used.
2</p>
        <sec id="sec-5-3-1">
          <title>Information Sources</title>
          <p>In this work we ran experiments using a subset of the
Netix prize data. Our subset consists of randomly
selected Net ix users from the training set, and each of these
user's probe ratings as a test set. To compliment this dataset
we crawled two different sources of rating pro les: Rotten
Tomatoes2 and Flixster3. Based on how ratings are input into
each of these systems, we call these sources experts and
enthusiasts respectively:</p>
          <p>Experts: The Rotten Tomatoes portal aggregates a number
of cinema critic reviews from a wide range of web sources,
including newspapers, specialized websites, and magazines.
The critics use different rating scales; some range from 1-10
stars, others 1-5, and some use a 100-point scale. However,
all of these ratings can be normalised. For example, a out
of star rating is the same as out of ; we adopt a
simple linear transpose to re-intepret ratings from one scale to
another.</p>
          <p>Enthusiasts: Flixster is one of the largest movie-oriented
social networks, and therefore contains ratings given by the
site's movie-enthusiast subscribers. We collected the pro les
of the top- users from Flixster. However, not all users set
their pro les to public: this reduced our collected dataset to
users. The Flixster users rate movies on a 1-5 star scale,
but also have a further two options available: want to see
(WS), and not interested (NI). In fact, the majority of
ratings in the data fall into one of these two latter categories.
2http://www.rottentomatoes.com
3http://www. ixster.com/
5000
0040
rsseU 3000
fr
o
ubeNm 2000
0.8
0.6
F
D
C0.4
0.2
00</p>
          <p>Due to string-matching inconsistencies between the movie
titles in Net ix and the crawled datasets, our datasets contain
out of the available in Net ix; furthermore, to
accomodate for this, we also cut any Net ix users who had
no training ratings within this set of movies. A summary of
the size of each dataset is given in Table 1. The analysis in
Section 3 is based on this number of movies, which could
be identi ed in all three datasets. We also compare the
predictive performance of the external data sources to two other
sets, each derived the the Net ix training subset we use:</p>
          <p>Neighbours: The benchmark performance that we
compare the above sources to is the approach adopted by
traditional user-based NN; there is no distinction made between
users, who all come from the same community.</p>
          <p>Power Users: This group is a subset of the neighbours
group, the user pro les that, based a simple measure of
prole size, are deemed to carry a signi cant amount of reliable
information, and represent a sub sample of the community
that may offer powerful predictions for the rest of the users.
The idea of power users has been explored in the past [Cho
et al., 2007], and usually relies on identifying users based on
a number of pre-de ned heuristics. In this work we focus
on pro le size; that is, we assume that users who are
proactively rating more items are following different behaviours to
the casual rater [Herlocker et al., 2004]. Note that all of the
above groups strongly differ to results that would be obtained
from clustering algorithms: users are grouped either based
on pro le attributes (rather than the value of their ratings), or
based on where their ratings were crawled from. It is thus
not guaranteed that users in the same group will agree with
each other, whereas clustering algorithms tend to group users
based on a notion of similarity.</p>
        </sec>
        <sec id="sec-5-3-2">
          <title>Comparing Information Sources</title>
          <p>In this section, we compare the three datasets. As introduced
above, they can already be differentiated from one another
according to a broad characterisation of the end-users of each
system; however, in this section we examine the extent that
ratings from different sources will differ in terms of summary
statistics: the number and distribution of ratings, the sparsity,
and rating deviation between per user.</p>
          <p>Number of Ratings: With less than of the users,
the Flixster data contains over times the number of ratings
than the Rotten Tomatoes data. The same feature can be
observed in Figure 1, which shows the cumulative distribution
(CDF) of the number of ratings in each dataset. As the plot
shows, 60% of the Rotten Tomatoes experts have about or
less ratings, 60% of the Net ix set users have ratings or
less, but the same proportion of Flixster reaches up to
ratings.</p>
          <p>Sparsity: Table 1 reports the sparsity values for each
dataset; once again, Rotten Tomatoes and Net ix share
similar sparsity values, while the Flixster dataset is a much denser
set of ratings. The table also reports two separate sparsity
values for the Flixster dataset. The rst value, though
excluding both WS and NI ratings, shows that the dataset is</p>
          <p>sparse. Including all these extra ratings reduces the
sparsity to : in both cases, the user-rating matrix
contains a much larger amount of ratings than Net ix alone. We
also measured how the sparsity uctuates as different sub
samples of users are selected, re ecting the evaluation of
power users we report in the Section 5. If we only select
users who have rated more than movies, both the number
of users and resulting dataset sparsity will change. Figure 2
shows how these changes are affected by the rating
threshold. The plots show that there is an uneven distribution of
ratings amongst the users themselves, reinforcing the notion
that users will behave differently as they interact with the
system. The plots also con rm what was observed in Figure 1:
the Net ix dataset, while having the highest number of users,
also is also the sparsest of the three datasets.</p>
          <p>Standard Deviation: Looking at this aspect aims to see
the extent that each group of users agrees with each other by
capturing the spread of ratings around each movie mean. As
Figure 1 shows, the distribution of standard deviation values
is very different from one community to the next. There are
also a very small proportion of Flixster pro les that appear
to be outliers: their pro les are full of the same rating for
nearly all content. This causes the standard deviation over
their ratings to be less than . These Flixster outliers will
not be able to contribute useful information to any prediction,
and can therefore be safely ignored.
0 1 ratings. The rst has a very high average, , and only
"!$#%&amp;(’ # &amp; ilarity: the similarity between two users and
)+*-,/. is scaled according to how many items their pro les
5 78 6 item being predicted for user (who has mean rating and
9 8 standard deviation ) would map to:
4 itive opinion, a Flixster neighbour 's NI and WS ratings for
23 NI ratings. The second has a very low average, ,
do better? In this light, and due to the semantics of
crossdataset prediction, we focus on a single method: the NN
algorithm.</p>
          <p>The NN can be built according to either the item-based
or user-based paradigm. Both methods operate in very
similar ways, and differ only in how they assume the underlying
data is structured. Here we only consider the user-based
approach. Once again, this makes our cross-dataset prediction
highly explainable and transparent, which is a key aspect in
building recommender systems that users trust [Herlocker et
al., 2004].</p>
          <p>Implementing a NN CF algorithm can be decomposed
into three steps: (a) neighbourhood formation, where the top
neighbours are computed for each user, (b) rating
aggregation, where ratings for an item are collected and used to make
a predicted rating, and nally (c) recommendation and
feedback, where users update their pro les by responding to the
recommendations they are given. When using external data
sources, neighbours for a user in the training dataset are found
from the source set; similarly, ratings for items in the test
set will be predicted using the ratings in the selected source
set. We identify neighbours using a weighted Cosine
Simshare in common. We base neighbour selection on a
similarity threshold: any neighbour with similarity greater than
zero is included. A predicted rating is then computed as a
weighted average of deviations from each neighbour's mean
[Herlocker et al., 2004]. The only limitation we impose is
a measure of prediction con dence: if less than
neighbour ratings have been found for a prediction, the prediction
is set to the user mean. Setting the similarity threshold at
zero and the con dence at may not be optimal values for
each dataset; in this work we focus on evaluating the
ability to use adaptive information sources when generating
predictions, rather than simply tweaking the algorithm itself for
optimal performance.</p>
          <p>We also noted that the Flixster dataset contains two
additional ratings, NI and WS. It is not immediately transparent
how these ratings should be transposed onto the scale in the
target dataset, since they are dif cult to place on an ordinal
scale of ratings; however, a relationship between the NI
rating and the movie average emerges in a few select cases. For
example, consider two different movies that each have
and NI ratings: it seems possible to assume that NI is
roughly equivalent to a form of negative feedback provided
by the user. However, we decided to ignore these ratings in
the neighbourhood formation part of our algorithm.</p>
          <p>We did test methods to include these ratings in the
prediction step. Drawing from the assumption that NI may act as a
form of negative feedback and WS represents a potential
posThere are a number of methods that can be implemented in
order to use the ratings of the above datasets to predict the test
set. In particular, one may simply combine all of the datasets
into a single, larger training set that can be fed into any
learning algorithm. However, in this work we aim to evaluate the
potential that disparate sources have to predict a common test
set: how well do experts predict the crowds? Do enthusiasts
(1)
(2)
4</p>
        </sec>
        <sec id="sec-5-3-3">
          <title>Collaborative Filtering With External Data</title>
          <p>Results when both using and ignoring these kind or ratings
are reported in the following section.
- improve the aggregate accuracy from to : users,
N shows the proportions of the test users dataset who
afL 's source set. From this matrix, we were able to compute
We measure the accuracy of predictions using the Root Mean
Squared Error (RMSE) [Herlocker et al., 2004]. We divide
our experiments into two parts; in the rst, we report the
results when using different power user groups from within
the Net ix dataset as sources. We then compare performance
across the three source datasets.</p>
          <p>Predicting With Power Users: Table 2 shows the
relationship the rating threshold used to de ne power users,
the number of power users found who match this criteria, and
the RMSE achieved when each group is used as the source set
for predictions. The table shows that as the rating-threshold is
increased, accuracy worsens. However, there are two points
to note here: (a) as above, the algorithm has not been fully
tuned for optimal performance (which is a dataset-dependent
problem, subject to both the similarity metric and rating
aggregation method implemented), and (b) relying on an
aggregate error measure like the RMSE does not highlight the
performance that is being achieved on a per-user basis. In
other words, a single RMSE value does not show whether
some users are better suited to certain information sources
than others.</p>
          <p>Based on the Table 2 alone, it seems that removing
nonpower users from the dataset results in a loss of prediction
accuracy. To explore the veracity of this impression, we built
a second matrix; in this case, each row corresponded to a user,
and each column represented a source set of power users
(according to a rating threshold). Each entry (BML ) in the matrix
is the RMSE achieved on user 's test ratings with column
the proportion of users who best af liated with each subset
of power users. In other words, we could determine which
source dataset was the most appropriate per user. Table 2 also
liated best with each subgroup of power users. As the plot
shows, there is in fact a very large proportion of users whose
best source of predictions is the set of power users who have
rated more than movies. An algorithm which could
manage to select the best power set group for each user would
therefore, associate differently with different sets of power
users, and simply targetting a global optimal does not achieve
the best possible accuracy. Conversely, it is also possible
Source
Rotten Tomatoes
Flixster-WI/NS</p>
          <p>Net ix</p>
          <p>Flixster
User-Matched</p>
          <p>Item-Matched
User-Item Matched
to infer from these experiments that the traditional
nearestneighbour model, based on selecting the best neighbours
for each user, is not optimal: improved accuracy is obtained
when user groups are made apriori, and users then nd their
neighbours within those groups.</p>
          <p>Predicting With External Data: The accuracy results
when using all the external datasets are reported in Table 3.
As the table shows, the overall accuracy when making
predictions with an external sources is worse than simply
using neighbours; from these values it would appear that
exclusively using external data sources is not a viable option when
designing a CF algorithm. The only point worth noting is that
including the NI/WS ratings when making predictions using
Flixster as a source provided an improvement. However, once
again we constructed the user-data source RMSE matrix, and
were able to extract the performance that a user-based NN
predictor would achieve if it were able to perfectly match
users, items, or user-item pairs to the correct source data set.
This time, each column of the error-matrix represented a
different source set. The improvement is remarkable: assigning
users to the correct data source provides (with this
subsample of the data) an accuracy below the target of the Net ix
prize. Repeating the above analysis across the three datasets
also shows that there is no absolutely dominant source, the
right column of Table 3 shows. The semantic interpretation
of these results is that the Net ix dataset optimally predicts
only half of the sampled population; movie critics and
enthusiasts are more accurate for the others.
6</p>
        </sec>
        <sec id="sec-5-3-4">
          <title>Benchmark Methods</title>
          <p>Based on the above work it becomes apparent that classi ers
like user-based NN can achieve very high accuracy in the
context of recommender systems if users are paired with the
correct source of data; the main problem is thus how to
infer, given a user pro le, the correct source. Our rst attempt
considered various qualities of each user's pro le, such as
pro le size, mean rating, rating standard deviation, and
average agreement of the user's pro le items with the movies'
mean ratings. However, none of the individual components
correlated strongly with each user's selection of optimal data
source.</p>
          <p>Our current work therefore focuses on how to infer what
the best data source is for each user. In this work we
propose and evaluate benchmark results derived from two
methods. The rst is based on linear combinations of the dataset</p>
          <p>RMSE
1.0215
0.9851
1.0253
1.013
0.9829
1.1243
1.208
0.9701
(3)
Weighted Combinations We rst tried a variety of linear
combinations of each sources' predictions for each user's test
items. As shown in Table 6, these ranged from weighting
each source equally, to weighting each source according to
the average shared similarity with the target user, weighting
according to how well the target's pro le ts the movie means
generated from each source, and weighting according to how
well each target ts the neighbourhood in each source. All
the linear combinations of each sources' predictions failed to
produce more accurate results on the test set than using the
Group
Min Avg Similarity</p>
          <p>Net ix 0.457
Rotten Tomatoes 0.269</p>
          <p>Flixster 0.277
Min Group-Mean RMSE</p>
          <p>Net ix 0.410
Rotten Tomatoes 0.261</p>
          <p>Flixster 0.255
Min Training Set RMSE</p>
          <p>Net ix 0.307
Rotten Tomatoes 0.247</p>
          <p>Flixster 0.263
Training Set Con dence</p>
          <p>Net ix 0.457
Rotten Tomatoes 0.0</p>
          <p>Flixster 0.2</p>
          <p>Method</p>
          <p>Equal (1/3)</p>
          <p>Avg Similarity
Group-Mean RMSE
Training Set RMSE</p>
          <p>Max Avg. Similarity
Min Group-Mean RMSE</p>
          <p>Best Training RMSE</p>
          <p>Most Training Con dence
F or weight the contribution of each source proportionally to
5 ing proportions of co-rated items with 's training set ratings,
OBQSR OBT lected (OBP , , ). Assuming that a relationship exists
be5 group is as follows. Any given user will have three
poten5 predictions, where a prediction of item for user is
generated by each source, and the nal prediction is computed as
a weighted average of each score. The second method
preclassi es each user to a particular dataset, and only computes
one prediction per user-item pair using the selected dataset.</p>
          <p>This way, we can evaluate the method both in terms of the
aggregate RMSE and the precision/recall metric related to how
well the method paired each user with the appropriate dataset.</p>
          <p>The results we report here can be broadly categorised into
two groups. The rst are structural properties: We weight
(or select) datasets based on emergent structural properties of
the NN algorithm; in particular, we measure the role that the
similarity function plays when correlating a user to each set,
by looking at the average positive similarity the users share
with each source set. The second, user- t RMSE, includes
methods that weight (or select) the best source set according
to how well each user ts the three sources, or how well each
source predicts the user's training pro le. This process
entails a two-fold use of each user's training set of ratings; it
is rst used to compute similarity weights with members of
each source set, and a second time to measure how well each
source predicts the user's pro le. The motivation for the latter
tial neighbourhoods: Net ix (N), Rotten Tomatoes (RT), and
Flixster (F). Each of these neighbourhoods will contain
varyand the RMSE between these co-rated movies can be
coltween each source's RMSE on the user's training pro le and
the predictive performance on the same user's test ratings, we
can either select the source that provides the lowest RMSE,
its accuracy on the training items:
Precision</p>
          <p>Recall
Net ix source alone. One of the primary reasons for this was
that, in many cases, each source produced diverging
predictions from the next: linear combinations of polarising
predictions therefore hurt the overall results.</p>
          <p>User-Source Classi cation The second set of experiments
were performed in two steps. The rst step assigns each user
to a source by generating a mapping from each user to the
the categorical set of sources, while the second step uses the
mapping to generate predictions for each user's test set with
the assigned source. As above, we tried classifying users
according to how well they t the source's mean ratings, their
neighbourhood in each source, or by selecting the group that
the user shares the highest amount of similarity with. We
complimented these with a classi er that operated on how
much con dence, or number of ratings, each source has about
the target user's pro le; the idea being that a user's behaviour
mimics that of a particular source if both consistently rate
the same items. Based on this methodology, we can measure
two results: (a) the RMSE achieved on the test set after
preclassifying each user, and (b) how well the pre-classi cation
step works. We evaluated the latter based on the precision
and recall metrics.</p>
          <p>In this case, we nd that the RMSE results are more
encouraging: they differ from when only using the Net ix
dataset by less than using the average-similarity
classi cation, and by when the con dence-based classi er
is implemented. However, exploring the precision and recall
metrics in Table 6 highlights why these results were obtained:
in the latter case, nearly all the users have been mapped to the
Net ix source, thus producing the same results. Majority of
the recall values are low, indicating a high proportion of
misclassi cations. Examining the results highlighted the fragility
of the pre-classi cation step, and the dependence it had on the
sources. In other words, users who were wrongly assigned to
the Net ix source did not contribute as much error as those
who were wrongly mapped to the smaller Flixster and Rotten</p>
          <p>Tomatoes datasets.
5 lated as follows: given a user pro le , what pro le features
performing a further crawl of Flixster would supply the NN
algorithm with a richer set of enthusiast movie raters. The
main focus of our future work will be combining the above
results, in order to match to the best subset of a source dataset.</p>
          <p>The potential of mining the web for rating information thus
shifts the focus of building an accurate CF algorithm away
from the algorithm and toward matching users to the
appropriate information sources. The problem can thus be
formuand emergent-structural properties of the NN algorithm can
be used to match the user to the best dataset? The
preliminary experiments we report in Section 6 are promising, but
still lack in the desired performance. In fact, alternative
classi cation methods, with varying levels of dependence on the
quality of the rating data, may perform better. We found that
the strongest improvement was measured when data sets were
adaptively selected for each user: the main result we observed
is that classi cation accuracy is strongly related to the data
source that is used, and improvement to the aggregate, global
RMSE is proportional to how well users and data sources are
linked. A viable option for building a CF system, therefore,
need not rely on a combination of predictors [Koren, 2008],
but rather on an optimal combination of data sources.</p>
          <p>The idea of using experts has been used before in CF research.</p>
          <p>The work by [Su et al., 2007] de nes experts as the
algorithms that can be used to produce predictions; the authors
construct a hybrid CF algorithm that outputs a weighted
average of multiple CF algorithms. This signi cantly departs
from the de nition we apply here, where expertise is a quality
of the data and not of the method applied to generate
predictions using it. On the other hand, [Cho et al., 2007] de ne
experts as a subset of the users of a community based on a
number of heuristics. In particular, expertise in an a domain
is based on how many items a user has rated in that domain.</p>
          <p>This de nition is closer to the way we identify power users,
based on rating frequency, although we do not differentiate
between domains within the items that can be rated.</p>
          <p>Previous work [Aciar et al., 2007] has also considered the
problem of source selection; however, Aciar et al address
the problems of identifying, selecting, and retrieving
unstructured information from the web in order to produce
recommendations. Sources are selected based on quanti able
relevance and considering how complete, diverse, and timely the
data the sources contain is. The authors therefore propose
a trust model to effectively select data sources. However,
they adopt the broader goal of producing recommendations,
while the work above centres on improving the accuracy of
recommender system algorithms with a basic model of
social in uence. Examining how the quality of data relates to
performance has also been discussed in the context of
computational trust. In particular, [O'Donovan and Smyth, 2005]
considers that users are more trustworthy sources of
information if they tend to provide ratings that are good predictors
of neighbour preferences. In our case we seek to identify the
most trustworthy source of data per user in the Net ix
community. Measuring trust based on a history of accurate
predictions is similar to the baseline experiment in Section 6 that
focused on how well users ts each source.</p>
          <p>EGF ratings ();: and ) on an ordinal scale. First, we examined</p>
          <p>The primary motivation of this work was to highlight the
dependence of CF algorithm's performance on the quality of the
data that is being used to predict user preferences. We
therefore explored the potential that a variety of datasets from the
web have to predict a sample set of Net ix users. In doing
so, we proposed a framework for cross dataset prediction,
including methods to normalise data and interpret non-numeric
the effect of learning to classify items based on a dense
subset of the available training data, by extracting power users
from the Net ix training set. We then analysed the
predictive potential of external data sources, based on a
collaborative method that generates a neighbourhood for a Net ix user
composed of Flixster or Rotten Tomatoes pro les. We
identi</p>
          <p>ed that the predictive power of both the power-user subsets
and external sources is user-dependent; there are some users
who are best predicted by power users, others by experts,
enthusiasts, or neighbours. The two experiments, however, are
not mutually exclusive. In fact, power users can also be
identi ed and exploited within the Rotten Tomatoes dataset, and
8</p>
        </sec>
        <sec id="sec-5-3-5">
          <title>Conclusion</title>
        </sec>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [Aciar et al.,
          <year>2007</year>
          ]
          <string-name>
            <given-names>S.</given-names>
            <surname>Aciar</surname>
          </string-name>
          , J. L.
          <article-title>de la Rosa i Esteva, and</article-title>
          <string-name>
            <given-names>J. L.</given-names>
            <surname>Herrera</surname>
          </string-name>
          .
          <article-title>Information Sources Selection Methodology for Recommender Systems Based on Intrinsic Characteristics and Trust Measure</article-title>
          .
          <source>In Proceedings of The 5th Workshop on Intelligent Techniques for Web Personalization</source>
          , Vancouver, Canada,
          <year>2007</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          <source>[Axelrod</source>
          ,
          <year>1997</year>
          ]
          <string-name>
            <given-names>R.</given-names>
            <surname>Axelrod</surname>
          </string-name>
          .
          <article-title>The dissemination of culture: A model with local convergence and global polarization</article-title>
          .
          <source>Journal of Con ict Resolution</source>
          , (
          <volume>41</volume>
          ):
          <volume>203</volume>
          
          <fpage>226</fpage>
          ,
          <year>1997</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [Cho et al.,
          <year>2007</year>
          ]
          <string-name>
            <given-names>Jinhyung</given-names>
            <surname>Cho</surname>
          </string-name>
          , Kwiseok Kwon, and Yongtae Park.
          <article-title>Collaborative ltering using dual information sources</article-title>
          .
          <source>IEEE Intelligent Systems</source>
          ,
          <volume>22</volume>
          (
          <issue>3</issue>
          ):
          <volume>30</volume>
          
          <fpage>38</fpage>
          ,
          <year>2007</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [Herlocker et al.,
          <year>2004</year>
          ]
          <string-name>
            <given-names>J.</given-names>
            <surname>Herlocker</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Konstan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Terveen</surname>
          </string-name>
          , and
          <string-name>
            <given-names>J.</given-names>
            <surname>Riedl</surname>
          </string-name>
          .
          <article-title>Evaluating collaborative ltering recommender systems</article-title>
          .
          <source>In ACM TOIS</source>
          , volume
          <volume>22</volume>
          , pages
          <fpage>5</fpage>
          <lpage></lpage>
          53. ACM Press,
          <year>2004</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          <source>[Koren</source>
          , 2008]
          <string-name>
            <given-names>Y.</given-names>
            <surname>Koren</surname>
          </string-name>
          .
          <article-title>Factorization meets the neighborhood: A multifaceted collaborative ltering model</article-title>
          .
          <source>In ACM SIG KDD Conference</source>
          ,
          <year>2008</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          <string-name>
            <surname>[McPherson</surname>
          </string-name>
          et al.,
          <year>2001</year>
          ]
          <string-name>
            <given-names>M.</given-names>
            <surname>McPherson</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Smith-Lovin</surname>
          </string-name>
          ,
          <article-title>and</article-title>
          <string-name>
            <given-names>J.M.</given-names>
            <surname>Cook</surname>
          </string-name>
          .
          <article-title>Homophily in social networks</article-title>
          .
          <source>Annual Review of Sociology</source>
          , (
          <volume>27</volume>
          ):
          <volume>415</volume>
          
          <fpage>444</fpage>
          ,
          <year>2001</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          <source>[O'Donovan and Smyth</source>
          , 2005
          <string-name>
            <surname>] J. O'Donovan</surname>
            and
            <given-names>B.</given-names>
          </string-name>
          <string-name>
            <surname>Smyth</surname>
          </string-name>
          .
          <article-title>Trust in recommender systems</article-title>
          .
          <source>In IUI '05: Proceedings of ACM IUI</source>
          , pages
          <volume>167</volume>
          
          <fpage>174</fpage>
          . ACM Press,
          <year>2005</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [Su et al.,
          <year>2007</year>
          ]
          <string-name>
            <given-names>Xiaoyuan</given-names>
            <surname>Su</surname>
          </string-name>
          , Russell Greiner,
          <string-name>
            <surname>Taghi M. Khoshgoftaar</surname>
            , and
            <given-names>Xingquan</given-names>
          </string-name>
          <string-name>
            <surname>Zhu</surname>
          </string-name>
          .
          <article-title>Hybrid collaborative ltering algorithms using a mixture of experts</article-title>
          .
          <source>In Web Intelligence</source>
          , pages
          <fpage>645</fpage>
          
          <fpage>649</fpage>
          ,
          <year>2007</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>