<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <journal-title-group>
        <journal-title>September</journal-title>
      </journal-title-group>
    </journal-meta>
    <article-meta>
      <title-group>
        <article-title>Case Study Evaluation of Mahout as a Recommender Platform</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Carlos E. Seminario</string-name>
          <email>cseminar@uncc.edu</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>David C. Wilson</string-name>
          <email>davils@uncc.edu</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Software and Information Systems Dept., University of North Carolina Charlotte</institution>
          ,
          <country country="US">USA</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2012</year>
      </pub-date>
      <volume>9</volume>
      <issue>2012</issue>
      <fpage>45</fpage>
      <lpage>50</lpage>
      <abstract>
        <p>Various libraries have been released to support the development of recommender systems for some time, but it is only relatively recently that larger scale, open-source platforms have become readily available. In the context of such platforms, evaluation tools are important both to verify and validate baseline platform functionality, as well as to provide support for testing new techniques and approaches developed on top of the platform. We have adopted Apache Mahout as an enabling platform for our research and have faced both of these issues in employing it as part of our work in collaborative ltering. This paper presents a case study of evaluation focusing on accuracy and coverage evaluation metrics in Apache Mahout, a recent platform tool that provides support for recommender system application development. As part of this case study, we developed a new metric combining accuracy and coverage in order to evaluate functional changes made to Mahout's collaborative ltering algorithms.</p>
      </abstract>
      <kwd-group>
        <kwd>Recommender systems</kwd>
        <kwd>Evaluation</kwd>
        <kwd>Mahout</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Categories and Subject Descriptors</title>
      <p>H.3.3 [Information Storage and Retrieval]: Information
Search and Retrieval{Information ltering</p>
    </sec>
    <sec id="sec-2">
      <title>1. INTRODUCTION</title>
      <p>Selecting a foundational platform is an important step in
developing recommender systems for personal, research, or
commercial purposes. This can be done in many di erent
ways: the platform may be developed from the ground up,
an existing recommender engine may be contracted (e.g.,
Permission to make digital or hard copies of all or part of this work for
personal or classroom use is granted without fee provided that copies are
not made or distributed for profit or commercial advantage and that copies
bear this notice and the full citation on the first page. To copy otherwise, to
republish, to post on servers or to redistribute to lists, requires prior specific
permission and/or a fee.</p>
      <p>OracleAS Personalization1), code libraries can be adapted,
or a platform may be selected and tailored to suit (e.g.,
LensKit2, MymediaLite3, Apache Mahout4, etc.). In some
cases, a combination of these approaches will be employed.</p>
      <p>For many projects, and particularly in the research
context, the ideal situation is to nd an open-source platform
with many active contributors that provides a rich and
varied set of recommender system functions that meets all or
most of the baseline development requirements. Short of
nding this ideal solution, some minor customization to an
already existing system may be the best approach to meet
the speci c development requirements. Various libraries have
been released to support the development of recommender
systems for some time, but it is only relatively recently
that larger scale, open-source platforms have become readily
available. In the context of such platforms, evaluation tools
are important both to verify and validate baseline platform
functionality, as well as to provide support for testing new
techniques and approaches developed on top of the platform.
We have adopted Apache Mahout as an enabling platform
for our research and have faced both of these issues in
employing it as part of our work in collaborative ltering
recommenders.</p>
      <p>
        This paper presents a case study of evaluation for
recommender systems in Apache Mahout, focusing on metrics
for accuracy and coverage. We have developed functional
changes to the baseline Mahout collaborative ltering
algorithms to meet our research purposes, and this paper
examines evaluation both from the standpoint of tools for baseline
platform functionality, as well as for enhancements and new
functionality. The objective of this case study is to evaluate
these functional changes made to the platform by comparing
the baseline collaborative ltering algorithms to the changed
algorithms using well known measures of accuracy and
coverage [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ]. Our goal is not to validate algorithms that have
already been tested previously, but to assess whether, and
to what extent, the functional enhancements have improved
the accuracy and coverage performance of the baseline
outof-the-box Mahout platform. Given the interplay between
accuracy and coverage in this context, we developed a
unied metric to assess accuracy vs. coverage trade-o s when
evaluating functional changes made to Mahout's
collaborative ltering algorithms.
1http://download.oracle.com/docs/cd/B10464 05/bi.904/
b12102/1intro.htm
2http://lenskit.grouplens.org/
3http://www.ismll.uni-hildesheim.de/mymedialite/
4http://mahout.apache.org
      </p>
    </sec>
    <sec id="sec-3">
      <title>RELATED WORK</title>
      <p>
        Revisiting evaluation in the context of recommender
platforms has received recent attention in the thorough
evaluation of the LensKit platform using previously tested
collaborative ltering algorithms and metrics, as reported in [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]. A
comprehensive set of guidelines for evaluating recommender
systems was provided by Herlocker et al [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ]; these guidelines
highlight the use of evaluation metrics such as accuracy and
coverage and suggest the need for an ideal \general
coverage metric" that would combine coverage with accuracy to
yield an overall \practical accuracy" measure. Many of these
evaluation metrics and techniques have also been covered
recently in [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ].
      </p>
      <p>
        Recommender system research has been primarily
concerned with improving recommendation accuracy [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ];
however, other metrics such as coverage [
        <xref ref-type="bibr" rid="ref10 ref4">10, 4</xref>
        ] and also novelty
and serendipity [
        <xref ref-type="bibr" rid="ref3 ref6">6, 3</xref>
        ] have been deemed necessary because
accuracy alone is not su cient to properly evaluate the
system. Mcnee et al [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ] states that recommendations that are
most accurate according to the standard metrics are
sometimes not the most useful to users and outlines a more
usercentric approach to evaluation. The interplay between
accuracy and other metrics such as coverage and serendipity
creates trade-o s for recommender system implementers and
this has been widely discussed in the literature, e.g., see [
        <xref ref-type="bibr" rid="ref3 ref4">4,
3</xref>
        ] and our previous work discussing trade-o s between
accuracy and robustness [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ].
      </p>
    </sec>
    <sec id="sec-4">
      <title>3. SELECTING APACHE MAHOUT</title>
      <p>To support our research in collaborative ltering,
several recommender system platforms were surveyed,
including LensKit, easyrec5, and MymediaLite. We selected
Mahout because it provides many of the desired characteristics
required for a recommender development workbench
platform. Mahout is a production-level, open-source, system
and consists of a wide range of applications that are useful
for a recommender system developer: collaborative ltering
algorithms, data clustering, and data classi cation. Mahout
is also highly scalable and is able to support distributed
processing of large data sets across clusters of computers using
Hadoop6. Mahout recommenders support various similarity
and neighborhood formation calculations, recommendation
prediction algorithms include user-based, item-based,
SlopeOne and Singular Value Decomposition (SVD), and it also
incorporates Root Mean Squared Error (RMSE) and Mean
Absolute Error (MAE) evaluation methods. Mahout is
readily extensible and provides a wide range of Java classes for
customization. As an open-source project, the Mahout
developer/contributor community is very active; the Mahout
wiki also provides a list of developers and a list of websites
that have implemented Mahout7.
3.1</p>
    </sec>
    <sec id="sec-5">
      <title>Uncovering Mahout Details</title>
      <p>Although Mahout is rich in documentation, there are
implementation details on how Mahout works that could only
be understood by looking at the source code. Thus, for
clarity in evaluation, we needed to verify the implementation
of baseline platform functionality. The following describes
some of these details for Mahout 0.4 `out-of-the-box':</p>
      <sec id="sec-5-1">
        <title>5http://easyrec.org/ 6http://hadoop.apache.org/ 7https://cwiki.apache.org/MAHOUT/mahout-wiki.html</title>
        <p>
          Similarity Weighting: Mahout implements the classic
Pearson Correlation as described in [
          <xref ref-type="bibr" rid="ref5 ref8">8, 5</xref>
          ]. Similarity weighting is
supported in Mahout and consists of the following method:
scaleFactor = 1.0 - count / (num + 1);
if (result &lt; 0.0)
        </p>
        <p>result = -1.0 + scaleFactor * (1.0 + result);
else</p>
        <p>result = 1.0 - scaleFactor * (1.0 - result);
where count is the number of co-rated items between two
users, num is the number of items in the dataset, and result
is the calculated Pearson Correlation coe cient.</p>
        <p>
          User-Based Prediction Algorithm: Mahout implements a
Weighted Average prediction method similar to the approach
described in [
          <xref ref-type="bibr" rid="ref1">1</xref>
          ], except that Mahout does not take the
absolute value of the individual similarities in the denominator,
however, it does ensure that the predicted ratings are within
the allowable range, e.g., between 1.0 and 5.0.
        </p>
        <p>
          Item-Based Prediction Algorithm: Mahout implements a
Weighted Average prediction method. This approach is
similar to the algorithm in [
          <xref ref-type="bibr" rid="ref9">9</xref>
          ], except that Mahout does not
take the absolute value of the individual similarities in the
denominator, however, it does ensure that the predicted
ratings are within the allowable range, e.g., between 1.0 and
5.0. Also, Mahout does not provide support for
neighborhood formation, e.g., similarity thresholding, for item-based
prediction.
        </p>
        <p>
          Accuracy Evaluation calculation: Mahout executes the
recommender system evaluator speci ed at run time (MAE
or RMSE) and implements traditional techniques found in
[
          <xref ref-type="bibr" rid="ref12 ref6">6, 12</xref>
          ]. For MAE, this would be,
        </p>
        <p>M AE =</p>
        <p>Pn
i=1 j ActualRatingi
n</p>
        <p>P redictedRatingi j (1)
where n is the total number of ratings predicted in the test
run.
3.2</p>
      </sec>
    </sec>
    <sec id="sec-6">
      <title>Making Mahout Fit for Purpose</title>
      <p>Through personal email communication with one of the
Mahout developers, we were informed that Mahout intended
to provide basic rating prediction and similarity weighting
capabilities for its recommenders and that it would be up
to developers to provide more elaborate approaches.
Several changes were made to the prediction algorithms and
the similarity weighting techniques for both the user-based
and item-based recommenders in order to meet our speci c
requirements and to match the best practices found in the
literature, as follows:</p>
      <p>
        Similarity weighting: De ned as Signi cance Weighting in
[
        <xref ref-type="bibr" rid="ref5">5</xref>
        ], this consists of the following method:
scaleFactor = count/50.0;
if (scaleFactor &gt; 1.0) scaleFactor = 1.0;
result = scaleFactor * result;
where count is the number of co-rated items between two
users, and result is the calculated Pearson Correlation
coe cient.
      </p>
      <p>
        User-user mean-centered prediction: After identifying a
neighborhood of similar users, a prediction, as documented
in [
        <xref ref-type="bibr" rid="ref1 ref5 ref8">8, 5, 1</xref>
        ], is computed for a target item i and target user
u as follows:
pu;i = ru +
      </p>
      <sec id="sec-6-1">
        <title>Pv V simu;v(rv;i</title>
        <p>Pv V j simu;v j
rv)
(2)
where V is the set of k similar users who have rated item i,
rv;i is the rating of those users who have rated item i, ru is
the average rating for the target user u over all rated items,
rv is the average rating for user v over all co-rated items,
and simu;v is the Pearson correlation coe cient.</p>
        <p>
          Item-item mean-centered prediction: A prediction, as
documented in [
          <xref ref-type="bibr" rid="ref1">1</xref>
          ], is computed for a target item i and target
user u as follows:
pu;i = ri +
        </p>
      </sec>
      <sec id="sec-6-2">
        <title>Pj Nu(i) simi;j (ru;j</title>
        <p>Pj Nu(i) j simi;j j
rj )
(3)
where Nu(i) is the set of items rated by user u most similar
to item i, ru;j is u's rating of item j, rj is the average rating
for item j over all users who rated item j, ri is the average
rating for target item i, and simi;j is the Pearson correlation
coe cient.</p>
        <p>
          Item-item similarity thresholding: This method was added
to Mahout and used in conjunction with the item-item
meancentered prediction described above. Similarity
thresholding, as described in [
          <xref ref-type="bibr" rid="ref5">5</xref>
          ], de nes a level of similarity that is
required for two items to be considered similar for purposes
of making a recommendation prediction; item-item
similarities that are less than the threshold are not used in the
prediction calculation.
        </p>
        <p>
          Coverage and combined accuracy/coverage metric: As
suggested in [
          <xref ref-type="bibr" rid="ref6">6</xref>
          ], the easiest way to measure coverage is to select
a random sample of user-item pairs, ask for a prediction for
each pair, and measure the percentage for which a
prediction was provided. To calculate coverage, code changes were
made to Mahout to provide, for each test run, the total
number of rating predictions requested that were unable to be
calculated as well as the total of number of rating
predictions requested that were actually calculated; the sum of
these two numbers is the total number of ratings requested.
Coverage was calculated as follows:
        </p>
        <p>Coverage =</p>
        <p>T otal#RatingsCalculated
T otal#RatingsRequested
(4)
Code changes were also made to calculate a combined
accuracy and coverage metric as de ned in Section 4.</p>
      </sec>
    </sec>
    <sec id="sec-7">
      <title>ACCURACY AND COVERAGE METRIC</title>
      <p>
        The metrics selected for this case study, accuracy and
coverage, were chosen because they are fundamental to the
utility of a recommender system [
        <xref ref-type="bibr" rid="ref10 ref6">10, 6</xref>
        ]. Although other metrics
such as novelty and serendipity can, and should, be used in
conjunction with accuracy and coverage, our objective was
to evaluate the very basic requirements of a recommender
system. Our implementation of coverage, referred to as
prediction coverage in [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ], measures the percentage of a dataset
for which the recommender system is able to provide
predictions. High coverage would indicate that the recommender
system is able to provide predictions for a large number of
items and is considered to be a desirable characteristic of
the recommender system [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ]. A combination of high
accuracy (low error rate) and high coverage are indeed desirable
by users and system operators because it improves the
utility or usefulness of the system from a user standpoint [
        <xref ref-type="bibr" rid="ref10 ref6">10,
6</xref>
        ].
      </p>
      <p>
        What constitutes `good' accuracy or coverage, however,
has not been well de ned in the literature: studies such
as [
        <xref ref-type="bibr" rid="ref10 ref4 ref5">10, 4, 5</xref>
        ] and many others, endeavor to maximize
accuracy (achieve lowest possible value) and/or coverage (achieve
highest possible value) and view these metrics on a
relative basis, i.e., how much the metric has increased or
decreased beyond a baseline value based on empirical results.
Furthermore, the interplay between accuracy and coverage,
i.e., coverage decreases as a function of accuracy [
        <xref ref-type="bibr" rid="ref3 ref4">4, 3</xref>
        ],
creates a trade-o for recommender system implementers that
has been discussed previously but not been developed
thoroughly. Inspired by the suggestion in [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ] to combine the
coverage and accuracy measures to yield an overall \practical
accuracy" measure for the recommender system, we
developed a straightforward \AC Measure" that combines both
accuracy and coverage into a single metric as follows:
ACi =
      </p>
      <p>Accuracyi
Coveragei
;
(5)
where i indicates the ith trial in an evaluation experiment.</p>
      <p>The AC Measure simply adjusts (upward) the Accuracy
according to the level of Coverage metrics found in an
experimental trial and is agnostic to the accuracy metric used,
e.g., MAE or RMSE. Using a family of curves for the Mean
Absolute Error (MAE) accuracy metric, Figure 1 illustrates
the relationship between accuracy, coverage, and the AC
Measure. As an example, following the \M AE : 0:5'' curve
we see that at 100% coverage, the AC Measure is 0.5, and
at 10% coverage, the AC Measure has increased to 5. The
intuition behind this metric is that when the recommender
system is able to provide predictions for a high percentage
of items in the dataset, the accuracy metric more closely
indicates the level of system performance; conversely, when
the coverage is low, the accuracy metric is \penalized" and is
adjusted upwards. We believe that the major bene t of the
AC Measure is that it formulates a solution for addressing
the trade-o between accuracy and coverage and can be used
to create a ranked list of results (low to high) from multiple
experimental trials to nd the best (lowest) AC Measure for
each set of test conditions. The simpli ed visualization of
the combined AC Measure shown in Figure 1 is an additional
bene t. For our evaluation purposes, the use of a combined
metric was ideal in addressing the inherent trade-o s
between accuracy and coverage, especially in the cases where
accuracy is found to be high when coverage is low; we posit
that the AC Measure will also be useful for other researchers
performing evaluations using accuracy and coverage.</p>
    </sec>
    <sec id="sec-8">
      <title>EXPERIMENTAL DESIGN</title>
      <p>The objective of this case study was to understand
Mahout's baseline collaborative ltering algorithms and
evaluate functional changes made to the platform using accuracy
and coverage metrics. The main intent of making functional
changes to Mahout recommender algorithms was to bring
the Mahout algorithms in line with best practices found in
the literature. Therefore, the overall hypothesis to be tested
in this case study was that the modi ed algorithms improve
Mahout's `out-of-the-box' prediction accuracy for both
userbased and item-based recommenders while maintaining
reasonable coverage.
5.1</p>
    </sec>
    <sec id="sec-9">
      <title>Datasets and Algorithms</title>
      <p>The data used in this study were the MovieLens datasets
downloaded from GroupLens Research8: the 100K dataset
with 100,000 ratings for 1,682 movies and 943 users
(referred to as ML100K in this study) and the 10M dataset
with 10,000,000 ratings for 10,681 movies and 69,878 users
(referred to as ML10M in this study). Ratings provided in
these datasets consist of integer values between 1 (did not
like) to 5 (liked very much).</p>
      <p>For User-based (see x3.1), Mahout uses Pearson
Correlation similarity (with and without similarity weighting),
Neighborhood formation (similarity thresholding or kNN),
and Weighted Average prediction. This was tested against
a modi ed algorithm (see x3.2) consisting of Pearson
Correlation similarity (with and without similarity weighting),
Neighborhood formation (similarity thresholding or kNN),
and Mean-centered prediction. For Item-based (see x3.1),
Mahout uses Pearson Correlation similarity (with and
without similarity weighting), no Neighborhood formation, and
Weighted Average prediction. This was tested against a
modi ed algorithm (see x3.2) consisting of Pearson
Correlation similarity (with and without similarity weighting),
Neighborhood formation (similarity thresholding), and
Meancentered prediction.
5.1.1</p>
      <p>Test Cases</p>
      <p>In order to test the overall hypothesis, the following test
cases were developed and executed for both user-based and
item-based recommenders using the ML100K and ML10M
datasets:
1. Mahout Prediction, No weighting
2. Mahout Prediction, Mahout weighted
3. Mahout Prediction, Signi cance weighted
4. Mean-Centered Prediction, No weighting
5. Mean-Centered Prediction, Mahout weighted
6. Mean-Centered Prediction, Signi cance weighted
5.1.2</p>
      <p>Accuracy and Coverage Metrics</p>
      <p>We used Mahout's MAE evaluator to measure the
accuracy of the rating predictions. For prediction coverage, we
used dataset training data to estimate the rating predictions
for the test set; the random sample of user-item pairs in our
testing was 30K pairs for ML100K and 25K pairs for ML10M
(see x3.2). AC Measures were calculated for all test cases.
5.1.3</p>
      <p>Dataset Partitioning</p>
      <p>The Mahout evaluator creates holdout 9 partitions
according to a set of run-time parameters. For the tests using the</p>
      <sec id="sec-9-1">
        <title>8http://www.grouplens.org 9Holdout is a method that splits a dataset into two parts, a</title>
        <p>ML100K dataset, the training set was 70% of the data, the
test set was 30% of the data, and 100% of the user data was
used; a total 30K rating predictions from 943 users were
requested for each test set. For the tests using the ML10M
dataset, the training set was 95% of the data, the test set
was 5% of the data, and 5% of the user data was used; a
total 25K rating predictions from 3180 users were requested
for each test set.
5.1.4</p>
        <p>Test Variations</p>
        <p>Various similarity thresholds and kNN neighborhood sizes
were executed for each test case in order to understand and
evaluate the corresponding behavior of the recommenders.
For User-based recommender testing, similarity thresholds
of 0.0, 0.1, 0.3, 0.5, and 0.7 and kNN neighborhood sizes of
600, 400, 200, 100, 50, 20, 10, 5, and 2 were tested. For
Item-based recommender testing, in addition to using no
similarity thresholding, similarity thresholds of 0.0, 0.1, 0.2,
0.3, 0.4, 0.5, 0.6, and 0.7 were tested.
6.
6.1</p>
      </sec>
    </sec>
    <sec id="sec-10">
      <title>RESULTS AND DISCUSSION</title>
    </sec>
    <sec id="sec-11">
      <title>ML10M Results</title>
      <p>Figures 2 and 3 show the results of test cases 1 through
6 for user and item-based algorithms, respectively10. The
key results of the experiment, for both user-based and
itembased algorithms unless otherwise noted, were as follows:
1. MAE for mean-centered prediction with signi cance
weighting is a signi cant improvement (p&lt;0.01) over MAE
for Mahout prediction, regardless of weighting, across
similarity thresholds (except item-based at similarity threshold
of 0.7) and kNN neighborhood sizes (except user-based at
kNN of 2, not shown).</p>
      <p>2. Mahout similarity weighting does not signi cantly
improve (p&lt;0.01) Mahout prediction MAE over prediction with
no similarity weighting (except Mahout prediction for
userbased and item-based at a similarity threshold of 0.4, not
shown). This would indicate that Mahout similarity
weighting is not very e ective as a weighting technique, especially
as compared to signi cance weighting.
6.2</p>
    </sec>
    <sec id="sec-12">
      <title>ML100K Results</title>
      <p>The results and trend lines for the ML100K experiment
are similar to ML10M. The key results, for both user-based
and item-based algorithms unless otherwise noted, were:
1. MAE for mean-centered prediction with signi cance
weighting is a signi cant improvement (p&lt;0.01) over MAE
for Mahout prediction, regardless of weighting, across
similarity thresholds and kNN neighborhood sizes (except
userbased at kNN of 400).</p>
      <p>2. Mahout similarity weighting does not signi cantly
improve (p&lt;0.01) Mahout prediction MAE over prediction with
training set and a test set, and the partitioning is performed
by randomly selecting some ratings from all, or some, of the
users. The selected ratings constitute the test set, while the
remaining ones are the training set.
10The following curves are superimposed over each other
because the values are very similar: MAE results for
meancentered prediction (no weighting and Mahout weighted),
MAE results for Mahout prediction (No weighting and
Mahout weighted), Coverage results for Mahout
prediction and mean-centered prediction (No weighting and
Mahout weighted), Coverage results for Mahout prediction and
mean-centered prediction (both Signi cance weighted).</p>
      <p>As hypothesized, results for both of the ML100K and
ML10M experiments show signi cant improvements in MAE
using the mean-centered prediction algorithm with signi
cance weighting compared to the Mahout baseline
prediction algorithm. However, when coverage is considered, the
\best" MAE results may need a second look. Can an MAE
of 0.5 or less be considered \good" when the associated
coverage is in the single digits? In this case, the recommender
system may only be able to provide recommendations to a
very small subset of its users and is a situation that must
be avoided by system operators. To help address the
accuracy vs. coverage trade-o , combined measures such as
the AC Measure (Section 4), can help by considering both
accuracy and coverage simultaneously. For the ML10M
experiment, we determined that the lowest MAE for the
Userbased algorithm using mean-centered prediction with
signi cance weighting was 0.578 at a similarity threshold of
0.7 and coverage of 0.833%; the AC Measure for this result
is calculated as 69.42. Similarly, the lowest MAE for the
Item-based algorithm using mean-centered prediction with
signi cance weighting was 0.371 at a similarity threshold of
0.7 and coverage of 1.02%; the AC Measure for this result is
calculated as 36.32. In each of these cases, the exceedingly
high values for the AC Measure indicate that these results
are not very desirable in a recommender system.</p>
      <p>Figures 4 and 5 show the AC Measure results for user and
item-based algorithms using ML10M, respectively. Rather
than show all 30 results for each algorithm (5 similarity
thresholds x 2 prediction methods x 3 weighting types), we
show only the results with calculated AC Measure values
less than 1.0; therefore, the lowest MAE results reported
above for user-based and item-based algorithms are clearly
beyond the range of this chart. We found that the best
combined accuracy/coverage results were found at higher
levels of coverage and lower levels of similarity threshold,
i.e., the best (lowest) AC Measure for user-based was 0.688
at a similarity threshold of 0.1 and for item-based was 0.665
at a similarity threshold of 0.0, both using mean-centered
prediction and signi cance weighting. We can also see that,
with few exceptions, mean-centered prediction is improved
over the Mahout prediction for the same similarity
weighting and similarity threshold. We observed similar results
using the ML100K dataset where the best (lowest) AC
Measure for user-based was 0.765 and for item-based was 0.746,
both at a similarity threshold of 0.0 and both using
meancentered prediction and signi cance weighting. These
results demonstrate that the \best" MAE may not always be
the lowest MAE, especially when coverage is also considered;
furthermore, recommender system settings such as similarity
weighting and neighborhood size also need to be considered
during system evaluation.</p>
      <p>
        Other observations of our experiments that match results
reported in [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ] and serve to validate our evaluation and
increase our con dence in the results are: (a) In general,
significance weighting improves prediction MAE, as compared to
predictions using Mahout similarity weighting or no
similarity weighting; (b) As the similarity threshold increases, MAE
for mean-centered prediction with signi cance weighting
improves and coverage degrades, whereas MAE and coverage
both degrade for Mahout prediction with Mahout weighting;
(c) Coverage decreases as neighborhood size decreases.
      </p>
    </sec>
    <sec id="sec-13">
      <title>CONCLUSION</title>
      <p>Our case study of Mahout as a recommender system
platform highlights evaluation considerations for developers and
also shows how straightforward functional enhancements
improves the performance of the baseline platform. We
evaluated our changes against current Mahout functionality
using accuracy and coverage metrics not only to assess
baseline results, but also to provide a view of the trade-o s
between accuracy and coverage resulting from using di erent
recommender algorithms. We reported cases where the
lowest MAE accuracy results were not necessarily always the
`best' when coverage results were also considered, and we
instrumented Mahout for a combined accuracy and
coverage metric (AC Measure) to evaluate these trade-o s more
directly. We believe that this case study will provide
useful guidance in using Mahout as a recommender platform,
and that our combined measure will prove useful in
evaluating algorithm changes for the inherent trade-o s between
accuracy and coverage.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>C.</given-names>
            <surname>Desrosiers</surname>
          </string-name>
          and
          <string-name>
            <given-names>G.</given-names>
            <surname>Karypis</surname>
          </string-name>
          .
          <article-title>A comprehensive survey of neighborhood-based recommendations methods</article-title>
          . In F. Ricci,
          <string-name>
            <given-names>L.</given-names>
            <surname>Rokach</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Shapira</surname>
          </string-name>
          , and P. B. Kantor, editors,
          <source>Recommender Systems Handbook</source>
          . Springer,
          <year>2011</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>M. D.</given-names>
            <surname>Ekstrand</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Ludwig</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. A.</given-names>
            <surname>Konnstan</surname>
          </string-name>
          , and
          <string-name>
            <given-names>J. T.</given-names>
            <surname>Riedl</surname>
          </string-name>
          .
          <article-title>Rethinking the recommender research ecosystem: Reproducibility, openness, and lenskit</article-title>
          .
          <source>In Proceedings of the 5th ACM Recommender Systems Conference (RecSys '11)</source>
          ,
          <year>October 2011</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>M.</given-names>
            <surname>Ge</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Delgado-Battenfeld</surname>
          </string-name>
          , and
          <string-name>
            <given-names>D.</given-names>
            <surname>Jannach</surname>
          </string-name>
          .
          <article-title>Beyond accuracy: Evaluating recommender systems by coverage and serendipity</article-title>
          .
          <source>In Proceedings of the 4th ACM Recommender Systems Conference (RecSys '10)</source>
          ,
          <year>September 2010</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>N.</given-names>
            <surname>Good</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. B.</given-names>
            <surname>Schafer</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. A.</given-names>
            <surname>Konstan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Borchers</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Sarwar</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Herlocker</surname>
          </string-name>
          , and
          <string-name>
            <surname>J. Riedl.</surname>
          </string-name>
          <article-title>Combining collaborative ltering with personal agents for better recommendations</article-title>
          .
          <source>In Proceedings of the 16th National Conference on Arti cial Intelligence (AAAI-99)</source>
          ,
          <year>July 1999</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>J. L.</given-names>
            <surname>Herlocker</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. A.</given-names>
            <surname>Konstan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Borchers</surname>
          </string-name>
          , and
          <string-name>
            <given-names>J.</given-names>
            <surname>Riedl</surname>
          </string-name>
          .
          <article-title>An algorithmic framework for performing collaborative ltering</article-title>
          .
          <source>In Proceedings of the ACM SIGIR Conference</source>
          ,
          <year>1999</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>J. L.</given-names>
            <surname>Herlocker</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. A.</given-names>
            <surname>Konstan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L. G.</given-names>
            <surname>Terveen</surname>
          </string-name>
          , and
          <string-name>
            <given-names>J.</given-names>
            <surname>Riedl</surname>
          </string-name>
          .
          <article-title>Evaluating collaborative ltering recommender systems</article-title>
          .
          <source>ACM Transactions on Information Systems</source>
          ,
          <volume>22</volume>
          (
          <issue>1</issue>
          ):5{
          <fpage>53</fpage>
          ,
          <year>2004</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>S.</given-names>
            <surname>Mcnee</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Riedl</surname>
          </string-name>
          , and
          <string-name>
            <given-names>J.</given-names>
            <surname>Konstan</surname>
          </string-name>
          .
          <article-title>Accurate is not always good: How accuracy metrics have hurt recommender systems</article-title>
          .
          <source>In Proceedings of the Conference on Human Factors in Computing Systems(CHI</source>
          <year>2006</year>
          ),
          <year>April 2006</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>P.</given-names>
            <surname>Resnick</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Iacovou</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Suchak</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Bergstrom</surname>
          </string-name>
          , and
          <string-name>
            <surname>J. Riedl.</surname>
          </string-name>
          <article-title>GroupLens: an open architecture for collaborative ltering of netnews</article-title>
          .
          <source>In Proceedings of the ACM CSCW Conference</source>
          ,
          <year>1994</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>B.</given-names>
            <surname>Sarwar</surname>
          </string-name>
          , G. Karypis,
          <string-name>
            <given-names>J.</given-names>
            <surname>Konstan</surname>
          </string-name>
          , and
          <string-name>
            <given-names>J.</given-names>
            <surname>Reidl</surname>
          </string-name>
          .
          <article-title>Item-based collaborative ltering recommendation algorithms</article-title>
          .
          <source>In Proceedings of the World Wide Web Conference</source>
          ,
          <year>2001</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>B. M.</given-names>
            <surname>Sarwar</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. A.</given-names>
            <surname>Konstan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Borchers</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Herlocker</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Miller</surname>
          </string-name>
          ,
          <string-name>
            <given-names>and J.</given-names>
            <surname>Riedl</surname>
          </string-name>
          .
          <article-title>Using ltering agents to improve prediction quality in the grouplens research collaborative ltering system</article-title>
          .
          <source>In Proceedings of the ACM 1998 Conference on Computer Supported Cooperative Work (CSCW '98)</source>
          ,
          <year>November 1998</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>C. E.</given-names>
            <surname>Seminario</surname>
          </string-name>
          and
          <string-name>
            <given-names>D. C.</given-names>
            <surname>Wilson</surname>
          </string-name>
          .
          <article-title>Robustness and accuracy tradeo s for recommender systems under attack</article-title>
          .
          <source>In Proceedings of the 25th Florida Arti cial Intelligence Research Society Conference (FLAIRS-25)</source>
          , May
          <year>2012</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>G.</given-names>
            <surname>Shani</surname>
          </string-name>
          and
          <string-name>
            <given-names>A.</given-names>
            <surname>Gunawardana</surname>
          </string-name>
          .
          <article-title>Evaluating recommendation systems</article-title>
          . In F. Ricci,
          <string-name>
            <given-names>L.</given-names>
            <surname>Rokach</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Shapira</surname>
          </string-name>
          , and P. B. Kantor, editors,
          <source>Recommender Systems Handbook</source>
          . Springer,
          <year>2011</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>