<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Data Pruning in Recommender Systems Research: Best-Practice or Malpractice?</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Joeran Beel∗</string-name>
          <email>beelj@tcd.ie</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Victor Brunel</string-name>
          <email>victor.brunel@etu.uca.fr</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Polytech Clermont-Ferrand, Department of, Mathematical Engineering and Modeling</institution>
          ,
          <addr-line>Clermont-Ferrand</addr-line>
          ,
          <country country="FR">France</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Trinity College Dublin, School of Computer, Science &amp; Statistics, ADAPT Centre</institution>
          ,
          <addr-line>Dublin</addr-line>
          ,
          <country country="IE">Ireland</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2019</year>
      </pub-date>
      <abstract>
        <p>Figure 1: Distribution of users and ratings in the MovieLens dataset. 16% of users rated less than 10 movies, 26% rated between 10 and 19 movies. In most MovieLens releases (100k, 1m, 10m, ...), these 42% of users and their ratings are not included. 3% of users have 500 or more ratings and contribute 28% of all ratings in MovieLens.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>ACM RecSys 2019 Late-breaking Results, 16th-20th September 2019, Copenhagen, Denmark
Copyright ©2019 for this paper by its authors. Use permited under Creative Commons License Atribution 4.0 International
(CC BY 4.0).</p>
      <p>INTRODUCTION
’Data pruning’ is common practice in recommender-systems research. We define data pruning as
the removal of instances from a dataset that would not be removed in the real-world, i.e. when used
by recommender systems in production environments. Reasons to prune datasets are manifold and
include user interests when publishing data (e.g. data privacy) or business interests. ’Data pruning’
difers from ’data cleaning’ as such that data cleaning (e.g. outlier removal) is typically a prerequisite
for the efective training of recommender-system and machine-learning algorithms, whereas data
pruning is not afecting the algorithm performance in itself.</p>
      <p>
        A prominent example is MovieLens in most of its variations1 [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]. MovieLens contains information
about how users of MovieLens.org rated movies, but the MovieLens team decided to exclude ratings of
users who rated less than 20 movies.1 The reasoning was as follows:2 "(1) [researchers] needed enough
ratings to evaluate algorithms, since most studies needed a mix of training and test data, and it is always
possible to use a subset of data when you want to study low-rating cases; and (2) the movies receiving the
first ratings for users during most of MovieLens’ history are biased based on whatever algorithm was in
place for new-user startup (for most of the site’s life, that was a mix of popularity and entropy), hence the
MovieLens team didn’t want to include users who hadn’t goten past the ’start-up’ stage [...]. "
      </p>
      <p>
        Not only the creators of datasets may prune data, but also individual researchers may do so. For
instance, Caragea et al. pruned the CiteSeer corpus for their research [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]. The corpus contains a large
number of research articles and their citations. Caragea et al. removed research papers with fewer
than ten and more than 100 citations as well as papers citing fewer than 15 and more than 50 research
papers. From originally 1.3 million papers in the corpus around 16,000 remained (1.2%). Similarly,
Pennock et al. removed many documents so that only 0.58% remained for their research [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ].
      </p>
      <p>
        We criticized the practice of data pruning previously, particularly when only a fraction of the
original data remains [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. We argued that evaluations based on a small fraction of the original data
are of litle significance. For instance, knowing that an algorithm performs well for 0.58% of users
is of litle significance if it remains unknown how the algorithm performs for the remaining 99.42%.
Also, it is well known that collaborative filtering tends to perform poorly for users with few ratings
[
        <xref ref-type="bibr" rid="ref4">4</xref>
        ]. Hence, when evaluating collaborative filtering algorithms, we would consider it crucial to not
ignore users with few ratings, i.e. those users for whom the algorithms presumably perform poorly.
      </p>
      <p>Our criticism was based more on ’gut feeling’ than scientific evidence. To the best of our knowledge,
no empirical data exists on how widely data pruning is applied, and how pruning afects
recommendersystems evaluations. Also, the recsys community has not discussed if, and to what extent a) datasets
should be pruned by their creators and b) whether individual researchers should prune data. We
conduct the first steps towards answering these questions. Ultimately, we hope to stimulate a discussion
that leads to widely accepted guidelines on pruning data for recommender-systems research.</p>
      <p>.
3We might have missed some relevant
information on pruning if that information was
provided in a section other than the Methodology
section.</p>
    </sec>
    <sec id="sec-2">
      <title>4For users with only 1 rating, ’Surprise’ uses a</title>
      <p>special technique for evaluations. Please also
note that using a random sample is not an
example of data pruning.</p>
      <sec id="sec-2-1">
        <title>METHODOLOGY</title>
        <p>To identify how widespread data pruning is, we analyzed all 112 full- and short papers published at
the ACM Conference on Recommender Systems 2017 and 2018. 88 papers (79%) used ofline datasets,
the remaining 24 papers (21%) conducted e.g. user studies or online evaluations. For the 88 papers, we
analyzed, which datasets the authors used, whether the datasets were pruned by the original creators
and whether authors conducted pruning. To identify the later part, we read the Methodology sections
of the manuscripts or similarly named sections.3 This analysis was done by a single person rather
quickly. Consequently, the reported numbers should be seen as ballpark figures.</p>
        <p>To identify the efect of data pruning on recommender-system evaluations, we run six collaborative
filtering algorithms from the Surprise library, namely SVD, SVD++, NMF, Slope One, Co-Clustering,
and the Baseline Estimator. We use the unpruned ’MovieLens Latest Full’ dataset (Sept. 2018), which
contains 27 million ratings by 280,000 users including data from users with less than 20 ratings. Due
to computing constraints, we use a random sample with 6,962,757 ratings made by 70,807 users.4</p>
        <p>We run the six algorithms on three sub-sets of the dataset, i.e. for a) the entire unpruned dataset,
b) the data that would be included in a ’normal’ version of MovieLens (users with 20+ ratings) c)
the data that would be ’normally’ not included in the MovieLens dataset (users with less than 20
ratings). We compare how algorithms perform on these diferent sets, and measure the performance of
algorithms by Mean Absolute Error (MAE) and Root Mean Square Error (RMSE). As the two metrics
led to almost identical results, we only report RMSE. Our source code and analysis of the manuscripts
is available at htps://github.com/BeelGroup/recsys-dataset-pruning.</p>
      </sec>
      <sec id="sec-2-2">
        <title>RESULTS</title>
      </sec>
      <sec id="sec-2-3">
        <title>Popularity of (Pruned) Recommender-Systems Datasets</title>
        <p>The authors of the 88 papers used a total of 64 unique datasets, whereas we counted diferent variations
of MovieLens, Amazon and Yahoo! as the same dataset. Our analysis empirically confirms what is
common wisdom in the recommender-system community already: MovieLens is the de-facto standard
dataset in recommender-systems research. 40% of the full- and short papers at RecSys 2017 and
2018 used the MovieLens dataset in at least one of its variations (Figure 3). The second most popular
dataset is Amazon, which was used by 35% of all authors. Other popular datasets are shown in Figure
3 and include Yelp (13%), Tripadvisor (8%), Yahoo! (6%), BookCrossing (5%), Epinions (5%), and LastFM
(5%). 11% of all researchers used a proprietary dataset, and 2% used a synthetic dataset.</p>
        <p>50% of the authors conducted research with a single dataset, 31% used two datasets, and only 2%
used six or more datasets (Figure 4). The highest number of datasets being used was 7. On average,
researchers used 1.88 datasets. 40% of the authors used a pruned dataset, and 15% pruned data
themselves. In total, 48% of all authors conducted research at least partially with pruned data.
5The actual numbers in the pruned MovieLens
versions may somewhat difer given that we
just used a sample, and the diferent MovieLens
versions difer due to the fact that they include
data from diferent time periods.</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>6It is actually not trivial to decide how to split</title>
      <p>the data in this case. Follow-up research is
needed to confirm the numbers, and
investigate diferent options.</p>
    </sec>
    <sec id="sec-4">
      <title>7The authors evaluated diferent algorithms.</title>
      <p>For users with few ratings (&lt;8), other
algorithms performed best than for users with more
ratings.</p>
      <sec id="sec-4-1">
        <title>The Efect of Data Pruning</title>
        <p>The user-rating distribution in MovieLens follows a long-tail distribution (Figure 1). 42% of the users
have less than 20 ratings, and these users contribute 5% of all ratings. The remaining 58% of users, with
20+ ratings, contribute 95% of all ratings. The top 3% of users – those with 500+ ratings – contribute
28% of all ratings in the dataset. Consequently, using a pruned MovieLens variation (100k, 1m, ...) is
equivalent to ignoring around 5% of the ratings and 42% of users.5</p>
        <p>There are notable diferences for the three data splits in terms of algorithm performance (Figure 2).
Over the entire unpruned data, RMSE of the six algorithms is 0.86 on average, with the best algorithm
being SVD++ (0.80), closely followed by SVD (0.81). The worst performing algorithm is Co-Clustering
(0.90). For the subset of ratings from users with 20+ ratings – that equals a ’normal’ MovieLens dataset
– RMSE over all algorithms is 0.84 on average (2.12% lower, i.e. beter). In other words, using a pruned
version of MovieLens will lead, on average, to a 2.12% beter RMSE compared to using the unpruned
data. But, to make this clear, the algorithms do not actually perform 2.12% beter. The results only
appear to be beter because data for which the algorithms tend to perform poorly was excluded in the
evaluation. The ranking of the algorithms remains the same when comparing the pruned with the
unpruned data (SVD++ performs best, followed by SVD, and Co-Clustering performs worst).</p>
        <p>We also looked at the users grouped by the number of ratings per user6. Figure 5 shows the RMSE
for users with 1-9 ratings, 10-19 ratings, ... 500+ ratings. There is a constant improvement (i.e. decrease)
in RMSE the more ratings user have. On average, the six algorithms achieve an RMSE of 1.03 for
users with less than 20 ratings (1.07 for users with &lt;=9 ratings; 1.02 for users with 10–19 ratings). This
contrasts an average RMSE of 0.84 for users with 20+ ratings. In other words, RMSE for users in a
pruned MovieLens dataset is 23% beter than RMSE for the excluded users. For SVD and SVD++, the
best performing algorithms, this efect is even stronger (+27% for SVD; +25% for SVD++).</p>
      </sec>
      <sec id="sec-4-2">
        <title>DISCUSSION &amp; FUTURE WORK</title>
        <p>Data Pruning is a widespread phenomenon with 48% of short- and full papers at RecSys 2017/2018
being based at least partially on pruned data. MovieLens nicely illustrates an issue that probably
applies to many datasets with user-rating data. In the pruned MovieLens datasets, the number of
removed ratings is rather small (5%). However, these 5% ratings were made by 42% of the users. For
researchers focusing on how well individual ratings can be predicted, the removed data has probably
litle impact. For researchers who focus on user-specific issues, ignoring 42% of users is probably not
ideal, particularly as their RMSE is 23% worse than the RMSE of the users with 20+ ratings.</p>
        <p>
          When discussing data pruning, the probably most important question is whether pruning changes
the ranking of the evaluated algorithms. Other research has already shown that the ranking may
change, though that research was not conducted in the context of data pruning [
          <xref ref-type="bibr" rid="ref4">4</xref>
          ].7 In our study,
the ranking of algorithms did not change. The algorithm that was best (second best...) on the pruned
MovieLens data was also best (second best...) on the unpruned data. However, it seems likely to us
that rankings may change if more diverse algorithms are compared, e.g. collaborative filtering vs.
content-based filtering. Also, the MovieLens dataset is relatively moderately pruned. We consider it
likely that heavy pruning, where only a small fraction remains (e.g. Pennock et al. [
          <xref ref-type="bibr" rid="ref5">5</xref>
          ]), might lead to a
change in the ranking of algorithms. More research with more diverse algorithms, diferent datasets,
and diferent degrees of pruning is needed to confirm or reject that assumption. A qualitative study
could help to identify details on the motivation of dataset creators and researchers to prune data.
        </p>
        <p>Given the current results, we propose that data pruning should be applied with great care, or, if
possible, be avoided. We would not generally consider pruning as a malpractice, but certainly not a
best-practice either. In some cases, especially when large parts of data are removed, data pruning
may become a malpractice, though the community yet has to determine how much removed data is
too much. As a starting point, we would recommend the following guidelines, though this is certainly
not a definite recommendation, and more discussion in the community is needed:
(1) Publishers of datasets should avoid pruning – if possible. If there are compelling reasons to
prune data (e.g. ensuring privacy), these should be clearly communicated in the documentation.
(2) Researchers using pruned datasets should discuss the implications in their manuscript.
(3) Individual researchers should not prune data. If they feel that an algorithm may perform well on
a subset of the data, they should report performance for both the entire dataset and the subset.
If researchers conduct pruning anyway, they should clearly indicate this in the manuscript,
provide reasoning, and discuss the implications.</p>
        <p>It may not always be obvious where data cleaning ends, legitimate data pruning begins, and when
data pruning becomes a malpractice. The community certainly needs more discussion about this issue.
We are confident that with widely agreed guidelines on data pruning, recommender-systems research
will become more reproducible, more comparable and more representative of ’the real world’.</p>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>Joeran</given-names>
            <surname>Beel</surname>
          </string-name>
          , Bela Gipp, Stefan Langer, and
          <string-name>
            <given-names>Corinna</given-names>
            <surname>Breitinger</surname>
          </string-name>
          .
          <year>2016</year>
          .
          <article-title>Research Paper Recommender Systems: A Literature Survey</article-title>
          .
          <source>International Journal on Digital Libraries</source>
          <volume>4</volume>
          (
          <year>2016</year>
          ),
          <fpage>305</fpage>
          -
          <lpage>338</lpage>
          . htps://doi.org/10.1007/s00799-015-0156-0
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>Cornelia</given-names>
            <surname>Caragea</surname>
          </string-name>
          , Adrian Silvescu, Prasenjit Mitra, and
          <string-name>
            <given-names>C Lee</given-names>
            <surname>Giles</surname>
          </string-name>
          .
          <year>2013</year>
          .
          <article-title>Can't See the Forest for the Trees? A Citation Recommendation System</article-title>
          . In iConference.
          <fpage>849</fpage>
          -
          <lpage>851</lpage>
          . htps://doi.org/10.9776/13434
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>F Maxwell</given-names>
            <surname>Harper and Joseph A Konstan</surname>
          </string-name>
          .
          <year>2016</year>
          .
          <article-title>The movielens datasets: History and context</article-title>
          .
          <source>ACM Transactions on Interactive Intelligent Systems (TiiS) 5</source>
          ,
          <issue>4</issue>
          (
          <year>2016</year>
          ),
          <fpage>19</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>Daniel</given-names>
            <surname>Kluver</surname>
          </string-name>
          and Joseph A Konstan.
          <year>2014</year>
          .
          <article-title>Evaluating recommender behavior for new users</article-title>
          .
          <source>In Proceedings of the 8th ACM Conference on Recommender Systems. ACM</source>
          ,
          <volume>121</volume>
          -
          <fpage>128</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <surname>David</surname>
            <given-names>M Pennock</given-names>
          </string-name>
          ,
          <string-name>
            <given-names>Eric</given-names>
            <surname>Horvitz</surname>
          </string-name>
          , Steve Lawrence, and
          <string-name>
            <given-names>C Lee</given-names>
            <surname>Giles</surname>
          </string-name>
          .
          <year>2000</year>
          .
          <article-title>Collaborative filtering by personality diagnosis: A hybrid memory-and model-based approach</article-title>
          .
          <source>In Sixteenth Conference on Uncertainty in Artificial Intelligence</source>
          .
          <fpage>473</fpage>
          -
          <lpage>480</lpage>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>