<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>30Music listening and playlists dataset</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Roberto Turrin</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Massimo Quadrana</string-name>
          <email>massimo.quadrana@polimi.it</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Roberto Pagano</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Paolo Cremonesi</string-name>
          <email>paolo.cremonesi@polimi.it</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>ContentWise R&amp;D</institution>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>DEIB</institution>
          ,
          <addr-line>Politecnico di Milano</addr-line>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2015</year>
      </pub-date>
      <abstract>
        <p>We introduce the 30Music dataset1, a collection of listening and playlists data retrieved from Internet radio stations through Last.fm API. In this paper we describe the creation process, its content, and its possible uses. Attractive features of the 30Music dataset that di erentiate it from existing public datasets include, among the others, (i) the user listening sessions complete of contextual time information, (ii) the user playlists, and (iii) the positive user ratings, key information to experiment with the task of modeling user taste and recommending playlists.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. INTRODUCTION</title>
      <p>Several challenges in the music domain have been only
partially explored due to the scarcity of data available to
researchers for experiments. For instance, tasks such as user
modeling and playlist recommendation require implicit
contextual information about listening events (e.g., user, track,
time, duration), explicit information about user preferences
(e.g., loved tracks, playlists), and user listening sessions.</p>
      <p>In this paper we introduce the 30Music dataset, a
freelyavailable music dataset designed to overcome these
problems. The main innovative aspects of the 30Music dataset
with respect to the existing public datasets are:
the dataset contains both implicit play events and
explicit user ratings (i.e., preferred tracks);
the dataset contains user-generated playlists;
play events are organized into listening sessions;
whenever a user plays a track from a playlist, the play
event is tagged.</p>
      <p>The rest of the paper is organized as follows. Section 2
discusses the existing music datasets. Section 3 presents the
process implemented to crawl the data from Last.fm and
create the dataset, whose main characteristics are explored
in Section 4. Finally, Section 5 draws the conclusions and
discusses future work.</p>
    </sec>
    <sec id="sec-2">
      <title>RELATED DATASETS</title>
      <p>There exist a number of publicly-available music datasets,
used in several music experiments. Most datasets provide
1http://recsys.deib.polimi.it/?page_id=54
content information (e.g., metadata, tags, acoustic features),
but only a few report some user-system interactions (e.g.,
ratings, play events) useful to pro le users and to experiment
with personalization tasks.</p>
      <p>
        The Million song dataset (MSD) [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ] is a public
collection well-known for its size. In fact, it contains audio
features (e.g., pitches, timbre, loudness, as provided by the
Echo Nest Analyze API2) and textual metadata (e.g.,
Musicbrainz3 tags, Echo Nest tags, Last.fm tags) about 1M
songs (related to 44K artists).
      </p>
      <p>
        Celma [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ] has published two music datasets collected from
Last.fm API: 1K-user and 360K-user. The smallest one -
1Kuser dataset - contains the user listening habits (20M play
events) of less than 1K users. On the other hand, the biggest
one - the 360K-user dataset - collects the information about
360K users, but it does not have any listening data other
than the number of times a user has listened to an artist.
Data are provided as downloaded from the Last.fm API.
      </p>
      <p>Yahoo! Labs have released several music datasets4. For
instance, the R1 and the R2 datasets provides ratings on
artists and songs, respectively, but not user play events.</p>
      <p>
        Some datasets have been extracted from microblogs, such
as the Million Musical Tweets Dataset [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]. Finally, The Art
of the Mix Playlist dataset5, consists of 29K user-contributed
playlists, containing 218K distinct songs for 60K distinct
artists. However, there are no user listening events.
3.
      </p>
    </sec>
    <sec id="sec-3">
      <title>DATASET CREATION</title>
      <p>The 30Music dataset has been obtained via Last.fm
public API6. Last.fm provides free API to track details of user
listening sessions. In the case a user has connected his
supported player to his Last.fm account, the player \scrobbles"
the user listening activity, i.e., it transfers the play event to
Last.fm that records such user interaction. It is worth
noting that only listening events are recorded, while pause/skip
events are not scrobbled from the user player to Last.fm,
as well as any playlist or explicit preference de ned or
expressed in the player. The main way for a user to create a
playlist in Last.fm is to access to the website; similarly, the
user can express explicit preferences (`love') about tracks
directly in the website. As a consequence, explicit ratings
and playlists stored in Last.fm are not biased by external
systems (e.g., the recommendations proposed in the player).
2http://the.echonest.com/
3https://musicbrainz.org/
4http://webscope.sandbox.yahoo.com/catalog.php
5http://labrosa.ee.columbia.edu/projects/musicsim/aotm.html
6http://www.last.fm/api
unfortunately, last.fm requires to scrobble only tracks played
for at least half their duration (or for at least 4 minutes), so
events not matching these conditions - such as skip events
are not con dent (although many events less than 5 seconds
have been found in the collected data).</p>
      <p>To build the 30Music dataset, we started from a list of 2M
Last.fm usernames from the Chris Meller dataset 7. Given
the list of users, we retrieved their playlists
(User.getPlaylists) together with the single tracks composing the playlists
(Playlist.fetch). Starting from users with at least one
playlist (about 90K users), we retrieved
(User.getRecentTracks) the user listening events over a 1-year time
window (from Mon, 20 Jan 2014 09:24:19). The raw playlists
and user listening events have been enriched with additional
information both about users (User.getInfo) and tracks
(Track.getInfo).</p>
      <p>Furthermore, the data downloaded with the Last.fm API
has been processed using Python scripts exploiting some
Apache Spark functions for a distributed processing of the
massive amount of data. In order to keep only complete and
reliable data, we discarded users with some missing data
(e.g., if the track scrobbled by the user has the wrong
metadata and it is not recognized by Last.fm, the whole user is
discarded). In this way, we maintained only the half of the
users that have complete information.</p>
      <p>Finally, we de ned a new entity, the user play session,
as an ordered list of play events that are assumed to be
consequently listened by the user with no interruptions. We
de ne a play event to be part of a session if it occurs no later
than 800 seconds after the previous user play event. This
processing required, for each user listening event, to compute
the play time, together with the ratio of track e ectively
listened by the user.
30Music format.</p>
      <p>The 30Music dataset is released in accordance with the
[anonimized for double blind review] data format, a
multigraph representation oriented to recommender system
evaluation that explicitly represents entities (i.e., nodes) and
relations (i.e., edges).</p>
      <p>Entity model any object that can be recommended. The
dataset is formed by 45K users, 5.6M tracks, 50K playlists,
600K artists, 200K albums, and 280K tags. Relations model
links between two (or more) entities. We de ned 31M user
play events, 2.7M user play sessions, and 4.1M user love
preferences.</p>
    </sec>
    <sec id="sec-4">
      <title>DATASET ANALYSIS</title>
      <p>The dataset contains 31,351,954 play events organized into
2,764,474 sessions (an average of 11 play events per session).
The dataset contains also 4,106,341 explicit ratings (loved
tracks), with an average of 33 ratings per user, and 57,561
user-created playlists. The number of events without track
duration is 1,277,893 (4.08%).</p>
      <p>We can observe that play events present a moderate
longtail distribution: the 20% most popular tracks collect 80%
of the play events. This long tail e ect is mitigated by
focusing on preferred tracks (i.e., loved tracks and tracks in
the playlists). We observe that the same percentage of play
events (80%) involves twice the tracks (40%) when
considering tracks in the playlists. We can deduce that users have
preferences spanning many di erent tracks, but their
listening behaviour is biased toward the most popular tracks.</p>
      <p>A similar analysis has been performed by aggregating the
tracks of the same artist. Di erently from tracks, these play
events present a strong long-tail distribution: the 20% most
popular artists collect more than 95% of the play events.
This long tail e ect is strongly mitigated when analyzing
preferred tracks. The same percentage of play events (95%)
involve 50% of the artists when considering tracks in the
playlists. We can deduce that users have preferences
spanning many di erent artists, but their listening behaviour is
strongly biased toward the most popular artist.</p>
      <p>An analysis of the empirical cumulative explicit like
distribution as a function of the (percentage) number of tracks
and artists shows that only the 14.73% of the tracks and
the 19.93% of the artists have received at least one explicit
preference. We observed that the distribution of the explicit
ratings within tracks does not exhibit a strong long-tail
behavior. The 5% of the most popular tracks collect the 75%
of the explicit ratings. On the other hand, the number of
explicit ratings is strongly skewed toward a few very
popular artists. The 5% of the of most popular artists collect
more than the 90% of the explicit ratings. These results
conrm our previous intuitions over users' listening behaviour.
Users tend to love (and to listen to) a few very popular
artists. However, their preference spans across several tracks
of these very popular artists. Still, they tend to provide
explicit rating for few of the tracks they have listened to. This
can be due to the mechanism adopted by Last.fm to
collect explicit feedback, which forces users to move from their
usual music player, to access to the Last.fm online service
and to provide their \love" to a track there. This clearly
imposes a heavy burden over users, but on the other hand
it enhances the strength of each explicit rating, because it
is a clear expression of the willingness of the speci c user to
provide that feedback.
5.</p>
    </sec>
    <sec id="sec-5">
      <title>CONCLUSIONS AND FUTURE WORK</title>
      <p>In this paper, we presented the 30Music dataset, a music
dataset consisting of both user interactions (i.e., user play
sessions) and user explicit preferences (i.e., playlists,
preferred tracks). The dataset is made available to the research
community and we expect it will foster the exploration of the
several challenges still open in the settings of online music
applications.</p>
      <p>Acknoledgements.</p>
      <p>The research leading to these results was performed in the
CrowdRec project, which has received funding from the
European Union Seventh Framework Programme
FP7/20072013 under grant agreement n. 610594.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>T.</given-names>
            <surname>Bertin-Mahieux</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D. P.</given-names>
            <surname>Ellis</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Whitman</surname>
          </string-name>
          , and
          <string-name>
            <given-names>P.</given-names>
            <surname>Lamere</surname>
          </string-name>
          .
          <article-title>The million song dataset</article-title>
          .
          <source>12th Int. Conf. on Music Information Retrieval (ISMIR)</source>
          ,
          <year>2011</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>O.</given-names>
            <surname>Celma</surname>
          </string-name>
          .
          <article-title>Music Recommendation and Discovery in the Long Tail</article-title>
          . Springer,
          <year>2010</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>D.</given-names>
            <surname>Hauger</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Schedl</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Kosir</surname>
          </string-name>
          , and
          <string-name>
            <given-names>M.</given-names>
            <surname>Tkalcic</surname>
          </string-name>
          .
          <article-title>The million musical tweet dataset - what we can learn from microblogs</article-title>
          .
          <source>In Proc. of the 14th Int. Society for Music Information Retrieval Conference, Nov 4-8</source>
          <year>2013</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>