<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Myusic: a Content-based Music Recommender System based on eVSM and Social Media</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Cataldo Musto</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Fedelucio Narducci</string-name>
          <email>narducci@disco.unimib.it</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Giovanni Semeraro</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Pasquale Lops</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Marco de Gemmis</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Department of Computer Science University of Bari Aldo Moro</institution>
          ,
          <country country="IT">Italy</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Department of Information Science</institution>
          ,
          <addr-line>Systems Theory</addr-line>
          ,
          <institution>and Communication University of Milano-Bicocca</institution>
          ,
          <country country="IT">Italy</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>This paper presents Myusic, a platform that leverages social media to produce content-based music recommendations. The design of the platform is based on the insight that user preferences in music can be extracted by mining Facebook pro les, thus providing a novel and effective way to sift in large music databases and overcome the cold-start problem as well. The content-based recommendation model implemented in Myusic is eVSM [4], an enhanced version of the vector space model based on distributional models, Random Indexing and Quantum Negation. The e ectiveness of the platform is evaluated through a preliminary user study performed on a sample of 50 persons. The results showed that 74% of users actually prefer recommendations computed by social mediabased pro les with respect to those computed by a simple heuristic based on the popularity of artists, and con rmed the usefulness of performing user studies because of the di erent outcomes they can provide with respect to o ine experiments.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>
        One of the main issues of the so-called personalization pipeline is preference
acquisition and elicitation. That step has always been considered the bottleneck
in recommendation process since classical approaches for gathering user
preferences are usually time consuming or intrusive. The widespread di usion of social
networks in the age of Web 2.0 o ers a new interesting chance to overcome that
problem, since users spend 22% of their time on social networks3 and 30
billion pieces of content are shared on Facebook every month [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]. In this scenario,
to harvest social media is a recent trend in the area of Recommender Systems
(RSs): it can merge the un-intrusiveness of implicit user modeling with the
accuracy of explicit techniques, since the information left by users is freely provided
and actually re ects real preferences.
      </p>
    </sec>
    <sec id="sec-2">
      <title>3 http://blog.nielsen.com/nielsenwire/social/</title>
      <p>
        This paper presents Myusic, a tool that provides users with music
recommendations. The goal of the system is to catch user preferences in music and
lter the huge amount of data stored in platforms such as iTunes or Amazon
in order to produce personalized suggestions about artists users could like. The
ltering model behind Myusic is eVSM, an enhanced extension of VSM based on
distributional models, Random Indexing and Quantum Negation. As introduced
in [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ], eVSM provides a lightweight semantic representation based on
distributional models, where each artist (and the user pro le, as well) is modeled as a
vector in a semantic vector space, according to the tags used to describe her
and the co-occurrences between the tags themselves. The model is based on the
assumption that a user pro le can be built by combining the tag-based
representation (obtained by crawling Last.fm platform) of the artists she is interested
in. Next, classical similarity measures can be exploited to match item
descriptions with content-based user pro les. A prototype version of Myusic was made
available online for two months in order to design a user study and evaluate the
e ectiveness of the model as well as its impact on real users.
      </p>
      <p>
        Generally speaking, this work concerns to the area of music recommendation.
The commonly used technique for providing recommendations is collaborative
ltering, implemented in very well known services, such as MyStrands4, Last.fm5
or iTunes Genius. An early attempt to recommend music using collaborative
ltering was done by Shardanand [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ]. Another trend is to use content-based
recommendation strategies, which analyze diverse sets of low-level features (e.g.
harmony, rhythm, melody), or high-level features (metadata or content-based
data available in social media) [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ] to provide recommendations. The use of Linked
Data for music recommendation is investigated in [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ]. Recently, Bu et al. [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]
followed the recent trend of harvesting information coming from social media for
personalization tasks and proposed its application for music recommendation.
Finally, Wang et al. [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ] showed the usefulness of tags with respect to other
content-based sources.
      </p>
      <p>The paper is organized as follows: the architecture of the systems is sketched
in Section 2; Section 3 focuses on the results of a preliminary experimental
evaluation and nally Section 4 contains conclusions and directions for future
research.
2</p>
      <sec id="sec-2-1">
        <title>Myusic: content-based music recommendations</title>
        <p>The general architecture of Myusic is sketched in Figure 1. We can identify four
main components:</p>
        <p>Crawler. The Crawler module queries Last.fm through its public APIs to
build a corpus of available artists. For each artist, the name, a picture, the title
of the most popular tracks, their playcount and a set of tags that describe that
artist are crawled. All the crawled data are locally stored.</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>4 http://www.mystrands.com 5 http://www.last.fm</title>
      <p>Extractor. The Extractor module connects to Facebook, extracts artists
the user likes (Favourite Music section in the Facebook pro le, see Figure 2), and
maps them to the data gathered from Last.fm in order to build a preliminary set
of artists the user likes. This information is locally modeled in her own pro le to
let her receive recommendations even in her rst interaction with Myusic, thus
avoiding the cold-start. Implicit information coming from the links posted by
the user and the events she attended are extracted, as well.</p>
      <p>
        Pro ler. The process of building user pro les is performed in two steps.
First, a weight is assigned to each artist returned by the Extractor. The
weight of a speci c artist is de ned according to a simple heuristic: if a user
posted a song, that information can be considered as a light evidence of her
preference for that artist, while the fact that she explicitly clicked on "Like"
on her Facebook page can be considered as a strong evidence. For example,
on a 5-point Likert scale, a score equal to 3 is assigned to the artists whose
name appear among the links posted by the user, while a score equal to 4 is
assigned to those occurring in her favorite Facebook pages. If an artist occurs
in both lists (that is to say, the user likes it and posted a song, as well), 5 out
of 5 is assigned as score. Next, a pro ling model has to be chosen. The eVSM
framework provides four di erent pro ling models [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ]: a basic pro le (referred to
as RI ), a simple variant that exploits negative user feedbacks (called QN ), and
two weighted counterparts which give greater weight to the artists a user liked
the most (respectively, W-RI and W-QN). Regardless the pro ling model, in
eVSM user pro les are de ned in eVSM by means of two vectors, p+u and p u,
which represent user preferences and negative feedbacks, respectively. They are
de ned as follows:
p+u =
jIu+j
X ai r(u; ai)
i=1
p u =
jIu j
X ai (M AX
i=1
r(u; ai))
(1)
(2)
where Iu+ is the set of user favorite artists, Iu is the set of artists the user
dislikes, MAX is the highest rating that can be assigned to an item, r(u; ai) is
the score assigned to the artist ai and ai is the vector space representation of the
artist. Since each artist is described through a set of tags t1 : : : tn extracted from
Last.fm, the vector space representation is a weighted vector ai = (wt1 ; : : : ; wtn )
where wti is the weight of the tag ti. Generally speaking, W-QN model combines
p+u with p u through a Quantum Negation operator implemented in eVSM
framework, while W-RI model exploits only the information coming from p+u
and does not take into account negative feedback. Finally, RI and QN follow the
same insight of their weighted counterpart with the di erence that they do not
exploit the user rating r(u; ai), thus a uniform weight is given to each artist.
      </p>
      <p>Recommender. Given a semantic vector space representation based on
distributional models for both artists and user pro les, through similarity measures
it is possibile to produce as output a ranked list of suggested artists. The cosine
similarity for all the possible couples (pu; a) is computed, where pu is the vector
space representation of user u, while a is the vector describing the artist a.
Figure 3 shows an example of recommendation list. The platform allows the user
to express feedbacks on recommendations. Positive and negative feedbacks are
used to respectively update positive and negative pro le vectors and to trigger
the recommendation process again.
3</p>
      <sec id="sec-3-1">
        <title>Experimental Evaluation</title>
        <p>The goal of the experimental evaluation is to validate the design of the platform
by carrying out a user study whose goal is to analyze the impact and the e
ectiveness of the di erent con gurations of eVSM implemented in Myusic. Speci cally,
a user study involving 50 users under 30, heterogeneously distributed by sex,
education and musical knowledge (according to the availability sampling strategy)
has been performed. They interacted for two months with the online version of
Myusic. A crawl of Last.fm was performed at the end of November, 2011 and
data about 228,878 artists were extracted. Each user explicitly granted the
access to her Facebook pro le to extract data about favourite artists. At the end of
the Extraction step, a set of 980 di erent artists the 50 users like were extracted
from Facebook pages. Generally speaking, 1,720 feedbacks were collected: 1,495
of them came from Facebook pro les, while 225 were explicitly provided by the
users (for example, expressing a feedback on their recommendations). The
collected feedbacks were highly unbalanced since only 116 (6.71%) on 1,720 were
negative. Last.fm APIs were exploited to extract the most popular tags
associated to each artist. The less expressive and meaningful ones (such as seenlive,
cool, and so on) were considered as noisy and ltered out. The design of the user
study was oriented to answer to the following questions:
{ Experiment 1: Does the cold-start problem can be mitigated by modeling
user pro les which integrate information coming from social media?
{ Experiment 2: Do the users actually perceive the utility of adopting
weighting schemes and negation when user pro les are represented?
{ Experiment 3: How does the platform perform in terms of novelty,
serendipity and diversity of the proposed recommendations?</p>
        <p>
          In the rst experiment, users were asked to login and to extract their data
from their own Facebook page. Next, a user pro le was built according to a
proling model randomly chosen among the 4 described above and a preliminary
set of recommendations was proposed to the target user. In order to evaluate
the e ectiveness of the Extractor we compared the recommendation list
generated through eVSM to a baseline represented by a list produced by simply
ranking the most popular artists. Next, we asked users to tell which list they
preferred. Obviously, they were not aware about which list was the baseline and
which one was built through eVSM. A plot that summarizes users' answers is
provided in Figure 4-a. It is straightforward to note that users actually prefer
social media-based recommendations, since 74% of them preferred that
strategy with respect to a simple heuristic based on popularity of the artists stored
in database. However, even if the results gained by this pro ling technique were
outstanding, it is necessary to understand why 26% of the users simply preferred
the most popular artists. Probably, there is a correlation between users'
knowledge in music and the list they choose. It is likely that users with very generic
tastes prefer a list of popular singers. Similarly, it is likely that users with a
poor knowledge in music might prefer a list of well-known singers with respect
to a list where most of the artists, even if related to their tastes, were unknown.
A larger evaluation with users, split according to their musical knowledge, may
be helpful to understand the dynamics behind users' choices. Similarly, it would
be good to investigate the impact of the amount of the information extracted
from Facebook pro les with the accuracy of the recommendations. The second
experiment was performed in two steps. In the rst step users were asked to login
and to extract their data from their own Facebook page, as in Experiment 1.
Next, two pro les were built by following the RI and the W-RI pro ling models,
respectively. Finally, recommendations were generated from both pro les, and
users were asked to choose the con guration they preferred. As in Experiment 1,
they were not aware about which recommendations were generated by
exploiting their weighted pro le and which ones were produced through its unweighted
counterpart. Results of this experiments are shown in Figure 4-b. Di erently
from the results obtained from an in-vitro experiment performed in a scenario
of movie recommendation [
          <xref ref-type="bibr" rid="ref5">5</xref>
          ], users did not perceive as useful the introduction
of a weighting scheme designed to give higher signi cance to the artists the user
likes the most. On the contrary, the RI pro ling model was the preferred one
for 70% of the users involved in the experiment. Similarly, in the second step of
the experiment the RI pro ling model was compared to the QN one, in order to
evaluate the impact on user perception of modeling negative preferences. Also
in this case the results were con icting with the outcomes that emerged from
the in-vitro experiment since 65% of the users preferred the recommendations
generated through the pro ling technique that does not model negative
preferences. Even if the results of Experiment 2 did not con rmed the outcomes
of the o ine evaluation of eVSM they are actually interesting. First, they
conrmed the usefulness of combining o ine experiments with user studies thanks
to the di erent outcomes they can provide. Indeed, in user-centered applications
such as content-based recommender systems, user perception and user feedbacks
play a central role and these factors need to be taken into account. In general,
further investigation is needed because most of these results may be due to a
speci c bias of the designed experiment. As stated above, the extraction of data
from Facebook pages crawls information about what a speci c user likes, so very
few negative feedback were collected (less than 7%). Consequently, the negative
part of the user pro le was very poor and this might justify the results. It is
likely that collecting more negative feedbacks would be enough to con rm the
usefulness of negative information. Finally, in Experiment 3 users were asked
to express their preference on the recommendations produced through the RI
pro ling model (since it emerged as the best one from the previous experiment)
in terms of novelty, accuracy and diversity. The results of this experiment are
sketched in Figure 4-c. In general, the results are encouraging since most of the
users expressed a positive opinion about the system. Speci cally, Myusic has a
positive impact on nal users in terms of trust, since the opinion of 92% of the
users ranges from Good to Very Good. This is likely due to the good accuracy
of the recommendations produced by the system. Indeed, more than 80% of
the users considered as accurate or very accurate the suggestions of the system.
Similarly, also the outcomes concerning diversity were positive, since more than
60% labeled the level of diversity among the recommendations as Very Good.
The only aspect that needs improvements regards the novelty of
recommendations since 34% of the users labeled as not novel the suggestions produced by
the system. This outcome was somehow expected since overspecialization it is a
typical problem of content-based recommender systems (CBRS). However, even
if these results lead us to carry on this research, they have to be considered as
preliminary since this evaluation needs to be extended by comparing results of
eVSM with other state of the art models, such as LSI, VSM or collaborative
ltering.
4
        </p>
      </sec>
      <sec id="sec-3-2">
        <title>Conclusions and Future Directions</title>
        <p>
          In this paper we proposed Myusic, a music recommendation platform. It
implements a content-based recommender system based on eVSM, an enhanced
version of classical VSM. The most distinguishing aspect of Myusic is the
exploitation of Facebook pro les for acquiring user preferences. An experimental
evaluation carried out by involving real users demonstrated that leveraging
social media is an e ective way for overcoming the cold-start problem of CBRS. On
the other hand, the exploitation of relevance feedback and user ratings generally
did not improve the predictive accuracy of Myusic. Users showed to trust the
system, and Myusic also achieved good results in terms of accuracy and diversity
of recommendations. Those results encouraged keeping on this research. In the
future we will investigate the adoption of recommendation strategies tailored on
the music background of each user, even by learning accurate interaction models
in order to classify users [
          <xref ref-type="bibr" rid="ref7">7</xref>
          ]. Furthermore, we will try to introduce more
unexpected suggestions. Experiments showed that novelty needs to be improved.
        </p>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <given-names>S.</given-names>
            <surname>Bu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            and
            <surname>Tan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Chen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Wu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Zhang</surname>
          </string-name>
          , and
          <string-name>
            <given-names>X.</given-names>
            <surname>He</surname>
          </string-name>
          .
          <article-title>Music recommendation by uni ed hypergraph: combining social media information and music content</article-title>
          .
          <source>In Proceedings of the international conference on Multimedia, MM '10</source>
          , pages
          <fpage>391</fpage>
          {
          <fpage>400</fpage>
          , New York, NY, USA,
          <year>2010</year>
          . ACM.
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <given-names>C.</given-names>
            <surname>Hahn</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Turlier</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Liebig</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Gebhardt</surname>
          </string-name>
          , and
          <string-name>
            <given-names>C.</given-names>
            <surname>Roelle</surname>
          </string-name>
          .
          <article-title>Metadata Aggregation for Personalized Music Playlists</article-title>
          .
          <source>HCI in Work and Learning, Life and Leisure</source>
          , pages
          <volume>427</volume>
          {
          <fpage>442</fpage>
          ,
          <year>2010</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <given-names>J.</given-names>
            <surname>Manyka</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Chui</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Brown</surname>
          </string-name>
          , J. Bughin,
          <string-name>
            <given-names>R.</given-names>
            <surname>Dobbs</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Roxburgh</surname>
          </string-name>
          ,
          <article-title>and</article-title>
          <string-name>
            <given-names>A. H.</given-names>
            <surname>Byers</surname>
          </string-name>
          .
          <article-title>Big data: The next frontier for innovation, competition, and productivity</article-title>
          .
          <source>Technical report, McKinsey Global Institute</source>
          ,
          <year>2011</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <given-names>C.</given-names>
            <surname>Musto</surname>
          </string-name>
          .
          <article-title>Enhanced vector space models for content-based recommender systems</article-title>
          .
          <source>In Proceedings of the fourth ACM conference on Recommender systems, RecSys '10</source>
          , pages
          <fpage>361</fpage>
          {
          <fpage>364</fpage>
          . ACM,
          <year>2010</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <given-names>C.</given-names>
            <surname>Musto</surname>
          </string-name>
          , G. Semeraro,
          <string-name>
            <given-names>P.</given-names>
            <surname>Lops</surname>
          </string-name>
          , and M. de Gemmis.
          <article-title>Random indexing and negative user preferences for enhancing content-based recommender systems</article-title>
          .
          <source>In EC-Web</source>
          , pages
          <volume>270</volume>
          {
          <fpage>281</fpage>
          ,
          <year>2011</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <given-names>A.</given-names>
            <surname>Passant</surname>
          </string-name>
          and
          <string-name>
            <given-names>Y.</given-names>
            <surname>Raimond</surname>
          </string-name>
          .
          <article-title>Combining Social Music and Semantic Web for MusicRelated Recommender Systems</article-title>
          .
          <source>In Social Data on the Web, Workshop of the 7th International Semantic Web Conference</source>
          , Karlsruhe, Deutschland,
          <year>Oktober 2008</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <given-names>Giovanni</given-names>
            <surname>Semeraro</surname>
          </string-name>
          , Stefano Ferilli, Nicola Fanizzi, and
          <string-name>
            <given-names>Fabio</given-names>
            <surname>Abbattista</surname>
          </string-name>
          .
          <article-title>Learning interaction models in a digital library service</article-title>
          . In Mathias Bauer,
          <string-name>
            <given-names>Piotr J.</given-names>
            <surname>Gmytrasiewicz</surname>
          </string-name>
          , and Julita Vassileva, editors,
          <source>User Modeling</source>
          , volume
          <volume>2109</volume>
          of Lecture Notes in Computer Science, pages
          <volume>44</volume>
          {
          <fpage>53</fpage>
          . Springer,
          <year>2001</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <given-names>U.</given-names>
            <surname>Shardanand</surname>
          </string-name>
          .
          <article-title>Social information ltering for music recommendation</article-title>
          .
          <source>Bachelor thesis</source>
          , Massachusetts Institute of Technology, Massachusetts,
          <year>1994</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <given-names>D.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Li</surname>
          </string-name>
          , and
          <string-name>
            <given-names>M.</given-names>
            <surname>Ogihara</surname>
          </string-name>
          .
          <article-title>Are tags better than audio? the e ect of joint use of tags and audio content features for artistic style clustering</article-title>
          .
          <source>In ISMIR</source>
          , pages
          <volume>57</volume>
          {
          <fpage>62</fpage>
          ,
          <year>2010</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>