<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Investigating the Decision Making Process of Users based on the PoliMovie Dataset</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Mona Naseri</string-name>
          <email>mona.naseri@polimi.it</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Mehdi Elahi</string-name>
          <email>mehdi.elahi@polimi.it</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Paolo Cremonesi</string-name>
          <email>paolo.cremonesi@polimi.it</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Politecnico di Milano</institution>
          ,
          <addr-line>Via Ponzio 34/5 20133, Milano</addr-line>
          ,
          <country country="IT">Italy</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>Making decision on which movie to watch is nowadays not an easy process for majority of people. Many people may decide to watch a movie based on the genre attribute of a movie, while for the others, the director can be the attribute that drives them to watch a movie. Hence, people may have di↵erent features, taken into account, when deciding which movie to watch. Recommender Systems can help people in making decision by allowing them to enter the attribute(s), that is the most important to them, and filter the movie catalog accordingly. In this paper we try to investigate the process that results in choosing a movie to watch by people. Hence we present an ongoing work that will ultimately lead to building a dataset (called PoliMovie) that will contain the preferences of users not only on movies but also on attributes of the movies, such as genre, director, and cast., that users selected as the most important attributes when choosing a movie to watch. We report some preliminary results based on the preferences collected from about 400 users, which confirm the di↵erence and complexity of decision making process for di↵erent users.</p>
      </abstract>
      <kwd-group>
        <kwd>Recommender systems</kwd>
        <kwd>attributes</kwd>
        <kwd>decision-making</kwd>
        <kwd>userprofile</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>Nowadays, choosing the right movie to watch is challenging for people due to
huge variety of the movies. By investigating the way people choose the movies
to watch, one may notice that there is no unique way that people may follow
in making decision. While some people may decide based on the genre of the
movies, others may prefer movies in which their favorite actor plays.</p>
      <p>
        Recommender Systems try to support this process by finding movies that
can match users’ need and interest. They analyze set of features (attributes) of
the items and create user profile for a user, that indicates the preferences and
interests of her for those features. Indeed, recommendations are generated by
matching up the features of the user profile (i.e., a structured representation of
her interests) against the features of the item. In order to do this, the
contentbased recommender systems build a Vector Space Model (VSM), where each
item is represented by an n-dimensional vector. Each dimension in this model
represents an attribute from the overall set of attributes used to describe the
item. Using this model, the system computes a relevance score that represents the
user’s degree of interest toward that item [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]. This also allows the recommender
systems to produce explanations to recommendations and to naturally solve the
new item problem [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ].
      </p>
      <p>
        These systems typically include also this side information describing the
attributes of movies (e.g., in the movie domain categories like genre, director, cast)
and build user models as prediction of users’ preference on features [
        <xref ref-type="bibr" rid="ref5 ref6">6,5</xref>
        ].
Incorporating such information is beneficial since it implicitly helps the system to
understand important attributes that users may base their reasoning in order to
choose the right movie to watch.
      </p>
      <p>Moreover, traditionally, such user models are tested using their ability to
provide relevant recommendations of movies. Hence, one of the publicly available
datasets such as Movielens is used as a ground truth for the preferences of users
on movies. However, as of our knowledge, none of the publicly available datasets
contain the explicit preferences of users on movie attributes, and hence, the
user models can not be evaluated for the true prediction of users preferences on
attributes, that they may choose movies to watch based on them.</p>
      <p>
        In this paper, we have analyzed our dataset, called PoliMovie [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ], and
obtained some preliminary results that can show clear di↵erences between the way
di↵erent users may form their reasoning process when making decision on what
to watch. These results confirm our initial intuition that traditional user models
based on implicit user preferences on attributes do not match well with explicit
user models in which users are explicitly elicited to provide their opinions on
attributes.
      </p>
      <p>We show that, in many cases, the favorite movies selected by a user have
attributes (e.g., actors, directors, genre) totally di↵erent from the attributes
selected as favorites by the same user. Hence, a user may like “The Dark Knight”
movie without choosing “Action, Crime, Drama” as favorite genre, “Christopher
Nolan” as favorite director, and “Christian Bale” as favorite actor. We show that
a big group of users have chosen attributes that they choose movies based on
them, that have not even contained in the movies they actually choose as their
favorite. These results are convincing us to even continue the data collection
procedure as well as a more extensive analysis of the data.
2</p>
    </sec>
    <sec id="sec-2">
      <title>Preliminary results</title>
      <p>In this section we report some preliminary results of analysis we have conducted
on this initial collection of data. Up to the date of writing this paper, PoliMovie
dataset contains approximately 1600 movies, 300 casts, 200 directors and all
genres (23) have been selected as favorite at least by one user.</p>
      <p>First of all, we have understood that the most important attributes, the
users take into account when making decision on which movie to watch are
“Genre” and the “Cast” of the movies, respectively. However, the “Director”
and the “Year of Production” play the least important role in decision making
for choosing which movie to watch.</p>
      <p>Accordingly, we have measured the popularity of the cast based on the users
explicit rating and also based on implicitly inferring from the movies they have
0.05</p>
      <p>0.1 0.15
%Match between explicit vs implicit user-cast profiles
Jaccard similarity between explicit vs implicit user-cast profiles
0.2</p>
      <p>0.25
0.15 0.2 0.25
%Match between explicit vs implicit user-genre profiles</p>
      <p>Jaccard similarity between explicit vs implicit user-genre profiles
0.05
0.1
0.3
0.35
0.4
added as favorite. Comparing these two lists, we have noticed that some
actors/actresses in the first list (based on explicit rating) are not actually presented
in the second (based on implicit preferences) or vice versa.</p>
      <p>Moreover, our observation shows that, based on users’ favorite movies, “Drama”
is the most popular genre. This is while, the most popular genre, selected by
users, is totally the opposite genre, i.e. “Comedy”.</p>
      <p>
        Additionally, we have computed the Jaccard similarity [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ] between the user
profiles, based on explicit and implicit user preferences. We considered only those
features that user selected as her most important feature(s). Each user could
select maximum two among five movie features (genre, cast, director, rating,
production-year). For instance, if user selected “Genre” and “Cast” as her most
important features, we measured similarity between her explicit and implicit
profiles only based on these two features.
      </p>
      <p>We have observed interesting results shown in figure 1, where x axis is the
percentage of match, and y axis is percentage of users with certain match
between their explicit and implicit profiles. The observations show that user models
based on implicit user preferences on attributes do not match well with explicit
user models in which users are explicitly elicited to provide their opinions on
attributes.</p>
      <p>Finally, it worth noting that, the distribution of similarities are totally
di↵erent for favorite casts and favorite genres. Indeed, the distribution of the match
between two user profiles based on cast looks exponential while the distribution
based on genre interestingly looks Gaussian with the mean about 0.2. This may
indicate that majority of users may have only 20% of match between the genre
they like and the genre of the movies they watch. Accordingly, there are people
with more match or less match between these two types of preferences. However,
for the user’s favorite cast, this may not be the case and most of the users may
watch a lot of movies where their favorite cast do not play.</p>
    </sec>
    <sec id="sec-3">
      <title>Conclusion</title>
      <p>This paper presents an ongoing work of data collection that will ultimately result
in a dataset called PoliMovie. The PoliMovie dataset is publicly available 1. In
spite of the currently available datasets, it will contain not only the preferences of
users provided for movies, but also preferences of users on features (attributes)
of movies. Such dataset can be very useful in investigating the di↵erence
between attributes that users consider when choosing movies to watch, as well as,
their final decision on which movies they actually watch. Hence, the researchers
and the practitioners in the community of Recommender Systems can use such
dataset to evaluate the quality of explanations provided by some types of
recommender systems as well as analyzing the complex process of decision making by
users while also benchmarking their feature-based recommendation algorithms.
For our future work, we are going to evaluate some of the state-of-the-art
algorithms on our dataset. This would be useful to see which algorithms can better
predict preference of users on features (attributes). Moreover, by having the
explicit opinion of users on features, we can evaluate the quality of explanations
provided by some types of recommender systems.
1 through the link: http://recsys.deib.polimi.it/polimovie-dataset</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <given-names>M.</given-names>
            <surname>Elahi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Ricci</surname>
          </string-name>
          , and
          <string-name>
            <given-names>N.</given-names>
            <surname>Rubens</surname>
          </string-name>
          .
          <article-title>Active learning strategies for rating elicitation in collaborative filtering: a system-wide perspective</article-title>
          .
          <source>ACM Transactions on Intelligent Systems and Technology (TIST)</source>
          ,
          <volume>5</volume>
          (
          <issue>1</issue>
          ):
          <fpage>13</fpage>
          ,
          <year>2013</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <given-names>M.</given-names>
            <surname>Levandowsky</surname>
          </string-name>
          and
          <string-name>
            <given-names>D.</given-names>
            <surname>Winter</surname>
          </string-name>
          .
          <article-title>Distance between sets</article-title>
          .
          <source>Nature</source>
          ,
          <volume>234</volume>
          (
          <issue>5323</issue>
          ):
          <fpage>34</fpage>
          -
          <lpage>35</lpage>
          ,
          <year>1971</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <given-names>P.</given-names>
            <surname>Lops</surname>
          </string-name>
          ,
          <string-name>
            <surname>M. De Gemmis</surname>
            , and
            <given-names>G.</given-names>
          </string-name>
          <string-name>
            <surname>Semeraro</surname>
          </string-name>
          .
          <article-title>Content-based recommender systems: State of the art and trends</article-title>
          .
          <source>In Recommender systems handbook</source>
          , pages
          <fpage>73</fpage>
          -
          <lpage>105</lpage>
          . Springer,
          <year>2011</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <given-names>M.</given-names>
            <surname>Nasery</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Elahi</surname>
          </string-name>
          , and
          <string-name>
            <given-names>P.</given-names>
            <surname>Cremonesi</surname>
          </string-name>
          .
          <article-title>Polimovie: a feature-based dataset for recommender systems</article-title>
          . In Workshop on Crowdsourcing and
          <article-title>human computation for recommender systems</article-title>
          ,
          <source>CrowdRec at RecSys</source>
          ,
          <year>2015</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <given-names>Y.</given-names>
            <surname>Shi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Larson</surname>
          </string-name>
          ,
          <article-title>and</article-title>
          <string-name>
            <given-names>A.</given-names>
            <surname>Hanjalic</surname>
          </string-name>
          .
          <article-title>Collaborative filtering beyond the user-item matrix: A survey of the state of the art and future challenges</article-title>
          .
          <source>ACM Comput. Surv.</source>
          ,
          <volume>47</volume>
          (
          <issue>1</issue>
          ):3:
          <fpage>1</fpage>
          -
          <lpage>3</lpage>
          :
          <fpage>45</fpage>
          , May
          <year>2014</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <given-names>P.</given-names>
            <surname>Symeonidis</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Nanopoulos</surname>
          </string-name>
          , and
          <string-name>
            <given-names>Y.</given-names>
            <surname>Manolopoulos</surname>
          </string-name>
          .
          <article-title>Feature-weighted user model for recommender systems</article-title>
          .
          <source>In Proceedings of the 11th International Conference on User Modeling</source>
          ,
          <source>UM '07</source>
          , pages
          <fpage>97</fpage>
          -
          <lpage>106</lpage>
          , Berlin, Heidelberg,
          <year>2007</year>
          . Springer-Verlag.
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>