<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Movie Rating Prediction using Multimedia Content and Modeling as a Classification Problem</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Fatemeh Nazary</string-name>
          <email>fatemeh.nazary01@universitadipavia.it</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Yashar Deldjoo</string-name>
          <email>deldjooy@acm.org</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>University of Milano-Bicocca</institution>
          ,
          <country country="IT">Italy</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>University of Pavia</institution>
          ,
          <country country="IT">Italy</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2018</year>
      </pub-date>
      <fpage>29</fpage>
      <lpage>31</lpage>
      <abstract>
        <p>This paper presents the method proposed for the recommender system task in Mediaeval 2018 on predicting user global ratings given to movies and their standard deviation through the audiovisual content and the associated metadata. In the proposed work, we model the rating prediction problem as a classification problem and employ diferent classifiers for the prediction task. Furthermore, in order to obtain a video-level representation of features from clip-level features, we employ statistical summarization functions. Results are promising and show the potential of leveraging the audiovisual content for improving the quality of existing movie recommendation systems in service.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>INTRODUCTION AND CONTEXT</title>
      <p>
        Video recordings are complex audiovisual signals. When we watch
a movie, a large amount of information is communicated to us
through diferent multimedia channels, in particular, the audio and
the visual channel. As a result, the video content can be described
in diferent manners since its consumption is not limited to one
type of perception. These multiple facets can be manifested by
descriptors of visual and audio content, but also in terms of metadata,
including information about the movie’s genre, actors, or plot of
a movie. The goal of movie recommendation systems (MRS) is to
provide personalized suggestions about movies that users would
likely find interesting. Collaborative filtering (CF) models lie at the
core of most MRS in service today and generate recommendation
by exploiting the items favored by other like-minded users [
        <xref ref-type="bibr" rid="ref10 ref2 ref9">2, 9, 10</xref>
        ].
Content-based filtering (CBF) methods on the other hand, base their
recommendations on the similarities between the target user’s
preferred or consumed items and other items in the catalog, where
this similarity is defined by computing a content-centric similarity
using descriptors (features) inferred or extracted from the item
content, typically by leveraging textual metadata, either editorial,
e.g., genre, cast, director, or user generated, e.g., tags, reviews [
        <xref ref-type="bibr" rid="ref1 ref8">1, 8</xref>
        ].
For instance, the authors in [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ] developed a heterogeneous
socialaware MRS that uses movie-poster images and textual description,
as well as user ratings and social relationships in order to generate
recommendations. Another example is [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ], in which a hybrid MRS
using tags and ratings is proposed, where user profiles are formed
based on users’ interaction in a social movie network.
      </p>
      <p>
        Regardless of the approach, metadata are prone to errors and
expensive to collect. Moreover, user feedback and user-generated
metadata are rare or absent for new movies, making it dificult
or even impossible to provide good quality recommendations a
scenario known as the cold-start problem. The goal of the current
MediaEval task [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ] is to bridge the gap between advances and
perspective in multimedia and recommender systems communities [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ].
In particular, participants are required to use the audiovisual
content and metadata in order predict global ratings of users provided
to movies (representing their appreciation/dis-appreciation) and
the corresponding standard deviation (characterizing users
agreement and disagreement). This task is novel in two regards. First, the
provided dataset uses movie clips instead of trailers [
        <xref ref-type="bibr" rid="ref4 ref5 ref6">4–6</xref>
        ], thereby
providing a wider variety of the movie’s aspects by showing
different kinds of scenes. Second, including information about the
ratings’ variance makes it possible to assess users’ agreement and
to uncover polarizing movies [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ].
2
      </p>
    </sec>
    <sec id="sec-2">
      <title>PROPOSED APPROACH</title>
      <p>
        The proposed framework can be divided into three phases:
(1) Multimodal feature fusion: This step is carried out in
the multimodal phase for hybridization of the features. It
aims to fuse two descriptors of diferent nature (e.g., audio
and visual) into a fixed-length descriptor. In this work, we
chose concatenation of features as a simple early fusion
approach toward multimodal fusion.
(2) Video-level representation building: The novelty of
this task is that it uses movie clips instead of movie
trailers [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ] in which each movie has several associated clips.
This step aims to aggregate clip-level representation of
features in order to build a video-level representation so it can
be used in the classification stage. In this work, we adopted
aggregation methods based on statistical summarization
including mean(), min() and max () to obtain video-level
representations of the features.
(3) Classification: The provided scores (global ratings and
their stds) are continuous values. Our approach to the
prediction problem consisted of treating it as a
classification problem. This means prior to classification, the
target scores are quantized to predefined values. We chose
2-level uniform quantization for global rating meaning
the ratings were mapped to one of the values in the set
{0.5, 1, 1.5, ..., 4.5, 5}. As for std, we chose 10-level plus
16-level uniform quantization where in the latter case, the
higher number of levels were chosen to provide a larger
resolution to the narrow distribution of std scores (std values
are quite compact around [0.5-1.5] whereas global ratings
are spread in the range [
        <xref ref-type="bibr" rid="ref1 ref2 ref3 ref4 ref5">0-5</xref>
        ]). Finally for classification,
we investigated three classification approaches: logistic
regression (LR), k-nearest neighbor (KNN) and random
forest (RF) classifier.
      </p>
    </sec>
    <sec id="sec-3">
      <title>RESULTS AND ANALYSIS</title>
      <p>The results of classification using the proposed approach are
presented in Table 1. Regarding the comparison of classifiers, we can
note that RF is the best classifier among others usually generating
the best performance for each feature or feature combination while
KNN is the worst (note that KNN classifier is a lazy classifier). Thus,
in reporting the results, we mostly base our judgment on results
obtained from RF and in some cases on LR. The final submitted
runs are selected based on the ones performing the best on the
development set, which are highlighted in bold in Table 1.
Predicting average ratings: From the result obtained it can be
seen that the performance of all audio and visual features,
regardless of their type i.e., traditional or state of the art, are closely similar
to each other. These results with a close margin look similar to the
performance of the genre descriptor. In fact the diference between
the best audio or visual feature and genre is 6-7% while this
diference with tag can reach up to 45%. These results are interesting
and confirm that user-generated tags assigned to movies contain
semantics that are well correlated with ratings given to movies by
user, even though the users of tags and ratings are not necessary the
same. For multimodal case, one can note that simple concatenation
of the features can not improve the final performance substantially
compared with unimodal audiovisual features. The best results are
obtained in cross modal fusion for i-vector + AVF (compare 0.54 v.s
Genre: 0.53) and BLF + Deep (0.54). However for metadata-based
multimodal fusion, the general observation is that audiovisual
features can slightly improve the performance of genre and tag (e.g.,
compare AVF+Tag: 0.44 v.s. Tag: 0.48 for LR and 0.38 v.s. 0.39 for
RF), hinting that they have a complementary nature which can be
better leveraged if right a fusion strategy is adopted.</p>
      <p>Predicting standard deviation of ratings: As for predicting
standard deviation of ratings, for unimodal case, it can be seen
that except genre feature with the worst performance, the rest of
audiovisual features and tag metadata have very similar results.
This indicates that genre is the weakest descriptor and compared to
others less capable of distinguishing diference in users’ opinions.
Note that under LR, genre descriptor performs similar to other
audiovisual features. For multimodal case, the results for majority
of combinations are pretty similar regardless of the classifier type.
The best performing combinations are AVF + Genre and i-vector +
Genre with the RMSE equal to 0.12 and 0.13.
4</p>
    </sec>
    <sec id="sec-4">
      <title>CONCLUSION</title>
      <p>
        This paper reports the description of the method for the
"Recommending movie Using Content: Which content is key" MediaEval
2018 task [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]. The proposed approach consists of three main steps:
(i) multimiodal fusion, (ii) video-level video representation building
and (iii) classification. Results of experiments using three
classification approaches are promising and show the eficacy of audiovisual
content in predicting user global ratings and to a lesser extent for
predicting rating variance.
      </p>
      <p>Recommending Movies Using Content: Which content is key?</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <surname>Charu</surname>
            <given-names>C</given-names>
          </string-name>
          <string-name>
            <surname>Aggarwal</surname>
          </string-name>
          .
          <year>2016</year>
          .
          <article-title>Content-based recommender systems</article-title>
          .
          <source>In Recommender systems</source>
          . Springer,
          <fpage>139</fpage>
          -
          <lpage>166</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <surname>Charu</surname>
            <given-names>C</given-names>
          </string-name>
          <string-name>
            <surname>Aggarwal</surname>
          </string-name>
          .
          <year>2016</year>
          .
          <article-title>Neighborhood-based collaborative filtering</article-title>
          .
          <source>In Recommender Systems</source>
          . Springer,
          <fpage>29</fpage>
          -
          <lpage>70</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>Yashar</given-names>
            <surname>Deldjoo</surname>
          </string-name>
          , Mihai Gabriel Constantin, Thanasis Dritsas, Markus Schedl, and
          <string-name>
            <given-names>Bogdan</given-names>
            <surname>Ionescu</surname>
          </string-name>
          .
          <year>2018</year>
          .
          <article-title>The MediaEval 2018 Movie Recommendation Task: Recommending Movies Using Content</article-title>
          . In MediaEval 2018 Workshop.
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>Yashar</given-names>
            <surname>Deldjoo</surname>
          </string-name>
          , Mihai Gabriel Constantin, Hamid Eghbal-Zadeh, Markus Schedl, Bogdan Ionescu, and
          <string-name>
            <given-names>Paolo</given-names>
            <surname>Cremonesi</surname>
          </string-name>
          .
          <year>2018</year>
          .
          <article-title>AudioVisual Encoding of Multimedia Content to Enhance Movie Recommendations</article-title>
          .
          <source>In Proceedings of the Twelfth ACM Conference on Recommender Systems</source>
          . ACM. https://doi.org/10.1145/3240323.3240407
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>Yashar</given-names>
            <surname>Deldjoo</surname>
          </string-name>
          , Mihai Gabriel Constantin, Bogdan Ionescu, Markus Schedl, and
          <string-name>
            <given-names>Paolo</given-names>
            <surname>Cremonesi</surname>
          </string-name>
          .
          <year>2018</year>
          .
          <article-title>MMTF-14K: A Multifaceted Movie Trailer Dataset for Recommendation and Retrieval</article-title>
          .
          <source>In Proceedings of the 9th ACM Multimedia Systems Conference (MMSys</source>
          <year>2018</year>
          ). Amsterdam, the Netherlands.
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>Yashar</given-names>
            <surname>Deldjoo</surname>
          </string-name>
          , Mehdi Elahi, Massimo Quadrana, and
          <string-name>
            <given-names>Paolo</given-names>
            <surname>Cremonesi</surname>
          </string-name>
          .
          <year>2018</year>
          .
          <article-title>Using Visual Features based on MPEG-7 and Deep Learning for Movie Recommendation</article-title>
          .
          <source>International Journal of Multimedia Information Retrieval</source>
          (
          <year>2018</year>
          ),
          <fpage>1</fpage>
          -
          <lpage>13</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>Yashar</given-names>
            <surname>Deldjoo</surname>
          </string-name>
          , Markus Schedl, Paolo Cremonesi, and
          <string-name>
            <given-names>Gabriella</given-names>
            <surname>Pasi</surname>
          </string-name>
          .
          <year>2018</year>
          .
          <article-title>Content-Based Multimedia Recommendation Systems: Definition and Application Domains</article-title>
          .
          <source>In Proceedings of the 9th Italian Information Retrieval Workshop (IIR</source>
          <year>2018</year>
          ). Rome, Italy.
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>Pasquale</given-names>
            <surname>Lops</surname>
          </string-name>
          , Marco De Gemmis, and
          <string-name>
            <given-names>Giovanni</given-names>
            <surname>Semeraro</surname>
          </string-name>
          .
          <year>2011</year>
          .
          <article-title>Content-based recommender systems: State of the art and trends</article-title>
          .
          <source>In Recommender systems handbook</source>
          . Springer,
          <fpage>73</fpage>
          -
          <lpage>105</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>Francesco</given-names>
            <surname>Ricci</surname>
          </string-name>
          , Lior Rokach, and
          <string-name>
            <given-names>Bracha</given-names>
            <surname>Shapira</surname>
          </string-name>
          .
          <year>2015</year>
          .
          <article-title>Recommender systems: introduction and challenges</article-title>
          .
          <source>In Recommender systems handbook. Springer</source>
          ,
          <fpage>1</fpage>
          -
          <lpage>34</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <surname>Yue</surname>
            <given-names>Shi</given-names>
          </string-name>
          ,
          <string-name>
            <given-names>Martha</given-names>
            <surname>Larson</surname>
          </string-name>
          , and
          <string-name>
            <given-names>Alan</given-names>
            <surname>Hanjalic</surname>
          </string-name>
          .
          <year>2014</year>
          .
          <article-title>Collaborative filtering beyond the user-item matrix: A survey of the state of the art and future challenges</article-title>
          .
          <source>ACM Computing Surveys (CSUR) 47</source>
          ,
          <issue>1</issue>
          (
          <year>2014</year>
          ),
          <fpage>3</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <surname>Shouxian</surname>
            <given-names>Wei</given-names>
          </string-name>
          , Xiaolin Zheng, Deren Chen, and
          <string-name>
            <given-names>Chaochao</given-names>
            <surname>Chen</surname>
          </string-name>
          .
          <year>2016</year>
          .
          <article-title>A hybrid approach for movie recommendation via tags and ratings</article-title>
          .
          <source>Electronic Commerce Research and Applications</source>
          <volume>18</volume>
          (
          <year>2016</year>
          ),
          <fpage>83</fpage>
          -
          <lpage>94</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <surname>Zhou</surname>
            <given-names>Zhao</given-names>
          </string-name>
          ,
          <string-name>
            <given-names>Qifan</given-names>
            <surname>Yang</surname>
          </string-name>
          , Hanqing Lu, Tim Weninger, Deng Cai, Xiaofei He, and
          <string-name>
            <given-names>Yueting</given-names>
            <surname>Zhuang</surname>
          </string-name>
          .
          <year>2017</year>
          .
          <article-title>Social-Aware Movie Recommendation via Multimodal Network Learning</article-title>
          .
          <source>IEEE Transactions on Multimedia</source>
          (
          <year>2017</year>
          ).
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>