<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>TUD-MMC at MediaEval 2016: Context of Experience task</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Bo Wang</string-name>
          <email>b.wang-6@student.tudelft.nl</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Cynthia C. S. Liem</string-name>
          <email>C.C.S.Liem@tudelft.nl</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Delft University of Technology</institution>
          ,
          <addr-line>Delft</addr-line>
          ,
          <country country="NL">The Netherlands</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2016</year>
      </pub-date>
      <fpage>20</fpage>
      <lpage>21</lpage>
      <abstract>
        <p>This paper provides a three-step framework to predict user assessment of the suitability of movies for an inflight viewing context. For this, we employed classifier stacking strategies. First of all, using the different modalities of training data, twenty-one classifiers were trained together with a feature selection algorithm. Final predictions were then obtained by applying three classifier stacking strategies. Our results reveal that different stacking strategies lead to different evaluation results. A considerable improvement can be found for the F1-score when using the label stacking strategy.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>INTRODUCTION</title>
      <p>
        A substantial amount of research has been conducted in
recommender systems that focus on user preference prediction. Here,
taking contextual information into account can have significant
positive impact on the performance of recommender systems [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ].
      </p>
      <p>The MediaEval Context of Experience task focuses on a specific
type of context: the viewing context of the user. The challenge
considers predicting the multimedia content that users find most
fitting to watch in a specific viewing condition, more specifically,
while being on a plane.</p>
    </sec>
    <sec id="sec-2">
      <title>DATASET DESCRIPTION AND INITIAL</title>
    </sec>
    <sec id="sec-3">
      <title>EXPERIMENTS</title>
      <p>The dataset for the Context of Experience (CoE) task[5] contains
metadata and pre-extracted features for 318 movies [6]. Features
are multimodal and include textual features, visual features and
audio features. The training set contains 95 labeled movies, which
are labeled as 0 (bad for airplane) or 1 (good for airplane).</p>
      <p>A set of initial experiments has been conducted in order to
evaluate the usefulness of the various modalities in the CoE dataset [6].
A rule-based PART classifier was employed to evaluate the feature
performance in terms of Precision, Recall and F1 Score, the result
can be found in Table 1.</p>
    </sec>
    <sec id="sec-4">
      <title>MULTIMODAL CLASSIFIER STACKING</title>
      <p>
        Ensemble learning uses a combination of different classifiers,
usually getting a much better generalization ability. This
particularly is the case for weak learners, which can be defined as learning
algorithms that perform just slightly better than random guessing
by themselves, but can be jointly grouped into an algorithm with
arbitrarily high accuracy [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ].
      </p>
      <sec id="sec-4-1">
        <title>Features used User rating Visual Metadata</title>
        <p>Metadata + user rating
Metadata + visual</p>
      </sec>
      <sec id="sec-4-2">
        <title>Precision</title>
        <p>0.371
0.447
0.524
0.581
0.584</p>
      </sec>
      <sec id="sec-4-3">
        <title>Recall</title>
        <p>0.609
0.476
0.516
0.6
0.6</p>
        <p>F1
0.461
0.458
0.519
0.583
0.586</p>
        <p>Therefore, we were interested in taking a multimodal classifier
stacking approach to the given problem, and use a combination of
multiple weak learners to ‘boost’ them into a strong learner.</p>
        <p>The process can be separated into three stages: classifier
selection, feature selection and classifier stacking.
3.1</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>Classifier Selection</title>
      <p>First of all , we want to select base classifiers that will be useful
candidates in a stacking approach. For this, we use the following
classifier selection procedure:
1. Initialize a list of candidate classifiers. For each modality,
we consider the following classifiers: k-nearest neighbor,
nearest mean, decision tree, logistic regression, SVM,
bagging, random forest, AdaBoost, gradient boosting, and naive
Bayes. We do not apply parameter tuning, but take the
default parameter values as offered by scikit-learn1.
2. Perform 10-fold cross-validation on the classifiers. As input
data, we use the training data set and its ground truth labels,
per single modality. For the audio MFCC features, we set
NaN values to 0, and calculate the average of each MFCC
coefficient over all frames.
3. If Precision and Recall and F1-Score &gt; 0.5, keep the
candidate classifier on the given modality as base classifier for our
stacking approach.</p>
      <p>The selected base classifiers and their relevant modalities can
be found in Table 2. It should be noted that the performance of
Bagging and Random forest is not stable. This is because
Bagging tries to use different subset of instances in each run and
RandomForest tries to use different subsets of instances and
features in each run.
3.2</p>
    </sec>
    <sec id="sec-6">
      <title>Feature Selection</title>
      <p>For each classifier and corresponding modality, a better-performing
subspace of features may optimize results further. Since we have</p>
      <sec id="sec-6-1">
        <title>Classifier</title>
        <p>k-Nearest neighbor
Nearest mean classifier</p>
        <p>Decision tree</p>
        <p>Logistic regression
SVM (Gaussian Kernel)</p>
        <p>Bagging
Random Forest</p>
        <p>AdaBoost
Gradient Boosting Tree</p>
        <p>Naive Bayes
k-Nearest neighbor
SVM (Gaussian Kernel)
k-Nearest neighbor</p>
        <p>Decision tree</p>
        <p>Logistic regression
SVM (Gaussian Kernel)</p>
        <p>Random Forest</p>
        <p>AdaBoost
Gradient Boosting Tree</p>
        <p>Logistic Regression
Gradient Boosting Tree</p>
        <p>Modality
metadata
metadata
metadata
metadata
metadata
metadata
metadata
metadata
metadata
textual
textual
textual
visual
visual
visual
visual
visual
visual
visual
audio
audio</p>
        <p>
          In previous research, classifier stacking (or metalearning) has
been proved beneficial for predictive performance by combining
different learning systems which each have different inductive bias
(e.g. representation, search heuristics, search space) [
          <xref ref-type="bibr" rid="ref4">4</xref>
          ]. By
combining separately learned concepts, meta-learning is expected to
derive a higher-level learned model that more accurately can
predict than any of the individual learners. In our work, we consider
three types of stacking strategies:
1. Majority Voting: this is the simplest case, where we select
classifiers and feature subspaces through the steps above, and
assign final predicted labels through majority voting on the
labels of the 21 classifiers.
2. Label Stacking: Assume we have n instances and T base
classifiers, then we can generate an n by T matrix consisting
of predictions (labels) given by each classifier. Label
combining strategy tries to build a second-level classifier based
on this label matrix, and return a final prediction result for
that.
3. Label-Feature Stacking: Similar to label stacking, label-feature
stacking strategy uses both base-classifier predictions and
features as training data to predict output.
4.
        </p>
        <p>We considered all prediction results by the 21 selected base
classifiers, and then applied the three different classifier stacking
strategies to the test data using 10-fold cross-validation. As results for
label stacking vs. label attribute stacking were comparable on the
training data, we only consider voting vs. label stacking on the test
data.</p>
        <p>All obtained results, on the training (development) and test dataset,
are given in Table 3. On the training data, we notice significant
improvement can be found in terms of Precision, Recall as well as F1
score in comparison to results obtained on individual modalities.
The voting strategy results in the best precision score, but has bad
performance in terms of recall. On the contrary, label stacking has
higher recall and the highest F1 score.</p>
        <p>Considering results obtained on the test dataset, we can conclude
that label stacking is more robust than the voting strategy. For
voting strategy, a significant decrease can be found in terms of
precision on test set. This is because majority vote (and Bayesian
averaging) tendency to over-fit derives from the likelihood’s
exponential sensitivity to random fluctuations in the sample, and increases
with the number of models considered. Meanwhile, label stacking
strategy performs reasonable well on test data.</p>
      </sec>
      <sec id="sec-6-2">
        <title>Stacking Strategy</title>
        <p>Voting (cv)</p>
        <p>Label Stacking (cv)
Label Attribute Stacking (cv)</p>
        <p>Voting (test)
Label Stacking (test)</p>
      </sec>
      <sec id="sec-6-3">
        <title>Precision</title>
        <p>0.94
0.72
0.71
0.62
0.62</p>
      </sec>
      <sec id="sec-6-4">
        <title>Recall</title>
        <p>0.57
0.86
0.79
0.80
0.90</p>
        <p>In our entry for the MediaEval CoE task, we aimed to improve
classifier performance by a combination of classifier selection,
feature selection and classifier stacking. Results reveal that employing
a ensemble approach can considerably increase the classification
performance, and is suitable for treating the multimodal Right
Inflight dataset.</p>
        <p>The larger diversity of base classifiers is able to produce a more
robust ensemble classifier. On the other hand, a blending of
multiple classifiers may also have some drawbacks, e.g computational
costs, and difficulty in traceable interpretation.</p>
        <p>We expect better results for our method can still be obtained
through parameter tuning, and by applying more robust classifier
stacking methods, such as feature weighted linear stacking [7].
Advances in distributed and parallel knowledge discovery,
pages 81–114. MIT/AAAI Press, 2000.
[5] M. Riegler, , C. Spampinato, M. Larson, P. Halvorsen, and
C. Griwodz. The mediaeval 2016 context of experience task:
Recommending videos suiting a watching situation. In
Proceedings of the MediaEval 2016 Workshop, 2016.
[6] M. Riegler, M. Larson, C. Spampinato, P. Halvorsen, M. Lux,
J. Markussen, K. Pogorelov, C. Griwodz, and H. Stensland.
Right inflight? A dataset for exploring the automatic
prediction of movies suitable for a watching situation. In
Proceedings of the 7th International Conference on</p>
        <p>Multimedia Systems, pages 45:1–45:6. ACM, 2016.
[7] J. Sill, G. Takacs, L. Mackey, and D. Lin. Feature-weighted
linear stacking. arXiv:0911.0460, 2009.</p>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>G.</given-names>
            <surname>Adomavicius</surname>
          </string-name>
          and
          <string-name>
            <given-names>A.</given-names>
            <surname>Tuzhilin</surname>
          </string-name>
          .
          <article-title>Context-aware recommender systems</article-title>
          .
          <source>In Recommender systems handbook</source>
          , pages
          <fpage>217</fpage>
          -
          <lpage>253</lpage>
          . Springer,
          <year>2011</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>Y.</given-names>
            <surname>Freund</surname>
          </string-name>
          and
          <string-name>
            <given-names>R. E.</given-names>
            <surname>Schapire</surname>
          </string-name>
          .
          <article-title>A Decision-theoretic Generalization of On-line Learning and an Application to Boosting</article-title>
          .
          <source>Journal of computer and system sciences</source>
          ,
          <volume>55</volume>
          :
          <fpage>119</fpage>
          -
          <lpage>139</lpage>
          ,
          <year>1997</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>H.</given-names>
            <surname>Liu</surname>
          </string-name>
          and
          <string-name>
            <given-names>R.</given-names>
            <surname>Setiono</surname>
          </string-name>
          .
          <article-title>Feature selection and classification-a probabilistic wrapper approach</article-title>
          .
          <source>In Proceedings of the 9th International Conference on Industrial and Engineering Applications of AI and ES</source>
          , pages
          <fpage>419</fpage>
          -
          <lpage>424</lpage>
          ,
          <year>1997</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>A.</given-names>
            <surname>Prodromidis</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Chan</surname>
          </string-name>
          , and
          <string-name>
            <given-names>S.</given-names>
            <surname>Stolfo</surname>
          </string-name>
          .
          <article-title>Meta-learning in distributed data mining systems: Issues and approaches</article-title>
          . In
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>