<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Music emotion recognition using Gaussian Processes</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Konstantin Markov, Motofumi Iwata</string-name>
          <email>markov@u-aizu.ac.jp</email>
          <email>{markov,s1180127}@u-aizu.ac.jp</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Tomoko Matsui</string-name>
          <email>tmatsui@ism.ac.jp</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Department of Statistical Modeling, Institute of Statistical Mathematics</institution>
          ,
          <addr-line>Tokyo</addr-line>
          ,
          <country country="JP">Japan</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Human Interface Laboratory, The University of Aizu</institution>
          ,
          <addr-line>Fukushima</addr-line>
          ,
          <country country="JP">Japan</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2013</year>
      </pub-date>
      <fpage>18</fpage>
      <lpage>19</lpage>
      <abstract>
        <p>This paper describes the music emotion recognition system developed at the University of Aizu for the Emotion in Music task of the MediaEval'2013 benchmark evaluation campaign. A set of standard feature types provided by the Marsyas toolkit was used to parametrize each music clip. Arousal and valence are modeled separately using Gaussian Process regression (GPR). We compared performances of the GPR and Support Vector regression (SVR) and found out that GPR gives better results than SVR for the static per song emotion estimation task. For the dynamic emotion estimation task GPR had some scalability problems and fair comparison was not possible.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. INTRODUCTION</title>
      <p>
        Gaussian Processes (GPs) [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ] are becoming more and more
popular in the Machine Learning community for their ability
to learn highly non-linear mappings between two continuous
data spaces, i.e. the feature space and the V/A space.
Previously, we have successfully applied GPs for music genre
classi cation task [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ] and this encouraged us to use GPs for
music emotion estimation. Many previous studies [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ] have
focused on Support Vector regression (SVR) since in most
cases it gives superior performance. In this study we
compare GP regression with SVR and show that in certain cases
GPR can signi cantly outperform SVR. In addition, GPR
produces probabilistic predictions, i.e. it outputs a Gaussian
distribution with mean which corresponds to the most
probable target value and variance which shows the certainty of
the prediction. As in the case of SVR, GPR also uses kernels,
but in contrast, it allows kernels parameters to be learned
from the training data.
      </p>
      <p>
        Database used in this evaluation is described in detail in
the Emotion in Music task overview paper [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ].
      </p>
    </sec>
    <sec id="sec-2">
      <title>GAUSSIAN PROCESS REGRESSION</title>
      <p>3.</p>
    </sec>
    <sec id="sec-3">
      <title>SYSTEM DESCRIPTION</title>
      <p>Dimensional music emotion recognition can be easily
decomposed into two independent classical regression
problems: one for the valence, and another for the arousal. Thus,
our system consists of two regression modules and a common
feature extraction module.
3.1</p>
    </sec>
    <sec id="sec-4">
      <title>Feature extraction</title>
      <p>Features are extracted only from the audio signal which is
rst downsampled to 22050 kHz. We tried various standard
features tailored for music processing such as MFCC,
Statistical Spectrum Descriptors (SSD), Chroma, Spectral Crest
Factor (SCF), and Spectral Flatness Measure (SFM)
separately as well as combinations of several of them. All feature
vectors were calculated using the Marsyas toolkit with 512
samples frames with no overlap. For the dynamic emotion
estimation task, rst order statistics (mean and std) of the
feature vectors are calculated for a window of about 1 sec.
giving 45 vectors per musical clip. For the static emotion
estimation, same statistics for these 45 vectors are
calculated resulting in a single high dimensional feature vector
per song. After extensive preliminary experimentation we
found that the best performing combination of features for
the per song emotion estimation is MFCC, SCF, and SFM.
Adding SSD features did not have any noticeable e ect, and
Chroma features actually hurt the performance. We refer
to this combination of features as UoA features. We have
also experimented with features released by the benchmark
organisers which we call MediaEval features.
3.2</p>
    </sec>
    <sec id="sec-5">
      <title>GPR implementation</title>
      <p>Valence and arousal are modeled by separate GPR. We
used standard Gaussian likelihood function which allows
exact inference to be performed. The GP mean was set to
zero and only the type of covariance kernel was varied. We
experimented with the following kernels:</p>
      <p>Linear (LIN): k(x; x0) = (xT x0 + 1)=l2
Squared Exponential (SE): k(x; x0) =
x0)T (x x0)=2l2)
Rational Quadratic (RQ): k(x; x0) =
x0)=2 l2)
2(1+(x x0)T (x
Matern 3 (MAT3): k(x; x0) =
p3(x x0)T (x x0)=l2)
2(1 + r) exp( r); r =
where and l are parameters learned from training data.
Sums or products of several kernels are also valid covariance
functions and often give better performance than single
kernels.</p>
    </sec>
    <sec id="sec-6">
      <title>RESULTS</title>
      <p>First, we present our results on the development data
obtained after 7-fold cross validation. In addition to GPR,
results from SVR using the same conditions are given in
the following tables. In the SVR case, the parameter C
was manually optimized using a grid search in the range
[0.01,100], and kernel parameters are set to their default
values (using LIBSVM package) since they cannot be learned.</p>
      <p>Table 1 shows the result for the static emotion estimation
in terms of R2 metrics for both MediaEval and UoA feature
sets. The last row of each feature set type shows the best
performing combination of GPR covariance kernels.</p>
      <p>Kernel Valence
MediaEval features</p>
      <p>Linear 0.112
RBF 0.017
LIN 0.132
SE 0.142</p>
      <p>RQ 0.150</p>
      <p>MAT3 0.143
LIN+RQ 0.170</p>
      <p>UoA features
Linear 0.314
RBF 0.367
LIN 0.322
SE 0.375</p>
      <p>RQ 0.430
MAT3 0.395
SExRQ 0.437</p>
      <p>Arousal
0.300
0.028
0.565
0.590
0.562
0.590
0.581
0.604
0.653
0.603
0.656
0.662
0.668
0.671
Algorithm</p>
      <p>SVR
GPR
SVR</p>
      <p>GPR</p>
      <p>In Table 2, we summarize results of the dynamic emotion
estimation task where Kendal measure is calculated
after pooling all arousal or valence estimates from all songs
together. We have to mention that since in this task the
amount of data was 40 times bigger, we ran into some
scalability problems with the GPR implementation and had to
resort to approximations of the kernel matrix using much
less data which, of course, decreased the performance
noticeably.</p>
      <p>Table 3 presents the results of the UoA submission runs:
two for the static and one for the dynamic emotion
estimation tasks. They are obtained using GPR with
corresponding best performing kernels. Direct comparison with Tables
1 and 2 is possible only for RSQ lines and it can be seen that
in contrast to MediaEval, UoA features give similar results.
cial results on the test data obtained
5.</p>
    </sec>
    <sec id="sec-7">
      <title>CONCLUSIONS</title>
      <p>We described the UoA emotion recognition system for the
"Emotion in Music" task of the MediaEval'2013 benchmark
evaluation which is based on the Gaussian Process
regression algorithm. Compared to the Support Vector
regression, GPR has several advantages, such as truly probabilistic
prediction, and ability to learn hyperparameters from data.
Performance wise, the GPR achieved better results for the
static per song emotion estimation, but failed for the
dynamic emotion estimation due to some scalability problems.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>K.</given-names>
            <surname>Markov</surname>
          </string-name>
          and
          <string-name>
            <given-names>T.</given-names>
            <surname>Matsui</surname>
          </string-name>
          .
          <article-title>Music genre classi cation using gaussian process models</article-title>
          .
          <source>In Proc. IEEE Workshop on Machine Learning for Signal Processing (MLSP)</source>
          ,
          <year>2013</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>C.</given-names>
            <surname>Rasmussen</surname>
          </string-name>
          and
          <string-name>
            <given-names>C.</given-names>
            <surname>Williams</surname>
          </string-name>
          .
          <article-title>Gaussian Processes for Machine Learning</article-title>
          .
          <source>Adaptive Computation and Machine Learning</source>
          . The MIT Press,
          <year>2006</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>M.</given-names>
            <surname>Soleymani</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Caro</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E. M.</given-names>
            <surname>Schmidt</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Sha</surname>
          </string-name>
          , and
          <string-name>
            <given-names>Y.</given-names>
            <surname>Yang</surname>
          </string-name>
          .
          <article-title>1000 songs for emotional analysis of music</article-title>
          .
          <source>In Proceedings of the ACM multimedia 2013 workshop on Crowdsourcing for Multimedia</source>
          ,
          <source>CrowdMM. ACM</source>
          , ACM,
          <year>2013</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>Y.-H.</given-names>
            <surname>Yang</surname>
          </string-name>
          and
          <string-name>
            <given-names>H.</given-names>
            <surname>Chen</surname>
          </string-name>
          .
          <article-title>Machine recognition of music emotion: A review</article-title>
          .
          <source>ACM Transactions on Intelligent Systems and Technology</source>
          ,
          <volume>3</volume>
          (
          <issue>3</issue>
          ):
          <volume>40</volume>
          :1{
          <fpage>40</fpage>
          :
          <fpage>30</fpage>
          , May
          <year>2012</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>