<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Emotion in Music task: lessons learned</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Yi-Hsuan Yang Academia Sinica Taipei</string-name>
          <email>yang@citi.sinica.edu.tw</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Taiwan</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Anna Aljanaki University of Geneva Geneva</institution>
          ,
          <country country="CH">Switzerland</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Mohammad Soleymani University of Geneva Geneva</institution>
          ,
          <country country="CH">Switzerland</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2016</year>
      </pub-date>
      <fpage>20</fpage>
      <lpage>21</lpage>
      <abstract>
        <p>The Emotion in Music task was organized within MediaEval benchmarking campaign during three consecutive years, from 2013 to 2015. In this paper we describe the challenges we faced and the solutions we found. We used crowdsourcing on Amazon Mechanical Turk to annotate a corpus of music pieces with continuous (per-second) emotion annotations. To assure su cient quality of the data, the annotation process on Mechanical Turk requires su cient attention. Labeling music with emotion continuously proved to be a very di cult task for listeners, where both time delay and demand for absolute ratings degraded the quality of the data. We suggest certain transformations to alleviate the problems. Finally, the length of the annotated segments (0.5-1s) led to task participants classifying music on the equally small time scale, which only allowed them to capture changes in dynamics and timbre, but not musically meaningful harmonic, melodic and other changes, which occur on a larger time scale. LSTM-RNN based methods, which allow to incorporate larger context, gave better results than other methods, but still the proposed methods did not show signi cant improvement over the years and the task was concluded.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. INTRODUCTION</title>
      <p>
        The Emotion in Music task was rst proposed by M.
Soleymani, M.N. Caro, E.M. Schmidt and Y.-H. Yang in 2013
[
        <xref ref-type="bibr" rid="ref9">9</xref>
        ]. The task was targeting music emotion recognition
algorithms for music indexing and recommendation,
predominantly for popular music. The most common paradigm for
music retrieval by emotion is the one when emotion is
assigned to an entire piece of music. However, a piece of music
exists in time and assigning just one emotion to a piece of
music is most of the time incorrect. Therefore, music
excerpts were annotated continuously using a paradigm that
was rst suggested for studying dynamics and general
emotionality in music | Continouos Response Digital Interface
(CRDI) [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]. The CRDI was later adapted by E. Schubert to
record emotional response on valence and arousal scale [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ].
In the rst year of the task, static and dynamic tasks existed
side by side. However, later the static task was dropped as
less challenging. The decision to focus on continuous
tracking on emotion had both pros and cons. On the bright side,
it made the Emotion in Music benchmark very distinct from
the existent Mood Classi cation task at MIREX1
benchmark. The continuous emotion recognition is also arguably
less researched than static emotion recognition. However,
the pragmatic utilitarian aspect of the task valued in the
MediaEval community became less prominent. There are
much less applications and much less interest (at least
currently) for automatic recognition of emotion varying over
time.
2.
      </p>
    </sec>
    <sec id="sec-2">
      <title>CROWDSOURCING THE DATA</title>
      <p>Music annotation with emotion is a time-consuming task,
which often generates very inconsistent responses even with
conscientious annotators. Therefore, it is di cult to
verify the responses, because many inconsistencies can be
attributed to individual perception. We used crowdsourcing
(on Amazon Mechanical Turk (AMT)) to annotate the data,
we paid the workers to annotate the music, and the workers
had to pass a test before being admitted to the task. In the
rst 2 years, we did not monitor the quality of the work after
the test was passed. We tried to estimate the lower bound of
the number low-quality workers by only counting the people
who did not move their mouse at all when annotating. Some
of the songs may not have any emotional change in them,
but at least some initial movement from the start position is
necessary before stabilizing. In year 2014, 99 annotators
annotated 1000 pieces of music, of them only 2 people did not
move their mouse at all, and they annotated only a small
amount of songs.</p>
      <p>
        However, the agreement between the annotators was not
very good both in 2013 and 2014 (less than 0.3 in Cronbach's
). In 2015, we changed the procedure to a more lab-like
setting by hiring 5 annotators to annotate all the dataset,
half of them in the lab and half on the AMT [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. The
quality was much better. This could also be attributed to the
other changes, such as choosing full length songs, choosing
the best annotators of the previous years, negotiating a fare
compensation in advance on a Turker forum2 and
introducing a preliminary listening stage.
      </p>
      <p>Despite having highly quali ed annotators, the following
problems were not resolved:
1. Absolute scale ratings. The ratings had to be given
on an absolute scale while estimating changes in arousal
and valence. Though the annotators often agreed on
the direction of change, the magnitude of change was
often di erent. We suggest shifting the annotation so
1http://www.music-ir.org/mirex
2http://www.mturkgrind.com/</p>
      <p>
        Method
2013, BLSTM-RNN [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ] :31
2014, LSTM [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ] :35
2015, BLSTM-RNN [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ] :66
Method
2013, BLSTM-RNN [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ] :19
2014, LSTM [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ] :20
2015, BLSTM-RNN [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ] :17
that its mean is where the mean was indicated by the
annotator (beforehand).
2. Human annotators have a reaction time. The biggest
time lag is observed in the beginning of the annotation
(around 13 seconds), but after every musical change a
small time lag is also present. The beginnings of the
annotations had to be deleted as unreliable.
3. The time scale. Some of the annotators would react
to every beat and every note, and some annotators
would only consider changing their arousal or valence
at section bounds.
      </p>
    </sec>
    <sec id="sec-3">
      <title>3. SUGGESTED SOLUTIONS</title>
      <p>Participants received the data as a sequence of features
and annotations with either half a second or one second
frame rate. Many participants extracted their own features,
but almost always the windows for feature extraction were
smaller than the 0.5-1s, and the features were very low-level,
mostly relating to timbral properties of the sound (energy
distribution across spectrum) and loudness.</p>
      <p>
        The best solutions in all the years were obtained using
Long-Short Term Memory Recurrent Neural Networks.
Although the arousal prediction performance improved over
the three years (see Table 1), the accuracy obtained when
predicting valence did not improve much (Table 2). It is a
known issue that modeling valence in music is more
challenging both due to the higher subjectivity associated with
valence perception and in part due to the absence of salient
valence-related audio features that can be reliably computed
by state-of-the-art music signal processing algorithms [
        <xref ref-type="bibr" rid="ref12 ref4 ref5">12, 5,
4</xref>
        ]. The almost twofold improvement in arousal can also be
attributed to the improvement in the quality and
consistency of the data. In year 2015, the situation with valence
became even worse, because we invested extra e ort to
assemble the data set in such a way, that valence and arousal
would not be correlated (by picking more songs from the
upper left (\angry") and lower right (\serene") quadrants).
Because of the di erence in the development and evaluation
sets' distributions, the evaluation results were inaccurate in
2015. We trained an LSTM-RNN with the features
supplied by the participants and evaluation set data. Using
20-fold cross-validation, we obtained more accurate
estimation of the state-of-the-art performance on valence. The
best result for valence detection on the test-set of 2015 was
achieved using JUNLP team's features ( = :27 :13 and
RM SE = :19 :35) [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ]. JUNLP team used feature reduction
to nd the features which were most important for valence.
However, the result is still much worse than the one obtained
for arousal. A very interesting nding was that even though
some sophisticated procedures for feature dimensionality
reduction and BLSTM-RNNs were suggested by the
participants, an almost equally good result could be obtained for
arousal by using just eight low-level timbral features, and
linear regression with smoothing.
      </p>
    </sec>
    <sec id="sec-4">
      <title>FUTURE OF THE TASK</title>
      <p>The major problem that we encountered when organizing
the task was assembling good quality data. The
improvement in performance over the years was partly dependent
on that. The problems arising when using the continuous
response annotation interface seem to be unsolvable unless
either the task or the interface change.</p>
      <p>One of the possible solutions is to change the
underlying task. It seems that the algorithms developed by the
teams can track musical dynamics rather well. Many
expressive means in music are characterized by gradual changes
(e.g., diminuendo, crescendo, rallentando). Tracking these
changes in tempo and dynamics could be useful as a
preliminary step to tracking emotional changes. Changes in timbre
can also be tracked in a similar way on a very short time
scale.</p>
      <p>
        Another possibility is changing the interface. One of the
alternative continuous annotation interfaces suggested by E.
Schubert uses categorical model instead of a dimensional one
[
        <xref ref-type="bibr" rid="ref8">8</xref>
        ]. Using categorical model would eliminate the problem
with absolute scaling.
      </p>
      <p>A more sophisticated interface could also allow to modify
the annotation afterwards by changing the scale (squeezing
or expanding), removing the time lags.</p>
      <p>At last, one of the major questions with the continuous
emotion tracking task is its practical applicability. In most
cases, the estimation of the overall emotion of the song,
or the most representative part of a song, is most useful
to users. Retrieval by continuous emotion tracking could
be useful when a song with a certain emotional trajectory
is necessary, for instance, for production music or
soundtracks. Another possible application would be nding the
most dramatic (emotionally charged) moment in a song to
be used as a snippet. Moreover, as music is often used as
a stimulus in the a ective computing community to study
emotion prediction by brain waves or physiological signals,
a model to predict dynamic emotion in music can be helpful
in this research. Departing from such bottom-up needs and
requirements, hopefully the problem could be reformulated
in a better way.
5.</p>
    </sec>
    <sec id="sec-5">
      <title>ACKNOWLEDGMENTS</title>
      <p>We would like to thank Erik M. Schmidt, Mike N. Caro,
Cheng-Ya Sha, Alexander Lansky, Sung-Yen Liu and
Eduardo Countinho for their contributions in the development
of this task. We also thank all the task participants and
anonymous turkers for their invaluable contributions.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>A.</given-names>
            <surname>Aljanaki</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.-H.</given-names>
            <surname>Yang</surname>
          </string-name>
          , and
          <string-name>
            <given-names>M.</given-names>
            <surname>Soleymani</surname>
          </string-name>
          .
          <article-title>Emotion in music task at mediaeval 2015</article-title>
          .
          <source>In Working Notes Proceedings of the MediaEval 2015 Workshop</source>
          ,
          <year>2015</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>E.</given-names>
            <surname>Coutinho</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Weninger</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Schuller</surname>
          </string-name>
          , and
          <string-name>
            <given-names>K. R.</given-names>
            <surname>Scherer</surname>
          </string-name>
          .
          <article-title>The munich lstm-rnn approach to the mediaeval 2014 emotion in music task</article-title>
          .
          <source>In Working Notes Proceedings of the MediaEval 2014 Workshop</source>
          ,
          <year>2014</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>D.</given-names>
            <surname>Gregory</surname>
          </string-name>
          .
          <article-title>Using computers to measure continuous music responses</article-title>
          .
          <source>Psychomusicology: A Journal of Research in Music Cognition</source>
          ,
          <volume>8</volume>
          (
          <issue>2</issue>
          ):
          <volume>127</volume>
          {
          <fpage>134</fpage>
          ,
          <year>1989</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>D.</given-names>
            <surname>Guan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Chen</surname>
          </string-name>
          , and
          <string-name>
            <given-names>D.</given-names>
            <surname>Yang</surname>
          </string-name>
          .
          <article-title>Music Emotion Regression Based on Multi-modal Features</article-title>
          .
          <source>In Symposium on Computer Music Multidisciplinary Research</source>
          , pages
          <volume>70</volume>
          {
          <fpage>77</fpage>
          ,
          <year>2012</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>C.</given-names>
            <surname>Laurier</surname>
          </string-name>
          ,
          <string-name>
            <given-names>O.</given-names>
            <surname>Lartillot</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Eerola</surname>
          </string-name>
          , and
          <string-name>
            <given-names>P.</given-names>
            <surname>Toiviainen</surname>
          </string-name>
          .
          <article-title>Exploring Relationships between Audio Features and Emotion in Music</article-title>
          .
          <source>In Proceedings of the 7th Triennal Conference of European Society for Cognitive Sciences of Music</source>
          , pages
          <volume>260</volume>
          {
          <fpage>264</fpage>
          ,
          <year>2009</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>B. G.</given-names>
            <surname>Patra</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Maitra</surname>
          </string-name>
          ,
          <string-name>
            <surname>D. Das</surname>
            , and
            <given-names>S.</given-names>
          </string-name>
          <string-name>
            <surname>Bandyopadhyay</surname>
          </string-name>
          .
          <source>Mediaeval</source>
          <year>2015</year>
          :
          <article-title>Music emotion recognition based on feed-forward neural network</article-title>
          .
          <source>In Working Notes Proceedings of the MediaEval 2015 Workshop</source>
          ,
          <year>2015</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>E.</given-names>
            <surname>Schubert</surname>
          </string-name>
          .
          <article-title>Continuous response to music using a two dimensional emotion space</article-title>
          .
          <source>In Proceedings of International Conference of Music Perception and Cognition</source>
          , pages
          <volume>263</volume>
          {
          <fpage>268</fpage>
          ,
          <year>1996</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>E.</given-names>
            <surname>Schubert</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Ferguson</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Farrar</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Taylor</surname>
          </string-name>
          , and
          <string-name>
            <given-names>G. E.</given-names>
            <surname>McPherson</surname>
          </string-name>
          .
          <article-title>Continuous Response to Music using Discrete Emotion Faces</article-title>
          .
          <source>In Proceedings of the 9th international symposium on computer music modeling and retrieval</source>
          ,
          <year>2012</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>M.</given-names>
            <surname>Soleymani</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. N.</given-names>
            <surname>Caro</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E. M.</given-names>
            <surname>Schmidt</surname>
          </string-name>
          , and
          <string-name>
            <given-names>Y.-H.</given-names>
            <surname>Yang</surname>
          </string-name>
          .
          <article-title>The mediaeval 2013 brave new task: Emotion in music</article-title>
          .
          <source>In Working Notes Proceedings of the MediaEval 2013 Workshop</source>
          ,
          <year>2013</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>F.</given-names>
            <surname>Weninger</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Eyben</surname>
          </string-name>
          , and
          <string-name>
            <given-names>B.</given-names>
            <surname>Schuller</surname>
          </string-name>
          .
          <article-title>The TUM approach to the mediaeval music emotion task using generic a ective audio features</article-title>
          .
          <source>In Working Notes Proceedings of the MediaEval 2013 Workshop</source>
          ,
          <year>2013</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>M.</given-names>
            <surname>Xu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Xianyu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Tian</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Meng</surname>
          </string-name>
          , and
          <string-name>
            <given-names>W.</given-names>
            <surname>Chen</surname>
          </string-name>
          <article-title>. Multi-scale approaches to the mediaeval 2015 \emotion in music" task</article-title>
          .
          <source>In Working Notes Proceedings of the MediaEval 2015 Workshop</source>
          ,
          <year>2015</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>Y.-H.</given-names>
            <surname>Yang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.-C.</given-names>
            <surname>Lin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.-F.</given-names>
            <surname>Su</surname>
          </string-name>
          , and
          <string-name>
            <given-names>H. H.</given-names>
            <surname>Chen</surname>
          </string-name>
          .
          <article-title>A Regression Approach to Music Emotion Recognition</article-title>
          .
          <source>IEEE Transactions on Audio, Speech, and Language Processing</source>
          ,
          <volume>16</volume>
          (
          <issue>2</issue>
          ):
          <volume>448</volume>
          {
          <fpage>457</fpage>
          ,
          <year>2008</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>