<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Emotion and Themes Recognition in Music with Convolutional and Recurrent Atention-Blocks</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Maurice Gerczuk</string-name>
          <email>maurice.gerczuk@uni-a.de</email>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Shahin Amiriparian</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Sandra Ottl</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Srividya Tirunellai Rajamani</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Björn Schuller</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Chair of Embedded Intelligence for Health Care</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Wellbeing</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Univeristy of Augsburg</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Germany</string-name>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>GLAM - Group on Language, Audio, &amp; Music, Imperial College London</institution>
          ,
          <country country="UK">U. K</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2020</year>
      </pub-date>
      <fpage>14</fpage>
      <lpage>15</lpage>
      <abstract>
        <p>Emotion is an essential aspect of music, and its recognition is a prevalent research topic in the field of computer audition. Machine learning-based Music Emotion Recognition (MER) systems could boost the accessibility of music collections by providing standardised methodologies of music categorisation. In this paper, we introduce our (team name: AugsBurger) machine learning architecture sequentially composed of a convolutional feature extractor with block attention modules and a recurrent stack with self-attention for automatic MER. We train 5 models and conduct various late fusion experiments. Utilising a Convolutional Recurrent Neural Network (CRNN) with convolutional block attention applied throughout a 18-layer ResNet and a single recurrent layer with a Gated Recurrent Unit cell, a ROC-AUC of 73.9 % can be achieved on the test partition of the MediaEval 2020 Emotion &amp; Themes in Music task. Applying late fusion on the individual model predictions and another challenge submission, this result is further increased to 75.3 % ROC-AUC.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>INTRODUCTION</title>
      <p>
        The ability of music to express emotions is a demonstrable and
eminent fact [
        <xref ref-type="bibr" rid="ref18">18</xref>
        ]. Emotional experiences of music are complex
and dependent on factors related to the states and traits of the
listener, the performer, and the listening context, with research
suggesting that musical structure alone is a key determinant of
the emotional indication of music [
        <xref ref-type="bibr" rid="ref25">25</xref>
        ]. Diferent music emotion
categories can induce emotional states, such as happiness, sadness,
hope, excitement, and joy in listeners [
        <xref ref-type="bibr" rid="ref16 ref19">16, 19</xref>
        ]. This is primarily
due to the afective information encoded in musical parameters,
including melody, timbre, rhythm, and dynamics which are
implicitly decoded by listeners [
        <xref ref-type="bibr" rid="ref13 ref17">13, 17</xref>
        ]. Conventional feature extraction
methods (e. g., openSMILE [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ]) have shown their suitability to
extract such features from music recordings [
        <xref ref-type="bibr" rid="ref28 ref32">28, 32</xref>
        ]. However, the
state-of-the-art for MER is defined by contemporary machine
learning approaches which utilise convolutional and recurrent neural
networks and learn data representations directly from the audio
signals (or spectrograms) instead of extraction of pre-defined
handcrafted features [
        <xref ref-type="bibr" rid="ref20 ref21 ref29 ref3">3, 20, 21, 29</xref>
        ]. Moreover, the integration of an
attention mechanism in such systems has shown promise for
various audio recognition tasks [
        <xref ref-type="bibr" rid="ref26 ref6">6, 26</xref>
        ]. Motivated by our previous
works with CRNNs [
        <xref ref-type="bibr" rid="ref2 ref3 ref5">2, 3, 5</xref>
        ] and the success of attention
mechanisms [
        <xref ref-type="bibr" rid="ref26 ref31 ref4 ref6">4, 6, 26, 31</xref>
        ], in this paper, we introduce an end-to-end
framework composed of two attention blocks: a Convolutional
Neural Network (CNN) with Convolutional Block Attention
Modules (CBAMs), and a recurrent block with self-attention for the task
of emotion and theme recognition in music [
        <xref ref-type="bibr" rid="ref10 ref8 ref9">8–10</xref>
        ].
2
      </p>
    </sec>
    <sec id="sec-2">
      <title>APPROACH</title>
      <p>
        A high-level overview of our approach is depicted in Figure 1. The
framework consists of a CNN feature extractor enhanced by CBAMs
and an RNN with self-attention. The convolutional block aims to
learn high-level shift-invariant features, whilst long(er)-term
temporal dependencies of music data [
        <xref ref-type="bibr" rid="ref10 ref8 ref9">8–10</xref>
        ] are mainly extracted by
the recurrent block [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. For all experiments in this paper, the
MTGJamendo dataset [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ] is solely used. The 18 486 audio tracks of
the MTG-Jamendo dataset [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ] are annotated in 56 distinct mood
and theme categories, with every track having at least one tag.
The dataset provides 60-20-20 % splits for training, validation, and
testing. A full description of the dataset can be found in [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ].
2.1
      </p>
    </sec>
    <sec id="sec-3">
      <title>Pre-Processing and Augmentation</title>
      <p>
        Our model uses the pre-computed mel-spectrograms that are part
of the challenge dataset. Furthermore, we only use random
windows of 8 seconds (500 timesteps) during training, reducing the
memory footprint of our models and also serving as a form of data
augmentation. Additionally, we apply SpecAugment [
        <xref ref-type="bibr" rid="ref22">22</xref>
        ], randomly
applying masks of a maximum width of 10 (timesteps or frequency
bands) to both the frequency and time domains of the spectrograms.
We do not, however, use warping. During validation and testing,
we take an 8 second chunk from the middle of each spectrogram
and do not apply SpecAugment.
2.2
In the CNN part of our modes, we use CBAMs [
        <xref ref-type="bibr" rid="ref30">30</xref>
        ] to refine the
learn feature maps. CBAMs sequentially apply channel and spatial
attention to the max and average pooled outputs of a convolutional
layer. The attention maps are applied by element-wise
multiplication. As these modules are a very lightweight extension, and
show consistent performance increases for a wide range of
popular image-recognition benchmarks [
        <xref ref-type="bibr" rid="ref30">30</xref>
        ], we evaluate their eficacy
when added to a CRNN for music emotion and theme recognition.
2.3
      </p>
    </sec>
    <sec id="sec-4">
      <title>Recurrent Self-Attention</title>
      <p>
        In the RNN head of our CRNN framework, we use an attention
mechanism to help the model focus on important parts of the feature
sequences extracted by the CNN. This is done by applying
selfattention [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ] to the RNN outputs and states at each time step,
      </p>
      <p>Raw Audios</p>
      <p>Melspectrograms</p>
      <p>
        SpecAugment
input
input
featuirnepsut
featuirnepsut
features
features
ifnally forming a compact representation for the whole sequence.
We use the scaled dot product attention of Vaswani et al. [
        <xref ref-type="bibr" rid="ref27">27</xref>
        ].
2.4
      </p>
    </sec>
    <sec id="sec-5">
      <title>Attention CRNN</title>
      <p>Combining a CNN feature extractor with CBAMs and an RNN head
with self-attention leads to our final attention CRNN. Specifically,
we use an 18-layer ResNet architecture and replace the global
pooling layer with an RNN stack. In the ResNet, we apply CBAMs in
every convolutional block right before adding the residual. For the
RNN, we evaluate using Gated Recurrent Unit (GRU) and Long
Short-Term Memory (LSTM) cells. We started with a single layer
with 256 units for both cell types. As we found GRUs to perform
better, we additionally trained a model containing two recurrent
layers of this type with 128 units each. Finally, the step-wise outputs
and the final hidden state are used in the self-attention mechanism
as keys and values, and query, respectively. Finally, a fully
connected layer with sigmoid activation is used to perform the theme
and mood tagging of the input audio samples. We train all of our
models for a maximum of 100 epochs with an Adam optimiser with
the learning rate set to 0.0003 but stop the training early if the
validation ROC-AUC does not improve for 20 epochs. We use the
weights from the best epoch (measured in validation ROC-AUC)
for evaluation on the held-out test set.
2.5</p>
    </sec>
    <sec id="sec-6">
      <title>Fusion Experiments</title>
      <p>
        We apply late fusion to the results achieved by our attention CRNNs
(models with CBAMs) through averaging prediction scores.
Furthermore, we fuse our predictions with another system submitted to the
challenge [
        <xref ref-type="bibr" rid="ref23">23</xref>
        ] which uses standalone self-attention an
attentionbased Rectified Linear Units ( ReLUs) [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ] added to the challenge
baseline’s vggish model [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ].
3
      </p>
    </sec>
    <sec id="sec-7">
      <title>RESULTS AND ANALYSIS</title>
      <p>
        The results of our experiments are shown in Table 1. The best
individual model can be found with a CBAM enhanced CRNN with
a single GRU layer. This model reaches 73.9 % ROC-AUC on the
test partition, compared to the challenge baseline of 72.5 %. Using
the same architecture but without applying CBAMs, only 69.4 %
are achieved on test. Furthermore, we observe that models with a
single GRUs layer outperform their LSTM counterparts. Fusing the
two best CRNNs – CBAMs in the CNN and GRU cells in the RNN –
further leads to a slight performance boost to 74.1 %. More efective
is fusing with the attention enhanced CNN from [
        <xref ref-type="bibr" rid="ref23">23</xref>
        ], achieving
our best result of 75.3 % ROC-AUC and 13.1 % PR-AUC on test. This
hints at complementarity of the two systems.
      </p>
      <p>A noteworthy characteristic of all the systems used in our
submission to the challenge is that they do not make use of any external
data and only train on short extracts of the songs (about 10 seconds
long). Furthermore, none of the models were trained for more than
50 epochs, against the challenge baseline’s 1 000 epochs. In this way,
our attention models reduce data, memory and time requirements
while achieving stronger performance than the baseline. Compared
to the other challenge submissions, only one submission that does
not rely on external data outperforms our best fusion model1.
4</p>
    </sec>
    <sec id="sec-8">
      <title>DISCUSSION AND OUTLOOK</title>
      <p>
        We have introduced a CRNN architecture with attention modules
for both convolutional (cf. Section 2.2) and recurrent blocks (cf.
Section 2.3) for emotions and themes recognition in music. In the
pre-processing step, in order to achieve a better model
generalisation, we have augmented the spectrograms from the music
recordings with SpecAugment [
        <xref ref-type="bibr" rid="ref22">22</xref>
        ] and trained our models with both
challenge and augmented data (cf. Section 2.1). Furthermore, as a
post-processing step, we have conducted a set of late (decision-level)
fusion experiments to check the complementary of the predictions
from each trained model (cf. Section 2.5). The results indicate the
eficacy of our applied methodologies for this challenge(cf.
Section 3). Considering the performance increase achieved by
utilising attention-based ReLUs with the baseline’s vggish architecture
in [
        <xref ref-type="bibr" rid="ref23">23</xref>
        ], it is worth investigating this activation mechanism in
combination with the CBAM enhanced CRNNs presented herein. As
our models only make use of the challenge data itself, one should
also consider exploiting external data, such as the Million Song
Dataset [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ], Music4All [
        <xref ref-type="bibr" rid="ref24">24</xref>
        ] or NSynth [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ] for possible
improvements in model accuracy.
1https://multimediaeval.github.io/2020-Emotion-and-Theme-Recognition-in-MusicTask/results
Emotion and Themes in Music
      </p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>Shahin</given-names>
            <surname>Amiriparian</surname>
          </string-name>
          .
          <year>2019</year>
          .
          <article-title>Deep representation learning techniques for audio signal processing</article-title>
          .
          <source>Ph.D. Dissertation</source>
          . Technische Universität München.
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>Shahin</given-names>
            <surname>Amiriparian</surname>
          </string-name>
          , Alice Baird, Sahib Julka, Alyssa Alcorn, Sandra Ottl, Suncica Petrović, Eloise Ainger, Nicholas Cummins,
          <string-name>
            <given-names>and Björn</given-names>
            <surname>Schuller</surname>
          </string-name>
          .
          <year>2018</year>
          .
          <article-title>Recognition of Echolalic Autistic Child Vocalisations Utilising Convolutional Recurrent Neural Networks</article-title>
          .
          <source>In Proceedings of INTERSPEECH</source>
          <year>2018</year>
          ,
          <article-title>19th Annual Conference of the International Speech Communication Association</article-title>
          . ISCA, Hyderabad, India,
          <fpage>2334</fpage>
          -
          <lpage>2338</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>Shahin</given-names>
            <surname>Amiriparian</surname>
          </string-name>
          , Maurice Gerczuk, Eduardo Coutinho, Alice Baird, Sandra Ottl, Manuel Milling, and
          <string-name>
            <given-names>Björn</given-names>
            <surname>Schuller</surname>
          </string-name>
          .
          <year>2019</year>
          .
          <article-title>Emotion and themes recognition in music utilising convolutional and recurrent neural networks</article-title>
          .
          <source>In MediaEval Benchmarking Initiative for Multimedia Evaluation. Sophia Antipolis</source>
          , France.
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>Shahin</given-names>
            <surname>Amiriparian</surname>
          </string-name>
          , Maurice Gerczuk, Sandra Ottl, Alice Baird, Lukas Stappen, Lukas Koebe, and
          <string-name>
            <given-names>Björn</given-names>
            <surname>Schuller</surname>
          </string-name>
          .
          <year>2020</year>
          .
          <article-title>Towards CrossModal Pre-Training and Learning Tempo-Spatial Characteristics for Audio Recognition with Convolutional and Recurrent Neural Networks</article-title>
          .
          <source>EURASIP Journal on Audio, Speech, and Music Processing</source>
          <year>2020</year>
          (
          <year>2020</year>
          ), to appear.
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>Shahin</given-names>
            <surname>Amiriparian</surname>
          </string-name>
          , Sahib Julka, Nicholas Cummins,
          <string-name>
            <given-names>and Björn</given-names>
            <surname>Schuller</surname>
          </string-name>
          .
          <year>2018</year>
          .
          <article-title>Deep Convolutional Recurrent Neural Networks for Rare Sound Event Detection</article-title>
          .
          <source>In Proceedings 44. Jahrestagung für Akustik</source>
          ,
          <string-name>
            <surname>DAGA</surname>
          </string-name>
          <year>2018</year>
          .
          <article-title>DEGA, Deutsche Gesellschaft für Akustik e</article-title>
          .V. (DEGA), Munich, Germany.
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>Shahin</given-names>
            <surname>Amiriparian</surname>
          </string-name>
          , Pawel Winokurow, Vincent Karas, Sandra Ottl, Maurice Gerczuk, and
          <string-name>
            <given-names>Björn</given-names>
            <surname>Schuller</surname>
          </string-name>
          .
          <year>2020</year>
          .
          <article-title>Unsupervised Representation Learning with Attention and Sequence to Sequence Autoencoders to Predict Sleepiness From Speech</article-title>
          .
          <source>In Proceedings of the 1st International on Multimodal Sentiment Analysis in Real-life Media Challenge and Workshop</source>
          . 11-
          <fpage>17</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>Thierry</given-names>
            <surname>Bertin-Mahieux</surname>
          </string-name>
          , Daniel PW Ellis, Brian Whitman, and
          <string-name>
            <given-names>Paul</given-names>
            <surname>Lamere</surname>
          </string-name>
          .
          <year>2011</year>
          .
          <article-title>The million song dataset</article-title>
          . (
          <year>2011</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>Dmitry</given-names>
            <surname>Bogdanov</surname>
          </string-name>
          , Alastair Porter,
          <string-name>
            <given-names>Philip</given-names>
            <surname>Tovstogan</surname>
          </string-name>
          , and
          <string-name>
            <given-names>Minz</given-names>
            <surname>Won</surname>
          </string-name>
          .
          <year>2019</year>
          .
          <article-title>MediaEval 2019: Emotion and Theme Recognition in Music Using Jamendo. In MediaEval Benchmarking Initiative for Multimedia Evaluation</article-title>
          . Sophia Antipolis, France.
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>Dmitry</given-names>
            <surname>Bogdanov</surname>
          </string-name>
          , Alastair Porter,
          <string-name>
            <given-names>Philip</given-names>
            <surname>Tovstogan</surname>
          </string-name>
          , and
          <string-name>
            <given-names>Minz</given-names>
            <surname>Won</surname>
          </string-name>
          .
          <year>2020</year>
          .
          <article-title>MediaEval 2020: Emotion and Theme Recognition in Music Using Jamendo. In MediaEval Benchmarking Initiative for Multimedia Evaluation</article-title>
          . Online.
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <surname>Dmitry</surname>
            <given-names>Bogdanov</given-names>
          </string-name>
          , Minz Won, Philip Tovstogan,
          <string-name>
            <given-names>Alastair</given-names>
            <surname>Porter</surname>
          </string-name>
          , and
          <string-name>
            <given-names>Xavier</given-names>
            <surname>Serra</surname>
          </string-name>
          .
          <year>2019</year>
          .
          <article-title>The MTG-Jamendo Dataset for Automatic Music Tagging</article-title>
          .
          <source>In Machine Learning for Music Discovery Workshop, International Conference on Machine Learning (ICML</source>
          <year>2019</year>
          ).
          <article-title>ICML, Long Beach</article-title>
          , CA, United States.
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>Dengsheng</given-names>
            <surname>Chen</surname>
          </string-name>
          and
          <string-name>
            <given-names>Kai</given-names>
            <surname>Xu</surname>
          </string-name>
          .
          <year>2020</year>
          .
          <article-title>AReLU: Attention-based Rectified Linear Unit</article-title>
          . (
          <year>2020</year>
          ). arXiv:arXiv:
          <year>2006</year>
          .13858
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <surname>Jianpeng</surname>
            <given-names>Cheng</given-names>
          </string-name>
          , Li Dong, and
          <string-name>
            <given-names>Mirella</given-names>
            <surname>Lapata</surname>
          </string-name>
          .
          <year>2016</year>
          .
          <article-title>Long ShortTerm Memory-Networks for Machine Reading</article-title>
          .
          <source>In Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics</source>
          , Austin, Texas,
          <fpage>551</fpage>
          -
          <lpage>561</lpage>
          . https://doi.org/10.18653/v1/
          <fpage>D16</fpage>
          -1053
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <given-names>Eduardo</given-names>
            <surname>Coutinho</surname>
          </string-name>
          and
          <string-name>
            <given-names>Björn</given-names>
            <surname>Schuller</surname>
          </string-name>
          .
          <year>2017</year>
          .
          <article-title>Shared acoustic codes underlie emotional communication in music and speech-Evidence from deep transfer learning</article-title>
          .
          <source>PloS one 12</source>
          ,
          <issue>6</issue>
          (
          <year>2017</year>
          ),
          <year>e0179289</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <surname>Jesse</surname>
            <given-names>Engel</given-names>
          </string-name>
          , Cinjon Resnick, Adam Roberts, Sander Dieleman, Mohammad Norouzi, Douglas Eck, and
          <string-name>
            <given-names>Karen</given-names>
            <surname>Simonyan</surname>
          </string-name>
          .
          <year>2017</year>
          .
          <article-title>Neural audio synthesis of musical notes with wavenet autoencoders</article-title>
          .
          <source>In International Conference on Machine Learning. PMLR</source>
          ,
          <fpage>1068</fpage>
          -
          <lpage>1077</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <surname>Florian</surname>
            <given-names>Eyben</given-names>
          </string-name>
          ,
          <string-name>
            <given-names>Martin</given-names>
            <surname>Wöllmer</surname>
          </string-name>
          , and
          <string-name>
            <given-names>Björn</given-names>
            <surname>Schuller</surname>
          </string-name>
          .
          <year>2010</year>
          .
          <article-title>Opensmile: the munich versatile and fast open-source audio feature extractor</article-title>
          .
          <source>In Proceedings of the 18th ACM international conference on Multimedia</source>
          .
          <volume>1459</volume>
          -
          <fpage>1462</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [16]
          <string-name>
            <given-names>Bruce</given-names>
            <surname>Ferwerda</surname>
          </string-name>
          and
          <string-name>
            <given-names>Markus</given-names>
            <surname>Schedl</surname>
          </string-name>
          .
          <year>2014</year>
          .
          <article-title>Enhancing Music Recommender Systems with Personality Information and Emotional States: A Proposal.</article-title>
          .
          <source>In Umap workshops.</source>
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [17]
          <string-name>
            <given-names>Alf</given-names>
            <surname>Gabrielsson</surname>
          </string-name>
          and
          <string-name>
            <given-names>Erik</given-names>
            <surname>Lindström</surname>
          </string-name>
          .
          <year>2010</year>
          .
          <article-title>The role of structure in the musical expression of emotions</article-title>
          .
          <source>In Handbook of music and emotion: Theory</source>
          , research, applications,
          <source>Patrik N. Juslin and John Sloboda (Eds.)</source>
          . Oxford University Press, Oxford,
          <fpage>367</fpage>
          -
          <lpage>400</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          [18]
          <string-name>
            <surname>Patrik</surname>
            <given-names>N.</given-names>
          </string-name>
          <string-name>
            <surname>Juslin</surname>
          </string-name>
          and John Sloboda (Eds.).
          <year>2011</year>
          .
          <article-title>Handbook of music and emotion: Theory, research, applications</article-title>
          . Oxford University Press.
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          [19]
          <string-name>
            <surname>Ai</surname>
            <given-names>Kawakami</given-names>
          </string-name>
          , Kiyoshi Furukawa, Kentaro Katahira, and
          <string-name>
            <given-names>Kazuo</given-names>
            <surname>Okanoya</surname>
          </string-name>
          .
          <year>2013</year>
          .
          <article-title>Sad music induces pleasant emotion</article-title>
          .
          <source>Frontiers in psychology 4</source>
          (
          <year>2013</year>
          ),
          <fpage>311</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          [20]
          <string-name>
            <surname>Khaled</surname>
            <given-names>Koutini</given-names>
          </string-name>
          , Shreyan Chowdhury, Verena Haunschmid,
          <article-title>Hamid Eghbal-zadeh, and</article-title>
          <string-name>
            <given-names>Gerhard</given-names>
            <surname>Widmer</surname>
          </string-name>
          .
          <year>2019</year>
          .
          <article-title>Emotion and Theme Recognition in Music with Frequency-Aware RF-Regularized CNNs</article-title>
          . arXiv preprint arXiv:
          <year>1911</year>
          .
          <volume>05833</volume>
          (
          <year>2019</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          [21]
          <string-name>
            <surname>Maximilian</surname>
            <given-names>Mayerl</given-names>
          </string-name>
          , Michael Vötter,
          <string-name>
            <surname>Hsiao-Tzu</surname>
            <given-names>Hung</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bo-Yu</surname>
            <given-names>Chen</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Yi-Hsuan Yang</surname>
            , and
            <given-names>Eva</given-names>
          </string-name>
          <string-name>
            <surname>Zangerle</surname>
          </string-name>
          .
          <year>2019</year>
          .
          <article-title>Recognizing Song Mood and Theme Using Convolutional Recurrent Neural Networks</article-title>
          .
          <article-title>In MediaEval Benchmarking Initiative for Multimedia Evaluation</article-title>
          . Sophia Antipolis, France.
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          [22]
          <string-name>
            <surname>Daniel S Park</surname>
            ,
            <given-names>William</given-names>
          </string-name>
          <string-name>
            <surname>Chan</surname>
          </string-name>
          , Yu Zhang, Chung-Cheng Chiu, Barret Zoph,
          <string-name>
            <surname>Ekin D Cubuk</surname>
          </string-name>
          , and
          <string-name>
            <surname>Quoc</surname>
            <given-names>V</given-names>
          </string-name>
          <string-name>
            <surname>Le</surname>
          </string-name>
          .
          <year>2019</year>
          .
          <article-title>Specaugment: A simple data augmentation method for automatic speech recognition</article-title>
          . arXiv preprint arXiv:
          <year>1904</year>
          .
          <volume>08779</volume>
          (
          <year>2019</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          [23]
          <string-name>
            <given-names>Srividya</given-names>
            <surname>Tirunellai</surname>
          </string-name>
          <string-name>
            <surname>Rajamani</surname>
          </string-name>
          , Kumar Rajamani, and
          <string-name>
            <given-names>Björn</given-names>
            <surname>Schuller</surname>
          </string-name>
          .
          <year>2020</year>
          .
          <article-title>Emotion and Theme Recognition in Music using Attentionbased Methods. In MediaEval Benchmarking Initiative for Multimedia Evaluation</article-title>
          . Online.
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          [24]
          <string-name>
            <given-names>Igor</given-names>
            <surname>André Pegoraro Santana</surname>
          </string-name>
          , Fabio Pinhelli, Juliano Donini, Leonardo Catharin, Rafael Biazus Mangolin, Valéria Delisandra Feltrim, Marcos Aurélio Domingues, and others.
          <year>2020</year>
          .
          <article-title>Music4All: A New Music Database and Its Applications</article-title>
          .
          <source>In 2020 International Conference on Systems, Signals and Image Processing (IWSSIP)</source>
          . IEEE,
          <fpage>399</fpage>
          -
          <lpage>404</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref25">
        <mixed-citation>
          [25]
          <string-name>
            <surname>Klaus</surname>
            <given-names>R.</given-names>
          </string-name>
          <string-name>
            <surname>Scherer</surname>
            and
            <given-names>Eduardo</given-names>
          </string-name>
          <string-name>
            <surname>Coutinho</surname>
          </string-name>
          .
          <year>2013</year>
          .
          <article-title>How music creates emotion: a multifactorial process approach</article-title>
          .
          <source>In The Emotional Power of Music: Multidisciplinary Perspectives on Musical Arousal</source>
          , Expression, and
          <string-name>
            <given-names>Social</given-names>
            <surname>Control</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T</given-names>
            <surname>Cochrane</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B</given-names>
            <surname>Fantini</surname>
          </string-name>
          ,
          <article-title>and</article-title>
          K R Scherer (Eds.).
          <source>Number 10</source>
          . Oxford University Press,
          <fpage>121</fpage>
          -
          <lpage>145</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref26">
        <mixed-citation>
          [26]
          <string-name>
            <surname>Lorenzo</surname>
            <given-names>Tarantino</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Philip N Garner</surname>
            , and
            <given-names>Alexandros</given-names>
          </string-name>
          <string-name>
            <surname>Lazaridis</surname>
          </string-name>
          .
          <year>2019</year>
          .
          <article-title>Self-Attention for Speech Emotion Recognition.</article-title>
          .
          <source>In INTERSPEECH</source>
          .
          <volume>2578</volume>
          -
          <fpage>2582</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref27">
        <mixed-citation>
          [27]
          <string-name>
            <surname>Ashish</surname>
            <given-names>Vaswani</given-names>
          </string-name>
          , Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez,
          <string-name>
            <surname>Łukasz Kaiser</surname>
            , and
            <given-names>Illia</given-names>
          </string-name>
          <string-name>
            <surname>Polosukhin</surname>
          </string-name>
          .
          <year>2017</year>
          .
          <article-title>Attention is all you need</article-title>
          .
          <source>In Advances in neural information processing systems</source>
          .
          <volume>5998</volume>
          -
          <fpage>6008</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref28">
        <mixed-citation>
          [28]
          <string-name>
            <surname>Felix</surname>
            <given-names>Weninger</given-names>
          </string-name>
          , Florian Eyben, and
          <string-name>
            <given-names>Björn</given-names>
            <surname>Schuller</surname>
          </string-name>
          .
          <year>2013</year>
          .
          <article-title>The TUM approach to the MediaEval music emotion task using generic afective audio features. In MediaEval Benchmarking Initiative for Multimedia Evaluation</article-title>
          . Barcelona, Spain.
        </mixed-citation>
      </ref>
      <ref id="ref29">
        <mixed-citation>
          [29]
          <string-name>
            <surname>Minz</surname>
            <given-names>Won</given-names>
          </string-name>
          , Andres Ferraro, Dmitry Bogdanov, and
          <string-name>
            <given-names>Xavier</given-names>
            <surname>Serra</surname>
          </string-name>
          .
          <year>2020</year>
          .
          <article-title>Evaluation of CNN-based Automatic Music Tagging Models</article-title>
          . arXiv preprint arXiv:
          <year>2006</year>
          .
          <volume>00751</volume>
          (
          <year>2020</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref30">
        <mixed-citation>
          [30]
          <string-name>
            <surname>Sanghyun</surname>
            <given-names>Woo</given-names>
          </string-name>
          , Jongchan Park,
          <string-name>
            <surname>Joon-Young Lee</surname>
          </string-name>
          , and In So Kweon.
          <year>2018</year>
          .
          <article-title>CBAM: Convolutional Block Attention Module</article-title>
          .
          <source>In Proceedings of the European Conference on Computer Vision</source>
          (ECCV).
        </mixed-citation>
      </ref>
      <ref id="ref31">
        <mixed-citation>
          [31]
          <string-name>
            <surname>Seunghyun</surname>
            <given-names>Yoon</given-names>
          </string-name>
          , Seokhyun Byun, Subhadeep Dey, and
          <string-name>
            <given-names>Kyomin</given-names>
            <surname>Jung</surname>
          </string-name>
          .
          <year>2019</year>
          .
          <article-title>Speech emotion recognition using multi-hop attention mechanism</article-title>
          .
          <source>In ICASSP 2019-2019 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)</source>
          . IEEE,
          <fpage>2822</fpage>
          -
          <lpage>2826</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref32">
        <mixed-citation>
          [32]
          <string-name>
            <surname>Fan</surname>
            <given-names>Zhang</given-names>
          </string-name>
          , Hongying Meng, and
          <string-name>
            <given-names>Maozhen</given-names>
            <surname>Li</surname>
          </string-name>
          .
          <year>2016</year>
          .
          <article-title>Emotion extraction and recognition from music</article-title>
          .
          <source>In 2016 12th International Conference on Natural Computation</source>
          ,
          <article-title>Fuzzy Systems and Knowledge Discovery (ICNC-FSKD)</article-title>
          . IEEE,
          <fpage>1728</fpage>
          -
          <lpage>1733</lpage>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>