<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Semi-Supervised Music Emotion Recognition using Noisy Student Training and Harmonic Pitch Class Profiles</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Hao Hao Tan helloharry</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>@gmail.com</string-name>
        </contrib>
      </contrib-group>
      <pub-date>
        <year>2021</year>
      </pub-date>
      <fpage>13</fpage>
      <lpage>15</lpage>
      <abstract>
        <p>We present Mirable's submission to the 2021 Emotions and Themes in Music challenge. In this work, we intend to address the question: can we leverage semi-supervised learning techniques on music emotion recognition? With that, we experiment with noisy student training, which has improved model performance in the image classification domain. As the noisy student method requires a strong teacher model, we further delve into the factors including (i) input training length and (ii) complementary music representations to further boost the performance of the teacher model. For (i), we ifnd that models trained with short input length perform better in PR-AUC, whereas those trained with long input length perform better in ROC-AUC. For (ii), we find that using harmonic pitch class profiles (HPCP) consistently improve tagging performance, which suggests that harmonic representation is useful for music emotion tagging. Finally, we find that noisy student method only improves tagging results for the case of long training length. Additionally, we ifnd that ensembling representations trained with diferent training lengths can improve tagging results significantly, which suggest a possible direction to explore incorporating multiple temporal resolutions in the network architecture for future work.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>INTRODUCTION</title>
      <p>
        Emotions and themes are high-level musical attributes that are
abstract and highly subjective. Obtaining emotion labels typically
require human annotation, which can be time consuming and
potentially costly. Is it possible to use semi-supervised learning
techniques, such that we can leverage on unlabelled music tracks to
learn emotion tags, while only using a small amount of labelled
data? Following this question, we intend to explore the usage of
noisy student training [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ] on music emotion recognition. Recently,
[
        <xref ref-type="bibr" rid="ref8">8</xref>
        ] proposed the music tagging transformer, which also uses noisy
student training, but it is applied to general music tagging and
does not focus on emotion and theme related tags. Additionally, we
explore two other factors to improve the tagging performance of
the teacher model: (i) the input training length; (ii) adding music
representations to complement the learning of music emotion.
2.1
      </p>
    </sec>
    <sec id="sec-2">
      <title>APPROACH</title>
    </sec>
    <sec id="sec-3">
      <title>Pre-Processing and Augmentation</title>
      <p>We extract Mel-spectrograms with 128 bins from raw audio using
a sampling rate of 44.1kHz, and the Mel-spectrograms are
downsampled with an averaging factor of 10 along the temporal
dimension. The number of time steps for each Mel-spectrogram vary
baseline
long-normal
long-hpcp
long-hpcp-noisy
short-normal
short-hpcp
short-hpcp-noisy
ensemble</p>
      <p>ROC-AUC</p>
      <p>PR-AUC</p>
      <p>
        F-Score
according to the training strategy, which will be discussed in
Section 2.3. For data augmentation, we perform time masking and
frequency masking, similar to the idea in SpecAugment [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ]. The
maximum possible length of both masks vary between 20 to 60,
and the value is being sampled randomly for each training batch.
2.2
      </p>
    </sec>
    <sec id="sec-4">
      <title>Model Training</title>
      <p>
        As shown in Figure 1, our base model architecture is similar to
CRNN [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ], with some revisions which include adding residual
connections to our ConvBlock, and using GeMPool [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ] instead of
MaxPool. We train all of our models for a maximum of 100 epochs, with
an Adam optimizer and learning rate of 0.0001. Early stopping is
performed when the validation ROC-AUC does not improve for 5
epochs, and we store the model weights from the epoch with the
best ROC-AUC evaluated on the validation set.
2.3
      </p>
    </sec>
    <sec id="sec-5">
      <title>Long VS Short Training Length</title>
      <p>For the long training length mode, we use the first ≈ 185 seconds
of the track, which corresponds to 1600 time steps in the
Melspectrogram after average pooling. For the short training length
mode, we chunk each track into samples of length ≈ 9.25 seconds,
which corresponds to 80 time steps in the Mel-spectrogram after
average pooling. During evaluation, we average the logits of all
chunks to obtain the final output for each track.
2.4</p>
    </sec>
    <sec id="sec-6">
      <title>Harmonic Pitch Class Profiles (HPCP)</title>
      <p>
        HPCP [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ] is a type of chroma feature that describes tonality and
harmonic content of a music track. We extract HPCP with 12 pitch
classes from raw audio using a sampling rate of 44.1kHz. We do
not apply average pooling along the temporal dimension of HPCP.
The corresponding number of time steps for HPCP are 4000 and
200 for both long and short training length mode respectively. We
concatenate the learnt latent features from the Mel-spectrogram
and the HPCP block, each with dimension  = 256, and pass through
two linear layers to obtain the fused output.
      </p>
    </sec>
    <sec id="sec-7">
      <title>Noisy Student Training</title>
      <p>
        Noisy student training [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ] is an extension of self-training, with
the usage of equal-or-larger student models and added noise to
improve the representation learnt from the teacher model. To add
noise, we enhance data augmentation by increasing the maximum
possible masking length to between 30 and 90 for both time and
frequency masking, as well as adding standard Gaussian noise
with a weight of 0.01. To implement stochastic depth [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ], we use
3 StochasticConvBlocks which are ConvBlocks that could be
randomly bypassed with a probability of 0.1 each. During evaluation,
all the layers will be passed through. StochasticConvBlock also has
an additional dropout of probability 0.1 after the ReLU layer.
      </p>
      <p>In this work, we use the corresponding HPCP models for each
long and short training length mode as the teacher model. We only
use the predictions which are &gt; 0.1 as positive pseudo-labels, and
those &lt; 1−6 as negative pseudo-labels. Both decision thresholds
are determined by conducting an empirical evaluation on the
predicted value distribution using the teacher model, carried out on
the training and validation set. We take the leftmost 5% percentile
for the negative label distribution, and the rightmost 5% percentile
for the positive label distribution to ensure better confidence.
2.6</p>
    </sec>
    <sec id="sec-8">
      <title>Model Ensemble</title>
      <p>Finally, we investigate the results of combining the output of both
long and short training length models, by simply taking the weighted
sum of their best models:   =  ·ℎ + (1 −  ) ·. We use
the validation set to find the ratio  which gives the best results.
3</p>
    </sec>
    <sec id="sec-9">
      <title>RESULTS AND ANALYSIS</title>
      <p>For the training length factor, we find that models trained with
long input length perform better in ROC-AUC, but models trained
with short input length perform significantly better in PR-AUC.
According to Table 2, this is because the former has a higher TNR,
while the latter has a higher TPR. Since PR-AUC focuses more on
the minority class (in this case the positive class) and ROC-AUC
focuses on both, the latter model scores better in PR-AUC. We
also find that adding HPCP improves tagging results consistently
for both cases, which suggests that harmonic representation is
important for music emotion recognition.</p>
      <p>For noisy student training, the results are rather inconclusive.
We find slight improvements in the long training length case, but
the result degrades for the short training length case. Also, we only
run noisy student training for 1 iteration, as we find the results
consistently degrade for subsequent iterations. Additionally, we try
to add more unlabelled tracks from the Lakh MP3 dataset (≈ 45, 000
30 seconds track) to increase the training dataset size, but we do
not observe any performance improvement. We infer that noisy
student method might not necessarily work well for music emotion
recognition tasks, due to the abstract nature and subjectivity of
emotion and theme labels. Hence, a small subset of emotion labels
might not be suficient to represent the full dataset.</p>
      <p>For model ensembling, we choose to ensemble the ‘long-noisy’
model and the ‘short-normal’ model. We find that  = 0.7 is optimal
through our validation set, hence suggesting that the final output
gives more weightage to the short training length model. From
the test set results, we can also see that this ensemble method
improves the tagging performance significantly, which suggest that
combining diferent views of audio in terms of temporal resolution
can produce better learnt representations.
4</p>
    </sec>
    <sec id="sec-10">
      <title>DISCUSSION AND OUTLOOK</title>
      <p>While investigating the related work, we find that this work still
uses a relatively long training length (even for short length we
use ≈ 9 seconds, as compared to previous works with ≈ 2 to 5
seconds), and low temporal resolution, which we intend to change
in our future work. For future work, we are interested in tweaking
the network architecture to capture views of diferent temporal
resolutions in the audio sample. We would also like to explore
using noisy student training with diferent model architectures and
datasets of a much larger scale.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>Dmitry</given-names>
            <surname>Bogdanov</surname>
          </string-name>
          , Minz Won, Philip Tovstogan,
          <string-name>
            <given-names>Alastair</given-names>
            <surname>Porter</surname>
          </string-name>
          , and
          <string-name>
            <given-names>Xavier</given-names>
            <surname>Serra</surname>
          </string-name>
          .
          <year>2019</year>
          .
          <article-title>The MTG-Jamendo Dataset for Automatic Music Tagging</article-title>
          .
          <source>In Machine Learning for Music Discovery Workshop, International Conference on Machine Learning (ICML</source>
          <year>2019</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>Keunwoo</given-names>
            <surname>Choi</surname>
          </string-name>
          , György Fazekas,
          <string-name>
            <given-names>Mark</given-names>
            <surname>Sandler</surname>
          </string-name>
          , and
          <string-name>
            <given-names>Kyunghyun</given-names>
            <surname>Cho</surname>
          </string-name>
          .
          <year>2017</year>
          .
          <article-title>Convolutional recurrent neural networks for music classification</article-title>
          .
          <source>In 2017 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)</source>
          . IEEE,
          <fpage>2392</fpage>
          -
          <lpage>2396</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>Emilia</given-names>
            <surname>Gómez</surname>
          </string-name>
          .
          <year>2006</year>
          .
          <article-title>Tonal description of polyphonic audio for music content processing</article-title>
          .
          <source>INFORMS Journal on Computing 18</source>
          ,
          <issue>3</issue>
          (
          <year>2006</year>
          ),
          <fpage>294</fpage>
          -
          <lpage>304</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>Gao</given-names>
            <surname>Huang</surname>
          </string-name>
          , Yu Sun, Zhuang Liu, Daniel Sedra, and
          <string-name>
            <surname>Kilian Q Weinberger</surname>
          </string-name>
          .
          <year>2016</year>
          .
          <article-title>Deep networks with stochastic depth</article-title>
          .
          <source>In European conference on computer vision</source>
          . Springer,
          <fpage>646</fpage>
          -
          <lpage>661</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <surname>Daniel</surname>
            <given-names>S Park</given-names>
          </string-name>
          ,
          <string-name>
            <given-names>William</given-names>
            <surname>Chan</surname>
          </string-name>
          , Yu Zhang, Chung-Cheng Chiu, Barret Zoph,
          <string-name>
            <surname>Ekin D Cubuk</surname>
          </string-name>
          , and
          <string-name>
            <surname>Quoc</surname>
            <given-names>V</given-names>
          </string-name>
          <string-name>
            <surname>Le</surname>
          </string-name>
          .
          <year>2019</year>
          .
          <article-title>Specaugment: A simple data augmentation method for automatic speech recognition</article-title>
          . arXiv preprint arXiv:
          <year>1904</year>
          .
          <volume>08779</volume>
          (
          <year>2019</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>Filip</given-names>
            <surname>Radenović</surname>
          </string-name>
          , Giorgos Tolias, and
          <string-name>
            <given-names>Ondřej</given-names>
            <surname>Chum</surname>
          </string-name>
          .
          <year>2018</year>
          .
          <article-title>Fine-tuning CNN image retrieval with no human annotation</article-title>
          .
          <source>IEEE transactions on pattern analysis and machine intelligence 41</source>
          ,
          <issue>7</issue>
          (
          <year>2018</year>
          ),
          <fpage>1655</fpage>
          -
          <lpage>1668</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>Philip</given-names>
            <surname>Tovstogan</surname>
          </string-name>
          , Dmitry Bogdanov, and
          <string-name>
            <given-names>Alastair</given-names>
            <surname>Porter</surname>
          </string-name>
          .
          <year>2021</year>
          .
          <article-title>MediaEval 2021: Emotion and Theme Recognition in Music Using Jamendo</article-title>
          .
          <source>In Proc. of the MediaEval 2021 Workshop</source>
          , Online,
          <fpage>13</fpage>
          -
          <lpage>15</lpage>
          December
          <year>2021</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>Minz</given-names>
            <surname>Won</surname>
          </string-name>
          , Keunwoo Choi, and
          <string-name>
            <given-names>Xavier</given-names>
            <surname>Serra</surname>
          </string-name>
          .
          <year>2021</year>
          .
          <article-title>Semi-supervised music tagging transformer</article-title>
          .
          <source>In Proc. of International Society for Music Information Retrieval.</source>
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>Qizhe</given-names>
            <surname>Xie</surname>
          </string-name>
          ,
          <string-name>
            <surname>Minh-Thang</surname>
            <given-names>Luong</given-names>
          </string-name>
          , Eduard Hovy, and
          <string-name>
            <surname>Quoc</surname>
            <given-names>V</given-names>
          </string-name>
          <string-name>
            <surname>Le</surname>
          </string-name>
          .
          <year>2020</year>
          .
          <article-title>Self-training with noisy student improves imagenet classification</article-title>
          .
          <source>In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition</source>
          .
          <fpage>10687</fpage>
          -
          <lpage>10698</lpage>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>