<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Single and Multi Column Neural Networks for Content-based Music Genre Recognition</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Chang Wook Kim</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Jaehun Kim</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Kwangsub Kim</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Minz Won</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Kakao Corp.</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Republic of Korea</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Kakao Brain</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Republic of Korea</string-name>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Delft University of Technology</institution>
          ,
          <country country="NL">Netherlands</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2017</year>
      </pub-date>
      <fpage>13</fpage>
      <lpage>15</lpage>
      <abstract>
        <p>This working note reports approaches of team KART to MediaEval2017 AcousticBrainz Genre Task and their results. To solve the problem, we mainly considered the sparsity and noise of data, network design for the multi-label classification, and implementation of successful Deep Neural Network (DNN) models. We propose three steps of preprocessing and depict two diferent approaches: a single-column model and a multi-column model.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>INTRODUCTION</title>
      <p>
        A music genre is a class, type or category that defined by
convention [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ]. However, taxonomies of music genres can difer by
communities. The MediaEval2017 AcousticBrainz Genre Task aims to
predict the genre and subgenre of unlabeled music recordings from four
diferent datasets which consist of four diferent genre/subgenre
taxonomies [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. Each dataset includes precomputed music audio
features using Essentia library [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ] and genre/subgenre annotations
that follow its own taxonomy.
      </p>
      <p>We approached the problem based on careful consideration of
following:
• How to handle the noisy and sparse data?
• How to solve the multi-label classification task?
• How to apply a variety of successful deep neural network
models to our task?</p>
    </sec>
    <sec id="sec-2">
      <title>PREPROCESSING</title>
      <p>Before starting the model training, we conducted three steps of
feature preprocessing: (i) feature vectorization, (ii) fitting outlier
feature values to the outlier boundaries, and (iii) selecting features
by feature value distribution analysis.</p>
      <p>Essentially, we tried to use features as raw as possible. Its
underlying assumption is that the deep neural network model can
learn useful representations of the raw data if there are suficient
amount of samples. We omitted all the information under ’meta
data’ keys. Further, we applied PCA to the covariance matrices
of filter banks. For the computation convenience, we only choose
the eigenvector whose corresponding eigenvalue is most large. In
addition, we encoded categorical features into binary vectors in
one-hot manner.</p>
      <p>For some features, there are outliers with extremely high or
low values in comparison with their medians (Figure 1 left). We
∗Author’s names are listed alphabetically; authors contributed equally to this work.
suppressed or boosted these outlier values to the outlier boundaries.
A lower limit and an upper limit boundaries for judging outliers
are defined as :</p>
      <p>Lower Limit = FirstQuartile − 1.5 ∗ IQR
U pper Limit = T hirdQuartile + 1.5 ∗ IQR
(1)
where Inter Quartile Range (IQR) is a diference between the first
quartile and the third quartile. Values bigger than the upper limit
were clipped at the upper limit, and values smaller than the lower
limit were boosted to the lower limit. Figure 1 (right) shows the
distribution after the fitting of the Figure 1 (left).</p>
      <p>Upper limits and lower limits of the training set were derived
from feature value distributions of each feature and each genre.
However, since the genre labels of the test set should not be
informed, we thresholded the test set at the maximum upper limit
and the minimum lower limit of the training set.</p>
      <p>After the outlier fitting, we defined features that concentrate
around same values for every genre as useless features and removed
133 features from the training and test sets.
3</p>
    </sec>
    <sec id="sec-3">
      <title>MODEL</title>
      <p>We implemented two Feed-forward Neural Networks (FNNs). The
main diference in architectures is whether the label hierarchy
between the genre and the sub-genre is considered explicitly. Since
the provided input features are already processed, the model is
designed for encoding interdependency among labels.
3.1</p>
    </sec>
    <sec id="sec-4">
      <title>Single Column Model</title>
      <p>As a baseline, we implemented an Single Column FNN (SCNN)
whose output dimensions correspond to the entire labels. The label
hierarchy between genre and sub-genre was not considered in this
model. The genre and sub-genre are equally treated as independent
labels. We applied the weight vector w with the loss function, to
give less penalty for more frequent labels as following:
wi = 1 +</p>
      <p>1
log (1 + fi )
where fi denote raw count of label i in the given training dataset.
In this way, the error from less frequent labels can be counted
relatively larger than more frequent labels. It leads to learning less
frequent labels more sensitively. The loss function of this model is:
1 Õ
LSC N N = M
m,i
wi H ( yˆi(m), yi(m))
where H denotes the binary cross-entropy of the true label and the
prediction, yi(m) is a binary vector for label i of the observation m,
yˆ (m) is a prediction for label i of the observation m, and M denotes
i
the size of the mini-batch. The inference of yˆi(m) is:</p>
      <p>
        yˆ (m) = f (x(m); θ )
where f (x(m); θ ) denotes an FNN that has a set of model
parameters denoted as θ , and x(m) is a feature vector of the observation
m. We applied ReLU[
        <xref ref-type="bibr" rid="ref5">5</xref>
        ] activation function for hidden layers and
used the sigmoid function for the output layer. We also used the
dropout[
        <xref ref-type="bibr" rid="ref7">7</xref>
        ] for every hidden layers with the dropout probability
0.5. We applied the batch normalization[
        <xref ref-type="bibr" rid="ref4">4</xref>
        ] to the hidden layers to
accelerate training process. The details are depicted in Figure 2.
      </p>
      <p>During the inference, we only predict labels whose probability
exceed the threshold α . We set α as 0.2 and 0.3 for SCNN, which
are found as optimal values by the cross-validation.
3.2</p>
    </sec>
    <sec id="sec-5">
      <title>Multi Column Model</title>
      <p>To explicitly reflect label dependency between sub-genre and genre,
we implemented a Multi Column FNN model (MCNN). It has a
parallel structure for each set of sub-genre and genre, and merged
by the Bayes rule on the top layer, as following:</p>
      <p>yˆ д(m) = f (x(m); θд )
yˆ s(mд) = yˆ д(m∗ ) f (x(m); θsд )
where yˆ д(m) and yˆ s(mд) denote the estimated probabilities of genre
and sub-genre from each model. The posterior probability of
subgenre yˆ s(mд) is conditioned by yˆ д(m∗ ). Here д∗ denotes the genre where
the sub-genre sд belongs. The loss function of this model is:
1 Õ
LMC N N = M m</p>
      <p>[wдH (yˆ д(m), yд(m)) + wsдH (yˆ s(mд), ys(mд))]
where wд and wsд are scalar weights to balance learning rate of
the genre column and the sub-genre column. We used the ratio of 1:9
between wд and wsд , considering sub-genre labels are more sparse
than genre labels. We used the batch normalization only after the
input layer. The dropout was not applied. We used threshold αsд =
0.25 for sub-genre and threshold αд = 0.4 for genre, respectively.</p>
      <p>
        Assuming given feature set is already suficiently processed, we
also applied a highway network[
        <xref ref-type="bibr" rid="ref8">8</xref>
        ] architecture. It controls the
gradient flow by the parametric gate at each layer, similar to the
Long Short-Term Memory[
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]. We applied this architecture for the
network not only to have a deeper structure, but also to use the
information more close to the input feature.
(2)
(3)
(4)
(5)
(6)
(7)
A partial result of our runs and baselines obtained from test set
presented in Table 1. The scores are mean F 1 scores per tracks and
per labels, respectively, which are averaged over all results from
datasets. Baseline1 is a random predictor and Baseline2 is a majority
predictor. SCNN presented in Table 1 uses threshold α = 0.2.
      </p>
      <p>When comparing the scores, we noticed that SCNN is overall
better than Baselines, but the MCNN is better than Baselines only in
per label scores. However, considering the recall scores, which are
not presented in the note due to space, it shows that both suggested
models score better recall than the Baseline 2. This shows both
models are working better for predicting sparse sub-genres than
baselines and suggesting our weighted losses work as intended.</p>
      <p>Compared to the validation accuracy, the test accuracy got worse.
Since the training set and the validation set are skewed and sparse
data, our models failed to learn generalized parameters.
Experiments with the data augmentation have to be explored to overcome
the drawback.</p>
      <p>Also, a large model size of MCNN can be another reason of its
worse test accuracy. The highway networks have 2 times larger than
the standard fully-connected layers and MCNN has two columns
of the network to model genre and sub-genre predictor separately.
This structure makes the model 4 times bigger than SCNN model.
A Multi-Column architecture with small units and standard
fullyconnected layers will be useful.</p>
      <p>Content-based Music Genre Recognition from Multiple Sources</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>Dmitry</given-names>
            <surname>Bogdanov</surname>
          </string-name>
          , Alastair Porter,
          <string-name>
            <given-names>Julian</given-names>
            <surname>Urbano</surname>
          </string-name>
          , and
          <string-name>
            <given-names>Hendrik</given-names>
            <surname>Schreiber</surname>
          </string-name>
          .
          <year>2017</year>
          .
          <article-title>MediaEval 2017 AcousticBrainz Genre Task: Contentbased Music Genre Recognition from Multiple Sources</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>Dmitry</given-names>
            <surname>Bogdanov</surname>
          </string-name>
          , Nicolas Wack, Emilia Gómez, Sankalp Gulati, Perfecto Herrera, Oscar Mayor, Gerard Roma, Justin Salamon, José R Zapata, Xavier Serra, and others.
          <source>2013</source>
          .
          <article-title>Essentia: An Audio Analysis Library for Music Information Retrieval</article-title>
          .
          <source>In International Society for Music Information Retrieval (ISMIR'13) Conference</source>
          . Curitiba, Brazil,
          <fpage>493</fpage>
          -
          <lpage>498</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>Sepp</given-names>
            <surname>Hochreiter and JÃĳrgen Schmidhuber</surname>
          </string-name>
          .
          <year>1997</year>
          .
          <article-title>Long Short-term Memory</article-title>
          .
          <volume>9</volume>
          (
          <issue>12</issue>
          <year>1997</year>
          ),
          <fpage>1735</fpage>
          -
          <lpage>80</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>Sergey</given-names>
            <surname>Iofe</surname>
          </string-name>
          and
          <string-name>
            <given-names>Christian</given-names>
            <surname>Szegedy</surname>
          </string-name>
          .
          <year>2015</year>
          .
          <article-title>Batch normalization: Accelerating deep network training by reducing internal covariate shift</article-title>
          .
          <source>In International Conference on Machine Learning</source>
          .
          <fpage>448</fpage>
          -
          <lpage>456</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>Vinod</given-names>
            <surname>Nair</surname>
          </string-name>
          and
          <string-name>
            <given-names>Geofrey E</given-names>
            <surname>Hinton</surname>
          </string-name>
          .
          <year>2010</year>
          .
          <article-title>Rectified linear units improve restricted boltzmann machines</article-title>
          .
          <source>In Proceedings of the 27th international conference on machine learning (ICML-10)</source>
          .
          <fpage>807</fpage>
          -
          <lpage>814</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>Jim</given-names>
            <surname>Samson</surname>
          </string-name>
          .
          <year>2017</year>
          .
          <article-title>Genre</article-title>
          .
          <source>In Grove Music Online. Oxford Music Online</source>
          . Oxford University Press. Web., http://www.oxfordmusiconline.com/subscriber/article/grove/music/40599.
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>Nitish</given-names>
            <surname>Srivastava</surname>
          </string-name>
          , Geofrey E Hinton, Alex Krizhevsky, Ilya Sutskever, and
          <string-name>
            <given-names>Ruslan</given-names>
            <surname>Salakhutdinov</surname>
          </string-name>
          .
          <year>2014</year>
          .
          <article-title>Dropout: a simple way to prevent neural networks from overfitting</article-title>
          .
          <source>Journal of machine learning research 15</source>
          ,
          <issue>1</issue>
          (
          <year>2014</year>
          ),
          <fpage>1929</fpage>
          -
          <lpage>1958</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>Rupesh</given-names>
            <surname>Kumar</surname>
          </string-name>
          <string-name>
            <surname>Srivastava</surname>
          </string-name>
          , Klaus Gref, and
          <string-name>
            <given-names>Jürgen</given-names>
            <surname>Schmidhuber</surname>
          </string-name>
          .
          <year>2015</year>
          .
          <article-title>Highway networks</article-title>
          .
          <source>arXiv preprint arXiv:1505.00387</source>
          (
          <year>2015</year>
          ).
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>