<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Novel Approach to Music Genre Classification using Clustering Augmented Learning Method (CALM)</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Soumya Suvra Ghosal</string-name>
          <email>fsoumyasuvraghosal@gmail.com</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Indranil Sarkar</string-name>
          <email>indranil.sarkar.nitdgp@gmail.comg</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>National Institute of Technology Durgapur</institution>
          <country country="IN">India</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2020</year>
      </pub-date>
      <fpage>23</fpage>
      <lpage>25</lpage>
      <abstract>
        <p>This paper proposes an automatic music genre-classification system using a deep learning model. The proposed model leverages Convolutional Neural Nets(CNN) to extract local features and LSTM Sequence to Sequence Autoencoders to learn representations of time series data by taking into account their temporal dynamics. The paper also introduces Clustering Augmented Learning Method (CALM) classifier which is based on the concept of simultaneous heterogeneous clustering and classification to learn deep feature representations of the features obtained from LSTM autoencoder. Computational Experiments using GTZAN dataset resulted in an overall test accuracy of 95.4% with a precision of 91.87%.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Introduction</title>
      <p>With the increasing amount of music available online, there
is automatically a growing demand for the symmetrical
organization of audio files and that has increased the interest
in music classification. To detect a group of the music of a
similar genre is the main work of the recommendation
system and playlist generators. Thus building a robust music
classifier using machine learning techniques is essential to
automate tagging unlabeled music and improve user’s
experience of media players and music libraries. In recent years,
convolutional neural networks(CNNs) have brought
revolutionary changes to the computer vision community.
Meanwhile, CNN’s have been widely used for music information
retrieval, especially music genre classification. Recently, it
became increasingly popular to combine CNNs with
recurrent networks(RNNs) to process audio signals, which
introduce time-sequential information to the model. In
convolutional recurrent networks(C-RNNs), the CNN
component is used to extract features while RNN plays the role
of summarizing temporal features. The inputs of C-RNNs
are soundtrack spectrograms and outputs are probabilities of
each genre at each timestep. Inspired by previous literature,
we propose to leverage the idea by augmenting LSTM
autoencoder with CNN and use a Clustering-based classifier to
predict the genre of music.</p>
    </sec>
    <sec id="sec-2">
      <title>Previous Works</title>
      <p>
        Music genre classification has been actively studied since
the early days. Tzanetakis and Cook [Tzanetakis and
Cook2002] used k-nearest neighbor classifier and Gaussian
Mixture models with a comprehensive set of features for
music classification. Those features could be summarized
into three categories: rhythm, pitch, and temporal
structure. Zhouyu Fu [Fu et al.2010] proposed a Naive Bayes
(NB) classifier framework, namely NB Nearest Neighbor
(NBNN) and NB Support Vector Machine (NBSVM) for
music genre classification. [Deshpande and Singh2001]
compared k-nearest neighbor, Gaussian Mixtures and SVM
to classify music into three genres which are rock, piano,
and jazz. In recent years, using an audio spectrogram has
become mainstream for music genre classification.
Spectrograms encode time and frequency information of given
music as a whole. Spectrograms can be considered as images
and used to train convolutional neural networks (CNNs)
        <xref ref-type="bibr" rid="ref14">( [Wyse2017])</xref>
        . [Li, Chan, and A2010] developed a CNN to
predict the music genre using the raw Mel Frequency
cepstral coefficients(MFCCs) as input.
      </p>
      <p>In this paper, we aim to combine convolutional nets with
LSTM Autoencoders to extract both spatial and temporal
features of the audio signal. Instead of baseline classifiers,
we propose a clustering-based classification model. In the
proposed classification approach we cluster the data based
on their inherent characteristics and in the process of
learning the best clustering solution we optimize the
hyperparameters of the classification model, thereby substantially
improving the learning process. We used the mel-spectrogram
as the only feature and compared the proposed model with
traditional classifiers and previous literature.</p>
    </sec>
    <sec id="sec-3">
      <title>Dataset and Representation</title>
      <p>Dataset In the paper, we have used the GTZAN dataset. It
contains 10 music genres, each genre has 100 audio clips in
.au format. The genres are - blues, classical, country, disco,
hip-hop, pop, jazz, reggae, rock, metal. Each audio clips has
a length 30 seconds, are 22050Hz Mono 16-bit files. The
dataset incorporates samples from a variety of sources like
CDs, radios, microphone recordings, etc. The training,
testing and validating sets are randomly partitioned following
proportion 8:1:1.</p>
      <p>Features A popular representation of sound is the
spectrogram which captures both time and frequency information.
In this study, we used the Mel spectrogram as the only
input to train our neural model. A mel spectrogram is a
spectrogram transformed to have frequencies in the mel scale,
which is logarithmic, more naturally representing how
human senses different sound frequencies. To convert raw
audio to Mel spectrogram, one must apply Short Time Fourier
Transforms(STFT) across sliding windows of audio, around
20ms wide.</p>
      <p>In this case, the music features are extracted using the
LibROSA library in Python using 128 mel filters, frame
length of 2048 samples and a hop size of 1024. We got a
spectrogram of size 647 128.</p>
    </sec>
    <sec id="sec-4">
      <title>Proposed architecture and methodology</title>
      <p>The model consists of a four-layer convolutional neural
network (CNN) which is followed by an LSTM Sequence to
Sequence Autoencoder(AE) and ultimately consists of the
proposed CALM classifier. Not only to make the network
unconstrained of any handcrafted features, but the
convolutional layers are also used to extract meaningful and useful
features from the song. The output of the CNN is a sequence
in which every timestep strongly relies on both the
immediate predecessors and long term structure of the entire song.
To capture both transient and overall characteristics, we use
LSTM Sequence to Sequence Autoencoder and for
classification, we propose CALM, which is explained in the
following sections. The assumption underlying this model is that
the temporal pattern can be aggregated better with LSTM
Autoencoders than CNNs while relying on CNNs on input
side for local feature extraction.</p>
      <p>The CNN architecture consists of 4 convolutional layers
of 64 feature maps, 3-by-3 convolution kernels and
maxpooling layers of dimensions (2×2)-(3×3)-(4×4)-(4×4). In
all convolutions, we pad zeros to each side of the input
to keep size fixed. Dropout(0.5) is applied to all
convolutional layers to increases generalization. The CNN
output has a feature map size of N×1×15 (number of
feature maps×frequency×time). For extracting temporal
pattern we use an LSTM-based architecture. The architecture
uses LSTM layers having f256,64,16g units as the encoder
LSTMenc and LSTM layers having f64,256g units as the
decoder LSTMdec.</p>
      <p>Methodology To start, features are extracted from the
spectrogram using convolutional layers. The output of
Convolutional Neural Networks is fed to an LSTM Seq to Seq
Autoencoder which collects key information about the
temporal properties of the input sequence in its hidden state. The
final hidden state of the LSTMenc is then passed through
some layers, the output of which is used to initialize the
hidden state of the LSTMdec. The function of the LSTMdec is
to reconstruct the input sequence based on the information
contained in its initial hidden state. The network is trained
to minimize the root mean squared error between the input
sequence and the reconstruction. Once the training is
complete, the activation of the fully connected encoded layer is
used as representations of the audio sequence and is fed as
input to Clustering Augmented Learning Method Classifier.
This system showed 98% accuracy at the end of the training.</p>
    </sec>
    <sec id="sec-5">
      <title>Clustering Augmented Learning Method (CALM)</title>
      <sec id="sec-5-1">
        <title>Proposed Approach</title>
        <p>Input augmentationAs in [Ghosal et al.2019], we consider
a matrix of input data D and a set of cluster centers C. Since
in this case study, there are 10 music genres, we keep C
as 10. In this paper, we use clustering to augment input data
x 2 D for better learning. To augment the input data, we add
a new set of features representing either an input example
belongs to a cluster or not. To distinguish input examples, we
introduce an additional index h 2 f1; : : : ; jDjg representing
the number of an input example (x1 is the first input example
of D). We define also a vector ch composed of chl, l 2 C
for each example xh 2 D. It is a one-hot representation
containing zeros except for the index of the cluster it belongs to
(e.g. c1 = [0; 0; 0; 1; 0; 0; 0; 0; 0; 0] means that the first
input example x1 belongs to the 4th cluster out of 10 clusters).
Finally, we augment input examples by concatenating the
vector xh with the vector ch for each h 2 f1; : : : ; jDjg.</p>
        <p>Cluster centers To determine the cluster centers, CALM
consists of a clustering model and a Feed-Forward Neural
Net(FNN) having a softmax output to classify the music
genres. For the clustering model, we propose to use a
Random Forest classifier to determine cluster centers. After the
FNN is trained using a state-of-the-art solver for data
belonging to a single cluster 2 f1; : : : ; jCjg, a Random Forest
Classifier is used to find the best cluster center. Hence we
repeat jCj instances of training the FNN to find the jCj
centers. For any instance l of the model, we use the one-hot
encoded vector of l as labels for all the input sample in that
cluster. In simple words, while predicting center of 4th
cluster (for example) we use [0; 0; 0; 1; 0; 0; 0; 0; 0; 0] as label for
all input samples, since jCj is 10.</p>
        <p>We propose that the input sample which has the lowest
error in predicting its cluster label is considered as the
center of that cluster in the subsequent iteration of the proposed
approach. In such a manner, the center would be the input
sample which is the most fitting representative of that
cluster. As a result, the clustering process would aggregate the
data having similar characteristics resulting in better
learning by the FNN classification model.</p>
      </sec>
      <sec id="sec-5-2">
        <title>Clustering Problem</title>
        <p>We have a distance/dissimilarity measure dil between input
examples i 2 D and cluster centers l 2 C. The clustering
problem aims to assign each input example to a cluster such
that the total distance between the elements of a cluster and
its center is minimized.</p>
        <p>In this paper we also propose a novel dissimilarity
measure based on the weights of the trained FNN classifier.
It uses the average of weights linked to each neuron of
the input layer. Assuming that the original input
(without the new clustering feature) has d dimensions (xh =
[x1h; : : : ; xdh]; h 2 f1; : : : ; jDjg) and the weight linking node
n of the input layer to node j 2 f1; : : : ; n1g of the
following layer is wjn, the two distances measures are formulated
as follows:
dil =</p>
        <p>P</p>
        <p>avg
n2f1:::dg j2f1;:::;n1g
wjnjxi
k</p>
        <p>k
xl j
Thus the distance measure computes the distance between
two examples based on how important is the contribution of
each input feature to the resulting prediction. Therefore, the
resulting clusters contain examples with similar potential to
improve the classification results.</p>
      </sec>
      <sec id="sec-5-3">
        <title>Proposed Algorithm</title>
        <p>We propose an approach (Algorithm 1) where we iteratively
train the FNN classifier, use its weights for input data
clustering thus changing the input vector, train again the FNN
classifier using the new input data, and so on until a
stopping criterion is attained. The stopping criterion is triggered
if the cluster assignment remains the same for consecutive
10 iterations, i.e., the clustering problem converges.
The configuration of the proposed model is given as:
where NC is the number of accurately predicted music
tracks, NF is the number of falsely predicted music tracks,
NM is the number of missed music tracks, totalc is the
number of all accurately predicted music tracks and totalm
is the number of all music tracks.</p>
        <p>To further interpret the results, we plotted the confusion
matrix(Table 3) of the proposed model. Looking more
closely at our confusion matrix, we see that our proposed
model managed to correctly classify 80% of rock audio
as rock, labeling the others as mainly country or blues.
Additionally, it incorrectly classified some country, as well
as a small fraction of blues and reggae, as rock music.</p>
        <p>Comparison with Baseline Classifiers We trained four
traditional classification models on the dataset as baseline
classifiers, including k-nearest neighbors, logistic
regression, random forest, multilayer perceptrons, and linear
support vector machine, using Mel Frequency Cepstral
Coefficients(MFCCs) by flattening them into a 1-D array. Apart
from baseline classifiers we also experimented by stacking
a Logistic Regression classifier with the features obtained
from Convolutional Net and LSTM Autoencoder to test the
performance of the CALM classifier. As evident from
Table 1, CALM outperforms the Logistic Regression
classifier when augmented with Convolutional nets and LSTM
autoencoder. Moreover, Fig.4 shows how the intra-cluster
variance decreases after approximately 75 iterations and then
stabilizes. To measure intra-cluster variance, we used
Euclidean distance in this case study. Similarly, it is evident
from Fig. 3 that testing loss starts decreasing after 80 epochs
and gradually as the clustering solution converges, the
accuracy begins to improve. This observation bolsters our initial
assumption that clustering data based on inherent
characteristics would improve the learning process of FNN.
For a fair comparison, all models are trained and tested
on the same dataset as the proposed model. The
hyperparameters are tuned by a grid search to ensure that the best
model configuration is adapted. In table 2 we have also
compared our model with relevant literature and it is evident that
proposed architecture performs strongly.</p>
      </sec>
    </sec>
    <sec id="sec-6">
      <title>Conclusion</title>
      <p>In this paper, we present a specially designed network for
accurately recognizing the music genre. The proposed model
Model
aims to take full advantage of low-level information of
Melspectrogram for making the classification decision. We have
shown how our model is effective by comparing the
stateof-art methods, including both hands crafted feature
approaches and deep learning models. In this work, we use
the GTZAN dataset which is a common benchmark dataset.
Our proposed model has achieved an impressive accuracy of
95.4% while testing, which outperforms all other models. In
the future, we will try to improve the model by improvising
some new distance metric methods to compute the similarity
between genres.</p>
      <p>Blues Classical Country</p>
      <p>Disco
0
0
82
3
0
0
0
0
0
7</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [Dai et al.2015]
          <string-name>
            <surname>Dai</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ; Liu,
          <string-name>
            <given-names>W.</given-names>
            ;
            <surname>Dong</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            ; and
            <surname>Yang</surname>
          </string-name>
          ,
          <string-name>
            <surname>H.</surname>
          </string-name>
          <year>2015</year>
          .
          <article-title>Multilingual deep neural network for music genre classification</article-title>
          .
          <source>In Sixteenth Annual Conference of the International Speech Communication Association.</source>
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [Deshpande and Singh2001]
          <string-name>
            <surname>Deshpande</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Singh</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          <year>2001</year>
          .
          <article-title>Classification of music signals in the visual domain</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [Fu et al.2010]
          <string-name>
            <surname>Fu</surname>
            ,
            <given-names>Z.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Lu</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Ting</surname>
            ,
            <given-names>K. M.</given-names>
          </string-name>
          ; and Zhang,
          <string-name>
            <surname>D.</surname>
          </string-name>
          <year>2010</year>
          .
          <article-title>Learning naive bayes classifiers for music classification and retrieval</article-title>
          .
          <source>In 2010 International Conference on Pattern Recognition.</source>
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [Ghosal et al.2019]
          <string-name>
            <surname>Ghosal</surname>
            ,
            <given-names>S. S.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Bani</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Amrouss</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ; and
          <string-name>
            <surname>El Hallaoui</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          <year>2019</year>
          .
          <article-title>A deep learning approach to predict parking occupancy using cluster augmented learning method</article-title>
          .
          <source>In 2019 International Conference on Data Mining Workshops (ICDMW)</source>
          ,
          <fpage>581</fpage>
          -
          <lpage>586</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [Karunakaran and Arya2018]
          <string-name>
            <surname>Karunakaran</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Arya</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <year>2018</year>
          .
          <article-title>A scalable hybrid classifier for music genre classification using machine learning concepts and spark</article-title>
          .
          <source>In International Conference on Intelligent Autonomous Systems (ICoIAS)</source>
          ,
          <fpage>128</fpage>
          -
          <lpage>135</lpage>
          . IEEE.
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [Kingma and Ba2014]
          <string-name>
            <surname>Kingma</surname>
            ,
            <given-names>D. P.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Ba</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          <year>2014</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          <article-title>Adam: A method for stochastic optimization</article-title>
          .
          <source>arXiv preprint arXiv:1412</source>
          .
          <fpage>6980</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [Li,
          <string-name>
            <surname>Chan</surname>
          </string-name>
          , and A2010]
          <string-name>
            <surname>Li</surname>
            ,
            <given-names>T. L.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Chan</surname>
            ,
            <given-names>A. B.</given-names>
          </string-name>
          ; and A,
          <string-name>
            <surname>C.</surname>
          </string-name>
          <year>2010</year>
          .
          <article-title>Automatic musical pattern feature extraction using convolutional neural network</article-title>
          .
          <source>In 2015 Data Mining and Applications</source>
          . IEEE.
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [Liu et al.2019] Liu,
          <string-name>
            <given-names>C.</given-names>
            ;
            <surname>Feng</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            ;
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <surname>H.</surname>
          </string-name>
          ; and Liu,
          <string-name>
            <surname>S.</surname>
          </string-name>
          <year>2019</year>
          .
          <article-title>Bottom-up broadcast neural network for music genre classification</article-title>
          . arXiv preprint arXiv:
          <year>1901</year>
          .08928v1.
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [Matityaho and Furst2006]
          <string-name>
            <surname>Matityaho</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Furst</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          <year>2006</year>
          .
          <article-title>Aggregate features and adaboost for music classification</article-title>
          .
          <source>Machine Learning 473-484.</source>
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [Nanni et al.2017]
          <string-name>
            <surname>Nanni</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Costa</surname>
            ,
            <given-names>Y. M.</given-names>
          </string-name>
          ;
          <string-name>
            <surname>Lucio</surname>
            ,
            <given-names>D. R.</given-names>
          </string-name>
          ; Silla Jr,
          <string-name>
            <given-names>C. N.</given-names>
            ; and
            <surname>Brahnam</surname>
          </string-name>
          ,
          <string-name>
            <surname>S.</surname>
          </string-name>
          <year>2017</year>
          .
          <article-title>Combining visual and acoustic features for audio classification tasks</article-title>
          .
          <source>Pattern Recognition Letters</source>
          <volume>49</volume>
          -56.
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [Prechelt1998]
          <string-name>
            <surname>Prechelt</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          <year>1998</year>
          .
          <article-title>Early stopping-but when? In Neural Networks: Tricks of the trade</article-title>
          . Springer.
          <fpage>55</fpage>
          -
          <lpage>69</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [Tzanetakis and Cook2002]
          <string-name>
            <surname>Tzanetakis</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Cook</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          <year>2002</year>
          .
          <article-title>Musical genre classification of audio signal</article-title>
          .
          <source>IEEE Transactions on Speech, and Audio Processing</source>
          <volume>10</volume>
          (
          <issue>3</issue>
          ):
          <fpage>293</fpage>
          -
          <lpage>302</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [Wyse2017]
          <string-name>
            <surname>Wyse</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          <year>2017</year>
          .
          <article-title>Audio spectrogram representations for processing with convolutional neural networks</article-title>
          .
          <source>arXiv preprint arXiv: 1706</source>
          .
          <fpage>09559</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [Zhang et al.
          <year>2016</year>
          ] Zhang, W.; Lei,
          <string-name>
            <given-names>W.</given-names>
            ;
            <surname>Xu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            ; and
            <surname>Xing</surname>
          </string-name>
          ,
          <string-name>
            <surname>X.</surname>
          </string-name>
          <year>2016</year>
          .
          <article-title>Improved music genre classifi- cation with convolutional neural networks</article-title>
          .
          <source>INTERSPEECH 3304-3308.</source>
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>