<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Nicolas Dauban</string-name>
          <email>nicolas.dauban@irit.fr</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>IRIT, Université de Toulouse</institution>
          ,
          <addr-line>CNRS, Toulouse</addr-line>
          ,
          <country country="FR">France</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2017</year>
      </pub-date>
      <fpage>13</fpage>
      <lpage>15</lpage>
      <abstract>
        <p>This paper presents a method of genre classification using deep neural networks for the AcousticBrainz genre classification task of MediaEval 2017. The AcousticBrainz Genre Task 2017 is a music genre recognition (MGR) task organised by MediaEval [1] where participants had to make predictions on genre and subgenres, based on audio features extracted from Essentia [2]. The target labels are provided by 4 diferent sources: Discogs, Allmusic, Tagtraum and Lastfm.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>INTRODUCTION</title>
    </sec>
    <sec id="sec-2">
      <title>RELATED WORK</title>
      <p>
        I participated in this challenge during my internship on
contentbased music recommendation. During this internship, I worked
on music genre recognition using music features with deep neural
network [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]. This approach has been tested on the GTZAN [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ]
corpus and then on the MagnaTagATune [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ] corpus. The accuracy
results obtained were about 90% on the GTZAN corpus and about
80% with MagnaTagATune. As these results were satisfying, we
decided to take part of the challenge to try the neural network
approach.
      </p>
    </sec>
    <sec id="sec-3">
      <title>APPROACH</title>
      <p>As the time evolution of the music features was not supplied,
the convolutional approach - which needs correlations between
successive frames of the input image - has been quickly aborted.
Thus, we used a classic deep neural network instead of a
convolutional one. The neural network has been implemented using the
Theano framework “Lasagne” 1. We tried to use diferent features
and diferent architectures of neural networks.</p>
    </sec>
    <sec id="sec-4">
      <title>Features</title>
      <p>During the development phase, the following features have been
tried: Mel bands, spectral rollof, zero crossing rate, spectral
entropy, Harmonic Pitch Class Profile (HPCP), tempo, danceability,
key strength, dissonance, tuning diatonic strength. The choice of
those features has been made by relying on their definitions and by
choosing those which seem to be the more relevant to characterize
the diferent music genres. In our preliminary experiments, Mel
bands yielded the best results, thus, only Mel bands were used as
input for our final submission.</p>
      <p>1. https://lasagne.readthedocs.io/en/latest/
3.2</p>
    </sec>
    <sec id="sec-5">
      <title>Architecture of the network</title>
      <p>To establish the final topology of the neural network, diferent
architectures have been tried. The final architecture comprises 8
full connected layers of 2000 neurons with ReLUs as activation
function (Figure 1). The input vector had size of all the data used
(for the 9 statistics on the 40 MEL bands, the input was a vector of
size 360), and an output with the total number of genre and
subgenres (e.g. 315 for the Discogs subset) with a sigmoid activation
function. With the sigmoid output function, it was common that
the network did not predict any genre for some tracks, because not
a single output neuron had a value superior to 0.5 for those tracks.
In order to have at least one genre prediction per track if no genre
was predicted, the genre with the maximum output value is chosen
as the predicted one. For the submission, the network was trained
with 40 epochs.</p>
    </sec>
    <sec id="sec-6">
      <title>RESULTS AND ANALYSIS</title>
      <p>Only one submission was made for each dataset, for the task 1
only.</p>
      <p>The Table 1 and Table 2 contain the results per tracks for all
labels and for genre labels:</p>
      <p>The results for sub-genres and per label are not shown here
because they were all between 0% and 10%. All the results are provided
on the AcousticBrainz Genre Task results page 2.</p>
      <p>2. https://multimediaeval.github.io/2017-AcousticBrainz-Genre-Task/results/</p>
      <p>F-score
0.3513</p>
      <p>During the development phase, the network obtained a 60%
FScore on genre-only on Discogs. During this phase, the data has
been split using the python script provided by organizers to ensure
album filtering: 80% to train the network and 20% to test it. However,
the best F-score obtained on genre labels of our final submission
was 35% per track maybe due to a lack of generalization power of
our network.
5</p>
    </sec>
    <sec id="sec-7">
      <title>DISCUSSION</title>
      <p>We plan to explore our initial idea to use combinations of
different acoustic feature types as we previously showed that these
were as successful as using Mel bands. Furthermore, some
modifications on the network can be done: adding batch-normalization
after each dense layer, residual blocks, try some diferent
initialization and decay strategies for the learning rate, and also lower the
threshold of the output sigmoid in order to give more predictions.</p>
    </sec>
    <sec id="sec-8">
      <title>RÉFÉRENCES</title>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>Dmitry</given-names>
            <surname>Bogdanov</surname>
          </string-name>
          , Alastair Porter,
          <string-name>
            <given-names>Julián</given-names>
            <surname>Urbano</surname>
          </string-name>
          , and
          <string-name>
            <given-names>Hendrik</given-names>
            <surname>Schreiber</surname>
          </string-name>
          .
          <article-title>The mediaeval 2017 acousticbrainz genre task: Content-based music genre recognition from multiple sources</article-title>
          .
          <source>In Working Notes Proceedings of the MediaEval Workshop</source>
          , Dublin, Ireland,
          <source>September 13-15</source>
          ,
          <year>2017</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>Dmitry</given-names>
            <surname>Bogdanov</surname>
          </string-name>
          , Nicolas Wack, Emilia Gómez, Sankalp Gulati, Perfecto Herrera, Oscar Mayor, Gerard Roma, Justin Salamon, José R Zapata,
          <string-name>
            <given-names>Xavier</given-names>
            <surname>Serra</surname>
          </string-name>
          , et al.
          <article-title>Essentia: An audio analysis library for music information retrieval</article-title>
          .
          <source>In ISMIR</source>
          , pages
          <fpage>493</fpage>
          -
          <lpage>498</lpage>
          ,
          <year>2013</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>Christine</given-names>
            <surname>Senac</surname>
          </string-name>
          , Thomas Pellegrini, Florian Mouret, and
          <string-name>
            <given-names>Julien</given-names>
            <surname>Pinquier</surname>
          </string-name>
          .
          <article-title>Music feature maps with convolutional neural networks for music genre classification</article-title>
          .
          <source>In Proceedings of ACM Workshop on Content Based Multimedia Indexing (CBMI)</source>
          .
          <source>ACM, June</source>
          <volume>19</volume>
          -21,
          <year>2017</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>George</given-names>
            <surname>Tzanetakis</surname>
          </string-name>
          and
          <string-name>
            <given-names>P</given-names>
            <surname>Cook</surname>
          </string-name>
          .
          <article-title>Gtzan genre collection</article-title>
          .
          <source>Music Analysis, Retrieval and Synthesis for Audio Signals</source>
          ,
          <year>2002</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>Edith</given-names>
            <surname>Law</surname>
          </string-name>
          , Kris West,
          <string-name>
            <surname>Michael I Mandel</surname>
          </string-name>
          , Mert Bay, and
          <string-name>
            <given-names>J Stephen</given-names>
            <surname>Downie</surname>
          </string-name>
          .
          <article-title>Evaluation of algorithms using games: The case of music tagging</article-title>
          .
          <source>In ISMIR</source>
          , pages
          <fpage>387</fpage>
          -
          <lpage>392</lpage>
          ,
          <year>2009</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>