<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <journal-title-group>
        <journal-title>N. Lazzari);</journal-title>
      </journal-title-group>
      <issn pub-type="ppub">1613-0073</issn>
    </journal-meta>
    <article-meta>
      <title-group>
        <article-title>Tempo estimation from symbolic annotations with periodic functions</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Nicolas Lazzari</string-name>
          <email>nicolas.lazzari3@unibo.it</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Valentina Presutti</string-name>
          <email>valentina.presutti@unibo.it</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="editor">
          <string-name>Pachet, Valentina Presutti, Luc Steels</string-name>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Bologna 40124</institution>
          ,
          <country country="IT">Italy</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Dipartimento di Informatica, University of Pisa</institution>
          ,
          <addr-line>Largo B. Pontecorvo, 3, Pisa 56127</addr-line>
          ,
          <country country="IT">Italy</country>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>Dipartimento di Lingue, Letterature e Culture Moderne, University of Bologna</institution>
          ,
          <addr-line>Via Cartoleria, 5</addr-line>
        </aff>
      </contrib-group>
      <volume>000</volume>
      <fpage>0</fpage>
      <lpage>0002</lpage>
      <abstract>
        <p>The tempo estimation task has been traditionally performed on musical compositions mostly represented as audio or MIDI. Recent methods obtain near-perfect results. Nevertheless, the same methods applied to symbolic representations, such as textual chord annotations, result in inaccurate estimations. This hampers the harmonisation of heterogeneous datasets composed of symbolic annotations since a conversion step towards a common representation is needed. In this paper, we propose a novel method to obtain accurate tempo estimation on musical compositions encoded using textual symbolic annotations, relying on relevant cognitive and musicological theories. All the code is available at https://github.com/n28div/TEwPF.</p>
      </abstract>
      <kwd-group>
        <kwd>Music Information Retrieval</kwd>
        <kwd>Music tempo estimation</kwd>
        <kwd>Symbolic music</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>CEUR
ceur-ws.org</p>
    </sec>
    <sec id="sec-2">
      <title>1. Introduction</title>
      <p>
        Estimating the tempo of music compositions is a well-researched area driven by real-world
applications, from recommender systems to similarity measures [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. For instance, Gouyon
and Dixon [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ] perform genre classification using only tempo information. The results are
comparable to the performances of the same algorithm when using audio representations.
Indeed, the tempo of a composition has a great influence on the cognitive perception of
listeners and composers [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ] as well as computational applications [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. Faster compositions
tend to be perceived as happier while slower compositions as sadder. Moreover, it has
been observed that neural activity modulates in the presence of music with a faster tempo
[
        <xref ref-type="bibr" rid="ref4">4</xref>
        ], enhancing performances in reactive tasks.
      </p>
      <p>
        The tempo estimation task is defined as the identification of the frequency that humans
tap to a musical composition [5]. It is characterized by two sub-tasks: global and
local tempo estimation. The global estimation assumes that a constant tempo can be
observed throughout the whole musical composition [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ] while local estimation relaxes
such constraint and takes into account time fluctuations [ 6]. Global tempo estimation is a
      </p>
      <p>
        2022 Copyright for this paper by its authors. Use permitted under Creative Commons License Attribution 4.0
International (CC BY 4.0).
subset of local tempo estimation, since it can be extracted from a set of local estimations,
for instance by taking the median of the local estimations [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ].
      </p>
      <p>
        Audio representation is the most used input representation for the tempo estimation
task. Despite the large number of datasets proposed, obtaining high-quality recordings
represents the main issue in optimising and measuring Music Information Retrieval (MIR)
methods [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. This is mainly due to copyright issues. It is not possible to directly share
high-quality audio samples and it is often impossible to obtain the same exact recording
version and reproduce results.
      </p>
      <p>This has led to a growing interest in the use of symbolic annotations, which can
be openly shared and have been shown to outperform audio-based methods in chord
recognition [7] and music generation [8, 9] tasks. In order to efectively exploit symbolic
data, however, annotations need to be provided in a coherent form. Symbolic datasets
annotated by experts (e.g. [10, 11]) are often dificult to combine. Among the many
open challenges, a prominent issue is that annotations are mostly provided using absolute
timing (in seconds) rather than the corresponding symbolic notation.</p>
      <p>To the best of our knowledge, there are no previous attempts to perform tempo
estimation based on symbolic annotations that use absolute timing. This hampers the
harmonisation of heterogeneous symbolic datasets, such as ChoCo [12], since it is not
directly possible to normalise all the annotations to rhythmic notation. This poses a
limit on the development of methods that can benefit from symbolic representations. An
accurate tempo estimations method enables the estimation of the meter of a composition
[13, 14] which trivially allows the inference of the rhythmic notation. Alongside the
enhancement of other MIR methods, it would also enable experts to analyse large music
corpora, such as done by De Clercq and Temperley [15].</p>
      <p>In this paper, we propose a novel method that estimates the tempo of a musical
composition from its symbolic annotation expressed in absolute timing. Our method
is based on works that model tempo from a cognitive [16, 17] or musicological [18, 19]
perspective. The core intuition is to formulate a set of hypotheses and identify the
one that best fits an annotation, similarly to the work of Grosche and Müller [20]. By
exploiting techniques from the Computer Vision field, we estimate the local tempo of
each annotation and extract a global estimation from them. Our method outperforms all
the related works in this task, reaching an accuracy of 71%.</p>
      <p>The rest of this paper is organised as follows: in Section 2 we analyse the most
relevant related works; in Section 3 we describe our model; in Section 4 we introduce
the experimental setting and in Section 5 we show the obtained results and how they
compare to existing methods. Finally, in Section 6 we summarise the proposed work and
highlight possible extensions in future works.</p>
    </sec>
    <sec id="sec-3">
      <title>2. Related Works</title>
      <p>
        The task of tempo estimation is a prolific research area in the Music Information Retrieval
(MIR) field, driven by multiple real-world applications [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ] that can be grouped in
performance analysis, perceptual modelling, audio analysis, and performance synchronization
[
        <xref ref-type="bibr" rid="ref1">5, 1</xref>
        ].
      </p>
      <p>Most of the existing literature focuses on the use of audio representations. Proposed
methods conceptually consist in a pipeline involving two steps: the first step processes
audio to produce a representation that is then fed to the second step, which extracts the
ifnal tempo estimation. Given our focus on symbolic annotations, we only address the
most relevant and recent methods, focusing on how tempo extraction (i.e. the second
step) is implemented.</p>
      <p>A common approach is to use a bank of resonating comb filters to extract the periodicity
from the input signal [21, 22, 23, 24] Informally, a resonating comb filter detects the
presence of a specific frequency in a signal by summing the signal with a scaled and shifted
version of that signal. In this way, the filter is able to resonate at diferent multiples of
the target frequency, thus resulting in a promising method to extract periodicities from a
signal. Multiple filters that are tuned to diferent frequencies, i.e. a bank of filters, is
used to detect the most prominent periodicity.</p>
      <p>A similar approach is to use autocorrelation with a shifted version of the original signal
[25], where the autocorrelation operator computes a self-similarity between all the input
time steps. This results in a signal whose peaks correspond to the period of prominent
rhythmic groups [25].</p>
      <p>Both approaches are efective in identifying frequencies that are repeated within
the piece, but encounter dificulties in capturing the near-periodic information that
characterises a whole composition,i.e. they neglect the fact that a stable pulse should
be assumed for the whole composition. To overcome this issue, the Predominant Local
Pulse (PLP) was introduced in Grosche and Müller [20]. A PLP is obtained by sliding
sinusoidal kernels all over the signal and accumulating the result. Through the use of
kernels with varying frequencies, a mid-level representation that captures local periodic
information is obtained. This allows noisy signals that display near-periodic characteristics
to be captured by the model. The original PLP proposal [20] uses a short-time Fourier
transform to identify periodicities. This makes the method less suited when symbolic
annotations are used: the distribution of annotations is uneven time-wise, i.e. some time
regions might be very dense of annotations while other regions are much less crowded.
Furthermore, it is dificult to identify an optimal trade-of between time and frequency
resolution when symbolic annotations are used as input. Our method overcomes the
limitation of PLP on symbolic annotations by avoiding the use of the Fourier transform.</p>
      <p>
        Recent solutions implement either the feature extraction step [
        <xref ref-type="bibr" rid="ref5">23, 26, 27, 24</xref>
        ] or provide
the whole tempo estimation [
        <xref ref-type="bibr" rid="ref6 ref7 ref8 ref9">28, 29, 30, 31</xref>
        ] using neural networks.
      </p>
      <p>
        Despite the accuracy obtained by recent methods [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ], they sufer from octave errors
[
        <xref ref-type="bibr" rid="ref6">28</xref>
        ] (described in Section 4). Extending the time estimation step by using ML algorithms
[25] or additional data, such as style [
        <xref ref-type="bibr" rid="ref10">32</xref>
        ], has proven to be efective to prevent these
errors.
      </p>
      <p>
        Despite an initial interest in the extraction of tempo estimations from symbolic
representations from early methods [
        <xref ref-type="bibr" rid="ref11">5, 33</xref>
        ], little interest has been devoted to this task in
recent years. This may be due to the representation format itself, since the most popular
ones (e.g. MIDI or MusicXML) are designed to represent tempo explicitly. Recent
proposals focus on the extraction of more sophisticated rhythmic structures instead, such
as meter detection [
        <xref ref-type="bibr" rid="ref12">14, 34</xref>
        ], where global tempo information is assumed to exist.
      </p>
    </sec>
    <sec id="sec-4">
      <title>3. Methodology</title>
      <p>We focus on tempo estimation of a musical composition, based on symbolic annotations
provided by experts. The underlying assumption in our model is that the tempo of a
composition tends to be locally consistent - i.e. neighbouring annotations have a similar
tempo. Regardless of tempo fluctuations, we assume that it is possible to identify an
over-arching tempo that approximates the neighbourhood of each annotation, similarly
to [20].</p>
      <p>We test how well each tempo hypothesis explains the annotations by computing the
value of a periodic function  at each timestamp, where  is defined as a linear combination
of cosine functions  .̂ The frequency of each cosine function  ̂ is set such that each peak
matches the frequency of the hypothesised beats per minute (BPM) or a multiple of it.
In practice, we convert a BPM  into  / using the equation
 ̂=
2
60
(1)
where  ∈ ℕ and  ̂ is  in  / .</p>
      <p>
        Since textual symbolic annotations are most commonly used for chords and sections
(e.g. [10, 11]), they require a low rhythmic resolution, such as whole notes, quarter
notes or eighth notes. We assume that the tatum - “the smallest time interval between
successive notes in a rhythmic phrase” [
        <xref ref-type="bibr" rid="ref13">35</xref>
        ] - corresponds to eighth notes. Depending on
the application at hand, other rhythmic figure might be more appropriate to be used
as tatum, such as sixteenth [
        <xref ref-type="bibr" rid="ref13">35</xref>
        ] or thirty-second notes [
        <xref ref-type="bibr" rid="ref12">34</xref>
        ]. We compute  as the linear
combination of three cosine functions with  ∈ [ 1/2, 1, 2] where  = 1 corresponds to
quarter notes,  = 1/2 to half notes and  = 2 to eighth notes.
      </p>
      <p>The fitness function  is hence defined as
 () =  cos4( 1) +  cos4( 2) +  cos4( 1 )
2
(2)
where   is the timing hypothesis converted using Equation (1) and , ,  are the
coeficients for the linear combination. This approach corresponds to the event rule in
the Generative Theory of Tonal Music (GTTM) [18, 19], which states that beats that
align with event onsets should be preferred over other beats. Here the peaks of  are the
beats and the time of each annotation are the event onsets.</p>
      <p>Figure 1 depicts the function  computed for  = 120 BPM. The onsets occurring at
correct beat positions have higher values when compared to neighbouring positions.</p>
      <sec id="sec-4-1">
        <title>3.1. Estimating tempo</title>
        <p>
          It is possible to optimise  by maximising the sum of  computed at each time step.
This requires an additional assumption that reduces the generality of the method: a
global BPM must exist for each composition. This method would struggle in those
cases in which tempo changes throughout the whole piece (e.g. live performances). A
straightforward solution is to optimise over sliding windows of a composition. Empirically
we observe that this process is too sensitive to the initial tempo hypothesis. While it
is possible to obtain an initial hypothesis using a data-driven approach, for instance by
using the composition’s genre [
          <xref ref-type="bibr" rid="ref10">32</xref>
          ], the result would be heavily influenced by the tempo
bias on the training data [
          <xref ref-type="bibr" rid="ref1">1</xref>
          ].
        </p>
        <p>To estimate local tempo, we identify a search space composed of a finite number of
hypotheses  ⊆ [  ,   ] where | | ∈ ℕ and   and   are the bounds of the
search space. We identify an hypothesis resolution Δ and sample   Δ−   equally
distant points in the search space  . The choice of Δ depends on the trade-of between
computational complexity and precision of the solution, since it influences the dimension
of the search space. Intuitively, lower values of Δ produce higher resolution results while
higher values of Δ result in lower complexity of the search procedure. We investigate the
influence of the parameters   ,   , and Δ , in Section 4.</p>
        <p>For each annotation  we compute the fitness  of each hypothesis ℎ ∈  . The result
is a matrix  ∈ ℛ | |×|| where || is the number of annotated timestamps in  .  can
be interpreted as an image describing the fitness of each hypothesis at each available
timestamp. So far, the described method is similar to PLP [20].</p>
      </sec>
      <sec id="sec-4-2">
        <title>3.2. Local tempo coherency</title>
        <p>
          Diferently from PLP and according to our initial assumption, we update  such that
neighbouring annotations and neighbouring BPMs influence each other. We use two
Gaussian filters [
          <xref ref-type="bibr" rid="ref14">36</xref>
          ] over  , one along the rows and one along the columns. A Gaussian
iflter is obtained by computing the convolution of an image, in our case  , with a Gaussian
kernel. Informally, each element  ∈  is updated by computing a weighted mean of the
neighbouring elements, where the weights are a 2
        </p>
        <sec id="sec-4-2-1">
          <title>Gaussian distribution centred at  .</title>
          <p>The standard deviation of the Gaussian distribution,  , is used to compute the size of the
neighbourhood, roughly 3 . This approach overcomes the limit of the PLP method when
applied to symbolic annotations. The application of two distinct filters results in two
additional parameters:  
and   . The first (  
) is used to compute the dimension of
the kernel along the timing dimension, and the second (  ) takes into account the BPM
resolution Δ :</p>
          <p>=   /Δ .
(a) Standard 
(b)  with Gaussian filter applied with</p>
          <p>=
1 and   = 3.
when compared to Figure a. In both images the global BPM is correctly identified, however, the local
predictions depicted in Figure b are more stable across annotations.</p>
          <p>In Figure 2 a visual comparison between an unfiltered  (a) and a filtered  (b) is
shown.</p>
          <p>We obtain a local tempo estimation from  by computing the cumulative sum over
each column and taking the maximum for each row. Formally, this is defined as
() =</p>
          <p>argmax ∑  ,
| |
=1
(3)
with  the annotations in a composition and  the annotations preceding  . We compute the
global tempo by taking either the median value of the local estimations or the peak from
Parncutt [16]  () =
van Noorden et. al [17]  () =</p>
          <p>Definition
 () = 1
 () =</p>
          <p>1 2
 √2 ⋅ exp(− (−)2 2 )
exp(− 12 ( 1 log10(   ))2)</p>
          <p>
            1
√((  60 )2−( 60 )2)2−⋅( 60 )2
.  ,  ,  are parameters that need
the histogram of the local estimations, as suggested in [
            <xref ref-type="bibr" rid="ref1">1</xref>
            ]. In Section 4 we experiment
with both methods.
          </p>
        </sec>
      </sec>
      <sec id="sec-4-3">
        <title>3.3. Perceptionally-based weight for BPM hypothesis</title>
        <p>A common issue among all tempo estimation methods is the octave error, i.e. the
estimation of a multiple of the actual BPM. An example can be seen on Figure 2b: the
local estimations are coherent only until the 50th annotation, where a BPM twice the
correct one is detected.</p>
        <p>Octave errors are intrinsic to the tempo estimation task: for example the piece in Figure
2b, Helter Skelter by The Beatles, has a final section which is faster and more upbeat
when compared to the previous sections. This can results in a denser time distribution of
the annotations that leads to the detection of a faster tempo.</p>
        <p>
          To overcome this problem we update the fitness function  to weight specific hypothesis
diferently. We experiment with 4 diferent weighting schemas: a uniform distribution, a
Gaussian distribution, the model proposed by Parncutt [16] and the model proposed by
van Noorden &amp; Moelants1 [17]. The fitness function  is hence updated to
  () =  () ⋅  ()
(4)
where  () represents the weighting schema, normalised in the range [
          <xref ref-type="bibr" rid="ref1">0, 1</xref>
          ] using
Laplace smoothing [
          <xref ref-type="bibr" rid="ref15">37</xref>
          ].
        </p>
        <p>The implemented weighting schemas are described in 1. In Section 4 we experiment
with diferent combinations of parameter values.</p>
      </sec>
      <sec id="sec-4-4">
        <title>3.4. Metric preferences</title>
        <p>
          In the GTTM, Lerdahl and Jackendof define a rule for rhythmic grouping: longer onsets
should align with strong beats [18, 19]. The definition of strong beats depends on the
meter of a composition - defined as the hierarchical organisation of beats at diferent
time scales [
          <xref ref-type="bibr" rid="ref3">3</xref>
          ]. We take in consideration this rule by extending Equation (4) as follows:
f (, ) = [   () + cos2(  ) ⋅  , ⋯ ]
(5)
1Our implementation is slightly diferent from the original proposal. See 1 for more details.
where f
        </p>
        <p>
          (, ) is a vector whose elements are hypotheses specialised to a metric preference
 ;  is the length of the annotation at time  , and   is the BPM hypothesis converted
using Equation (1) with  =  . In the experiment presented in Section 4 we set  ∈ [
          <xref ref-type="bibr" rid="ref3 ref4">3, 4</xref>
          ]
to represent ternary and binary meters respectively, but it is trivial to support additional
groupings as well. To detect the meter that best fits the composition we maximise  in
Equation (5). Formally, this is defined as follows

||
=0
argmax ∑ f () + cos2( 1) ⋅ 
(6)
        </p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>4. Experimental Setup</title>
      <p>
        In this section, we experiment with the diferent combinations of methods and parameters,
described in Section 3. We perform an extensive set of experiments relying on Bayesian
Search [
        <xref ref-type="bibr" rid="ref16">38</xref>
        ] to find the best combination of parameters in a complete and eficient way.
The global tempo estimation methods described in Section 3.1 and the weighting schemas
of Section 3.3 are treated as parameters of the model.
      </p>
      <p>Component
 ()
Smoothing
Gaussian weight
Parncutt [16]
van Noorden et al. [17]</p>
      <sec id="sec-5-1">
        <title>2, provides an overview of the identified search spaces.</title>
        <p>
          We compare our model with the optimisation methods sketched in Section 3. We
use the Nelder-Mead algorithm [
          <xref ref-type="bibr" rid="ref17">39</xref>
          ] (based on gradients) with the initial solution set to
0.5 ∗ ( 
−
        </p>
        <p>
          ), and the Particle Swarm Optimisation (PSO, free from gradients)
method [
          <xref ref-type="bibr" rid="ref18">40</xref>
          ] to minimise the objective function. To provide a fair comparison, we search
for the best parameter ( 
∈ [10, 50],  
∈ [180, 300] and Δ ∈ [
          <xref ref-type="bibr" rid="ref1">1, 10</xref>
          ] , with Δ
the sliding window size) on the PSO model as well, including the hyper-parameters
(, ,  ∈ [
          <xref ref-type="bibr" rid="ref1">0, 1</xref>
          ]
        </p>
        <p>) in the search space as well.</p>
        <p>
          Finally, we compare our results with other publicly available methods: Böck et al.
[23], Böck et al. [
          <xref ref-type="bibr" rid="ref19">41</xref>
          ], and Grosche et al. [20]. They use, respectively, comb filters,
autocorrelation and PLP. Each method is either implemented using the madmom [
          <xref ref-type="bibr" rid="ref20">42</xref>
          ] or
essentia [
          <xref ref-type="bibr" rid="ref21">43</xref>
          ] libraries. All the related methods expect a signal representation as input.
Given our setting, we construct a signal with sampling rate   = 200  and manually add
peaks at the samples corresponding to each annotation.
        </p>
        <p>
          Each result is compared using the standard measures of Accuracy and Formal Octave
Errors (FOE). Accuracy is divided into two measures, Accuracy 1 and 2, defined as:
where  is the correct BPM,  the estimation, and  = 1 for Accuracy 1 and
 ∈ [
          <xref ref-type="bibr" rid="ref1 ref2 ref3">13 , 21 , 1, 2, 3</xref>
          ] for Accuracy 2. Both accuracy measures are binary measures: the
estimate is correct if it is within a 4% tolerance with respect to the true BPM. Diferently
from Accuracy 1, Accuracy 2 considers octave errors as correct estimations. As pointed
out by Schreiber et al. in [
          <xref ref-type="bibr" rid="ref1">1</xref>
          ], Accuracy 1 and 2 are of dificult interpretation: important
information, such as the most common octave errors, are hidden by the binary result.
We address this issue by analysing the FOE measures:
        </p>
        <p>
          )
argmin  1( ⋅ , )
1(, )|
2(, )|
where  is the same as the one for Accuracy 2. FOE measures are complementary to
Accuracy 1 and 2 and are easier to be visually interpreted. The search procedure optimises
Accuracy 1 over a subset of the Beatles [10] and RWC Pop [11] datasets provided by
the mir_data library [
          <xref ref-type="bibr" rid="ref22">44</xref>
          ]. We use 3-fold cross-validation on a subset of the data (80%)
randomly sampled and use the remaining data to evaluate the method in Section 5.
        </p>
      </sec>
    </sec>
    <sec id="sec-6">
      <title>5. Results</title>
      <p>
        In Table 3 the best results obtained from the experiments described in Section 4 are
described. Our method outperforms existing techniques on Accuracy 1, providing
estimations that are less flawed by octave errors. We obtain our best results by estimating
global tempo using the median operator. This provides additional evidence that this
operator is best suited to estimate global tempo from a list of local tempos, as also
argued by others [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. Interestingly, when it comes to Accuracy 2, the approach from
Böck et al. [23] (comb filter-based) achieves good results, outperforming some of our
experiments. The lower Accuracy 1 score, however, indicates that it is not a reliable
method for symbolic annotations. Using a weighting schema, as described in Section
3.3, correctly biases the method towards octave-correct estimations, as in the case of
Gaussian weighting, which largely improves Accuracy 2 results. When using the work
Weighting schema
      </p>
      <p>Nelder-MeadΔ=7
represented in bold. The best results overall are also underlined. A1 and A2 are used to refer to Accuracy
1 and Accuracy 2. Parncutt weighting schema refers to the work of Parncutt [16] while Resonance to the
work of van Noorden et al. [17].
from van Noorden et al. [17] and Parncutt [16], Accuracy 1 improvements over a uniform
distribution also result in a degraded Accuracy 2 score. This might happen because
these methods bias the tested hypotheses in a more aggressive way. A possible approach
to overcome this issue is to dampen the amount of added bias through an additional
parameter.</p>
      <p>In general, we consider the method that uses Parncutt weighting and the median
operator to be our best-performing experiment, given the Accuracy 1 result of 0.71. We
remark that regardless of the use of a numerical method (Nelder-Mead) or a meta-heuristic
one (PSO), the use of optimisation methods show worse performance in comparison to
other approaches.</p>
      <p>In Figure 3 the measures describing octave errors are reported. The distribution of
 1 in the best model is centred towards 0 when compared to other models. Analogously
in</p>
      <p>2, the distribution is more accurate when octave errors are also considered correct.
From the</p>
      <p>1 graph, it can be seen that related works are completely skewed towards
octave errors since the plot distribution is distributed along all the x-axis, while all of our
models are more prone to estimate BPMs that are 1/2 of the target BPM, since denser
clusters can be identified around the point</p>
      <p>
        −1 [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ].
      </p>
      <p>In Figure 4 the accuracy of our best methods, represented in Table 3, is plotted as a
function of the tolerance (4% in Equation Equation 7). The median approach converges
much quicker to the best results, providing further evidence that it is best suited to
extract a global tempo estimation.</p>
    </sec>
    <sec id="sec-7">
      <title>6. Conclusion</title>
      <p>We propose a novel method for tempo estimation of music composition, based on symbolic
text annotations, as input. The core idea is to exploit the linear combination of periodic
functions and techniques from the computer vision field, to find a BPM that best explains
the annotations. By relying on existing works from computational musicology and
cognitive perception of music, we devise a methodology that reaches an accuracy of 71%,
outperforming all existing approaches, applicable to this task.</p>
      <p>In Section 3 we propose a variation of our method using optimisation techniques to
obtain a tempo estimation. Even though the results are not comparable with our other
approaches, we argue that framing the tempo estimation task as an optimisation and
carefully designing an objective function can lead to robust and accurate methods.</p>
      <p>Regardless of the method used to obtain local tempo estimations, our results provide
additional evidence that the median operator is the best way to estimate global tempo
from a list of local estimations. The histogram method, however, can still be incorporated
into our approach when estimating local tempo. Instead of solving the maximisation
problem formulated in Equation (6), the histogram operator can be used to retrieve a list
of top-k candidates for each time step. We will investigate this option in future works.</p>
      <p>Finally, given the promising results from Böck et al. [23] in Table 3, an interesting
approach to explore is the combination of comb filters with our proposed method, to
enhance the performance on audio-representation as well.</p>
    </sec>
    <sec id="sec-8">
      <title>7. Acknowledgments</title>
      <p>This project has received funding from the FAIR Future Artificial Intelligence Research
foundation as part of the grant agreement MUR n. 341.
music modulate neural activity during reactive task performance, Psychology of
Music 42 (2014) 714–727. URL: https://doi.org/10.1177/0305735613490595. doi:10.1177/
0305735613490595. arXiv:https://doi.org/10.1177/0305735613490595.
[5] S. Dixon, Automatic extraction of tempo and beat from expressive
performances, Journal of New Music Research 30 (2001) 39–58. URL: https://www.
tandfonline.com/doi/abs/10.1076/jnmr.30.1.39.7119. doi:10.1076/jnmr.30.1.39.7119.
arXiv:https://www.tandfonline.com/doi/pdf/10.1076/jnmr.30.1.39.7119.
[6] G. Peeters, Time variable tempo detection and beat marking, in: Proceedings of
the 2005 International Computer Music Conference, ICMC 2005, Barcelona, Spain,
September 4-10, 2005, Michigan Publishing, 2005. URL: https://hdl.handle.net/2027/
spo.bbp2372.2005.186.
[7] T.-P. Chen, L. Su, Attend to chords: Improving harmonic analysis of symbolic
music using transformer-based models, Transactions of the International Society for
Music Information Retrieval (2021). doi:10.5334/tismir.65.
[8] M. Zeng, X. Tan, R. Wang, Z. Ju, T. Qin, T.-Y. Liu, MusicBERT: Symbolic music
understanding with large-scale pre-training, in: Findings of the Association for
Computational Linguistics: ACL-IJCNLP 2021, Association for Computational
Linguistics, Online, 2021, pp. 791–800. URL: https://aclanthology.org/2021.findings-acl.
70. doi:10.18653/v1/2021.findings-acl.70.
[9] D. von Rütte, L. Biggio, Y. Kilcher, T. Hofmann, FIGARO: controllable music
generation using learned and expert features, in: The Eleventh International
Conference on Learning Representations, ICLR 2023, Kigali, Rwanda, May 1-5,
2023, OpenReview.net, 2023. URL: https://openreview.net/pdf?id=NyR8OZFHw6i.
[10] M. Mauch, C. Cannam, M. Davies, S. Dixon, C. Harte, S. Kolozali, D. Tidhar,
M. Sandler, Omras2 metadata project 2009, in: 12th International Society for Music
Information Retrieval Conference, ISMIR, 2009.
[11] T. Cho, J. P. Bello, A feature smoothing method for chord recognition using
recurrence plots, in: 12th International Society for Music Information Retrieval
Conference, ISMIR, 2011.
[12] J. de Berardinis, A. Meroño-Peñuela, A. Poltronieri, V. Presutti, Choco: a chord
corpus and a data transformation workflow for musical harmony knowledge graphs,
Scientific Data 10 (2023) 641. URL: https://doi.org/10.1038/s41597-023-02410-w. doi:10.
1038/s41597-023-02410-w.
[13] A. Volk, The study of syncopation using inner metric analysis: Linking theoretical
and experimental analysis of metre in music, Journal of New Music Research
37 (2008) 259–273. URL: https://doi.org/10.1080/09298210802680758. doi:10.1080/
09298210802680758. arXiv:https://doi.org/10.1080/09298210802680758.
[14] W. B. de Haas, A. Volk, Meter detection in symbolic music using inner metric analysis,
in: M. I. Mandel, J. Devaney, D. Turnbull, G. Tzanetakis (Eds.), Proceedings of
the 17th International Society for Music Information Retrieval Conference, ISMIR
2016, New York City, United States, August 7-11, 2016, 2016, pp. 441–447. URL:
https://wp.nyu.edu/ismir2016/wp-content/uploads/sites/2294/2016/07/033_Paper.pdf.
[15] T. De Clercq, D. Temperley, A corpus analysis of rock harmony, Popular Music 30
(2011) 47–70.
[16] R. Parncutt, A Perceptual Model of Pulse Salience and Metrical Accent in
Musical Rhythms, Music Perception 11 (1994) 409–464. URL: https://doi.org/10.2307/
40285633. doi:10.2307/40285633.
arXiv:https://online.ucpress.edu/mp/articlepdf/11/4/409/145282/40285633.pdf.
[17] L. van Noorden, D. Moelants, Resonance in the perception of musical
pulse, Journal of New Music Research 28 (1999) 43–66. URL: https://www.
tandfonline.com/doi/abs/10.1076/jnmr.28.1.43.3122. doi:10.1076/jnmr.28.1.43.3122.
arXiv:https://www.tandfonline.com/doi/pdf/10.1076/jnmr.28.1.43.3122.
[18] D. Temperley, D. D. Sleator, Modeling meter and harmony: A
preferencerule approach, Comput. Music. J. 23 (1999) 10–27. URL: https://doi.org/10.1162/
014892699559616. doi:10.1162/014892699559616.
[19] F. Lerdahl, R. S. Jackendof, A Generative Theory of Tonal Music, The MIT Press,
1996. URL: https://doi.org/10.7551/mitpress/12513.001.0001. doi:10.7551/mitpress/
12513.001.0001.
[20] P. Grosche, M. Müller, A mid-level representation for capturing dominant tempo
and pulse information in music recordings, in: K. Hirata, G. Tzanetakis, K. Yoshii
(Eds.), Proceedings of the 10th International Society for Music Information Retrieval
Conference, ISMIR 2009, Kobe International Conference Center, Kobe, Japan,
October 26-30, 2009, International Society for Music Information Retrieval, 2009,
pp. 189–194. URL: http://ismir2009.ismir.net/proceedings/OS2-3.pdf.
[21] E. D. Scheirer, Tempo and beat analysis of acoustic musical signals, The Journal of
the Acoustical Society of America 103 (1998) 588–601.
[22] A. Klapuri, A. J. Eronen, J. Astola, Analysis of the meter of acoustic musical signals,
IEEE Trans. Speech Audio Process. 14 (2006) 342–355. URL: https://doi.org/10.1109/
TSA.2005.854090. doi:10.1109/TSA.2005.854090.
[23] S. Böck, F. Krebs, G. Widmer, Accurate tempo estimation based on recurrent
neural networks and resonating comb filters, in: M. Müller, F. Wiering (Eds.),
Proceedings of the 16th International Society for Music Information Retrieval
Conference, ISMIR 2015, Málaga, Spain, October 26-30, 2015, 2015, pp. 625–631.</p>
      <p>URL: http://ismir2015.uma.es/articles/196_Paper.pdf.
[24] S. Böck, M. E. P. Davies, Deconstruct, analyse, reconstruct: How to improve tempo,
beat, and downbeat estimation, in: J. Cumming, J. H. Lee, B. McFee, M. Schedl,
J. Devaney, C. McKay, E. Zangerle, T. de Reuse (Eds.), Proceedings of the 21th
International Society for Music Information Retrieval Conference, ISMIR 2020,
Montreal, Canada, October 11-16, 2020, 2020, pp. 574–582. URL: http://archives.
ismir.net/ismir2020/paper/000223.pdf.
[25] G. Percival, G. Tzanetakis, Streamlined tempo estimation based on autocorrelation
and cross-correlation with pulses, IEEE ACM Trans. Audio Speech Lang. Process.
22 (2014) 1765–1776. URL: https://doi.org/10.1109/TASLP.2014.2348916. doi:10.1109/
TASLP.2014.2348916.
[26] S. Böck, M. E. P. Davies, P. Knees, Multi-task learning of tempo and beat:
Learning one to improve the other, in: A. Flexer, G. Peeters, J. Urbano, A. Volk
(Eds.), Proceedings of the 20th International Society for Music Information Retrieval
Conference, ISMIR 2019, Delft, The Netherlands, November 4-8, 2019, 2019, pp.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>H.</given-names>
            <surname>Schreiber</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Urbano</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Müller</surname>
          </string-name>
          ,
          <article-title>Music tempo estimation: Are we done yet?</article-title>
          ,
          <source>Trans. Int. Soc. Music. Inf. Retr</source>
          .
          <volume>3</volume>
          (
          <year>2020</year>
          )
          <article-title>111</article-title>
          . URL: https://doi.org/10.5334/tismir.43. doi:
          <volume>10</volume>
          .5334/tismir.43.
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>F.</given-names>
            <surname>Gouyon</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Dixon</surname>
          </string-name>
          ,
          <article-title>Dance music classification: A tempo-based approach</article-title>
          ,
          <source>in: ISMIR</source>
          <year>2004</year>
          , 5th International Conference on Music Information Retrieval, Barcelona, Spain,
          <source>October 10-14</source>
          ,
          <year>2004</year>
          , Proceedings,
          <year>2004</year>
          . URL: http://ismir2004.ismir.
          <source>net/proceedings/ p091-page-501-paper151.pdf.</source>
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>J. D.</given-names>
            <surname>McAuley</surname>
          </string-name>
          , Tempo and Rhythm, University of California Press,
          <year>2010</year>
          , pp.
          <fpage>165</fpage>
          -
          <lpage>199</lpage>
          . doi:
          <volume>10</volume>
          .1007/978- 1-
          <fpage>4419</fpage>
          - 6114-
          <issue>3</issue>
          _
          <fpage>6</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>D. T.</given-names>
            <surname>Bishop</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. J.</given-names>
            <surname>Wright</surname>
          </string-name>
          ,
          <string-name>
            <surname>C. I. Karageorghis</surname>
          </string-name>
          ,
          <article-title>Tempo and intensity of pre-task 486-493</article-title>
          . URL: http://archives.ismir.net/ismir2019/paper/000058.pdf.
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [27]
          <string-name>
            <given-names>H. F.</given-names>
            <surname>Aarabi</surname>
          </string-name>
          , G. Peeters,
          <article-title>Deep-rhythm for global tempo estimation in music</article-title>
          , in: A.
          <string-name>
            <surname>Flexer</surname>
            , G. Peeters,
            <given-names>J.</given-names>
          </string-name>
          <string-name>
            <surname>Urbano</surname>
            ,
            <given-names>A</given-names>
          </string-name>
          . Volk (Eds.),
          <source>Proceedings of the 20th International Society for Music Information Retrieval Conference</source>
          ,
          <string-name>
            <surname>ISMIR</surname>
          </string-name>
          <year>2019</year>
          , Delft,
          <source>The Netherlands, November 4-8</source>
          ,
          <year>2019</year>
          ,
          <year>2019</year>
          , pp.
          <fpage>636</fpage>
          -
          <lpage>643</lpage>
          . URL: http://archives.ismir. net/ismir2019/paper/000077.pdf.
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [28]
          <string-name>
            <given-names>H.</given-names>
            <surname>Schreiber</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Müller</surname>
          </string-name>
          ,
          <article-title>A single-step approach to musical tempo estimation using a convolutional neural network</article-title>
          , in: E.
          <string-name>
            <surname>Gómez</surname>
            ,
            <given-names>X.</given-names>
          </string-name>
          <string-name>
            <surname>Hu</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          <string-name>
            <surname>Humphrey</surname>
          </string-name>
          , E. Benetos (Eds.),
          <source>Proceedings of the 19th International Society for Music Information Retrieval Conference</source>
          ,
          <source>ISMIR 2018</source>
          , Paris, France,
          <source>September 23-27</source>
          ,
          <year>2018</year>
          ,
          <year>2018</year>
          , pp.
          <fpage>98</fpage>
          -
          <lpage>105</lpage>
          . URL: http://ismir2018.ircam.fr/doc/pdfs/141_Paper.pdf.
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [29]
          <string-name>
            <given-names>H.</given-names>
            <surname>Schreiber</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Müller</surname>
          </string-name>
          ,
          <article-title>Musical tempo and key estimation using convolutional neural networks with directional filters</article-title>
          , CoRR abs/
          <year>1903</year>
          .10839 (
          <year>2019</year>
          ). URL: http: //arxiv.org/abs/
          <year>1903</year>
          .10839. arXiv:
          <year>1903</year>
          .10839.
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [30]
          <string-name>
            <surname>M. S. de Oliveira de Souza</surname>
            , P. N. de Souza Moura,
            <given-names>J.</given-names>
          </string-name>
          <string-name>
            <surname>Briot</surname>
          </string-name>
          ,
          <article-title>Music tempo estimation via neural networks - A comparative analysis</article-title>
          ,
          <source>CoRR abs/2107</source>
          .09208 (
          <year>2021</year>
          ). URL: https://arxiv.org/abs/2107.09208. arXiv:
          <volume>2107</volume>
          .
          <fpage>09208</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [31]
          <string-name>
            <given-names>X.</given-names>
            <surname>Sun</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Q.</given-names>
            <surname>He</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Gao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <article-title>Musical tempo estimation using a multi-scale network</article-title>
          , in: J. H.
          <string-name>
            <surname>Lee</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Lerch</surname>
            ,
            <given-names>Z.</given-names>
          </string-name>
          <string-name>
            <surname>Duan</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          <string-name>
            <surname>Nam</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          <string-name>
            <surname>Rao</surname>
            ,
            <given-names>P. van Kranenburg</given-names>
          </string-name>
          ,
          <string-name>
            <surname>A</surname>
          </string-name>
          . Srinivasamurthy (Eds.),
          <source>Proceedings of the 22nd International Society for Music Information Retrieval Conference</source>
          ,
          <string-name>
            <surname>ISMIR</surname>
          </string-name>
          <year>2021</year>
          , Online, November 7-
          <issue>12</issue>
          ,
          <year>2021</year>
          ,
          <year>2021</year>
          , pp.
          <fpage>682</fpage>
          -
          <lpage>689</lpage>
          . URL: https://archives.ismir.net/ismir2021/paper/000085.pdf.
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [32]
          <string-name>
            <given-names>F.</given-names>
            <surname>Hörschläger</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Vogl</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Böck</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Knees</surname>
          </string-name>
          ,
          <article-title>Addressing tempo estimation octave errors in electronic music by incorporating style information extracted from wikipedia</article-title>
          ,
          <source>in: Proceedings of the Sound and Music Computing Conference (SMC)</source>
          , Maynooth, Ireland,
          <year>2015</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [33]
          <string-name>
            <given-names>H.</given-names>
            <surname>Takeda</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Nishimoto</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Sagayama</surname>
          </string-name>
          ,
          <article-title>Rhythm and tempo analysis toward automatic music transcription</article-title>
          ,
          <source>in: Proceedings of the IEEE International Conference on Acoustics, Speech, and Signal Processing</source>
          ,
          <string-name>
            <surname>ICASSP</surname>
          </string-name>
          <year>2007</year>
          , Honolulu, Hawaii, USA, April
          <volume>15</volume>
          -
          <issue>20</issue>
          ,
          <year>2007</year>
          , IEEE,
          <year>2007</year>
          , pp.
          <fpage>1317</fpage>
          -
          <lpage>1320</lpage>
          . URL: https://doi.org/10.1109/ICASSP.
          <year>2007</year>
          .
          <volume>367320</volume>
          . doi:
          <volume>10</volume>
          .1109/ICASSP.
          <year>2007</year>
          .
          <volume>367320</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [34]
          <string-name>
            <given-names>A.</given-names>
            <surname>McLeod</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Steedman</surname>
          </string-name>
          ,
          <article-title>Meter detection and alignment of MIDI performance</article-title>
          , in: E.
          <string-name>
            <surname>Gómez</surname>
            ,
            <given-names>X.</given-names>
          </string-name>
          <string-name>
            <surname>Hu</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          <string-name>
            <surname>Humphrey</surname>
          </string-name>
          , E. Benetos (Eds.),
          <source>Proceedings of the 19th International Society for Music Information Retrieval Conference</source>
          ,
          <source>ISMIR 2018</source>
          , Paris, France,
          <source>September 23-27</source>
          ,
          <year>2018</year>
          ,
          <year>2018</year>
          , pp.
          <fpage>113</fpage>
          -
          <lpage>119</lpage>
          . URL: http://ismir2018.ircam.fr/ doc/pdfs/136_Paper.pdf.
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [35]
          <string-name>
            <given-names>A.</given-names>
            <surname>Klapuri</surname>
          </string-name>
          ,
          <article-title>Musical meter estimation and music transcription</article-title>
          , in: Cambridge Music Processing Colloquium, Citeseer,
          <year>2003</year>
          , pp.
          <fpage>40</fpage>
          -
          <lpage>45</lpage>
          . doi:https://citeseerx.ist.psu. edu/viewdoc/summary?doi=10.1.1.77.8559.
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [36]
          <string-name>
            <given-names>L. G.</given-names>
            <surname>Shapiro</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G. C.</given-names>
            <surname>Stockman</surname>
          </string-name>
          , Computer vision, Pearson,
          <year>2001</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [37]
          <string-name>
            <surname>C. D. Manning</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          <string-name>
            <surname>Raghavan</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          <string-name>
            <surname>Schütze</surname>
          </string-name>
          , Introduction to information retrieval, Cambridge University Press,
          <year>2008</year>
          . URL: https://nlp.stanford.edu/IR-book/pdf/irbookprint. pdf. doi:
          <volume>10</volume>
          .1017/CBO9780511809071.
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [38]
          <string-name>
            <given-names>J.</given-names>
            <surname>Snoek</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Larochelle</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R. P.</given-names>
            <surname>Adams</surname>
          </string-name>
          ,
          <article-title>Practical bayesian optimization of machine learning algorithms</article-title>
          , in: P. L.
          <string-name>
            <surname>Bartlett</surname>
            ,
            <given-names>F. C. N.</given-names>
          </string-name>
          <string-name>
            <surname>Pereira</surname>
            ,
            <given-names>C. J. C.</given-names>
          </string-name>
          <string-name>
            <surname>Burges</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          <string-name>
            <surname>Bottou</surname>
            ,
            <given-names>K. Q.</given-names>
          </string-name>
          <string-name>
            <surname>Weinberger</surname>
          </string-name>
          (Eds.),
          <source>Advances in Neural Information Processing Systems 25: 26th Annual Conference on Neural Information Processing Systems 2012. Proceedings of a meeting held December 3-6</source>
          ,
          <year>2012</year>
          ,
          <string-name>
            <given-names>Lake</given-names>
            <surname>Tahoe</surname>
          </string-name>
          , Nevada, United States,
          <year>2012</year>
          , pp.
          <fpage>2960</fpage>
          -
          <lpage>2968</lpage>
          . URL: https://proceedings.neurips.cc/paper/2012/ hash/05311655a15b75fab86956663e1819cd-Abstract.html.
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [39]
          <string-name>
            <given-names>F.</given-names>
            <surname>Gao</surname>
          </string-name>
          , L. Han,
          <article-title>Implementing the nelder-mead simplex algorithm with adaptive parameters</article-title>
          ,
          <source>Comput. Optim. Appl</source>
          .
          <volume>51</volume>
          (
          <year>2012</year>
          )
          <fpage>259</fpage>
          -
          <lpage>277</lpage>
          . URL: https://doi.org/10.1007/ s10589-010-9329-3. doi:
          <volume>10</volume>
          .1007/s10589-010-9329-3.
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          [40]
          <string-name>
            <given-names>K.</given-names>
            <surname>Hussain</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. N. M.</given-names>
            <surname>Salleh</surname>
          </string-name>
          , S. Cheng, Y. Shi,
          <article-title>Metaheuristic research: a comprehensive survey</article-title>
          ,
          <source>Artif. Intell. Rev</source>
          .
          <volume>52</volume>
          (
          <year>2019</year>
          )
          <fpage>2191</fpage>
          -
          <lpage>2233</lpage>
          . URL: https://doi.org/10.1007/ s10462-017-9605-z. doi:
          <volume>10</volume>
          .1007/s10462-017-9605-z.
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          [41]
          <string-name>
            <given-names>S.</given-names>
            <surname>Böck</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Schedl</surname>
          </string-name>
          ,
          <article-title>Enhanced beat tracking with context-aware neural networks</article-title>
          ,
          <source>in: Proc. Int. Conf. Digital Audio Efects</source>
          ,
          <year>2011</year>
          , pp.
          <fpage>135</fpage>
          -
          <lpage>139</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          [42]
          <string-name>
            <given-names>S.</given-names>
            <surname>Böck</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Korzeniowski</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Schlüter</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Krebs</surname>
          </string-name>
          ,
          <string-name>
            <surname>G.</surname>
          </string-name>
          <article-title>Widmer, madmom: a new Python Audio and Music Signal Processing Library</article-title>
          ,
          <source>in: Proceedings of the 24th ACM International Conference on Multimedia, Amsterdam</source>
          , The Netherlands,
          <year>2016</year>
          , pp.
          <fpage>1174</fpage>
          -
          <lpage>1178</lpage>
          . doi:
          <volume>10</volume>
          .1145/2964284.2973795.
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          [43]
          <string-name>
            <given-names>D.</given-names>
            <surname>Bogdanov</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Wack</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Gómez</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Gulati</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Herrera</surname>
          </string-name>
          ,
          <string-name>
            <given-names>O.</given-names>
            <surname>Mayor</surname>
          </string-name>
          , G. Roma, J. Salamon,
          <string-name>
            <given-names>J. R.</given-names>
            <surname>Zapata</surname>
          </string-name>
          ,
          <string-name>
            <surname>X.</surname>
          </string-name>
          <article-title>Serra, ESSENTIA: an open-source library for sound and music analysis</article-title>
          , in: A.
          <string-name>
            <surname>Jaimes</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          <string-name>
            <surname>Sebe</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          <string-name>
            <surname>Boujemaa</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          <string-name>
            <surname>Gatica-Perez</surname>
            ,
            <given-names>D. A.</given-names>
          </string-name>
          <string-name>
            <surname>Shamma</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          <string-name>
            <surname>Worring</surname>
          </string-name>
          , R. Zimmermann (Eds.), ACM Multimedia Conference, MM '
          <fpage>13</fpage>
          ,
          <string-name>
            <surname>Barcelona</surname>
          </string-name>
          , Spain,
          <source>October 21-25</source>
          ,
          <year>2013</year>
          , ACM,
          <year>2013</year>
          , pp.
          <fpage>855</fpage>
          -
          <lpage>858</lpage>
          . URL: https: //doi.org/10.1145/2502081.2502229. doi:
          <volume>10</volume>
          .1145/2502081.2502229.
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          [44]
          <string-name>
            <surname>R. M. Bittner</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          <string-name>
            <surname>Fuentes</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          <string-name>
            <surname>Rubinstein</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Jansson</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          <string-name>
            <surname>Choi</surname>
          </string-name>
          , T. Kell,
          <article-title>mirdata: Software for reproducible usage of datasets</article-title>
          , in: A.
          <string-name>
            <surname>Flexer</surname>
            , G. Peeters,
            <given-names>J.</given-names>
          </string-name>
          <string-name>
            <surname>Urbano</surname>
            ,
            <given-names>A</given-names>
          </string-name>
          . Volk (Eds.),
          <source>Proceedings of the 20th International Society for Music Information Retrieval Conference</source>
          ,
          <string-name>
            <surname>ISMIR</surname>
          </string-name>
          <year>2019</year>
          , Delft,
          <source>The Netherlands, November 4-8</source>
          ,
          <year>2019</year>
          ,
          <year>2019</year>
          , pp.
          <fpage>99</fpage>
          -
          <lpage>106</lpage>
          . URL: http://archives.ismir.net/ismir2019/paper/000009.pdf.
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>