<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <journal-title-group>
        <journal-title>March</journal-title>
      </journal-title-group>
    </journal-meta>
    <article-meta>
      <title-group>
        <article-title>Towards a Hybrid Recommendation System for a Sound Library</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Jason Smith</string-name>
          <email>jsmith775@gatech.edu</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Dillon Weeks</string-name>
          <email>dweeks7@gatech.edu</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Mikhail Jacob</string-name>
          <email>mikhail.jacob@gatech.edu</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Jason Freeman</string-name>
          <email>jason.freeman@gatech.edu</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Brian Magerko</string-name>
          <email>magerko@gatech.edu</email>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Center for Music Technology, Georgia Institute of Technology</institution>
          ,
          <addr-line>Atlanta, GA</addr-line>
          ,
          <country country="US">USA</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>School of Interactive Computing, Georgia Institute of Technology</institution>
          ,
          <addr-line>Atlanta, GA</addr-line>
          ,
          <country country="US">USA</country>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>School of Literature</institution>
          ,
          <addr-line>Media, and, Communication</addr-line>
          ,
          <institution>Georgia Institute of Technology</institution>
          ,
          <addr-line>Atlanta, GA</addr-line>
          ,
          <country country="US">USA</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2019</year>
      </pub-date>
      <volume>20</volume>
      <issue>2019</issue>
      <abstract>
        <p>Recommendation systems are widespread in music distribution and discovery services but far less common in music production software such as EarSketch, an online learning environment that engages learners in writing code to create music. The EarSketch interface contains a sound library that learners can access through a browser pane. The current implementation of the sound browser includes basic search and filtering functionality but no mechanism for sound discovery, such as a recommendation system. As a result, users have historically selected a small subsection of sounds in high frequencies, leading to lower compositional diversity. In this paper, we propose a recommendation system for the EarSketch sound browser which uses collaborative filtering and audio features to suggest sounds.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>CCS CONCEPTS</title>
      <p>• Human-centered computing → User interface design; •
Applied computing → Sound and music computing.
recommendation systems, interface design, music</p>
      <p>IUI Workshops’19, March 20, 2019, Los Angeles, USA
Copyright ©2019 for the individual papers by the papers’ authors. Copying permitted
for private and academic purposes. This volume is published and copyrighted by its
editors.</p>
    </sec>
    <sec id="sec-2">
      <title>INTRODUCTION</title>
      <p>
        EarSketch [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ] is an online environment for learning computer
programming and audio loop-based music composition. Students write
JavaScript or Python scripts to algorithmically generate musical
compositions. The user interface borrows design cues from both
integrated development environments (IDEs) and digital audio
workstation (DAW) software, combining a code editor and console with
a multi-track audio timeline and sound browser. EarSketch has
primarily been used in high school and college computer science
classrooms, with over 300,000 users to date [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ].
      </p>
      <p>
        In previous research in EarSketch classrooms, significant
relationships have been found between student perceptions of
authenticity – including their desire to share personally expressive work
with others – and student attitudes towards computing [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ].
Exploration of a larger number of musical ideas – including the sounds
that form the building blocks of student compositions in EarSketch
– may magnify a student’s capacity to create personally expressive
compositions.
      </p>
      <p>EarSketch contains a library of over 3,500 sounds for students to
use in their compositions. The sounds were created by musicians
Richard Devine and Young Guru specifically for EarSketch and
consist of multi-measure audio loops that are separated by instrument
and span over 20 popular musical genres. However, a statistical
analysis of scripts written by users showed that the vast majority of
user projects used only a small subset of the sound library. Feedback
from EarSketch users (found in the interviews section) showed that
their lack of exploration was primarily the result of the dificulty
in finding sounds that appealed to them. We propose, therefore,
that providing users with an easier mechanism for exploring the
sound library will enable them to find and use audio loops that
spur further musical creativity and personal expression, while
ultimately furthering their learning about music and coding through
EarSketch.</p>
      <p>
        We have explored the addition of a recommendation (or
recommender) system, after conducting user studies, as a method of
encouraging users to explore more of the EarSketch sound library
in their scripts. Recommendation systems are widespread in music
distribution and discovery platforms (where they operate at the
song level) but far less common in music production workflows
(where they could operate at the sound clip level). Recommendation
systems suggest content to users that is most likely to appeal to
them based on profiles of their preferences as well as content that
they would most likely find novel, diverse, and unexpectedly useful
(serendipitous) [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. EarSketch could use such a recommendation
system to automatically search through its sound library to find
relevant sounds that encourage the user to explore novel, diverse,
and serendipitous regions of the sound library.
      </p>
      <p>
        Recommendation generation techniques include collaborative
ifltering, content-based filtering, and hybrid techniques.
Collaborative filtering [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ] involves comparing the current user to previous
users in order to generate recommendations from what similar
users in the past selected (for example [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ]). Content-based
filtering compares inherent properties of content to recommend items,
such as with the use of audio feature-based deep learning [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ] and
calculation of short sample similarity metrics [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ]. We can use a
hybrid approach that combines both techniques to generate
recommendations.
      </p>
      <p>
        Some previous recommendation systems for sounds employed
the Freesound sample library [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ]. These projects used feature
similarity calculations without co-usage statistics [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ] or textual
metadata to augment recommendations [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ]. The proposed system for
EarSketch difers from these examples by combining only audio
similarity and co-usage to generate recommendations, reserving
genre labels for manual user filtering.
      </p>
      <p>
        In this article, we present our initial research on a
recommendation system for discovering new sounds for use in EarSketch. The
main contributions discussed are:
• An initial user-centered design process for systematically
understanding how best to add an audio loop
recommendation system into the EarSketch environment, including, the
way users currently use the sound browser, the challenges
to using it successfully, the kinds of recommendations users
desire, and the best way to present users with
recommendations.
• The initial application of a hybrid (collaborative and
contentbased filtering) recommendation system for sounds in a
digital audio workstation, in contrast to song recommendation
systems. This is a first step towards improving user
exploration of the EarSketch sound library according to the user
requirements and design principles arising from the initial
user-centered design process.
• A proposed methodology for evaluating both the success of
the recommendation system in providing users with
relevant, novel, diverse, and serendipitous recommendations [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]
and the relative importance of the diferent factors used to
generate recommendations, as well as the usability of the
sound browser with the recommendation system added.
The remainder of the paper describes the details of the user-centered
design process for adding a recommendation system to the
EarSketch sound browser and the initial prototype of the hybrid
recommendation system resulting from that design process. The paper
concludes by discussing the planned evaluation methodology,
limitations of the current prototype, and future work.
2
      </p>
    </sec>
    <sec id="sec-3">
      <title>USER RESEARCH AND INTERFACE DESIGN</title>
      <p>An initial user study was conducted in order to gain a systematic
understanding of how best to add a recommendation system into
the EarSketch sound browser. This included understanding the
diferent ways that users used the sound browser, the challenges
they faced to use it successfully, the kinds of recommendations users
desired, and the best ways to present recommendations to users.
The study resulted in a set of requirements for the recommendation
system and a redesign of the sound browser interface integrating
the generated recommendations.
The sound browser experience prior to the addition of a
recommendation system included sound folders that consisted of a title and a
list of sounds corresponding to that title. For example, the sound
folder titled "DUBSTEP 140 BPM DUBBASS WOBBLE" included a
list of "DUBSTEP BASS WOBBLE" sounds underneath it, followed
by other sound folders and their associated sounds. This list was
navigated via scrolling and sounds were distributed across
multiple pages within the browser. The user had the ability to favorite
and preview these sounds from within the browser as well. The
user also had the ability to discover sounds in the library via text
search from the search bar along with the functionality to filter
these sounds by artists, instruments, and genre.
2.2</p>
    </sec>
    <sec id="sec-4">
      <title>Interviews and Survey of EarSketch</title>
    </sec>
    <sec id="sec-5">
      <title>Students</title>
      <p>Four qualitative interviews were conducted with undergraduate
students in an introductory programming course at a four-year
college to explore current EarSketch users’ challenges, behaviors,
and interactions with the sound browser. This was done to identify
the best opportunities for the recommendation system to fit their
needs. These interviews were utilized to gather qualitative data
such as reported behaviors, motivations behind those behaviors,
opportunities for future designs and a recommendation integration.
A quantitative survey was sent out to the same undergraduate class
and received 55 responses. The survey was used to determine the
prevalence of identified behaviors and preferences.</p>
      <p>Participants reported being more inclined to use the Instrument
and Genre filters than the Artist filter. In addition, users expressed
their desire for a Key and Beats-Per-Minute (BPM) filter. This
suggested the need to prioritize recommendations based on
instruments, genres, keys, and BPM in the future.</p>
      <p>
        Users reported that it was hard to discover groups of sounds
they considered to be a good recommendation. They considered
strong recommendations to be sounds that they liked that also fit
in their script (relevant) and that they had not heard before (novel)
or were not expecting (serendipitous). Discovering sounds similar
to previously used sounds was of lesser importance to them. This
confirmed that those users desired recommendations that were in
accordance with the recommendation system goals defined by [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ].
3
      </p>
    </sec>
    <sec id="sec-6">
      <title>HYBRID RECOMMENDATION SYSTEM</title>
      <p>A set of design principles arose as a result of these user studies.
Recommendations were to be relevant, novel, diverse, and
serendipitous. Additionally, users were interested in getting
recommendations in the interface separated into diferent categories (e.g.
"Sounds That Fit Your Tastes" and "Discover Diferent Kinds of
Sounds"). Users were also interested in getting recommendations
matching with semantic features of the sounds in their
work-inprogress compositions (e.g. instrument, genre, key, and BPM).</p>
      <p>The initial recommendation system we have developed does not
yet support the entire set of user requirements that were
illuminated by the user studies. It does combine collaborative filtering
(using a statistical analysis of sound usage in past user scripts) and
content-based filtering (using extracted audio features) to increase
the relevance and novelty of the generated recommendations.
Recommendations are generated as follows:
(1) The algorithm takes in one or more sounds as its input. This
input is the set of sounds that are already a part of a user’s
work-in-progress script/composition.
(2) The algorithm then generates a first list of sounds from
the EarSketch sound library (the co-usage list) that have
commonly been used in the past with the input sounds in
scripts by any user.
(3) The algorithm then uses audio features of the sounds in
the co-usage list to create a second list containing other
sounds in the sound library that are acoustically similar to
the sounds in the co-usage list (the similarity list).
(4) The algorithm removes sounds from the similarity list that
have been commonly used with sounds in the co-usage list.
(5) Finally, the algorithm chooses sounds from the similarity list
to present to the user as recommendations.</p>
      <p>The co-usage list is an example of collaborative filtering (see
the collaborative filtering section) and adds relevance to the
generated recommendations by ensuring that recommendations are
compatible with the set of sounds in the user’s work-in-progress
script/composition. The usage of the similarity list (rather than just
the co-usage list) is an example of content-based filtering (see the
content filtering section). The removal of sounds from the similarity
list (that are commonly used with the co-usage list) adds novelty
to the recommendations. The approach described here attempts to
address diversity and serendipity of the generated
recommendations, but explicit measures to ensure and evaluate these qualities
is planned for future work (see future work).
3.1</p>
    </sec>
    <sec id="sec-7">
      <title>Collaborative Filtering</title>
      <p>
        The input to the collaborative filtering is the collection of sounds
already being used in an active script at the time of recommendation
generation. We take an item-based approach involving only an
analysis of previous co-usage between sounds [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ]. We take this
approach to impose minimal collection of user information, such as
user demographics and profile usage history, protecting EarSketch’s
primarily school-aged user base and conforming with its privacy
policy [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ]. The system returns a co-usage list of sounds in order of
co-usage frequency. This co-usage is calculated using a sample set
of 20,000 user scripts. Any sounds that are also in the input list are
excluded to ensure that commonly co-used input sounds do not
simply recommend each other.
3.2
      </p>
    </sec>
    <sec id="sec-8">
      <title>Content-based Filtering</title>
      <p>We compare two audio features to find sounds acoustically
similar to the items in the co-usage list. These recommendations are
the final output of the system. Recommended sounds are chosen
based on their similarity to the most commonly co-used sounds
comparing two properties of the audio signal — Short-Time Fourier
Transform features and Mel-Frequency Cepstral Coeficients. The
sounds are compared using the euclidean distance between their
feature vectors, taken from the first 2 seconds of 48000 sample rate
audio with a 1024-point Hann window and normalized for tempo.</p>
      <sec id="sec-8-1">
        <title>Short-Time Fourier Transform Features DST FT is the eu</title>
        <p>
          clidean distance between the spectral density of two sounds,
calculated using the librosa STFT function [
          <xref ref-type="bibr" rid="ref8">8</xref>
          ]. This function
allows us to evaluate time-based similarities between sounds,
and recommend sounds with similar function in a rhythmic
context.
        </p>
      </sec>
      <sec id="sec-8-2">
        <title>Mel-Frequency Cepstral Coeficients DM F CC is the euclidean</title>
        <p>
          distance between the short-term power spectrum of two
sounds, using the librosa MFCC function [
          <xref ref-type="bibr" rid="ref8">8</xref>
          ] [
          <xref ref-type="bibr" rid="ref10">10</xref>
          ]. This
compares sounds in terms of temporally-independent energy,
and acts as genre or instrument groupings.
        </p>
        <p>
          Both features have been chosen due to their common usage in
music information retrieval [
          <xref ref-type="bibr" rid="ref6">6</xref>
          ].
3.3
        </p>
      </sec>
    </sec>
    <sec id="sec-9">
      <title>Recommendation Algorithms</title>
      <p>This design aims to generate recommendations of sounds that are
serendipitous to the user by not having high co-usage, and relevant
through acoustical similarity to sounds that do. Diversity in
recommendations is possible by including a high number of co-used
sounds of a variety of styles. The multiple stages of randomness in
both models, while not guaranteeing novelty, allow for diferent
recommendations to be generated for the same combinations of
inputs.</p>
      <p>N represents an arbitrary factor limiting the amount of results
gathered at diferent steps in the algorithms, and will be empirically
determined during evaluation. The value of each variable labeled N
in the below sections can be manipulated separately. This includes
the lengths of the list of final recommendations, the co-usage list,
and the similarity list.</p>
      <p>The initial prototype of the recommender system is designed for
use in standalone ofline applications in addition to integration with
the main EarSketch browser. Two recommendation algorithms were
developed: one for live, real-time recommendation calculations
and the other for faster server-side calculations. The first model,
the dynamic model, conducts all calculations ofline using
precomputed audio features to generate a list of recommendations
for any combination of sounds. The static model, intended for
online use, combines pre-computed lists of recommendations for
any individual sounds to generate a single recommendation list.
3.3.1 Dynamic. The most commonly used sounds in conjunction
with any of the input sounds parsed from a user script are found
collectively using the collaborative filtering paradigm in the
collaborative filtering section. Each commonly co-used sound is then
compared to all other sounds in the EarSketch library, and a
recommendation score for each is generated as the following equation:
where DST FT = normalized STFT euclidean distance, DM F CC =
normalized normalized STFT euclidean distance, and U =
normalized co-usage.</p>
      <p>
        Additionally, STFT and MFCC distance from the original input
samples are added or subtracted from the final recommendation
score. This to generate recommendations that are either
acoustically similar or diferent from the ones already found in the user
script at the time recommendation. The sounds with the highest N
recommendation scores are stored and joined together in a single
similarity list. A random selection of N recommendations is chosen
from the highest N normalized recommendation scores in the
master list, with higher priority given to the highest recommendations
through fitness proportionate selection [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ].
      </p>
      <p>Static. The static model difers from the dynamic model in that
it uses a pre-computed list of similarity lists generated for each
individual sound in EarSketch, in order to make the recommendation
algorithm less computationally intensive for server-side
deployment. The lists for any combination of input sounds are joined
together into a master list, and any duplicate sounds have their
recommendation scores added and balanced by a factor of the square
root of the number of lists. This method of balancing is in order
to assign higher value to the strongest recommendations without
drowning out the others, and is another scalable parameter that
will be evaluated in future work (see the future work section). A
random selection of N recommendations is chosen with higher
priority given to the highest recommendations as with the dynamic
model.</p>
    </sec>
    <sec id="sec-10">
      <title>FUTURE WORK</title>
      <p>This algorithm is an exploratory stage of development and we plan
to expand it along with the interface design with respect to current
limitations information gained from user testing.
4.1</p>
    </sec>
    <sec id="sec-11">
      <title>Recommendation System</title>
      <p>The recommendation generation process will be modified to
improve how it explicitly addresses its goals of relevance, novelty,
diversity, and serendipity. Recommendation relevance will be
improved by adding semantic metadata tags to the sounds, like
instrument, genre, key, and BPM, and using those parameters (in addition
to co-usage statistics and feature similarity) to select sounds.
Novelty will be explicitly optimized for by measuring the distance
between sounds in the lists and ensuring that recommendations
are intentionally selected to be diferent from previously generated
recommendations by some threshold novelty value N. Additionally,
the calculations between audio features will be performed with
operations and statistical measures other than euclidean distance, and
will incorporate higher-level features such as rhythm. Similarly for
a threshold diversity value D, recommendations would be chosen
by adding sounds to a candidate set such that each new addition is
at least D distance from every other item already in the set. Finally,
serendipity will be explicitly optimized for by first collecting data
searching for recommendations that are relevant but with low
cousage frequencies (indicating that they are rarely used together).
Finally, each of the four recommendation generation goals will be
weighted in order to tailor recommendations to diferent situations
or diferent recommendation folders.
4.2</p>
      <p>
        Proposed Evaluation
4.2.1 Recommendation System. Participants in a user study will
empirically refine the various iterations of the recommendation
system using diferent output-limiting values of N, and diferent
relative weighting of DM F CC and DST FT . Additionally, they will
be asked to choose sounds from the recommendation system and
rate them in terms of relevance, novelty, diversity, and serendipity
[
        <xref ref-type="bibr" rid="ref13">13</xref>
        ] for a combination of input sounds. The sounds they choose will
be represented by the recommendation scores generated by each
system iteration, in order to evaluate the weightings independently.
Additionally, qualitative questions will reveal user opinions on
other design aspects, like how many recommendations users want
to see at once.
4.2.2 Interface Redesign. The current redesign has not been
properly tested in a real world scenario, thus potential usability issues
may arise with the navigation, language, and recommendation
types. We will conduct moderated usability testing and record
users’ sessions interacting with a high-fidelity prototype while
a researcher prompts them with tasks to complete. This testing will
allow more information regarding EarSketch users’ perceptions
of a ’good’ recommendation and how users will actually utilize
these recommendations. As we move toward understanding how
to recommend sounds to our user and better facilitate the
exploration and discovery of sounds within EarSketch, our near-term
goal is to iterate and improve on the proposed EarSketch redesign
to accommodate recommendations.
      </p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <surname>Charu</surname>
            <given-names>C Aggarwal</given-names>
          </string-name>
          et al.
          <year>2016</year>
          .
          <article-title>Recommender systems</article-title>
          . Springer.
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>Thomas</given-names>
            <surname>Bäck</surname>
          </string-name>
          .
          <year>1996</year>
          .
          <article-title>Evolutionary Algorithms in Theory and Practice: Evolution Strategies, Evolutionary Programming, Genetic Algorithms</article-title>
          . Oxford University Press, Inc., New York, NY, USA.
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>S.</given-names>
            <surname>Chang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Abdul</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Chen</surname>
          </string-name>
          , and
          <string-name>
            <given-names>H.</given-names>
            <surname>Liao</surname>
          </string-name>
          .
          <year>2018</year>
          .
          <article-title>A Personalized Music Recommendation System using Convolutional Neural Networks Approach</article-title>
          .
          <source>In IEEE International Conference on Applied System Invention (ICASI)</source>
          .
          <source>IEEE, St. Petersburg Russia</source>
          ,
          <fpage>47</fpage>
          -
          <lpage>49</lpage>
          . https://doi.org/10.1109/ICASI.
          <year>2018</year>
          .8394293
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4] Bram de Jong.
          <year>2005</year>
          . Freesound. https://freesound.org
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>Brian</given-names>
            <surname>Magerko Jason Freeman</surname>
          </string-name>
          .
          <year>2011</year>
          . EarSketch. http://earsketch.gatech.edu/ landing/
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>Alexander</given-names>
            <surname>Lerch</surname>
          </string-name>
          .
          <year>2012</year>
          .
          <article-title>An Introduction to Audio Content Analysis: Applications in Signal Processing</article-title>
          and Music
          <string-name>
            <surname>Informatics</surname>
          </string-name>
          (1st ed.). Wiley-IEEE Press.
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>Brian</given-names>
            <surname>Magerko</surname>
          </string-name>
          , Jason Freeman, Tom
          <string-name>
            <surname>Mcklin</surname>
            ,
            <given-names>Mike Reilly</given-names>
          </string-name>
          ,
          <string-name>
            <given-names>Elise</given-names>
            <surname>Livingston</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Scott</given-names>
            <surname>Mccoid</surname>
          </string-name>
          ,
          <string-name>
            <given-names>and Andrea</given-names>
            <surname>Crews-Brown</surname>
          </string-name>
          .
          <year>2016</year>
          .
          <article-title>Earsketch: A steam-based approach for underrepresented populations in high school computer science education</article-title>
          .
          <source>ACM Transactions on Computing Education (TOCE) 16</source>
          ,
          <issue>4</issue>
          (
          <year>2016</year>
          ),
          <fpage>14</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <surname>Brian</surname>
            <given-names>McFee</given-names>
          </string-name>
          ,
          <string-name>
            <given-names>Colin</given-names>
            <surname>Rafel</surname>
          </string-name>
          , Dawen Liang, Daniel PW Ellis,
          <string-name>
            <surname>Matt</surname>
            <given-names>McVicar</given-names>
          </string-name>
          ,
          <string-name>
            <given-names>Eric</given-names>
            <surname>Battenberg</surname>
          </string-name>
          , and
          <string-name>
            <given-names>Oriol</given-names>
            <surname>Nieto</surname>
          </string-name>
          .
          <year>2015</year>
          .
          <article-title>librosa: Audio and music signal analysis in python</article-title>
          .
          <source>In Proceedings of the 14th python in science conference</source>
          .
          <volume>18</volume>
          -
          <fpage>25</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <surname>Tom</surname>
            <given-names>McKlin</given-names>
          </string-name>
          ,
          <string-name>
            <given-names>Brian</given-names>
            <surname>Magerko</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Taneisha</given-names>
            <surname>Lee</surname>
          </string-name>
          , Dana Wanzer, Doug Edwards, and
          <string-name>
            <given-names>Jason</given-names>
            <surname>Freeman</surname>
          </string-name>
          .
          <year>2018</year>
          .
          <article-title>Authenticity and Personal Creativity: How EarSketch Afects Student Persistence</article-title>
          .
          <source>In Proceedings of the 49th ACM Technical Symposium on Computer Science Education. ACM</source>
          ,
          <volume>987</volume>
          -
          <fpage>992</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>Paul</given-names>
            <surname>Mermelstein</surname>
          </string-name>
          .
          <year>1976</year>
          .
          <article-title>Distance measures for speech recognition, psychological and instrumental</article-title>
          .
          <source>Pattern recognition and artificial intelligence 116</source>
          (
          <year>1976</year>
          ),
          <fpage>374</fpage>
          -
          <lpage>388</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>Sergio</given-names>
            <surname>Oramas</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.C.</given-names>
            <surname>Ostuni</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Di</surname>
          </string-name>
          <string-name>
            <surname>Noia</surname>
          </string-name>
          , Xavier Serra, and
          <string-name>
            <given-names>E. Di</given-names>
            <surname>Sciascio</surname>
          </string-name>
          .
          <year>2016</year>
          .
          <article-title>Sound and Music Recommendation with Knowledge Graphs</article-title>
          .
          <source>ACM Transactions on Intelligent Systems and Technology (TIST) 8</source>
          (
          <issue>10</issue>
          /
          <year>2016</year>
          2016),
          <fpage>1</fpage>
          -
          <lpage>21</lpage>
          . https: //doi.org/10.1145/2926718
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>Gerard</given-names>
            <surname>Roma</surname>
          </string-name>
          and
          <string-name>
            <given-names>Xavier</given-names>
            <surname>Serra</surname>
          </string-name>
          .
          <year>2015</year>
          .
          <article-title>Music performance by discovering community loops</article-title>
          .
          <source>In Proceedings of the Web Audio Conference (WAC)</source>
          , Paris.
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <given-names>E.</given-names>
            <surname>Shakirova</surname>
          </string-name>
          .
          <year>2017</year>
          .
          <article-title>Collaborative Filtering for Music Recommender System</article-title>
          .
          <source>In IEEE Conference of Russian Young Researchers in Electrical and Electronic</source>
          Engineering (EIConRus).
          <source>IEEE, St. Petersburg Russia</source>
          ,
          <fpage>548</fpage>
          -
          <lpage>550</lpage>
          . https://doi.org/10.1109/ EIConRus.
          <year>2017</year>
          .7910613
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <given-names>Kai</given-names>
            <surname>Siedenburg</surname>
          </string-name>
          and
          <string-name>
            <given-names>Daniel</given-names>
            <surname>Müllensiefen</surname>
          </string-name>
          .
          <year>2007</year>
          .
          <article-title>Modeling Timbre Similarity of Short Music Clips</article-title>
          .
          <source>Frontiers in psychology 8</source>
          ,
          <issue>1</issue>
          (April
          <year>2007</year>
          ),
          <fpage>36</fpage>
          -
          <lpage>44</lpage>
          . https: //doi.org/10.3389/fpsyg.
          <year>2017</year>
          .00639
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>