<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Linked Data Collection and Analysis Platform of Audio Features</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Yuri Uehara</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Takahiro Kawamura</string-name>
          <email>kawamura@ohsuga.is.uec.ac.jp</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Shusaku Egami</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Yuichi Sei</string-name>
          <email>seiuny@uec.ac.jp</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Yasuyuki Tahara</string-name>
          <email>tahara@uec.ac.jp</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Akihiko Ohsuga</string-name>
          <email>ohsuga@uec.ac.jp</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Graduate School of Information Systems, University of Electro-Communications</institution>
          ,
          <addr-line>Tokyo</addr-line>
          ,
          <country country="JP">Japan</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2015</year>
      </pub-date>
      <fpage>2</fpage>
      <lpage>5</lpage>
      <abstract>
        <p>Audio features extracted from music are commonly used in music information retrieval (MIR), but there is no open platform for data collection and analysis of audio features. Therefore, we build the platform for the data collection and analysis for MIR research. On the platform, we represent the music data with Linked Data. In this paper, we rst investigate the frequency of the audio features used in previous studies on MIR for designing the Linked Data schema. Then, we build a platform, that automatically extracts the audio features and music metadata from YouTube URIs designated by users, and adds them to our Linked Data DB. Finally, the sample queries for music analysis and the current record of music registrations in the DB are presented.</p>
      </abstract>
      <kwd-group>
        <kwd>Linked Data</kwd>
        <kwd>audio features</kwd>
        <kwd>music information retrieval</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>
        Recently, there are a large number of studies on music. Music Information
Retrieval (MIR) deals with music on computers and has been studied in various
ways [
        <xref ref-type="bibr" rid="ref2">1</xref>
        ]. In these studies, audio features extracted from music are frequently
used, however, there is no open platform for collecting data including the audio
features for music analysis. Therefore, we propose the platform for MIR research
in this paper.
      </p>
      <p>On the platform, we used Linked Data format, since it is suitable for complex
searches for audio features and songs-related metadata. Note that this platform
is designed for music-related researchers and developers, who intend to
analyze music information and create their own applications, e.g., recommendation
mechanism. Use of a listener is beyond the scope of this paper.</p>
      <p>Schema Design of Music Information
In this section, designing Linked Data schema, including audio features and
music metadata is described.
2.1</p>
    </sec>
    <sec id="sec-2">
      <title>Selection of audio features</title>
      <p>
        Audio features refer to the characteristics of the music, such as Tempo
representing the speed of the track, the features used in MIR studies vary. For example,
Osmalskyj et al. used Tempo and Loudness to identify cover songs [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]. Luo et
al. used the audio features Pitch, Zero crossing rate, etc. to detect of common
mistakes in novice violin playing [4].
      </p>
      <p>Thus, we investigated the frequency of the audio features used in previous
MIR studies. We collected 114 papers published in the International Society of
Music Information retrieval (ISMIR)1 in 2015, which is the top conference in
the eld of MIR. Then, we selected some of the audio features according to the
policies: Features appeared just once in the publications can be ignored, etc.
2.2</p>
    </sec>
    <sec id="sec-3">
      <title>Design of the schema</title>
      <p>We de ned original properties for selected audio features, excluding Key and
Mode, since there were no existing properties for them, or the properties are
not appropriate for our purpose. Then, we classi ed properties of audio features
into some classes for making them easy to use. Table 1 shows the classes and
properties corresponding to the audio features.</p>
      <p>We designed the music schema with the video id (URI) of YouTube. In Fig.
1, the id: dvgZkm1xWPE indicates a song \Viva La Vida" by Coldplay. In the
graph, the id node links to the classes of audio features and then links to each
audio feature. Also, we added some degrees for categorizing numerical values in
the features. The tempo has tmarks based on tempo values2, which is a measure
of the speed marks: Slow means 39 or less bpm, Largo means 40 { 49 bpm, etc.</p>
      <p>In addition, we extended the schema of music metadata. In the graph of
metadata, the video id node links to the class of metadata, and then links to
the detailed value, as well as the graph for the audio features. Also, some nodes
such as the artist name are linked to the external DBs like DBpedia.</p>
      <p>There is the graphs of \Viva La Vida" by Coldplay, and thus the audio
features graph and the metadata graph can be linked with the video id of YouTube.
1 http://www.ismir.net/
2 http://www.sii.co.jp/music/try/metronome/01.html</p>
      <p>AFlatMajor
graph, the id node li nrrddff:s:hhtttptp:/://w/wwwww. w.w3.3o.orgrg/1/929090/00/20/12/r2d-rfd-sf-cshyenmtaadx#-niso# features and then linkskeytso.owel#aKecyh</p>
      <p>ks to the classes of au
in the features. The lowenergy, the rmsenergy, and the brightness have radf:tcypleass
audio feature. Also, wmmuoe:s-hvtotcp::/h/ptdtupre:l/./odwrgw/wosn.oothomlsouggya/.mis.odu/eec.gacr.jpe/meussicf/voocrabuclaaryt#egorizinrdgfs:lanbeul merical values 213.35^^xsd:double
ad e
means 192 – 208 bpm, Fast means over 209 bpm.
bTyhe0.t1e,mtphoehzaesrotcmroasrskshba1as38s.0ea5d5^c^olxasnds:dstoeubbmleypo10v0a,luaAelnlesgd3ro, twhheicrholilosmffoa:AFmhlataMesaajosraurcelaosfs tbhye20s01p0e0e0d. rdf:value
mus-voc:
marks: Slow means 39 orrdfl:evasluse bpm, Largo means 40 – 49 bpm, Lentoclamssifeicaatinons 5 0zerocross</p>
      <p>mus-voc:tmark mo:key
– 55 bpm, Adagio means 56 – 62 bpm, Andante means 63 – 75 bpm, mMus-ovodc:zeerroacrtooss</p>
      <p>tempo mus-voc:tempo
means 76 – 95 bpm, Allegretto means 96 – 119 bpm, Allegro means 120 – 151</p>
      <p>audiofeatures mus-voc:timbre
bpm, Vivace means 152 – 175 bpm, Presto means 176 – 191 bpm, Prestissimo
timbre
mus-voc:dynamics
In addition, we extended the schmeums-voc:featuremsusic metadata, that includes notmus-voc:brightness
a of</p>
      <p>dynamics
only song title, artist name, etc. in the mMus-voc:loweOnerngytology, but also lyricist name,</p>
      <p>usic
dvgZkm1xWPE
cd name for the complex search for music information. In thmues-vgocr:rampsehnerogyf metadbarigthatness</p>
      <p>lowenergy
, the video id node links to the class of metadata, and thermnselnienrgkys to the detailed
value, as well as the graph for rtdfh:vaelueaudio features. Also, some nodecmslaussss-ivfuoiccac:tihon as rtdfh:vealue</p>
      <p>mus-voc: rdf:value
artist name are linked to the external DBs likcleassDificBatiopn edia.</p>
      <p>There is the graphs of 0“.5V5^i^vxsad:dLouableVida” by Coldplay0.,31 a^^nxsdd:dtouhblue s the tw0.3o graph0s.447773 ^^xsd:double</p>
      <p>0.5
can be linked with the video id of YouTube.</p>
      <p>mus-voc:rolloff
http://purl.org/NET/c4dm/
3564.23^^xsd:double
rdf:value
rolloff</p>
      <p>3000
mus-voc:
classification
0.4
3</p>
      <p>3 Music Information Extracting
The system architecture for our music information extraction is shown in Fig.
2, and its workflow is indicated by the number 1 to 11 as follows.</p>
      <p>The system architecture for our music information extraction is shown in Fig.
3, and its work ow is indicated by the number 1 to 11.</p>
      <p>Client</p>
      <p>Add songs
Browser</p>
      <p>Input
the YouTube video URL
Songs retrieval
Browser</p>
      <p>SPARQL query
1
9
10
11</p>
      <p>Server</p>
      <p>Servlet 1
- Cache the
data of
YouTube video
- Start the
MATLAB
program</p>
      <p>MATLAB
MIRtoolbox</p>
      <p>4</p>
      <p>Analysis of
2 audio features 3 MySQL</p>
      <p>RDB</p>
      <p>7</p>
      <p>Servlet 2
- Acquire the meta data
RDF DB -Acquire the audio features</p>
      <p>- Create RDF
Virtuoso 8 -Add to Virtuoso
6</p>
      <p>5
Last.fm</p>
      <p>YouTube
1. Download the video data from the YouTube video</p>
      <p>URI designated by a user in a web browser.
2. Call the MATLAB process that analyzes audio</p>
      <p>features in the video file.
3. Store the obtained audio features in an RDB,</p>
      <p>MySQL.
4. Call the RDF create program.
5. Obtain the music information for the video from</p>
      <p>the YouTube website.
6. Search the music metadata using Last.fm API.
7. Query the audio features of the video for MySQL.
8. Convert the metadata and audio features to RDF</p>
      <p>graphs, and store them in an RDF store, Virtuoso.
9. Notify the completion to the user.
10. Submit a simple SPARQL query for confirmation.
11. Returns the evidence of the inclusion of new
sub</p>
      <p>graphs corresponding to the video.
1. Download the video data from the YouTube video URI designated by a user in a web browser.
2. Call thepMAuTbLAlBicproucessesrtshatcaannalyezaessialuydioefxetateunreds itnhtehe mvidueosifcilei.nformation on the platform. However,
3. Store the obtained audio features in an RDB, MySQL.
4. Call thewRDeF cdriesatceaprrdogrtahm.e video les after extracting the audio features, and thus we believe
5. Obtain the music information for the video from the YouTube website.
6. Search thtehmiussipcrmeotcadeastas udsiongesLansto.ftm cAPaIu.se any legal or moral problems.</p>
      <p>Our system obtains videos to analyze audio features from YouTube, and so
The work ow is divided into several phases. The rst phase is for analyzing
3 http://www.sii.co.jp/music/try/metronome/01.html
the audio features of the YouTube video, and the second phase is for acquiring
the metadata of the YouTube video. Then, the third phase converts the metadata
and audio features to RDF graphs, and the RDF graphs are stored in Virtuoso
database. Then, the</p>
      <p>nal phase is for the con rmation of newly added graphs.
4</p>
      <p>Example of Music Analysis
The current number of music registered in the platform is 1073 and the
number of triples automatically extracted for representing the audio features and</p>
      <p>paper.
! A copyright form, signed by one author on behalf of all of the authors of the
! A readme giving the name and email address of the corresponding author.
[SPARQL Query]
PREFIX mus-voc:&lt;http://www.ohsuga.is.uec.ac.jp/music/
vocabulary#&gt;
PREFIX mo:&lt;http://purl.org/ontology/mo/&gt;
SELECT ?artist_x ?title_x ?brightness_x
WHERE { ?metadata rdfs:label ?title .</p>
      <p>?resource mus-voc:meta ?metadata .</p>
      <p>6 Checklist of Items to be?rSeseounrcte tmuos-vVoco:fleuatmurees ?Efedatiutreosr.s
the metadata for that music is?fe2at0ur8e5s8mu.s-vToch:teimbprela?ttimfborre m. is publicly available at
?timbre mus-voc:brightness ?brightnessc .
httpH:/e/rewiswawch.oechksliustgoaf.iesv.euryetch.iangc.tjhpe/v?mborliuugmhstinecese/sdc.itrodrf:rvealquueir?ebrsifgrhotnmessy.ou:</p>
      <p>?brightnessc_x rdf:value ?brightness_x.</p>
      <p>I!ntThhise fisenacltiLAoTnE,Xwsoeurscheofiwlesthe r?etismubrlet_sx mousf-vsooc:mbrieghetnxesas m?brpiglhetnqesusce_xri.es on the platform,
?features_x mus-voc:timbre ?timbre_x .
and how the music Linked Data?recsoaurnce_bx emusu-vsoec:dfeaftourres M?feIatRur.eIsn_x .the SPARQL Query,
! A final PDF file ?resource_x mus-voc:meta ?metadata_x .
we speci ed the audio feature B?rmeitgahdattna_exsrsdfso:lfabtehle?tistloen_xg. \Hello, Goodbye" by The
! A copyright form, signed by one a?umetthaodartao_nx bmoe:hMuaslifcoArftaisllt o?fMutshiecAartuitsht_oxrs. of the
Beatles, and search other songs,?MiunsicwArthisitc_hxrdtfhs:elabvelal?uaretisot_fx t.he Brightness is similar
paper. FILTER regex(?title, "Hello Goodby") . }
to the speci ed song. As thOeRDErReBsYu(lt, we get 5 songs, in which the Brightness has
! A readme giving the name aInF(d?ebmrigahitlnaesdsd&lt;re?sbsriogfhttnheses_cxo,rresponding author.
the similar degree in Table 2. ?brightness_x - ?brightness,</p>
      <p>?brightness - ?brightness_x )
[SPARQL Query] ) LIMIT 5
PREFIX mus-voc:&lt;http://www.ohsuga.is.uec.ac.jp/music/
vocabulary#&gt;
PREFIX mo:&lt;http://purl.org/ontology/mo/&gt;
SELECT ?artist_x ?title_x ?brightness_x
WHERE { ?metadata rdfs:label ?title .</p>
      <p>?resource mus-voc:meta ?metadata .
?resource mus-voc:features ?features .
?features mus-voc:timbre ?timbre .
?timbre mus-voc:brightness ?brightnessc .
?brightnessc rdf:value ?brightness.
?brightnessc_x rdf:value ?brightness_x.
?timbre_x mus-voc:brightness ?brightnessc_x .
?features_x mus-voc:timbre ?timbre_x .
?resource_x mus-voc:features ?features_x .
?resource_x mus-voc:meta ?metadata_x .
?metadata_x rdfs:label ?title_x .
?metadata_x mo:MusicArtist ?MusicArtist_x .
?MusicArtist_x rdfs:label ?artist_x .</p>
      <p>FILTER regex(?title, "Hello Goodby") . }
ORDER BY (</p>
      <p>IF( ?brightness &lt; ?brightness_x,
?brightness_x - ?brightness,
?brightness - ?brightness_x )
) LIMIT 5
5</p>
      <p>Conclusion and Future Work</p>
      <p>Table 1. Result of submitting the SPARQL query
In this paper, we prop oarstisetdx a platittlefxorm forbrpighrtnoesvsxiding audio features and the music</p>
      <p>The Beatles Can’t Buy Me Love 0.553889
metadata to MIR reseCWaohrlditcpnleahyy H.ouston PNreivnecresGsiveOUfpChina 00..556509073896</p>
      <p>Ft. Rihanna
In future, we plan Ltaody Gpagraovi dJuedas more sop0.5h502i7s9ticated examples and applications</p>
      <p>The Beatles Penny Lane 0.550221
of music information analysis, which will encourage the expansion of the music
Linked Data to music researchers and developers.</p>
    </sec>
    <sec id="sec-4">
      <title>Acknowledgments. artTist hxis</title>
      <p>The Beatles
wotitrle kx was</p>
      <p>Can’t Buy Me Love
sburightnesosxrted
pp
0.553889
by JSPS</p>
      <p>KAKENHI Grant
Research on</p>
      <p>Music:0. Foreword. IPSJ magazine \Joho Shori". 57 (6), 504{505. (2016)
aware Music Recommendation with Serendipity Using Semantic Relations.
Proceedings of 3rd Joint International Semantic Technology Conference. 17{32. (2013)
Song Identi cation. Proceedings of the 16th International Society for Music
Infor</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          <source>Numbers 16K12411, 16Whitn0ey0Ho4us1ton Nev1er6GiKve Up K</source>
          <volume>9</volume>
          ,
          <fpage>125330</fpage>
          ..559786
          <source>Coldplay Princess Of China 0.560039 Ft. Rihanna Lady Gaga Judas 0.550279 The Beatles Penny Lane</source>
          <volume>0</volume>
          .
          <fpage>550221</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          1.
          <string-name>
            <surname>Kitahara</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Nagano</surname>
          </string-name>
          , H.: Advancing Information Sciences through
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Osmalskyj</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Foster</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Dixon</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Embrechts</surname>
            ,
            <given-names>J.J.</given-names>
          </string-name>
          :Combining Features for Cover
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>