<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Automatic composition of descriptive music: A case study of the relationship between image and sound</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Luc a Mart n-Gomez</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Javier Perez-Marcos</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Mar a Navarro Caceres</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>BISITE Research Group, University of Salamanca</institution>
          ,
          <addr-line>Calle Espejo, S/N, 37007 Salamanca</addr-line>
          ,
          <country country="ES">Spain</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>Human beings establish relationships with the environment mainly through sight and hearing. This work focuses on the concept of descriptive music, which makes use of sound resources to narrate a story. The Fantasia lm, produced by Walt Disney was used in the case study. One of its musical pieces is analyzed in order to obtain the relationship between image and music. This connection is subsequently used to create a descriptive musical composition from a new video. Naive Bayes, Support Vector Machine and Random Forest are the three classi ers studied for the model induction process. After an analysis of their performance, it was concluded that Random Forest provided the best solution; the produced musical composition had a considerably high descriptive quality.</p>
      </abstract>
      <kwd-group>
        <kwd>Descriptive music</kwd>
        <kwd>automatic composition</kwd>
        <kwd>image</kwd>
        <kwd>video</kwd>
        <kwd>classi cation</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>
        Human beings establish all kind of relationships with their environment. These
relationships are made thanks to the human senses, such as sight and hearing,
which extract information from the surroundings. Cognitive processes allow us
to assimilate the information [
        <xref ref-type="bibr" rid="ref32">32</xref>
        ]. Throughout history, images and music have
been two of the most common means of interaction between humans and their
surroundings.
      </p>
      <p>
        Some authors have already studied the way the human being uses the senses
of vision and audition to establish relationships. Thus, numerous approaches
concerning the perception and cognition of music can be found in the literature.
There are researches in the eld of Psychology that study human reactions when
a person is listening to music [
        <xref ref-type="bibr" rid="ref1 ref24">1, 24</xref>
        ]. Furthermore, some techniques such as
sentiment analysis [
        <xref ref-type="bibr" rid="ref29">29</xref>
        ] and Brain Computer Interfaces (BCIs) [
        <xref ref-type="bibr" rid="ref34">34</xref>
        ] are used with
the same purpose. On the other hand, there are some techniques that study the
human perception of images. Speci cally, a work called Eyetracking [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ] measures
eye positions and analyzes its movements with the purpose of understanding the
way a person processes an image. In [
        <xref ref-type="bibr" rid="ref33">33</xref>
        ], some attentional regions are discovered
in each image after the extraction of some descriptors.
      </p>
      <p>
        Some works have already intended to capture human behavior in this regard,
by using di erent approaches to compose music, with the aid of techniques such
as synesthesia [
        <xref ref-type="bibr" rid="ref11 ref20">11,20</xref>
        ]. After a careful examination of the literature, we developed
a new approach that aims to generate a sequence of sounds from a preliminary
video. The nal goal of the proposed system is to obtain a musical result that
will describe the image provided by a user. To achieve this, we propose to divide
the video into frames. Afterwards, the visual and auditory characteristics of
these frames are extracted. Finally, a pattern must be established from this
information by applying some data mining techniques.
      </p>
      <p>
        The animated lm Fantasia [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ] produced by Walt Disney has been chosen
for the case study. This lm is made up of a concatenation of eight pieces of
classical music, described with the illustrations of professional animators. Due
to the deep analysis made by expert people, Fantasia will be used in this work
in order to create a model that relates some characteristics of the images to
sound. This pattern will then be applied to a new image in order to compose
music. A deep analysis of 483 movie frames was carried out in order to extract a
set of image descriptors; it will be used to perform the translation in the reverse
direction. Furthermore, a musical analysis has been performed in order to extract
information about the relationship between each of the frames studied and its
sound. Then, the selected classi ers were applied to the data set, and due to
the quality of the results Random Forest (RF) [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ] was chosen for this proposal.
Thus, a model which represented the relationship between the characteristics of
the image and the sound was obtained. This model will then be applied in the
last stage of this work, where a new video is divided into frames, and each one
of them is translated into a sound. The concatenation of all the resulting sounds
will compose the nal melody.
      </p>
      <p>Section 2 reviews some image-to-sound approaches found in the literature.
Section 3 outlines the work ow of the system and Section 4 analyzes di erent
techniques used for image processing and descriptive music composition that are
used in this work. In Section 5, all the details of this case study are explained:
rstly, a deep analysis of the animated lm Fantasia is made and the obtained
data are presented. Then, classi cation techniques are applied for the purpose
of de ning the relationship between the characteristics of image and sound.
Different classi cation techniques are applied in this work, and their results are
discussed in Section 6. This section also presents the obtained musical results,
where the previously created model is used to compose a sequence of sounds
that describes new images provided by users. Finally, the last section presents
the conclusions drawn upon the completion of this work. Moreover, future lines
of research are discussed.
2</p>
    </sec>
    <sec id="sec-2">
      <title>Image to music conversion</title>
      <p>The interest to fuse visual and musical art is not new. Some researchers have tried
to nd a psychological link between images and sounds through synesthesia or
cognitive audition. Drawing on these phenomena, di erent approaches have been
developed, some related to the synesthesia, the spectograms or the descriptive
music.</p>
      <p>
        Synesthesia has been widely used for this purpose over the course of
history [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ]. It is a neurological phenomenon that occurs when one sense is being
stimulated and an automatic experience arises in another one. This perceptual
process has led to many creations that relate di erent elds of art such as
poetry, dance, painting and music [
        <xref ref-type="bibr" rid="ref15 ref23">15, 23</xref>
        ]. One of the most recurrent associations
in the synesthesia phenomenon is the one which unites color and sound.
Computationally, there are many works that propose an automatic translation; in [
        <xref ref-type="bibr" rid="ref22">22</xref>
        ] a
visual color notation for music is proposed and the Monalisa App [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ] is a
standalone software that converts images into sounds and vice versa with the aid of
an intermediate step in which the information is treated as binary numbers.
      </p>
      <p>
        Newton established a model called Musical Color Wheel [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ] where each one of
the 7 colors of the prism was related to one of the 7 musical notes. However, this
theory is weakened by the fact that the chromatic scale has 12 notes. Lagresille
used this model as a starting point and proposed a new broaded translation
system by establishing a relationship between the properties of color and those
of sound [
        <xref ref-type="bibr" rid="ref25">25</xref>
        ]. Speci cally, he proposed that the saturation of color and the
volume of sound were directly proportional, as well as the brightness of color
and the sharpness of sound. His theory also involves obtaining notes on the
basis of synesthesia. He divides the chromatic circle into 12 sectors and assigns
a musical note to each one [
        <xref ref-type="bibr" rid="ref20">20</xref>
        ].
      </p>
      <p>However, the use of synesthesia as a mean of conversion between images and
music leads to several problems. On the one hand, this phenomenon entails a
high level of subjectivity due to the huge di erences in the way that people
perceive information. On the other hand, only 1% of population experiences
synesthesia, which makes it di cult to establish a proven relationship between
color and sound.</p>
      <p>
        In order to address the subjectivity problems, there is a visualization
technique that derives from the signal processing theory: spectrograms, which are
visual representations of the spectrum of some signal frequencies such as sound [
        <xref ref-type="bibr" rid="ref19">19</xref>
        ].
It is widely used in music classi cation problems because of the large amount
of information that it provides. In [
        <xref ref-type="bibr" rid="ref26">26</xref>
        ] a musical onset detection problem is
solved by applying a Convolutional Neural Network (CNN) to a set of
spectrograms. [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ] proposes to extract information from the spectrogram to increase the
robustness of convolutional neural networks and to decrease noisy and degraded
channel conditions. Nevertheless, although this technique makes use of visual
representations and sound, it would be very di cult to use it as a tool for the
translation of images to music.
      </p>
      <p>
        Another approach that involves the concept of descriptive music, [
        <xref ref-type="bibr" rid="ref28">28</xref>
        ], which
could be considered as a music genre that tries to create certain images, scenes
or moods in the listener's mind. Throughout history, many relevant composers
have contributed with their creations to this genre [
        <xref ref-type="bibr" rid="ref28">28</xref>
        ]. A classic example could
be Vivaldi's The Four Seasons, where each one of the four violin concerts is
related to one season (spring, summer, autumn and winter) and it attempts to
musically describe some typical details and climatic conditions such as
thunders, the song of the birds, the rain and the wind. Another example could be
Saint-Saens'The Carnival of the Animals, where each movement characterizes
an animal (the swan, the elephant, some tortoises, etc.). A set of this kind of
musical compositions has been compiled in the lm Fantasia, produced by Walt
Disney [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]. Some professional animators have been chosen for the creative task;
they created a model that related colors to musical emotions and analyzed how
to design the characters and their movements on the basis of musical details [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ].
Figure 1 shows a frame of the video fragment related to the musical piece \The
Sorcerer's Apprentice" where Mickey Mouse learns to do magic. The creation of
the model, as well as the process of composing music will be explained in more
detail in Section 5.
The nal goal of this work is the automatic composition of descriptive music.
Figure 2 shows a visual scheme of this creative process in which the work ow is
divided into two stages. On the one hand, in the training stage, the preliminary
video is divided into a set of frames. A set of image descriptors is extracted
from each one of them, as well as the main note that sounds at the moment
when the frame is being played. Afterwards, a classi er is applied to the data
in order to extract a model that relates both, image descriptors and sounds. On
the other hand, the test stage where the musical composition is created takes
place. A new video is chosen and divided into a set of frames. All the considered
image descriptors are extracted again, and the model is applied to each one of
them in order to obtain a sound. The concatenation of all the sounds makes up
a sequence of notes which describes the characteristics of the frames.
The characteristics that de ne each frame have been divided into two groups:
shape features and color features. To obtain the rst ones, Scale-Invariant
Feature Transform method has been used together with Bag of Visual Words in
order to achieve a dimensional reduction of the SIFT descriptor vector. These
techniques are described in sections 4.1 and 4.2 respectively. For the second
group of characteristics, color histogram has been used. It is described in section
4.3.
Scale-Invariant Feature Transform (SIFT) is a method for extracting distinctive
invariant features from images that can be used to perform reliable matching
between di erent views of an object or scene [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ]. SIFT descriptors are invariant
to translations, rotations and scaling transformations and partially perspective
transformations and illumination variations.
      </p>
      <p>
        SIFT descriptors comprise a method for detecting points of interest from a
grey-level image at which statistics of local gradient directions of image
intensities were accumulated to give a summarizing description of the local image
structures in a local neighborhood around each interest point. The steps for
obtaining these image descriptors are as follows [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ]:
1. Scale-space extrema detection: First step consist in searching over all
scales and image locations. In this phase, keypoints are detected using a
cascade ltering method in order to identify candidate locations that are
invariant to scale and orientation.
2. Keypoint localization: The next step is to perform a detailed t to the
nearby data for location, scale, and ratio of principal curvatures. This
information allows to reject points that have low contrast or are poorly localized
along an edge.
3. Orientation assignment: In this stage one or more orientations are
assigned to each keypoint location based on local image gradient directions.
By assigning a consistent orientation to each keypoint based on local image
properties. The keypoint descriptor can be represented relative to a
consistent orientation, achieving invariance to image rotation.
4. Keypoint descriptor: The last step is to compute a descriptor for the
local image region that is highly distinctive yet is as invariant as possible to
remaining variations. The local image gradients are measured at the selected
scale in the region around each keypoint. These are transformed into a
representation that allows to be invariant to signi cant levels of local shape
distortion and change in illumination.
4.2
      </p>
      <sec id="sec-2-1">
        <title>Bag of Visual Words</title>
        <p>
          As seen in the previous section, an image can be de ned from a series of
descriptors that contain important information about the features of the image
(like SIFT descriptors). These keypoints can be grouped into clusters, so that
each cluster is considered a visual word. Bag of Visual Words (BoVW) is a
technique by which images are represented from visual words that symbolize the
clusters where keypoints are grouped together, by obtaining a vector containing
the (weighted) count of each visual word in that image [
          <xref ref-type="bibr" rid="ref35">35</xref>
          ]. This characteristic
vector is used in the classi cation task. BoVW is a representation of images
analogous to the Bag of Words (BoW) representation of text documents.
4.3
        </p>
      </sec>
      <sec id="sec-2-2">
        <title>Color Histogram</title>
        <p>
          A color histogram H(M ) is a vector (hl; h2; : : : ; hn), where each element hJ
represents the number of pixels falling in bin j in image M [
          <xref ref-type="bibr" rid="ref17">17</xref>
          ]. The color space
chosen for this work has been RGB, one of the most used in color histograms.
Each of the channels has been divided into 256 bins, which correspond to the
color intensity in the range [0 255]. For each of the channels a histogram has
been obtained, i. e. R, G and B. In this way, a representation of the color of the
image is obtained as three vectors, one for each channel of the RGB space.
5
        </p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>Case study</title>
      <p>We will detail our proposal by making use of the Fantasia lm as a preliminary
case study. As we explained before, the methodology follows a two stage ow. In
Section 5.1 the task of extracting data from the initial video and its processing
are explained. Section 5.2 describes the process in which the model that relates
color, shape and the disposition of elements with sound is obtained.
5.1</p>
      <sec id="sec-3-1">
        <title>Data description</title>
        <p>
          When a composer creates music he needs a source of inspiration. In
Computational Creativity, the generation of creative content is carried out by making
machines to imitate the way in which people create art [
          <xref ref-type="bibr" rid="ref31">31</xref>
          ]. In this work we
design a system that composes descriptive music on the basis of the visual
characteristics of a preliminary video. The rst step was to analyze the fragment of
"The Nutcracker Suite" played in the Fantasia lm. The aim of this proposal is
to carry out a reverse process to how the Fantasia lm was created: we create
a model that relates sounds to some features of the images such as color, shape
and disposition of the elements.
        </p>
        <p>
          First, an analysis of the video frames is carried out. The only seconds of
the lm that are considered are those that present changes in color, shape or
disposition of elements. The lm is shot at 24 frames per second and for our study
the rst one out of every 8 is chosen. Thus, three frames are considered per each
selected second. This results in a collection of 483 images for analysis. From each
one of them we obtain 1168 image descriptors which gather up their main visual
features. In addition, an auditory analysis is performed by a musician to extract
the most important note from all those that sound simultaneously at the moment
the frame is being played in the video. Data relative to image descriptors and
the most important sound of each frame can be found at [
          <xref ref-type="bibr" rid="ref18">18</xref>
          ].
        </p>
        <p>We establish a relationship between image and sound by extracting patterns
from their basic features. The image descriptors that are considered in this work
have been divided into two groups: the rst one describes the characteristics
of the shape and disposition of the elements in the image, and the second one
collects information related to color.</p>
        <p>The rst group of data is comprised of the rst 400 values and they
correspond to the frequency vector of the visual words obtained from the application
of BoVW to the SIFT descriptors. 400 visual words have been chosen to form the
visual vocabulary. The second group of data is comprised of the last 798 values
and they correspond to the color histogram for RGB color space, 256 values for
each channel. Figure 3 shows the SIFT descriptors and color histogram of frame
1442:11. Each SIFT descriptor represents a border or shape of the elements that
make up the image, such as the head and arms of the owers. The color
histogram shows that the three RGB channels have a high number of pixels in the
50 to 75 bins, which corresponds to the colors of the owers where the three R,
G and B channels are present.</p>
        <p>
          Finally, the class of the data set is the most important note that is heard in
each frame. Sometimes, this sound will be the fundamental pitch of the chord
that is being reproduced, but in other cases it will be the note that stands out
from the sounds cloud. The 12 notes of the chromatic scale are considered for
this task. The format used for its speci cation is MIDI [
          <xref ref-type="bibr" rid="ref8">8</xref>
          ], restricting the record
to the octave that corresponds to the central C of the piano to simplify the task.
Table 1 shows each one of the 12 notes and its corresponding MIDI encoding.
        </p>
        <p>As a last step, after obtaining the visual and sound information of each
one of the selected frames and collecting it in the data set, a brief analysis is
performed. The number of attributes, which are all descriptors of the image, are
a total of 1168 for each instance. Additionally, in Table 2 the number of frames
that is classi ed for each label is shown. This information is also presented as the
percentage of frames per label out of the total number of the cases considered.</p>
        <p>From the values shown in the Table 2, two conclusions can be drawn. On
the one hand, there is no data pertaining to class \G#". This means that, in
the video fragment that has been studied, there is no frame in which the most
important note is \G#". Thus, no appraisals can be made for this note in future
cases. On the other hand, the number of frames related to each class or note is
not the same. Some classes like \D" and \A" have more occurrences than others
like \C" and \G#". In data mining this problem is known as a class imbalance,
and in this case it is due to the analysis of a single musical piece with a de ned
tonality that determines the important notes over the others.
5.2</p>
      </sec>
      <sec id="sec-3-2">
        <title>Model induction</title>
        <p>
          Supervised Machine Learning [
          <xref ref-type="bibr" rid="ref13">13</xref>
          ] infers a mathematical function that relates a
set of attributes. This technique builds a model that classi es instances based
on a set of previously labeled data. In this work, classi cation is used to extract
a pattern that relates the image descriptors to the main sound in a video frame.
        </p>
        <p>After a theoretical analysis of the classi cation task, in search of those
techniques that would provide the best results, three algorithms were selected: Naive
Bayes (NB), Support Vector Machine (SVM) and Random Forest (RF).</p>
        <p>
          NB classi er is a probabilistic classi er based on the Bayes theorem and the
hypothesis of independence between predictive variables. It assumes that the
predictive attributes are conditionally independent given the class, and it posits
that no hidden or latent attributes in uence the prediction process [
          <xref ref-type="bibr" rid="ref12">12</xref>
          ].
        </p>
        <p>
          SVM classi er is a supervised technique that nds a linear separating
hyperplane with the maximal margin in a higher dimensional space [
          <xref ref-type="bibr" rid="ref10">10</xref>
          ]. In this work,
the Minimum Sequential Optimization (SMO) training algorithm has been used
[
          <xref ref-type="bibr" rid="ref27">27</xref>
          ]. This algorithm divides the global problem into a series of problems that are
as small as possible and which are solved independently.
        </p>
        <p>
          RF classi er is an ensemble algorithm which consists of tree predictors, such
that each tree depends on the values of a random vector sampled independently
and with the same distribution for all trees in the forest [
          <xref ref-type="bibr" rid="ref2">2</xref>
          ].
        </p>
        <p>In order to determine the nal choice, three classi ers were applied to the
data set and the obtained results were discussed as detailed in Section 6. In
conclusion, RF outperformed the other classi ers and for this reason it was used
to create the model.
6</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>Discussion</title>
      <p>In this section two types of results will be discussed. On the one hand,
Section 6.1 presents a comparison of the performance of the algorithms applied in
the classi cation task. On the other hand, the descriptive quality of the music
obtained is analyzed in Section 6.2.
6.1</p>
      <sec id="sec-4-1">
        <title>Performance of the classi ers</title>
        <p>
          The performance indicators of the three classi ers (NB, SVM and RF) applied in
the model induction process are shown in Table 3. Each row shows all the
information on the performance of a classi er and the columns correspond to di erent
quality measures [
          <xref ref-type="bibr" rid="ref9">9</xref>
          ]. Precision and recall are measures of exactness in the
classi cation task: they look at how many of the returned instances are correct and
how many positives the model returned respectively. F-score combines precision
and recall information and determine a weighted single value and Kappa
measures the agreement of the evaluations on the same samples. Root-Mean-Square
Error (RMSE) is a metric that gives information about the concentration of
the data around its best t. Finally, Receiver Operating Characteristic (ROC)
represents the exchanges between true positives and false positives.
NB
SVM
RF
        </p>
        <p>
          As shown in Table 3, the accuracy of the classi ers varies greatly. Based on
the metrics of precision, recall and F-score, NB is the classi er with the least
exactness. Conversely, RF has a good performance. The Kappa metric shows
that NB (0:3489) and SVM (0:7627) classify a ected by a random agreement
factor unlike RF (0:807). The RF obtains the best value for the RMSE metric
(0:1855); it means that the misclassi ed sounds are close to the good one. The
optima value for ROC is 1; thus, RF almost reaches the maximum value for this
metric too, outperforming the other classi ers again. Since it performed better
than the other classi ers, RF has been chosen for this work.
To evaluate the descriptive quality of the musical composition process, a new
video is selected. The video fragment, that corresponds to the musical piece
known as \The Firebird" of Igor Stravinski in the lm Fantasia 2000 [
          <xref ref-type="bibr" rid="ref5">5</xref>
          ], is
chosen for this task due to its di erent color ranges and the movement of its
characters. Three frames per second were extracted and feature extraction was
applied. This data can be found at [
          <xref ref-type="bibr" rid="ref18">18</xref>
          ]. Finally, the model obtained in the
classi cation task was applied to its image descriptors. As a result, a sequence
of notes was created. To do a better analysis, the sound of the preliminary video
was removed and the sequence of sounds composed by the system was played
instead. The sequence of sounds can also be found at [
          <xref ref-type="bibr" rid="ref18">18</xref>
          ].
        </p>
        <p>As in the previous case, the musical composition was not inspired by the video
but by the images that compose it. For this purpose, the video was divided into
a set of frames; speci cally, in this case three frames per second were extracted.
The next step is feature extraction for each one of the frames. In the same way
as in the creation of the data set, the information about the shape, disposition
and color was obtained.</p>
        <p>In a nal step, the model extracted by the classi er is applied to the new
data. As a result, each one of the frames is translated into a sound that describes
it according to the previously de ned pattern. The concatenation of the obtained
sounds gives rise to a sequence of notes. The duration of each of these sounds
is 0:33 seconds. When two or more consecutive frames are translated into the
same note, the musical result is a single sound with a duration equal to the sum
of those of the isolated sounds. Thus, the musical composition is endowed with
rhythm. Moreover, when there is no visual change in a set of consecutive frames,
the sound that describes them is a continuous note.</p>
        <p>From the conducted evaluation we can identify two aspects. On the one hand,
the sequence of sounds is substantially di erent depending on the classi er that
is used in the model induction process. On the other hand, in all cases, the
relationship between the sound and the visual characteristics of the frames can be
spotted easily. When two or more frames are visually similar, their corresponding
sounds are the same; however, when two frames have some di erences in color,
shape or disposition the sound obtained is also di erent.
7</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>Conclusions and future work</title>
      <p>This proposal successfully builds a system for automatic musical composition
that describes a preliminary video. First, a deep analysis of the lm Fantasia was
conducted and the relationship between the color, the shape and the disposition
of the elements of an image and its sound was established. Afterwards, a new
video is divided into a set of frames, and the previously extracted model is
applied to their image descriptors for the creation of a sequence of sounds.</p>
      <p>
        One of the steps of the data set creation process is the labeling task: each
frame is related to an isolated sound, which is the most important of all those
that are sounding at the moment when it is being visually played. Despite the
existence of some automated techniques for obtaining the main pitch of a sounds
cloud or a chord which are based on the Fast Fourier Transform [
        <xref ref-type="bibr" rid="ref21 ref30">21, 30</xref>
        ], in our
data set the label is analytically obtained by a music expert. This makes the
labeling a slower, more expensive and more complex process. However, according
to the nal results it is a well-invested e ort.
      </p>
      <p>After the creation of the data set that is used to build the model, a brief study
on the distribution of information is conducted. Due to the analysis of a single
musical piece there is a class with no representation in the data set and a class
imbalance problem: in music, each composition has a note which is considered
the most important, and there is a set of sounds that are more related to it
than the others. However, the accuracy of 83% proves that the classi er behaves
appropriately despite the di erence of examples for each class considered. It
means that there is a clear relationship between the image descriptors and the
sound in this case study. With regard to the classi ers, RF is the one that has
the best result, outperforming NB and SVM. Additionally, despite their similar
success rates, the three models produce substantially di erent musical results.</p>
      <p>
        Furthermore, the fragment of \The Firebird" from the lm Fantasia 2000 [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ]
was used in the test stage. It belongs to Animation genre as well as the lm
Fantasia. For this reason, there is a considerable similarity between the frames
of both videos, what entails that the descriptors extracted from the images are
also analogous in both cases. Thus, due to the application of the model to a data
set which is similar to the training one, the results are quite good.
      </p>
      <p>In this work, the movement of objects and characters is not speci cally
analyzed. However, a video could be considered as a concatenation of several frames
or images. Thus, an isolated frame was translated into an isolated sound, and as
the movement is a continuous change of position, it was implicitly studied in the
progression of several consecutive frames. Additionally, the result is not a melody
with a great musical quality based on some harmony rules, but a sequence of
sounds that describes the concatenation of frames.</p>
      <p>In this proposal, only the video section corresponding to "The Nutcracker
Suite" is analyzed. An extension of the data set could be performed by studying
some more fragments of the lm. This deeper analysis could lead to the solution
of the unbalanced data problem; the tonality is di erent for each musical piece
that sound in the lm, and consequently the frequency of the notes also varies.
Moreover, the data set would consist of several pieces of di erent styles, giving
rise to a more robust model constructed with the data of all classes.</p>
    </sec>
    <sec id="sec-6">
      <title>Acknowledgments</title>
      <p>This work was supported by the Spanish Ministry, Ministerio de Econom a y
Competitividad and FEDER funds. Project. SURF: Intelligent System for
integrated and sustainable management of urban eets TIN2015-65515-C4-3-R.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Babtrakinova</surname>
            ,
            <given-names>O.I.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Voloshko</surname>
            ,
            <given-names>A.A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Snistaryova</surname>
            ,
            <given-names>P.A.</given-names>
          </string-name>
          :
          <article-title>In uence of modern music on young generation</article-title>
          .
          <source>Human and society (2)</source>
          , 4{
          <issue>6</issue>
          (
          <year>2017</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Breiman</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          :
          <article-title>Random forests</article-title>
          .
          <source>Machine learning 45(1)</source>
          ,
          <volume>5</volume>
          {
          <fpage>32</fpage>
          (
          <year>2001</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Clague</surname>
            ,
            <given-names>M.:</given-names>
          </string-name>
          <article-title>Playing in'toon: Walt disney's" fantasia"(1940) and the imagineering of classical music</article-title>
          .
          <source>American Music</source>
          <volume>22</volume>
          (
          <issue>1</issue>
          ),
          <volume>91</volume>
          {
          <fpage>109</fpage>
          (
          <year>2004</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Collopy</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          :
          <article-title>Color, form, and motion: Dimensions of a musical art of light</article-title>
          .
          <source>Leonardo</source>
          <volume>33</volume>
          (
          <issue>5</issue>
          ),
          <volume>355</volume>
          {
          <fpage>360</fpage>
          (
          <year>2000</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Culhane</surname>
          </string-name>
          , J.:
          <source>Fantasia</source>
          <year>2000</year>
          :
          <article-title>Visions of Hope</article-title>
          . Disney
          <string-name>
            <surname>Editions</surname>
          </string-name>
          (
          <year>1999</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Cytowic</surname>
          </string-name>
          , R.E.:
          <article-title>Synesthesia: A union of the senses</article-title>
          . MIT press (
          <year>2002</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>Duchowski</surname>
          </string-name>
          , A.T.:
          <article-title>Eye tracking methodology</article-title>
          .
          <source>Theory and practice 328</source>
          (
          <year>2007</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>England</surname>
          </string-name>
          , R.:
          <article-title>Standard midi le production as the focus of a broad computer science course</article-title>
          .
          <source>Journal of Computing Sciences in Colleges</source>
          <volume>32</volume>
          (
          <issue>5</issue>
          ),
          <volume>4</volume>
          {
          <fpage>10</fpage>
          (
          <year>2017</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <surname>Guillet</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hamilton</surname>
            ,
            <given-names>H.J.:</given-names>
          </string-name>
          <article-title>Quality measures in data mining</article-title>
          , vol.
          <volume>43</volume>
          . Springer (
          <year>2007</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <surname>Hsu</surname>
            ,
            <given-names>C.W.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chang</surname>
            ,
            <given-names>C.C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lin</surname>
            ,
            <given-names>C.J.</given-names>
          </string-name>
          , et al.:
          <article-title>A practical guide to support vector classi cation (</article-title>
          <year>2003</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <string-name>
            <surname>Jo</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Nagano</surname>
          </string-name>
          , N.:
          <article-title>Monalisa:" see the sound, hear the image"</article-title>
          .
          <source>In: NIME</source>
          . pp.
          <volume>315</volume>
          {
          <issue>318</issue>
          (
          <year>2008</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          12. John,
          <string-name>
            <given-names>G.H.</given-names>
            ,
            <surname>Langley</surname>
          </string-name>
          ,
          <string-name>
            <surname>P.</surname>
          </string-name>
          :
          <article-title>Estimating continuous distributions in bayesian classi ers</article-title>
          .
          <source>In: Proceedings of the Eleventh conference on Uncertainty in arti cial intelligence</source>
          . pp.
          <volume>338</volume>
          {
          <fpage>345</fpage>
          . Morgan Kaufmann Publishers Inc. (
          <year>1995</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          13.
          <string-name>
            <surname>Kotsiantis</surname>
            ,
            <given-names>S.B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zaharakis</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Pintelas</surname>
            ,
            <given-names>P.:</given-names>
          </string-name>
          <article-title>Supervised machine learning: A review of classi cation techniques (</article-title>
          <year>2007</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          14.
          <string-name>
            <surname>Kovacs</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Toth</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Van Compernolle</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ganapathy</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          :
          <article-title>Increasing the robustness of cnn acoustic models using autoregressive moving average spectrogram features and channel dropout</article-title>
          .
          <source>Pattern Recognition Letters</source>
          (
          <year>2017</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          15.
          <string-name>
            <surname>Lerdahl</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          :
          <article-title>The sounds of poetry viewed as music</article-title>
          .
          <source>Annals of the New York Academy of Sciences</source>
          <volume>930</volume>
          (
          <issue>1</issue>
          ),
          <volume>337</volume>
          {
          <fpage>354</fpage>
          (
          <year>2001</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          16.
          <string-name>
            <surname>Lowe</surname>
            ,
            <given-names>D.G.</given-names>
          </string-name>
          :
          <article-title>Distinctive image features from scale-invariant keypoints</article-title>
          .
          <source>International journal of computer vision 60(2)</source>
          ,
          <volume>91</volume>
          {
          <fpage>110</fpage>
          (
          <year>2004</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          17.
          <string-name>
            <surname>Lu</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Phillips</surname>
            ,
            <given-names>J.:</given-names>
          </string-name>
          <article-title>Using perceptually weighted histograms for colour-based image retrieval</article-title>
          .
          <source>In: Signal Processing Proceedings</source>
          ,
          <year>1998</year>
          . ICSP'
          <volume>98</volume>
          . 1998 Fourth International Conference on. vol.
          <volume>2</volume>
          , pp.
          <volume>1150</volume>
          {
          <fpage>1153</fpage>
          .
          <string-name>
            <surname>IEEE</surname>
          </string-name>
          (
          <year>1998</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          18.
          <string-name>
            <surname>Mart</surname>
            n-Gomez,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Perez-Marcos</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          :
          <source>Data repository of fantasia case study (Nov</source>
          <year>2017</year>
          ), https://github.com/lumg/FantasiaDisney_data
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          19. Muller, M.:
          <article-title>Book: Fundamentals of music processing</article-title>
          .
          <source>Signal 1</source>
          ,
          <issue>0</issue>
          {
          <fpage>5</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          20.
          <string-name>
            <surname>Navarro-Caceres</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bajo</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Corchado</surname>
            ,
            <given-names>J.M.:</given-names>
          </string-name>
          <article-title>Applying social computing to generate sound clouds</article-title>
          .
          <source>Engineering Applications of Arti cial Intelligence</source>
          <volume>57</volume>
          ,
          <fpage>171</fpage>
          {
          <fpage>183</fpage>
          (
          <year>2017</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          21.
          <string-name>
            <surname>Peeters</surname>
          </string-name>
          , G.:
          <article-title>Music pitch representation by periodicity measures based on combined temporal and spectral representations</article-title>
          .
          <source>In: Acoustics, Speech and Signal Processing</source>
          ,
          <year>2006</year>
          .
          <source>ICASSP 2006 Proceedings. 2006 IEEE International Conference on. vol. 5</source>
          ,
          <string-name>
            <given-names>pp. V{V.</given-names>
            <surname>IEEE</surname>
          </string-name>
          (
          <year>2006</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          22.
          <string-name>
            <surname>Poast</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          :
          <article-title>Color music: Visual color notation for musical expression</article-title>
          .
          <source>Leonardo</source>
          <volume>33</volume>
          (
          <issue>3</issue>
          ),
          <volume>215</volume>
          {
          <fpage>221</fpage>
          (
          <year>2000</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          23.
          <string-name>
            <surname>Ranjan</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gabora</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>OConnor</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          :
          <article-title>The cross-domain re-interpretation of artistic ideas</article-title>
          .
          <source>arXiv preprint arXiv:1308.4706</source>
          (
          <year>2013</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          24.
          <string-name>
            <surname>Sakka</surname>
            ,
            <given-names>L.S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Juslin</surname>
            ,
            <given-names>P.N.</given-names>
          </string-name>
          :
          <article-title>Emotional reactions to music in depressed individuals</article-title>
          . Psychology of Music p.
          <volume>0305735617730425</volume>
          (
          <year>2017</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref25">
        <mixed-citation>
          25.
          <string-name>
            <surname>Sanz</surname>
            ,
            <given-names>J.C.</given-names>
          </string-name>
          :
          <article-title>Lenguaje del color:(sinestesia cromatica en poes a y arte visual)</article-title>
          .
          <source>El autor</source>
          (
          <year>1981</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref26">
        <mixed-citation>
          26.
          <string-name>
            <surname>Schluter</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bock</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          :
          <article-title>Improved musical onset detection with convolutional neural networks</article-title>
          .
          <source>In: Acoustics, speech and signal processing (icassp)</source>
          ,
          <year>2014</year>
          ieee international conference on. pp.
          <volume>6979</volume>
          {
          <fpage>6983</fpage>
          .
          <string-name>
            <surname>IEEE</surname>
          </string-name>
          (
          <year>2014</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref27">
        <mixed-citation>
          27. Scholkopf,
          <string-name>
            <given-names>B.</given-names>
            ,
            <surname>Burges</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.J.</given-names>
            ,
            <surname>Smola</surname>
          </string-name>
          ,
          <string-name>
            <surname>A.J.:</surname>
          </string-name>
          <article-title>Advances in kernel methods: support vector learning</article-title>
          . MIT press (
          <year>1999</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref28">
        <mixed-citation>
          28.
          <string-name>
            <surname>Seeger</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          :
          <article-title>Prescriptive and descriptive music-writing</article-title>
          .
          <source>The Musical Quarterly</source>
          <volume>44</volume>
          (
          <issue>2</issue>
          ),
          <volume>184</volume>
          {
          <fpage>195</fpage>
          (
          <year>1958</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref29">
        <mixed-citation>
          29.
          <string-name>
            <surname>Tsoumakas</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Katakis</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Vlahavas</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          :
          <article-title>Mining multi-label data</article-title>
          .
          <source>In: Data mining and knowledge discovery handbook</source>
          , pp.
          <volume>667</volume>
          {
          <fpage>685</fpage>
          . Springer (
          <year>2009</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref30">
        <mixed-citation>
          30.
          <string-name>
            <surname>Tzanetakis</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ermolinskyi</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Cook</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          :
          <article-title>Pitch histograms in audio and symbolic music information retrieval</article-title>
          .
          <source>Journal of New Music Research</source>
          <volume>32</volume>
          (
          <issue>2</issue>
          ),
          <volume>143</volume>
          {
          <fpage>152</fpage>
          (
          <year>2003</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref31">
        <mixed-citation>
          31.
          <string-name>
            <surname>Varshney</surname>
            ,
            <given-names>L.R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Pinel</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Varshney</surname>
            ,
            <given-names>K.R.</given-names>
          </string-name>
          , Schorgendorfer,
          <string-name>
            <given-names>A.</given-names>
            ,
            <surname>Chee</surname>
          </string-name>
          ,
          <string-name>
            <surname>Y.M.</surname>
          </string-name>
          :
          <article-title>Cognition as a part of computational creativity</article-title>
          .
          <source>In: Cognitive Informatics &amp; Cognitive Computing (ICCI* CC)</source>
          ,
          <year>2013</year>
          12th IEEE International Conference on. pp.
          <volume>36</volume>
          {
          <fpage>43</fpage>
          .
          <string-name>
            <surname>IEEE</surname>
          </string-name>
          (
          <year>2013</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref32">
        <mixed-citation>
          32.
          <string-name>
            <surname>Vishton</surname>
            ,
            <given-names>P.M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Vishton</surname>
            ,
            <given-names>P.M.</given-names>
          </string-name>
          :
          <article-title>Understanding the secrets of human perception</article-title>
          .
          <source>Teaching Company</source>
          (
          <year>2011</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref33">
        <mixed-citation>
          33.
          <string-name>
            <surname>Wang</surname>
            ,
            <given-names>Z.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chen</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Li</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Xu</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lin</surname>
            ,
            <given-names>L.:</given-names>
          </string-name>
          <article-title>Multi-label image recognition by recurrently discovering attentional regions</article-title>
          .
          <source>In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition</source>
          . pp.
          <volume>464</volume>
          {
          <issue>472</issue>
          (
          <year>2017</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref34">
        <mixed-citation>
          34.
          <string-name>
            <surname>Yanagimoto</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sugimoto</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          :
          <article-title>Convolutional neural networks using supervised pre-training for eeg-based emotion recognition</article-title>
          .
          <source>In: 8th International Workshop on Biosignal Interpretation (BSI)</source>
          (
          <year>2016</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref35">
        <mixed-citation>
          35.
          <string-name>
            <surname>Yang</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Jiang</surname>
            ,
            <given-names>Y.G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hauptmann</surname>
            ,
            <given-names>A.G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ngo</surname>
            ,
            <given-names>C.W.</given-names>
          </string-name>
          :
          <article-title>Evaluating bag-of-visualwords representations in scene classi cation</article-title>
          .
          <source>In: Proceedings of the international workshop on Workshop on multimedia information retrieval</source>
          . pp.
          <volume>197</volume>
          {
          <fpage>206</fpage>
          .
          <string-name>
            <surname>ACM</surname>
          </string-name>
          (
          <year>2007</year>
          )
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>