<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Towards a Cognitive Architecture for Music Perception</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Antonio Chella</string-name>
          <email>antonio.chella@unipa.it</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Department of Chemical</institution>
          ,
          <addr-line>Management, Computer</addr-line>
          ,
          <institution>Mechanical Engineering University of Palermo</institution>
          ,
          <addr-line>Viale delle Scienze, building 6 90128 Palermo</addr-line>
          ,
          <country country="IT">Italy</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>The framework of a cognitive architecture for music perception is presented. The architecture extends and completes a similar architecture for computer vision developed during the years. The extended architecture takes into account many relationships between vision and music perception. The focus of the architecture resides in the intermediate area between the subsymbolic and the linguistic areas, based on conceptual spaces. A conceptual space for the perception of notes and chords is discussed along with its generalization for the perception of music phrases. A focus of attention mechanism scanning the conceptual space is also outlined. The focus of attention is driven by suitable linguistic and associative expectations on notes, chords and music phrases. Some problems and future works of the proposed approach are also outlined.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Introduction</title>
      <p>
        Garderfors [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ], in his paper on \Semantics, Conceptual Spaces and Music"
discusses a program for musical spaces analysis directly inspired to the framework
of vision proposed by Marr [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]. More in details, the rst level that feeds input
to all the subsequent levels is related with pitch identi cation. The second level
is related with the identi cation of musical intervals ; this level takes also into
account the cultural background of the listener. The third level is related with
tonality, where scales are identi ed and the concepts of chromaticity and
modulation arise. The fourth level of analysis is related with the interplay of pitch
and time. According to Gardenfors, time is concurrently processed by means of
di erent levels related with temporal intervals, beats, rhythmic patterns, and at
this level the analysis of pitch and the analysis of time merge together.
      </p>
      <p>
        The correspondences between vision and music perception have been
discussed in details by Tanguiane [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]. He considers three di erent levels of analysis
distinguishing between statics and dynamics perception in vision and music. The
rst visual level in statics perception is the level of pixels, in analogy of the
image level of Marr, that corresponds to the perception of partials in music. At
the second level, the perception of simple patterns in vision corresponds to the
perception of single notes. Finally at the third level, the perception of structured
patterns (as patterns of patterns), corresponds to the perception of chords.
Concerning dynamic perception, the rst level is the same as in the case of static
perception, i.e., pixels vs. partials, while at the second level the perception of
visual objects corresponds to the perception of musical notes, and at the third
nal level the perception of visual trajectories corresponds to the perception of
music melodies.
      </p>
      <p>
        Several cognitive models of music cognition have been proposed in the
literature based on di erent symbolic or subsymbolic approaches, see Pearce and
Wiggins [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ] and Temperley [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ] for recent reviews. Interesting systems, representative
of these approaches are: MUSACT [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ][
        <xref ref-type="bibr" rid="ref7">7</xref>
        ] based on various kinds of neural
networks; the IDyOM project based on probabilistic models of perception [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ][
        <xref ref-type="bibr" rid="ref8">8</xref>
        ][
        <xref ref-type="bibr" rid="ref9">9</xref>
        ];
the Melisma system [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ] based on preference rules of symbolic nature; the HARP
system, aimed at integrating symbolic and subsymbolic levels [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ][
        <xref ref-type="bibr" rid="ref12">12</xref>
        ].
      </p>
      <p>
        Here, we sketch a cognitive architecture for music perception that extends
and completes an architecture for computer vision developed during the years.
The proposed cognitive architecture integrates the symbolic and the sub
symbolic approaches and it has been employed for static scenes analysis [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ][
        <xref ref-type="bibr" rid="ref14">14</xref>
        ],
dynamic scenes analysis [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ], reasoning about robot actions [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ], robot
recognition of self [
        <xref ref-type="bibr" rid="ref17">17</xref>
        ] and robot self-consciousness [
        <xref ref-type="bibr" rid="ref18">18</xref>
        ]. The extended architecture
takes into account many of the above outlined relationships between vision and
music perception.
      </p>
      <p>In analogy with Tanguiane, we distinguish between \static" perception
related with the perception of chords in analogy with perception of static scenes,
and \dynamic" perception related with the perception of musical phrases, in
analogy with perception of dynamic scenes.</p>
      <p>
        The considered cognitive architecture for music perception is organized in
three computational areas - a term which is reminiscent of the cortical areas
in the brain - that follows the Gardenfors theory of conceptual spaces [
        <xref ref-type="bibr" rid="ref19">19</xref>
        ] (see
Forth et al. [
        <xref ref-type="bibr" rid="ref20">20</xref>
        ] for a discussion on conceptual spaces and musical systems).
      </p>
      <p>In the following, Section 2 outlines the cognitive architecture for music
perception, while Section 3 describes the adopted music conceptual space for the
perception of tones. Section 4 presents the linguistic area of the cognitive
architecture and Section 5 presents the related operations of the focus of attention.
Section 6 outlines the generalization of the conceptual space for tones perception
to the case of perception of music phrases, and nally Section 7 discusses some
problems of the proposed approach and future works.
2</p>
    </sec>
    <sec id="sec-2">
      <title>The Cognitive Architecture</title>
      <p>The proposed cognitive architecture for music perception is sketched in Figure 1.
The areas of the architecture are concurrent computational components working
together on di erent commitments. There is no privileged direction in the ow of
information among them: some computations are strictly bottom-up, with data
owing from the subconceptual up to the linguistic through the conceptual area;
other computations combine top-down with bottom-up processing.</p>
      <p>Conceptual</p>
      <p>Area</p>
      <p>Linguistic</p>
      <p>Area
Subconceptual</p>
      <p>Area</p>
      <p>The subconceptual area of the proposed architecture is concerned with the
processing of data directly coming from the sensors. Here, information is not
yet organized in terms of conceptual structures and categories. In the linguistic
area, representation and processing are based on a logic-oriented formalism.</p>
      <p>The conceptual area is an intermediate level of representation between the
subconceptual and the linguistic areas and based on conceptual spaces. Here,
data is organized in conceptual structures, that are still independent of linguistic
description. The symbolic formalism of the linguistic area is then interpreted on
aggregation of these structures.</p>
      <p>It is to be remarked that the proposed architecture cannot be considered as
a model of human perception. No hypotheses concerning its cognitive adequacy
from a psychological point of view have been made. However, various cognitive
results have been taken as sources of inspiration.
3</p>
    </sec>
    <sec id="sec-3">
      <title>Music Conceptual Space</title>
      <p>The conceptual area, as previously stated, is the area between the subconceptual
and the linguistic area, and it is based on conceptual spaces. We adopt the term
knoxel (in analogy with the term pixel ) to denote a point in a conceptual space
CS. The choice of this term stresses the fact that a point in CS is the knowledge
primitive element at the considered level of analysis.</p>
      <p>The conceptual space acts as a workspace in which low-level and high-level
processes access and exchange information respectively from bottom to top and
from top to bottom. However, the conceptual space has a precise geometric
structure of metric space and also the operations in CS are geometric ones: this
structure allows us to describe the functionalities of the cognitive architecture
in terms of the language of geometry.</p>
      <p>
        In particular, inspired by many empirical investigations on the perception of
tones (see Oxenham [
        <xref ref-type="bibr" rid="ref21">21</xref>
        ] for a review) we adopt as a knoxel of a music
conceptual space the set of partials of a perceived tone. A knoxel k of the music CS is
therefore a vector of the main perceived partials of a tone in terms of the Fourier
Transform analysis. A similar choice has been carried out by Tanguiane [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]
concerning his proposed correlativity model of perception.
      </p>
      <p>
        It should be noticed that the partials of a tone are related both with the
pitch and the timbre of the perceived note. Roughly, the fundamental frequency
is related with the pitch, while the amplitude of the remaining partials are also
related with the timbre of the note. By an analogy with the case of static scenes
analysis, a knoxel changes its position in CS when a perceived 3D primitive
changes its position in space or its shape [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ]; in the case of music perception,
the knoxel in the music CS changes its position either when the perceived sound
changes its pitch or its timbre changes as well. Moreover, considering the partials
of a tone allows us to deal also with microtonal tones, trills, embellished notes,
rough notes, and so on.
      </p>
      <p>A chord is a set of two or more tones perceived at the same time. The chord
is treated as a complex object, in analogy with static scenes analysis where a
complex object is an object made up by two or more 3D primitives. A chord is
then represented in music CS as the set of the knoxels [ka; kb; : : : ] related with
the constituent tones. It should be noticed that the tones of a chord may di er
not only in pitch, but also in timbre. Figure 2 is an evocative representation of
a chord in the music CS made up by knoxel ka corresponding to tone C and
knoxel kb corresponding to the tone G.</p>
      <p>In the case of perception of complex objects in vision, their mutual positions
and shapes are important in order to describe the perceived object: e.g., in the
case of an hammer, the mutual positions and the mutual shapes of the handle
and the head are obviously important to classify the composite object as an
hammer. In the same way, the mutual relationships between the pitches (and the
timbres) of the perceived tones are important in order to describe the perceived
chord. Therefore, spatial relationships in static scenes analysis are in some sense
analogous to sounds relationships in music CS.</p>
      <p>It is to be noticed that this approach allows us to represent a chord as a set
of knoxels in music CS. In this way, the cardinality of the conceptual space does
not change with the number of tones forming the chord. In facts, all the tones of
the chord are perceived at the same time but they are represented as di erent
points in the same music CS; that is, the music CS is a sort of snapshot of the
set of the perceived tones of the chord.</p>
      <p>In the case of a temporal progression of chords, a scattering occur in the
music CS: some knoxels which are related with the same tones between chords
will remain in the same position, while other knoxels will change their position
in CS, see Figure 3 for an evocative representation of scattering in the music CS.
In the gure, the knoxels ka, corresponding to C, and kb, corresponding to E,</p>
      <p>Ax(3)</p>
      <p>Ay(0)
Ay(1)
Ay(2)
Ay(3)
change their position in the new chord: they becomes A and D, while knoxel kc,
corresponding to G, maintains its position. The relationships between mutual
positions in music CS could then be employed to analyze the chords progression
and the relationships between subsequent chords.</p>
      <p>
        A problem may arise at this point. In facts, in order to analyze the
progression of chords, the system should be able to nd the correct correspondences
between subsequent knoxels: i.e., k0a should correspond to ka and not to, e.g.,
kb. This is a problem similar to the correspondence problem in stereo and in
visual motion analysis: a vision system analyzing subsequent frames of a moving
object should be able to nd the correct corresponding object tokens among the
motion frames; see the seminal book by Ullman [
        <xref ref-type="bibr" rid="ref22">22</xref>
        ] or Chap. 11 of the recent
book by Szeliski [
        <xref ref-type="bibr" rid="ref23">23</xref>
        ] for a review. However, it should be noticed that the
expectation generation mechanism described in Section 5 could greatly help facing
this di cult problem.
      </p>
      <p>
        The described representation is well suited for the recognition of chords: for
example we may adopt the algorithms proposed by Tanguiane [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]. However,
Tanguiane hypothesizes, at the basis of his correlativity principle, that all the
notes of a chord have the same shifted partials, while we consider the possibility
that a chord could be made by tones with di erent partials.
      </p>
      <p>
        The proposed representation is also suitable for the analysis of the e ciency
in voice leading, as described by Tymoczko [
        <xref ref-type="bibr" rid="ref24">24</xref>
        ]. Tymoczko describes a
geometrical analysis of chords by considering several spaces with di erent cardinalities,
      </p>
      <p>Ax(3)</p>
      <p>Ay(0)
Ay(1)
Ay(2)
Ay(3)
i.e., a one note circular space, a two note space, a three note space, and so on.
Instead, the cardinality of the considered conceptual space does not change, as
previously remarked.
4</p>
    </sec>
    <sec id="sec-4">
      <title>Linguistic area</title>
      <p>
        In the linguistic area, the representation of perceived tones is based on a high
level, logic oriented formalism. The linguistic area acts as a sort of long term
memory, in the sense that it is a semantic network of symbols and their
relationships related with musical perceptions. The linguistic area also performs
inferences of symbolic nature. In preliminary experiments, we adopted a
linguistic area based on a hybrid KB in the KL-ONE tradition [
        <xref ref-type="bibr" rid="ref25">25</xref>
        ]. A hybrid formalism
in this sense is constituted by two di erent components: a terminological
component for the description of concepts, and an assertional component, that stores
information concerning a speci c context. A similar formalism has been adopted
by Camurri et al. in the HARP system [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ][
        <xref ref-type="bibr" rid="ref12">12</xref>
        ].
      </p>
      <p>In the domain of perception of tones, the terminological component contains
the description of relevant concepts such as chords, tonic, dominant and so on.
The assertional component stores the assertions describing speci c situations.
Figure 4 shows a fragment of the terminological knowledge base along with its
mapping into the corresponding entities in the conceptual space.
has-dominant</p>
      <p>has-tonic
Dominant</p>
      <p>Tonic</p>
      <p>A generic Chord is described as composed of at least two knoxels. A
SimpleChord is a chord composed by two knoxels; a Complex-Chord is a chord composed
of more than two knoxels. In the considered case, the concept Chord has two
roles: a role has-dominant, and a role has-tonic both lled with speci c tones.</p>
      <p>In general, we assume that the description of the concepts in the symbolic
KB is not exhaustive. We symbolically represent the information necessary to
make suitable inferences.</p>
      <p>The assertional component contains facts expressed as assertions in a
predicative language, in which the concepts of the terminological components
correspond to one argument predicates, and the roles (e.g., part of) correspond to
two argument relations. For example, the following predicates describe that the
instance f7#1 of the F7 chord has a dominant which is the constant ka
corresponding to a knoxel ka and a tonic which is the constant k#b corresponding to
a knoxel kb of the current CS:
ChordF7(f7#1)
has-dominant(f7#1,ka)
has-tonic(f7#1,kb)</p>
      <p>By means of the mapping between symbolic KB and conceptual spaces, the
linguistic area assigns names (symbols) to perceived entities, describing their
structure with a logical-structural language. As a result, all the symbols in the
linguistic area nd their meaning in the conceptual space which is inside the
system itself.</p>
      <p>
        A deeper account of these aspects can be found in Chella et at. [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ].
5
      </p>
    </sec>
    <sec id="sec-5">
      <title>Focus of Attention</title>
      <p>A cognitive architecture with bounded resources cannot carry out a one-shot,
exhaustive, and uniform analysis of the perceived data within reasonable resource
constraints. Some of the perceived data (and of the relations among them) are
more relevant than others, and it should be a waste of time and of computational
resources to detect true but useless details.</p>
      <p>In order to avoid the waste of computational resources, the association
between symbolic representations and con gurations of knoxels in CS is driven
by a sequential scanning mechanism that acts as some sort of internal focus of
attention, and inspired by the attentive processes in human perception.</p>
      <p>In the considered cognitive architecture for music perception, the perception
model is based on a focus of attention that selects the relevant aspects of a sound
by sequentially scanning the corresponding knoxels in the conceptual space. It is
crucial in determining which assertions must be added to the linguistic knowledge
base: not all true (and possibly useless) assertions are generated, but only those
that are judged to be relevant on the basis of the attentive process.</p>
      <p>The recognition of a certain component of a perceived con guration of
knoxels in music CS will elicit the expectation of other possible components of the
same chord in the perceived conceptual space con guration. In this case, the
mechanism seeks for the corresponding knoxels in the current CS con guration.
We call this type of expectation synchronic because it refers to a single con
guration in CS.</p>
      <p>
        The recognition of a certain con guration in CS could also elicit the
expectation of a scattering in the arrangement of the knoxels in CS; i.e., the
mechanism generates the expectations for another set of knoxels in a subsequent CS
con guration. We call this expectation diachronic, in the sense that it involves
subsequent con gurations of CS. Diachronic expectations can be related with
progression of chords. For example, in the case of jazz music, when the system
recognized the Cmajor key (see Rowe [
        <xref ref-type="bibr" rid="ref26">26</xref>
        ] for a catalogue of key induction
algorithms) and a Dm chord is perceived, then the focus of attention will generate
the expectations of G and C chords in order to search for the well known chord
progression ii V I (see Chap. 10 of Tymoczko [
        <xref ref-type="bibr" rid="ref24">24</xref>
        ]).
      </p>
      <p>
        Actually, we take into account two main sources of expectations. On the one
side, expectations could be generated on the basis of the structural
information stored in the symbolic knowledge base, as in the previous example of the
jazz chord sequence. We call these expectations linguistic. Several sources may
be taken into account in order to generate linguistic expectations, for example
the ITPRA theory of expectation proposed by Huron [
        <xref ref-type="bibr" rid="ref27">27</xref>
        ], the preference rules
systems discussed by Temperley [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ] or the rules of harmony and voice leading
discussed in Tymoczko [
        <xref ref-type="bibr" rid="ref24">24</xref>
        ], just to cite a few. As an example, as soon as a
particular con guration of knoxel is recognized as a possible chord lling the role
of the rst chord of the progression ii V I, the symbolic KB generates the
expectation of the remaining chords of the sequence.
      </p>
      <p>On the other side, expectations could be generated by purely Hebbian,
associative mechanisms. Suppose that the system learnt that typically a jazz player
adopts the tritone substitution when performing the previous described jazz
progression. The system could learn to associate this substitution to the progression:
in this case, when a compatible chord is recognized, the system will generate also
expectations for the sequence ii [II I. We call these expectations associative.</p>
      <p>Therefore, synchronic expectations refer to the same con guration of knoxels
at the same time; diachronic expectations involve subsequent con gurations of
knoxels. The linguistic and associative mechanisms let the cognitive architecture
generate suitable expectations related to the perceived chords progressions.
6</p>
    </sec>
    <sec id="sec-6">
      <title>Perception of Music Phrases</title>
      <p>So far we adopted a \static" conceptual space where a knoxel represents the
partials of a perceived tone. In order to generalize this concept and in analogy
with the di erences between static and dynamic vision, in order to represent a
music phrase, we now adopt a \dynamic" conceptual space in which each knoxel
represents the whole set of partials of the Short Time Fourier Transform of the
corresponding music phrase. In other words, a knoxel in the dynamic CS now
represents all the parameters of the spectrogram of the perceived phrase.</p>
      <p>
        Therefore, inspired by empirical results (see Deutsch [
        <xref ref-type="bibr" rid="ref28">28</xref>
        ] for a review) we
hypothesize that a musical phrase is perceived as a whole \Gestaltic" group, in
the same way as a movement could be visually perceived as a whole and not
as a sequence of single frames. It should be noticed that, similarly to the static
case, a knoxel represents the sequence of pitches and durations of the perceived
phrase and also its timbre: the same phrase played by two di erent instruments
corresponds to two di erent knoxels in the dynamic CS.
      </p>
      <p>The operations in the dynamic CS are largely similar to the static CS, with
the main di erence that now a knoxel is a whole perceived phrase.</p>
      <p>A con guration of knoxels in CS occurs when two or more phrases are
perceived at the same time. The two phrases may be related with two di erent
sequences of pitches or it may be the same sequence played for example, by two
di erent instruments. This is similar to the situation depicted in Figure 2, where
the knoxels ka and kb are interpreted as music phrases perceived at the same
time.</p>
      <p>A scattering of knoxels occurs when a change occurs in a perceived phrase.
We may represent this scattering in a similar way to the situation depicted in
Figure 3, where the knoxels also in this case are interpreted as music phrases:
knoxels ka and kb are interpreted as changed music phrases while knoxels kc
corresponds to the same perceived phrase.</p>
      <p>
        As an example, let us consider the well known piece In C by Terry Riley.
The piece is composed by 53 small phrases to be performed sequentially; each
player may decide when to start playing, how many times to repeat the same
phrase, and when to move to the next phrase (see the performing directions of
In C [
        <xref ref-type="bibr" rid="ref29">29</xref>
        ]).
      </p>
      <p>Let us consider the case in which two players, with two di erent instruments,
start with the rst phrase. In this case, two knoxels ka and kb will be activated
in the dynamic CS. We remark that, although the phrase is the same in terms of
pitch and duration, it corresponds to two di erent knoxels because of di erent
timbres of the two instruments. When a player will decide at some time to move
to next phrase, a scattering occur in the dynamic CS, analogously with the
previous analyzed static CS: the corresponding knoxel, say ka, will change its
position to k0 .</p>
      <p>a</p>
      <p>The focus of attention mechanism will operate in a similar way as in the
static case: the synchronous modality of the focus of attention will take care of
generation of expectations among phrases occurring at the same time, by taking
into account, e.g., the rules of counterpoint. Instead, the asynchronous modality
will generate expectations concerning, e.g., the continuation of phrases.</p>
      <p>Moreover, the static CS and the dynamic CS could generate mutual
expectations: for example, when the focus of attention recognizes a progression of chords
in the static CS, this recognized progression will constraint the expectations of
phrases in the dynamic CS. As another example, the recognition of a phrase in
the dynamic CS could constraint as well the recognition of the corresponding
progression of chords in the static CS.
7</p>
    </sec>
    <sec id="sec-7">
      <title>Discussion and Conclusions</title>
      <p>The paper sketched a cognitive architecture for music perception extending and
completing a computer vision cognitive architecture. The architecture integrates
symbolic and the sub symbolic approaches by means of conceptual spaces and it
takes into account many relationships between vision and music perception.</p>
      <p>
        Several problems arise concerning the proposed approach. A rst problem,
analogously with the case of computer vision, concerns the segmentation step.
In the case of static CS, the cognitive architecture should be able to segment
the Fourier Transform signal coming from the microphone in order to
individuate the perceived tones; in the case of dynamic CS the architecture should be
able to individuate the perceived phrases. Although many algorithms for music
segmentation have been proposed in the computer music literature and some
of them are also available as commercial program, as the AudioSculpt program
developed by IRCAM1, this is a main problem in perception. Interestingly,
empirical studies concur in indicating that the same Gestalt principles at the basis
of visual perception operate in similar ways in music perception, as discussed by
Deutsch [
        <xref ref-type="bibr" rid="ref28">28</xref>
        ].
      </p>
      <p>
        The expectation generation process at the basis of the focus of attention
mechanism can be employed to help solving the segmentation problem: the
linguistic information and the associative mechanism can provide interpretation
1 http://forumnet.ircam.fr/product/audiosculpt/
contexts and high level hypotheses that help segmenting the audio signal, as
e.g., in the IPUS system [
        <xref ref-type="bibr" rid="ref30">30</xref>
        ].
      </p>
      <p>
        Another problem is related with the analysis of time. Currently, the proposed
architecture does not take into account the metrical structure of the perceived
music. Successive development of the described architecture will concern a
metrical conceptual space; interesting starting points are the geometric models of
metrical-rhythmic structure discussed by Forth et al. [
        <xref ref-type="bibr" rid="ref20">20</xref>
        ].
      </p>
      <p>However, we maintain that an intermediate level based on conceptual spaces
could be a great help towards the integration between the music cognitive
systems based on subsymbolic representations, and the class of systems based on
symbolic models of knowledge representation and reasoning. In facts,
conceptual spaces could o er a theoretically well founded approach to the integration
of symbolic musical knowledge with musical neural networks.</p>
      <p>Finally, as stated during the paper, the synergies between music and vision
are multiple and multifaceted. Future works will deal with the exploitation of
conceptual spaces as a framework towards a sort of uni ed theory of perception
able to integrate in a principled way vision and music perception.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1. Gardenfors, P.:
          <article-title>Semantics, conceptual spaces andthe dimensions of music</article-title>
          . In Rantala, V.,
          <string-name>
            <surname>Rowell</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Tarasti</surname>
          </string-name>
          , E., eds.
          <source>: Essays on the Philosophy of Music. Philosophical Society of Finland</source>
          , Helsinki (
          <year>1988</year>
          )
          <volume>9</volume>
          {
          <fpage>27</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Marr</surname>
            ,
            <given-names>D.: Vision. W.H.</given-names>
          </string-name>
          <string-name>
            <surname>Freeman</surname>
          </string-name>
          and Co., New York (
          <year>1982</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Tanguiane</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          :
          <article-title>Arti cial Perception and Music Recognition</article-title>
          .
          <source>Number 746 in Lecture Notes in Arti cial Intelligence</source>
          . Springer-Verlag, Berlin Heidelberg (
          <year>1993</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Wiggins</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Pearce</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <article-title>Mullensiefen: Computational modelling of music cognition and musical creativity</article-title>
          . In
          <string-name>
            <surname>Dean</surname>
          </string-name>
          , R., ed.:
          <source>The Oxford Handbook of Computer Music</source>
          . Oxford University Press, Oxford (
          <year>2009</year>
          )
          <volume>387</volume>
          {
          <fpage>414</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Temperley</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          :
          <article-title>Computational models of music cognition</article-title>
          . In Deutsch, D., ed.:
          <article-title>The Psychology of Music. Third edn</article-title>
          . Academic Press, Amsterdam, The Netherlands (
          <year>2012</year>
          )
          <volume>327</volume>
          {
          <fpage>368</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Bharucha</surname>
          </string-name>
          , J.:
          <article-title>Music cognition and perceptual facilitation: A connectionist framework</article-title>
          .
          <source>Music Perception: An Interdisciplinary Journal</source>
          <volume>5</volume>
          (
          <issue>1</issue>
          ) (
          <year>1987</year>
          )
          <volume>1</volume>
          {
          <fpage>30</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>Bharucha</surname>
          </string-name>
          , J.:
          <article-title>Pitch, harmony and neural nets: A psychological perspective</article-title>
          . In Todd, P.,
          <string-name>
            <surname>Loy</surname>
          </string-name>
          , D., eds.: Music and Connectionism. MIT Press, Cambridge, MA (
          <year>1991</year>
          )
          <volume>84</volume>
          {
          <fpage>99</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>Pearce</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Wiggins</surname>
          </string-name>
          , G.:
          <article-title>Improved methods for statistical modelling of monophonic music</article-title>
          .
          <source>Journal of New Music Research</source>
          <volume>33</volume>
          (
          <issue>4</issue>
          ) (
          <year>2004</year>
          )
          <volume>367</volume>
          {
          <fpage>385</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <surname>Pearce</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Wiggins</surname>
          </string-name>
          , G.:
          <article-title>Expectation in melody: The in uence of context and learning</article-title>
          .
          <source>Music Perception: An Interdisciplinary Journal</source>
          <volume>23</volume>
          (
          <issue>5</issue>
          ) (
          <year>2006</year>
          )
          <volume>377</volume>
          {
          <fpage>406</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <surname>Temperley</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          :
          <article-title>The Cognition of Basic Musical Structures</article-title>
          . MIT Press, Cambridge, MA (
          <year>2001</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <string-name>
            <surname>Camurri</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Frixione</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Innocenti</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          :
          <article-title>A cognitive model and a knowledge representation system for music and multimedia</article-title>
          .
          <source>Journal of New Music Research</source>
          <volume>23</volume>
          (
          <year>1994</year>
          )
          <volume>317</volume>
          {
          <fpage>347</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          12.
          <string-name>
            <surname>Camurri</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Catorcini</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Innocenti</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Massari</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          :
          <article-title>Music and multimedia knowledge representation and reasoning: the HARP system</article-title>
          .
          <source>Computer Music Journal</source>
          <volume>19</volume>
          (
          <issue>2</issue>
          ) (
          <year>1995</year>
          )
          <volume>34</volume>
          {
          <fpage>58</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          13.
          <string-name>
            <surname>Chella</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Frixione</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gaglio</surname>
            ,
            <given-names>S.:</given-names>
          </string-name>
          <article-title>A cognitive architecture for arti cial vision</article-title>
          .
          <source>Arti cial Intelligence</source>
          <volume>89</volume>
          (
          <year>1997</year>
          )
          <volume>73</volume>
          {
          <fpage>111</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          14.
          <string-name>
            <surname>Chella</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Frixione</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gaglio</surname>
            ,
            <given-names>S.:</given-names>
          </string-name>
          <article-title>An architecture for autonomous agents exploiting conceptual representations</article-title>
          .
          <source>Robotics and Autonomous Systems</source>
          <volume>25</volume>
          (
          <issue>3-4</issue>
          ) (
          <year>1998</year>
          )
          <volume>231</volume>
          {
          <fpage>240</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          15.
          <string-name>
            <surname>Chella</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Frixione</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gaglio</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          :
          <article-title>Understanding dynamic scenes</article-title>
          .
          <source>Arti cial Intelligence</source>
          <volume>123</volume>
          (
          <year>2000</year>
          )
          <volume>89</volume>
          {
          <fpage>132</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          16.
          <string-name>
            <surname>Chella</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gaglio</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Pirrone</surname>
          </string-name>
          , R.:
          <article-title>Conceptual representations of actions for autonomous robots</article-title>
          .
          <source>Robotics and Autonomous Systems</source>
          <volume>34</volume>
          (
          <year>2001</year>
          )
          <volume>251</volume>
          {
          <fpage>263</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          17.
          <string-name>
            <surname>Chella</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Frixione</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gaglio</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          :
          <article-title>Anchoring symbols to conceptual spaces: the case of dynamic scenarios</article-title>
          .
          <source>Robotics and Autonomous Systems</source>
          <volume>43</volume>
          (
          <issue>2-3</issue>
          ) (
          <year>2003</year>
          )
          <volume>175</volume>
          {
          <fpage>188</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          18.
          <string-name>
            <surname>Chella</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Frixione</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gaglio</surname>
            ,
            <given-names>S.:</given-names>
          </string-name>
          <article-title>A cognitive architecture for robot selfconsciousness</article-title>
          .
          <source>Arti cial Intelligence in Medicine</source>
          <volume>44</volume>
          (
          <year>2008</year>
          )
          <volume>147</volume>
          {
          <fpage>154</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          19. Gardenfors, P.: Conceptual Spaces. MIT Press, Bradford Books, Cambridge, MA (
          <year>2000</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          20.
          <string-name>
            <surname>Forth</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Wiggins</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>McLean</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          :
          <article-title>Unifying conceptual spaces: Concept formation in musical creative systems</article-title>
          .
          <source>Minds and Machines</source>
          <volume>20</volume>
          (
          <year>2010</year>
          )
          <volume>503</volume>
          {
          <fpage>532</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          21.
          <string-name>
            <surname>Oxenham</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          :
          <article-title>The perception of musical tones</article-title>
          . In Deutsch, D., ed.:
          <article-title>The Psychology of Music. Third edn</article-title>
          . Academic Press, Amsterdam, The Netherlands (
          <year>2013</year>
          )
          <volume>1</volume>
          {
          <fpage>33</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          22.
          <string-name>
            <surname>Ullman</surname>
            ,
            <given-names>S.:</given-names>
          </string-name>
          <article-title>The Interpretation of Visual Motion</article-title>
          . MIT Press, Cambridge,MA (
          <year>1979</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          23.
          <string-name>
            <surname>Szeliski</surname>
          </string-name>
          , R.:
          <source>Computer Vision: Algorithms and Applications</source>
          . Springer, London (
          <year>2011</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          24.
          <string-name>
            <surname>Tymoczko</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          :
          <article-title>A Geometry of Music. Harmony and Counterpoint in the Extended Common Practice</article-title>
          . Oxford University Press, Oxford (
          <year>2011</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref25">
        <mixed-citation>
          25.
          <string-name>
            <surname>Brachman</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Schmoltze</surname>
            ,
            <given-names>J.:</given-names>
          </string-name>
          <article-title>An overview of the KL-ONE knowledge representation system</article-title>
          .
          <source>Cognitive Science</source>
          <volume>9</volume>
          (
          <issue>2</issue>
          ) (
          <year>1985</year>
          )
          <volume>171</volume>
          {
          <fpage>216</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref26">
        <mixed-citation>
          26.
          <string-name>
            <surname>Rowe</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          <article-title>: Machine Musicianship</article-title>
          . MIT Press, Cambridge, MA (
          <year>2001</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref27">
        <mixed-citation>
          27.
          <string-name>
            <surname>Huron</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          : Sweet Anticipation.
          <article-title>Music and the Psychology of Expectation</article-title>
          . MIT Press, Cambridge, MA (
          <year>2006</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref28">
        <mixed-citation>
          28.
          <string-name>
            <surname>Deutsch</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          :
          <article-title>Grouping mechanisms in music</article-title>
          . In Deutsch, D., ed.:
          <article-title>The Psychology of Music. Third edn</article-title>
          . Academic Press, Amsterdam, The Netherlands (
          <year>2013</year>
          )
          <volume>183</volume>
          {
          <fpage>248</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref29">
        <mixed-citation>
          29.
          <string-name>
            <surname>Riley</surname>
          </string-name>
          , T.: In C :
          <article-title>Performing directions</article-title>
          .
          <source>Celestial Harmonies</source>
          (
          <year>1964</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref30">
        <mixed-citation>
          30.
          <string-name>
            <surname>Lesser</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Nawab</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Klassner</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          :
          <article-title>IPUS: An architecture for the integrated processing and understanding of signals</article-title>
          .
          <source>Arti cial Intelligence</source>
          <volume>77</volume>
          (
          <year>1995</year>
          )
          <volume>129</volume>
          {
          <fpage>171</fpage>
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>