<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Breathing position influences speech perception</article-title>
      </title-group>
      <contrib-group>
        <aff id="aff0">
          <label>0</label>
          <institution>CNRS &amp; Université Paris Descartes</institution>
          ,
          <addr-line>Paris</addr-line>
          ,
          <country country="FR">France</country>
        </aff>
      </contrib-group>
      <fpage>425</fpage>
      <lpage>429</lpage>
      <abstract>
        <p>Participants were asked to breath through their mouth or their nose, forcing them to adopt a particular position of their velum (up or down). While breathing in each of these positions, they categorized sounds from an /ada/ to /ana/ continuum. The position of the speech articulators, even though adopted for the purposes of breathing, altered participants' perception of external speech sounds, so that when the velum was down (to breath through the nose), they tended to hear the consonant as the nasal /n/ - a sound necessarily produced with a lowered velum, rather than as /d/ in which the velum must be raised.</p>
      </abstract>
      <kwd-group>
        <kwd>speech perception</kwd>
        <kwd>articulation</kwd>
        <kwd>corollary discharge</kwd>
        <kwd>efference copy</kwd>
        <kwd>forward model</kwd>
        <kwd>motor control</kwd>
        <kwd>breathing</kwd>
        <kwd>motor theory</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Introduction</title>
      <p>
        If the organs of speech production are moved during speech
perception, that movement can influence how perception
unfolds. For example, if hearing a sound ambiguous between
/aba/ and /ava/ while mouthing (or even imagining) /ava/,
a person will tend to hear the sound as /ava/ (and
contrariwise for mouthing /aba/)
        <xref ref-type="bibr" rid="ref24">(Scott, Yeung, Gick, &amp; Werker,
2013)</xref>
        . This phenomenon is purportedly due to the
perceptual anticipation (in the form of corollary discharge) caused
by mouthing/imagining. This experiment explores whether a
similar perceptual capture can occur even when the organs of
speech production are not being engaged in a speech task —
specifically, whether the position of the velum (up for
breathing through the mouth, down for breathing through the nose)
can influence the perception of nasal vs. non-nasal stop
consonants.
      </p>
      <sec id="sec-1-1">
        <title>Motor involvement in speech perception</title>
        <p>
          The question of whether (or to what degree) the motor system
influences speech perception is old and strongly contested.
The best known version of this idea is the Motor Theory
of Speech Perception
          <xref ref-type="bibr" rid="ref14">(Liberman &amp; Mattingly, 1985)</xref>
          which
claims that speech perception is achieved via a specialized
module which extracts the intended speech gestures from the
acoustic signal.
        </p>
        <p>
          A somewhat similar view is proposed by
          <xref ref-type="bibr" rid="ref4">Fowler (1986)</xref>
          who offers a Gibsonian approach to speech perception in
which it is speech gestures that are recovered in perception
but, in contrast to the Motor Theory, this recovery is through
general auditory mechanisms rather than a biologically
specialized speech-perception mechanism. The view that the
fundamental units of speech perception (and production) are
gestures is also shared by the Gestural Phonology approach
          <xref ref-type="bibr" rid="ref6">(Goldstein &amp; Fowler, 2003)</xref>
          .
        </p>
        <p>
          More recent alternatives to the Motor Theory (and its
variants) maintain a role for the motor system, but do not claim
that it is an obligatory component of speech perception. For
example
          <xref ref-type="bibr" rid="ref19">Pickering and Garrod (2007)</xref>
          , propose that when
perception is faced with a difficult task, top-down information in
the form of motor-predictions can be used to ‘fill in the gaps’.
According to this theory, speech perception would primarily
be an auditory process but when the auditory signal is
particularly unclear the motor system makes predictions about what
is to come and, in so doing, constrains the possibilities that
the auditory system must entertain, thus easing the
computational load. This would be a hybrid motor/auditory view of
speech perception. Skipper, Nusbaum, and Small (2006) and
Skipper, van Wassenhove, Nusbaum, and Small (2007) have
proposed a similar theory, again arguing that speech
perception is not necessarily a matter of coding the incoming sound
into a gestural code, but that engagement of predictions from
the motor system can be used to aid in speech-perception.
          <xref ref-type="bibr" rid="ref26">Skipper et al. (2007)</xref>
          argue that such predictions are what
underlie the influence of vision on speech perception, such as in
the McGurk effect
          <xref ref-type="bibr" rid="ref15">(McGurk &amp; MacDonald, 1976)</xref>
          .
        </p>
        <p>
          These motor-helping-hearing theories are quite
similar to the Perception-for-Action-Control Theory
          <xref ref-type="bibr" rid="ref22">(Schwartz,
Basirat, Ménard, &amp; Sato, 2010)</xref>
          . This theory argues that
speech perception is not motor-based, but that speech
gestures do define equivalence classes for speech sounds. The
idea is that the motor system helps establish which sounds
count as members of the same category, membership
being determined by sharing a common method of production.
However, once the sound classes are set, the motor system is
normally not used online in the act of perceiving the members
of these classes, such online perception being achieved by the
auditory system.
          <xref ref-type="bibr" rid="ref22">Schwartz et al. (2010)</xref>
          argue for one
exception to the independence of sensory and motor processes —
when auditory perception is made difficult because of
missing information, the motor system can be used to ‘fill in’ that
missing information.
        </p>
      </sec>
      <sec id="sec-1-2">
        <title>Corollary Discharge</title>
        <p>These recent alternatives to the Motor Theory propose a
specific mechanism by which the motor system influences
speech perception: corollary discharge.1</p>
        <p>
          Corollary discharge is an internal sensory signal generated
by one’s own motor system whenever one acts
          <xref ref-type="bibr" rid="ref1">(Aliu, Houde,
&amp; Nagarajan, 2009)</xref>
          . One primary function of corollary
discharge is to provide pseudo feedback in situations where
regular sensory feedback is too slow to guide one’s actions.
There is an unavoidable time delay in real sensory feedback
— our senses do not operate instantaneously, it takes time for
a change in the environment (or in our body) to be transduced
by our end-organs, then transmitted to and processed by the
central nervous system and for a motor correction on the
basis of this information to be issued. This delay can be quite
considerable – for auditory speech perception it has been
estimated at around 130 ms
          <xref ref-type="bibr" rid="ref11">(Jones &amp; Munhall, 2002)</xref>
          . This
means that feedback is not available (or minimally available)
for speech movements that are faster than 130ms (which is in
fact many speech sounds). Corollary discharge can be
generated by the motor system before sensory feedback is available
and thus can serve the feedback role and avoid the time-lag
problem
          <xref ref-type="bibr" rid="ref27">(Wolpert &amp; Flanagan, 2001)</xref>
          .
        </p>
        <p>
          Another role of corollary discharge is to tag concurrent
matching external sensations as “self-caused” and so
unworthy of intense perceptual processing
          <xref ref-type="bibr" rid="ref3">(Eliades &amp; Wang,
2008)</xref>
          . This function means that corollary discharge functions
to anticipate perceptions and so can pull ambiguous
stimuli into alignment with the anticipated percept.
          <xref ref-type="bibr" rid="ref24">Scott et al.
(2013)</xref>
          hypothesized that corollary discharge, generated
during mouthing of speech, channels the perception of external
speech into matching the corollary discharge prediction, thus
performing a perceptual capture function.
        </p>
      </sec>
      <sec id="sec-1-3">
        <title>Perceptual Capture</title>
        <p>
          Perceptual capture is a shift in perception caused by the
fact that corollary discharge is an anticipation, and as such
can pull ambiguous stimuli into alignment with the
anticipated percept. Hickok,
          <xref ref-type="bibr" rid="ref9">Houde, and Rong (2011)</xref>
          provide an
overview of how corollary discharge influences perception
through its role as an anticipation.
        </p>
        <p>
          <xref ref-type="bibr" rid="ref20">Repp and Knoblich (2009)</xref>
          demonstrated perceptual
capture from motor-induced anticipations by having pianists
perform hand motions for a rising or falling sequence of notes.
These hand motions were performed in synchrony with a
sequence of notes that could be heard (thanks to a perceptual
illusion) as rising or falling. When performing the hand motion
consistent with a rising sequence, pianists tended to hear the
ambiguous sound as ascending.
          <xref ref-type="bibr" rid="ref23">Schütz-Bosbach and Prinz
(2007)</xref>
          review several such perceptual capture effects across
sensory modalities.
        </p>
        <p>
          In terms of hearing, several studies have found evidence of
anticipations altering the perception of sounds. A compelling
1Corollary discharge is the sensory prediction generated by a
‘forward model’ – which is a system that takes motor commands as
input and predicts the sensory consequences. Efference copy refers
to the motor command received by the forward model, but
sometimes the terms corollary discharge and efference copy are used
interchangeably.
example of this is the “White Christmas” effect
          <xref ref-type="bibr" rid="ref16">(Merckelbach
&amp; Ven, 2001)</xref>
          in which people are induced to hear the song
“White Christmas” when presented with white noise,
simply by telling them that the song might be buried under the
noise (but is not). Such perceptual shifts in speech
perception arising from the influence of the motor system have been
shown in other studies such as Sams, Möttönen, and Sihvonen
(2005), Ito, Tiede, and Ostry (2009) and
          <xref ref-type="bibr" rid="ref24">Scott et al. (2013)</xref>
          .
        </p>
        <p>
          A consequence of theories which propose a perceptual ‘fill
in the gap’ role for corollary discharge is that the position of
the perceiver’s own articulators should matter for what
information gets filled in. For a prediction of the sensory
consequences of an action to be accurate it must take into account
the starting point of the action, as the sensory consequences
of an action can be vastly different depending on where the
effector is starting — think of the tactile sensory difference
between slamming your jaw shut when your tongue is in its
normal resting position vs. extended out between your teeth
(ouch!). Thus corollary discharge is necessarily generated
using the current position of the articulators as the basis for
prediction
          <xref ref-type="bibr" rid="ref7 ref8 ref9">(Houde &amp; Nagarajan, 2011; Hickok, 2012)</xref>
          .
        </p>
      </sec>
    </sec>
    <sec id="sec-2">
      <title>Prediction</title>
      <p>Given that the predictions of corollary discharge take into
account the position of the effectors, the perceptual channelling
discussed above should be sensitive to the positions of one’s
own speech articulators (even if one is not speaking). Thus
when a person’s velum is down in order to allow breathing
through the nose, then the perceptual channelling (or ‘filling
in’) done by the motor system should be biased towards a
prediction of nasality and the person should thus be more likely
to hear an ambiguous external sound as nasal. In the
context of this experiment, this means that people should hear an
/ada/~/aga/ ambiguous sound more often as /ana/ when they
are breathing through their nose in comparison to when they
are breathing through their mouth.</p>
      <p>The sounds /d/ and /n/ were chosen as these sounds have
the same place of articulation and are both voiced and are
both stops, differing almost exclusively in whether the velum
is up (for /d/ and so no airflow through the nose) or down (for
/n/ with airflow through the nose). Thus the primary
difference between these sounds is mirrored in the position of the
velum for breathing — up for breathing through the mouth,
down for breathing through the nose.</p>
      <p>
        The prediction that people should hear more /n/ when their
velum is down for breathing through their nose is similar
to the perceptual capture effect demonstrated in
        <xref ref-type="bibr" rid="ref21">Sams et al.
(2005)</xref>
        or
        <xref ref-type="bibr" rid="ref24">Scott et al. (2013)</xref>
        but, unlike those experiments, in
the current experiment the articulators are not being used in a
speech task by the perceiver.
      </p>
      <p>
        This experiment is also similar to that of
        <xref ref-type="bibr" rid="ref10">Ito et al. (2009)</xref>
        , in
which they showed that dynamic deformation of a perceiver’s
face (by a robot device) can alter perception in line with the
movement, but only if the movement is timed appropriately
with the percept. In contrast, the current experiment asks
whether a static articulator position can also induce a shift
in perception.
      </p>
    </sec>
    <sec id="sec-3">
      <title>Methods</title>
      <p>Participants were asked to breath through their mouth or nose
while they categorized sounds as /ada/ or /ana/. The
prediction is that when participants are breathing through their
nose, the necessarily lowered velum will influence their
perception so that they hear the sounds as more similar to the
nasal /ana/. In a control experiment, participants categorized
/ada/ vs. /aga/ while breathing through their mouth or nose.
No difference was predicted for this control experiment.</p>
      <sec id="sec-3-1">
        <title>Stimuli</title>
        <p>
          A female native speaker of standard European French was
recorded saying /ada/ and /ana/. A 10 000 step
continuum between these sounds was created using STRAIGHT
          <xref ref-type="bibr" rid="ref12 ref13">(Kawahara et al., 2008; Kawahara, Irino, &amp; Morise, 2011)</xref>
          .
While this may seem like a large number of continuum steps,
it should be kept in mind that participants only heard a small
subset of these sounds and the large number of steps is
simple to generate and allows for very fine-grained precision in
estimating phoneme boundaries.
        </p>
      </sec>
      <sec id="sec-3-2">
        <title>Procedures</title>
        <p>There were two conditions:</p>
      </sec>
      <sec id="sec-3-3">
        <title>1. Breathing through the mouth (velum up)</title>
      </sec>
      <sec id="sec-3-4">
        <title>2. Breathing through the nose (velum down)</title>
        <p>
          In each condition, stimuli were presented using the
staircase method
          <xref ref-type="bibr" rid="ref2">(Cornsweet, 1962)</xref>
          . The staircase method
presents points along a continuum and shifts the subsequent
target for presentation based on previous responses. Thus it is
able to ‘search out’ the perceptual boundary between sounds
quite quickly. Two interleaved staircases with random
switching (to prevent participants being able to predict the
upcoming sound) were used for each condition, and participants
alternated back and forth between conditions so that each
participant performed the interleaved staircase procedure for
each condition twice (a total of four staircases per condition).
Thus the experiment determined each participant’s perceptual
boundary between /ada/ and /ana/ while participants breathed
through their mouth or nose.
        </p>
        <p>Each staircase consisted of thirteen reversals with
decreasing stepsize after each reversal. The step sizes were: 1250,
1000, 800, 650, 500, 350, 250, 150, 80, 50, 30, 20, 10 (from
a 10 000 step continuum). The two interleaved staircases
started at points 2400 and 7600 of the continuum.</p>
        <p>An abstract example of the layout of a staircase procedure
is shown in Figure 1. A sound from one end of the continuum
(for example the /ada/ end of the continuum) is played to the
participant. If the participant categorizes it as belonging to
the category consistent with that end of the continuum (/ada/
in our example), then the computer selects the next sound to
be closer to the other (/ana/) end of the continuum. If the
participant hears this also as /ada/, the computer chooses the next
sound to be even closer to /ana/ and so on until the participant
reports hearing /ana/. At this point, the computer reverses
direction of sound selection (hence this is called a ‘reversal’)
and selects the next sound to be closer to the /ada/ end of
the continuum, however the stepsize along the continuum is
made smaller (the computer makes smaller jumps along the
continuum between sound selections, so that there is more
precision). When the participant starts to hear the sound as
/ada/ again, the computer reverses again (and again makes
the stepsizes along the continuum smaller) moving back
toward /ana/. This back and forth continues as the computer
homes in on the participant’s boundary between /ada/ and
/ana/, changing direction of movement along the continuum
and getting more precise with each reversal. This is a robust
and relatively quick method of estimating a person’s
perceptual boundary between two sounds.
The structure of each trial was very simple, participants
were presented with an audio stimulus that was somewhat
ambiguous between /ada/ or /ana/ and they pressed (with their
right hand) a keyboard button (right or left arrow key) to
indicate their perception of the sound as /ada/ or /ana/.</p>
        <p>Participants were given instructions on the task and
familiarized with the software before conducting the experiment.
The experiment itself took about 20 minutes for each
participant to complete.</p>
        <p>The order of breathing conditions (half the participants
starting with breathing through the mouth half starting with
breathing through the nose) was counterbalanced across
participants as was the correspondence of response button (left
arrow vs. right arrow on the computer keyboard) to sound.</p>
        <p>
          The experiment was run on the PsychoPy
experimentplatform
          <xref ref-type="bibr" rid="ref17 ref18">(Peirce, 2007, 2009)</xref>
          .
        </p>
      </sec>
      <sec id="sec-3-5">
        <title>Participants</title>
        <p>Thirty-nine native French speaking participants (31 female,
35 right-handed) were run at Université Paris Descartes
(average age 22.28, standard deviation 2.37).</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>Results</title>
      <p>For each breathing position, each participant’s data from the
four staircases was submitted to a logistic regression to
determine what point on the continuum corresponded to their
perceptual boundary between /ada/ and /ana/ (the point at which
they would hear the sound equally often as /ada/ and /ana/).
These calculated boundaries were the dependent measure for
the experiment. As this was a within-subjects design, the
individual variability in perceptual boundaries (which are often
highly variable between individuals) is not an issue here –
each participant served as his/her own control.</p>
      <p>As predicted, participants heard significantly more /ana/
when breathing through their nose than when breathing
through their mouth, as determined by a paired t-test [t(38)
= 2.08, p = .044, d = .33].</p>
      <p>In the control experiment (categorizing /aga/ vs. /ada/),
no such difference was found between breathing conditions
[t(38) = 1.36, p = .18, d = .21]. However, the interaction
between experiment and control versions did not reach
significance. We believe this is due to a lack of power and are
creating a new version of the experiment to address this issue.</p>
    </sec>
    <sec id="sec-5">
      <title>Discussion &amp; Conclusion</title>
      <p>Components of our motor systems, known as forward models,
constantly predict the sensory effects of our actions — this
prediction is corollary discharge. Corollary discharge serves
a variety of crucial roles, including providing feedback for
actions performed too quickly to use ‘regular’ sensory feedback
and tagging self-produced sensations as such, thus preventing
sensory confusion. It is this second role that allows corollary
discharge to influence the concurrent perception of external
sensations.</p>
      <p>
        A recent group of theories
        <xref ref-type="bibr" rid="ref22 ref25 ref26">(Skipper et al., 2006, 2007;
Schwartz et al., 2010)</xref>
        have suggested that this function of
corollary discharge may regularly be used to supplement
perception in cases of perceptual uncertainty; generating a
prediction, on the basis of one’s own motor system, to guide
the sensory processing. Forward models necessarily consult
the current position of a person’s articulators when
generating corollary discharge, as sensory consequences are strongly
dependent on the starting point of the effectors. This leads
to the prediction that the position of one’s own articulators
should influence the perception of external speech sounds
when those sounds are ambiguous and thus draw on the motor
system’s prediction abilities.
      </p>
      <p>This experiment tested that prediction and has shown that
the position of one’s own articulators does influence the
perception of the speech — even when the position of the
articulators is adopted for a non-speech activity (breathing). These
results support theories which argue for a role of the motor
system (and corollary discharge) in speech perception and
makes a unique contribution in showing that the static
position of the articulators can have this effect even when their
position is not intended to produce speech.</p>
      <p>
        These results are relevant to the ongoing debate about
embodied cognition — the degree to which the body and motor
control systems are used in cognition. In the realm of
semantic processing of language a similar debate is ongoing
about the degree of motor involvement in the processing of
the meaning of sentences. For example, the Action Sentence
Compatibility effect demonstrates that movements of the arm
are faster when a person reads a sentence implying arm
movements, suggesting that the person’s motor plan for arm
movements was triggered by reading the sentence
        <xref ref-type="bibr" rid="ref5">(Glenberg &amp;
Kaschak, 2002)</xref>
        . The current experiment demonstrates a
related example of embodied cognition, but at a ‘lower’,
perceptual level of language processing.
      </p>
      <p>Ongoing research is currently exploring the extent of
this effect, examining how widespread (in terms of speech
sounds) such effects are.</p>
    </sec>
    <sec id="sec-6">
      <title>Acknowledgements</title>
      <p>Funding for this project was provided by a United Arab
Emirates University Research Start-Up Grant (Perceptual-Motor
Linkages in Speech and Cognition - 31h060) to Mark Scott.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          <string-name>
            <surname>Aliu</surname>
            ,
            <given-names>S. O.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Houde</surname>
            ,
            <given-names>J. F.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>Nagarajan</surname>
            ,
            <given-names>S. S.</given-names>
          </string-name>
          (
          <year>2009</year>
          , April).
          <article-title>Motor-induced suppression of the auditory cortex</article-title>
          .
          <source>Journal of Cognitive Neuroscience</source>
          ,
          <volume>21</volume>
          (
          <issue>4</issue>
          ),
          <fpage>791</fpage>
          -
          <lpage>802</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          <string-name>
            <surname>Cornsweet</surname>
            ,
            <given-names>T. N.</given-names>
          </string-name>
          (
          <year>1962</year>
          ,
          <article-title>September)</article-title>
          .
          <source>The Staircase-Method in Psychophysics. The American Journal of Psychology</source>
          ,
          <volume>75</volume>
          (
          <issue>3</issue>
          ),
          <fpage>485</fpage>
          -
          <lpage>491</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          <string-name>
            <surname>Eliades</surname>
            ,
            <given-names>S. J.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>Wang</surname>
            ,
            <given-names>X.</given-names>
          </string-name>
          (
          <year>2008</year>
          ).
          <article-title>Neural substrates of vocalization feedback monitoring in primate auditory cortex</article-title>
          .
          <source>Nature</source>
          ,
          <volume>453</volume>
          (
          <issue>7198</issue>
          ),
          <fpage>1102</fpage>
          -
          <lpage>1106</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          <string-name>
            <surname>Fowler</surname>
            ,
            <given-names>C. A.</given-names>
          </string-name>
          (
          <year>1986</year>
          ).
          <article-title>An event approach to the study of speech perception from a direct-realist perspective</article-title>
          .
          <source>Journal of Phonetics</source>
          ,
          <volume>14</volume>
          (
          <issue>1</issue>
          ),
          <fpage>3</fpage>
          -
          <lpage>28</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          <string-name>
            <surname>Glenberg</surname>
            ,
            <given-names>A. M.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>Kaschak</surname>
            ,
            <given-names>M. P.</given-names>
          </string-name>
          (
          <year>2002</year>
          ).
          <article-title>Grounding language in action</article-title>
          .
          <source>Psychonomic Bulletin &amp; Review</source>
          ,
          <volume>9</volume>
          (
          <issue>3</issue>
          ),
          <fpage>558</fpage>
          -
          <lpage>565</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          <string-name>
            <surname>Goldstein</surname>
            ,
            <given-names>L. M.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>Fowler</surname>
            ,
            <given-names>C. A.</given-names>
          </string-name>
          (
          <year>2003</year>
          ).
          <article-title>Articulatory phonology: A phonology for public language use. In Phonetics and phonology in language comprehension and production: Differences and similarities</article-title>
          (pp.
          <fpage>159</fpage>
          -
          <lpage>207</lpage>
          ). Berlin: Mouton de Gruyter.
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          <string-name>
            <surname>Hickok</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          (
          <year>2012</year>
          ).
          <article-title>Computational neuroanatomy of speech production</article-title>
          .
          <source>Nature Reviews Neuroscience</source>
          ,
          <volume>13</volume>
          (
          <issue>2</issue>
          ),
          <fpage>135</fpage>
          -
          <lpage>145</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          <string-name>
            <surname>Hickok</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Houde</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>Rong</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          (
          <year>2011</year>
          ,
          <article-title>February)</article-title>
          .
          <article-title>Sensorimotor integration in speech processing: computational basis and neural organization</article-title>
          .
          <source>Neuron</source>
          ,
          <volume>69</volume>
          (
          <issue>3</issue>
          ),
          <fpage>407</fpage>
          -
          <lpage>22</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          <string-name>
            <surname>Houde</surname>
            ,
            <given-names>J. F.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>Nagarajan</surname>
            ,
            <given-names>S. S.</given-names>
          </string-name>
          (
          <year>2011</year>
          ).
          <article-title>Speech Production as State Feedback Control</article-title>
          .
          <source>Frontiers in Human Neuroscience</source>
          ,
          <volume>5</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          <string-name>
            <surname>Ito</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Tiede</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>Ostry</surname>
            ,
            <given-names>D. J.</given-names>
          </string-name>
          (
          <year>2009</year>
          , January).
          <article-title>Somatosensory function in speech perception</article-title>
          .
          <source>Proceedings of the National Academy of Sciences of the United States of America</source>
          ,
          <volume>106</volume>
          (
          <issue>4</issue>
          ),
          <fpage>1245</fpage>
          -
          <lpage>8</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          <string-name>
            <surname>Jones</surname>
            ,
            <given-names>J. A.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>Munhall</surname>
            ,
            <given-names>K. G.</given-names>
          </string-name>
          (
          <year>2002</year>
          ).
          <article-title>The role of auditory feedback during phonation: studies of Mandarin tone production</article-title>
          (Vol.
          <volume>30</volume>
          )
          <article-title>(No. 3).</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          <string-name>
            <surname>Kawahara</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Irino</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>Morise</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          (
          <year>2011</year>
          ).
          <article-title>An interference-free representation of instantaneous frequency of periodic signals and its application to F0 extraction</article-title>
          .
          <source>In 2011 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)</source>
          (pp.
          <fpage>5420</fpage>
          -
          <lpage>5423</lpage>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          <string-name>
            <surname>Kawahara</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Morise</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Takahashi</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Nisimura</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Irino</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>Banno</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          (
          <year>2008</year>
          ).
          <article-title>Tandem-STRAIGHT: A temporally stable power spectral representation for periodic signals and applications to interference-free spectrum, F0, and aperiodicity estimation</article-title>
          .
          <source>In IEEE International Conference on Acoustics, Speech and Signal Processing</source>
          ,
          <year>2008</year>
          . ICASSP 2008 (pp.
          <fpage>3933</fpage>
          -
          <lpage>3936</lpage>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          <string-name>
            <surname>Liberman</surname>
            ,
            <given-names>A. M.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>Mattingly</surname>
            ,
            <given-names>I. G.</given-names>
          </string-name>
          (
          <year>1985</year>
          ).
          <article-title>The motor theory of speech perception revised</article-title>
          .
          <source>Cognition</source>
          ,
          <volume>21</volume>
          (
          <issue>1</issue>
          ),
          <fpage>1</fpage>
          -
          <lpage>36</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          <string-name>
            <surname>McGurk</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>MacDonald</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          (
          <year>1976</year>
          ).
          <article-title>Hearing lips and seeing voices</article-title>
          .
          <source>Nature</source>
          ,
          <volume>264</volume>
          ,
          <fpage>746</fpage>
          -
          <lpage>748</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          <string-name>
            <surname>Merckelbach</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>Ven</surname>
            ,
            <given-names>V. V. D.</given-names>
          </string-name>
          (
          <year>2001</year>
          ).
          <article-title>Another White Christmas: fantasy proneness and reports of 'hallucinatory experiences' in undergraduate students</article-title>
          .
          <source>Journal of Behavior Therapy and Experimental Psychiatry</source>
          ,
          <volume>32</volume>
          ,
          <fpage>137</fpage>
          -
          <lpage>144</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          <string-name>
            <surname>Peirce</surname>
            ,
            <given-names>J. W.</given-names>
          </string-name>
          (
          <year>2007</year>
          , May).
          <source>PsychoPy-Psychophysics software in Python. Journal of Neuroscience Methods</source>
          ,
          <volume>162</volume>
          (
          <issue>1-2</issue>
          ),
          <fpage>8</fpage>
          -
          <lpage>13</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          <string-name>
            <surname>Peirce</surname>
            ,
            <given-names>J. W.</given-names>
          </string-name>
          (
          <year>2009</year>
          ).
          <article-title>Generating stimuli for neuroscience using PsychoPy</article-title>
          . Frontiers in Neuroinformatics,
          <volume>2</volume>
          ,
          <fpage>10</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          <string-name>
            <surname>Pickering</surname>
            ,
            <given-names>M. J.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>Garrod</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          (
          <year>2007</year>
          , March).
          <article-title>Do people use language production to make predictions during comprehension? Trends in Cognitive Sciences</article-title>
          ,
          <volume>11</volume>
          (
          <issue>3</issue>
          ),
          <fpage>105</fpage>
          -
          <lpage>10</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          <string-name>
            <surname>Repp</surname>
            ,
            <given-names>B. H.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>Knoblich</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          (
          <year>2009</year>
          ).
          <article-title>Performed or observed keyboard actions affect pianists' judgements of relative pitch</article-title>
          .
          <source>The Quarterly Journal of Experimental Psychology: Human Experimental Psychology</source>
          ,
          <volume>62</volume>
          (
          <issue>11</issue>
          ),
          <fpage>2156</fpage>
          -
          <lpage>2170</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          <string-name>
            <surname>Sams</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Möttönen</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>Sihvonen</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          (
          <year>2005</year>
          ).
          <article-title>Seeing and hearing others and oneself talk</article-title>
          .
          <source>Cognitive Brain Research</source>
          ,
          <volume>23</volume>
          (
          <issue>2-3</issue>
          ),
          <fpage>429</fpage>
          -
          <lpage>435</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          <string-name>
            <surname>Schwartz</surname>
            ,
            <given-names>J.-L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Basirat</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ménard</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>Sato</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          (
          <year>2010</year>
          ,
          <article-title>January). The Perception-for-Action-Control Theory (PACT): A perceptuo-motor theory of speech perception</article-title>
          .
          <source>Journal of Neurolinguistics</source>
          ,
          <fpage>1</fpage>
          -
          <lpage>19</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          <string-name>
            <surname>Schütz-Bosbach</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>Prinz</surname>
            ,
            <given-names>W.</given-names>
          </string-name>
          (
          <year>2007</year>
          ).
          <article-title>Perceptual resonance: action-induced modulation of perception</article-title>
          .
          <source>Trends in Cognitive Sciences</source>
          ,
          <volume>11</volume>
          (
          <issue>8</issue>
          ),
          <fpage>349</fpage>
          -
          <lpage>355</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          <string-name>
            <surname>Scott</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Yeung</surname>
            ,
            <given-names>H. H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gick</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>Werker</surname>
            ,
            <given-names>J. F.</given-names>
          </string-name>
          (
          <year>2013</year>
          ).
          <article-title>Inner speech captures the perception of external speech</article-title>
          .
          <source>Journal of the Acoustical Society of America Express Letters</source>
          ,
          <volume>133</volume>
          (
          <issue>4</issue>
          ),
          <fpage>286</fpage>
          -
          <lpage>293</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref25">
        <mixed-citation>
          <string-name>
            <surname>Skipper</surname>
            ,
            <given-names>J. I.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Nusbaum</surname>
            ,
            <given-names>H. C.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>Small</surname>
            ,
            <given-names>S. L.</given-names>
          </string-name>
          (
          <year>2006</year>
          ).
          <article-title>Lending a helping hand to hearing: another motor theory of speech perception</article-title>
          . In M. A.
          <string-name>
            <surname>Arbib</surname>
          </string-name>
          (Ed.),
          <article-title>Action to Language via the Mirror Neuron System</article-title>
          (pp.
          <fpage>250</fpage>
          -
          <lpage>285</lpage>
          ). Cambridge: Cambridge University Press.
        </mixed-citation>
      </ref>
      <ref id="ref26">
        <mixed-citation>
          <string-name>
            <surname>Skipper</surname>
            ,
            <given-names>J. I.</given-names>
          </string-name>
          , van
          <string-name>
            <surname>Wassenhove</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Nusbaum</surname>
            ,
            <given-names>H. C.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>Small</surname>
            ,
            <given-names>S. L.</given-names>
          </string-name>
          (
          <year>2007</year>
          ,
          <article-title>October)</article-title>
          .
          <article-title>Hearing lips and seeing voices: how cortical areas supporting speech production mediate audiovisual speech perception</article-title>
          .
          <source>Cerebral</source>
          cortex (New York, N.Y. :
          <year>1991</year>
          ),
          <volume>17</volume>
          (
          <issue>10</issue>
          ),
          <fpage>2387</fpage>
          -
          <lpage>99</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref27">
        <mixed-citation>
          <string-name>
            <surname>Wolpert</surname>
            ,
            <given-names>D. M.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>Flanagan</surname>
            ,
            <given-names>J. R.</given-names>
          </string-name>
          (
          <year>2001</year>
          ).
          <article-title>Motor prediction</article-title>
          .
          <source>Current Biology</source>
          ,
          <volume>11</volume>
          (
          <issue>18</issue>
          ),
          <fpage>729</fpage>
          -
          <lpage>732</lpage>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>