<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Perception Deception: Audio-Visual Mismatch in Virtual Reality Using the McGurk E ect</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>AbuBakr Siddig</string-name>
          <email>abubakr.siddig@ucd.ie</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Pheobe Wenyi Sun</string-name>
          <email>wenyi.sun@ucdconnect.ie</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Matthew Parker</string-name>
          <email>matthew.parker@ttu.edu</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Andrew Hines</string-name>
          <email>andrew.hines@ucd.ie</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>School of Computer Science, University College Dublin</institution>
          ,
          <country country="IE">Ireland</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Texas Tech University</institution>
          ,
          <country country="US">USA</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>The audio-visual synchronisation is a big challenge in the Virtual Reality (VR) industry. Studies investigating the e ect of incongruent multisensory stimuli will have a direct impact on the design of immersive experience. In this paper, we explored the e ect of audio-visual mismatch on the sensory integration in a VR context. Inspired by the McGurk e ect, we designed an experiment addressing a few critical VR content production concerns today including sound spatialisation and unisensory signal quality. The results con rm previous studies using 2D videos where audio spatial separation has no signi cant impact on the McGurk e ect, yet the ndings raise new thoughts regarding the future compression and multisensory signal design strategies to optimise the perceptual immersion in the 3D context.</p>
      </abstract>
      <kwd-group>
        <kwd>Ambisonics</kwd>
        <kwd>McGurk e ect</kwd>
        <kwd>Virtual Reality</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>
        A virtual reality (VR) system places a participant in a computer-generated 3D
environment that mimics or expands the physical world. Creating a better
immersive experience has been an ongoing goal in VR research. Humans immersive
experience is a result of multisensory integration. VR researchers are developing
technologies that deliver accurate sensory cues providing increasingly natural
sensorimotor contingencies [
        <xref ref-type="bibr" rid="ref17">17</xref>
        ] increasing the plausibility of virtual events. The
sensory cues in state of the art VR systems (e.g. Oculus, Vive) primarily focus
on visual and auditory immersion with some haptic feedback in hand controllers.
Although the technology to create immersive visual content has become more
powerful, accessible, and a ordable, attention on auditory content creation and
the audio-visual synchronisation has lagged which impacts the overall fused
immersive experience [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ].
      </p>
      <p>
        To ensure the immersive feeling, audio and vision should match. If the sounds
and audio cues are not plausibly aligned with the associated visual experience,
the resulting incongruity causes the virtual immersion to collapse [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ]. Multi
signal mismatch is still a common problem due to issues related to latency,
head-tracking technology, headphone, and immersive media content production.
Despite the potential of breaking the immersive feeling due to unavoidable
multisensory disparity, an interesting phenomenon saying that humans can
perceive a uni ed precept in the event of incongruent stimuli [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ] rst presented by
psychologists McGurk and McDonald in 1976 inspires new concerns of
audiovisual design for VR. This phenomenon proposes that human minds can join
together mismatched sensory information to form an acceptable conclusion about
an experience, and is studied using a classical research tool called the McGurk
e ect experiment. The McGurk e ect highlights the importance of knowing the
qualities of the unisensory stimuli (i.e., the clarity of the visual components, and
the resolution of auditory components) and the coherence between the input
sensory information before analysing the overall integrated experience [
        <xref ref-type="bibr" rid="ref19">19</xref>
        ]. The
immersive experience in VR relies on multisensory integration. In this paper, we
focus on the spatial qualities of audio-visual stimuli, speci cally the relationship
between where sounds are perceived to be coming from relative to their
corresponding visual source. A better understanding of the importance of spatial
audio localisation and image resolution on the immersive experience can guide
the development and application of compression from a quality of experience
perspective.
      </p>
      <p>In this paper, we explore the interactions between disparities between the
auditory and visual signals in a VR context. We experiment on speech perception
inspired by the McGurk e ect and we investigate how di erent factors (audio
directionality and quality of visuals) a ect the strength of audiovisual integration.
The results are anticipated to guide to the candidacy of audio localisation for
data compression in VR as a means to limit the bandwidth used by streaming
media in a virtual environment.
2</p>
    </sec>
    <sec id="sec-2">
      <title>Background</title>
      <p>2.1</p>
      <sec id="sec-2-1">
        <title>Immersive Experience in VR</title>
        <p>
          Immersive experience is a result of multisensory processing. Sensory inputs for
immersive experience commonly include vision, audio, touch and force feedback
and less often smell and taste. In the context of VR, the goal is to simulate as
many human natural sensory inputs as possible with the help of computer-based
technology. The critical tools developed to reach this goal include wide
eldof-view vision, stereo, head tracking, low-latency from head move to display,
and high-resolution displays [
          <xref ref-type="bibr" rid="ref17">17</xref>
          ]. However, due to the multisensory property of
the immersive experience, a single improvement in one aspect is not su cient
to improve the overall immersive experience in VR. When a virtually generated
multisensory illusion gives an impression of a high plausibility, it is thus believed
that the designed scene is actually happening [
          <xref ref-type="bibr" rid="ref17">17</xref>
          ]. Any mismatch between
different senses would make a scene less convincing and resulting in the immersive
sense of presence in a virtual scene collapse [
          <xref ref-type="bibr" rid="ref13">13</xref>
          ].
        </p>
      </sec>
      <sec id="sec-2-2">
        <title>Spatial Audio</title>
        <p>
          In practice, designers have been increasingly adopting 3D sound techniques to
enrich the scenes contained in the auditory information to overcome the
shortcomings of the traditionally-used stereo recordings. However, quality 3D audio is
technically challenging to deliver for a variety of reasons. These include
compromises in current content capture, production and delivery. For example, many
existing a ordable 360 cameras still use mono or stereo as their audio signals [
          <xref ref-type="bibr" rid="ref14">14</xref>
          ].
Without extra e ort spent in capturing and rendering the spatial audio, the
camera-recorded audio-visual content is incapable of creating a truly immersive
sense of being there in the VR environment [
          <xref ref-type="bibr" rid="ref14">14</xref>
          ]. Dynamically updating both
the visual and the auditory signal based on head-position and orientation pose
further challenges to both network and compression technologies used to deliver
immersive VR.
2.3
        </p>
      </sec>
      <sec id="sec-2-3">
        <title>Spatial Separation in VR</title>
        <p>
          The disparity between modalities can occur as a result of a timing mismatch
between network and tracking technology. It can also be due to a mismatch
of localisation information as a result of audio production and rendering
technology. A detailed discussion of hardware and network factors that can result
in disparities is beyond the scope of this paper. This paper will focus on the
perception of spatial separation. The rami cations of audio-visual spatial
separation are still not de nite [
          <xref ref-type="bibr" rid="ref16">16</xref>
          ]. A misalignment of auditory and visual signals
can negatively in uence VR immersive experience [
          <xref ref-type="bibr" rid="ref13">13</xref>
          ]. However, psychologists
highlight the human mind's capacity to form illusions given incongruent stimuli.
For example, the `ventriloquism e ect` says given a visual e ect, the location of
the auditory sources can be overlooked, therefore forming an illusion as if the
sound source is coming from the same place as the visual [
          <xref ref-type="bibr" rid="ref18">18</xref>
          ]. A similar idea
was also re ected in the `unity assumption` e ect [
          <xref ref-type="bibr" rid="ref2">2</xref>
          ]. These theories drive the
motivation to investigate whether spatial separation is perceptually detrimental
in VR. The extent that spatial separation is tolerated will be important to both
VR content and technology developers.
2.4
        </p>
      </sec>
      <sec id="sec-2-4">
        <title>The McGurk E ect</title>
        <p>
          To investigate whether a fusion e ect exists in the event of audio-visual spatial
separation in a VR context, this study builds upon a classical research tool
called the McGurk e ect that explores a phenomenon of an altered perception of
auditory speech signal given incongruent audiovisual pairing [
          <xref ref-type="bibr" rid="ref11">11</xref>
          ] [
          <xref ref-type="bibr" rid="ref4">4</xref>
          ]. It is chosen
due to the strength of the McGurk e ect to `re ect the strength of audiovisual
integration' [
          <xref ref-type="bibr" rid="ref19">19</xref>
          ]. In a classical McGurk e ect experiment, a participant is usually
presented with an auditory and a visual signal simultaneously where each signal
carries di erent speech tokens (i.e., audio of /ba/ together with a vision of /ga/).
The participant's reported subjective auditory precept is expected to deviate
from what is actually presented acoustically as a result of unconscious integration
of phonetic information (i.e., hearing /da/ when being presented with audio /ba/
and visual /ga/). Such categorical change of speech perception is described as
the McGurk e ect [
          <xref ref-type="bibr" rid="ref19">19</xref>
          ].
        </p>
        <p>
          Since the initial publication of this lab controlled illusion experiment in 1976,
the McGurk e ect has attracted a lot of attention in the eld of cognitive science.
Psychologists believe the `laws of common fate' and `spatial proximity' are the
two underlying theories for this phenomenon. These theories are based on the
fundamental principles of perceptual information fusion [
          <xref ref-type="bibr" rid="ref4">4</xref>
          ]. Further down the
line, researchers have been actively investigating factors that contribute to this
phenomenon. Aside from the interpersonal di erence and age factors [
          <xref ref-type="bibr" rid="ref11">11</xref>
          ][
          <xref ref-type="bibr" rid="ref9">9</xref>
          ],
those factors describing stimuli including visual degradation [
          <xref ref-type="bibr" rid="ref7">7</xref>
          ], talker voice[
          <xref ref-type="bibr" rid="ref8">8</xref>
          ],
choice of utterance [
          <xref ref-type="bibr" rid="ref6">6</xref>
          ], and time lag in synchronisation [
          <xref ref-type="bibr" rid="ref10">10</xref>
          ][
          <xref ref-type="bibr" rid="ref12">12</xref>
          ] all have been
proven to have an impact to the strength of McGurk e ect. All these in uencing
factors lead to a further abstraction of the conditions for the McGurk e ect to
occur: the quality of the unisensory signal and the coherence of the multiple
sensory signals [
          <xref ref-type="bibr" rid="ref19">19</xref>
          ].
2.5
        </p>
      </sec>
      <sec id="sec-2-5">
        <title>The McGurk E ect in VR</title>
        <p>The spatial separation of audio and the visual information, as discussed in 2.3,
is one of the most prominent concerns in the VR content creation process. Based
on the current ndings of the conditions for the occurrence of the McGurk e ect,
factors including the quality of sound localisation information, the quality of the
visual information, and how far apart the separation between the auditory and
the visual scenes are all critical concerns to provide users with an immersive
experience. Ideally, the sound localisation information should be designed in
the utmost detailed and accurate manner to match the visual scene so that the
combined e ect can generate a higher plausibility (Psi) to trick a participant into
believing a virtual scene is real. Meanwhile, the `ventriloquism e ect`, familiar
to children where a puppet appears to speak, implies some tolerance of the
coherence of the multisensory signals for the fusion e ect to take place.</p>
        <p>
          The existing literature, however, presents inconsistent conclusions regarding
the rami cations on the integrated speech perception in the event of audio-visual
disparity. There was no obvious impact found on the audiovisual integration
when the spatial separation was up to 37.5 degrees [
          <xref ref-type="bibr" rid="ref1">1</xref>
          ]. A weak separation e ect
was found when the separation angle reached 60 [
          <xref ref-type="bibr" rid="ref15">15</xref>
          ]. Jones and Munhall found
little di erence in the impact on the McGurk e ect when increasing the spatial
separation of the auditory and the visual scenes [
          <xref ref-type="bibr" rid="ref5">5</xref>
          ] [
          <xref ref-type="bibr" rid="ref4">4</xref>
          ]. These results were also
obtained by Siddig et al. [
          <xref ref-type="bibr" rid="ref16">16</xref>
          ] using 3D audio virtually binaurally rendered spatial
audio for up to 90 . Jones and Jarick also concluded that the e ect of the spatial
separation was only signi cant when the sound was 180 away from the visual [
          <xref ref-type="bibr" rid="ref4">4</xref>
          ].
All of the above experiments were conducted based on 2D visuals only, although
spatial sound was used to generate the auditory signals. We are interested in
investigating whether the occurrence of the fusion phenomenon would change
in a 3D scenario where participants are placed in a surrounding environment
stimulated by both 3D audio and 3D visuals. Because in a 3D environment a
participant has an expectation of the source of the audio according to the visual
stimuli, and the mind in this scenario is more likely to fuse multisensory signals
to trick them into believing a scene is real. Therefore we postulate that a change
of visual scenario from 2D to 3D could lead to a change in the integration e ect.
Following the current reasoning of the McGurk e ect discussed in Section 2.4,
it is reasonable to hypothesise that the strength of the McGurk e ect in a VR
context are a result of intelligibility of speech, resolution of the visual signals,
and di erent degrees of spatial disparities between the sources of the auditory
and visual signals.
        </p>
        <p>
          If the psychological theories regarding perceptual phenomena hold true in
a VR context, we will tolerate an amount of audio-spatial separation resulting
from streaming and digital content processing as it will not result in a negative
integrated experience. Building on the experiment described in [
          <xref ref-type="bibr" rid="ref16">16</xref>
          ] we aim to
provide insights on multisensory data compression factors to consider for VR
streaming applications.
3
3.1
        </p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>Methodology</title>
      <sec id="sec-3-1">
        <title>Experiment Design</title>
        <p>The study was carried out in a quiet meeting room with a meeting room setup.
Participants were asked to sit in a chair wearing a VR headset (HTC Oculus
Quest) and stereo closed-back headphones (Audio-Technica ATH-M70x).
Multiple virtual scenes of a speaker sitting across them were presented to the
participants in sequence. The participants were asked to report what they heard after
each playback. The experiment consisted of two tests: the audio-only utterance
discrimination test and the audio-visual utterance discrimination test.</p>
        <p>The collection of the perceived utterance takes the form of self-reported
measures. A VR controller was used for participants to choose what they perceived
given 8 possible confusing options. Namely /ba/, /ga/, /da/, /ma/, /ka/, /pa/,
/va/, and /ta/.
3.2</p>
      </sec>
      <sec id="sec-3-2">
        <title>Control</title>
        <p>To detect the genuine McGurk e ect in the experiment, we have to ensure the
quality of the test unisensory stimuli controlled before we experiment. From the
literature, the choice of utterance, the talker voice, the intelligibility of
utterance, and the synchronisation of multisensory signals are factors contributing
to the McGurk e ect. We adopted what was used in the classical McGurk
effect experiment, the simplest consonant-vowel combination (/ga/, /ka/, /ma/,
and /pa/) as test utterances, and used di erent techniques to keep the levels
of intelligibility of each enunciation of test utterances consistent. We selected
talkers' voices that generate the highest intelligibility result among the multiple
recorded talkers. In addition, the resolutions of each video clip capturing the lip
movements are controlled. To set the same expectation in all participants that
they are placed in an immerse environment, all participants were asked to be
seated, and the VR scene was preloaded before the headset was given to the
participants. The same headset, headphones, and controllers were used for all
participants.</p>
        <p>The primary factor of interest in this experiment is the degree of spatial
separation between the auditory and the visual signals. Therefore we xed the
position of the video to the front of the participants and rendered auditory
signals with multiple directions of arrival (DOA). To further investigate the
strength of the McGurk e ect, two di erent qualities of the visual presentations
(camera placed at 0.8 m and 2.3 m away from the talker) were used to test the
robustness of the nding.
3.3</p>
      </sec>
      <sec id="sec-3-3">
        <title>Stimuli</title>
        <p>
          Audio and visual recordings The goal of designing a McGurk e ect is to
have 'incongruent' stimuli where 'each modality would present di erent speech
tokens' [
          <xref ref-type="bibr" rid="ref6">6</xref>
          ]. We have recorded two male talkers uttering /ba/, /ma/, /pa/, /ga/,
/ka/ and captured 360 videos of the same talkers mouthing /ga/ and /ka/.
During the experiment, di erent combinations of the stand-alone visual and
audio les were used to generate di erent pairs of audio-visual presentations. In
addition to a complex of possible confusing combinations of visual and audio
presentation, the e ect of distance was also taken into account using the near
and far scenes respectively. A close-up scene was taken at a distance of 2 metres,
and the far scene was taken from 4 metres away.
Virtual environment creation The 360 visual stimuli featuring a talker's lip
movements were produced using the Insta 360 One X camera. The spatial e ect
of the recorded auditory stimuli was created using Google ambisonics API. The
virtual scenes of a talker uttering di erent syllables were dubbed with various
con icting auditory utterances in HitFilm. Unity was used to develop and build
the VR experiment on Oculus Quest. To present the participants with a higher
standard of immersive experience, the experiment was designed to be carried
out using over-the-ear closed-back headphones in addition to a VR headset.
A total of seventeen participants from University College Dublin took part in this
experiment. All participants spoke English uently, having normal or corrected
to normal vision and reporting no hearing problems.
4
4.1
        </p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>Experiment Analysis</title>
      <sec id="sec-4-1">
        <title>Audio-Only Utterance Discrimination Test</title>
        <p>
          Procedure In the rst test, the stimuli were played to the participants without
visual display to learn participants' ability to discriminate the auditory
utterances without any visual interference. To put the audio discrimination test in a
spatial context, the audio rendered as if it is coming right in front of the
participants. A xed DOA, an azimuth of 0 degree, was used because previous literature
showed that speech intelligibility did not change with azimuth angle [
          <xref ref-type="bibr" rid="ref5">5</xref>
          ][
          <xref ref-type="bibr" rid="ref16">16</xref>
          ].
Table 1 shows a list of the audio stimuli. After listening to each audio utterance,
the participants were then asked to choose what they heard from the 8 options
(namely /ba/, /ga/, /da/, /ma/, /ka/, /pa/, /va/, and /ta/) given on the screen
using the VR controller.
        </p>
        <p>
          Data Processing The purpose of this test is two-fold. It is used to lter out
unreliable responses for the following test. Those who had di culty
discriminating between utterances from the audio will not be able to re ect the hearing
No. Utterance Talker
1 /ba/ 1
2 /ma/ 1
3 /pa/ 1
4 /ba/ 2
5 /ma/ 2
6 /pa/ 2
confusion as a result of the visual interference [
          <xref ref-type="bibr" rid="ref3">3</xref>
          ] in the Audio-Visual Utterance
Discrimination Test. This test is also used to gauge the quality of unisensory
input signal. The accuracy of utterance discrimination based on the unisensory
signal here will be taken into account when analysing the integration e ect in
the later test to gauge the genuinity of the McGurk e ect and the strength of
it [
          <xref ref-type="bibr" rid="ref19">19</xref>
          ].
        </p>
        <p>
          Results Fifteen out of seventeen participants could discriminate most of the
spoken token (/ba/, /ma/, /pa/) from the auditory utterances recorded from two
di erent talkers. Two participants fell below 50% and were excluded from the
second experiment due to consideration of the reliability of the responses [
          <xref ref-type="bibr" rid="ref3">3</xref>
          ] as
discussed above. There was a slight di erence in the rate of successful recognition
between di erent utterances (see Table 2). Some of the languages do not have
both sounds of /b/ and /p/, this is re ected on the lower accuracy rate for these
sounds as seen in Table 2. Authors think this confusion is the main reason for
two participants to fell below 50% in audio-only tests. Further investigation can
be done to cover this issue. This insight was used in evaluating the strength of
the McGurk e ect in the second test.
        </p>
        <p>Utterance Accuracy Rate
/ba/ 73.5%
/ma/ 88.2%
/pa/ 70.6%</p>
      </sec>
      <sec id="sec-4-2">
        <title>Audio-visual Utterance Discrimination Test</title>
        <p>Procedure In the second test, both audio and visual stimuli were presented
to the participants. The audio and video recordings of di erent utterances were
put together in a synchronised manner to test the McGurk e ect. The visual
stimuli were presented at the xed azimuth position, as if the virtual talker were
sitting in front of the participant (see Fig 2). However, the far and near scenes
capturing two di erent levels of details of the talker's lip movements were used
to test factors' impact on the strength of the McGurk e ect. The audio stimuli
were played at di erent azimuth angles to test the e ect of spatial separation on
the occurrence of the McGurk e ect. Every audio-visual test stimuli was played
only once for each participant (see Table 3).</p>
        <p>
          Paring Audio Visual Audio
Type Signal Signal Directionality
1 /ba/ /ga/ azi 0 , 90 , 120 , 180
1 /ba/ /ga/ azi 0 , 90 , 120 , 180
1 /ba/ /ga/ azi 0 , elevation 60
1 /ba/ /ga/ azi 0 , elevation 60
2 /ma/ /ka/ azi 0 , 90 , 120 , 180
2 /ma/ /ka/ azi 0 , 90 , 120 , 180
3 /pa/ /ka/ azi 0 , 90 , 120 , 180
3 /pa/ /ka/ azi 0 , 90 , 120 , 180
Data Processing In this paper, the McGurk e ect is labelled as positive when
the reported percept deviates from the auditory signal. In the case where the
reported percept is a third consonant other than the auditory or the visual signal,
the response is also labelled as a strong McGurk e ect. Because hearing a third
consonant is a result of the fusion e ect where the brain merges the speech tokens
from both the auditory signal and the visual signal [
          <xref ref-type="bibr" rid="ref19">19</xref>
          ].
        </p>
        <p>
          The test stimuli were grouped in various ways in the ANOVA test to evaluate
the e ect of sound directionality on the integration e ect. Speci cally, we were
interested in exploring the front-rear di erence, directionality di erence, and the
elevation di erence in contributing to the McGurk e ect in a VR scenario. the
left-right di erence was not explored because the previous research using 2D
visual signals [
          <xref ref-type="bibr" rid="ref16">16</xref>
          ] concluded there was no signi cant e ect between the left and
right auditory signals.
        </p>
        <p>
          Results In this experiment, the percentage of the reported the McGurk e ect
given front 3D visuals and audio at azimuth 0 was 30%, much lower than what
was reported in the literature for experiments that were carried out with 2D
videos. This result indicates that it is harder for the McGurk e ect to take place
in the event of mismatching multisensory signals in an immersive environment.
Among all the reported occurrence of the McGurk e ect, 97.7% showed signs
of a strong McGurk e ect. Such a high percentage of reported strong McGurk
e ect gave more con dence to say that the reported confused answers were as
a result of a fusion e ect. Jones and Jarick previously concluded the spatial
separation e ect on the McGurk e ect was only signi cant when the sound was
coming from the rear [
          <xref ref-type="bibr" rid="ref4">4</xref>
          ]. However, in this experiment, there showed no signi cant
di erence between the signals generating from the front (0 ) and the rear (180 )
[F (1; 270) = 0, p &gt; 0:5]. Unlike some results drown from the spatial separation
experiment in a 2D environment, the change of DOAs seemed to have little e ect
on the formation of illusion in a 3D environment [F (5; 408) = 1:9, p &gt; 0:1]. No
talker's e ect was found in this experiment [F (1; 408) = 0:72, p &lt; 0:5].
        </p>
        <p>Test Independent Variables p</p>
        <p>Front &amp; rear azi 0 &amp; 180 &gt;0.5
Azimuth angles azi 60 ; 90 ; 120 &gt;0.1</p>
        <p>
          Talkers talker 1 &amp; talker 2 &lt;0.5
Several recent studies did not nd a signi cant di erence in the e ect of sound
directionalities on the audio-visual immersion e ect [
          <xref ref-type="bibr" rid="ref16">16</xref>
          ]. Following previous
researchers ndings, this paper speci cally analysed the di erence between the
sound coming from the front and from the back. However, no di erence was found
between the front and rear sounds in this experiment. This could be accounted
for by poor externalisation from binaural ambisonic rendering but requires more
experimental testing. This paper also controlled the quality of unisensory
input signal, attempting to see how much contribution a single modal signal can
contribute to the strength of the McGurk e ect. The controls used in this
experiment were the resolution of talkers' utterance video and the recognition rate of
the utterance sounds. The result did not show a signi cant statistical di erence
that altered the McGurk e ect.
        </p>
        <p>
          Comparing the occurrence of the McGurk e ect in 3D with the ndings from
the 2D experiments [
          <xref ref-type="bibr" rid="ref7">7</xref>
          ], contrary to our expectation, we realised the conditions
for a McGurk e ect to occur is stricter in the immersive environment. This
nding could be due to the change of expectation. Participants in a VR setting
are by default expecting a higher level of coherence of the multisensory signals to
form a belief that a real event occurs, thus hindering the chances for an illusion
to occur given mismatched audio-visual signals.
6
        </p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>Conclusions and Future Work</title>
      <p>
        The experimental results in the VR environment are in line with the results seen
by [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ] using loudspeakers and with binaurally rendered spatial audio over
headphones [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ]. In an immersive environment with the source behind your head,
no statistical di erence in the McGurk e ect was seen although the
experimental protocol did not capture the participants feedback on sound externalisation
which can be covered in future work. The visual lip cue distance did not show
a signi cant statistical di erence point to an altered McGurk e ect experience.
The future work can further investigate the audio-visual mismatch along the
elevation level to complement the current study on the e ect of spatial separation
of multisensory cues.
      </p>
      <p>Spatial separation is considered an important quality issue for VR
environments. However, these results show that people may not be very sensitive to
separation for multisensory inputs { the fusion phenomenon was not signi
cantly impacted by the changes to separation or visual cue resolution in VR.
Our current nding indicates that, unless the audio and the visual are perfectly
synchronised, a considerable extent of spatial separation between auditory and
visual signals can be tolerated as the level of mismatch does not in uence the
level of immersion. As the 3D content streaming services rise in popularity, given
limited bandwidth in a streaming scenario, a less strict audio localisation
requirement can be adopted without deteriorating the overall immersive experience. As
VR has gained popularity in today's society, this nding might be helpful for
the VR content creators to optimise the usage of bandwidth when streaming 3D
media content in a virtual environment. However, more experiments are needed
to draw a nal conclusion.
7</p>
    </sec>
    <sec id="sec-6">
      <title>Acknowledgements</title>
      <p>This publication has emanated from research supported in part by a research
grant from Science Foundation Ireland (SFI) and is co-funded under the
European Regional Development Fund under Grant Number 13/RC/2289 and Grant
Number SFI/12/RC/2077. The experimental work was supported by the
University College Dublin STEM Summer Research Project. Thanks to Hamed Z.
Jahromi and Alessandro Ragano for assistance with statistical analysis of the
results.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Bertelson</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Vroomen</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Wiegeraad</surname>
            , G., de Gelder,
            <given-names>B.</given-names>
          </string-name>
          :
          <article-title>Exploring the relation between McGurk interference and ventriloquism</article-title>
          .
          <source>In: Third International Conference on Spoken Language Processing</source>
          . pp.
          <volume>559</volume>
          {
          <fpage>562</fpage>
          . ISCA, Yokohama, Japan (
          <year>1994</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Chen</surname>
            ,
            <given-names>Y.C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Spence</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          :
          <article-title>Assessing the role of the unity assumptionon multisensory integration: A review</article-title>
          .
          <source>Frontiers in psychology 8</source>
          ,
          <issue>445</issue>
          (
          <year>2017</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Colin</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Radeau</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Deltenre</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Demolin</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Soquet</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          :
          <article-title>The role of sound intensity and stop-consonant voicing on McGurk fusions and combinations</article-title>
          .
          <source>European Journal of Cognitive Psychology</source>
          <volume>14</volume>
          (
          <issue>4</issue>
          ),
          <volume>475</volume>
          {
          <fpage>491</fpage>
          (
          <year>2002</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Jones</surname>
            ,
            <given-names>J.A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Jarick</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          :
          <article-title>Multisensory integration of speech signals: The relationship between space and time</article-title>
          .
          <source>Experimental Brain Research</source>
          <volume>174</volume>
          (
          <issue>3</issue>
          ),
          <volume>588</volume>
          {
          <fpage>594</fpage>
          (
          <year>2006</year>
          ). https://doi.org/10.1007/s00221-006-0634-0, https://link.springer.com/content/pdf/10.1007%
          <fpage>2Fs00221</fpage>
          -
          <fpage>006</fpage>
          -0634-0.pdf
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Jones</surname>
            ,
            <given-names>J.A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Munhall</surname>
            ,
            <given-names>K.G.</given-names>
          </string-name>
          :
          <article-title>The e ects of separating auditory and visual sources on audiovisual integration of speech</article-title>
          .
          <source>Canadian Acoustics - Acoustique Canadienne</source>
          <volume>25</volume>
          (
          <issue>4</issue>
          ),
          <volume>13</volume>
          {
          <fpage>19</fpage>
          (
          <year>1997</year>
          ), https://jcaa.caaaca.ca/index.php/jcaa/article/viewFile/1106/836
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>MacDonald</surname>
          </string-name>
          , J.:
          <article-title>Hearing Lips and Seeing Voices: The Origins and Development of the 'McGurk E ect' and Re ections on AudioVisual Speech Perception over the Last 40 Years</article-title>
          .
          <source>Multisensory Research</source>
          <volume>31</volume>
          (
          <issue>1-2</issue>
          ),
          <volume>7</volume>
          {
          <fpage>18</fpage>
          (
          <year>2018</year>
          ). https://doi.org/10.1163/
          <fpage>22134808</fpage>
          -00002548, https://brill.com/view/journals/msr/31/1-2/article-p7 2.xml
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>MacDonald</surname>
          </string-name>
          , J.,
          <string-name>
            <surname>Andersen</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bachmann</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          :
          <article-title>Hearing by eye: How much spatial degradation can be tolerated?</article-title>
          <source>Perception</source>
          <volume>29</volume>
          (
          <issue>10</issue>
          ),
          <volume>1155</volume>
          {
          <fpage>1168</fpage>
          (
          <year>2000</year>
          ). https://doi.org/10.1068/p3020
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>Mallick</surname>
            ,
            <given-names>D.B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Magnotti</surname>
            ,
            <given-names>J.F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Beauchamp</surname>
            ,
            <given-names>M.S.</given-names>
          </string-name>
          :
          <article-title>Variability and stability in the McGurk e ect: contributions of participants, stimuli, time, and response type</article-title>
          .
          <source>Psychonomic bulletin &amp; review 22(5)</source>
          ,
          <volume>1299</volume>
          {
          <fpage>1307</fpage>
          (
          <year>2015</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <surname>Massaro</surname>
            ,
            <given-names>D.W.</given-names>
          </string-name>
          :
          <article-title>Children's perception of visual and auditory speech</article-title>
          .
          <source>Child development 55(5)</source>
          ,
          <volume>1777</volume>
          {
          <fpage>1788</fpage>
          (
          <year>1984</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <surname>Massaro</surname>
            ,
            <given-names>D.W.</given-names>
          </string-name>
          , Cohen,
          <string-name>
            <surname>M.M.:</surname>
          </string-name>
          <article-title>Perceiving asynchronous bimodal speech in consonant-vowel and vowel syllables</article-title>
          .
          <source>Speech Communication</source>
          <volume>13</volume>
          (
          <issue>1-2</issue>
          ),
          <volume>127</volume>
          {
          <fpage>134</fpage>
          (
          <year>1993</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <string-name>
            <surname>McGurk</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          , MacDonald, J.:
          <article-title>Hearing lips and seeing voices</article-title>
          .
          <source>Nature</source>
          <volume>264</volume>
          (
          <issue>5588</issue>
          ),
          <volume>746</volume>
          {
          <fpage>748</fpage>
          (
          <year>1976</year>
          ). https://doi.org/10.1038/264746a0, http://www.nature.com/articles/264746a0
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          12.
          <string-name>
            <surname>Munhall</surname>
            ,
            <given-names>K.G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gribble</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sacco</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ward</surname>
            ,
            <given-names>M.:</given-names>
          </string-name>
          <article-title>Temporal constraints on the McGurk e ect</article-title>
          .
          <source>Perception &amp; Psychophysics</source>
          <volume>58</volume>
          (
          <issue>3</issue>
          ),
          <volume>351</volume>
          {
          <fpage>362</fpage>
          (
          <year>1996</year>
          ). https://doi.org/10.3758/BF03206811, http://www.springerlink.com/index/10.3758/BF03206811
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          13.
          <string-name>
            <surname>Narbutt</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <given-names>O</given-names>
            <surname>'Leary</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            ,
            <surname>Allen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            ,
            <surname>Skoglund</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            ,
            <surname>Hines</surname>
          </string-name>
          ,
          <string-name>
            <surname>A.</surname>
          </string-name>
          :
          <article-title>Streaming VR for immersion: Quality aspects of compressed spatial audio</article-title>
          .
          <source>In: 23rd International Conference on Virtual System &amp; Multimedia (VSMM)</source>
          . pp.
          <volume>1</volume>
          {
          <issue>6</issue>
          . IEEE,
          <string-name>
            <surname>Ireland</surname>
          </string-name>
          (
          <year>2017</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          14.
          <string-name>
            <surname>Rana</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ozcinar</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Smolic</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          :
          <article-title>Towards Generating Ambisonics Using Audiovisual Cue for Virtual Reality</article-title>
          .
          <source>In: ICASSP 2019-2019 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)</source>
          . pp.
          <year>2012</year>
          {
          <year>2016</year>
          (
          <year>2019</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          15.
          <string-name>
            <surname>Sharma</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          :
          <article-title>Audio-visual speech integration and perceived location</article-title>
          .
          <source>Ph.D. thesis</source>
          , University of Reading (
          <year>1989</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          16.
          <string-name>
            <surname>Siddig</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ragano</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Jahromi</surname>
            ,
            <given-names>H.Z.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hines</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          :
          <article-title>Fusion confusion: exploring ambisonic spatial localisation for audio-visual immersion using the McGurk e ect</article-title>
          .
          <source>In: Proceedings of the 11th ACM Workshop on Immersive Mixed and Virtual Environment Systems</source>
          . pp.
          <volume>28</volume>
          {
          <issue>33</issue>
          (
          <year>2019</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          17.
          <string-name>
            <surname>Slater</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sanchez-Vives</surname>
            ,
            <given-names>M.V.</given-names>
          </string-name>
          :
          <article-title>Enhancing Our Lives with Immersive Virtual Reality</article-title>
          .
          <source>Frontiers in Robotics and AI</source>
          <volume>3</volume>
          (
          <issue>December</issue>
          ),
          <volume>1</volume>
          {
          <fpage>47</fpage>
          (
          <year>2016</year>
          ). https://doi.org/10.3389/frobt.
          <year>2016</year>
          .00074
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          18.
          <string-name>
            <surname>Stein</surname>
            ,
            <given-names>B.E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Meredith</surname>
            ,
            <given-names>M.A.</given-names>
          </string-name>
          :
          <article-title>The merging of the senses</article-title>
          . The MIT Press (
          <year>1993</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          19.
          <string-name>
            <surname>Tiippana</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          :
          <article-title>What is the McGurk e ect? Frontiers in Psychology 5</article-title>
          ,
          <issue>725</issue>
          (jul
          <year>2014</year>
          ). https://doi.org/10.3389/fpsyg.
          <year>2014</year>
          .
          <volume>00725</volume>
          , http://journal.frontiersin.org/article/10.3389/fpsyg.
          <year>2014</year>
          .00725/abstract
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>