<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Browser Application for Virtual Audio Walkthrough</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Thomas Deppisch</string-name>
          <email>thomas.deppisch@student.tugraz.at</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Alois Sontacchi</string-name>
          <email>sontacchi@iem.at</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Student, Graz University of Technology, and University of Music and Performing Arts</institution>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>University of Music and Performing Arts, Institute of Electronic Music and Acoustics</institution>
          ,
          <addr-line>Inffeldgasse 10, 8010 Graz</addr-line>
          ,
          <country country="AT">Austria</country>
        </aff>
      </contrib-group>
      <fpage>145</fpage>
      <lpage>150</lpage>
      <abstract>
        <p>-We present an application allowing an interactive virtualization of auditory scenes. It enables the user to navigate through the virtual scene inside a web browser. Audio signals are spatialized for headphone playback using a binaural Ambisonics approach. A mixture of cues is used to activate and enhance distance perception. Customized scenes are created using a simple text file which contains meta data regarding properties of the virtual room and the audio objects. In order to scale the audio reproduction quality corresponding to available computational power, parameters like Ambisonics order and image source order are used to adjust the virtualization during runtime. The source code is provided online1.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>I. INTRODUCTION</title>
      <p>
        Hitherto in conventional and classical audio recordings the
acoustic perspective within the recording has been defined by
the tonmeister. However, new developments [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ] provide the
possibility to follow new practices in media/audio immersion:
Listeners can navigate throughout a production visiting any
favored position of interest. The addressed invention [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ] relates
to an audio production, processing, and playback apparatus to
convey a multichannel interactive audio experience, allowing
the listener to traverse an entire sound scene. Hereinafter, we
present a web based implementation of this approach.
Before going into implementation details, the following
introduction states how direction and distance of acoustic sources
are perceived and reproduced. Basic concepts of the Web
Audio Application Programming Interface (API) for audio
processing in a browser environment are shown as well.
      </p>
      <sec id="sec-1-1">
        <title>A. Perception of direction</title>
        <p>
          Cues for the perception of an acoustic source direction
are classified into monaural and binaural cues [
          <xref ref-type="bibr" rid="ref2">2</xref>
          ]. Binaural
cues utilize information from differences in both ear signals
while monaural cues utilize equivalent parts of both ear signals
to determine the direction of a sound source [
          <xref ref-type="bibr" rid="ref2">2</xref>
          ]. Binaural
cues can be further divided into interaural level differences
(ILDs) and interaural time differences (ITDs). ILDs arise
due to head shadowing effects for signals with small wave
lengths compared to the diameter of the head. Hence, lateral
sources produce higher levels on the ipsilateral ear than on the
contralateral ear [
          <xref ref-type="bibr" rid="ref2">2</xref>
          ]. The delayed arrival of a sound signal at
the contralateral ear in comparison to the ipsilateral ear results
in an interaural time difference. Such a delay is evaluated
using the phase difference in both ear signals. For wave
        </p>
      </sec>
    </sec>
    <sec id="sec-2">
      <title>1https://git.iem.at/thomasdeppisch/walkthrough</title>
      <p>
        lengths smaller than the diameter of the head these phase
differences do not contain useful information. Therefore, ITDs
are predominantly used for localization of signals with low
frequency content [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]. Still, evaluation of the signal envelope
allows localization based on ITDs for higher frequency signal
components [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ].
      </p>
      <p>
        Monaural cues are manifested in direction dependent
spectral changes of the ear signals’ frequency responses. These
spectral changes emerge due to reflections on pinna and torso,
resulting in constructive and destructive interferences. Spectral
localization cues are predominantly important for localization
of elevated sources in sagittal planes, to prevent confusions
and ambiguities [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ].
      </p>
      <sec id="sec-2-1">
        <title>B. Head-related transfer function</title>
        <p>
          Both, monaural and binaural cues are incorporated in the
head-related transfer function (HRTF) and its time domain
representative, the head-related impulse response (HRIR) [
          <xref ref-type="bibr" rid="ref2">2</xref>
          ].
The HRIR can be obtained by placing microphone probes
inside the ear channels of a test person or dummy head
and measuring the impulse response for a number of source
directions [
          <xref ref-type="bibr" rid="ref4">4</xref>
          ]. The HRIR is generally direction-dependent and
hence can be used to simulate direction of a source in binaural
synthesis. For distances smaller than 1 m the HRTF also shows
distance-dependent spectral variations. For non-static sources
or when head movements are incorporated, interpolation of
a finite number of measured HRTFs is essential [
          <xref ref-type="bibr" rid="ref5">5</xref>
          ]. The
anthropometric differences between human individuals result
in individual spectral differences in HRTFs which can lead
to an impairment of the binaural experience when using
nonindividualized HRTFs.
        </p>
      </sec>
      <sec id="sec-2-2">
        <title>C. Perception of distance</title>
        <p>
          Distance perception for acoustic sources is generally less
accurate than the perception of direction [
          <xref ref-type="bibr" rid="ref6">6</xref>
          ]. There are several
acoustic cues which allow a distance estimation for sound
sources but also non-acoustic cues that play a big role in
overall distance perception. The most prominent acoustic
distance cue is the inverse distance law for sound pressure
which states a 6 dB reduction of sound pressure level when
doubling the source distance in free field conditions [
          <xref ref-type="bibr" rid="ref6">6</xref>
          ].
Another acoustic distance cue is the direct-to-reverberant
energy ratio in reflective environments. Here, close sources
provide a greater amount of direct energy in comparison to
reverberant energy [
          <xref ref-type="bibr" rid="ref6">6</xref>
          ]. For sources further away than 15 m, air
absorption results in high frequency attenuation and therefore
in spectral distance cues [
          <xref ref-type="bibr" rid="ref2">2</xref>
          ]. Furthermore, for sources closer
than 1 m an increase in low frequency ILDs has a strong
impact on distance perception for close sources [
          <xref ref-type="bibr" rid="ref7">7</xref>
          ].
        </p>
      </sec>
      <sec id="sec-2-3">
        <title>D. The Web Audio API</title>
        <p>
          The Web Audio API2 (WAA) allows modular audio
processing in a web browser environment. Audio signals thereby
are sent through an audio routing graph consisting of audio
nodes which can be connected arbitrarily. A source node such
as the MediaElementAudioSourceNode allows the integration
of audio files into the routing graph. Several predefined
audio nodes such as BiquadFilterNode, DelayNode, GainNode
and ConvolverNode provide the possibility of realtime audio
processing. The AudioDestinationNode connects the audio
routing graph to the audio hardware. The WAA also allows
basic spatialization by providing a SpatialListenerNode and
a SpatialPannerNode. Customization of settings like HRTF
set, distance function and directivity function are currently not
possible [
          <xref ref-type="bibr" rid="ref8">8</xref>
          ].
        </p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>II. RELATED WORK</title>
      <p>
        So far, traversing a sound scene in reproduction could be
realized by audio spatialization based on isolated recordings
combined with additional spatial recordings or rendering of
reverberation (object-based). Although the listener is meant to
be located at a central position, by changing the arrangement
of the virtual sources the playback perspective at the
reproduction side can be adapted. There are several products allowing
the use of this approach, e.g. Fraunhofer Spatial Sound Wave3,
or the Ambix Plugin Suite [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ].
      </p>
      <p>
        Moreover, Pihlajama¨ki and Pulkki [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ] presented a different
approach based on the DirAC [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ] method. There the sound
field is decomposed into a non-diffuse and diffuse part. Then
the non-diffuse part gets resynthesized by assigning a direction
to each frequency band. Transformations of the direction
vectors, gain control and diffuseness control are used to simulate
translations of the listener.
      </p>
      <p>
        A method for sound field navigation using Ambisonics was
presented by Allen and Kleijn [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ]. After the directional
decomposition of a signal, an adjustment for the translated
origin is performed by filtering. Re-encoding is done in respect
to the new angles based on the translation vector.
BogJS is a JavaScript framework for object-based audio
rendering in browsers4. A demo5 shows the use case of auditory
scene virtualization in a web browser. As the spatialization is
done solely with Web Audio API functionalities, the
possibilities of personalization (e.g. change of the HRTF set) and
flexible adjustments (e.g. of the distance gain function) are
restricted.
      </p>
    </sec>
    <sec id="sec-4">
      <title>2https://www.w3.org/TR/webaudio/</title>
      <p>3https://www.idmt.fraunhofer.de/en/institute/projects products/q t/
spatialsound wave.html
4https://github.com/IRT-Open-Source/bogJS
5https://lab.irt.de/demos/object-based-audio/interactive/</p>
    </sec>
    <sec id="sec-5">
      <title>III. RECORDING AN AUDITORY SCENE FOR</title>
      <p>VIRTUALIZATION</p>
      <p>
        For recording of auditory scenes with the goal of later
virtualization two approaches are feasible: Virtualization of
sound objects recorded through spot microphones or
virtualization of the scene recorded by multichannel microphone
arrays (cf. figure 1). In the first case every microphone signal
represents an acoustic object in the virtual space, e.g. a musical
instrument. In the second case the signals of one microphone
array represent a part of the sound field spatially sampled at
one point in the room. Hence, the overall sound intensity of
the multichannel microphone arrays needs to be normalized,
so a higher density of microphone arrays in one part of the
room does not result in a higher intensity. A hybrid approach
combining spot microphones and multichannel microphone
arrays is also feasible. During playback every microphone
capsule is interpreted as a virtual speaker object which then
gets placed in the room according to its original position.
Multiplication of the signal s(t) with the spherical
harmonics evaluated at the desired source position θ~0 contained in
~yN , yields the Ambisonics encoded signals χ~N (t) (eq. (1)).
The order at which the evaluation of spherical harmonics is
truncated is called Ambisonics order N . The encoded signals
are decoded to speaker signals ~sls(t) by multiplication with a
suitable decoder matrix D (eq. (2)). The decoder matrix can be
obtained in several ways such as mode-matching, sampling or
AllRAD [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ]. The vector ~aN can contain psychoacoustically
motivated optimization factors, e.g. for max ~rE optimization.
      </p>
      <p>
        Max ~rE optimization reduces sidelobes and therefore leads to
a more distinct source localization (cf. [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ], [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ]).
Apart from full periphonic (3D) Ambisonics, circular
harmonics can be employed to obtain planar (2D) Ambisonics.
Further, mixed-order schemes are used to encode horizontal
source information in higher order than vertical
information [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ]. Rotation of a sound field is done efficiently in
the Ambisonics domain by matrix multiplication as described
in [
        <xref ref-type="bibr" rid="ref17">17</xref>
        ].
      </p>
      <sec id="sec-5-1">
        <title>B. Virtual Ambisonics approach</title>
        <p>
          For binaural synthesis a virtual Ambisonics approach is
used [
          <xref ref-type="bibr" rid="ref18">18</xref>
          ]. A regular distribution of virtual speakers is placed
around the virtual listener. Decoding of the encoded
Ambisonics signals χ~N (t) at the virtual speaker positions θ~q is achieved
by multiplication with the decoding matrix Dvls. The binaural
signals for the left and right ear (eq. (3), (4)) are obtained by
convolving the resulting virtual loudspeaker signals with their
corresponding HRIRs and summing them up for each ear.
q=1
q=1
m
sl(t) = X H RIRl,q(θ~q) ∗ (~eqT Dvls χ~N (t))
        </p>
        <p>
          m
sr(t) = X H RIRr,q(θ~q) ∗ (~eqT Dvls χ~N (t))
(3)
(4)
This approach, in contrast to HRTF interpolation methods [
          <xref ref-type="bibr" rid="ref5">5</xref>
          ],
allows a rotation of the encoded sound field in the Ambisonics
domain [
          <xref ref-type="bibr" rid="ref17">17</xref>
          ] instead of interpolation of HRTFs for every
sound object. Therefore, the number of needed HRTFs is only
depending on the number of virtual speakers and not on the
number of virtual sound objects. This can reduce the amount
of convolutions needed and hence reduce the computational
effort.
        </p>
      </sec>
    </sec>
    <sec id="sec-6">
      <title>V. IMPLEMENTATION</title>
      <sec id="sec-6-1">
        <title>A. General functionality</title>
        <p>
          The application uses an interaction of JavaScript code and
Web Audio API (WAA) audio nodes based on C++
implementations. Background signal processing such as convolutions,
filtering and gain adjustments are accomplished by WAA audio
nodes. The calculations to retrieve the values for spatialization
are done in JavaScript code. For Ambisonics processing,
classes of the open source JavaScript library JSAmbisonics6
[
          <xref ref-type="bibr" rid="ref19">19</xref>
          ] were adapted to provide periphonic as well as planar
Ambisonics processing. As JSAmbisonics is built on top of
the WAA as well, a seamless integration is possible.
In the following the construction of auditory scenes allowing
a virtual walkthrough is explained step by step.
        </p>
      </sec>
      <sec id="sec-6-2">
        <title>B. Scene File</title>
        <p>To construct the virtual scene, meta data needs to be
provided in a simple text file, the scene file (cf. figure 2). A
valid scene file needs to follow the JSON 7 (JavaScript Object
Notation) standard.</p>
        <p>6https://github.com/polarch/JSAmbisonics
7http://json.org/
[{
} ,{
} ,{
}]
” t y p e ” : ” room ” ,
” w i d t h ” : 4 . 5 ,
” l e n g t h ” : 5 . 5 ,
” h e i g h t ” : 4 ,
” l i s t e n e r S t a r t ” :
{”x ” : 2 , ” y ” : 1</p>
        <p>}
” t y p e ” : ”mono ” ,
” name ” : ” Noise ” ,
” p o s i t i o n ” : {”x ” : 1 , ” y ” : 1 , ” z ” : 1
” g a i n ” : 0 . 8 ,
”NFC” : 1 ,
” o r i e n t a t i o n ” : {” azim ” : 9 0 , ” e l e v ”:
” d i s t G a i n ” : {” a ” : 1 . 4 , ” g0 ” : 1 } ,
” f i l e ” : ” s o u n d s / n o i s e . wav”
} ,
−45} ,
” t y p e ” : ” f o u r C h a n n e l A r r a y ” ,
” name ” : ” Oktava ” ,
” c e n t e r ” : {”x ” : 4 , ” y ” : 4 } ,
” c e n t e r D i s t a n c e ” : 0 . 5 ,
” d i r e c t i v i t y ” : 0 . 5 ,
” f i l e ” : ” s o u n d s / o k t a v a 1 . ogg ” ,
” channelMapping ” : {” s p e a k e r 1 ” : 1 , ” s p e a k e r 2 ” : 2 , ”</p>
        <p>
          s p e a k e r 3 ” : 3 , ” s p e a k e r 4 ” : 4 }
In the first section the scene file provides information of the
room as well as coordinates for the starting point of the virtual
listener. Below the room data an arbitrary number of audio
objects can be defined. Objects of type mono are based on a
mono audio track, e.g. a spot microphone recording. Objects
of type fourChannelArray represent a spatially sampled part of
the sound field recorded by a microphone array consisting of
four capsules. Each defined audio object has parameters like
position, gain, orientation, distance gain function, directivity
and reference to a sound file. Optionally, near field
compensation filters (NFC) which approximate the filters given
in [
          <xref ref-type="bibr" rid="ref20">20</xref>
          ] can be activated for mono objects. Objects of type
fourChannelArray are defined by a center position and a center
distance for each of the four corresponding virtual speakers.
The sound file of a fourChannelArray object contains four
separate mono channels. To map these four mono channels to
the corresponding virtual speaker object, a channel mapping
parameter is provided.
        </p>
        <p>By using the directivity parameter a virtual speaker radiation
directivity can be controlled. The directivity gain follows
equation (5) and hence enables interpolation between
omnidirectional (γ = 1), cardioid (γ = 0.5) and figure of eight
(γ = 0). These directivity patterns are also valid for the three
dimensional space as the angle ϕ is calculated as the angle
between the vector pointing from the virtual speaker to the
listener and the vector pointing in the same direction as the
speaker.</p>
        <p>gdir = γ + (1 − γ)cos(ϕ)
The distance gain function follows equation (6) and can be
adjusted by using the parameters α and g0. The resulting
distance gain equals 1 for a distance r = 1, linearily interpolates
to g0 for distances r &lt; 1 and decreases by 1/rα for distances
r &gt; 1.</p>
        <p>gdist
=
(g0 + (1 − g0)r
1
rα
, if r ≤ 1
, if r &gt; 1
(5)
(6)</p>
        <p>
          In the next step, when activated, mirror image sources for
each audio object are built. These additional copies of sources
follow the concept of simulating room reflections by mirroring
sources along the room boundaries as explained in [
          <xref ref-type="bibr" rid="ref21">21</xref>
          ].
Image sources of first and second order are provided in the
application and can be activated during runtime. Activation of
image sources in big auditory scenes can lead to performance
impairments due to the fact that the number of sources and
hence all calculations for spatialization are multiplied.
From this point onwards the signal processing steps are
displayed in a block diagram (cf. figure 4). Relating to the
position of the virtual listener, angle and distance to each
virtual speaker are calculated dynamically. From this data
directivity gain and distance gain as described in eqations
(5), (6), as well as a dynamic delay line (equation (7)) are
adjusted.
        </p>
        <p>Δt =</p>
        <p>r
343 ms
As the delay is adjusted dynamically to fit the distance r
between listener and speaker, it is able to reproduce the
Doppler shift.</p>
        <p>Image sources are then lowpass filtered simulating a high
frequency loss caused by absorption during reflections on
room boundaries. In the last step before Ambisonics
encoding, loudspeaker objects corresponding to fourChannelArray
sources are intensity normalized as described in section III.
Before binaural headphone signals are obtained, Ambisonics
encoding, rotation and decoding takes place. The encoder
evaluates spherical or circular harmonics at the speaker directions
relative to the virtual listener. Mono as well as
fourChannelArray sources are encoded in an adjustable Ambisonics
order N . If fourChannelArray sources (first order Ambisonics
microhones) are encoded in a higher order than first order, the
sound field does not get reproduced accurately. Yet, due to
the superposition of several sound field sample points, audio
information from the four room directions get reproduced
more sharply when higher order encoding is enforced. For
close distances all Ambisonics channels but the W-Channel
(contains omnidirectional information) get interpolated to zero
to avoid discontinuities when passing through a virtual speaker
object.
(7)
The Ambisonics rotator is able to rotate the whole sound field
in the Ambisonics domain. It enables head rotations of the
virtual listener.</p>
        <p>
          At the decoding stage Ambisonics signals get decoded to a
regular distribution of virtual speakers. The number of virtual
speakers depends on the Ambisonics order N : For periphonic
Ambisonics a t-design [
          <xref ref-type="bibr" rid="ref22">22</xref>
          ] of degree t = 2N +1 and for planar
Ambisonics a circular distribution of 2N + 2 speakers are
used. For HRTF individualization arbitrary SOFA8 (Spatially
Oriented Format for Acoustics) HRTFs are supported.
        </p>
      </sec>
      <sec id="sec-6-3">
        <title>D. User interface</title>
        <p>
          Figure 5 shows the user interface of the application. Before
scene playback can be started, an auditory scene and an HRTF
set need to be chosen by using the blue dropdown menus.
Optionally, the Ambisonics type, Ambisonics order N and
image source order can be adjusted to fit the scene-specific
needs and computational possibilities. The Ambisonics type
can be switched between 2D, 3D and 2D in combination with
first order 3D components. The restriction of a maximum
of 32 channels per audio node by the WAA and a highest
supported t-design of degree t = 21 by JSAmbisonics yields
maximum Ambisonics orders of N = 4 for 3D, N = 15 for
2D and N = 10 for 2D with first order 3D components. The
navigation of the listener (depicted by a head, cf. figure 5) is
accomplished by mouse dragging or using the up and down
arrow keys. The left and right arrow keys as well as the
azimuth slider are used to turn the head of the listener in
the horizontal plane. The elevation slider is used to perform
up and down head movements which are not graphically
depicted as the scene is represented from a 2D perspective.
Alternatively, head movements can be controlled via a
lowcost open-source MIDI headtracker9 [
          <xref ref-type="bibr" rid="ref23">23</xref>
          ] which is integrated
using the Web MIDI API10 and WebMidi.js11. The usage of
a headtracker is currently only possible either using Chrome
or Opera browsers, supporting the Web MIDI API. A volume
slider and a start/pause toggle allow controlling the playback.
A grey canvas below the settings section represents the room.
Inside the canvas the listener and the virtual speaker objects
are depicted. A virtual speaker object of type mono is depicted
by a single speaker symbol. Virtual speaker objects of type
fourChannelArray are represented by four speaker symbols
arranged in a circle.
        </p>
      </sec>
    </sec>
    <sec id="sec-7">
      <title>VI. CONCLUSION AND OUTLOOK</title>
      <p>After informal listening tests the presented application
creates a promising impression of a virtual concert scene.
Localization of single sound sources works well, especially if
mono sources (corresponding to spot microphones at recording
stage) are used. The use of fourChannelArray sources
(corresponding to microphone arrays sampling a part of the sound
field) enhances the immersion. Therefore, a combination of</p>
    </sec>
    <sec id="sec-8">
      <title>8https://www.sofaconventions.org/ 9https://git.iem.at/DIY/MrHeadTracker 10https://www.w3.org/TR/webmidi/ 11https://github.com/cotejp/webmidi</title>
      <p>fourChannelArray and mono sources leads to the best results.
Crosstalk between spot microphones should be avoided as
much as possible as it may split the perceived direction of
a sound source. The perceived immersion due to a valid room
impression can be further improved by using image sources.
Unfortunately, for big auditory scenes it is often not possible
to activate image sources as the number of simultaneously
processed audio channels rises by a multiple for every image
source order. For big auditory scenes containing a high number
of audio channels the computational power of an average
personal computer may then be insufficient. The efficiency of
the program might be improved by using an underlying C++
implementation integrated through a JavaScript wrapper like
in WAA audio nodes. New drafts of the WAA also contain
AudioWorkerNode classes which might be able to enhance the
performance. Further challenges occur when embedding the
application into a website: The limited download speed might
prohibit the playback of big auditory scenes due to the big
amount of audio data which needs to be downloaded.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          <article-title>[1] “Method and apparatus acoustic scene playback,” European patent application: PCT/EP2016/075595, applicant institution: HUAWEI Technologies CO</article-title>
          . LTD. (China),
          <source>inventors: Scho¨rkhuber Christian</source>
          , Zotter Franz, Frank Matthias, Ho¨ldrich Robert, and Grosche Peter.
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>J.</given-names>
            <surname>Blauert</surname>
          </string-name>
          , Spatial Hearing:
          <article-title>The psychophysics of human sound localization</article-title>
          . MIT Press,
          <year>1997</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>E.</given-names>
            <surname>Macpherson</surname>
          </string-name>
          and
          <string-name>
            <given-names>J.</given-names>
            <surname>Middlebrooks</surname>
          </string-name>
          , “
          <article-title>Listener weighting of cues for lateral angle: The duplex theory of sound localization revisited</article-title>
          ,
          <source>” Journal of the Acoustical Society of America</source>
          , vol.
          <volume>111</volume>
          , no.
          <issue>5</issue>
          ,
          <year>2002</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>H.</given-names>
            <surname>Moller</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. F.</given-names>
            <surname>Sorensen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Hammershi</surname>
          </string-name>
          , and
          <string-name>
            <given-names>C. B.</given-names>
            <surname>Jensen</surname>
          </string-name>
          , “
          <article-title>Headrelated transfer functions of human subjects</article-title>
          ,
          <source>” J. Audio Eng. Soc</source>
          , vol.
          <volume>43</volume>
          , no.
          <issue>5</issue>
          , pp.
          <fpage>300</fpage>
          -
          <lpage>321</lpage>
          ,
          <year>1995</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>K.</given-names>
            <surname>Hartung</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Braasch</surname>
          </string-name>
          , and
          <string-name>
            <given-names>S. J.</given-names>
            <surname>Sterbing</surname>
          </string-name>
          , “
          <article-title>Comparison of different methods for the interpolation of head-related transfer functions,”</article-title>
          <source>in AES 16th International Conference on Spatial Sound Reproduction</source>
          ,
          <year>1999</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>P.</given-names>
            <surname>Zahorik</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Brungart</surname>
          </string-name>
          ,
          <article-title>and</article-title>
          <string-name>
            <given-names>A.</given-names>
            <surname>Bronkhorst</surname>
          </string-name>
          ,
          <article-title>A“uditory distance perception in humans: A summary of past and present research,” Acta Acustica united with Acustica</article-title>
          , vol.
          <volume>91</volume>
          , pp.
          <fpage>409</fpage>
          -
          <lpage>420</lpage>
          ,
          <year>05 2005</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>D. S.</given-names>
            <surname>Brungart</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N. I.</given-names>
            <surname>Durlach</surname>
          </string-name>
          , and
          <string-name>
            <given-names>W. M.</given-names>
            <surname>Rabinowitz</surname>
          </string-name>
          ,
          <article-title>A“uditory localization of nearby sources. II. Localization of a broadband source,”</article-title>
          <source>The Journal of the Acoustical Society of America</source>
          , vol.
          <volume>106</volume>
          , no.
          <issue>4</issue>
          , pp.
          <fpage>1956</fpage>
          -
          <lpage>1968</lpage>
          ,
          <year>1999</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>T.</given-names>
            <surname>Carpentier</surname>
          </string-name>
          , “
          <article-title>Binaural synthesis with the Web Audio API,” in 1st Web Audio Conference</article-title>
          (WAC), Paris, France,
          <year>2015</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>M.</given-names>
            <surname>Kronlachner</surname>
          </string-name>
          ,
          <article-title>A“mbisonics plug-in suite for production and performance usage</article-title>
          ,” in Linux Audio Conference,
          <year>2013</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>T.</given-names>
            <surname>Pihlajama</surname>
          </string-name>
          <article-title>¨ki and V. Pulkki, “Synthesis of complex sound scenes with transformation of recorded spatial sound in virtual reality</article-title>
          ,
          <source>” J. Audio Eng. Soc</source>
          , vol.
          <volume>63</volume>
          , no.
          <issue>7</issue>
          /8, pp.
          <fpage>542</fpage>
          -
          <lpage>551</lpage>
          ,
          <year>2015</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>V.</given-names>
            <surname>Pulkki</surname>
          </string-name>
          , “
          <article-title>Spatial sound reproduction with directional audio coding</article-title>
          ,
          <source>” J. Audio Eng. Soc</source>
          , vol.
          <volume>55</volume>
          , no.
          <issue>6</issue>
          , pp.
          <fpage>503</fpage>
          -
          <lpage>516</lpage>
          ,
          <year>2007</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>A.</given-names>
            <surname>Allen</surname>
          </string-name>
          and
          <string-name>
            <given-names>W. B.</given-names>
            <surname>Kleijn</surname>
          </string-name>
          ,
          <article-title>A“mbisonic soundfield navigation using directional decomposition and path distance estimation</article-title>
          ,” in 4th International Conference on Spatial Audio, Graz, Austria,
          <year>2017</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <given-names>F.</given-names>
            <surname>Zotter</surname>
          </string-name>
          ,
          <article-title>A“nalysis and synthesis of sound-radiation with spherical arrays</article-title>
          ,
          <source>” Dissertation</source>
          , University of Music and Performing Arts, Graz,
          <string-name>
            <surname>A</surname>
          </string-name>
          ,
          <year>2009</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <given-names>F.</given-names>
            <surname>Zotter</surname>
          </string-name>
          and
          <string-name>
            <given-names>M.</given-names>
            <surname>Frank</surname>
          </string-name>
          ,
          <article-title>A“ll-round ambisonic panning and decoding</article-title>
          ,
          <source>” J. Audio Eng. Soc.</source>
          , Vol.
          <volume>60</volume>
          , No.
          <volume>10</volume>
          ,
          <year>2012</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <given-names>M.</given-names>
            <surname>Gerzon</surname>
          </string-name>
          , “
          <article-title>General metatheory of auditory localisation</article-title>
          ,
          <source>” in Preprint 3306, 92nd Conv. Audio Eng. Soc.</source>
          ,
          <year>1992</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [16]
          <string-name>
            <given-names>C.</given-names>
            <surname>Travis</surname>
          </string-name>
          ,
          <article-title>A“ new mixed-order scheme for ambisonic signals</article-title>
          ,” in Ambisonics Symposium, Graz,
          <string-name>
            <surname>A</surname>
          </string-name>
          ,
          <year>2009</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [17]
          <string-name>
            <given-names>J.</given-names>
            <surname>Ivanic</surname>
          </string-name>
          and
          <string-name>
            <given-names>K.</given-names>
            <surname>Ruedenberg</surname>
          </string-name>
          , “
          <article-title>Rotation Matrices for Real Spherical Harmonics</article-title>
          . Direct Determination by Recursion,”
          <source>The Journal of Physical Chemistry</source>
          , vol.
          <volume>100</volume>
          , no.
          <issue>15</issue>
          , pp.
          <fpage>6342</fpage>
          -
          <lpage>6347</lpage>
          ,
          <year>1996</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          [18]
          <string-name>
            <given-names>M.</given-names>
            <surname>Noisternig</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Musil</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Sontacchi</surname>
          </string-name>
          , and
          <string-name>
            <given-names>R.</given-names>
            <surname>Ho</surname>
          </string-name>
          <article-title>¨ldrich, “3d binaural sound reproduction using a virtual ambisonic approach</article-title>
          ,” in IEEE International Symposium on Virtual Environments, Human-Computer
          <source>Interfaces and Measurement Systems</source>
          ,
          <year>2003</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          [19]
          <string-name>
            <given-names>A.</given-names>
            <surname>Politis</surname>
          </string-name>
          and
          <string-name>
            <surname>D.</surname>
          </string-name>
          Poirier-Quinot, “
          <article-title>JSAmbisonics: A Web Audio library for interactive spatial sound processing on the web,” Interactive Audio Systems Symposium</article-title>
          , York, UK,
          <year>2016</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          [20]
          <string-name>
            <given-names>J.</given-names>
            <surname>Daniel</surname>
          </string-name>
          and S. Moreau, “
          <article-title>Further study of sound field coding with higher order ambisonics,” in 116th Conv</article-title>
          .
          <source>Audio Eng. Soc.</source>
          ,
          <year>2004</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          [21]
          <string-name>
            <given-names>J.</given-names>
            <surname>Allen</surname>
          </string-name>
          and
          <string-name>
            <given-names>D.</given-names>
            <surname>Berkley</surname>
          </string-name>
          , “
          <article-title>Image method for efficiently simulating smallroom acoustics</article-title>
          ,
          <source>” Journal of the Acoustical Society of America</source>
          , vol.
          <volume>65</volume>
          , no.
          <issue>4</issue>
          , pp.
          <fpage>943</fpage>
          -
          <lpage>950</lpage>
          ,
          <year>1979</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          [22]
          <string-name>
            <given-names>F.</given-names>
            <surname>Zotter</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Frank</surname>
          </string-name>
          ,
          <article-title>and</article-title>
          <string-name>
            <given-names>A.</given-names>
            <surname>Sontacchi</surname>
          </string-name>
          , “
          <article-title>The virtual t-design ambisonicsrig using vbap,” in 1st EAA Euroregio Ljubljana</article-title>
          ,
          <year>2010</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          [23]
          <string-name>
            <given-names>M.</given-names>
            <surname>Romanov</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Berghold</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Rudrich</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Zaunschirm</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Frank</surname>
          </string-name>
          , and
          <string-name>
            <given-names>F.</given-names>
            <surname>Zotter</surname>
          </string-name>
          , “
          <article-title>Implementation and evaluation of a low-cost head-tracker for binaural synthesis</article-title>
          ,
          <source>” in Paper 9689, 142nd Conv. Audio Eng. Soc.</source>
          ,
          <year>2017</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>