<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Indoor Localization Using Multiple Stereo Speakers for Smartphones</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Masanari Nakamura</string-name>
          <email>Nakamura.Masanari@db.MitsubishiElectric.co.jp</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Hiroshi Kameda</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Information Technology R&amp;D Center</institution>
          ,
          <addr-line>Mitsubishi Electric Corp., Kanagawa</addr-line>
          ,
          <country country="JP">Japan</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>In this paper, we propose an acoustic indoor localization method for an embedded microphone of a smartphone using multiple stereo speakers installed indoors. In this configuration, although stereo speakers are synchronized, speakers not producing stereo sound are not synchronized. Therefore, a moving microphone could receive acoustic signals required for localization at different locations. This causes bias error in the conventional method. To address this issue, we propose a method that utilizes an asynchronous tracking filter and compensates the differences in locations of received signals using motion modelling. Through experiments, we verify that our proposed method can effectively reduce the bias error.</p>
      </abstract>
      <kwd-group>
        <kwd>Acoustic Signal</kwd>
        <kwd>Asynchronous Tracking Filter</kwd>
        <kwd>Smartphone</kwd>
        <kwd>Time Difference of Arrival</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Introduction</title>
      <p>
        These days, mobile devices like smartphones are highly popular, and localization
using such devices are gaining attention [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. Although global navigation satellite
system (GNSS) is a common localization method, it cannot be used indoors because
receiving GNSS signals is difficult in such environments. Hence, other approaches
employing the embedded sensors of smartphones are required.
      </p>
      <p>
        Acoustic signals are suitable for accurate indoor localization using smartphones.
In a related work [
        <xref ref-type="bibr" rid="ref2 ref3 ref4 ref5">2-5</xref>
        ], such signals were produced by speakers and received by a
microphone; subsequently, the microphone’s location was calculated. The time
difference of arrival (TDoA) between the received signals provides elliptic hyperboloid,
which indicates the area where the microphone exists. The estimation of location is
done by calculating the intersection between multiple elliptic hyperboloids
      </p>
      <p>In these methods, highly precise time synchronization of all the speakers is
required. Certain audio players that can synchronize several speakers are available.
However, they are more expensive than commercial off the shelf (COTS) stereo
speakers, which have only two channels owing to their limited application.</p>
      <p>Therefore, we utilize COTS stereo speakers in this study. These speakers can
produce two acoustic signals simultaneously from two sides. The receiver obtains the
elliptical hyperboloid from the TDoA of the speaker signals. Using the TDoAs of
multiple stereo speakers, the receiver’s location can be calculated. However, when the
microphone is moving, TDoA can be observed at different locations, thereby
rendering the intersection of elliptical hyperboloid to drift away from the true location. This
is because speakers not producing stereo sound cannot produce signals
simultaneously.</p>
      <p>We herein propose an asynchronous tracking filter to compensate the difference of
observation locations, and therefore reduce the bias error.</p>
      <p>The rest of this paper is organized as follows. Chapter 2 describes the conventional
method and its problem. Chapter 3 presents the proposed method for dealing with the
above-mentioned problem of bias error. In Chapter 4, simulation experiments were
conducted to evaluate the proposed method. Chapter 5 concludes this paper.
2</p>
    </sec>
    <sec id="sec-2">
      <title>Related work</title>
      <p>
        In indoor environments, some localization methods using acoustic signals utilize
the embedded speakers of smartphones and microphones installed at indoor
environment [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ]. Whereas, other methods use smartphone’s embedded microphone and
speakers installed at indoor environment [
        <xref ref-type="bibr" rid="ref2 ref3 ref4 ref5">2-5</xref>
        ]. Herein, these two types of methods are
referred to as active-tag system and passive-tag system, respectively. In the former,
multiple signals produced by speakers of multiple smartphones collide at the
microphone. Although difficult, if the transmission time of speakers is precisely
synchronized, this collision is avoidable. Therefore, the latter being more preferable for
localization of multiple devices, is the focus of this section.
      </p>
      <p>
        Acoustic indoor localization methods are of two types: ranging-based and
TDoAbased. The former utilizes range measurements between speakers and microphones.
The latter utilizes the difference in signal received times. Ranging-based localization
is typically more precise than TDoA-based localization. However, it requires high
accurate time synchronization (e.g., μ order) between speakers and microphones and
is therefore used for systems employing designated devices [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ]. TDoA-based
localization does not require such synchronizations. As accurate time synchronization is
difficult in smartphones, TDoA-based localization more preferable. We herein describe
the localization for a two-dimensional space for our convenience.
      </p>
      <p>The relationship between the location and received time is represented as
 
 
 1 −  ( 1) −  
 1 −  ( 1) −  
 2 −  ( 2) =  ⋅ ( 1 −  2)
 3 −  ( 3) =  ⋅ ( 1 −  3)
(1)
where the term   = (  ,   ) denotes the position vector of speaker  , and  ( )
denotes the position vector of microphone at time  .   ( = 1,2,3) denotes the
received time of signal produced by speaker  ,c is the sound speed, and   (⋅) is the
Euclidean distance. Equations in (1) represent hyperbolic curves of the microphone’s
position  ( ).</p>
      <p>When the smartphone is stationary, the positions,  ( 1),  ( 2),  ( 3) are the same
and location can be calculated by solving equation (1). When the smartphone is
moving,  ( 1),  ( 2),  ( 3)are not the same. However, the differences are insignificant as
bias error.
is
all the speakers are synchronized. Therefore, location can be obtained without large</p>
      <p>To simultaneously produce signals from all speakers, a special device such as a
multi-channel audio player is required. As such devices are expensive, we instead use
a multiple COTS stereo speaker, which hereinafter is referred to as a unit.</p>
      <p>In this configuration, although speakers from a same unit are synchronized, those
from different units are not. The relationship between the location and received time
 
 
( 1 −  ( 1)) −  
( 2 −  ( 2)) −</p>
      <p>(( 21 −−  (( 21)))) ==  ⋅⋅ (( 12 −−  12)).</p>
      <p>where  
spectively; 
 ，

 ，</p>
      <p>are locations of speakers R and L, belonging to unit 

(
= 1,2),
re are received times of signals produced by R and L, respectively.</p>
      <p>When the smartphone is stationary, microphone positions,  ( 1),  ( 1),  ( 2),
and  ( 2) are same, and the location can be easily estimated without bias error. When
the smartphone is moving, equations  ( 1) ≈  ( 1) and  ( 2) ≈  ( 2) are satisfied
owing to speaker synchronization. However, the differences between  ( 1) and
 ( 2) can increase if the smartphone starts moving. In this case, the intersection of the
hyperbolic curve is drifted away from the true position (refer Figure 1-(a)), and
location estimated by solving (2) incurs bias error.
(2)
collide with each other at the microphone, thereby resulting in large errors. Therefore,
the transmission time of each unit should be controlled using general communication
systems such as Wi-Fi and Bluetooth.</p>
      <p>
        Speaker time lags in conventional methods produce bias errors, therefore
necessitating a multiple-channel audio player to synchronize all speakers [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ].
      </p>
      <p>The smartphone’s current position is estimated using an asynchronous tracking
filter when a TDoA is measured. The estimation is done by calculating the intersection
between the predicted position and hyperbolic curve of TDoA, as shown in Figure
1(b), and it reduces the bias error as shown in Figure 1-(a), the details of which are
described in the next section.
3.2</p>
      <p>Asynchronous tracking filter.</p>
      <sec id="sec-2-1">
        <title>TDoA observation model and motion model.</title>
        <p>state and position is represented as
The state vector,  comprises position and velocity [

 ̇  ̇ ] where  ,  are
positions,  ̇,  ̇ are velocities, and  is the transpose. The relationship between the

  
0
1
the TDoA observation time.   denotes as the observation time of  -th TDoA.
TDoA measurement, comprising   and   . We herein define the lesser of   or t as

The TDoA measurement,   received at   , corresponding to unit  is represented as
  = 
 +   


− 
 +   


= ℎ (  ) +    −   


where  
 ,    are the measured noises of received times produced by R and L of</p>
        <p>equation (2), ℎ (  ) represents
the unit  , respectively.   is the state vector [ 
 
 ̇
 ̇ ] at time   . From
ℎ (  ) =
 
(  −   ) −  
(  −    )
distribution has the reproductive property.</p>
        <p>The observation noises  
 ,    are white Gaussian noise with zero-mean and</p>
        <p>known variance</p>
        <p>2
 ,  2 , respectively. The measurement error of TDoA   is</p>
        <p>−    . Notably, we assume that  


  nd    are independent. In this case,  


is white Gaussian noise with zero-mean and variance   2 +  2 because Gaussian



As the motion model of the microphone, we utilize the constant velocity model.
  =
1
0
0
1
  −1 +
     =  (   )  −1 +   
(6)
(3)
(4)
(5)
where   is the process noise vector   
 
Δ  is the difference of the observation times   −  −1
.
the motion.    and  
 are white Gaussian noise with zero-mean and variance   2 .</p>
        <p>which represents ambiguity of
Algorithm
state   = 





measurement is inputted to the asynchronous tracking filter ( = 1), particles  
representing the microphone state are generated ( = 1, … ,  ). Each element of the
is generated by uniform random numbers. The
ated ( ≥
ticle.
weight   is updated.
weight  k of the state   is given as  k = 1/ . When   and  k were already
gener
2), the prediction based on the motion model (6) is conducted for each
parIn the updating step, the likelihood of each particle is calculated using   and the
(7)
  =</p>
        <p>1
√2
2
exp
−</p>
        <p>− ℎ (  )
2 2</p>
        <p>2


The estimated state   is obtained by the weighted sum ∑=1   ⋅   .</p>
        <p>
          When the next measurement is inputted, the above-mentioned process is repeated
with the resampled particles [
          <xref ref-type="bibr" rid="ref9">9</xref>
          ].
4
4.1
        </p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>Numerical result</title>
      <sec id="sec-3-1">
        <title>Scenarios</title>
        <p>We conducted following simulation experiments for the evaluation of our proposed
method. Figure 2 shows the microphone route and the four speaker locations. The
microphone moves at a speed of 1 m/sec and receives its first TDoA measurement at
the route’s starting position (-3.0, 1.5).
nal length and the reverberation time of multipath. Regarding the signal length, a</p>
        <p>
          ̇  
short one with large amplitude is preferable for the precision and the updating rate of
localization. However, the amplitude is generally restricted due to speaker’s
inaudibility. To obtain the desired signal-to-noise ratio (SNR) with the restricted amplitude, a
longer signal can be utilized. A reverberation time of 100 ms is sufficient to attenuate
the multipath [
          <xref ref-type="bibr" rid="ref4">4</xref>
          ]. Considering the above-mentioned reasons, transmission intervals of
these simulations are set to 100 ms and 300 ms for the short and the long signal cases,
respectively.
        </p>
        <p>
          The observation noise of the received time  depends on the bandwidth and SNR.
The bandwidth and SNR herein are set to 1 kHz and 30 dB, respectively. These
parameters were determined in consideration of the characteristic of microphones
emHence, the observation noise of TDoA is 2 × 1.58 × 10−5 = 3.16 × 10−5 sec.
bedded in smartphones [
          <xref ref-type="bibr" rid="ref4 ref5">4-5</xref>
          ]. In this case, the observation noise  is 1.58 × 10−5 sec.
        </p>
        <p>
          The conventional method numerically solves equation (2) using the current and
previous measurements. The observation noise, process noise, and number of particles
ly. The range of uniform random numbers to generate these particles
of the proposed method were set to 3.16 × 10−5 sec, 0.5 m/sec, and 5000,
respectiveat  = 1 were [
          <xref ref-type="bibr" rid="ref3">-3, 3</xref>
          ]，[
          <xref ref-type="bibr" rid="ref3">-1, 3</xref>
          ]，[-1.5, 1.5]，[-1.5, 1.5] m,
reas the conventional method cannot estimate it.
        </p>
        <p>Root mean square error (RMSE) was used for evaluating each trial with the
number of trials set to 100. The starting position  = 1 was excluded from this evaluation
4.2</p>
      </sec>
      <sec id="sec-3-2">
        <title>Results and Discussion</title>
        <p>intersections of hyperbolic curves, far from true positions, which represent the bias
errors. Figure 4-(b) indicates that the proposed method can reduce these errors using
the asynchronous tracking filter.</p>
        <p>In Figures 3-(a) and 3-(b), RMSE of the proposed method was minimum around  =
0 m, turning slightly worse in  &gt; 0 m. For clarification, predicted particles at  =
0 m and 2.5 m were plotted in Figures 5-(a) and 5-(b). The transmission interval was
set to 300 ms. In these figures, the particles are scattered in an ellipse, and the major
axis was roughly parallel to the previously plotted hyperbolic curve. This is because
these particles were generated by extrapolating the resampled particles based on the
previously obtained hyperbolic curve.</p>
        <p>When the microphone’s position was  = 0 m (Figure 5-(a)), current hyperbolic
curve crossed the major axis of the ellipse nearly at right angles. In this case, the area
of resampled particles was narrow, leading to high precision. When the microphone’s
position was  = 2.5 m (Figure 5-(b)), the minor axis of the ellipse crossed the
current hyperbolic curve at roughly right angles, and the area of resampled particles was
relatively broad, leading to reduction in precision.</p>
        <p>In this paper, we described an acoustic indoor localization system with multiple
COTS stereo speakers. We proposed the asynchronous tracking filter to reduce the
bias error caused by the asynchrony of speakers that do not produce stereo sound
when the smartphone is moving. The simulation experiments showed that the
proposed method can effectively reduce the bias error.</p>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <given-names>J.</given-names>
            <surname>Xiao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Zhou</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Yi</surname>
          </string-name>
          and
          <string-name>
            <given-names>L. M.</given-names>
            <surname>Ni</surname>
          </string-name>
          , “
          <article-title>A survey on wireless indoor localization from the device perspective,” ACM Computing Surveys</article-title>
          , vol.
          <volume>49</volume>
          , no.
          <issue>2</issue>
          , pp.
          <volume>25</volume>
          :
          <fpage>1</fpage>
          -
          <lpage>25</lpage>
          :
          <fpage>31</fpage>
          , (
          <year>2016</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <given-names>K.</given-names>
            <surname>Liu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Liu</surname>
          </string-name>
          and
          <string-name>
            <given-names>X.</given-names>
            <surname>Li</surname>
          </string-name>
          , “
          <article-title>Guoguo: Enabling Fine-Grained Smartphone Localization via Acoustic Anchors,”</article-title>
          <source>IEEE Trans. on Mobile Computing</source>
          , vol.
          <volume>15</volume>
          , no.
          <issue>5</issue>
          , pp.
          <fpage>1144</fpage>
          -
          <lpage>1156</lpage>
          , (
          <year>2016</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <given-names>F. J.</given-names>
            <surname>Álvarez</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Aguilera</surname>
          </string-name>
          and
          <string-name>
            <given-names>R. L.</given-names>
            <surname>Valcarce</surname>
          </string-name>
          , “
          <article-title>CDMA-based acoustic local positioning system for portable devices with multipath cancellation,” Digital Signal Processing</article-title>
          , vol.
          <volume>62</volume>
          , pp.
          <fpage>38</fpage>
          -
          <lpage>51</lpage>
          , (
          <year>2017</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <given-names>T.</given-names>
            <surname>Akiyama</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Nakamura</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Sugimoto</surname>
          </string-name>
          and
          <string-name>
            <given-names>H.</given-names>
            <surname>Hashizume</surname>
          </string-name>
          , “
          <article-title>Smart phone localization method using dual-carrier acoustic waves</article-title>
          ,
          <source>” Proc. of IPIN</source>
          <year>2013</year>
          , pp.
          <fpage>1</fpage>
          -
          <lpage>9</lpage>
          ,
          <year>2013</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <given-names>M.</given-names>
            <surname>Nakamura</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Akiyama</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Sugimoto</surname>
          </string-name>
          and
          <string-name>
            <given-names>H.</given-names>
            <surname>Hashizume</surname>
          </string-name>
          , “
          <article-title>3D FDM-PAM: rapid and precise indoor 3D localization using acoustic signal for smartphone</article-title>
          ,
          <source>” Proc. of ACM Ubicomp</source>
          <year>2014</year>
          , pp.
          <fpage>123</fpage>
          -
          <lpage>126</lpage>
          , (
          <year>2014</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <given-names>F.</given-names>
            <surname>Höflinger</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Zhang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Hoppe</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Bannoura</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Reindl</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Wendeberg</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Buhrer</surname>
          </string-name>
          and
          <string-name>
            <given-names>C.</given-names>
            <surname>Schindelhauer</surname>
          </string-name>
          , “
          <article-title>Acoustic Self-calibrating System for Indoor Smartphone Tracking (ASSIST</article-title>
          ),
          <source>” Proc. of IEEE IPIN</source>
          <year>2012</year>
          , pp.
          <fpage>1</fpage>
          -
          <lpage>9</lpage>
          , (
          <year>2012</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <given-names>N.</given-names>
            <surname>Priyantha</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Chakraborty</surname>
          </string-name>
          and
          <string-name>
            <given-names>H.</given-names>
            <surname>Balakrishnan</surname>
          </string-name>
          , '
          <article-title>The Cricket Location Support System,”</article-title>
          <source>Proc. of ACM MobiCom</source>
          <year>2000</year>
          , pp.
          <fpage>32</fpage>
          -
          <lpage>43</lpage>
          , (
          <year>2000</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <given-names>Y.</given-names>
            <surname>Takabayashi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Matsuzaki</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Kameda</surname>
          </string-name>
          and
          <string-name>
            <given-names>M.</given-names>
            <surname>Ito</surname>
          </string-name>
          , “
          <article-title>Target tracking using TDOA/FDOA measurements in the distributed sensor network</article-title>
          ,
          <source>” Proc. of SICE</source>
          <year>2008</year>
          , pp.
          <fpage>3441</fpage>
          -
          <lpage>3446</lpage>
          , (
          <year>2008</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <given-names>S. M.</given-names>
            <surname>Arulampalam</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Maskell</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Gordon</surname>
          </string-name>
          and
          <string-name>
            <given-names>T.</given-names>
            <surname>Clapp</surname>
          </string-name>
          , “
          <article-title>A Tutorial on Particle Filters for Online Nonlinear/Non-Gaussian Bayesian Tracking,”</article-title>
          <source>IEEE Trans. on Signal Processing</source>
          , vol.
          <volume>50</volume>
          , no.
          <issue>2</issue>
          , pp.
          <fpage>174</fpage>
          -
          <lpage>188</lpage>
          , (
          <year>2002</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <given-names>R. A.</given-names>
            <surname>Mark</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. A.</given-names>
            <surname>Scheer</surname>
          </string-name>
          and
          <string-name>
            <given-names>W. A.</given-names>
            <surname>Holm</surname>
          </string-name>
          , “Principles of Modern Radar: Basic Principles,” SciTech Publishing, (
          <year>2010</year>
          ).
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>