<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>An Improved MVDR Filter Using Speech Presence Probability</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Quan Trong The</string-name>
          <email>quantrongthe@itmo.ru</email>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>University ITMO</institution>
          ,
          <addr-line>St.Petersburg</addr-line>
          ,
          <country country="RU">Russia</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>This paper describes an improved minimum variance distortionless response lter in a two-microphone speech enhancement system. Dual-microphone system, which is one of the most basically form of microphone array, has potentially ability of easy implementation, low cost of computation, exploiting of a priori spatial information. The proposed algorithm uses a current estimation information of target speech activity for calculating more precisely auto and cross power spectral densities. Due to the disadvantage of the conventional algorithm is still existing speech distortion, the author introduces a new adaptive technique signal processing, which is suitable for dual-microphone system. The proposal technique evaluated in noisy environments and compared with the conventional algorithm. The results show the reduction of target speech suppression about 8dB, and the quality of estimated speech in the term of the signal-to-noise ratio increased about 15.2 (dB). The enhanced performance provided that suggested algorithm can be incorporated into multi microphone signal processing system. Furthermore, speech presence probability intends to combine with various pre- ltering, post- ltering technique to obtain a certain of noise reduction. The rest of paper is organized as follow. In the Section 2, the scenario of dualmicrophone system and combination with speech presence probability are introduced. Section 3 includes the experiments, discussion of signi cant achievement of the e cient proposed technique. Finally, Section 4 gives the future of the above algorithm's development in di erent condition of noise.</p>
      </abstract>
      <kwd-group>
        <kwd>noise reduction</kwd>
        <kwd>microphone array</kwd>
        <kwd>dual-microphone</kwd>
        <kwd>minimum variance distortion response</kwd>
        <kwd>speech presence probability</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>
        Almost single-channel algorithm aim using spectral subtraction at reduce
background noise while maintaining useful speech component. In many speech
application, which associated with human life, such as speech coding, communication
Copyright ' 2019 for this paper by its authors. Use permitted under Creative
Commons License Attribution 4.0 International (CC BY 4.0).
system, distant conference require a high speech quality without speech
distortion or delay. In real environment, the target speaker always interfered by
coherent, incoherent, di use noise and others unwanted acoustic sound. Due to
highly noisy environments, the recorded signals can be corrupted and it's speech
intelligibility is a ected. With single algorithm approach, the limitation is
existing of original speech suppression and musical noise. Microphone array [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]
has been studied in many research articles. Since the crucial spatial information
has been exploited, many unresolved problem include attenuation desired signal,
residual noise can be easily removed. Microphone array signal processing give us
more advantages than mono system. In such scenario, the most important factor
is the spatial diversity, which obtained by geometry distribution of microphones.
The diversity is combined with some appropriate signal processing techniques to
improve captured signals, which contain unknown interferences and additional
noises.
      </p>
      <p>
        Minimum Variance Distortionless Response (MVDR) [
        <xref ref-type="bibr" rid="ref2 ref3 ref4 ref5 ref6">2-6</xref>
        ] is the most e
ective algorithm in term of noise reduction while saving the target speaker. MVDR
lter processes the input diversity, that is the direction of arrival (DOA) of useful
signal, and based on a constraint condition of minimization total output power
noise and una ectedness on desired signal.
      </p>
      <p>
        However, in real application; due to the rapidly change of undetermined type
of noise or complex surroundings, MVDR lter's evaluation has the limitation. In
this paper, that author proposes to incorporate the speech presence probability
[
        <xref ref-type="bibr" rid="ref7 ref8">7-8</xref>
        ] (SPP) into the MVDR lter to reduce speech distortion and increase the
speech quality of suggested algorithm. Objective measure used for comparing to
the conventional MVDR lter. The promising preliminary results provided that,
the suggested algorithm can be considered as pre- ltering method in various
complex equipments.
2
      </p>
    </sec>
    <sec id="sec-2">
      <title>Combination of MVDR</title>
    </sec>
    <sec id="sec-3">
      <title>Probability lter and Speech Presence</title>
      <p>In this section, the author presented signal processing principles of combination
between conventional MVDR lter and speech presence probability.</p>
      <p>A dual - microphone system (MA2), in which placed two omnidirectional
microphones, is the basic form of microphone array signal processing. The
received diversity such as coherence between two noisy signals, direction of arrival,
phase di erence, power level di erence, are easily need to processed to achieve
signi cant noise reduction in compared to single-channel method. The scheme
of digital signal processing by MA2 show in Fig 1. The desired target speaker
source in the same workplace with dual microphone and relates to axis MA2 an
angle s.</p>
      <p>We're after here denote the distance between microphones is d, the sound
speed is c (343m=s), 0 = d=c is the sound delay. General algorithm is
considered in frequency-domain with current frame k, frequency f , desired signal
X(f; k), two noisy signals were recorded Y1(f; k); Y2(f; k), and additive noise</p>
      <p>N1(f; k); N2(f; k). In the form of vector, the representation of short-term Fourier
transform can be expressed as:
(1)
(2)
(3)
(4)
Y1(f; k) =
Y2(f; k) =</p>
      <p>X(f; k)ej s + V1(f; k)</p>
      <p>X(f; k)e j s + V2(f; k)</p>
      <p>Let's start with Y (f; k) = [Y1(f; k) Y2(f; k)]T , V (f; k) = [V1(f; k) V2(f; k)]T
and D(f; s) = [ej s e j s ]T with ()T indicates transpose operator, and D(f; s)
is phase shift vector, where s = f 0cos( s). The equation (1-2) can be
rewritten as:</p>
      <p>Y (f; k) = X(f; k)D(f; s) + V (f; k)</p>
      <p>All signal processing algorithm aim nding an optimal solution W (f; k) at
ensuring noise reduction and maintaining target speaker. The output signal is
obtained by multiplying the coe cients of solution with vector input signals
Y (f; k). The estimated signal X^ (f; k) given by:</p>
      <p>X^ (f; k) = W H (f; k)Y (f; k)
where ()H is the symbol of Hermitian conjugation.</p>
      <p>With inverse short-term Fourier and add-overlap, the output signal is
transformed into time domain.</p>
      <p>The purpose of dual-microphone system is to extract the interest signal.
Exploiting of priori diversity is main advantage of MVDR lter. MVDR ensures
minimization of the total noise power, while maintaining the undistorted desired
signal from given determined direction s. The constraint problem leads to the
optimal solution, which can be expressed in form of vector coe cients as follows:
where P V V (f; k) is a cross spectral matrix of noise signals, P V V (f; k) =
EfV (f; k)V (f; k)g.</p>
      <p>Unfortunately, it's always di cult to calculate spectral matrix P V V (f; k), so
spectral matrix of observed signals used instead of. The cross spectral matrix of
observed signals: P Y Y (f; k) = EfY (f; k)Y (f; k)g.</p>
      <p>Matrix P Y Y (f; k) can be computed as:</p>
      <p>P Y Y (f; k) =</p>
      <p>PY1Y1 (f; k) 1:001</p>
      <p>PY2Y1 (f; k)</p>
      <p>PY1Y2 (f; k)
PY2Y2 (f; k) 1:001
where PYiYi (f; k); PYiYj (f; k) are the smoothed cross-spectra:</p>
      <p>PYiYj (f; k) =</p>
      <p>PXiXj (f; k
1) + (1
)Yi (f; k)Yj (f; k)
i; j 2 f1; 2g
where is the smoothing parameter, which in the range f0:::1g.</p>
      <p>So in conventional MVDR tler, the coe cients become:</p>
      <p>
        In practical implementations, the target speaker may not stay precisely, the
captured signals can be in uenced by unwanted interference and can not give
accuracy direction of arrival of useful signal; furthermore, the di erent
sensitivities, spatial location, frequency response, mismatch of microphones or errors of
calculation of steering vector can negative a ect on remaining desired speech at
the output of system. This produces may led to both poor interference reduction
and target speech distortion, and hence cause performance degradation [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ].
      </p>
      <p>Requirement of knowledge of speech presence probability is an essential
information to estimate and control the updating rate.</p>
      <p>The author proposed an current estimation of speech presence probability to
adjust the auto and cross power spectral densities of observed signals.
PXiXj (f; k) = SP P (f; k)PXiXj (f; k 1)+(1 SP P (f; k))Xi (f; k)Xj (f; k) i; j 2 f1; 2g
(9)</p>
      <p>The new adaptive suggested algorithm ensures accuracy, exactly and
immediately calculating according to the presence or absence of speech
components. This approach leads to decrease the speech distortion when compared to
conventional MVDR lter. Experiments have con rmed the e ectiveness of the
proposed solution.
3</p>
    </sec>
    <sec id="sec-4">
      <title>Experiments and results</title>
      <p>
        In this section, the suggested algorithm (MVDR-SPP) is performed to deal
speech enhancement and reduce speech distortion problem in an anechoeic
chamber. Dual microphone were placed on a table at the center of room, distance
between two microphones was set 5(cm); a speaker stood at distance 2(m) from
dual-microphone. The purpose of the experiment was to test the MVDR-SPP
algorithm on real signals and verify the improvement of reducing speech
suppression when compared to conventional MVDR lter (MVDR-CONV). The
objective measure NIST STNR [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ] used to measure the signal-to-noise ratio (SNR).
The scheme of the experiment is shown in Fig. 3. Two noisy recorded signals
was sampled at sampling rate 16(kHz). For calculating PSD estimation, these
necessary parameters: 512 point FFT, a Hamming window, overlap 50% were
set.
      </p>
      <p>Figure 8 shows RMS between original and processed signal by MVDR-SPP.</p>
      <p>The adaptive algorithm MVDR SPP allows to suppress nonstationary noise,
and prove the ability of algorithm. The noise reduction was about 33.5 dB. The
target speaker was remained.</p>
      <p>From Figure 9; as we can see that algorithm MVDR SPP can save the target
speech due to at these frames, the auto and cross spectral were updated according
speech presence probability; while conventional MVDR doesn't take in account.
The advantage of MVDR SPP is increasing capability of saving speech up to
8dB. The improvement in speech quality presented in Table 1, the increasing of
SNR from 26.8 to 42 (dB) provided the capability of suggested algorithm.
Method Estimation Original signal MVDR-CONV MVDR-SPP</p>
      <p>NIST STNR 4.0 26.8 42
4</p>
    </sec>
    <sec id="sec-5">
      <title>Conclusions</title>
      <p>This paper addresses the problem of enhancing a speech signal corrupted with
additive noise when observations from two microphones are available. The
experimental results indicate when the spectral components of the noisy speech
changes rapidly, we need an information of speech presence probability to
calculate accurately the auto and cross power spectral densities. The algorithm
achieves better noise cancellation, less speech distortion, increasing e ciency up
to 8 dB and can be used as an e cient front end of speech application. The
challenge of time-varying environment is always available, the author continues
using other priori spatial diversities to enhance Minimum Variance Distortionless
Response lter in di erent type of noise.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Brandstein</surname>
            <given-names>M.</given-names>
          </string-name>
          and
          <string-name>
            <surname>Ward D</surname>
          </string-name>
          . (Eds.).
          <source>Microphone Arrays: Signal Processing Techniques and Applications</source>
          , Springer,
          <year>2001</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <surname>Ehrenberg L</surname>
          </string-name>
          . et al.:
          <source>Sensitivity Analysis of MVDR and MPDR Beamformers/ IEEE 26-th Convention of Electrical and Electronics Engineers in Israel</source>
          ,
          <year>2010</year>
          , pp.
          <fpage>416</fpage>
          -
          <lpage>420</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Lockwood</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          et al.:
          <article-title>Performance of time- and frequency-domain binaural beamformers based on recorded signals from real rooms</article-title>
          .
          <source>J. Acoust. Soc. Am</source>
          .
          <volume>115</volume>
          (
          <issue>1</issue>
          ), pp.
          <fpage>379</fpage>
          -
          <lpage>391</lpage>
          , (
          <year>2004</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Stolbov</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>The</surname>
            ,
            <given-names>Q.</given-names>
          </string-name>
          <article-title>Study of MVDR dual-microphone algorithm for speech enhancement in coherent noise presence</article-title>
          . Scienti c and
          <source>Technical Journal of Information Technologies, Mechanics and Optics</source>
          ,
          <year>2019</year>
          , vol.
          <volume>19</volume>
          , no.
          <issue>1</issue>
          , pp.
          <volume>180</volume>
          {
          <issue>183</issue>
          (in Russian).
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Souden</surname>
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Benesty</surname>
            <given-names>J.,</given-names>
          </string-name>
          <article-title>A es S., A study of the LCMV and MVDR noise reduction lters</article-title>
          ,
          <source>IEEE Trans.Signal Process.</source>
          , vol.
          <volume>58</volume>
          , pp.
          <volume>4925</volume>
          {
          <issue>4935</issue>
          ,
          <string-name>
            <surname>Sept</surname>
          </string-name>
          .
          <year>2010</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>Stolbov</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          , Quan Trong The.:
          <article-title>Dual-Microphone Speech Enhancement System Attenuating both Coherent and Di use Background Noise In: A. A</article-title>
          .
          <string-name>
            <surname>Salah</surname>
          </string-name>
          et al.(Eds.)
          <string-name>
            <surname>Proc</surname>
            <given-names>SPECOM</given-names>
          </string-name>
          <year>2019</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>Gerkmann</surname>
            <given-names>T. Unbiased</given-names>
          </string-name>
          <article-title>MMSE-Based Noise Power Estimation with Low Complexity and Low Tracking Delay</article-title>
          , IEEE TASL,
          <year>2012</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <surname>Gerkmann</surname>
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hendriks</surname>
            <given-names>R</given-names>
          </string-name>
          .
          <article-title>Noise Power Estimation Based on the Probability of Speech Presence</article-title>
          ,
          <string-name>
            <surname>WASPAA</surname>
          </string-name>
          <year>2011</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>9. https://labrosa.ee.columbia.edu/projects/snreval/.</mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>