<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Bayesian estimators-based microphone array speech enhancement in adverse environment</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Quan Trong The</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Falculty of Information Security, Posts and Telecommunications Institute of Technology</institution>
          ,
          <addr-line>122 Hoang Quoc Viet, Cau Giay District, Hanoi</addr-line>
          ,
          <country country="VN">Vietnam</country>
        </aff>
      </contrib-group>
      <fpage>127</fpage>
      <lpage>136</lpage>
      <abstract>
        <p>Speech enhancement aims at noise reduction and extracting the desired target speaker from the noisy mixture in a complex recording environment. Microphone array (MA) beamforming is commonly used in almost all acoustic equipment, such as hearing aids, surveillance device, teleconference system, mobile phone, voice controlled, smart home. MA beamforming techniques use the a priori information about the properties of surrounding environment, the designed MA distribution, the direction of arrival (DoA) of interest useful signal to achieve a better noise reduction and speech enhancement at the same time. Generalized Sidelobe Canceller (GSC) beamformer eficiently remove background noise while saving the original clean speech component in an annoying recording scenario. However, due to undetermined factors, the overall GSC beamformer's performance often degraded in realistic recording schemes. In this paper, the authors proposed exploiting the Bayesian estimator of short time spectral amplitude (STSA) to gain the amplitude of beamformer's output signal. The results obtained showed that the suggested method can use the Bayesian estimator to improve the speech quality in the term of the signal-to-noise ratio (SNR) from 7.9 to 15.3 dB and reduce the speech distortion to 14.1 dB. The numerical results indicate the advantage of the author's approach to overcome the drawback of GSC beamformer in real-life application as compared to state-of-the-art solutions.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;microphone array</kwd>
        <kwd>speech enhancement</kwd>
        <kwd>gain function</kwd>
        <kwd>the signal-to-noise ratio</kwd>
        <kwd>beamformer</kwd>
        <kwd>short-time spectral amplitude</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>
        Nowadays, the perceptual metric listener, the speech quality or speech intelligibility are afected by
numerous types of noise, as shown in figure 1. Speech enhancement aims at the precise estimation of
the clean speech component from its noisy mixture of interference, background noise, non-directional
noise source, competing talker in an adverse environment. The single-channel algorithm is often based
on the spectral subtraction method, which has the simplicity of performing. However, this approach
leads to speech distortion in complex and annoying recording situations, where the non-stationary noise
can not be exactly calculated. Therefore, the microphone array (MA) beamforming [
        <xref ref-type="bibr" rid="ref1 ref2 ref3">1, 2, 3</xref>
        ] has received
more considerable attention due to its superiority of efectiveness of preserving the original speech
component while suppressing interfering noise at the same time. MA technology exploits the spatial
diversity, the a priori information about geometry of MA, the characteristic of captured situations to
achieve a better noise reduction, as shown in figure 2.
      </p>
      <p>
        GSC beamformer [
        <xref ref-type="bibr" rid="ref4 ref5 ref6">4, 5, 6</xref>
        ] is one of the most efective methods for steering the designed beampattern
toward the sound source while mitigating background noise and other signals, which come from other
directions. In practical speech application, GSC beamformer’s often implemented in the frequency
domain, because of its efectiveness in providing better source separation of the clean speech from
the observed mixture. The scheme of principal MA technology is described in figure 3 by utilizing 
microphones, with the recorded signals on each microphone 1(), ...,  () and the final output signal
().
      </p>
      <p>Unfortunately, due to the complex undetermined environment, the diferent MA sensitivities, the
error of MA displacement, the inaccurate calculation of DoA seriously afect the GSC beamformer’s
evaluation, that decreases the speech quality, speech intelligibility and listener perception.</p>
      <p>
        In this article, the author proposes using the STSA estimator for gaining the output signal to recover
the obtained signal approximate to the original speech signal. Ephraim and Malah [
        <xref ref-type="bibr" rid="ref7 ref8">7, 8</xref>
        ] proposed an
STSA estimator with superior performance compared to the other conventional methods like Wiener
ifltering and spectral subtraction. This approach is based on the constrained criteria of minimum cost
function, which describes the error between the original clean speech and the estimated speech spectral
amplitude. Several improved modifications [
        <xref ref-type="bibr" rid="ref10 ref11 ref9">9, 10, 11</xref>
        ] include the heavy-tailed non-Gaussian prior
distributions for expressing the formulation of the clean speech STSA.
      </p>
      <p>The rest of this contribution is organized as follows. The first section introduces the problem of speech
enhancement and the MA beamforming technology. The section 2 presents the principal working of GSC
beamformer in the frequency-domain. The author proposed the using of STSA estimator for improving
the GSC beamformer’s evaluation in reducing the speech distortion. The perspective experiment was
conducted in section 4. The section 5 concludes the numerical results and the author’s direction of
research in the future.</p>
    </sec>
    <sec id="sec-2">
      <title>2. Generalized sidelobe canceller beamformer</title>
      <p>
        In this section, the author presents the scheme of traditional structure of GSC beamformer [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ] in
frequency domain to recover the desired target speech component while removing the background
noise, interference and other competing talkers, as shown in figure 4. In general cases, the authors use
the model of dual-microphone array (DMA2) for illustrating the model signals.
      </p>
      <p>The representations of observed MA signals in the short-time Fourier transform (STFT) can be
expressed as follows:
1(, ) =(, )Φ + 1(, ),
2(, ) =(, )− Φ + 2(, ),
(1)
(2)
where  = 2 and ,  denote the current considered frequency and frame; (, ) is the original
speech component; 1(, ) and 2(, ) present the additive noise component, unwanted
interferences; Φ  =   0 cos( ),  0 = / and  is the range between two microphones;  = 343 (m/s) is the
sound speed propagation in the air;  is the direction of arrival of useful talker relative to the DMA2
geometry.
(3)
(4)
(5)
(6)
(7)
(8)</p>
      <p>The traditional GSC beamformer own three parts: (1) the fixed beamformer (FBF), which concerns
the steerable beampattern towards the direction of sound source, (2) the block matrix (BM) that aims at
suppressing the speech component to achieve the only noise, (3) adaptive noise canceller (ANC) that it
used to extract the target speaker from FBF’s output with using BM’s output as reference signal. The
upper branch usually implements delay and sum beamformer and lower branch uses signal subtraction.</p>
      <p>The upper branch signal (, ) and lower branch signal (, ) can be calculated as:
(, ) =
(, ) =
1(, )− Φ + 2(, )Φ</p>
      <p>2
1(, )− Φ − 2(, )Φ
2
,
.</p>
      <p>The auto-cross power spectral densities (PSD) between (, ), (, ) can be derived in the
following way:
  (, ) = (1 −  )  (,  − 1) +  (, )*(, )
  (, ) = (1 −  )  (,  − 1) +  (, )*(, )
where  is a smoothing parameter, which is in the range of {0..1} and * denote the conjugate operator.</p>
      <p>The determined Wiener filter’s coeficient yields as:
(, ) =
  (, )
  (, )
.</p>
      <p>The obtained signal by applying GSC beamformer is given by:</p>
      <p>(, ) = (, ) − (, ) * (, ).</p>
      <p>
        Because of the undetermined recording conditions, as well as the complex and annoying environment,
the displacement of MA’s configuration, the error of sampling rate, the inaccurate estimation of preferred
steering vector, the diference of microphone quality, the overall GSC beamformer’s performance in
adverse noisy situations often degraded. There is still existing speech distortion, unacceptable noise level
or musical noise, which decreases the speech quality for perceptual metric listener. Consequently, in
the next section, the author proposes using the observed phase diference of MA to form an appropriate
gain function to preserve the clean speech data at the GSC beamformer’s output.
where ( |,  ) denotes the captured MA signals conditional PDF and () expresses the STSA prior
distribution and  means for the possible value of spectral amplitude. As in single-channel approach
the STSA prior, the presented formulation of these factors can be derived as:
[
        <xref ref-type="bibr" rid="ref6">6</xref>
        ] with the assumption of complex Gaussian distribution for ( |,  ) and Rayleigh distribution for
ˆ


= 0
∞
∫︀ ( |,  )()
∞
0
∫︀ ( |,  )()
      </p>
      <p>,
( |,  ) =
︂(</p>
      <p>1
  
2 exp −  2 |  −   |2 ,</p>
      <p>︂)
() =</p>
      <p>︃(
2
 2 exp −  2
2 )︃
,</p>
    </sec>
    <sec id="sec-3">
      <title>3. STSA estimator based on phase diference</title>
      <p>In the broadside recording situation, the observed noisy MA signals can be represented as:
 (, ) = (, ) +  (, ),  = 1, 2, ..., 
at the current frequency-frame (, ).  = 2 . The speech spectral component (, ) can be
expressed as (, ) = (, ) (,) with (, ) ≥
{[−  
]} the spectral phase,  is the number of microphones.
0 being the spectral amplitude,  , ∈</p>
      <p>From the received noisy spectral  (, ), STSA estimator calculates the original amplitude
(, ).</p>
      <p>
        By applying the Bayesian rule [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ], which utilizes a minimum mean-square error (MMSE) for the
spectral amplitude ˆ
 , the posteriori probability density function (PDF) is given by:
(9)
(10)
(11)
(12)
(14)
(15)
(16)

=1
ˆ


=
1
1
4
where  2 and  2 are the speech STSA variances and spectral noise, respectively. Under the criteria
that , = , , , , is the spectral amplitude and ( |,  ) can be represented as the
product of all PDF of each microphones [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ], the formulation of conditional joint can be computed in the
following way:
( |,  ) = ∏︁ (,|,  ) = ∏︁

=1   
2 · exp
      </p>
      <p>=1
︃( ∑︁ 2, cos(  −  ,) − 2 − 2, )︃</p>
      <p>. (13)
 2</p>
      <p>The described equation of MMSE STSA estimator is given by:
where  (.) expresses the Gaussian error and
  =</p>
      <p>1
 2</p>
      <p>+ ∏︁</p>
      <p>1
=1  2,
,   = −
=1  2,
∑︁ , cos(  −  ,).</p>
      <p>With the introduced MMSE STSA estimator, the author’s proposed gain function for enhancing GSC
beamformer’s performance as:
− 2  2 +
1
2  −
︁( 2 2+  )︁ √︁ 
 5 exp  
︁(  2 )︁ (︁
1
− 
︁(   ︁)
√</p>
      <p>︁(   ︁) √︁  exp  
2   
︁(  2 )︁ (︁
1
− 
︁(   ︁)
√
 
 (, ) =
︁( ˆ

 )︁ 2</p>
      <p>1
 (, )Φ− 1 (,)(, )
where  is the Hermitian operator, (,  ) = [Φ − Φ ] and Φ  (, )
= {  (, ) (, )} is the covariance matrix of observed MA signals.</p>
      <p>And the enhanced GSC beamformer’s output signal is given by:</p>
      <p>ˆ (, ) =  (, ) ×  (, ).</p>
      <p>The method proposed by the author uses the additive gain function, which based on the STSA
estimator of phase diference. Therefore, it is referred as aaSTSA. In the next section, an experiment to
verify the efectiveness of the aaSTSA technique in realistic recording scenario is described.
=
(17)</p>
    </sec>
    <sec id="sec-4">
      <title>4. Experiments</title>
      <p>In this section, the author illustrate the performance of the proposed multi-channel microphone speech
enhancement technique for reducing the speech distortion, improving the speech quality and speech
intelligibility of GSC beamformer’s output signal.</p>
      <p>
        The author chooses the clean speech from the TIMIT database [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ] and additive street noise from
NOISEX-92 [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ]. TIMIT Acoustic-Phonetic Continous Speech Corpus contains of recordings of 630
speakers, which is considered a standard dataset for implementing various types of signal processing.
NOISEX-92 addresses several problems of estimating the efects of noise on modern automatic digital
speech processing systems. The authors used the MA with 8 microphones, with the range  = 5 (cm)
between two mounted microphones, the distance  = 3 (m) from speaker to the axis of MA, the direction
of arrival of interest useful signal  = 90∘ . The scheme of the experiment is shown in figure 5. The
sampling rate is 16, overlap 50%,    denotes the Fast Fourier Transform and    is the
Inverse Fast Fourier Transform. The purpose of the experiment is to compare the efectiveness of
increasing the speech quality in the term of SNR and reducing the speech distortion. An objective
measurement [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ] was applied to calculate SNR between the received array signals, the processed
signals by GSC beamformer (coGSbe) and the author’s approach (aaSTSA).
      </p>
      <p>The observed MA signals are presented in figure 6.</p>
      <p>By applying coGSbe, the output signal is shown in figure 7.</p>
      <p>It is seen from figure 9 that, the speech distortion occurred, due to the diferent microphone
mismatches, the inaccurate estimation of preferred DoA of target talker, the displacement of MA significantly
degrade the GSC beamformer. As a result, the amplitude of output signal at frames, where the speech
component exists, was degraded.</p>
      <p>Using the additive gain function, which is based on the amplitude Bayesian estimator, the final output
signal is given in figure 8.</p>
      <p>The comparison of energy and speech quality between the original MA signals, the signals processed
by coGSbe, and aaSTSA was shown in figure 9 and table 1.</p>
      <p>
        Based on the assumption that, the speech component is distributed as the Gamma distribution and
the additive noise is distributed according to the Gaussian model, the WADA (Waveform Amplitude
Distribution Analysis) SNR [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ] was used for computing the SNR. NIST denotes for National Institute of
Standards and Technology and STNR (Signal-To-Noise Ratio) [
        <xref ref-type="bibr" rid="ref17">17</xref>
        ] in an optimum solution for computing
the ratio SNR by applying Gaussian model.
      </p>
      <p>From the noisy mixture of MA signals, the suggested method aaSTSA calculated the a priori
information about the amplitude of the original speech signal for forming an efective gain function to reduce
speech distortion to 14.1 dB, recover the amplitude of output signal and increase the speech quality
from 7.9 to 15.3 dB. Compared to coGSbe, aaSTSA allows obtaining better speech enhancement in
preserving the speech component while suppressing background noise. The illustrated experiment
was performed in the realistic living room with MA configuration broadside. The numerical results
demonstrate that, the author’s approach can be integrated into other multi-channel signal processing
to solve several complicated problems of MA beamforming.</p>
    </sec>
    <sec id="sec-5">
      <title>5. Conclusion</title>
      <p>In this paper, the author have proposed the exploiting STSA estimator, which uses the spectral phase
estimation for improving the GSC beamformer’s speech enhancement by adding a gain function for
recovering clean speech components. The suggested method addressed the problem of speech distortion
in realistic annoying recording scenarios for increasing the SNR ratio. In addition to providing superior
evaluation in comparison with the traditional GSC beamformer, the author’s proposed method is found
to enhance performance in a complex noisy environment. The efectiveness of the suggested method
allows decreasing the speech distortion to 14.1 dB and improves the speech quality from 7.9 to 15.3
dB. From the numerical results, the suggested technique can be installed in various types of acoustic
equipment to obtain better noise reduction, speech enhancement. The future work will be aimed at
enhancing signal processing by incorporating STSA estimators with diferent types of noise properties.
Declaration on Generative AI: The author have not employed any generative AI tools.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>Q.</given-names>
            <surname>Zhang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <surname>L. Zhang,</surname>
          </string-name>
          <article-title>A robust speech enhancement method based on microphone array</article-title>
          ,
          <source>in: 2017 IEEE 17th International Conference on Communication Technology (ICCT)</source>
          ,
          <year>2017</year>
          , pp.
          <fpage>1673</fpage>
          -
          <lpage>1678</lpage>
          . doi:
          <volume>10</volume>
          .1109/ICCT.
          <year>2017</year>
          .
          <volume>8359915</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>C.</given-names>
            <surname>Xu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Zhou</surname>
          </string-name>
          , L. Xu,
          <article-title>Adaptive Speech Enhancement Algorithm Based on First-order Diferential Microphone Array</article-title>
          ,
          <source>in: 2021 IEEE 2nd International Conference on Big Data, Artificial Intelligence and Internet of Things Engineering (ICBAIE)</source>
          ,
          <year>2021</year>
          , pp.
          <fpage>41</fpage>
          -
          <lpage>44</lpage>
          . doi:
          <volume>10</volume>
          .1109/ICBAIE52039.
          <year>2021</year>
          .
          <volume>9390004</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>G.</given-names>
            <surname>Junlong</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Dongmei</surname>
          </string-name>
          , L. Runsheng,
          <article-title>Study of Speech Enhancement Based on the SecondOrder Diferential Microphone Array</article-title>
          ,
          <source>in: 2018 2nd International Conference on Imaging, Signal Processing and Communication (ICISPC)</source>
          ,
          <year>2018</year>
          , pp.
          <fpage>155</fpage>
          -
          <lpage>159</lpage>
          . doi:
          <volume>10</volume>
          .1109/ICISPC44900.
          <year>2018</year>
          .
          <volume>9006721</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>L.</given-names>
            <surname>Siping</surname>
          </string-name>
          ,
          <string-name>
            <surname>Z. Linghua,</surname>
          </string-name>
          <article-title>Research on generalized sidelobe canceller based on modified wavelet threshold function</article-title>
          ,
          <source>in: 2017 IEEE 3rd Information Technology and Mechatronics Engineering Conference (ITOEC)</source>
          ,
          <year>2017</year>
          , pp.
          <fpage>339</fpage>
          -
          <lpage>343</lpage>
          . doi:
          <volume>10</volume>
          .1109/ITOEC.
          <year>2017</year>
          .
          <volume>8122311</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>B.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.-H.</given-names>
            <surname>Zhang</surname>
          </string-name>
          ,
          <article-title>An improved speech enhancement algorithm based on generalized sidelobe canceller</article-title>
          ,
          <source>in: 2016 International Conference on Audio, Language and Image Processing (ICALIP)</source>
          ,
          <year>2016</year>
          , pp.
          <fpage>463</fpage>
          -
          <lpage>468</lpage>
          . doi:
          <volume>10</volume>
          .1109/ICALIP.
          <year>2016</year>
          .
          <volume>7846528</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>S.</given-names>
            <surname>Dai</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Q. H.</given-names>
            <surname>Abbasi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. A.</given-names>
            <surname>Imran</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A Fast</given-names>
            <surname>Blocking</surname>
          </string-name>
          <article-title>Matrix Generating Algorithm for Generalized Sidelobe Canceller Beamformer in High Speed Rail Like Scenario</article-title>
          ,
          <source>IEEE Sensors Journal</source>
          <volume>21</volume>
          (
          <year>2021</year>
          )
          <fpage>15775</fpage>
          -
          <lpage>15783</lpage>
          . doi:
          <volume>10</volume>
          .1109/JSEN.
          <year>2020</year>
          .
          <volume>3002699</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>Y.</given-names>
            <surname>Ephraim</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Malah</surname>
          </string-name>
          ,
          <article-title>Speech enhancement using a minimum-mean square error short-time spectral amplitude estimator</article-title>
          ,
          <source>IEEE Transactions on Acoustics, Speech, and Signal Processing</source>
          <volume>32</volume>
          (
          <year>1984</year>
          )
          <fpage>1109</fpage>
          -
          <lpage>1121</lpage>
          . doi:
          <volume>10</volume>
          .1109/TASSP.
          <year>1984</year>
          .
          <volume>1164453</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>Y.</given-names>
            <surname>Ephraim</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Malah</surname>
          </string-name>
          ,
          <article-title>Speech enhancement using a minimum mean-square error log-spectral amplitude estimator</article-title>
          ,
          <source>IEEE Transactions on Acoustics, Speech, and Signal Processing</source>
          <volume>33</volume>
          (
          <year>1985</year>
          )
          <fpage>443</fpage>
          -
          <lpage>445</lpage>
          . doi:
          <volume>10</volume>
          .1109/TASSP.
          <year>1985</year>
          .
          <volume>1164550</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>E.</given-names>
            <surname>Plourde</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Champagne</surname>
          </string-name>
          ,
          <article-title>Auditory-Based Spectral Amplitude Estimators for Speech Enhancement</article-title>
          ,
          <source>IEEE Transactions on Audio, Speech, and Language Processing</source>
          <volume>16</volume>
          (
          <year>2008</year>
          )
          <fpage>1614</fpage>
          -
          <lpage>1623</lpage>
          . doi:
          <volume>10</volume>
          .1109/TASL.
          <year>2008</year>
          .
          <volume>2004304</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>J. S.</given-names>
            <surname>Erkelens</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R. C.</given-names>
            <surname>Hendriks</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Heusdens</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Jensen</surname>
          </string-name>
          ,
          <article-title>Minimum Mean-Square Error Estimation of Discrete Fourier Coeficients With Generalized Gamma Priors</article-title>
          ,
          <source>IEEE Transactions on Audio, Speech, and Language Processing</source>
          <volume>15</volume>
          (
          <year>2007</year>
          )
          <fpage>1741</fpage>
          -
          <lpage>1752</lpage>
          . doi:
          <volume>10</volume>
          .1109/TASL.
          <year>2007</year>
          .
          <volume>899233</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <surname>M. B. Trawicki</surname>
          </string-name>
          , M. T. Johnson,
          <article-title>Distributed multichannel speech enhancement with minimum mean-square error short-time spectral amplitude, log-spectral amplitude, and spectral phase estimation</article-title>
          ,
          <source>Signal Processing 92</source>
          (
          <year>2012</year>
          )
          <fpage>345</fpage>
          -
          <lpage>356</lpage>
          . doi:
          <volume>10</volume>
          .1016/j.sigpro.
          <year>2011</year>
          .
          <volume>07</volume>
          .021.
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>M. E.</given-names>
            <surname>Lockwood</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D. L.</given-names>
            <surname>Jones</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R. C.</given-names>
            <surname>Bilger</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C. R.</given-names>
            <surname>Lansing</surname>
          </string-name>
          ,
          <string-name>
            <surname>W. D. O'Brien</surname>
          </string-name>
          ,
          <string-name>
            <surname>Jr.</surname>
            ,
            <given-names>B. C.</given-names>
          </string-name>
          <string-name>
            <surname>Wheeler</surname>
            ,
            <given-names>A. S.</given-names>
          </string-name>
          <string-name>
            <surname>Feng</surname>
          </string-name>
          ,
          <article-title>Performance of time- and frequency-domain binaural beamformers based on recorded signals from real rooms</article-title>
          ,
          <source>The Journal of the Acoustical Society of America</source>
          <volume>115</volume>
          (
          <year>2003</year>
          )
          <fpage>379</fpage>
          -
          <lpage>391</lpage>
          . doi:
          <volume>10</volume>
          .1121/1.1624064.
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <given-names>J. S.</given-names>
            <surname>Garofolo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L. F.</given-names>
            <surname>Lamel</surname>
          </string-name>
          , W. M. Fisher,
          <string-name>
            <given-names>J. G.</given-names>
            <surname>Fiscus</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D. S.</given-names>
            <surname>Pallett</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N. L.</given-names>
            <surname>Dahlgren</surname>
          </string-name>
          ,
          <string-name>
            <given-names>DARPA TIMIT</given-names>
            <surname>Acoustic-Phonetic Continuous Speech Corpus</surname>
          </string-name>
          ,
          <year>1993</year>
          . URL: https://nvlpubs.nist.gov/nistpubs/ Legacy/IR/nistir4930.pdf.
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <given-names>A.</given-names>
            <surname>Varga</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H. J. M.</given-names>
            <surname>Steeneken</surname>
          </string-name>
          ,
          <article-title>Assessment for automatic speech recognition: II. NOISEX-92: A database and an experiment to study the efect of additive noise on speech recognition systems</article-title>
          ,
          <source>Speech Communication</source>
          <volume>12</volume>
          (
          <year>1993</year>
          )
          <fpage>247</fpage>
          -
          <lpage>251</lpage>
          . doi:
          <volume>10</volume>
          .1016/
          <fpage>0167</fpage>
          -
          <lpage>6393</lpage>
          (
          <issue>93</issue>
          )
          <fpage>90095</fpage>
          -
          <lpage>3</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <given-names>D.</given-names>
            <surname>Ellis</surname>
          </string-name>
          ,
          <source>Objective measures of speech quality/SNR</source>
          ,
          <year>2011</year>
          . URL: https://labrosa.ee.columbia.edu/ projects/snreval/.
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [16]
          <string-name>
            <given-names>C.</given-names>
            <surname>Kim</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R. M.</given-names>
            <surname>Stern</surname>
          </string-name>
          ,
          <article-title>Robust signal-to-noise ratio estimation based on waveform amplitude distribution analysis</article-title>
          ,
          <source>in: 9th Annual Conference of the International Speech Communication Association, INTERSPEECH</source>
          <year>2008</year>
          , Brisbane, Australia,
          <source>September 22-26</source>
          ,
          <year>2008</year>
          , ISCA,
          <year>2008</year>
          , pp.
          <fpage>2598</fpage>
          -
          <lpage>2601</lpage>
          . doi:
          <volume>10</volume>
          .21437/INTERSPEECH.2008-
          <volume>644</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [17]
          <string-name>
            <given-names>A. W.</given-names>
            <surname>Rix</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. G.</given-names>
            <surname>Beerends</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. P.</given-names>
            <surname>Hollier</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A. P.</given-names>
            <surname>Hekstra</surname>
          </string-name>
          ,
          <article-title>Perceptual evaluation of speech quality (PESQ)-a new method for speech quality assessment of telephone networks and codecs</article-title>
          , in: IEEE International Conference on Acoustics, Speech, and Signal Processing,
          <string-name>
            <surname>ICASSP</surname>
          </string-name>
          <year>2001</year>
          ,
          <volume>7</volume>
          -11 May,
          <year>2001</year>
          , Salt Palace Convention Center, Salt Lake City, Utah, USA, Proceedings, IEEE,
          <year>2001</year>
          , pp.
          <fpage>749</fpage>
          -
          <lpage>752</lpage>
          . doi:
          <volume>10</volume>
          .1109/ICASSP.
          <year>2001</year>
          .
          <volume>941023</volume>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>