<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <journal-title-group>
        <journal-title>COLINS-</journal-title>
      </journal-title-group>
    </journal-meta>
    <article-meta>
      <title-group>
        <article-title>Improvement of MVDR Beamformer's Performance Based on Spectral Mask</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Quan Trong The</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Digital Agriculture Cooperative</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Cau Giay</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Ha Noi</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Viet Nam.</string-name>
        </contrib>
      </contrib-group>
      <pub-date>
        <year>2023</year>
      </pub-date>
      <volume>7</volume>
      <fpage>20</fpage>
      <lpage>21</lpage>
      <abstract>
        <p>In many speech applications, such as source tracking, hearing aids, augmented reality, teleconferencing, robot audition; acoustic beamforming is routinely implemented to enhance the speech quality, speech intelligibility of captured microphone array signals in many realworld recording situations. The designed beamformer uses priori information to form a spatial beampattern, which moves towards the target sound source while eliminating all surrounding noise and interferences. However, robust performance in annoying scenarios still exists as a challenging task, due to several reasons. In this article, the author proposed a spectral mask, which applied to Minimum Variance Distortionless Response beamformer to improve the speech enhancement. The resulting experiment shows that the advantage of suggested technique was confirmed in increasing the signal-to-noise ratio from 5.2 (dB) to 6.2 (dB) and reduce speech distortion to 3.2 (dB). The author's proposed approach consistently ensures enhancing perceptual quality metrics compared to the conventional beamformer.</p>
      </abstract>
      <kwd-group>
        <kwd>1 microphone array</kwd>
        <kwd>minimum variance distortionless response</kwd>
        <kwd>speech enhancement</kwd>
        <kwd>the signalto-noise ratio (SNR)</kwd>
        <kwd>perceptual quality</kwd>
        <kwd>robust performance</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>The utilizing of microphone arrays (MA) [1-9] and its technique beamforming has become widely
commonly used in almost speech applications, such as robot audition, teleconferencing, mobile phones,
hearing aids, surveillances devices, virtual assistants. These devices require acquiring desired speech
from a target direction in presence of third-party talker, complex annoying noise, and unwanted
interferences from the other directions. In a special recording scenario, when the talker is far from
microphones, the received signal - to - noise ratio (SNR) will be inadequate for further signal processing
and in these cases the spatial filtering can’t provide high speech quality or little distortion. The existing
beamformers outperform well in laboratory conditions but may less well in real-world situations, which
contains multiple undetermined noise source, interfering sound sources with locations and
characteristics vary with times and non-stationary.</p>
      <p>Acoustics beamforming are conveniently installed in the short time Fourier transform (STFT)
domain. In each time - frequency cell, the complex value of final output signal is derived by    ,
where  is the optimum coefficients that related to the designed beamformer’s properties. When
choosing  , a common purpose of the constrained criteria is to maximize the SNR of the beamformer
output signal with minimizing the total output noise power. For obtaining this goal, it is convenient to
calculate the direction of arrival (DOA) of interest signal   , the steering vector of target speaker
  ( ,   ), which indicates the frequency response of the target sound source and each element of MA,
and MA’s geometry distribution.</p>
      <p>Minimum Variance Distortionless Response (MVDR) [10-17] beamformer is one of the most
importance MA beamforming, which use the a priori information of   ,   ( ,   ) and the covariance
matrix of observed MA signals to find the optimum solution  . Consequently, MVDR beamformer
probably the most commerce beamforming technique. A lot of research, which referred to robust
MVDR, has been proposed, evaluated in real-world experimental conditions to avoid speech distortion.
As a rule, these algorithms are performed by extending the spatial region. Nevertheless, even assuming
perfect the DOA of useful talker or sound source localization, the different microphone sensitivities and
directional responses make the performance of MVDR beamformer is not handle well. Therefore,
speech distortion is the existing problem of MA.</p>
      <p>In this paper, the author considers the problem of preserving the original speech acquisition in noisy
environment. Since surrounding noise greatly corrupts the speech enhancement, high quality noise
reduction is an essential problem in MVDR’s performance. While precise estimation of steering vector
plays a major role for robust MA beamforming, in practical situations, the priori information of steering
vector is often based on the knowledge of MA geometry and plan wave propagation of sound source.
To overcome this limitation, recently, a time - frequency mask - based research direction has been
proposed that enhances the MVDR beamformer’s evaluation. The central idea is suppressing the speech
component in the microphone array signal.</p>
      <p>In this paper, the author suggested using a suitable spectral mask, which uses an appropriate
modified coherence - valued of surrounding noise, and desired signal. The illustrated experiments have
confirmed the effectiveness of the proposed method through comparison of the conventional MVDR
beamformer (MVDR-conventional) and the suggested technique (SLM) in terms of SNR.</p>
      <p>This contribution is organized as follows. The second section describes the principal working of
MVDR beamformer. Section III will analyze the suggested ideal of SLM and the experiments will be
evaluated in Section IV. Finally, the Conclusion and the direction of the author’s research.</p>
    </sec>
    <sec id="sec-2">
      <title>2. The model signal</title>
      <p>In this section, the principal working of MVDR beamformer is presented in Figure 4. MVDR
beamforming uses the spatial information about the direction - of - arrival of useful talker and minimizes
the total noise power output for preserving the target speech component. Consequently, MVDR
beamformer is based on the constrained problem to extracting desired speaker while suppressing all
background noise without speech distortion. The scheme of the implementation of MVDR beamformer
with dual – microphone system (DMA2) [19-25, 28] can be written as the following way in the
frequency domain.</p>
      <p>
        Two captured microphone array signals are denoted by  1( , ),  2( , ) with the frequency index
 and frame index  , respectively. The representation in short - time Fourier transform as:
 1( , ) =  ( , )    +  1( , )
 2( , ) =  ( , ) −   +  2( , )
(
        <xref ref-type="bibr" rid="ref1">1</xref>
        )
(
        <xref ref-type="bibr" rid="ref2">2</xref>
        )
      </p>
      <p>Where  ( , ): the desired speech component, additive noise  1( , ), 2( , ),   direction of
arrival of interest talker, the distance between two microphones  , speed propagation of sound in the
fresh air is  (343 m/s),  0 =  / is the sound delay and   =    0 (  ).</p>
      <p>
        Without generality, we can denote  ( ,  ) is the steering vector,  ( ,  ) = [     −   ] ,
 ( , ) = [ 1( , )  2( , )] and  ( , ) = [ 1( , )  2( , )] with symbol 
indicates transpose operator. The equations (
        <xref ref-type="bibr" rid="ref1 ref2">1-2</xref>
        ) can be expressed as the above formulation:
 −1 ( ,  )
 ( , ) =   ( ,  ) −1 ( ,  )
 ( , ) =  ( , ) ( ,  ) +  ( , )
      </p>
      <p>̂( , ) =   ( , ) ( , )</p>
      <p>In almost digital signal processing algorithm, the important requirements is finding an optimum
appropriate solution  ( , ), which adjust the final output signal  ̂( , ) is approximately the original
 ( , ):</p>
      <p>Where symbol  is Hermitian conjugation.</p>
      <p>
        The constrained of saving the desired target speech while alleviating, minimizing the total output
noise power without speech distortion can be expressed in a mathematical formulation as:

 ( , )  ( , )   ( , ) ( ,  )  . .  ( , ) ( ,  ) = 1
where   ( , ) =  { ( , ) ∗( , )} is a covariance matrix of noise signals. (
        <xref ref-type="bibr" rid="ref5">5</xref>
        ) leads to the
coefficients of MVDR beamformer:
      </p>
      <p>Unfortunately, in real - life recording situations, the information about noise often can’t be precisely
calculated or correctly estimated. And the covariance matrix of observed microphone arrays signals is
used instead of.   ( , ) =  { ( , ) ∗( , )}of received microphone signals are determined by:
where      ( , ),     ( , ), , ∈ {1,2}computed as:</p>
      <p>1 1( , ) ∗ 1.001
  ( , ) = {
  2 1( , )</p>
      <p>
        1 2( , )
  2 2( , ) ∗ 1.001}
     ( , ) = (1 −  )     ( , − 1) +   ∗( , )  ( , )
(
        <xref ref-type="bibr" rid="ref3">3</xref>
        )
(
        <xref ref-type="bibr" rid="ref4">4</xref>
        )
(
        <xref ref-type="bibr" rid="ref5">5</xref>
        )
(
        <xref ref-type="bibr" rid="ref6">6</xref>
        )
(
        <xref ref-type="bibr" rid="ref7">7</xref>
        )
(
        <xref ref-type="bibr" rid="ref8">8</xref>
        )
Where  is the smoothing parameter, which in the range {0 … 1}.
      </p>
      <p>Finally, the received optimized solution of conventional MVDR beamformer is:
 ( ,  ) =</p>
      <p>−1 ( ,   )
  ( ,   ) −1 ( ,   )</p>
    </sec>
    <sec id="sec-3">
      <title>3. The suggested spectral mask</title>
      <p>is derived in the following equation:</p>
      <p>The ideal of spectral mask 
( ,  ) is based on the estimation of a priori SNR. And the 
( ,  )
In [26], an estimation of the signal - to - noise ratio is derived by:
coherence function of the desired signal and the coherence of surrounding noisy environment.</p>
      <p>Where   ,   ,   is the coherence function between two microphone array signals, the complex
We can predict the appropriate model, which presents exactly these coherence functions due to many
factors. Based on the working [27], the authors use the formulation as:
suppress the speech component.</p>
      <p>Therefore, microphone array signal,  1( ,  ),  2( ,  ) are pre - processed as the following way to
The spectral mask allows outperforming the MVDR’s evaluation more robust. In the next section,
the authors demonstrated an experiment in coherence noise field.</p>
    </sec>
    <sec id="sec-4">
      <title>4. Experiments</title>
      <p>stand at distance  = 2( ) related to a DMA2 at direction   = 900. The distance between two
microphones is  = 5( ). The recording situation in a living room, where still exists coherence noise
field.</p>
      <p>The purpose is verifying the effectiveness of the proposed spectral mask (SLM) in comparison
with the MVDR-conventional in terms of increasing the speech quality and reducing speech distortion.
An objective measurement [18] is used for calculating the speech quality. The noisy signal is captured
with DAM2 at  = 16 . For further signal processing, these necessarily parameters are used:
 = 512, overlap 50%, smoothing parameter  = 0.5. Figure 6 shows the waveform of
microphone array signal.
7.</p>
      <p>By applying the conventional MVDR beamformer, the resulting output signal is derived in Figure</p>
      <p>The spectral mask allows removing the speech component at the MVDR beamformer’s input
and enhances the overall performance. The received signal is shown in Figure 8.</p>
      <p>In the comparison the energy of microphone array signal, the processed signals by MVDR –
conventional and SLM, we can see that SLM reduced speech diction to 3.2 (dB), and increase the speech
quality in terms of the signal-to-noise ratio (SNR) from 5.2 (dB) to 6.2 (dB).</p>
      <p>In this demonstrated experiment, the advantage of the suggested spectral mask has been proven. The
obtained result is very promising in improvement of speech enhancement by MVDR beamformer,
which is the most widely common installed MA configuration in almost acoustic device. Speech
degradation or corrupted the output signal still a problem with digital signal processing algorithms, the
author exploits the priori information about the direction of arrival of interest signal, the properties of
surrounding environment to form an appropriate spectral mask to suppress the speech component and
improve MVDR beamformer’s performance. The proposed method, which is easy to implement and
owns low computation, can be applied into multi - microphones system.</p>
    </sec>
    <sec id="sec-5">
      <title>5. Conclusion</title>
      <p>Target speech separation methods extract desired speaker from noisy mixture of speech, background
noise when interfering sources and third - party talker exits. These designed algorithms serve as
essential front - ends for many speech communication systems, such as speech recognition, digital
hearing aid devices, surveillance, smart home, speaker verification, teleconferencing systems.
Consequently, digital signal processing by MA beamforming is an important part in almost speech
applications. In this contribution, the author demonstrated an additive useful spectral mask, which
suppresses the speech component in the MA signals to enhance the MVDR beamformer’s performance.
The numerical results confirmed the suggested technique in terms of increasing the speech quality and
perceptual quality metric of the final output signal from 5.2 (dB) to 6.2 (dB) and reducing speech
distortion to 3.2 (dB). The author’s future working is combination with surrounding properties of
recording situations to improve the MVDR beamformer’s enhancement.</p>
    </sec>
    <sec id="sec-6">
      <title>6. Acknowledgements</title>
      <p>This research was supported by Digital Agriculture Cooperative. The author thanks our colleagues
from Digital Agriculture Cooperative, who provided insight and expertise that greatly assisted the
research.</p>
    </sec>
    <sec id="sec-7">
      <title>7. References</title>
      <p>[10] Zhang Z., Xu Y., Yu M. Multi-Channel Multi-Frame ADL-MVDR for Target Speech Separation.</p>
      <p>IEEE/ACM Trans. Audio Speech and Language Processing, vol. 29, pp. 3526-3540, Nov.2021.
https://doi.org/10.48550/arXiv.2012.13442.
[11] Tammen M., Doclo S. Deep Multi-Frame MVDR Filtering for Single-Microphone Speech
Enhancement // Proc. IEEE International Conference on Acoustics Speech and Signal Processing
(ICASSP), pp. 8443-8447, Jun. 2021.
[12] Fengqi T., Changchun B., Liu T. An Effective Dereverberation Algorithm by Fusing MVDR and
MCLP // 2022 IEEE International Conference on Signal Processing, Communications and
Computing (ICSPCC). DOI: 10.1109/ICSPCC55723.2022.9984583.
[13] Schreibman A., Hadad E., Barnov A., Tzirkel-Hancock E. Dual MVDR Architecture for Adaptive
Cancellation of Dynamic Interference // 2022 30th European Signal Processing Conference
(EUSIPCO). DOI: 10.23919/EUSIPCO55093.2022.9909959.
[14] Hadad E., Doclo S., Nordholm S., Gannot S. Pareto Optimal Binaural MVDR Beamformer with
Controllable Interference Suppression // 2022 International Workshop on Acoustic Signal
Enhancement (IWAENC). DOI: 10.1109/IWAENC53105.2022.9914759.
[15] Tammen M., Doclo S. Deep Multi-Frame MVDR Filtering for Binaural Noise Reduction // 2022
International Workshop on Acoustic Signal Enhancement (IWAENC).</p>
      <p>DOI: 10.1109/IWAENC53105.2022.9914742.
[16] Piyushkumar K., Shreya S., Ankur T., Hemant A. Robustness of
DAS Beamformer Over MVDR for Replay Attack Detection On Voice Assistants // 2022 IEEE
International Conference on Signal Processing and Communications (SPCOM).</p>
      <p>DOI: 10.1109/SPCOM55316.2022.9840757.
[17] Alastair H., Hafezi S., Rebecca R., Patrick A., Brookes M. A Compact Noise Covariance Matrix
Model for MVDR Beamforming. IEEE/ACM Transactions on Audio, Speech, and Language
Processing. Pp: 2049 - 2061. DOI: 10.1109/TASLP.2022.3180671.
[18] https://labrosa.ee.columbia.edu/projects/snreval/
[19] Won K., Yeoum S., Kang B., Kim M., Yeji Shin Y., Hyunseung Choo H. Inaudible
Transmission System with elective Dual Frequencies Robust to Noisy Surroundings. 2020 IEEE
International Conference on Consumer Electronics (ICCE).</p>
      <p>DOI: 10.1109/ICCE46568.2020.9042989.
[20] Zhou J., He S., Mo H., Tian X., Li Z.. A Modified Dual Microphone Adaptive Filter for
Auscultation. 2019 IEEE 14th International Conference on Intelligent Systems and Knowledge
Engineering (ISKE). DOI: 10.1109/ISKE47853.2019.9170433.
[21] Kotus J., Szwoch G. Localization of sound sources ith dual acoustic vector sensor. 2019 Signal
Processing: Algorithms, Architectures, Arrangements, and Applications (SPA).</p>
      <p>DOI: 10.23919/SPA.2019.8936724.
[22] KimS.M. Hearing Aid Speech Enhancement Using Phase Difference-Controlled
Dual</p>
      <p>Microphone Generalized Sidelobe Canceller. IEEE Access. DOI: 10.1109/ACCESS.2019.2940047
[23] Tan K., Zhang X., Wang D.L. Real-time Speech Enhancement Using an Efficient Convolutional
Recurrent Network for Dual-microphone Mobile Phones in Close-talk Scenarios. ICASSP 2019
2019 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP).</p>
      <p>DOI: 10.1109/ICASSP.2019.8683385.
[24] Huang Y.A., Shabestary T.Z., Gruenstein A. Hotword Cleaner: Dual-microphone Adaptive Noise
Cancellation with Deferred Filter Coefficients for Robust Keyword Spotting. ICASSP 2019 - 2019
IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP).</p>
      <p>DOI: 10.1109/ICASSP.2019.8682682.
[25] Bagekar S., Tank V. Dual Channel Coherence Based Speech Enhancement with Wavelet
Denoising. 2018 Second International Conference on Intelligent Computing and
Control Systems (ICICCS). DOI: 10.1109/ICCONS.2018.8662885.
[26] Schwarz A., Kellermann W. Coherent-to-Diffuse Power Ratio Estimation for Dereverberation.</p>
      <p>Page(s): 1006 - 1018. IEEE/ACM Transactions on Audio, Speech, and Language
Processing (Volume: 23, Issue: 6, June 2015). DOI: 10.1109/TASLP.2015.2418571.
[27] Jeub M., Schafer M., Esch T., Vary P. Model-based dereverberation preserving binaural cues.</p>
      <p>IEEE Trans. Audio, Speech, and Language Process., vol. 18, no. 7, pp. 1732–1745, 2010.
DOI: 10.1109/TASL.2010.2052156.
[28] Pu Y., Butterfield D., Garcia J., Xie J., Lin M., Sauhta R., Farley R., Shellhammer S.,
Derkalousdian M., Newham A., Shi C., Shenoy R., Gousev E., Attar R. An Ultra-low-power 28nm
CMOS Dual-die ASIC Platform for Smart Hearables. 2018 IEEE Biomedical Circuits
and Systems Conference (BioCAS). DOI: 10.1109/BIOCAS.2018.8584806.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <surname>Dietzen</surname>
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Doclo</surname>
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Moonen</surname>
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Waterschoot</surname>
            <given-names>T.</given-names>
          </string-name>
          <string-name>
            <surname>Integrated</surname>
          </string-name>
          <article-title>Sidelobe Cancellation and Linear Prediction Kalman Filter for Joint Multi-Microphone Speech Dereverberation Interfering Speech Cancellation and Noise Reduction</article-title>
          .
          <source>IEEE/ACM Trans. Audio Speech Lang. Process</source>
          , vol.
          <volume>28</volume>
          , pp.
          <fpage>740</fpage>
          -
          <lpage>754</lpage>
          ,
          <year>2020</year>
          . DOI:
          <volume>10</volume>
          .1109/TASLP.
          <year>2020</year>
          .
          <volume>2966869</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <surname>Wei</surname>
            <given-names>Wang W.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chen</surname>
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Wang R. A Fast Irregular</surname>
          </string-name>
          <article-title>Microphone Array Design Method Based on Acoustic Beamforming</article-title>
          .
          <source>IEEE Sensors Journal. DOI: 10.1109/JSEN</source>
          .
          <year>2023</year>
          .
          <volume>3240888</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <surname>Albertini</surname>
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bernardini</surname>
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Borra</surname>
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Antonacci</surname>
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sarti</surname>
            <given-names>A</given-names>
          </string-name>
          .
          <article-title>Two-Stage Beamforming With Arbitrary Planar Arrays of Differential Microphone Array Units</article-title>
          .
          <source>IEEE/ACM Transactions on Audio, Speech, and Language Processing</source>
          . pp:
          <fpage>590</fpage>
          -
          <lpage>602</lpage>
          , DOI: 10.1109/TASLP.
          <year>2022</year>
          .
          <volume>3231719</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <surname>Yang</surname>
            <given-names>W.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Huang</surname>
            <given-names>G.</given-names>
          </string-name>
          , Zhang W.,
          <string-name>
            <surname>Chen</surname>
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Benesty</surname>
            <given-names>J</given-names>
          </string-name>
          .
          <article-title>Dereverberation with differential microphone arrays and the weighted-prediction-error method</article-title>
          .
          <source>2018 16th International Workshop on Acoustic Signal Enhancement (IWAENC)</source>
          , pp.
          <fpage>376</fpage>
          -
          <lpage>380</lpage>
          ,
          <year>2018</year>
          . DOI:
          <volume>10</volume>
          .1109/IWAENC.
          <year>2018</year>
          .
          <volume>8521286</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <surname>Xiao</surname>
            <given-names>Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zhu</surname>
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Song</surname>
            <given-names>W.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Wan</surname>
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gu</surname>
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Li</surname>
            <given-names>T.</given-names>
          </string-name>
          <string-name>
            <surname>Acoustic</surname>
          </string-name>
          <article-title>Beamforming via Interference-Plus-Noise Covariance Matrix Construction for Interferences and Noise Attenuation</article-title>
          .
          <source>2022 IEEE International Conference on Robotics and Biomimetics (ROBIO)</source>
          .
          <source>DOI: 10.1109/ROBIO55434</source>
          .
          <year>2022</year>
          .
          <volume>10012011</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <surname>Kagimoto</surname>
            <given-names>Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Itoyama</surname>
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Nishida</surname>
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Nakadai</surname>
            <given-names>K.</given-names>
          </string-name>
          <article-title>Spotforming by NMF Using Multiple Microphone Arrays</article-title>
          .
          <source>2022 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS)</source>
          .
          <source>DOI: 10.1109/IROS47612</source>
          .
          <year>2022</year>
          .
          <volume>9981808</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <surname>Kodrasi</surname>
            <given-names>I.</given-names>
          </string-name>
          , Doclo S.
          <source>Joint Late Reverberation and Noise Power Spectral Density Estimation in a Spatially Homogeneous Noise Field // 2018 IEEE International Conference on Acoustics Speech and Signal Processing (ICASSP) IEEE</source>
          , pp.
          <fpage>441</fpage>
          -
          <lpage>445</lpage>
          ,
          <year>2018</year>
          . DOI:
          <volume>10</volume>
          .1109/ICASSP.
          <year>2018</year>
          .
          <volume>8462142</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <surname>Braun</surname>
            <given-names>S.</given-names>
          </string-name>
          <article-title>Evaluation and Comparison of Late Reverberation Power Spectral Density Estimators</article-title>
          .
          <source>IEEE/ACM Trans. Audio Speech Lang. Process</source>
          , vol.
          <volume>26</volume>
          , no.
          <issue>6</issue>
          , pp.
          <fpage>1056</fpage>
          -
          <lpage>1071</lpage>
          ,
          <year>June 2018</year>
          . DOI:
          <volume>10</volume>
          .1109/TASLP.
          <year>2018</year>
          .
          <volume>2804172</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <surname>Cheng</surname>
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bao</surname>
            <given-names>C.</given-names>
          </string-name>
          , Cui Z.
          <article-title>Mass: Microphone array speech simulator in room acoustic environment for multi-channel speech coding and enhancement</article-title>
          .
          <source>Applied Sciences</source>
          , vol.
          <volume>10</volume>
          , no.
          <issue>4</issue>
          , pp.
          <fpage>1484</fpage>
          ,
          <year>2020</year>
          . https://doi.org/10.3390/app10041484.
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>