<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Microphone Array Spectral Mask Using Artificial Neural Network for Enhancing Minimum Variance Distortionless Response Beamformer ⋆</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>QuanTrong The</string-name>
          <email>theqt@ptit.edu.vn</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="editor">
          <string-name>CPWErooUrckResehdoinpgs IhStSpN:/c1e6u1r3-w-0s.o7r3g CEUR Workshop Proceedings (CEUR-WS.org)</string-name>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Post and Telecommunication Institute of Technology</institution>
          ,
          <addr-line>Hanoi</addr-line>
          ,
          <country country="VN">Vietnam</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>ProfIT AI 2024: 4</institution>
        </aff>
      </contrib-group>
      <abstract>
        <p>In several speech applications, the requirement of high perceptual quality and intelligibility of speech is an important role in signal processing to reduce the level of noise. Therefore, improving the captured speech signal is an essential challenging task in speech applications, such as hearing aid, smart-home, mobile phone, voice - controlled device, teleconference system, smart vehicle. Microphone array (MA) technology has been commonly installed in almost all acoustic equipment to achieve better performance and noise reduction in adverse and complex recording environments. MA are typically placed at a distance from the desired speaker and provides a high directional beampattern towards the sound source while suppressing or minimizing total output noise power. MA beamforming forms directional gain at certain direction of speech source and attenuation background noise, interference. However, in the complex and adverse environment, the MA beamforming's performance is often corrupted due to the existence of transport, nondirectional noise, diffuse noise field, competing talker. Therefore, an additive spectral mask, which blocks the speech component at the observed MA signals to increase the Minimum Variance Distortionless Response (MVDR) beamformer's performance is an optimum solution. In this paper, the author proposed a new structure of artificial neural networks (ANN) for learning MA signal's characteristics to remove speech component for enhancing MVDR beamformer's evaluation. The numerical simulations show the effectiveness of the author's proposed method in improving speech quality in the term of signal-to-noise (SNR) ratio from 5.5 (dB) to 6.0 (dB) and reduce the speech distortion to 5.5 (dB). Objective scores were utilized to measure the overall performance of the conventional MVDR beamformer and the author's suggested technique.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;Microphone array</kwd>
        <kwd>beamforming</kwd>
        <kwd>the signal-to-noise ratio</kwd>
        <kwd>speech quality</kwd>
        <kwd>artificial neural network</kwd>
        <kwd>spectral mask</kwd>
        <kwd>1</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>In many applications such as mobile phones, hearing aids, teleconference system, speech acquisition,
smart – home, voice – controlled device, speech enhancement is applied to improve the captured
MA signals. When the target speaker is distant from the capturing microphone, speech quality and
speech intelligibility are usually significantly degraded due to the reverberation, surrounding noise,
interference factor. The single channel approach often leads to speech distortion due to it usually
based on spectral subtraction method, which properly estimate noise power in stationary situations
and not sufficient in complex, annoying and non – stationary recording scenario. Therefore, the
using the MA beamforming attracted more scholars, engineering and researchers to exploit the
designed spatial information to obtain better speech enhancement, noise reduction.</p>
      <p>
        MA technique use the priori information about the designed MA distribution, the direction of
arrival of interest useful signal, the characteristics of background noise, the acoustical features, the
preprocessing and post – Filtering techniques to process observed MA signals. MA beamforming
algorithms can be categorized into two groups: fixed beamformer with delay-and-sum DAS [
        <xref ref-type="bibr" rid="ref1 ref2 ref3">1-3</xref>
        ];
adaptive beamformer: differential microphone array DIF [
        <xref ref-type="bibr" rid="ref4 ref5 ref6">4-6</xref>
        ], generalized sidelobe canceller GSC
[
        <xref ref-type="bibr" rid="ref7 ref8 ref9">7-9</xref>
        ], linearly constraint minimum variance LCMV [
        <xref ref-type="bibr" rid="ref10 ref11 ref12">10-12</xref>
        ], minimum variance distortionless
response MVDR [
        <xref ref-type="bibr" rid="ref13 ref14 ref15">13-15</xref>
        ].
      </p>
      <p>MVDR beamformer is one of the most useful beamforming, which is applied into numerous
speech applications to extract the target speech component while suppressing background noise.
However, due to the microphone mismatches, the displacement of MA, the error of estimation of
steering vector, the difference between microphone sensitivities, MVDR beamformer’s evaluation
often corrupted. For overcoming the drawback, spectral mask is an effective solution for improving
MVDR beamformer’s performance in complex environment.</p>
      <p>
        Time - frequency (T-F) masking is based on the windowing - disjoint orthogonality assumption
of concentration of speech energy only to a few time - frequency points [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ]. A T - F mask typically
approximates the ideal binary mask (IBM) and multiplied with the received MA signals, thus passing
only the designed component to improve the speech enhancement [
        <xref ref-type="bibr" rid="ref17">17</xref>
        ]. Machine learning methods
are popular installed in speech enhancement algorithms for properly learning acoustic database to
form effective signal processing systems. In [
        <xref ref-type="bibr" rid="ref18">18</xref>
        ], several types of noise were learned by a
nonnegative matrix factorization (NMF), which is used to denoise the received MA signals. The authors
of [
        <xref ref-type="bibr" rid="ref19">19</xref>
        ] train a long short-term memory (LSTM) to obtain T-F mask for removing background noise,
enhancing speech quality. The denoising and automatic speech recognition (ASR) are performed by
using a long short-term memory recurrent neural network [20].
      </p>
      <p>Beamforming is linear filtering used to MA signals to extract the desired speaker from the noisy
mixture, amplify the useful signal and attenuate background noise. In this paper, the author proposed
an ANN, which uses MA characteristic to train ANN for enhancing MVDR beamformer’s
performance in complex, adverse recording scenario. The numerical results have confirmed the
effectiveness of the author’s suggested method in increasing the speech quality from 5.5 dB to 6.0 dB
and reduce speech distortion to 5.5 dB.</p>
      <p>The rest of this article is organized as follows. The next section introduces the principal working
MVDR beamformer. Section III describes the artificial neural network for training MA features.
Section IV demonstrates a perspective experiment to illustrate the advantages of the author’s
proposed technique. Section V concludes the promising result.</p>
    </sec>
    <sec id="sec-2">
      <title>2. Minimum Variance Distortionless Response Beamformer</title>
      <p>As in Figure 4, the scheme of MVDR beamformer’s implementation was depicted. In general case,
the author uses dual - microphone array (DMA2) model to express the principal working of MVDR
beamformer in the frequency - domain. At considered frame , frequency , two captured
microphone arrays signals can be described as:
!(, ) = (, )"#! + !(, )
$(, ) = (, )%"#! + $(, )
(1)
(2)</p>
      <p>Where &amp; = '(&amp;), &amp; is the direction of arrival of interest useful signal relative to the axis
of DMA2, ' = ⁄ is time delay, with  is the range between two mounted microphones,  =
343 (⁄) is the speed of propagation of sound in the air, (, ) is the clean speech component,
!(, ), $(, ) is the additive noise.</p>
      <p>We denote: (, ) = [!(, ) $(, )]( , (, ) = [!(, ) $(, )]( , &amp;(, &amp;) =
["#! %"#! ]( , the equations (1) – (2) can be expressed as:</p>
      <p>(, ) = (, )&amp;(, &amp;) + (, )
Where //(, ) = {) (, )(, )} = G</p>
      <p>{|!(, )|$}
{$∗(, )!(, )}
{!∗(, )$(, )}
{|$(, )|$}</p>
      <p>I
(⬚)) is conjugate operator.</p>
      <p>By using recursive equation, the auto – cross power spectral densities can be calculated as:
/"/"(, ) = /"/" (,  − 1) + (1 − )1∗(, )1(, )</p>
      <p>/"/#(, ) = /"/#(, ) + (1 − )1∗(, )"(, )
Where  is the smoothing parameter, which in the range {0 … 1}.</p>
      <p>In several speech applications, because of the complex and annoying environment, the
microphone mismatches, the different microphone sensitivities, the error of estimation of steering
vector, the displacement of MA or undetermined environmental factors, the MVDR beamformer’s
performance often corrupted. To overcome this drawback, the author proposed using spectral mask</p>
      <p>The constrained criteria of MVDR beamformer assumes that minimum the total output noise
power while preserving the speech data without distortion. The formulation of MVDR beamformer
can be defined as:</p>
      <p>(, )) (, )**(, )(, )  ) (, )&amp;(, &amp;) = 1
And the formulation of MVDR beamformer’s coefficient yields:</p>
      <p>Unfortunately, the spectral matrix of noise is not available and still a challenging task in almost
acoustic equipment. Therefore, the matrix of observed microphone array signals used instead of. The
final optimum coefficient can be derived as:
+,-.(, ) =</p>
      <p>%*!*(, )&amp;(, &amp;)
)&amp; (, &amp;)%*!*(, )&amp;(, &amp;)
+,-.(, ) =</p>
      <p>/%/! (, )&amp;(, &amp;)
)&amp; (, &amp;)/%/! (, )&amp;(, &amp;)
(3)
(4)
(5)
(6)
(7)
(8)
to suppress the speech component in observed MA signals to enhance beamforming’s evaluation.
The author uses ANN for training to obtain an effective spectral mask.</p>
    </sec>
    <sec id="sec-3">
      <title>3. The using of spectral mask and multilayer perceptrons for training data</title>
      <sec id="sec-3-1">
        <title>3.1. The spectral mask</title>
        <p>The author’s ideal is using the spectral mask, which based on the temporal the signal - to - noise
ratio Q(, ). The proposed spectral mask yields as:
*(, ) =</p>
        <p>1
1 + Q(, )</p>
        <p>From the captured MA signals, the relation between two microphone signals can represented by
the coherence /$/%(, ) = /$/%(, )VT/$/$(, ) × /%/%(, ).</p>
        <p>In realistic recording environment, this coherence derived with the information presence of
speech component and noise as the following equation:
/$/%(, ) =</p>
        <p>Q(, )
1 + Q(, )
"$#! +</p>
        <p>1
1 + Q(, ) *
Where * = 1 in coherent noise field, and * = (')⁄' in diffuse noise field.
Substituting (9), (10) can be rewritten as:</p>
        <p>/$/%(, ) = X1 − *(, )Y"$#! + *(, )*
And:</p>
        <p>Before implementing MVDR beamformer, the observed MA signals multiplied with the spectral
mask as the following way:
(10)
(11)
(12)
(13)
(14)
(9)
*(, ) =
/$/%(, ) − "$#!</p>
        <p>* − "$#!
Z!(, ) = !(, ) × *(, )
Z$(, ) = $(, ) × *(, )</p>
      </sec>
      <sec id="sec-3-2">
        <title>3.2. Multilayer perceptrons</title>
        <p>In a living room, a dual - microphone system with 5 cm was used to simulate audio data with
sampling frequency 16  with added coherent noise. The distance from speaker to the axis of
DMA2 is  = 3(). The direction of arrival of interest useful signal is &amp; = 90(). From TIMIT
database speech, the author uses 150 randomly selected sentences for recording and training data to
obtain proper coherence /$/%(, ). The purpose of multilayer perceptrons (MLP) is suppress
incoherent, diffuse noise field and non-directional noise to achieve microphone array signal, which
contain the only coherence noise.</p>
        <p>With the observed microphone array signal, the temporal coherence b/$/%(, ) between two
mounted microphones can be computed. With  = 512, b/$/%(, ) is the vector input of MLP.
 = [/$/%(!, ) ⋯ /$/%(233( , )]( with !, … , 233( is the frequency band. The  output
layer values are the promising coherence.</p>
        <p>The value of th node of the output layer can be expressed as:
45(, ) =  jk
$233(
"8!</p>
        <p>233(
4($") j k "(1!)15 + "('!)m + 4($')m
18!
(15)
Where (. ) is the sigmoid function, "('!), 4($') are the bias weights, "(1!) mean the weight for
($) denotes the coefficient for the output of the th hidden
the th input by hidden layer neutron , 4"
layer neuron by the th node of the output layer, and  mean the set of coefficients.</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>4. Experiment results</title>
      <p>The purpose of the conducted experiment is illustrating the effectiveness of the author’s spectral
mask in increasing the perceptual metric listener, the speech intelligibility and the speech quality of
the MVDR beamformer’s output signal. The advantage of spectral mask, which suppresses the speech
component to pass only the noise component for MVDR beamformer’s input, according to the theory
of MVDR beamformer. For capturing the clean speech data, these parameters  = 512, overlap
50% were set. The observed MA signals can be depicted in Figure 6.</p>
      <p>By using the traditional MVDR beamformer, the output signal yields as:</p>
      <p>By applying the proposed spectral mask, the processed signal can be derived in Figure 8.</p>
      <p>An objective measurement [21] is used for calculating the speech quality of the original MA
signals, the processed signals by MVDR beamformer and the author’s suggested technique. NIST
(National Institute of Standards and Technology) STNR (Signal-To-Noise Ratio) [22] is an efficient
method for measuring the SNR based on the estimation of sequential Gaussian mixture. WADA
(Waveform Amplitude Distribution Analysis) SNR [23] based on the model of additive Gaussian
noise signal and Gamma distribution of useful speech signal of talker to compute the SNR.</p>
      <p>NIST SNR
WADA SNR
6.2
4.8</p>
      <sec id="sec-4-1">
        <title>MVDR Beamformer 20.2 20.8</title>
      </sec>
      <sec id="sec-4-2">
        <title>The proposed spectral mask 25.7 26.8</title>
        <p>Table 1 and Figure 9 compare the speech quality and energy of microphone array signal, the
processed signal by MVDR beamformer and using the spectral mask. The effectiveness of the
author’s suggested method has been confirmed with increasing the speech quality from 5.5 dB to 6.0
db and reducing speech distortion to 6.0 dB. Because of the original constrained criteria of MVDR
beamformer is using the only noise component, therefore, the author’s proposed technique blocked
the speech component at the observed MA signals, which contains the only noise component for
ideally performing the MVDR beamformer to preserving the target speaker while suppressing
background noise.</p>
        <p>Due to the microphone mismatches, the different microphone sensitivities, the displacement of
MA, the error of estimation of preferred steering vector or non-directional noise, MVDR
beamformer’s evaluation often corrupted. Using spectral mask, which is based on training data from
the ANN is an optimum solution for enhancing beamforming’s performance for recovering the clean
speech data while removing interference, background noise. The author’s structure of ANN can be
integrated into multi-channel system for resolving other complicated problems.</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>5. Conclusion</title>
      <p>Speech enhancement plays an important role in almost acoustic speech applications to enhance
perceptual listener, speech intelligibility and speech quality while removing interference,
background noise or competing talker. In this article, the author suggested a structure of ANN for
forming an effective spectral mask, which multiplies with the observed microphone array signals to
enhance MVDR beamformer’s evaluation and extract the desired speech signal. The obtained result
has confirmed the effectiveness of the author’s spectral mask, which was trained by the ANN in
improvement of considered beamformer. The promising result has shown that the speech quality in
the term of signal - to - noise increased from 5.5 (dB) to 6.0 (dB) and reduced speech distortion to 5.5
(dB). The proposed configuration of ANN can be integrated into a multi-channel system for solving
complicated problems.
[20] J. Oruh, S. Viriri and A. Adegun, "Long Short-Term Memory Recurrent Neural Network for
Automatic Speech Recognition," in IEEE Access, vol. 10, pp. 30069-30079, 2022, doi:
10.1109/ACCESS.2022.3159339.
[21] SNRVAD. [Online]. Available: https://labrosa.ee.columbia.edu/projects/snreval/.
[22] A. W. Rix, J. G. Beerends, M. P. Hollier and A. P. Hekstra, "Perceptual evaluation of speech
quality - A new method for speech quality assessment," Proc. Int. Conf. Acoust., Speech Signal
Process. (ICASSP), 2001, pp. 749–752.
[23] C. Kim and R.M. Stern, "Robust signal-to-noise ratio estimation based on waveform amplitude
distribution analysis," Proc. Interspeech 2008, 2598-2601, doi: 10.21437/Interspeech.2008-644.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>S.</given-names>
            <surname>SeyedShah KaramFard</surname>
          </string-name>
          , B. Mohammadzadeh Asl, “
          <article-title>Fast Delay-Multiply-and-Sum Beamformer: Application to Confocal Microwave Imaging,” IEEE Antennas and Wireless Propagation Letters</article-title>
          , vol.
          <volume>19</volume>
          , no.
          <issue>1</issue>
          , pp.
          <fpage>14</fpage>
          -
          <lpage>18</lpage>
          , Jan.
          <year>2020</year>
          , doi: 10.1109/LAWP.
          <year>2019</year>
          .
          <volume>2951575</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>P. K.</given-names>
            <surname>Chodingala</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S. S.</given-names>
            <surname>Chaturvedi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A. T.</given-names>
            <surname>Patil</surname>
          </string-name>
          and
          <string-name>
            <given-names>H. A.</given-names>
            <surname>Patil</surname>
          </string-name>
          , “
          <article-title>Robustness of DAS Beamformer Over MVDR for Replay Attack Detection On Voice Assistants</article-title>
          ,”
          <source>2022 IEEE International Conference on Signal Processing and Communications (SPCOM)</source>
          , Bangalore, India,
          <year>2022</year>
          , pp.
          <fpage>1</fpage>
          -
          <lpage>5</lpage>
          , doi: 10.1109/SPCOM55316.
          <year>2022</year>
          .
          <volume>9840757</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>M.</given-names>
            <surname>Papez</surname>
          </string-name>
          and
          <string-name>
            <given-names>K.</given-names>
            <surname>Vlcek</surname>
          </string-name>
          , “
          <source>Recognition System Based on DTW and DAS Beamforming,” 2015 Second International Conference on Mathematics and Computers in Sciences and in Industry (MCSI)</source>
          , Sliema, Malta,
          <year>2015</year>
          , pp.
          <fpage>176</fpage>
          -
          <lpage>181</lpage>
          , doi: 10.1109/
          <string-name>
            <surname>MCSI</surname>
          </string-name>
          .
          <year>2015</year>
          .
          <volume>49</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>X.</given-names>
            <surname>Zhao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Luo</surname>
          </string-name>
          , G. Huang,
          <string-name>
            <given-names>J.</given-names>
            <surname>Chen</surname>
          </string-name>
          and
          <string-name>
            <given-names>J.</given-names>
            <surname>Benesty</surname>
          </string-name>
          ,
          <article-title>"Differential Beamforming with Null Constraints for Spherical Microphone Arrays,"</article-title>
          <source>ICASSP 2024 - 2024 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)</source>
          , Seoul, Korea, Republic of,
          <year>2024</year>
          , pp.
          <fpage>776</fpage>
          -
          <lpage>780</lpage>
          , doi: 10.1109/ICASSP48485.
          <year>2024</year>
          .
          <volume>10446768</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>J.</given-names>
            <surname>Liang</surname>
          </string-name>
          and
          <string-name>
            <given-names>Q.</given-names>
            <surname>Zeng</surname>
          </string-name>
          ,
          <article-title>"</article-title>
          <source>Improved Spectral Subtraction Based on Second-Order Differential Array and Phase Spectrum Compensation," 2023 3rd International Symposium on Computer Technology and Information Science (ISCTIS)</source>
          , Chengdu, China,
          <year>2023</year>
          , pp.
          <fpage>658</fpage>
          -
          <lpage>661</lpage>
          , doi: 10.1109/ISCTIS58954.
          <year>2023</year>
          .
          <volume>10212993</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>X.</given-names>
            <surname>Luo</surname>
          </string-name>
          , G. Huang,
          <string-name>
            <given-names>J.</given-names>
            <surname>Chen</surname>
          </string-name>
          and
          <string-name>
            <given-names>J.</given-names>
            <surname>Benesty</surname>
          </string-name>
          ,
          <article-title>"On the Design of Robust Differential Beamformers with Uniform Circular Microphone Arrays,"</article-title>
          <source>2023 31st European Signal Processing Conference (EUSIPCO)</source>
          , Helsinki, Finland,
          <year>2023</year>
          , pp.
          <fpage>1</fpage>
          -
          <lpage>5</lpage>
          , doi: 10.23919/EUSIPCO58844.
          <year>2023</year>
          .
          <volume>10289970</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>J.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Yang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Guo</surname>
          </string-name>
          and
          <string-name>
            <given-names>J.</given-names>
            <surname>Yang</surname>
          </string-name>
          ,
          <article-title>"Robust Adaptation Control for Generalized Sidelobe Canceller with Time-Varying Gaussian Source Model,"</article-title>
          <source>2023 31st European Signal Processing Conference (EUSIPCO)</source>
          , Helsinki, Finland,
          <year>2023</year>
          , pp.
          <fpage>16</fpage>
          -
          <lpage>20</lpage>
          , doi: 10.23919/EUSIPCO58844.
          <year>2023</year>
          .
          <volume>10289801</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>S.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Q.</given-names>
            <surname>Liu</surname>
          </string-name>
          and
          <string-name>
            <given-names>W.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <article-title>"Generalized Sidelobe Canceller with Variable Step-Size Least Mean Square Algorithm Controlled by Signal-to-</article-title>
          <string-name>
            <surname>Noise</surname>
            <given-names>Ratio</given-names>
          </string-name>
          ,
          <article-title>"</article-title>
          <source>2022 5th International Conference on Data Science and Information Technology (DSIT)</source>
          , Shanghai, China,
          <year>2022</year>
          , pp.
          <fpage>1</fpage>
          -
          <lpage>6</lpage>
          , doi: 10.1109/DSIT55514.
          <year>2022</year>
          .
          <volume>9943855</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>J.</given-names>
            <surname>Park</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Hong</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. -W.</given-names>
            <surname>Choi</surname>
          </string-name>
          and
          <string-name>
            <given-names>M.</given-names>
            <surname>Hahn</surname>
          </string-name>
          ,
          <article-title>"Determinant-Based Generalized Sidelobe Canceller for Dual-Sensor Noise Reduction,"</article-title>
          <source>in IEEE Sensors Journal</source>
          , vol.
          <volume>22</volume>
          , no.
          <issue>9</issue>
          , pp.
          <fpage>8858</fpage>
          -
          <issue>8868</issue>
          ,
          <fpage>1</fpage>
          <lpage>May1</lpage>
          ,
          <year>2022</year>
          , doi: 10.1109/JSEN.
          <year>2022</year>
          .
          <volume>3162619</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <surname>As'ad H</surname>
          </string-name>
          .,
          <string-name>
            <surname>Bouchard</surname>
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kamkar-Parsi H. A Robust Target</surname>
          </string-name>
          <article-title>Linearly Constrained Minimum Variance Beamformer With Spatial Cues Preservation for Binaural Hearing Aids</article-title>
          .
          <source>IEEE/ACM Transactions on Audio, Speech, and Language Processing</source>
          , vol.
          <volume>27</volume>
          , no.
          <issue>10</issue>
          , pp.
          <fpage>1549</fpage>
          -
          <lpage>1563</lpage>
          , Oct.
          <year>2019</year>
          , doi: 10.1109/TASLP.
          <year>2019</year>
          .
          <volume>2924321</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <surname>Sherson</surname>
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kleijn</surname>
            <given-names>W. B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Heusdens</surname>
            <given-names>R.</given-names>
          </string-name>
          <article-title>A distributed algorithm for robust LCMV beamforming /</article-title>
          / Proc 2016 IEEE International Conference on Acoustics,
          <source>Speech and Signal Processing (ICASSP)</source>
          , Shanghai, China,
          <year>2016</year>
          , pp.
          <fpage>101</fpage>
          -
          <lpage>105</lpage>
          , doi: 10.1109/ICASSP.
          <year>2016</year>
          .
          <volume>7471645</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <surname>Chazan</surname>
            <given-names>S. E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Goldberger</surname>
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gannot</surname>
            <given-names>S.</given-names>
          </string-name>
          <article-title>DNN-Based Concurrent Speakers Detector and its Application to Speaker Extraction with LCMV Beamforming /</article-title>
          / Proc 2018 IEEE International Conference on Acoustics,
          <source>Speech and Signal Processing (ICASSP)</source>
          , Calgary, AB, Canada,
          <year>2018</year>
          , pp.
          <fpage>6712</fpage>
          -
          <lpage>6716</lpage>
          , doi: 10.1109/ICASSP.
          <year>2018</year>
          .
          <volume>8462407</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <surname>P. -O. Lagacé</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          <string-name>
            <surname>Ferland</surname>
            and
            <given-names>F.</given-names>
          </string-name>
          <string-name>
            <surname>Grondin</surname>
          </string-name>
          ,
          <article-title>"Ego-Noise Reduction of a Mobile Robot Using Noise Spatial Covariance Matrix Learning and Minimum Variance Distortionless Response,"</article-title>
          <source>2023 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS)</source>
          , Detroit, MI, USA,
          <year>2023</year>
          , pp.
          <fpage>3533</fpage>
          -
          <lpage>3538</lpage>
          , doi: 10.1109/IROS55552.
          <year>2023</year>
          .
          <volume>10342193</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <given-names>S.</given-names>
            <surname>Yadav</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Pal</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Kumar</surname>
          </string-name>
          and
          <string-name>
            <given-names>M.</given-names>
            <surname>Aggarwal</surname>
          </string-name>
          ,
          <article-title>"Study of MVDR Beamformer for a single Acoustic Vector Sensor,"</article-title>
          <source>2023 International Symposium on Ocean Technology (SYMPOL)</source>
          , Kochi, India,
          <year>2023</year>
          , pp.
          <fpage>1</fpage>
          -
          <lpage>6</lpage>
          , doi: 10.1109/SYMPOL59195.
          <year>2023</year>
          .
          <volume>10455004</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <given-names>S.</given-names>
            <surname>Erdim</surname>
          </string-name>
          and
          <string-name>
            <given-names>J. R.</given-names>
            <surname>Buck</surname>
          </string-name>
          ,
          <article-title>"Mitigating Multiple Moving Interferers With the Hybrid Double Zero MVDR Beamformer,"</article-title>
          <source>in IEEE Access</source>
          , vol.
          <volume>12</volume>
          , pp.
          <fpage>111206</fpage>
          -
          <lpage>111217</lpage>
          ,
          <year>2024</year>
          , doi: 10.1109/ACCESS.
          <year>2024</year>
          .
          <volume>3437749</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [16]
          <string-name>
            <given-names>D.</given-names>
            <surname>Skariah</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Rajan</surname>
          </string-name>
          and
          <string-name>
            <given-names>J.</given-names>
            <surname>Thomas</surname>
          </string-name>
          ,
          <article-title>"CycleGAN based Speech Enhancement Using Time Frequency Masking,"</article-title>
          2023 Second International Conference on Electrical, Electronics, Information and Communication
          <string-name>
            <surname>Technologies</surname>
          </string-name>
          (ICEEICT), Trichirappalli, India,
          <year>2023</year>
          , pp.
          <fpage>1</fpage>
          -
          <lpage>6</lpage>
          , doi: 10.1109/ICEEICT56924.
          <year>2023</year>
          .
          <volume>10157492</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [17]
          <string-name>
            <given-names>L.</given-names>
            <surname>Wang</surname>
          </string-name>
          and
          <string-name>
            <given-names>F.</given-names>
            <surname>Chen</surname>
          </string-name>
          ,
          <article-title>"EEG-based Auditory Attention Detection with Estimated Speech Sources Separated from an Ideal - binary -</article-title>
          masking
          <string-name>
            <surname>Process</surname>
          </string-name>
          ,
          <article-title>" 2022 Asia-Pacific Signal and Information Processing Association Annual Summit and Conference (APSIPA ASC)</article-title>
          ,
          <source>Chiang Mai, Thailand</source>
          ,
          <year>2022</year>
          , pp.
          <fpage>1545</fpage>
          -
          <lpage>1549</lpage>
          , doi: 10.23919/APSIPAASC55919.
          <year>2022</year>
          .
          <volume>9980112</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          [18]
          <string-name>
            <given-names>A.</given-names>
            <surname>Sack</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W.</given-names>
            <surname>Jiang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Perlmutter</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Salanevich</surname>
          </string-name>
          and
          <string-name>
            <given-names>D.</given-names>
            <surname>Needell</surname>
          </string-name>
          ,
          <article-title>"</article-title>
          <source>On Audio Enhancement via Online Non-Negative Matrix Factorization," 2022 56th Annual Conference on Information Sciences and Systems (CISS)</source>
          , Princeton, NJ, USA,
          <year>2022</year>
          , pp.
          <fpage>287</fpage>
          -
          <lpage>291</lpage>
          , doi: 10.1109/CISS53076.
          <year>2022</year>
          .
          <volume>9751157</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          [19]
          <string-name>
            <given-names>J.</given-names>
            <surname>Chen</surname>
          </string-name>
          , C. Liu,
          <string-name>
            <given-names>J.</given-names>
            <surname>Xie</surname>
          </string-name>
          and
          <string-name>
            <given-names>N.</given-names>
            <surname>Huang</surname>
          </string-name>
          , “
          <string-name>
            <surname>Time-Frequency Mask-Aware Bidirectional</surname>
            <given-names>LSTM</given-names>
          </string-name>
          :
          <article-title>A Deep Learning Approach for Underwater Acoustic Signal Separation</article-title>
          ,”
          <source>Sensors</source>
          <year>2022</year>
          ,
          <volume>22</volume>
          (
          <issue>15</issue>
          ),
          <volume>5598</volume>
          ; https://doi.org/10.3390/s22155598.
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>