<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Speech Source Separation Based on Dual - Microphone System</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Quan Trong The</string-name>
          <email>theqt@nongnghiepso.com</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Digital Agriculture Cooperative</institution>
          ,
          <addr-line>Cau Giay, Ha Noi</addr-line>
          ,
          <country country="VN">Viet Nam</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2023</year>
      </pub-date>
      <fpage>21</fpage>
      <lpage>23</lpage>
      <abstract>
        <p>The fundamental intention of extracting the desired target speech source is saving the speech components of useful interesting signal at a certain direction while suppressing interference and unwanted signal, which come from other directions. Microphone array is one of the most widely technology used for improving speech enhancement in scenarios with multiple speech sources. Microphone arrays use a priori information of the direction of arrival to extract wanted signal, while eliminating background noise or different speech signal. In this article, dualmicrophone system was used for recording two speech sources at two opposite directions; the authors proposed a method, which developed from the author's previous work for extracting the correct only target speaker at each direction, uses the ratio of power between two directions. The experiments have proven the capability of suggested method in comparison with the author's previous work.</p>
      </abstract>
      <kwd-group>
        <kwd>Microphone arrays</kwd>
        <kwd>dual - microphone system</kwd>
        <kwd>power</kwd>
        <kwd>direction</kwd>
        <kwd>speech source separation</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>One of the most complicated tasks in speech communication system is speech signal acquisition
without distortion and suppression background noise. This challenging problem still exist in real-life
and requires scholar to investigate and solve it in such applications: speech recognition, hearing aids,
telephony, human-machine interface, hands-free instrument. High quality of obtained output signal
processing system is prerequisite order to restrict the effect of surrounding noise from a mixture of
noisy signals and extract the clean desired speech (figure 1).</p>
      <p>2023 Copyright for this paper by its authors.</p>
      <p>
        Ambiguous speech enhancement is more clearly in a complex complicated acoustic surrounding
environment when using a single-microphone approach, due to observed signal is corrupted by
unwanted interference, diff use noise, coherence noise or complex condition. In order to obtain high
speech quality and eliminate background noise, the technology of microphone array (MA) [
        <xref ref-type="bibr" rid="ref1 ref2 ref3 ref4">1-4</xref>
        ] has
been developed for extracting satisfactory useful signal. This approach uses spatial information, such
as distributed geometry MA or direction of arrival of desired signal for enhancing one signal in certain
direction while suppressing other noise. Therefore, exploiting the properties of MA and environment
background noise are the mainstream of current technology beamforming.
      </p>
      <p>At present time, technology beamforming, which use microphone array, has three approaches, such
as: fixed beamforming with fixed coefficient weights, adaptive beamforming algorithm, that adapts to
observed array signals and post-filtering, which improve the speech enhancement. Adaptive
beamforming and post-filtering have entrained many scholars to put some improvements for enhancing
performance. These technologies have ability of saving desired signal and suppressing background
noise while obtaining high speech quality. The combination with a justifiable post-filtering allows
achieving an acceptable result in performance signal processing system and satisfied speech
enhancement (figure 2).</p>
      <p>Because of possibly rapidly changing position of speaker, microphone array mismatch, or
nonstationary acoustical environment; microphone array algorithm doesn’t sufficiently deliver effective
performance in terms of speech distortion, signal-to-noise ratio.</p>
      <p>
        Dual-microphone array (DMA2) [
        <xref ref-type="bibr" rid="ref10 ref8 ref9">8-10</xref>
        ], that is one of the most widely MA, has been exploited for
many purposes, such as speech enhancement, estimation of direction of arrival (DOA), reverberation.
Due to it’s compact, DMA2 exists in almost speech equipment. And problem of separation desired
target speaker is an essence reality approach. In DMA2, the differential algorithm is often widely used
to focusing on one certain direction along axis DMA2. Despite its simplicity, low computational
performance, easily null forming towards noise source, the acoustic environments are often very
complex and noisy. The noisy scenario severely reduces the effectiveness or speech quality of output
signal. Therefore, the goal in this study is presented a new possible post-filtering to clearly pick-up
speech in a certain direction.
      </p>
      <p>In this paper, the authors considered problem of extracting desired target speaker when using DMA2
for recording two speakers, which their positions opposite each other. This work is the development of
the author’s paper for purpose of separation of target speaker. For more robust adaptive beamforming,
the authors proposed a post-filtering, which based on estimation power each desired speech, for saving
necessary speech while eliminating other. The experiments have illustrated the effectiveness and the
ability of the proposed post-filtering in comparison with the author’s previous work. The observed
amplitude and spectrogram of processed signals has proven the capability and the received enhancement
of the proposed additional post-filtering (figure 3).</p>
    </sec>
    <sec id="sec-2">
      <title>2. Separation of the Speakers</title>
      <p>
        Natural Speech Processing (NSP) [
        <xref ref-type="bibr" rid="ref18 ref19">18-19</xref>
        ] methods has been utilized Automatic meeting
transcription to aim extracting desired speaker in a long-duration recording speech signal. This
challenging problem is considerable, because of using Speaker Identification and Diarization [
        <xref ref-type="bibr" rid="ref20 ref21">20-21</xref>
        ]
and Automatic Speech Recognition (ASR) [
        <xref ref-type="bibr" rid="ref22 ref23 ref24">22-24</xref>
        ] technique to solving commonly speech applications.
The existing available commercial method for speaker diarization can be categorized into two mainly
approaches: (1) technology for distant speech acquisition and processing; (2) technology for
closetalking processing. Distant acquisition processing tends to the scenarios, where a microphone array are
distributed following an optimal geometry, situated at a long distance to record all attending desired
speakers, and consequently interfering speech signals incoming from other directions and surrounding
noise. Close-talking microphone uses each speech signal per each speaker to process and enhance the
recording signal.
      </p>
      <p>On the market, the promising automatic meeting transcription method are available, which ensure
several perspective solutions for solving separation of speakers. Close-talking speech technique require
less complicated and computational performance than distant speech processing. However, this
situation almost is a trivial condition, that each attending speaker must be equipped each personal audio
device.</p>
      <p>On the other hand, distant speech processing is a useful method to pose complex task for speaker
diarization. All negatived unwanted effect: mixture of different signals; attenuation due to microphone
mismatch, microphone sensitivities, long distant; reverberation and influence of environmental factors
on the accuracy, effectiveness of separation of speakers. The state-of-the-art of in separation of speaker
has many attractive achievements on speech acquisition quality and recognition, there still exits
limitation of prominent extraction methods fitted for separating. An Artificial Neural Network (ANN)
is used for training model by using acoustical, lexical features. Distribution of microphone array also
ensure an additional technique to eliminate background noise and extract spatial diversity of target
speaker. The biometric properties provide a prominent approach of research.</p>
      <p>The expected application of the proposed method in this paper to extract useful speech signal, that
from a determined direction of target speaker in scenario of presence of two active speakers. The
conversation is recorded by using a DMA2, which place in a table or a desk. The recording audio device
DMA2 is appropriate for discussion, interviews, service provision; because of it’s compact, convenient,
and low complexity.</p>
      <p>The necessary task of speech enhancement by using DMA2 in realistic situation is extraction and
estimation speech or target speaker, which incoming from a certain direction, each active speaker per
each utterance. In this article, the authors primarily address the conversation, which involves speech of
two active speakers, their speech doesn’t appear simultaneously but continuously.</p>
      <p>The rest of this article is organized as follows: the next section is introduced DMA2 and differential
algorithm, which is used to forming beampattern to desired speech source. In section IV, the authors
suggested using the estimation of speech power to form an additional post-filtering. Section V, the
scheme and comparison’s performance are shown. Final section VI, conclusion, and the author’s future
work.</p>
    </sec>
    <sec id="sec-3">
      <title>3. Differential microphone array</title>
      <p>
        Differential algorithm [
        <xref ref-type="bibr" rid="ref11 ref12 ref5 ref6 ref7">5-7,11-12</xref>
        ] is one of the most popular algorithms for processing microphone
array signals. A structure of DMA2 is presented in Figure 4. Two microphones arranged toward two
different speakers. DMA2 allow steering beampattern to a certain direction for obtaining high diversity
and suppressing signals, interference, or noise, which come from other directions. Speech enhancement
in DMA2 [
        <xref ref-type="bibr" rid="ref13 ref14 ref15">13-15</xref>
        ] is easily implemented cause of incorporating a suitable equalizer or post-filtering for
attenuating background noise and saving the desired speech source.
      </p>
      <p>Here, the sound speech is  (343  / ), distane between microphones is  ,  0 =  / , useful signal
or speech source propagates with direction   ,   =   0   ,  ,  are the current frequency and
current frame. Two noisy observed microphone signals in STFT-domain are represented as:
 1( ,  ) =  ( ,  )   
 2( ,  ) =  ( ,  ) −  
(1)
(2)</p>
      <p>The adjustable value  is added for controlling null - pattern when extracting the useful signal. So
DMA2 is used for reducing surrounding noise and extracting the desired signal. Two output signals,
which derived from operator subtraction signal between two recored microphone signals, are calculated
by the following equation:
 1
 2
=  ( , )
= − ( , )
( , ) =  1( , )−  2( , ) −
( , ) =  2( , )−  1( , ) −
( 20 (
+</p>
      <p>))

 0
2
2
( 20 (
−</p>
      <p>))

 0
Equations (3,4,5,6) yield two steered beampattern, which can be expressed (figure 5):
 1( , ) = | 1
 2( , ) = | 2
 ( , ) | = |
 ( , ) | = |
( 20 (
( 20 (
+
−


 0
 0
))|
))|
(3)
(4)
(5)
(6)
(7)
(8)
(9)</p>
      <p>
        An additional equalizer [
        <xref ref-type="bibr" rid="ref17">17</xref>
        ] is added for improving the received signal:
  ( ) =  (2  )
{
6
1
 
1
0
0  &lt;  ≤ 200( )
200 &lt;  ≤ 
 &lt;  ≤ 2 ∗
      </p>
      <p>2 ∗  &lt; 
equalizer ensures deriving desired signal.</p>
      <sec id="sec-3-1">
        <title>Where</title>
        <p>= 4∗1 0. The value of   ( ) is limited with a determined threshold 12( ). This</p>
      </sec>
      <sec id="sec-3-2">
        <title>So the output two desired target speakers are:</title>
        <p>1( , ) =  1
 2( , ) =  2
( , )×   ( )
( , )×   ( )</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>4. The proposed Post-Filtering</title>
      <p>The designed post-filtering is proposed for microphone array system for increasing implementation
in various types of background noise. Post-filtering method often uses a priori information of
environment’s characteristic or coherence between microphones. The authors proposed a post-filtering,
which uses a determined power speech at a determined direction for forming post-filtering.</p>
      <p>We have two directions: 
=  1 = 0
0 and 
=  2 = 1800. With certain two steering vector
  ( , 1
) = [    1
 −   1] and   ( , 2
) = [    2
 −   2] , where   1 =   0
  0
( 2), we can easily calculated power speech at two directions as the following equations:
( 1),  2 =
Where   ( , ) =  { ( , ) ∗( , )}can be determined as:
  1 1( , ) = (  ( , 1) −1( ,  )  ( , 1))
  2 2( , ) = (  ( , 2) −1( ,  )  ( , 2))
−1
−1
{  1 1</p>
      <p>( , )∗ 1.001
  2 1
Where  is the smoothing parameter, which in the range {0…1}.</p>
      <p>The suggested post-filtering for enhancing speech quality is formulated as:
And the final signals of interest are given by:
  1( , ) =
  2( , ) =</p>
      <p>1 1( , )
  1 1( , )+   2 2( , )</p>
      <p>2 2( , )
  1 1( , )+   2 2( , )




1( , ) =  1( , )×</p>
      <p>1( , )
2( , ) =  2( , )×   2( , )
(10)
(11)
(12)
(13)
(14)
(15)
(16)
(17)
(18)
(19)</p>
      <p>
        In the next section, the authors will compare the effectiveness of using the proposed post-filtering
(DIF-PF-Pro) and the known algorithm [
        <xref ref-type="bibr" rid="ref17">17</xref>
        ].
      </p>
    </sec>
    <sec id="sec-5">
      <title>5. Experiments</title>
      <p>In this section, a DMA2 system was placed in an anechoic room to capture two microphone array
signals from two independent speakers, which stand opposite each other. The distance from speaker to
DMA2 is 3 ( ), the distance between microphones is  = 2,5 (  ). So, the directions of two interest
useful signals are opposite (figure 6).</p>
      <p>
        All two noisy microphone array signals are sampled 16 ( ). For processing the method [
        <xref ref-type="bibr" rid="ref17">17</xref>
        ] and
suggested method (DIF-PF-Pro), a Hamming window, overlap 50%,  = 512 and a constant
smoothing parameter  = 0,1 were chosen to compute PSD and additive post-filtering.
      </p>
      <p>
        For estimating speech quality, an objective measure NIST SNR [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ] was utilized to verify the
effectiveness of the proposed method (DIF-PF-Pro) in comparison with known differential algorithm
[
        <xref ref-type="bibr" rid="ref17">17</xref>
        ] (figure 7).
      </p>
      <p>
        Figure 7 shows the captured microphone array signals with two directions of two desired speakers.
The next figure presents the obtained signal by using method [
        <xref ref-type="bibr" rid="ref17">17</xref>
        ] for extracting speech of speaker, who
stand at direction  2 = 180 ( ). As we can see that, the large amount of first speaker’s speech still
exits.
      </p>
      <p>By using the suggested method, the target useful of second speaker was saved while reducing the
amount of first speaker’s speech. Result of DIF-PF-Pro is shown in Figure 9, that verifies the
effectiveness of the proposed additive post-filtering.</p>
    </sec>
    <sec id="sec-6">
      <title>6. Conclusion</title>
      <p>The goal of microphone array is to increase the evaluation of speech enhancement system with the
purpose of saving useful target speech signal while eliminating interference and noise, which come
from other directions. Therefore, improvement of MA’s performance is prerequisite order in front-end
almost speech applications. A lot of noise reduction and preserving speech signal methods are
investigated, evaluated. Because of undetermined noise source, beamforming technique can not
sufficiently exclude almost all noise component; so, an additive post-filtering can help further keep
speech in desired direction and reduce considered noise component.</p>
      <p>An important essence in signal processing is exploiting’ characteristic of dual-microphone system
to obtain useful speech target speaker and improve speech enhancement. The dual-microphone system
always provides a high noise reduction, high spatial diversity for extracting desired speech in noisy
complicated conditions. This paper proposes a new technology post-filtering to improve the capability
of extracting the correct desired target speaker. With priori information of direction of arrival, the
suggested method has proven in increasing performance dual-microphone system in comparison with
other work. While the components of useful signal at a certain direction are saved, and the other speech
source is suppressed. Dual - microphone is a basic element in microphone array, so the proposed method
can be further integrated into system multi-microphone for resolving more complicated problem speech
enhancement.</p>
    </sec>
    <sec id="sec-7">
      <title>7. References</title>
      <p>[26] A. Sarafnia, M. O. Ahmad, M. N. Swamy, A Beam Steerable Speaker Tracking-based First-order
Differential Microphone Array, in: Proc 2023 21st IEEE Interregional NEWCAS Conference
(NEWCAS), Edinburgh, United Kingdom, 2023, pp. 1-5, doi:
10.1109/NEWCAS57931.2023.10198069.
[27] T. Liu, Z. Lu, T. Fei, A Hybrid Reverberation Model and Its Application to Joint Speech
Dereverberation and Separation. IEEE/ACM Transactions on Audio, Speech, and Language
Processing, 31 (2023) 3000-3014. doi: 10.1109/TASLP.2023.3301227.
[28] J. Jin, J. Benesty, J. Chen, G. Huang, Differential Beamforming From a Geometric Perspective.</p>
      <p>IEEE/ACM Transactions on Audio, Speech, and Language Processing, 31 (2023) 3042-3054. doi:
10.1109/TASLP.2023.3301245.
[29] Z. Wang, W. Zou, H. Su, Y. Guo, D. Li, Multiple Sound Source Localization Exploiting Robot
Motion and Approaching Control. IEEE Transactions on Instrumentation and Measurement,
(2021). doi: 10.1109/TIM.2023.3298406.
[30] D. M. Baruah, S. Senchowa, B. K. Das, Impact of Reference Microphone Selection on Binaural
Cue Preservation for Directional Sources in Multi-microphone Hearing Aids, in: Proc 2023 IEEE
Guwahati Subsection Conference (GCON), Guwahati, India, 2023, pp. 1-5. doi:
10.1109/GCON58516.2023.10183407.
[31] G. Li, Audio-Visual End-to-End Multi-Channel Speech Separation, Dereverberation and
Recognition, IEEE/ACM Transactions on Audio, Speech, and Language Processing, 31 (2017)
2707-2723. doi: 10.1109/TASLP.2023.3294705.
[32] J. Wang, H. He, Y. Yu, Y. Zhou, L. Zhang, Robust Time-Delay Estimation for Speaker
Localization Using Mutual Information Among Multiple Microphone Signals. IEEE Sensors
Journal, (2023). doi: 10.1109/JSEN.2023.3293499.
[33] S. Chen, R. Tan, Z. Wang, X. Tong, K. Li, VoiceMap: Autonomous Mapping of Microphone Array
for Voice Localization, IEEE Internet of Things Journal, (2021). doi: 10.1109/JIOT.2023.3294937.
[34] X. Fang, Y. Liu, X. Zhang, Q. Sun, Online Monitoring of Converter Station Using an Acoustic
Signal Analysis Method Based on The Mobile Microphone Array, in: Proc 2023 IEEE 6th
International Electrical and Energy Conference (CIEEC), Hefei, China, 2023, pp. 3918-3922. doi:
10.1109/CIEEC58067.2023.10166859.
[35] Y. Shen Research on the Design of Omni-Directional Mobile Robot for Sound Seeking,
Positioning and Navigation, in: Proc 2023 4th International Conference for Emerging Technology
(INCET), Belgaum, India, 2023, pp. 1-6. doi: 10.1109/INCET57972.2023.10170055.
[36] F. Yang, R. Song, A Review of Sound Source Localization Research in Three-Dimensional Space,
in: Proc 2023 IEEE 12th Data Driven Control and Learning Systems Conference (DDCLS),
Xiangtan, China, 2023, pp. 579-584. doi: 10.1109/DDCLS58216.2023.10165974.
[37] J. Wang, F. Yang, J. Yang, A General Approach to the Design of the Fractional-Order
Superdirective Beamformer, IEEE Transactions on Circuits and Systems II: Express Briefs,
(2022). doi: 10.1109/TCSII.2023.3287918.
[38] Y. Wakabayashi, K. Yamaoka, N. Ono, Sound Field Interpolation for Rotation-Invariant
Multichannel Array Signal Processing, IEEE/ACM Transactions on Audio, Speech, and Language
Processing, 31 (2023) 2286-2298. doi: 10.1109/TASLP.2023.3282098
[39] C. Zhang, J. Liu, H. Li, X. Zhang, Neural Multi-Channel and Multi-Microphone Acoustic Echo
Cancellation, IEEE/ACM Transactions on Audio, Speech, and Language Processing, 31 (2023)
2181-2192. doi: 10.1109/TASLP.2023.3282103.
[40] J. Wang, F. Yang, J. Yang, A Perspective on Fully Steerable Differential Beamformers for Circular</p>
      <p>Arrays, IEEE Signal Processing Letters, 30 (2023) 648-652. doi: 10.1109/LSP.2023.3280852.
[41] X. Zhao, G. Huang, J. Chen, J. Benesty, Design of 2D and 3D Differential Microphone Arrays
With a Multistage Framework, IEEE/ACM Transactions on Audio, Speech, and Language
Processing, 31 (2023) 2016-2031. doi: 10.1109/TASLP.2023.3278182.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1] Microphone Arrays / ed. by
          <string-name>
            <given-names>M.</given-names>
            <surname>Brandstein</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Ward</surname>
          </string-name>
          . Heidelberg, Germany: Springer-Verlag,
          <year>2001</year>
          . doi:
          <volume>10</volume>
          .1007/978-3-
          <fpage>662</fpage>
          -04619-7.
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>J.</given-names>
            <surname>Benesty</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Chen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Huang</surname>
          </string-name>
          ,
          <source>Microphone Array Signal Processing</source>
          , Berlin, Germany: SpringerVerlag,
          <year>2008</year>
          . doi:
          <volume>10</volume>
          .1007/978-3-
          <fpage>540</fpage>
          -78612-2.
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>J.</given-names>
            <surname>Benesty</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Chen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Pan</surname>
          </string-name>
          ,
          <source>Fundaments of Diff erential Beamforming</source>
          . Springer,
          <year>2016</year>
          . doi:
          <volume>10</volume>
          .1007/
          <fpage>978</fpage>
          -981-10-1046-0.
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>J.</given-names>
            <surname>Benesty</surname>
          </string-name>
          , I. Cohen,
          <string-name>
            <given-names>J.</given-names>
            <surname>Chen</surname>
          </string-name>
          ,
          <source>Fundamentals of Signal Enhancement and Array Signal Processing</source>
          , Wiley, IEEE Press,
          <year>2017</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>G.W.</given-names>
            <surname>Elko</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.-T.N.</given-names>
            <surname>Pong</surname>
          </string-name>
          ,
          <article-title>A steerable and variable fi rst-order diff erential microphone array</article-title>
          ,
          <source>in: Proc. IEEE International Conference on Acoustics, Speech, and Signal Processing (ICASSP)</source>
          .
          <year>1997</year>
          , p.
          <fpage>223</fpage>
          -
          <lpage>226</lpage>
          . doi:
          <volume>10</volume>
          .1109/ICASSP.
          <year>1997</year>
          .
          <volume>599609</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>M.</given-names>
            <surname>Buck</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Rößler</surname>
          </string-name>
          ,
          <article-title>First order diff erential microphone arrays for automotive applications</article-title>
          ,
          <source>in: Proc. 7th International Workshop on Acoustic Echo and Noise Control, IWAENC</source>
          .
          <year>2001</year>
          , pp.
          <fpage>19</fpage>
          -
          <lpage>22</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>M.</given-names>
            <surname>Buck</surname>
          </string-name>
          ,
          <article-title>Aspects of first-order diff erential microphone arrays in the presence of sensor imperfections</article-title>
          .
          <source>European Transactions on Telecommunications, 13</source>
          <volume>2</volume>
          (
          <year>2002</year>
          )
          <fpage>115</fpage>
          -
          <lpage>122</lpage>
          . doi:
          <volume>10</volume>
          .1002/ett.4460130206.
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>H.</given-names>
            <surname>Zhang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Q.</given-names>
            <surname>Fu</surname>
          </string-name>
          .,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Yan</surname>
          </string-name>
          ,
          <article-title>A frequency domain approach for speech enhancement with directionality using compact microphone array</article-title>
          ,
          <source>in: Proc. 9th Annual Conference of the International Speech Communication Association (INTERSPEECH</source>
          <year>2008</year>
          ),
          <year>2008</year>
          . pp.
          <fpage>447</fpage>
          -
          <lpage>450</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>S.</given-names>
            <surname>Fischer</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Simmer</surname>
          </string-name>
          ,
          <article-title>An adaptive microphone array for hands-free communication</article-title>
          ,
          <source>in: Proc. IWAENC-95</source>
          ,
          <year>1995</year>
          , pp.
          <fpage>1</fpage>
          -
          <lpage>4</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>J.</given-names>
            <surname>Bitzer</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Kammeyer</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Simmer</surname>
          </string-name>
          ,
          <article-title>An alternative implementation of the superdirective beamformer</article-title>
          ,
          <source>in: Proc. IEEE Workshop on Applications of Signal Processing to Audio and Acoustics (WASPLAA'99)</source>
          ,
          <year>1999</year>
          , pp.
          <fpage>7</fpage>
          -
          <lpage>9</lpage>
          . doi:
          <volume>10</volume>
          .1109/ASPAA.
          <year>1999</year>
          .
          <volume>810836</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>M.</given-names>
            <surname>Buck</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Wolff</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Haulick</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Schmidt</surname>
          </string-name>
          ,
          <article-title>A compact microphone array system with spatial postfiltering for automotive applications</article-title>
          ,
          <source>in: Proc. IEEE International Conference on Acoustics, Speech, and Signal Processing (ICASSP</source>
          <year>2009</year>
          ),
          <year>2009</year>
          , pp.
          <fpage>221</fpage>
          -
          <lpage>224</lpage>
          . doi:
          <volume>10</volume>
          .1109/ICASSP.
          <year>2009</year>
          .
          <volume>4959560</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <surname>Microphone</surname>
            <given-names>Array</given-names>
          </string-name>
          <string-name>
            <surname>Beamforming</surname>
          </string-name>
          .
          <source>Application Note AN-1140</source>
          . URL: http://www.invensense.com/wp- content/uploads/2015/02/Microphone-Array-Beamforming.pdf.
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <given-names>T.</given-names>
            <surname>Lotter</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Vary</surname>
          </string-name>
          ,
          <article-title>Dual-channel speech enhancement by superdirective beamforming</article-title>
          .
          <source>Eurasip Journal on Applied Signal Processing</source>
          , (
          <year>2006</year>
          )
          <fpage>1</fpage>
          -
          <lpage>14</lpage>
          . doi:
          <volume>10</volume>
          .1155/ASP/
          <year>2006</year>
          /63297.
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <given-names>J.</given-names>
            <surname>Bitzer</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.U.</given-names>
            <surname>Simmer</surname>
          </string-name>
          ,
          <string-name>
            <surname>K.-D. Kammeyer</surname>
          </string-name>
          <article-title>, Multi-microphone noise reduction techniques for handsfree speech recognition - a comparative study</article-title>
          ,
          <source>in Proc. of Workshop on Robust Methods forSpeech Recognition in Adverse Conditions (ROBUST 99)</source>
          ,
          <year>1999</year>
          , pp.
          <fpage>171</fpage>
          -
          <lpage>174</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <given-names>Y.</given-names>
            <surname>Ji</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Byun</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Park</surname>
          </string-name>
          ,
          <article-title>Coherence-based dual-channel noise reduction algorithm in a complex noisy environment</article-title>
          ,
          <source>in: INTERSPEECH</source>
          <year>2017</year>
          . Stockholm, Sweden,
          <source>August 20-24</source>
          ,
          <year>2017</year>
          , doi:10.21437/INTERSPEECH.2017-
          <volume>1464</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [16]
          <string-name>
            <surname>Snreval</surname>
          </string-name>
          . URL: https://labrosa.ee.columbia.edu/projects/snreval/.
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [17]
          <string-name>
            <given-names>M.</given-names>
            <surname>Stolbov</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Tatarnikova</surname>
          </string-name>
          ,
          <string-name>
            <surname>The Q.T.</surname>
          </string-name>
          <article-title>Using dual-element microphone arrays for automatic keyword recognition</article-title>
          .
          <source>Lecture Notes in Computer Science (including subseries Lecture Notes in Artifi cial Intelligence and Lecture Notes in Bioinformatics)</source>
          .
          <volume>11096</volume>
          (
          <year>2018</year>
          )
          <fpage>667</fpage>
          -
          <lpage>675</lpage>
          . DOI:
          <volume>10</volume>
          .1007/978-3-
          <fpage>319</fpage>
          -99579-3-68.
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          [18]
          <string-name>
            <given-names>S.</given-names>
            <surname>Liu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Yao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Liu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Zhu</surname>
          </string-name>
          ,
          <string-name>
            <surname>J.</surname>
          </string-name>
          <article-title>Ren, Multi-band discriminant speech synthesis analysis based on Natural Language Processing</article-title>
          ,
          <source>in: 2022 7th International Conference on Intelligent Computing and Signal Processing (ICSP)</source>
          ,
          <year>2022</year>
          , pp.
          <fpage>1933</fpage>
          -
          <lpage>1936</lpage>
          . doi:
          <volume>10</volume>
          .1109/ICSP54964.
          <year>2022</year>
          .
          <volume>9778506</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          [19]
          <string-name>
            <given-names>Y.</given-names>
            <surname>Lu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Harati</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Rutowski</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Oliveira</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Chlebek</surname>
          </string-name>
          , E. Shriberg,
          <article-title>Robust Speech and Natural Language Processing Models for Depression Screening</article-title>
          , in: 2020
          <source>IEEE Signal Processing in Medicine and Biology Symposium (SPMB)</source>
          , Philadelphia, PA, USA,
          <year>2020</year>
          , pp.
          <fpage>1</fpage>
          -
          <lpage>5</lpage>
          . doi:
          <volume>10</volume>
          .1109/SPMB50085.
          <year>2020</year>
          .
          <volume>9353611</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          [20]
          <string-name>
            <given-names>Z.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Whitehill</surname>
          </string-name>
          ,
          <article-title>Compositional Embedding Models for Speaker Identifi cation and Diarization with Simultaneous Speech From 2+ Speakers</article-title>
          , in: ICASSP 2021
          <article-title>-</article-title>
          2021 IEEE International Conference on Acoustics,
          <source>Speech and Signal Processing (ICASSP)</source>
          , Toronto, ON, Canada,
          <year>2021</year>
          , pp.
          <fpage>7163</fpage>
          -
          <lpage>7167</lpage>
          . doi:
          <volume>10</volume>
          .1109/ICASSP39728.
          <year>2021</year>
          .
          <volume>9413752</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          [21]
          <string-name>
            <given-names>F.</given-names>
            <surname>Roumiassa</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Chelali</surname>
          </string-name>
          ,
          <article-title>Speaker Identifi cation and Verifi cation System for Arabic and Berber Language</article-title>
          ,
          <source>in: 2020 1st International Conference on Communications, Control Systems and Signal Processing (CCSSP)</source>
          ,
          <source>El Oued, Algeria</source>
          ,
          <year>2020</year>
          , pp.
          <fpage>242</fpage>
          -
          <lpage>247</lpage>
          . doi:
          <volume>10</volume>
          .1109/CCSSP49278.
          <year>2020</year>
          .
          <volume>9151633</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          [22]
          <string-name>
            <given-names>K.</given-names>
            <surname>Chu</surname>
          </string-name>
          , L. Collins,
          <string-name>
            <given-names>B.</given-names>
            <surname>Mainsah</surname>
          </string-name>
          ,
          <article-title>Using Automatic Speech Recognition and Speech Synthesis to Improve the Intelligibility of Cochlear Implant users in Reverberant Listening Environments</article-title>
          , in: ICASSP 2020
          <article-title>-</article-title>
          2020 IEEE International Conference on Acoustics,
          <source>Speech and Signal Processing (ICASSP)</source>
          , Barcelona, Spain,
          <year>2020</year>
          , pp.
          <fpage>6929</fpage>
          -
          <lpage>6933</lpage>
          . doi:
          <volume>10</volume>
          .1109/ICASSP40776.
          <year>2020</year>
          .
          <volume>9054450</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          [23]
          <string-name>
            <given-names>M.</given-names>
            <surname>Yousefi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Hansen</surname>
          </string-name>
          ,
          <article-title>Speaker Conditioning of Acoustic Models Using Affine Transformation for Multi-Speaker Speech Recognition, in: 2021 IEEE Automatic Speech Recognition</article-title>
          and Understanding Workshop (ASRU), Cartagena, Colombia,
          <year>2021</year>
          , pp.
          <fpage>283</fpage>
          -
          <lpage>288</lpage>
          . doi:
          <volume>10</volume>
          .1109/ASRU51503.
          <year>2021</year>
          .
          <volume>9688231</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          [24]
          <string-name>
            <given-names>M.</given-names>
            <surname>Soleymanpour</surname>
          </string-name>
          , M. T. Johnson,
          <string-name>
            <given-names>R.</given-names>
            <surname>Soleymanpour</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Berry</surname>
          </string-name>
          ,
          <article-title>Synthesizing Dysarthric Speech Using Multi-Speaker Tts For Dysarthric Speech Recognition</article-title>
          , in: ICASSP 2022
          <article-title>-</article-title>
          2022 IEEE International Conference on Acoustics,
          <source>Speech and Signal Processing (ICASSP)</source>
          , Singapore, Singapore,
          <year>2022</year>
          , pp.
          <fpage>7382</fpage>
          -
          <lpage>7386</lpage>
          . doi:
          <volume>10</volume>
          .1109/ICASSP43922.
          <year>2022</year>
          .
          <volume>9746585</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref25">
        <mixed-citation>
          [25]
          <string-name>
            <given-names>D.</given-names>
            <surname>Sundström</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Lindström</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Jakobsson</surname>
          </string-name>
          ,
          <article-title>Recursive Spatial Covariance Estimation with Sparse Priors for Sound Field Interpolation</article-title>
          ,
          <source>in: Proc 2023 IEEE Statistical Signal Processing Workshop</source>
          (SSP), Hanoi, Vietnam,
          <year>2023</year>
          , pp.
          <fpage>517</fpage>
          -
          <lpage>521</lpage>
          , doi: 10.1109/SSP53291.
          <year>2023</year>
          .
          <volume>10208010</volume>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>