<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <article-id pub-id-type="doi">10.1155/2017/6782176</article-id>
      <title-group>
        <article-title>Edge computing applications: using a linear MEMS microphone array for UAV position detection through sound source localization</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Andrii V. Riabko</string-name>
          <xref ref-type="aff" rid="aff4">4</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Tetiana A. Vakaliuk</string-name>
          <email>tetianavakaliuk@acnsci.org</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
          <xref ref-type="aff" rid="aff5">5</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Oksana V. Zaika</string-name>
          <email>ksuwazaika@gmail.com</email>
          <xref ref-type="aff" rid="aff4">4</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Roman P. Kukharchuk</string-name>
          <xref ref-type="aff" rid="aff4">4</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Valerii V. Kontsedailo</string-name>
          <email>valerakontsedailo@gmail.com</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Academy of Cognitive and Natural Sciences</institution>
          ,
          <addr-line>54 Universytetskyi Ave., Kryvyi Rih, 50086</addr-line>
          ,
          <country country="UA">Ukraine</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Inner Circle</institution>
          ,
          <addr-line>Nieuwendijk 40, 1012 MB Amsterdam</addr-line>
          ,
          <country country="NL">Netherlands</country>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>Institute for Digitalisation of Education of the NAES of Ukraine</institution>
          ,
          <addr-line>9 M. Berlynskoho Str., Kyiv, 04060</addr-line>
          ,
          <country country="UA">Ukraine</country>
        </aff>
        <aff id="aff3">
          <label>3</label>
          <institution>Kryvyi Rih State Pedagogical University</institution>
          ,
          <addr-line>54 Universytetskyi Ave., Kryvyi Rih, 50086</addr-line>
          ,
          <country country="UA">Ukraine</country>
        </aff>
        <aff id="aff4">
          <label>4</label>
          <institution>Oleksandr Dovzhenko Hlukhiv National Pedagogical University</institution>
          ,
          <addr-line>24 Kyivska Str., Hlukhiv, 41400</addr-line>
          ,
          <country country="UA">Ukraine</country>
        </aff>
        <aff id="aff5">
          <label>5</label>
          <institution>Zhytomyr Polytechnic State University</institution>
          ,
          <addr-line>103 Chudnivsyka Str., Zhytomyr, 10005</addr-line>
          ,
          <country country="UA">Ukraine</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2017</year>
      </pub-date>
      <abstract>
        <p>This study explores the use of a microphone array to determine the position of an unmanned aerial vehicle (UAV) based solely on the sound of its engines. The accuracy of localization depends crucially on the arrangement of the microphones. The study also considers a mathematical model of pulse density modulation for a digital MEMS microphone. It demonstrates the frequency dependence of the eficiency of a diferential array of first-order microphones. Based on this frequency dependence of directivity and the instability model of the microphone parameters, a rational operating frequency range for the normal functioning of the microphone array can be established. The study proposes a model of a linear microphone array based on MEMS omnidirectional microphones. With a specific geometrical arrangement, this array produces a bidirectional pattern, which can be easily transformed into a unidirectional pattern using specialized algorithms or hardware (e.g., ADAU1761 codecs).</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;edge computing</kwd>
        <kwd>UAV</kwd>
        <kwd>sound source localization</kwd>
        <kwd>MEMS microphone</kwd>
        <kwd>microphone array</kwd>
        <kwd>frequency</kwd>
        <kwd>directivity</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>Determining the position of a UAV (Unmanned Aerial Vehicle) by the sound of its engines can be
important for several reasons. In military or security applications, being able to identify and locate UAVs
by their engine sounds can help in detecting potential threats, including hostile drones or unauthorized
surveillance. Sound-based localization can aid in the development of countermeasures to mitigate the
risks posed by UAVs in sensitive areas.</p>
      <p>Sound-based UAV detection can complement existing air trafic management systems, providing
additional situational awareness for managing airspace and preventing collisions with manned aircraft.
In search and rescue operations or in case of lost or malfunctioning drones, sound-based tracking can
assist in locating and recovering UAVs.</p>
      <p>In conservation eforts, it can help monitor UAVs used for illegal activities like poaching or wildlife
disturbance. In urban areas or regions with dense UAV trafic, sound-based tracking can be useful for
enforcing regulations related to UAV flight paths, altitudes, and no-fly zones. For protecting privacy,
sound-based detection can help identify UAVs flying near private properties, providing a means to take
legal action against intrusive drones.</p>
      <p>Studying the acoustic signatures of UAVs can aid in research and development eforts to design quieter
and more environmentally friendly drones. During natural disasters or emergencies, knowing the
positions of UAVs, such as those used for aerial surveys or damage assessment, can assist in coordinating
response eforts. Sound-based UAV detection can be employed in border control to monitor and respond
to unauthorized drone incursions.</p>
      <p>As drone delivery and urban air mobility concepts develop, sound-based localization can contribute
to managing UAV trafic in urban environments. While sound-based UAV localization ofers several
advantages, it also has limitations, such as accuracy challenges in noisy environments and the need
for specialized equipment. Therefore, it is often used in conjunction with other tracking and detection
methods, such as radar, visual recognition, and GPS, to provide comprehensive situational awareness
and enhance safety and security in various applications.</p>
      <p>The goal of our work is to develop a software and hardware system for capturing
hardwaresynchronized sound using digital MEMS microphones (Microelectromechanical Systems, MEMS) for
further use in sound source localization systems. This system is intended for further use in sound
source localization systems, marking a significant advancement in the field of edge computing.</p>
    </sec>
    <sec id="sec-2">
      <title>2. Theoretical background</title>
      <p>
        Over the past few decades, acoustic source localization has emerged as a focal point of interest within the
research community [
        <xref ref-type="bibr" rid="ref1 ref2">1, 2</xref>
        ]. Most studies of sound source identification are based on the analysis of the
physiological mechanism of human hearing [
        <xref ref-type="bibr" rid="ref3 ref4">3, 4</xref>
        ]. It is common practice to use arrays of microphones
[
        <xref ref-type="bibr" rid="ref5">5</xref>
        ]. An actual problem is acoustic beam formation for sound source localization and its application [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ].
      </p>
      <p>
        Microphone array processing represents a well-established methodology employed in the estimation
of sound source direction. In a groundbreaking contribution by Yamada et al. [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ], they introduce an
innovative approach referred to as Multiple Triangulation and Gaussian Sum Filter Tracking (MT-GSFT).
This advanced technique adeptly derives the precise location of sound sources through triangulation,
utilizing microphone arrays seamlessly integrated into a fleet of multiple drones [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ]. The domain of
speech signal processing encompasses several critical areas, and among them, multiple sound source
localization (SSL) stands out as a notable and relevant field. A notable contribution to this field comes
from Firoozabadi et al. [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ], who introduced a two-step approach for the localization of multiple sound
sources in three dimensions (3D). This method relies on the precise estimation of time delays (TDE) and
strategically leverages distributed microphone arrays (DMA) to enhance the accuracy and efectiveness
of the localization process [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ].
      </p>
      <p>
        Sasaki et al. [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ] present a method designed to map the 3D coordinates of a sound source by leveraging
data gathered from an array of microphones, with each microphone providing an autonomous directional
estimate. Additionally, LiDAR technology is employed to create a comprehensive 3D representation of
the surroundings and accurately determine the sensor’s position with six degrees of freedom (6-DoF).
      </p>
      <p>
        Catalbas et al. [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ] conduct a comparative analysis, assessing the efectiveness of generalized
crosscorrelation techniques in contrast to noise reduction filters concerning the estimation of sound source
trajectory. Throughout the entire movement, they calculate the azimuth angle between the sound
source and the receiver. This calculation relies on the parameter of Interaural Time Diference (ITD) to
determine the azimuth angle. They then evaluate the accuracy of the estimated delay using various
types of Generalized Cross-Correlation (GCC) algorithms for comparison.
      </p>
      <p>
        It is possible for unmanned aerial vehicles (UAVs) to use audio information to compensate for poor
visual information. Hoshiba et al. [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ] developed a microphone array system built into the UAV to
localize the sound source in flight. They developed the Spherical Microphone Array System (SMAS),
consisting of a microphone array, a stable wireless network communication system, and intuitive
visualization tools.
      </p>
      <p>
        Tachikawa et al. [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ] introduced an innovative approach that involves estimating positions by utilizing
a modified variant of the convex clustering method in conjunction with sparse coeficients estimation.
Additionally, they put forth a technique for constructing a well-suited monopole dictionary, which is
based on coherence, ensuring that the convex clustering-based method can accurately estimate the
distances of sound sources. The study involved conducting a series of numerical and measurement
experiments aimed at assessing the efectiveness and performance of this novel methodology.
      </p>
      <p>
        When dealing with multiple sound sources, establishing a reliable data association between
localization information and the corresponding sound sources becomes paramount for achieving optimal
performance. To address the challenges posed by data association uncertainty, Wakabayashi et al.
[
        <xref ref-type="bibr" rid="ref13">13</xref>
        ] extended the Global Nearest Neighbor (GNN) approach, introducing a modified version known
as GNN-c, specifically tailored to meet the real-time and low-latency requirements of drone audio
applications. The outcome of their eforts showcases a system capable of accurately estimating the
positions of multiple sound sources, achieving an impressive accuracy level of approximately 3 meters.
      </p>
      <p>
        Many acoustic image-based sound source diagnosis systems sufer from spatial stationary limitations,
making it challenging to integrate information from various capture positions, thereby leading to
unreliable and incomplete diagnostics. In their paper, Carneiro and Berry [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ] introduce a novel measurement
methodology called Acoustic Imaging Structure From Motion (AISFM). This approach utilizes a mobile
spherical microphone array to create acoustic images through beamforming, seamlessly integrating data
from multiple capture positions. Their method is not only proposed but also meticulously developed
and rigorously validated, ofering a promising solution to enhance the accuracy and comprehensiveness
of sound source diagnostics.
      </p>
      <p>
        In a research conducted by Kita and Kajikawa [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ] a sound source localization (SSL) technique is
introduced, specifically designed for the localization of sources situated within structures, including
mechanical equipment and buildings.
      </p>
      <p>
        The registration of acoustic signals with cross-shaped antennas is widely discussed in the literature
[
        <xref ref-type="bibr" rid="ref16">16</xref>
        ].
      </p>
      <p>
        Advanced signal processing methods involving multiple microphones can enhance noise resilience.
However, as the quantity of microphones employed escalates, the computational overhead rises
concomitantly. This, in turn, curtails response time and hinders their extensive adoption across various
categories of mobile robotic platforms [
        <xref ref-type="bibr" rid="ref17">17</xref>
        ]. Within the realm of robot audition, sound source localization
(SSL) holds a pivotal role, serving as a fundamental component. SSL empowers a robotic platform to
pinpoint the origin of sound using auditory cues exclusively. Its significance extends beyond mere sound
localization, as it significantly influences other facets of robot audition, including source separation.
Moreover, SSL contributes to elevating the quality of human-robot interaction by augmenting the
robot’s perceptual prowess [
        <xref ref-type="bibr" rid="ref18">18</xref>
        ].
      </p>
      <p>
        In general, machine learning is widely used in acoustics [
        <xref ref-type="bibr" rid="ref19 ref20">19, 20</xref>
        ]. In the realm of human-robot
interaction, He et al. [
        <xref ref-type="bibr" rid="ref21">21</xref>
        ] have introduced a pioneering approach. Their proposal involves harnessing
neural networks for the simultaneous detection and localization of multiple sound sources. This
innovative method represents a departure from conventional signal processing techniques by ofering a
distinct advantage: it necessitates fewer stringent assumptions about the environmental conditions,
thereby enhancing its adaptability and efectiveness [
        <xref ref-type="bibr" rid="ref21">21</xref>
        ]. Ebrahimkhanlou and Salamone [
        <xref ref-type="bibr" rid="ref22">22</xref>
        ] have
put forth an advanced methodology for localizing acoustic emissions (AE) sources within metallic
plates, especially those with intricate geometric features like rivet-connected stifeners. This innovative
approach leverages two deep learning techniques: a stack of autoencoders and a convolutional neural
network (CNN), strategically employed to enhance the accuracy and precision of the localization process
[
        <xref ref-type="bibr" rid="ref22">22</xref>
        ].
      </p>
      <p>
        In their pioneering work, Adavanne et al. [
        <xref ref-type="bibr" rid="ref23">23</xref>
        ] have introduced an innovative solution – a
convolutional recurrent neural network (CRNN) – designed to address the intricate task of joint sound event
localization and detection (SELD) within three-dimensional (3-D) space. This method represents a
significant advancement in the field, enabling the simultaneous identification and spatial localization of
multiple overlapping sound events with remarkable precision.
      </p>
      <p>
        Let’s summarize the theoretical review. Localizing a sound source means determining the direction
or location from which a sound is emanating. There are several algorithms and techniques used for
sound source localization, and the choice of method often depends on the specific application and
available hardware. Here are some commonly used algorithms. Time Diference of Arrival (TDOA) is
based on measuring the time it takes for a sound to reach multiple microphones. By comparing the
time diferences, it’s possible to triangulate the source’s position. Cross-correlation or beamforming
techniques are often used to calculate the time diferences accurately. Generalized Cross-Correlation
(GCC) is a technique used in conjunction with TDOA. It involves cross-correlating the signals from
two or more microphones to find the delay between them. GCC-PHAT (GCC with Phase Transform) is
a commonly used variant that works well in reverberant environments. Steering Vector Methods are
commonly used in microphone arrays or beamforming applications [
        <xref ref-type="bibr" rid="ref24">24</xref>
        ]. They estimate the direction of
arrival (DOA) by analyzing the phase diferences between signals received by diferent microphones.
Popular algorithms include Multiple Signal Classification (MUSIC) and Estimation of Signal Parameters
via Rotational Invariance Techniques (ESPRIT). Acoustic Intensity Methods measure the sound intensity
at multiple microphone positions and use this information to estimate the source direction. The Steered
Response Power (SRP) algorithm is an example of this approach. Machine learning and deep learning
techniques, such as neural networks and support vector machines, can be used to train models for
sound source localization. These models can take input from multiple microphones and learn to predict
the source location based on training data. Particle filtering is a probabilistic method that estimates
the source location using a Bayesian filtering approach. It is useful when dealing with complex and
dynamic environments. Some methods use time-frequency analysis techniques like the Short-Time
Fourier Transform (STFT) or Wavelet Transform to analyze the spectral content of audio signals and
infer the source location. In mobile sound source localization, the Doppler efect can be used to estimate
the source’s speed and direction based on the frequency shift in the received signal. Many practical
systems use a combination of the above techniques to improve accuracy and robustness, especially in
real-world scenarios with noise and reverberation.
      </p>
      <p>The choice of algorithm depends on factors like the number and arrangement of microphones,
environmental conditions, computational resources, and the desired level of accuracy. Diferent applications,
such as robotics, audio conferencing, surveillance, and hearing aids, may employ diferent algorithms
tailored to their specific requirements.</p>
      <p>Determining the position of a UAV (Unmanned Aerial Vehicle) based solely on the sound of its
engines can be challenging but is feasible using a combination of sound source localization techniques
and signal processing. Here’s a high-level overview of the process:
1. Microphone Array Setup: Set up a microphone array on the ground. The microphones should
be strategically placed to capture the UAV’s sound from diferent angles. The arrangement of
microphones plays a crucial role in accurate localization. The response of microphone arrays
depends, first of all, on the number of microphones working on the array [25].
2. Sound Data Collection: Record the sound generated by the UAV’s engines as it flies overhead.</p>
      <p>Ensure that the recording system has a high sampling rate to capture the sound accurately.
3. Time Delay of Arrival (TDOA): Analyze the recorded audio data to calculate the time delay of
arrival (TDOA) of the sound at each microphone. TDOA is the time diference between when the
sound reaches diferent microphones. This information is critical for triangulation.
4. Triangulation: Use the TDOA data from multiple microphones to triangulate the UAV’s
position. Several algorithms, such as multilateration or beamforming, can help estimate the UAV’s
coordinates based on the TDOA information.
5. UAV Sound Signature: To improve accuracy, consider using machine learning techniques to create
a database of UAV sound signatures. This involves training a model to recognize the unique
sound characteristics of diferent UAVs. When a new sound recording is obtained, the model can
help identify the specific UAV type.
6. Integration with Other Sensors: For real-time tracking, integrate sound-based localization with
other sensors like GPS, radar, or visual cameras. This fusion of data sources can provide more
accurate and robust positioning.
7. Calibration and Testing: Regularly calibrate and test the microphone array and signal processing
algorithms to ensure accurate and reliable results.</p>
      <p>It’s important to note that the accuracy of sound-based UAV localization depends on various factors,
including the UAV’s altitude, speed, engine type, and background noise. Additionally, environmental
conditions, such as wind and temperature, can afect sound propagation and localization accuracy.
Therefore, this method may work best in controlled environments or in conjunction with other tracking
methods for enhanced precision and reliability.</p>
    </sec>
    <sec id="sec-3">
      <title>3. Research methods</title>
      <p>The goal of our work is to develop a software and hardware system for capturing hardware-synchronized
sound using digital MEMS microphones (Microelectromechanical Systems, MEMS) for further use in
sound source localization systems.</p>
      <p>Despite the fact that the use of radar equipment has become part of everyday practice when monitoring
UAVs, there is some interest in assessing the possibility of using airborne acoustic signals for this purpose.
The above applies mainly to receiving hydroacoustic antennas, i.e. to conditions when the speed of
the source is much lower than the speed of sound  = / ≪ 1. In the case of receiving air-acoustic
signals propagating at a speed of sound significantly lower than that of hydroacoustic signals in water
and created by fairly fast moving sources (passenger cars on autobahns, racing cars, the movement of
airliners along runways during takeof and landing, UAVs), there is a diferent situation. Research into
the features of recording these signals with phased arrays remains relevant, since on their basis data
can be obtained on the current coordinates and speed of movement of a moving object. The purpose of
this work is to analyze the angular dependencies in the signal at the output of a receiving air-acoustic
antenna and those qualitative changes in their nature that are introduced due to a combination of
factors such as the Doppler efect and the sharp directivity of the antenna array.</p>
      <p>A special case is considered, which is widespread in everyday practice, when the trajectory of an
object is rectilinear, lies in a horizontal plane, close and parallel to the Earth’s surface, and the speed of
its movement is constant.</p>
      <p>As previously stated, the arrangement of microphones plays a crucial role in accurate localization.
The purpose of the study is to find the optimal configuration of a microphone array for localizing a
moving sound source (UAV).</p>
      <p>Directivity is the sensitivity of a microphone to sound depending on the direction or angle from
which the sound is coming. Directionality or sound pickup angle is considered to be the area of possible
location of the sound signal source, within which there is no significant loss of microphone eficiency.
Microphones use diferent directivity characteristics. They are most often depicted as polar diagrams.
This is done to graphically display sensitivity variations around the microphone over a 360-degree
range, where the microphone is the center of the circle and the angular reference point is placed in
front of the microphone. The polar pattern shows how a microphone’s sensitivity to a sound signal
depends on the location of its source.</p>
      <p>Microphone arrays are an array of several microphones combined by joint digital signal processing.
Microphone arrays provide the following advantages over single-channel systems: 1) directionality of
sound reception; 2) noise suppression of point sources; 3) suppression of non-stationary environmental
noise; 4) partial weakening of reverberation; 5) the possibility of spatial localization of the sound source;
6) the ability to accompany a moving point sound source.</p>
      <p>A microphone array is one of the types of directional microphones, implemented as a set of sound
receivers operating in concert (in phase or with certain phase delays). Geometrically, gratings can
be implemented in diferent configurations – one-dimensional (linear, arc-shaped), two-dimensional
(flat, spherical), three-dimensional, spiral, with uniform or non-equidistant pitch. The array’s radiation
pattern is created by changing the ratio of phase delays for diferent channels (in the simplest case, an
in-phase array with a fixed position of the main lobe; in more complex and expensive implementations,
a scanning system). The implementation of phase delays can be hardware (for example, on analog delay
lines) or software (digital).</p>
      <p>The basic microphone array structures are Broadside and Endfire (figure 1).</p>
      <p>These structures use omnidirectional microphones (microphones that, regardless of their orientation,
receive signals from any direction). The figure 2 shows signal reception versus direction for various
frequencies with a single omnidirectional microphone. For one microphone, frequency invariance is
observed.</p>
      <p>The Broadside structure is an array of omnidirectional microphones positioned perpendicular to
the direction of the desired signal. Such arrays have an axis of symmetry, relative to which the sound
is released without attenuation both “in front” of the array and “behind”. Such structures are widely
used in applications where sound pressure waves enter the sensor array from one side. Consider a
Broadside structure consisting of two microphones spaced 7.5 cm apart. The minimum response is
observed when the signal is incident at an angle of 90∘ or 270∘ (in this case, the angle between the
direction of the useful signal and the normal to the line of elements is taken as 0∘ ). But this response
strongly depends on the frequency of the received signal. Theoretically, such a system has a perfect
zero at a frequency of 2.3 kHz. Above this frequency, depending on the direction of arrival, there are
zeros at other angles (figure 3). The microphone array shows a clear directional characteristic at 4 kHz,
and at 1 kHz its pattern is essentially omnidirectional. As a result, at lower frequencies the array cannot
achieve significant spatial filtering.</p>
      <p>The Endfire structure consists of several microphones located in the direction of the useful acoustic
signal. This design is called a diferential array of microphones. The delayed signal from the first
microphone is summed with the signal from the next microphone. To create a cardioid polar pattern,
the signal from the rear microphones must be delayed by the same amount of time that the sound
waves travel between the two microphone elements. Such structures are used to produce cardioid,
hypercardioid or supercardioid directional response and theoretically completely eliminate sound
incident on the array at an angle of 180∘ . A unidirectional microphone is more sensitive to sound
coming from one direction and less sensitive to sounds from other directions. The most typical for such
microphones is the cardioid characteristic, representing a peculiar diagram in the shape of a heart. At
the same time, the peak of sensitivity is reached in the direction along the axis of the microphone, and
the decline is in the opposite direction (figure 4).</p>
      <p>To generate a cardioid response in direction, the signal from the omnidirectional microphones must be
delayed for a time equal to the propagation of the acoustic wave between the two elements. Developers
of such systems have two degrees of freedom to change the output signal of the speaker system:
changing the distance between microphones and changing the delay time. Figure 5 shows the signal
reception versus direction for various frequencies by the Endfire structure with two elements and a
distance between them of 2.1 cm.</p>
      <p>The distance between the microphones is crucial for the formation of a cardioid response. Figure 6
shows the same microphones, but placed at a distance of 15 cm.</p>
      <p>The structures considered have the following advantages and disadvantages. Advantages of
Broadside: flat geometry, simple processing implementation, ability to control the direction of the beam.
Disadvantages of the Broadside: less of-axis rejection, close microphone spacing, and a large number
of microphones needed to prevent spatial leakage.</p>
      <p>Advantages of Endfire: Better of-axis suppression, smaller overall size. Disadvantages of Endfire:
non-flat (volumetric) geometry, more complex processing, suppression of the useful signal in the low
frequency range, the direction of the source of the useful signal must coincide with the axis of the
microphone array; For two-dimensional gratings, beam formation is possible only in the horizontal
direction (the grating array).</p>
      <p>To form a diferential array of higher orders, you need to add additional microphones. Since the petals
will deviate more back and to the side in the directional diagram, the distance between the microphones
will have to be increased. The figure 7 shows an array of 4 microphones (third order), which forms
a supercardioid pattern. Consider how beam formation depends on the number of microphones and
the distance between them. It is worth noting that the sensitivity and frequency response of all array
microphones must be precisely matched.</p>
      <p>Diferential microphone arrays make it possible to obtain high directivity characteristics of the system
with its small size. But with such a construction, the problem arises of a significant change in the
characteristics of the entire system with a slight deviation of the parameters of an individual microphone
from its nominal values. If this approach is used for critical applications, measures must be taken
to reduce deviations of microphone parameters from nominal values. As stated earlier, a first-order
diferential microphone array consists of two omnidirectional sensors separated by  (figure 8).</p>
    </sec>
    <sec id="sec-4">
      <title>4. Results</title>
      <p>When sound arrives from the main direction  = 0, a delay appears between these sensors:
  =


,
(1)
(2)
(3)
(4)
(5)
where  is the speed of sound.</p>
      <p>A plane wave, which is characterized by wave numbe→r−  , arrives at the input of the diferential
grating. Due to radial symmetry, the output signals of sensors 1() and 2() can be expressed
by a function depending on the angle  and frequency . There is a relationship→|−  | =  =  
between the wave number and the frequency of the signal. At the central point of the array, you can
place a virtual microphone with an output signal 0(). A plane wave incident at an angle  with
wave number causes  = 2/</p>
      <p>the appearance of signals at the output of microphones 1 and 2:
1() = 0() 2 cos  , 2() = 0()−  2 cos  .</p>
      <p>At the output of the diferential lattice we get
1
2
() =</p>
      <p>(1() − 2() .
(,  ) = −  2 sin
︂(  (︂ 
+ cos</p>
      <p>The directivity function of the diferential array
 is the ratio of the signal at the output of the
array () to the signal at the output of the virtual microphone 0():
sin  ≈  . In this case, the idealized directivity function  has the form:</p>
      <sec id="sec-4-1">
        <title>Usually very small values of  ≪</title>
        <sec id="sec-4-1-1">
          <title>1 are considered, which makes it possible to use the approximation</title>
          <p>(,  ) ≈ ̃︀( ) = 
+ cos</p>
          <p>With this view, the main characteristics of diferential microphone arrays are obvious: 1) the form of
̃︀( ) is determined by the expression  /  + cos  , which does not depend on frequency; 2) due to
the subtraction of the signal a phase shift occurs by / 2; 3) the frequency response of the directivity
function () has the form of a first-order high-pass filter.</p>
          <p>At low frequencies, the output signal () becomes highly susceptible to any changes in the shape
of the characteristic (). For this reason, the distance d should not be chosen too small, which may
lead to a conflict with the condition  ≪</p>
          <p>The exact expression for the directivity function (4) contains a sine function that scales the amplitude.
It is rational to limit the operating range of the diferential grating in the low frequency range to the
ifrst maximum of the sine. This first maximum fixes the cutof frequency
:
 =</p>
          <p>+ 
.</p>
          <p>For low frequencies, the directivity characteristics are practically independent of frequency. However,
as the frequency increases, the shape of the frequency response becomes more and more deformed. In
addition, at some frequencies the signal is completely suppressed.</p>
          <p>In order to compensate for the high-frequency nature of the behavior (,  ) it is necessary to
develop a filter</p>
          <p>(). For the main direction  = 0, the adjusted frequency response(,  =
0)() must be constant and equal to 0 dB, and for frequencies below :
() =</p>
          <p>2
{︃</p>
          <p>1
sin(︁  ︁) , 0 &lt;  &lt; ,
1,
in other cases.</p>
        </sec>
      </sec>
      <sec id="sec-4-2">
        <title>For low frequencies  → 0, the filter gain</title>
        <p>has very large values. This means that any noise
present in the input signal will be greatly amplified. The level of this noise is determined by the specific
sensor. This circumstance limits the frequency range of the signal for processing using a diferential
microphone array.</p>
        <p>The directional properties of a microphone array are characterized by the directional coeficient (DI).
It can be expressed as the ratio of the squared modulus of the directivity function in the main direction
to the average value of the squared modulus in all directions:
(6)
(7)
(8)
(9)
̃︀
(10)
(11)
(12)
() =
2 
1 ∫︀ ∫︀ |(,  )|2 sin</p>
        <p>.
() =</p>
        <p>22 (︀ 2 (  +  ))︀
1</p>
        <p>−  ( ) cos ( )
→0
lim () = ̃︀ =
3(  +  )2
3 2 + 3 2
.
where () = 1 sin().
according to expression (5):</p>
        <p>The eficiency factor for low frequencies is obtained similarly to the result of approximation
Taking into account the exact expression for the directivity function (4), we can obtain a new
expression for the dependence of the directivity on frequency:</p>
        <p>The directivity function ^ for a diferential array, taking into account the instability of the
microphone parameters, can be obtained similarly to expression (4). But now there are additional conditions</p>
        <p>Let us study the influence of microphone parameter mismatch for first-order diferential arrays. We
use a model of instability of microphone parameters in the form of a transfer function  =  +Δ .
The nominal transfer function of the sensor  in this case is normalized to the value 1. It is assumed
that the deviation Δ is an independent random variable with variance:</p>
        <p>2 = {|Δ |2},
follows:
where {|Δ |2} expectation operator. Signals from two sensors in figure 8 will then be written as
^1() = 0()(1 + Δ1) 2 cos  , ^2() = 0()(1 + Δ2)−  2 cos  .
that depend on Δ ( = 1, 2). For random numbers, the quadratic terms remain, and the linear ones
are set to zero, so we get:</p>
        <p>{|^(,  )|2} = |(,  )|2 + 2 2 .</p>
        <p>As a result, we can obtain a modified expression for DI:
(13)
(14)
{|^()|} =
2 sin2 (︀ 2 (  +  ))︀ +  2
1 −  ( ) cos ( )  2</p>
        <p>It is important to understand that in expression (13) the eficiency factor (,  ) characterizes the
behavior of the system at high frequencies. While the  equalization filter takes into account the
efects of microphone instability and enhances them for low frequencies.</p>
        <p>Thus, this work shows the dependence of the eficiency of a diferential array of first-order
microphones on frequency. It can be supplemented using a model of instability of microphone parameters at
low frequencies.</p>
        <p>Based on the presented dependence of the directivity on frequency and the instability model of
the microphone parameters, a rational operating frequency range for the normal functioning of the
microphone array can be determined. The lower limit of this range is limited by the instability of the
microphone parameters, and the upper cutof frequency is determined by the geometry of the array .</p>
        <p>Currently, there are many applications in which acoustic signals are processed.
Microelectromechanical microphones (MEMS) are increasingly being used for these purposes. The use of such microphones
allows the construction of diferential microphone arrays. Microelectromechanical systems (MEMS) are
a variety of microdevices of a wide variety of designs and purposes, in the production of which modified
microelectronics technological techniques are used. Typically, all elements of such systems are placed
on a common silicon base, the size of which is only a couple of millimeters. A MEMS microphone is an
electro-acoustic device for converting sound vibrations into electrical waves, which is small enough to
be installed in a tightly integrated product, for example: a smartphone, headset, speakerphone, laptop or
any other device. There are two fundamentally important elements in such microphones: an integrated
circuit (ASIC) and a MEMS sensor. It is the latter that ensures the capture and subsequent transmission
of sound. The MEMS sensor itself consists of a flexible membrane and a rigidly fixed cover. Under the
influence of air pressure, the membrane moves, changing the capacitance between the plates. This
data is recalculated and output as an electrical signal to an integrated circuit. It is this signal that is
converted into the sound that we hear.</p>
        <p>Thanks to their design, MEMS microphones have the following advantages. Greater resistance to noise,
vibration and temperature changes due to the absence of unnecessary connecting elements. Multiple
MEMS microphones can be combined together to create a single array. Thanks to capacitive technology,
these microphone arrays can capture sound from a precisely defined direction, efectively canceling
echoes and background noise. Unlike other small microphones, such as electrets, MEMS microphones
include more additional elements, such as preamplifiers, various filters and analog-to-digital converters.
This means greater functionality while maintaining microscopic dimensions. Possibility of mounting
such devices on the board using soldering.</p>
        <p>Despite their many advantages, MEMS microphones are also not without their disadvantages. As we
wrote above, MEMS microphones are often used as part of arrays, which increases the sound capture
area, but at the same time reduces the service life of the devices. To work correctly, all microphones
must work in unison, but the likelihood of one of them breaking is much higher than an individual
device. Worse protection from moisture and dust than other microphones.</p>
        <p>Microphone arrays include two or more built-in microphones, to which is added a programmable
microprocessor designed to continuously determine the primary source of audio input and optimally
adjust the output to achieve the best sound quality.</p>
        <p>Let us highlight the most significant quality indicators of sound capture systems:
• useful signal/noise ratio, where the useful signal is the sound of the drone engine, and the noise
is background noise, the microphone’s own noise, and sounds from non-target sources;
• the shape of the radiation pattern and the ability of the system to change it depending on the
environment;
• ability to localize the source of a useful signal and measurement accuracy parameters.</p>
        <p>The most common way to build a signal capture unit is based on analog microphone arrays. A
description of the problems that arise when developing analog microphone arrays, as well as the
rationale for reducing the importance of the problem when using digital microphones in audio capture
systems, is given in table 1.</p>
        <p>Problem when designing an analog ar- Rationale for using digital mirophones
ray microphones
Sharp increase in cost Lack of a large number of auxiliary analog components
Reduced yield of suitable products due Reducing the total number of microcircuits and topological
comto the large number of components plexity leads to an increase in the percentage of usable products
due to the general laws of statistics
Increased development and debugging Implementation of algorithms in code and digital interface blocks,
costs which allows you to attract developers with less qualifications and
experience
High sensitivity to electromagnetic ra- The use of digital components is less sensitive to static failures and
diation and power quality degradation of power supply quality
Increased production cycle Less topological complexity guarantees the ability to produce a
product according to almost any modern technological standards,
making the launch process faster and cheaper
Increasing the testing cycle The digital implementation allows you to write synthetic tests and
generate input signals in the same way. Digital generators are more
flexible and low cost, and the testing and debugging process is
reduced to working with code</p>
        <p>As an alternative to existing approaches that have the disadvantages outlined above, the authors
of this work proposed to use an architecture built using digital MEMS microphones, which have
become widespread recently. The meter for these microphones is located on-chip, so its digital output
is minimally afected by the components that surround it. The most simple, inexpensive and perfect
solution in terms of signal capture parameters was developed, which consists of using digital MEMS
microphones and an Arduino microcomputer.</p>
        <p>When choosing a digital microphone for use in a linear diferential microphone array, it is important
to consider the following factors:
• Sensitivity: The sensitivity of the microphone is a measure of how well it can convert sound
waves into electrical signals. A higher sensitivity microphone will be able to pick up quieter
sounds, but it may also be more susceptible to noise.
• Signal-to-Noise Ratio (SNR): The SNR of the microphone is a measure of the ratio of the desired
signal (sound) to the undesired signal (noise). A higher SNR microphone will have less noise,
resulting in cleaner recordings.
• Dynamic range: The dynamic range of the microphone is the range of sound pressure levels that
it can accurately measure. A wider dynamic range microphone will be able to capture both very
loud and very quiet sounds without distortion.
• Linearity: The linearity of the microphone is a measure of how accurately it can reproduce the
input signal. A more linear microphone will produce recordings that are more faithful to the
original sound. In addition to the above factors, it is also important to consider the cost and
availability of the microphone when making a selection.</p>
        <p>Each digital MEMS microphone can be simplified into the model shown in figure 9. Input sound
vibrations are converted through a MEMS membrane into a weak electrical signal, which is then fed to
the input of amplifier A. The pre-amplified signal then passes through an analog low-pass filter (LPF),
which is necessary to protect against aliasing. The final element of signal processing in the microphone
is a 4th order Σ − Δ modulator, which converts the input analog signal into a one-bit digital stream.
The frequency of data bits from the output of the Σ − Δ modulator is equal to the frequency of the
input timing signal CLK and, as a rule, lies in the range from 1 to 4 MHz.</p>
        <p>In the time domain, the output of a Σ − Δ modulator is a jumbled collection of ones and zeros.
However, if we assign a value of 1.0 to each high logical level of the microphone output, and a value of
–1.0 to each low level and then perform a Fourier transform, we will obtain a spectrogram of the output
data from the microphone.</p>
        <p>Let’s look at the pins of a digital microphone. VDD – microphone power supply, GND – Ground,
CLK – input clock signal, synchronously with which the DATA line switches its DATA states. During
one half of the CLK cycle this pin is in a high impedance state, and during the second half it serves
as a pin for reading data from the Σ − Δ output of the microphone modulator. / – this pin is
used to control switching of the DATA line. If / is connected to VDD, then after some time after
detecting the rising edge of the CLK signal, the DATA pin goes into a high impedance state, and after
the arrival of the falling edge of the CLK signal, the DATA pin is connected to the Σ − Δ output of the
microphone modulator. If / is connected to GND, the edges of the CLK signal, along which the
DATA line switches, are reversed (figure 10).</p>
        <p>To isolate the audio frequency band signal, the data from the microphone must be filtered and
resampled at a lower frequency (usually 50–128 times lower than the sampling frequency of the Σ − Δ
modulator). A digital low-pass filter filters out external noise and the microphone’s own noise outside
the operating band ( &gt;  /2 ) to protect against aliasing, and also makes it possible to reduce
the data repetition rate. In figure 11 presents one of the possible options for processing a one-bit data
stream from a microphone, implemented in software on a DSP or in hardware in audio codecs. Shown
in figure 11, the sampling frequency compression circuit (compressor) lowers the sampling frequency
due to the fact that from every  samples of the filtered signal ( ),  ˘1 sample is discarded. The
input and output of the converter shown in figure 8 are related by the following expression:
() = ( ) =
∞
∑︁
=−∞
ℎ()( − ).</p>
        <p>(15)</p>
        <p>[] = − (− 1[]),
where [] contains in each term the relative change in the signal in the form of 1 bit with a sign, which
is specified by the transition. A negative increment is a transition from 1 to 0, a positive increment is
from 0 to 1. Repeating ones increases the overall amplitude of the signal, and repeating zeros decreases
(figure 12).</p>
        <p>A mathematical model for pulse density modulation can be obtained using a delta-sigma modulator
model. In the discrete frequency domain, the operation of a delta-sigma modulator can be described by
the formula</p>
        <p>() =  () + ()(1 − − 1),
where (),  () are the signal spectra at the input and output of the modulator; () is the sampling
error of the delta-sigma modulator; 1 − − 1 is high-pass filter. As a result of transforming the formula,
we get
() = ()[ () − ()− 1)]
1 − − 1</p>
        <p>According to this formula, the error () reduces the value of the signal at the output () in the
low-frequency region and increases it in the high-frequency region, as a result of which the quantization
noise spectrum shifts predominantly to the high-frequency region.
1</p>
        <p>.
(16)
(17)
(18)</p>
        <p>Let [] be a sample of the signal at the input of the modulator in the time domain, and [] be a
sample of the output signal, then, using the inverse -transform, we can proceed to the expression
where
(19)
(20)
(21)
[] = [] + [] − [ − 1],
[] = [] − [] + [ − 1].</p>
        <p>The signal from the output signal sample [] is represented as 1 bit and takes values ± 1, and is
implemented so that the value of the current quantization error [] is minimal. In this case, the
quantization error [] of each sample appears at the device input during the subsequent sample.</p>
        <p>When implementing frequency converters in software, a finite impulse response (FIR) filter or an
Infinite impulse response (IIR) filter can be used as a digital LPF. Developers should be very careful
when choosing the type of filter, its length and bit depth, since the performance of the entire system as a
whole directly depends on this. A correctly calculated and implemented decimator (frequency converter)
in some cases will significantly reduce the cost of products and increase its technical characteristics.</p>
        <p>As a second option, audio codecs adapted for this can be used to convert data from the output of a
digital microphone, which will significantly reduce product development time. For example, Analog
Devices ofers the ADAU1361 and ADAU1761 codecs, which are suitable for the ADMP521 microphones.
In our work we used a microphone ADMP521. However, the process of creating digital audio devices
becomes simple in terms of hardware implementation and complex in terms of writing programs for
the microcontrollers used.</p>
        <p>Next, we conducted a simulation and computational experiment of a uniform linear array of 2
omnidirectional microphones using Matlab Sensor Array Analyzer.</p>
        <p>The following model parameters were used. The distance between the microphones is 20 mm. The
board has two MEMS microphones spaced 20 mm apart. This spacing is ideal for detecting acoustic
events. Additionally, the 20 mm spacing is equivalent to 8 · 2.54 mm, which makes it suitable for
DIP (Dual In-line Package) – a type of housing for microcircuits, electronic modules and some other
electronic components. Experimentally, a distance of 0.017 m was determined for the formation of a
bi-directional pattern.</p>
        <p>Next, we conducted a simulation and computational experiment of a uniform linear array of 2
omnidirectional microphones using Matlab Sensor Array Analyzer. The following model parameters
were used. The distance between the microphones is 17 mm. The speed of sound is 343 m/s, the signal
frequency is 10 kHz. As a result, we obtained the parameters listed in table 2.</p>
        <p>The Matlab script is listed below:
% Create a Uniform Linear Array Object
Array = phased.ULA(’NumElements’,2, ’ArrayAxis’,’y’);
Array.ElementSpacing = 0.017;</p>
        <p>The resulting pattern has the shape of a bi-directional (figure 13). As you can see, the design of
the grille allows you to create a grille with the main lobes directed at -90 and 90 degrees. To form a
cardiode radiation pattern, as mentioned above, it is necessary to use delay-and-sum and filter-and-sum
algorithms. The meaning of these algorithms is that microphone signals are added with diferent
delays (diferent phase shifts), aligning the phases of signals coming from the selected direction (source
localization) for each frequency. In this case, the beamforming algorithm makes it possible to amplify
the signals generated by sound coming from the selected direction, i.e. performs a kind of focusing of
sounds.</p>
        <p>ADMP521 microphones were connected to the ADAU1761 codec in accordance with the technical
specifications of both products (figure 14).</p>
        <p>A model was also created based on an array of four omnidirectional microphones located at a distance
of 20 mm (figure 15). The following model parameters were used. The distance between the microphones
is 17 mm. The speed of sound is 343 m/s, the signal frequency is 10 kHz. As a result, we obtained the
parameters listed in table 3.</p>
        <p>Array characteristic</p>
        <sec id="sec-4-2-1">
          <title>The Matlab script has the following form:</title>
          <p>% Create a Uniform Linear Array Object
Array = phased.ULA(’NumElements’,4, ’ArrayAxis’,’y’);
Array.ElementSpacing = 0.017;
Array.Taper = ones(1,4).’;
% Create an omnidirectional microphone element
Elem = phased.OmnidirectionalMicrophoneElement;
Elem.FrequencyRange = [0 10000];
Array.Element = Elem;</p>
          <p>So, during the computational experiment, we built 2 linear microphone arrays with bi-directionality.
The directionality of these arrays can be easily converted to unidirectional (cardioid) using known
algorithms or hardware (codecs). The tuning of the circuit to create cardioid directivity will be considered
in further studies.</p>
        </sec>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>5. Discussion</title>
      <p>The following questions require additional discussion and clarification: the dependence of the azimuthal
pattern of the proposed linear microphone arrays on the source frequency; the choice of analog or
digital MEMS microphones and labor-intensiveness in the development of a microphone array; the use
of directional microphones instead of omnidirectional; peculiarities of localization of a moving sound
source (Doppler efect, reflection from obstacles, etc.); higher-order diferential beam array formers;
signal processing algorithms of microphone arrays.</p>
    </sec>
    <sec id="sec-6">
      <title>6. Conclusions</title>
      <p>The study looks at setting up a microphone array to determine the position of a UAV (unmanned aerial
vehicle) based solely on the sound of its engines. The location of the microphones plays a crucial role for
accurate localization. A mathematical model of pulse density modulation of a digital MEMS microphone
is also considered. This work shows the dependence of the eficiency of a diferential array of first-order
microphones on frequency. Based on the presented dependence of the directivity on frequency and the
instability model of the microphone parameters, a rational operating frequency range for the normal
functioning of the microphone array can be determined.</p>
      <p>A model of a linear microphone array based on MEMS omnidirectional microphones is proposed,
which with a certain geometrical arrangement give a bi-directional pattern, which, in principle, can
be easily transformed into a unidirectional one with the use of special algorithms or hardware (for
example, ADAU1761 codecs). Refinement of the circuit to achieve cardioid directivity will be addressed
in forthcoming research.</p>
    </sec>
    <sec id="sec-7">
      <title>7. Author contributions</title>
      <p>Conceptualization, methodology – Andrii V. Riabko, Oksana V. Zaika; setting tasks, conceptual analysis
– Tetiana A. Vakaliuk, Oksana V. Zaika; development of the model – Andrii V. Riabko, Valerii V.
Kontsedailo; software development, verification – Andrii V. Riabko, Roman P. Kukharchuk; analysis of
results, visualization – Roman P. Kukharchuk, Tetiana A. Vakaliuk; drafting of the manuscript – Valerii
V. Kontsedailo, reviewing and editing – Tetiana A. Vakaliuk.</p>
      <p>All authors have read and approved the published version of this manuscript.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>M.</given-names>
            <surname>Cobos</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Antonacci</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Alexandridis</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Mouchtaris</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Lee</surname>
          </string-name>
          ,
          <article-title>A Survey of Sound Source Localization Methods in Wireless Acoustic Sensor Networks</article-title>
          ,
          <source>Wireless Communications and Mobile Computing</source>
          <year>2017</year>
          (
          <year>2017</year>
          )
          <article-title>3956282</article-title>
          . doi:
          <volume>10</volume>
          .1155/
          <year>2017</year>
          /3956282.
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>A. R.</given-names>
            <surname>Petrosian</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R. V.</given-names>
            <surname>Petrosyan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>I. A.</given-names>
            <surname>Pilkevych</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. S.</given-names>
            <surname>Graf</surname>
          </string-name>
          ,
          <article-title>Eficient model of PID controller of unmanned aerial vehicle</article-title>
          ,
          <source>Journal of Edge Computing</source>
          <volume>2</volume>
          (
          <year>2023</year>
          )
          <fpage>104</fpage>
          -
          <lpage>124</lpage>
          . doi:
          <volume>10</volume>
          .55056/jec.593.
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>M.</given-names>
            <surname>Risoud</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.-N.</given-names>
            <surname>Hanson</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Gauvrit</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Renard</surname>
          </string-name>
          , P.-E. Lemesre,
          <string-name>
            <given-names>N.-X.</given-names>
            <surname>Bonne</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Vincent</surname>
          </string-name>
          , Sound source localization,
          <source>European Annals of Otorhinolaryngology, Head and Neck Diseases</source>
          <volume>135</volume>
          (
          <year>2018</year>
          )
          <fpage>259</fpage>
          -
          <lpage>264</lpage>
          . doi:
          <volume>10</volume>
          .1016/j.anorl.
          <year>2018</year>
          .
          <volume>04</volume>
          .009.
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>W. A.</given-names>
            <surname>Yost</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. T.</given-names>
            <surname>Pastore</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Zhou</surname>
          </string-name>
          ,
          <source>Sound Source Localization Is a Multisystem Process</source>
          , Springer International Publishing, Cham,
          <year>2021</year>
          , pp.
          <fpage>47</fpage>
          -
          <lpage>79</lpage>
          . doi:
          <volume>10</volume>
          .1007/978-3-
          <fpage>030</fpage>
          -57100-
          <issue>9</issue>
          _
          <fpage>3</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>E.</given-names>
            <surname>King</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Tatoglu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Iglesias</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Matriss</surname>
          </string-name>
          ,
          <article-title>Audio-visual based non-line-of-sight sound source localization: A feasibility study</article-title>
          ,
          <source>Applied Acoustics</source>
          <volume>171</volume>
          (
          <year>2021</year>
          )
          <article-title>107674</article-title>
          . doi:
          <volume>10</volume>
          .1016/j.apacoust.
          <year>2020</year>
          .
          <volume>107674</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>P.</given-names>
            <surname>Chiariotti</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Martarelli</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Castellini</surname>
          </string-name>
          ,
          <article-title>Acoustic beamforming for noise source localization - reviews, methodology and applications</article-title>
          ,
          <source>Mechanical Systems and Signal Processing</source>
          <volume>120</volume>
          (
          <year>2019</year>
          )
          <fpage>422</fpage>
          -
          <lpage>448</lpage>
          . doi:
          <volume>10</volume>
          .1016/j.ymssp.
          <year>2018</year>
          .
          <volume>09</volume>
          .019.
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>T.</given-names>
            <surname>Yamada</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Itoyama</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Nishida</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Nakadai</surname>
          </string-name>
          ,
          <article-title>Sound Source Tracking by Drones with Microphone Arrays</article-title>
          , in: 2020
          <source>IEEE/SICE International Symposium on System Integration (SII)</source>
          ,
          <year>2020</year>
          , pp.
          <fpage>796</fpage>
          -
          <lpage>801</lpage>
          . doi:
          <volume>10</volume>
          .1109/SII46433.
          <year>2020</year>
          .
          <volume>9026185</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>A. D.</given-names>
            <surname>Firoozabadi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Irarrazaval</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Adasme</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Zabala-Blanco</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Palacios-Játiva</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Durney</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Sanhueza</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Azurdia-Meza</surname>
          </string-name>
          ,
          <article-title>Three-dimensional sound source localization by distributed microphone arrays</article-title>
          ,
          <source>in: 2021 29th European Signal Processing Conference (EUSIPCO)</source>
          ,
          <year>2021</year>
          , pp.
          <fpage>196</fpage>
          -
          <lpage>200</lpage>
          . doi:
          <volume>10</volume>
          .23919/EUSIPCO54536.
          <year>2021</year>
          .
          <volume>9616326</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>Y.</given-names>
            <surname>Sasaki</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Tanabe</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Takemura</surname>
          </string-name>
          ,
          <article-title>Probabilistic 3D sound source mapping using moving microphone array</article-title>
          ,
          <source>in: 2016 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS)</source>
          ,
          <year>2016</year>
          , pp.
          <fpage>1293</fpage>
          -
          <lpage>1298</lpage>
          . doi:
          <volume>10</volume>
          .1109/IROS.
          <year>2016</year>
          .
          <volume>7759214</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <surname>M. C. Catalbas</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          <string-name>
            <surname>Yildirim</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Gulten</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          <string-name>
            <surname>Kurum</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          <string-name>
            <surname>Dobrišek</surname>
          </string-name>
          ,
          <article-title>Estimation of Trajectory and Location for Mobile Sound Source</article-title>
          ,
          <source>International Journal of Advanced Computer Science and Applications</source>
          <volume>7</volume>
          (
          <year>2016</year>
          ). doi:
          <volume>10</volume>
          .14569/IJACSA.
          <year>2016</year>
          .
          <volume>070934</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>K.</given-names>
            <surname>Hoshiba</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Washizaki</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Wakabayashi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Ishiki</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Kumon</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Bando</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Gabriel</surname>
          </string-name>
          , K. Nakadai,
          <string-name>
            <given-names>H. G.</given-names>
            <surname>Okuno</surname>
          </string-name>
          ,
          <article-title>Design of UAV-Embedded Microphone Array System for Sound Source Localization in Outdoor Environments</article-title>
          ,
          <source>Sensors</source>
          <volume>17</volume>
          (
          <year>2017</year>
          )
          <article-title>2535</article-title>
          . doi:
          <volume>10</volume>
          .3390/s17112535.
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>T.</given-names>
            <surname>Tachikawa</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Yatabe</surname>
          </string-name>
          ,
          <string-name>
            <surname>Y. Oikawa,</surname>
          </string-name>
          <article-title>3D sound source localization based on coherence-adjusted monopole dictionary and modified convex clustering</article-title>
          ,
          <source>Applied Acoustics</source>
          <volume>139</volume>
          (
          <year>2018</year>
          )
          <fpage>267</fpage>
          -
          <lpage>281</lpage>
          . doi:
          <volume>10</volume>
          .1016/j.apacoust.
          <year>2018</year>
          .
          <volume>04</volume>
          .033.
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <given-names>M.</given-names>
            <surname>Wakabayashi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H. G.</given-names>
            <surname>Okuno</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Kumon</surname>
          </string-name>
          ,
          <article-title>Drone audition listening from the sky estimates multiple sound source positions by integrating sound source localization and data association</article-title>
          ,
          <source>Advanced Robotics</source>
          <volume>34</volume>
          (
          <year>2020</year>
          )
          <fpage>744</fpage>
          -
          <lpage>755</lpage>
          . doi:
          <volume>10</volume>
          .1080/01691864.
          <year>2020</year>
          .
          <volume>1757506</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <given-names>L.</given-names>
            <surname>Carneiro</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Berry</surname>
          </string-name>
          ,
          <article-title>Three-dimensional sound source diagnostic using a spherical microphone array from multiple capture positions</article-title>
          ,
          <source>Mechanical Systems and Signal Processing</source>
          <volume>199</volume>
          (
          <year>2023</year>
          )
          <article-title>110455</article-title>
          . doi:
          <volume>10</volume>
          .1016/j.ymssp.
          <year>2023</year>
          .
          <volume>110455</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <given-names>S.</given-names>
            <surname>Kita</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Kajikawa</surname>
          </string-name>
          ,
          <article-title>Fundamental study on sound source localization inside a structure using a deep neural network and computer-aided engineering</article-title>
          ,
          <source>Journal of Sound and Vibration</source>
          <volume>513</volume>
          (
          <year>2021</year>
          )
          <article-title>116400</article-title>
          . doi:
          <volume>10</volume>
          .1016/j.jsv.
          <year>2021</year>
          .
          <volume>116400</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [16]
          <string-name>
            <surname>F. R.</surname>
          </string-name>
          do Amaral,
          <string-name>
            <given-names>J.</given-names>
            <surname>Rico</surname>
          </string-name>
          ,
          <string-name>
            <surname>M. A. F. de Medeiros</surname>
          </string-name>
          ,
          <article-title>Design of microphone phased arrays for acoustic beamforming</article-title>
          ,
          <source>Journal of the Brazilian Society of Mechanical Sciences and Engineering</source>
          <volume>40</volume>
          (
          <year>2018</year>
          )
          <article-title>354</article-title>
          . doi:
          <volume>10</volume>
          .1007/s40430-018-1275-5.
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [17]
          <string-name>
            <given-names>F.</given-names>
            <surname>Grondin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Michaud</surname>
          </string-name>
          ,
          <article-title>Lightweight and optimized sound source localization and tracking methods for open and closed microphone array configurations</article-title>
          ,
          <source>Robotics and Autonomous Systems</source>
          <volume>113</volume>
          (
          <year>2019</year>
          )
          <fpage>63</fpage>
          -
          <lpage>80</lpage>
          . doi:
          <volume>10</volume>
          .1016/j.robot.
          <year>2019</year>
          .
          <volume>01</volume>
          .002.
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          [18]
          <string-name>
            <given-names>C.</given-names>
            <surname>Rascon</surname>
          </string-name>
          ,
          <string-name>
            <surname>I. Meza</surname>
          </string-name>
          ,
          <article-title>Localization of sound sources in robotics: A review</article-title>
          ,
          <source>Robotics and Autonomous Systems</source>
          <volume>96</volume>
          (
          <year>2017</year>
          )
          <fpage>184</fpage>
          -
          <lpage>210</lpage>
          . doi:
          <volume>10</volume>
          .1016/j.robot.
          <year>2017</year>
          .
          <volume>07</volume>
          .011.
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          [19]
          <string-name>
            <surname>M. J. Bianco</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          <string-name>
            <surname>Gerstoft</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          <string-name>
            <surname>Traer</surname>
            , E. Ozanich,
            <given-names>M. A.</given-names>
          </string-name>
          <string-name>
            <surname>Roch</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          <string-name>
            <surname>Gannot</surname>
          </string-name>
          , C.
          <article-title>-A. Deledalle, Machine learning in acoustics: Theory and applications</article-title>
          ,
          <source>The Journal of the Acoustical Society of America</source>
          <volume>146</volume>
          (
          <year>2019</year>
          )
          <fpage>3590</fpage>
          -
          <lpage>3628</lpage>
          . doi:
          <volume>10</volume>
          .1121/1.5133944.
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          [20]
          <string-name>
            <given-names>H.</given-names>
            <surname>Niu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Gong</surname>
          </string-name>
          , E. Ozanich,
          <string-name>
            <given-names>P.</given-names>
            <surname>Gerstoft</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <article-title>Deep-learning source localization using multi-frequency magnitude-only data</article-title>
          ,
          <source>The Journal of the Acoustical Society of America</source>
          <volume>146</volume>
          (
          <year>2019</year>
          )
          <fpage>211</fpage>
          -
          <lpage>222</lpage>
          . doi:
          <volume>10</volume>
          .1121/1.5116016.
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          [21]
          <string-name>
            <given-names>W.</given-names>
            <surname>He</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Motlicek</surname>
          </string-name>
          ,
          <string-name>
            <surname>J.-M. Odobez</surname>
          </string-name>
          ,
          <article-title>Deep Neural Networks for Multiple Speaker Detection and Localization</article-title>
          , in: 2018
          <source>IEEE International Conference on Robotics and Automation (ICRA)</source>
          ,
          <year>2018</year>
          , pp.
          <fpage>74</fpage>
          -
          <lpage>79</lpage>
          . doi:
          <volume>10</volume>
          .1109/ICRA.
          <year>2018</year>
          .
          <volume>8461267</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          [22]
          <string-name>
            <given-names>A.</given-names>
            <surname>Ebrahimkhanlou</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Salamone</surname>
          </string-name>
          ,
          <article-title>Single-Sensor Acoustic Emission Source Localization in PlateLike Structures Using Deep Learning, Aerospace 5 (</article-title>
          <year>2018</year>
          )
          <article-title>50</article-title>
          . doi:
          <volume>10</volume>
          .3390/aerospace5020050.
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          [23]
          <string-name>
            <given-names>S.</given-names>
            <surname>Adavanne</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Politis</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Nikunen</surname>
          </string-name>
          , T. Virtanen,
          <article-title>Sound Event Localization and Detection of Overlapping Sources Using Convolutional Recurrent Neural Networks</article-title>
          ,
          <source>IEEE Journal of Selected Topics in Signal Processing</source>
          <volume>13</volume>
          (
          <year>2019</year>
          )
          <fpage>34</fpage>
          -
          <lpage>48</lpage>
          . doi:
          <volume>10</volume>
          .1109/JSTSP.
          <year>2018</year>
          .
          <volume>2885636</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          [24]
          <string-name>
            <given-names>G.</given-names>
            <surname>Chardon</surname>
          </string-name>
          ,
          <article-title>Theoretical analysis of beamforming steering vector formulations for acoustic source localization</article-title>
          ,
          <source>Journal of Sound and Vibration</source>
          <volume>517</volume>
          (
          <year>2022</year>
          )
          <article-title>116544</article-title>
          . doi:
          <volume>10</volume>
          .1016/j.jsv.
          <year>2021</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>