<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Verification of Animal External Ear Model-based Microphone Sensor for Hazardous Sound Detection</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Miho Yamada</string-name>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Takeshi Kumaki</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Kyosuke Kageyama</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Dept. of Electrical, Electronic and Communication Engineering, Kindai University</institution>
          ,
          <addr-line>3-4-1 Kowakae, Higashi-Osaka, Osaka</addr-line>
          ,
          <country country="JP">Japan</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Dept. of Electronic and Computer Engineering, Ritsumeikan University</institution>
          ,
          <addr-line>1-1-1 Noji-Higashi, Kusatsu, Shiga</addr-line>
          ,
          <country country="JP">Japan</country>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>Graduate School of Science and Engineering, Kindai University</institution>
          ,
          <addr-line>3-4-1 Kowakae, Higashi-Osaka, Osaka</addr-line>
          ,
          <country country="JP">Japan</country>
        </aff>
      </contrib-group>
      <fpage>69</fpage>
      <lpage>75</lpage>
      <abstract>
        <p>Hazardous sounds represent a particular type of danger signal and should be heard clearly to ensure safety. Hearing-impaired people can be involved in accidents. This is exacerbated by the fact that they may have dificulty hearing hazardous sounds. Therefore, they spend time with a hearing-assistance dog in their daily lives. However, the number of hearing-assistance dogs has decreased recently, and they are less known in society and they may even be prevented from entering public facilities. Therefore, an animal external ear model-based mobile device that can detect hazardous sounds has been proposed. This device can alert the hearing-impaired person. In this paper, the animal external attachment with a humanand a cat-external ear model-verified for hazardous sound detection accuracy by microphone sensor is presented. Microphone sensors with and without a human- and a cat-external ear model are used. An ambulance siren and a bicycle bell are used as the hazardous sounds, and the spectral envelopes of each recorded sound are compared. The results show that the animal external ear model-based microphone sensor can more efectively detect the frequency features of the ambulance siren and the bicycle bell. A microphone sensor attached an external attachment with a human-external ear model can detect 2 ∼ 3 kHz frequency features with the highest sound pressure compared to the other microphones, regardless of the distance between the microphone sensor and the sound source. Also, a microphone sensor attached an external attachment with a cat-external ear model can detect wide range of frequencies with the highest sound pressure compared to the other microphone sensors, regardless of the distance between the microphone sensor and the sound source.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;Hazardous sound detection</kwd>
        <kwd>External ear</kwd>
        <kwd>Hearing-impaired people</kwd>
        <kwd>Formant</kwd>
        <kwd>Spectral envelope</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>
        Various sounds are encountered in our daily lives. Among them, hazardous sounds pose possible
danger signals. Our safety might depend on being able to hear hazardous sounds such as an
emergency bell or the siren of an emergency vehicle. However, hearing-impaired people cannot
hear these hazardous sounds. In 2016, there were about 300,000 hearing-impaired people in
Japan [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. The degree of hearing is diferent for each hearing-impaired person. But many
hearing-impaired people can be impacted by accidents. Therefore, some hearing-impaired
people spend time with a hearing-assistance dog. This dog supports hearing-impaired people
and warns them of danger by touching or pulling them [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]. The hearing-assistance dogs are
identified by wearing a cape with words like “Hearing-assistance dog”. Thus, people can
understand that there is hearing-impaired person in the area [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]. However, the number of
hearing-assistance dogs has decreased recently, and only 52 hearing-assistance dogs work in
Japan currently [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ]. Additionally, hearing-assistance dogs aren’t recognized well in society, and
they are prevented from entering public facilities such as restaurants [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ]. Thanks to having two
ears, animals can judge the direction and source of sounds. The external ear consists of two
structures: a pinna and an ear canal. The pinna has a sound collection efect and the canal has a
resonance efect [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ]. The evolution of animal’s ears over time has supported them to enable the
identification of hazardous sounds. Therefore, animals can detect enemies or danger in real
time. Generally, a human, a dog, or a cat can hear sounds of approximately 0.020 ∼ 20 kHz,
0.065 ∼ 50 kHz, and 0.05 ∼ 65 kHz sounds, respectively [
        <xref ref-type="bibr" rid="ref7 ref8">7, 8</xref>
        ]. A cat can hear frequencies three
times wider than humans.
      </p>
      <p>
        Our new method can detect more hazardous sounds than only a microphone sensor, and alert
a hearing-impaired person to hazardous sounds, like the hearing-assistance dog. An animal
external ear model-based mobile device for hazardous sound detection has been proposed [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ].
This proposed device implements an animal external attachment to a microphone sensor to
pick up more detail of sounds. This device can detect hazardous sounds and warn the
hearingimpaired persons who are wearing it. In this paper, the animal external attachment follows a
human- and cat-external ear model, and has a verified hazardous sounds detection accuracy by
the microphone sensor.
      </p>
      <p>The rest of the paper is organized as follows. Section 2 outlines existing technologies.
Section 3 explains the animal external ear model-based mobile device for hazardous sound
detection. Section 4 describes how to process recorded sounds and its spectral envelope. Section
5 explains the experimental method. Section 6 shows the spectral envelopes of recorded sounds
as experimental results, and Section 7 gives our conclusions.</p>
    </sec>
    <sec id="sec-2">
      <title>2. Existing technologies</title>
      <p>
        There are four existing technologies to detect hazardous sounds. The first system detects
hazardous sounds and situations using clustering by the complete linkage method and probability
modeling of daily sounds [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ]. These systems need to perform a lot of calculations to process
a huge amount of data. And, the second system that monitor the elderly using sound can
manage health and detect danger for the elderly to monitor indoor living environment mainly
from acoustic signal observed using a microphone sensor. In this system, acoustic events are
detected from audio signals observed by microphone sensor using machine learning and obtains
information about daily life [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ]. However, incorporating machine learning may increase
processing requirements and power consumption. Therefore, These systems making it dificult
to implement on a mobile device. The third system can detect and determine daily sounds and
non-daily sounds which are possibly hazardous using a microphone array [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ]. This system is
stationary and has a height of about 3 m. So, it is dificult to carry and use. The fourth system is
hearing aid. A hearing aid helps the hearing-impaired people hear better. However, they have
some disadvantages, such as high performance at a very high price, and the mufled sound and
echoes of one’s own voice can be bothered.
      </p>
    </sec>
    <sec id="sec-3">
      <title>3. The animal external ear model-based mobile device for hazardous sound detection</title>
      <p>The animal external ear model-based mobile device for hazardous sound detection is explained
in this section. Overview of this device is shown in Fig. 1 (i). The hazardous sound is an
alarm sound, siren, and emergency bell of an emergency vehicle which is designed to make
people notice some danger. This device continuously pics up the sounds from the microphone
sensor attached to an animal external attachment. The animal external attachment is shaped
somewhat like the external ear of an animal (Fig. 1 (i) (a)). The picked-up sound is analyzed
with a microcomputer and is determined to be hazardous sound (Fig. 1 (i) (b)). If the sound is
judged to be a hazardous sound, the device notifies the hearing-impaired person using LED or
other methods (Fig. 1 (i) (c)). This device is implemented as a mobile device.</p>
    </sec>
    <sec id="sec-4">
      <title>4. Flow of recorded sound processing</title>
      <p>
        Figure 1 (ii) shows the flow of recorded sounds processing in this experiment. The purpose
of this process is to derive the spectral envelope. A sound is recorded by the microphone
sensor. And, the recorded data are processed by a short-time Fourier transform (STFT) to find
a spectrogram [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ]. The power spectrum defined as log|X(k)|2, is extracted from one row of
the spectrogram. And a discrete Fourier transform (DFT) is applied to each row to obtain the
cepstrum. The cepstrum defined as CEP(k), is considered to be in the form of
      </p>
      <p>CEP(k) = DFT(log|X(k)|2)
From the cepstrum, the quefrency above a certain threshold is set to 0 to extract the low
quefrency components of the cepstrum. Quefrency is a twisted phrase of “frequency” and
horizontal axis of cepstrum. Such quefrency includes components of the spectral envelope,
which is defined as log(|highpass(k)|2). Then, an inverse discrete Fourier transform (IDFT) is
applied to the resulting data to find the spectral envelope, which configured as follows SPE(k):</p>
      <p>
        SPE(k) = IDFT(log|highpass(k)|2)
The spectral envelope includes formants, which are frequency characteristics of a specific
frequency region that characterize the sound. In particular, the spectral envelope is useful
when indicating vocal-tract characteristics of a human voice. By observing formants can get
the frequency characteristics of a vocal tract [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ]. In speech recognition, Japanese vowels have
long been discriminated by their formant diferences. When formants are sorted from lowest to
highest frequency, the first and second formants are important clues to discriminate Japanese
vowels. Non-human sound sources do not have a human vocal tract. However, it is considered
that by observing the spectral envelope, it is possible to confirm the frequency characteristics
of any sound.
      </p>
    </sec>
    <sec id="sec-5">
      <title>5. Experimental method</title>
      <p>
        intervals of 0.65 s. The bicycle bell is rung by hand. The sound sources are recorded for 15 s,
and the recorded sounds are processed in section 3. Three microphone sensors are connected to
a SPRESENSE [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ], which has a main board (CXD5602PWBMAIN1) and an extension board
(CXD5602PWBEXT1). SPRESENSE can record sounds at a sampling frequency of 48 kHz.
      </p>
    </sec>
    <sec id="sec-6">
      <title>6. Experimental result</title>
      <sec id="sec-6-1">
        <title>6.1. The ambulance siren</title>
        <p>Figure 3 shows the spectral envelope of the ambulance siren for distances of 1 m and 3 m,
respectively. The horizontal axis is frequency and the vertical axis is power. The power is
decreased by about 3 dB due to change the distance between the microphone sensors and the
sound source from 1 m to 3 m in Fig. 3. It is shown that each formant of the ambulance siren
can be detected at around 0.77 kHz and 0.96 kHz at intervals of 0.65 s regardless of the distance
between the microphone sensors and the sound source from Fig. 3. Also, Fig. 3 shows that
MIC_human detects formant of 2 ∼ 3 kHz frequency features with the highest sound pressure
compared to the other microphone sensors, regardless of the distance between the microphone
sensor and the sound source. On the other hand, MIC_cat detects formants of wide range of
frequencies with the highest sound pressure in the range of 0 ∼ 7.5 kHz compared to the other
microphone sensors, regardless of the distance between the microphone sensor and the sound
source. And the sound pressure detected by MIC_human and MIC_cat are almost greater than
the sound pressure detected by MIC_only.</p>
      </sec>
      <sec id="sec-6-2">
        <title>6.2. The bicycle bell</title>
        <p>The spectral envelope of the bicycle bell is shown in Fig. 4 for distances of 1 m and 3 m,
respectively. The horizontal axis is frequency and the vertical axis is power. The power is
decreased by about 1 dB due to change the distance between the microphone sensors and the
sound source from 1 m to 3 m in Fig. 4. Figure 4 shows that each formant of the bicycle bell can
be detected five places at around 2.0 kHz, 3.6 kHz, 5.0 kHz, 8.5 kHz, and 15 kHz, regardless of
the distance between the microphone sensors and the sound source. Also, Fig. 4 shows that
MIC_human detects formant of 3.6 kHz with the highest sound pressure compared to the other
microphones regardless of the distance between the microphone sensor and the sound source.
On the other hand, MIC_cat detects formants of 5.0 kHz, 8.5 kHz, and 15 kHz with the highest
sound pressure compared to the other microphones, regardless of the distance between the
microphone sensor and the sound source. In other words, MIC_human can detect approximately
2 ∼ 3 kHz frequency features and MIC_cat can detect a wide range of frequency features in the
range of 0 ∼ 18 kHz from Fig. 4. And, the sound pressure detected by MIC_human and MIC_cat
are almost greater than the sound pressure detected by MIC_only.</p>
      </sec>
    </sec>
    <sec id="sec-7">
      <title>7. Conclusion</title>
      <p>Our proposed method can detect more hazardous sounds more than only a microphone sensor,
and alert a hearing-impaired person to hazardous sounds, like the hearing-assistance dog. In
this paper, the animal external attachment with a human- and a cat-external ear model-verified
hazardous sounds detection accuracy by the microphone sensor. Three diferent microphone
sensors are compared, and the spectral envelopes of two recorded sounds are detected. The
results show a microphone sensor attached an external attachment with a human-external ear
model can detect 2 ∼ 3 kHz frequency features with the highest sound pressure compared to the
other microphone sensors, regardless of the distance between the microphone sensor and the
sound source. Also, a microphone sensor attached an external attachment with a cat-external
ear model can detect wide range of frequencies. This enables the highest sound pressure possible
compared to the other microphone sensors, regardless of the distance between the microphone
sensor and the sound source. Thus, the animal external ear model-based microphone sensor
can more efectively detect the frequency features of the ambulance siren and the bicycle bell.
Therefore, it makes easier to detect hazardous sounds by adapting the external attachment of
the microphone sensor to match each hazardous sound in our proposed device. In the future,
experiments will be conducted in various environments, including noisy environments and
changing temperatures. In addition, various external ear models will be verified to detect
hazardous sounds more efectively by changing the shape and material of external attachment.
Furthermore, efectiveness of our proposed method which detects other hazardous sounds that
are common in daily life is verified.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1] https://www.mhlw.go.jp/toukei/list/dl/seikatsu_chousa_c_h28.
          <string-name>
            <surname>pdf</surname>
          </string-name>
          (Accessed, Apr.
          <year>2024</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2] https://www.hojyoken.or.jp/what/hearingdog (Accessed, Apr.
          <year>2024</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3] https://www.hearingdogjp.org/about/workandrole (Accessed, Apr.
          <year>2024</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4] https://www.jsdrc.jp/hojoken/chodoken_suu/ (Accessed, Apr.
          <year>2024</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5] https://grapee.jp/99344 (Accessed, May.
          <year>2024</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>J.</given-names>
            <surname>Melloui</surname>
          </string-name>
          ,
          <string-name>
            <given-names>O.</given-names>
            <surname>Bouattane</surname>
          </string-name>
          , and
          <string-name>
            <given-names>J.</given-names>
            <surname>Bakkoury</surname>
          </string-name>
          , “
          <article-title>Study of the efect of a cause of tinnitus on the resonant frequency of the outer ear</article-title>
          ,
          <source>” 2020 1st International Conference on Innovative Research in Applied Science, Engineering and Technology (IRASET)</source>
          , pp.
          <fpage>1</fpage>
          -
          <lpage>4</lpage>
          ,
          <year>2020</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>S. L.</given-names>
            <surname>Mak</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S. L.</given-names>
            <surname>Au</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W. F.</given-names>
            <surname>Tang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C. H.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C. C.</given-names>
            <surname>Lee</surname>
          </string-name>
          , and
          <string-name>
            <given-names>W. H.</given-names>
            <surname>Chiu</surname>
          </string-name>
          , “
          <article-title>A Study on Hearing Hazards and sound measurement for Dogs,” 2022 IEEE International Symposium on Product Compliance Engineering - Asia (ISPCE-ASIA)</article-title>
          , pp.
          <fpage>1</fpage>
          -
          <lpage>4</lpage>
          ,
          <year>2022</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8] https://www.signia.net/ja-jp/blog/local/ja-jp/
          <article-title>cat-dog-human-hearing/ (Accessed</article-title>
          , Apr.
          <year>2024</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>M.</given-names>
            <surname>Yamada</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Kumaki</surname>
          </string-name>
          , and
          <string-name>
            <given-names>K.</given-names>
            <surname>Kageyama</surname>
          </string-name>
          , “
          <article-title>Proposal of hazard sound detection system based on the external ear model</article-title>
          ,
          <source>” IPSJ SIG Technical Report</source>
          , Vol.
          <volume>2023</volume>
          <source>-EMB-64, No. 9</source>
          , pp.
          <fpage>1</fpage>
          -
          <lpage>4</lpage>
          , Nov.
          <year>2023</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>K.</given-names>
            <surname>Suzuki</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Tanaka</surname>
          </string-name>
          , and
          <string-name>
            <given-names>H.</given-names>
            <surname>Inoue</surname>
          </string-name>
          , “
          <article-title>Clustering and modeling of hazardous or nonhazardous sound in daily life</article-title>
          ,”
          <source>2012 Proceedings of SICE Annual Conference (SICE)</source>
          , Japan, pp.
          <fpage>1991</fpage>
          -
          <lpage>1995</lpage>
          ,
          <year>2012</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11] https://www.fujitsu.com/jp/about/resources/case-studies/trends/pdf/cs-201710-connecte dlife.
          <source>pdf. (Accessed</source>
          , Jun.
          <year>2024</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>M.</given-names>
            <surname>Kawamoto</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Asano</surname>
          </string-name>
          , and
          <string-name>
            <given-names>K.</given-names>
            <surname>Kurumatani</surname>
          </string-name>
          , “
          <article-title>A Security Monitoring System of Detecting Unusual Sounds and Hazardous Situations by Sound Environment Measurement Using Microphone Arrays,”</article-title>
          <source>IPSJ SIG Technical Report</source>
          , Vol. 2008-UBI-19, No.
          <volume>66</volume>
          , pp.
          <fpage>19</fpage>
          -
          <lpage>26</lpage>
          , Jul.
          <year>2008</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <given-names>J.</given-names>
            <surname>Allen</surname>
          </string-name>
          , “
          <article-title>Short term spectral analysis, synthesis, and modification by discrete Fourier transform</article-title>
          ,
          <source>” IEEE Transactions on Acoustics, Speech, and Signal Processing</source>
          , Vol.
          <volume>25</volume>
          , No.
          <issue>3</issue>
          , pp.
          <fpage>235</fpage>
          -
          <lpage>238</lpage>
          , Jun.
          <year>1977</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <given-names>M.</given-names>
            <surname>Yokoyama</surname>
          </string-name>
          , “
          <article-title>Distinction of timbre of violin by spectral envelope,”</article-title>
          <source>IPSJ SIG Technical Report</source>
          , Vol. 2020-SLP-132, No.
          <volume>26</volume>
          , pp.
          <fpage>1</fpage>
          -
          <lpage>6</lpage>
          , May.
          <year>2020</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15] https://www.sony-semicon.com/ja/products/spresense/index.html (Accessed, Apr.
          <year>2024</year>
          )
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>