<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Entropy-based detection of the words boundaries of continuous speech</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Andrey S. Karpov</string-name>
          <email>andrey revol125@mail.ru</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Galina V. Shagrova</string-name>
          <email>g shagrova@mail.ru</email>
          <email>shagrova@mail.ru</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Victoria I. Drozdova</string-name>
          <email>drozdova@rambler.ru</email>
          <email>victoria drozdova@rambler.ru</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Aleksey V. Shevchenko</string-name>
          <email>luckyleo769@mail.ru</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Information Systems Technologies Dept.</institution>
          ,
          <addr-line>Stavropol, NCFU</addr-line>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2018</year>
      </pub-date>
      <abstract>
        <p>An algorithm for nding the word boundaries in a merged speech is proposed on the basis of a method using the de nition of the entropy of a speech signal. The di erence between the proposed algorithm and the known ones is the comparison of the speech signal entropy value with the entropy threshold in two stages. The work of the known and proposed algorithm is compared.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Introduction</title>
      <p>Description of the object and methods of research
De ning the boundaries of words in a speech signal is a key aspect of the human speech recognition task, as
at this stage speech data is separated from unnecessary noise and speech artifacts (a cough, speech harmonics,
microphone echo, etc.).</p>
      <p>The use of the method based on the value of the entropy of the speech signal gives high indicators of the
de nition of word boundaries for the task of recognizing isolated commands [Alu14, Alu16, Boz11, Wah02].</p>
      <p>The essence of the method is that the input voice data is preliminarily processed using a bandpass lter. This
lter removes the constant and low-frequency components of the background, as well as high-frequency noise and
speech harmonics, arising from the spectral properties of the voice path. The pre-processed speech is normalized
so that the amplitude values of the signal lie in the range from 1 to -1. Then, the normalized signal is divided into
frames of approximately 25 milliseconds of speech. To avoid loss of information, these frames have an overlap of
25 - 50%.</p>
      <p>Then the entropy value in each frame of the analyzed sound sequence is calculated:
where Hj (j = 1; 2; : : : ; m) { the value of entropy of the j-th frame, m { the number of frames;;
pi { the probability of i-th signal count, in j-th frame;
N { the number of counts within the frame.</p>
      <p>As a result, the entropy pro le is determined for the incoming speech signal, which is a histogram of the
entropy values of all the frames, the recognizable fragment of speech:</p>
      <p>In the case of recognizing isolated instructions, the entropy pro le of the signal is used to calculate the entropy
threshold .</p>
      <p>=
max( )
min( )
+
min( );</p>
      <p>&gt; 0
2
where { the noise ratio, which is selected experimentally [Boz11].</p>
      <p>However, for example, in [Alu14, Alu16] the value of the entropy threshold is not calculated but is taken equal
to a constant value: =0,1. But this approach does not give good results for cases of a noisy signal.</p>
      <p>After determining the threshold, the value of the entropy of each frame Hj is compared with the entropy
threshold . Any value equal to or greater than the entropy threshold is considered a speech and all that is less
is silence or noise.</p>
      <p>H =</p>
      <p>N
X piln(pi)
i=1
= [H1; H2; Hm]
=
(</p>
      <p>Hj ; Hj
0;</p>
      <p>Hj &lt;
;</p>
      <p>However, due to the vocal characteristics of the speech signal, the entropy index may be too small in the area
of the recognizable speech signal that carries the information. Or, conversely, because of instantaneous noise, a
signal segment that does not carry speech data is recognized as speech [Naz15]. In order to avoid the erroneous
de nition of the word boundaries in the speech signal, the concepts of the minimum word length (k) and the
minimum distance between words ( ) [Alu14, Alu16, Obi12]. Both these quantities are measured in the number
of frames.</p>
      <p>The rst criterion is that each recognized speech segment ( i, j ) must have a certain minimum length, which
is indicated as a constant. That is i &lt; k and dij &gt; , the i-th segment is discarded as a segment that does not
contain voice information. Also, if j &lt; k and dij &gt; , the j-th segment is discarded.
(1)
(2)
(3)
(4)</p>
      <p>The second criterion is based on the minimum distance between words It consists in the fact that two segments
of the analyzed speech signal, de ned as speech, are combined into one if the distance between them (dij ) is less
than the speci ed number of frames. This means that if ( i or j ) &gt; k and dij &lt; , then the two segments are
combined into one.</p>
      <p>This approach gives a high result of detecting the boundaries of isolated words. In order to use this approach
in the recognition of the continuous speech, an algorithm is proposed, according to which the entropy threshold
was determined by the formula (5):
= min( ) + (max( )
min( )) k
(5)
where k { the coe cient that was selected experimentally, the word boundaries were determined in two stages.
At each stage, the minimum distance between words (dij ) was used.</p>
      <p>The result of the proposed algorithm is given in the work by the example of separating the boundaries of the
words of the merged speech, which is a speech signal containing the phrase "Dear passengers, please keep calm,
the train will soon leave" pronounced in a woman's voice.</p>
      <p>The analyzed phrase consisting of eight words was recorded with a sampling frequency of 22kHz, the number
of channels 2 (stereo) and 16 bits. The duration of the speech signal was 6,583 seconds. The boundaries of the
words of this phrase were in two stages.</p>
      <p>The minimum distance between words in the rst stage was 12 frames, and k = 0,9. This means that all
frame groups de ned as speech, but less than 12 frames in length, are discarded as non-verbal data. The results
of the rst stage of the algorithm are shown in Figure 2.</p>
      <p>As shown in Figure 2, as a result of the rst stage of the algorithm, three large groups of frames carrying
the voice information were formed. At the second stage, only those frame groups that were formed after the
rst stage were considered. For them, the minimum distance between words was 3 frames, and k = 0,75. Let us
consider the work of the second stage of the algorithm for frame groups formed after the rst stage (Figure 3).</p>
      <p>As can be seen from Figure 3, in the second stage, the algorithm divided the rst large group of frames into
two smaller ones, which are separate words. Similarly, both the second and third large groups were divided into
three smaller ones. As a result, the boundaries of all eight words were found.</p>
      <p>An example of a comparison of the work of the known and proposed algorithms is given for the speech signal,
which is the phrase "Today is good weather". The analyzed phrase is pronounced by a man and recorded with
a sampling frequency of 16 kHz, the number of channels 1 (mono) and 8 bits. The duration of this phrase was
2,535 seconds. The results are shown in Figure 4.</p>
      <p>As can be seen from Figure 4, the known algorithm de nes the entire phrase as a group of frames that carry
information. Whereas the proposed algorithm determines the boundaries of all three words quite accurately.</p>
    </sec>
    <sec id="sec-2">
      <title>Summary</title>
      <p>A new algorithm for determining the boundaries of words in a merged speech is proposed, which di ers from the
known, based on the calculation of the entropy value of a speech signal, in that the process of separating the
boundaries of words is performed in two stages. At the rst stage, a rough selection of large groups of frames
containing verbal information is carried out. At the second stage, there is a more detailed segmentation of the
speech fragments obtained in the rst stage.</p>
      <p>Due to the use of the method, based on the de nition of the entropy of the speech signal in speech recognition
systems, much higher recognition rates of the word boundaries can be achieved, both in isolated and in the
combined speech.</p>
      <p>Figure 4: The result of the known (a) and proposed (b) algorithm for the phrase "Today is ne weather"</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [Alu14]
          <string-name>
            <given-names>D.Yu</given-names>
            <surname>Alunov</surname>
          </string-name>
          .
          <article-title>On Methods for Estimation of the Signal Parameters</article-title>
          .
          <source>Current Problems of Science and Education No. 6</source>
          ,
          <year>2014</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [Alu16]
          <string-name>
            <given-names>D.</given-names>
            <surname>Yu</surname>
          </string-name>
          . Alunov,
          <string-name>
            <given-names>E.S.</given-names>
            <surname>Sergeev</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.V.</given-names>
            <surname>Pigachev</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.N.</given-names>
            <surname>Mytnikov</surname>
          </string-name>
          .
          <article-title>Implementation of the algorithm for processing and recognizing speech</article-title>
          .
          <source>Modern high technology No. 3-2</source>
          , pp.
          <fpage>225</fpage>
          -
          <lpage>230</lpage>
          ,
          <year>2016</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [Boz11]
          <string-name>
            <given-names>A.S.</given-names>
            <surname>Bozhdai</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.A.</given-names>
            <surname>Gudkov</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.A.</given-names>
            <surname>Gudkov</surname>
          </string-name>
          .
          <article-title>Embedded identi cation system by voice biometric indicators</article-title>
          .
          <source>Open Education No 2-2</source>
          ,
          <year>2011</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [Koc15]
          <string-name>
            <given-names>A.V.</given-names>
            <surname>Kochetkov</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.V.</given-names>
            <surname>Fedotov</surname>
          </string-name>
          .
          <article-title>About various meanings of the concept "entropy"</article-title>
          .
          <source>Internet-journal Naukovedenie</source>
          , Vol.
          <volume>6</volume>
          ,
          <year>2015</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [Naz15]
          <string-name>
            <given-names>A.V.</given-names>
            <surname>Nazarov</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.L.</given-names>
            <surname>Yakimov</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.F.</given-names>
            <surname>Avdeev</surname>
          </string-name>
          .
          <article-title>The algorithm for maximizing the entropy of the training sample and its use in the synthesis of forecast models for discrete states of nonlinear dynamical systems</article-title>
          .
          <source>Scienti c Journal "Information Control Systems": Issue</source>
          <volume>2</volume>
          ,
          <string-name>
            <surname>St. Petersburg</surname>
          </string-name>
          ,
          <year>2015</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [Obi12]
          <string-name>
            <given-names>N.</given-names>
            <surname>Obin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Liuni</surname>
          </string-name>
          .
          <article-title>On the generalization of shannon entropy for speech recognition</article-title>
          .
          <source>IEEE workshop on Spoken Language Technology, United States</source>
          ,
          <year>2012</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [Wah02]
          <string-name>
            <given-names>K.</given-names>
            <surname>Waheed</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Weaver</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.M.</given-names>
            <surname>Salam</surname>
          </string-name>
          .
          <article-title>A robust algorithm for detecting speech segments using an entropic contrast</article-title>
          .
          <source>Midwest Symposium on Circuits and Systems 3</source>
          ,
          <year>2002</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>