<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <journal-title-group>
        <journal-title>Workshop, Stavropol and Arkhyz, Russian Federation</journal-title>
      </journal-title-group>
    </journal-meta>
    <article-meta>
      <title-group>
        <article-title>Research on Dependences of Speech Pitch Parameters on Pulse and Heartbeat Signals</article-title>
      </title-group>
      <contrib-group>
        <aff id="aff0">
          <label>0</label>
          <institution>Dmitry Poleshenkov ITMO University 49 Kronverksky Pr.</institution>
          ,
          <addr-line>St.Petersburg, 197101</addr-line>
          ,
          <institution>Russia Ekaterina Pakulova Southern Federal University 105/42 Bolshaya Sadovaya Str., Rostov-on-Don, 344006, Russia Oleg Basov ITMO University 49 Kronverksky Pr.</institution>
          ,
          <addr-line>St.Petersburg, 197101</addr-line>
          ,
          <country country="RU">Russia</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2019</year>
      </pub-date>
      <volume>1</volume>
      <fpage>7</fpage>
      <lpage>09</lpage>
      <abstract>
        <p>In this paper, we consider the in uence of the cardiovascular system to the process of speech production. We propose the mathematical model of the impact of cardiovascular system elements on the process of speech pitch synthesis. This model takes into account the e ect of the functioning of the heart and the pulsations of the large vessels of the thorax on the intensity of the air ow from the lungs during breathing, as well as the e ect of the pulsations of the blood vessels of the vocal cords on the speech synthesis process. We made the initial check of approximate properties of the model in practice. In the future, the proposed model allows to construct the precision model of speech signal forming. The solution of the direct synthesis problem allows to increase the speech quality in the synthesis, recognition and coding processes. The inverse problem aims to increase the e ciency of physiologic and psycho-emotional state estimation of a person. In conclusion, we discussed further research directions that may improve the quality of the estimates obtained on the basis of the formulated model.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Introduction</title>
      <p>Analysis of mathematical models of the processes of speech production [FF14, LS17, LZ16, RN19, Sor85, Sor16]
in the tasks of recognition of a physiological and psycho-emotional state of a user of a socio-cyber-physical system
shows that they are very speci c. The consideration of functioning of vocal tract uncoupled from other human
body systems leads to signi cant deviation of modelling results from the real physical process.</p>
      <p>One of such human body systems that has a signi cant in uence on speech production is the cardiovascular
system. In [Bab05, BO14] is shown that speech pitch period contains periodic and random components. In
[BO14] authors also suppose that periodic changes of pitch may be caused by blood ow pulsation.</p>
      <p>The contribution of the present work is detection of dependency between changes of pitch parameters and
functioning of the cardiovascular system in the process of speech production
2</p>
    </sec>
    <sec id="sec-2">
      <title>Analysis of human pulse and heartbeat</title>
      <p>The cardiovascular system in uences on all body systems. Its work is mainly characterized by heart and distal
vessels. Periodic atrial and ventricular contraction of a heart with vessels guarantees blood motion.</p>
      <p>A normal heart rate for adults ranges from 65 to 85 beats per minute. The cardiac cycle consists of two periods:
the period of atrial and ventricular contraction (systole) and the period of their relax (diastole). Duration of
systole is on average 0,33 sec, duration of diastole is 0,47 sec [PV03].</p>
      <p>The temporal representation of the pulse (pulsogram) s1(t) (see Fig. 1a) has two peaks in the period,
corresponding to the moments of release of blood from the ventricles and the re ection of blood ow from the
closed semilunar heart valves, while the pause between these peaks is xed. Pulse spectral bandwidth A1(f )(see
Fig. 1b) on average doesn't exceed 20 Hz [PA12].</p>
      <p>Cardiophonography is one of the methods for the analysis of heart functioning which records the heart sounds.
The temporal representation of heartbeat (phonocardiogram) s2(t) is di erent from pulsogram (see Fig. 1a). It
caused by the complex process of heart functioning. The signal is periodic with distinct peaks corresponding to
the rst (systole) and second (diastole) tones of heart rate. The maxima of these tones fall at the closure of the
atrioventricular and semilunar valves, respectively. In the spectral representation of the heartbeat signal A2(f )
(see Fig. 1b), an increase in the amplitudes of the spectral components is observed in the band from 4 to 20 Hz
with respect to the spectral representation of the pulse signal.</p>
      <p>The phonocardiograms are mainly characterized by the sounds of the heart valves and the blood ow in heart
chambers, the aorta and the pulmonary artery. In this regard, the mechanical vibrations transmitted to the
surface of the lungs will di er signi cantly from the corresponding acoustic signal of the heartbeat. It is proved
by a change in the structure of recorded phonocardiograms depending on the signal recording point [Tsa18].
Window function w(m)
and average Fa</p>
      <p>calculation
s(n) filtering in a band 
[Fa-Fa/2:Fa+Fa/2]
i = 1
i
1
length(s)–i+1≤m</p>
      <p>Yes</p>
      <p>No</p>
      <p>9
10
11
12
13
14
8</p>
      <sec id="sec-2-1">
        <title>Selection of an analysis segment seg(1:m) = s(i:i+m–1)</title>
      </sec>
      <sec id="sec-2-2">
        <title>Calculation of S(f), F, A</title>
        <p>[fmax, amax] = max(S(f))
F(i:i+m–1)=fmax
A(i:i+m–1)=amax</p>
        <p>i = i + st
i length(s)</p>
        <p>F, A</p>
        <p>End</p>
      </sec>
      <sec id="sec-2-3">
        <title>Selection of an analysis segment</title>
        <p>seg(1:length(s)–i+1) = s(i:end)
• the e ect of heart contractions on the intensity of the air ow from the lungs by physical impact on their
surface;
• the e ect of changes in blood ow in the aorta, pulmonary artery and blood vessels of the lungs on the
intensity of the air ow;
• the e ect of pulsations of the blood vessels of the vocal cords on the phase relationships of the acoustic wave
process.</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>Description of mechanisms of cardiovascular system in uence on the synthesis process of speech pitch</title>
      <p>Based on the physiology of the blood circulation process and the structure of the organs of the circulatory system
[PM85, Kab19], we may assume the following basic mechanisms of the in uence of the circulatory system on the
synthesis of the pitch of speech:</p>
      <p>Since the heart is located close to the surface of the lungs, heart contractions have a periodic mechanical
impact on them. It leads to the intensity uctuation of the air stream during the process of breathing. By the
physical model of speech production [Sor92], a change in the ow intensity leads to a change in the frequency
and amplitude of the pitch during speech synthesis. It is worth also to note that both atrial contraction and
ventricular contraction will cause uctuations in the intensity of the air stream.</p>
      <p>Similarly, changes in air ow during breathing are a ected by pulsations of the aorta, pulmonary artery, and
pulmonary blood vessels. Given the close relative positioning of the organs of the cardiovascular system inside
the thoracic cavity, the total mechanical e ect on the lungs is complex.</p>
      <p>During the assessing the in uence of the pulsation of the small blood vessels of the vocal cords on the process
of speech production, one should take into account their small size and low blood pressure in them [Sud00, VS17].
From this, we can conclude that the pulsations of these blood vessels cause a periodic change in the oscillation
phase of the vocal cords. Herewith we should take into account the delay of pressure uctuations in the vessels
of the vocal cords in relation to the heart contractions.</p>
      <p>a)
b)</p>
    </sec>
    <sec id="sec-4">
      <title>Analysis of frequency and amplitude variations of pitch</title>
      <p>In order to get the most accurate conclusion about the e ect of the cardiovascular system on the change of pitch
parameters, it is necessary to minimize the distortion of the analysed speech signals in the pitch frequency band
(50 350 Hz). To do this we study sound records of long-spoken vocalized phonemes (which were synchronously
recorded with pulsograms and phonocardiograms) with the following parameters: sampling frequency fd = 48
kHz; bit rate B = 1,536 bps; microphone bandwidth 2 f = 20 kHz; duration t = 10 ... 14 s. The use of such
initial signal allows us to take the in uence of the articulation component beyond the scope of the study, which,
makes it possible to abandon the complex ways of recording the speech signal.</p>
      <p>During the analysis of common approaches to the selection of pitch parameters [DHKM18, PPS18, SSGZ16,
VM16, ZL18], the algorithm based on the spectral separation method (see Fig. 2) was selected. This algorithm
allows to obtain an estimate of the pitch frequency and amplitude variations with su cient time resolution
determined by the dimension of the discrete transform Fourier, transform window width and signal sampling
rate. The principle of the algorithm is based on the samples allocation of the frequency trajectories and the
amplitude of the pitch calculated at intervals of analysis frames based on the spectral representation S(f ). The
proposed algorithm is simpli ed and does not imply the extraction of phonemes from the speech signal since
there are no phoneme transitions in the studied signals.</p>
      <p>The algorithm uses the following input data: s(n) is the analysed speech signal; N is the dimension of the
Fourier transform; m is the duration of the analysis time window; st is the step shift of the analysis window.
In order to increase the accuracy of trajectory extraction, the shift step is equal to 1. The shift step of the
analysis window should be less than the length of the pitch period. This eliminates the necessity to synchronize
the analysis window with the pitch period and removes the need to use electroglotograms. The output data of
the algorithm are samples sequences of the frequency trajectories F (n) (see Fig. 3a) and amplitudes A(n) (see
Fig. 3b) of the pitch. During our research of vocalized phonemes pronounced by di erent speakers, the result
of the algorithm (in terms of pitch frequency trajectories) completely coincides with the result of frequency
demodulation of the 1st harmonic of the speech signal.</p>
      <p>The received signal of the pitch frequency trajectory carries a large constant component (average pitch
frequency), which makes it di cult to assess the dependence of the considered oscillation on the processes of the
cardiovascular system. Additionally, there is the presence of low-frequency components caused by the intonation
of the spoken segments, as well as a change of the lungs volume as it exhales, which negatively a ects the accurate
analysis. To compensate for the in uence of these factors, we apply the correction algorithm to minimize the
distortion (see Fig. 4).</p>
      <p>The input data for the algorithm is only corrected signal s(n). The principle of the correction algorithm is
based on calculating the intervals of the correction curve pc on the signal duration between the extreme points
x(1 : N ) and the derivative s0(n) of the corrected signal, succeeded by subtracting the generated curve from the
original signal. The algorithm introduces additional distortions caused by the uneven (piecewise linear) nature
of the correction curve. However, they lead to an insigni cant spreading of the periodic components of the
spectrum of the signal, since the position of the zero crossing points and the position of the signal extremes are
not changed relative to the time scale.</p>
      <p>For the analysis of the samples sequence of the trajectories of the frequency F 0(n) (see Fig. 5a) and the
amplitude A0(n) (see Fig. 5b) of the pitch for the presence of components caused by the cardiovascular system,
we made an assessment using a matched ltering mechanism. The re ected periods of the pulse signal were used
as the impulse response of the lters. We also analyse these signals with the wavelet transform.
5</p>
    </sec>
    <sec id="sec-5">
      <title>Results</title>
      <p>As a result of matched ltering, we obtain the periodic signals s4(n) (response of the matched lter (MF) to
the signal of the pitch frequency trajectory), s5(n) (response of the MF to the signal of the pitch amplitude
trajectory (see Fig. 6a). In general, the position of the corresponding local extremes of these signals (with a
small delay) corresponds to the responses s3(n) of the MF. These responses are getting by ltering the signals
of the temporal representation of the pulse.</p>
      <p>Since the total mechanical e ect transmitted to the surface of the lungs is complex and di ers from the
recorded pulsograms and phonocardiograms, the responses of the MF to the signals of the frequency trajectories
and amplitude of the pitch have a complex structure.</p>
      <p>Additional experiments aimed at research on the changes in the pressure of the air stream during breathing
will allow to evaluate the total impact. We used wavelet transform (Daubechies wavelet) for the analysis that
showed the presence of periodic changes in the trajectories of the frequency and amplitude of pitch, corresponding
to heart contractions (see Fig. 6b, Fig. 6c).</p>
      <p>Based on the studied mechanisms of the in uence of the functioning of the cardiovascular system on the process
of pitch synthesis and the obtained experimental data, the pitch frequency trajectory for vocalized phonemes
can be determined as follows.</p>
      <p>f (t) = F0 + f0(t) + m1
dp(t</p>
      <p>1)
dt
+ m2h(t
2)
(1)
where f (t) is the pitch frequency trajectory, F0 is the average frequency of the pitch,f0(t) is the intonation
component, m1, m2 are proportionality factors, 1, 2 are the time delays, p(t) is the e ect on the vocal cords,
h(t) is the total e ect on the lungs from the organs of the cardiovascular system located inside the thoracic
cavity.</p>
      <p>In accordance with expression 1 (see Fig. 6d), the reconstructed signal s6(n) repeats the extremes of the pitch
frequency trajectory s7(n). A model of aperiodic oscillations of the resonant system with a large attenuation
coe cient was used as a model of impact on the lungs from the organs of the cardiovascular system located
inside the thoracic cavity. Moments of oscillations correspond to ventricular and atrial systoles. An unmodi ed
pulse signal was used as a signal to in uence the vocal cords.</p>
      <p>In order to assess the repeatability of the presented data, we check them with the speech signal (male and
female) of long-spoken vocalized phonemes. We estimate the correlation coe cient corrected in accordance with
the correction algorithm (see Fig.5) and the signal restored by the expression (1). The summarized results are
shown in table 1.
6</p>
    </sec>
    <sec id="sec-6">
      <title>Conclusion</title>
      <p>The results presented in this paper suggest that the cardiovascular system makes the main contribution to the
change of the frequency and amplitude of the pitch period. It allows us to construct the precision model of a
speech signal forming that takes into account the in uence of the considered factors. The solution of the direct
synthesis problem allows to increase the speech quality in the synthesis, recognition and coding processes. The
inverse problem aims to increase the e ciency of the physiologic and psycho-emotional state estimation of a
person.
7</p>
    </sec>
    <sec id="sec-7">
      <title>Acknowledgements</title>
      <p>The authors would like to thank the anonymous referees for their valuable comments and helpful suggestions.
This work is supported by the Russian Foundation For Basic Research (grant №18-07-00380A).</p>
      <p>V.V. Babkin. Noise immune speech pitch isolator. In Proceedingsof the 7th international conference
"Digital signal processing and its application", volume X-1. IPU RAS, 2005.</p>
      <p>Shalaginov V.A. Basov O.O., Nosov M.V. Pitch-jitter analysis of the speech signal. In SPIIRAS
Proceedings, volume 32, 2014. on russian.
[DHKM18] Thomas Drugman, Goeric Huybrechts, Viacheslav Klimkov, and Alexis Moinet. Traditional machine
learning for pitch detection. IEEE Signal Processing Letters, 25(11):1745{1749, 2018.</p>
      <p>Mohamed Hesham Farouk and Farouk. Application of wavelets in speech processing. Springer, 2014.
N.A. Kabanov. Human Anatomy. Urait Publishing House, 2019.</p>
      <p>AS Leonov and VN Sorokin. Upper bound of errors in solving the inverse problem of identifying a
voice source. Acoustical Physics, 63(5):570{582, 2017.</p>
      <p>NA Lyubimov and EV Zakharov. Mathematical model of acoustic speech production with mobile
walls of the vocal tract. Acoustical Physics, 62(2):225{234, 2016. on russian.</p>
      <p>Ompokov V.D. Pavlov AE, Boronoev V.V. The research of the level of tness of sportsman organism
at the diagnostic complex apdk. Journal of Buryat State University, 4, 2012.</p>
      <p>Bushkovich V.I. Prives M.G., Lisenkov N.K. Human Anatomy. Izdatelstvo Meditsina, 1985.
Monisankha Pal, Dipjyoti Paul, and Goutam Saha. Synthetic speech detection using fundamental
frequency variation and spectral features. Computer Speech &amp; Language, 48:31{50, 2018.
[Bab05]
[SSGZ16]
[Tsa18]
[VM16]
[VS17]
[ZL18]</p>
      <p>K Sreenivasa Rao and NP Narendra. Source Modeling Techniques for Quality Enhancement in
Statistical Parametric Speech Synthesis. Springer, 2019.</p>
      <p>V.N. Sorokin. Theory of speech production. Radio i sviaz, 1985. on russian.</p>
      <p>V.N. Sorokin. Speech synthesis. Izdatelstvo Nauka, 1992. on russian.</p>
      <p>VN Sorokin. Segmentation of the period of the fundamental tone of a voice source. Acoustical
Physics, 62(2):244{254, 2016. on russian.</p>
      <p>Michael Staudacher, Viktor Steixner, Andreas Griessner, and Clemens Zierhofer. Fast fundamental
frequency determination via adaptive autocorrelation. EURASIP Journal on Audio, Speech, and
Music Processing, 2016(1):17, 2016.</p>
      <p>K.V. Sudakov. Physiology.Basics and functional systems. Izdatelstvo Meditsina, 2000. on russian.
V.P. Tsarev. Auscultation of the heart. Belarussian state medical university, 2018. on russian.
DA Volf and RV Meshsheryakov. Model of process of singular estimation of the primary tone of a
speech signal. Acoustical Physics, 62(2):215{224, 2016.</p>
      <p>D.S. Sveshnikov V.M. Smirnov, V.A. Pavdivtseva. Physiology: A textbook for students of medical
and pediatric faculties. Izdatelstvo MIA, 2017. on russian.</p>
      <p>Xiaoheng Zhang and Yongming Li. Pitch tracking algorithm based on evolutionary computing with
regularisation in very low snr. The Journal of Engineering, 2018(16):1509{1514, 2018.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          <source>[Sor85] [Sor92] [Sor16] [Sud00]</source>
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>