<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>FONTI 4.0: Evaluating Speech-to-Text Automatic Transcription of Digitized Historical Oral Sources</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Roberta Bianca Luzietti</string-name>
          <email>robertabianca.luzietti@unipd.it</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Niccol o` Pretto</string-name>
          <email>niccolo.pretto@dei.unipd.it</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Fre´de´ric Kaplan</string-name>
          <email>frederic.kaplan@epfl.ch</email>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Alain Dufaux</string-name>
          <email>alain.dufaux@epfl.ch</email>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Sergio Canazza</string-name>
          <email>canazza@dei.unipd.it</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>. University of Padua</institution>
          ,
          <country country="IT">Italy</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>Conducting “manual” transcriptions and analyses is unsustainable for most historical oral archives because they require a remarkable amount of funds and time. The FONTI 4.0 project aims at exploring the suitability of automatic transcription and information extraction technologies for making historical oral sources available. In this work, we conducted an experiment to test the performance of two commercial speech-to-text services (Google Cloud Speech-to-text and Amazon Transcribe) on digitized oral sources. We created an eight-hour corpus made of manually transcribed and annotated historical speech recordings in TEI format. The results clearly show how audio quality and disturbing elements (e.g., overlaps, foreign words, etc.) impact on the automatic transcription, showing what needs to be improved for implementing an unsupervised transcription chain.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1 Introduction</title>
      <p>FONTI 4.01 is a project aiming at exploring the
suitability of automatic transcription and analysis
tools for the preservation of historical oral sources
recorded on analog carriers, in particular magnetic
tapes. The digitization of an audio archive is a
long and expensive task that can require several
years. Furthermore, the content of audio
recordings needs to be listened and cataloged for making
audio recordings retrievable. Archives composed
by hundreds or thousands of hours of audio require
a huge amount of time, people and funds for
making the content accessible and preventing their
exploitation. Therefore, automatizing the
transcription and the analysis task could drastically reduce
the time for making digitized audio recordings
accessible.</p>
      <p>
        The project consists in a transcription-chain
(Tchain), firstly defined in (van Hessen et al., 2020),
that differs in two main aspects: (a) in FONTI
4.0, the transcription obtained with
speech-totext (STT) algorithms should not be corrected by
human; (b) an additional restoration step could
be required for digitized audio recordings.
Furthermore, differently from STT evaluation
experiments conducted by
        <xref ref-type="bibr" rid="ref2 ref4 ref6">(Moore et al., 2019;
Kostuchenko et al., 2019; Filippidou and
Moussiades, 2020)</xref>
        , we decided to employ two
commercial software, namely Google Cloud Platform and
Amazon Web Services, to test their ability to
transcribe historical analog recordings, and to
eventually include in our pipeline.
      </p>
      <p>During the digitization process, speed and
equalization errors can occur, especially when
different speed and equalization configurations are
used in different part of the same tape (Pretto et
al., 2020). This leads to distortions of the recorded
signal that becomes unlistenable. By using the
correction workflow and digital filters described in
(Pretto et al., 2021a; Pretto et al., 2021b) these
errors can be corrected and at least parts of the
signal can be saved. This task is essential for
making the speech signal suitable for STT algorithms.
This paper aims at evaluating the transcription
performance of two commercial software on a real
use case and identifying potential problems or
limitations concerning peculiarities of analog audio
recordings. Section 2 describes the corpus, used
as ground-truth for the experiment. Section 3
outlines the methodology adopted for this
experiment, whereas results are reported in Section 4.
Finally, Section 5 presents the authors’
conclusions.
2</p>
    </sec>
    <sec id="sec-2">
      <title>Corpus</title>
      <p>
        The Cinema &amp; Civilta` (C&amp;C) corpus was
conceived within the FONTI 4.0 as ground-truth for
evaluating the performance of STT services on a
real case study. To build the corpus, we
transcribed speech recorded on four magnetic tapes
made available by the Giorgio Cini Foundation
in Venice and digitized at the Centro di
Sonologia Computazionale - CSC
        <xref ref-type="bibr" rid="ref2 ref4">(Canazza and De Poli,
2020)</xref>
        . The recordings are parts of the Cinema
&amp; Civilta` conference for the awarding of the San
Giorgio prize, part of the Venice Film Festival, that
took place between the 7th and 9th of September
1959, attended by important figures of the history
of cinema such as Roberto Rossellini and
representatives of the Italian literary critics such as
Vittore Branca. Each reel of magnetic tape is
composed of two sides: each side counting 60
minutes of recorded speech for a total of eight hours
of recording. The C&amp;C corpus is also a
multilingual corpus of 64,930 tokens and three
subcorpora: Italian 49,772 tokens, French 9,555
tokens (L1 and L2), and Spanish 5,603 tokens. This
corpus was manually transcribed and annotated as
described in the following subsections and is
available at this link2.
2.1
      </p>
      <sec id="sec-2-1">
        <title>Transcription</title>
        <p>Defining the methodology for the transcription is
an important step for the preservation, analysis and
access of oral sources. The main difficulty consists
in making decisions on how to represent and
convey both verbal and non-verbal elements in
written form. Because of the absence of a universal
standard of transcription (Schorrsidt, 2011), the
methodology usually depends on the research aim.</p>
        <p>In this research, we decided to complete a
verbatim transcription, by reporting every word
spoken in the recording including errors, false starts,
truncations, and overlaps in Italian, French and
Spanish. Using the software ELAN (Lausberg
and Sloetjes, 2016), we first segmented audio files
extracted from the digitized tapes, making the
start and end of each segment coincide with the
2DOI: 10.5281/zenodo.5645827
speaker’s turn of talk. Then, we transcribed each
segment while listening to the corresponding part
of audio in slow motion. Eventually, we opted for
employing automatic transcriptions from Google
Cloud Speech-to-text (GCS) and Amazon
Transcribe (AT)3, later used in the STT experiment,
and correcting the text playing the audio at
normal speed. This allowed us to save half the time
for each transcription, which previously required
a full day of work. Moreover, we were able to
retrace and match the identity of the speakers to the
voices in the recordings, through the consultation
of historical documentation on the conference, and
also by comparing voices across the recordings.
2.2</p>
      </sec>
      <sec id="sec-2-2">
        <title>Annotation</title>
        <p>
          The annotation was employed for the addition of
important metadata to the C&amp;C corpus
regarding different levels of audio quality and the
presence of disturbing elements in the recordings. Our
methodology is in compliance with the Text
Encoding Initiative (TEI) standard guidelines4 for
transcribed spoken material
          <xref ref-type="bibr" rid="ref1">(Burnard and
Bauman, 2007)</xref>
          . To proceed with the annotation,
we first converted the transcription files from the
ELAN .eaf into the XML TEI standard using the
EXMARaLDA (Schmidt and Wo¨rner, 2014) tool
TEI Drop (Schorrsidt, 2011). Subsequently, we
used Oxygen5 to assign TEI tags to the relevant
tokens. The list of tags together with a brief
description and examples is given below:
&lt;pause&gt; marks a pause either between or within
utterances in the same segment, e.g.: unica
fisionomia. &lt;pause/&gt; Parte dell’architettura;
&lt;unclear&gt; contains a word, phrase, or
passage that could not be transcribed with
certainty because it is illegible or
inaudible in the source, e.g.: gli stessi
&lt;unclear reason=”inaudible”&gt; strumenti
&lt;/unclear&gt;, volti agli stessi fini ;
&lt;gap&gt; indicates a point where material has been
omitted in the transcription because it is
inaudible, e.g.: erba che sorgera` &lt;gap
reason=”inaudible”/&gt; quell’asfalto.;
3Automatic transcriptions were obtained on the 16th,
17th, 19th and 24th of March 2021.
        </p>
        <p>4tei-c.org/release/doc/tei-p5-doc/en/
html/TS.html (last accessed September 3rd, 2021
5oxygenxml.com (last accessed September 3rd, 2021)
&lt;foreign&gt; identifies a word or phrase as
belonging to some language other than that
of the surrounding text, e.g.: &lt;foreign
xml:lang=”fr-FR”&gt; Mesdames, messieurs
&lt;/foreign&gt;;
&lt;shift&gt; marks the point at which some
paralinguistic feature of a series of
utterances by any one speaker changes, e.g.:
Io credo che questo argomento sia &lt;shift
feature=”tempo” new=”a”/&gt;
particolarmente importante &lt;shift feature=”tempo”
new=”normal”/&gt; per vedere;
&lt;del&gt; contains a letter, word, or passage
indicated as superfluous by the
annotator, in this case it was used for false
starts, repetitions and truncations, e.g.: in
questo &lt;del type=”falseStart”&gt; moden
&lt;/del&gt; momento (false start) momento
di &lt;del type=”repetition”&gt; di &lt;/del&gt;
crisi (repetition) suggestione di &lt;del
type=”truncation”&gt; spettaco &lt;/del&gt; di
spettacolo (truncation);
&lt;anchor&gt; was used to mark overlaps by
attaching an identifier to a point within a
text, e.g.: a contatto di un &lt;anchor
synch=”ovrl6” xml:id=”S06”/&gt; pensiero
&lt;anchor synch=”ovrl6e” xml:id=”S06e”/&gt;
lo inducono a (interrupted speaker) &lt;anchor
xml:id=”ovrl6”/&gt; Io non lo vedo. Chi
e` questo? Chi e` questo? &lt;anchor
xml:id=”ovrl6e”/&gt; (interrupting speaker);
&lt;note&gt; contains notes or citations, and, for the
purpose of this research, it was used to
annotate the audio quality at the beginning of each
segment, e.g.: &lt;note&gt;good &lt;/note&gt;;
Audio quality annotations (&lt;note&gt;) were
assigned to each segment using the the following
scale (Samar and Metz, 1988):
excellent: speech is completely intelligible;
good: speech is intelligible with the exception of
a few words or phrases;
fair: with difficulty, the listener can understand
about half the content of the message;
poor: speech is very difficult to understand, only
isolated words or phrases are intelligible;
bad: speech is completely unintelligible.</p>
        <p>The distribution of words (without punctuation
and events) for each audio quality annotation is
reported in Table 1.</p>
      </sec>
      <sec id="sec-2-3">
        <title>Scale</title>
      </sec>
      <sec id="sec-2-4">
        <title>Excel.</title>
      </sec>
      <sec id="sec-2-5">
        <title>Good</title>
      </sec>
      <sec id="sec-2-6">
        <title>Fair</title>
      </sec>
      <sec id="sec-2-7">
        <title>Poor</title>
        <p>TOT
it-IT
9,075
30,571
2,919
1,417
43,984
fr-FR
5,930
2,514
83
23
8,550
es-ES
4,097
800
0
0
4,897
&lt;distinct&gt; identifies any word or phrase which
is regarded as linguistically distinct, as in the
case of prosodically unified units, e.g.:
staccarsi da &lt;distinct type=”pcu”&gt; questa
estetica &lt;/distinct&gt; e dai pregiudizi;</p>
        <p>
          The STT experiment consisted in testing the
ability of GCS and AT to correctly transcribe
historical recordings. Furthermore, we decided to
investigate the performance of STT transcriptions
obtained from GCS and AT at different levels of
au&lt;vocal&gt; marks any vocalized but not necessarily dio quality and in presence of disturbing elements
lexical phenomenon, e.g.: del nostro mondo in the recordings such as background noise,
over&lt;vocal&gt; &lt;desc&gt;cough&lt;/desc&gt; &lt;/vocal&gt; laps, code switching etc. (see Section 2.2).
che direi postmoderno.; To analyze the performance of the two STT
systems, we developed a Jupyter notebook able to
&lt;incident&gt; marks any phenomenon or iflter the text by language, audio quality,
disturboccurrence, not necessarily com- ing elements, etc., and select several options, such
municative, for example incidental as tokenization rules. In this experiment, we
denoises or other events affecting com- cided to use only lower case characters, split
aposmunication, e.g.: e` attivita` creatrice, trophes and remove punctuation from both
man&lt;incident&gt;&lt;desc&gt;noise&lt;/desc&gt;&lt;/incident&gt; ual and automatic transcriptions. The
groundma non propriamente l’artista; truth and the resulting transcription of the STT
services were canonicalized. The alignment
algorithm works on single utterances and minimizes
the Levenshtein distance
          <xref ref-type="bibr" rid="ref5">(Jurafsky and Martin,
2008)</xref>
          . The obtained metrics were: the number of
correct matches (COR) and mismatches, i.e.:
deletions (DEL), substitutions (SUB) and insertions
(INS), and the word error rate (WER), which is
the ratio between the number of mismatches and
words in the reference text (Morris et al., 2004).
        </p>
        <p>
          It is important to note that we did not employ this
metric to tell how good a system is, but only that
one is better than the other
          <xref ref-type="bibr" rid="ref3">(Errattahi et al., 2018)</xref>
          .
        </p>
        <p>In order to avoid the introduction of errors not
due to the transcription task, we decided not to
use the automatic language recognition feature
because it could drastically impact on the
performance. Therefore, we cut and divided the audio
ifles in different languages and automatically
transcribed them separately.
In this preliminary work we illustrate and
compare mainly WER trends between the two STT
systems, calculated on the entire corpus as well as
each sub-corpora in relation to audio quality levels
and the presence of disturbing elements.</p>
        <p>Figure 1 illustrates that the performance of AT
are better than GCS in all corpora. The
difference between the two systems is small in the
Italian sub-corpus, but much wider in the French.</p>
        <p>A possible explanation could be the presence of
L2 speakers of French whose pronunciation could
have negatively affected the recognition
performance. Nevertheless, it should be also considered
that the Italian sub-corpus is more than vfie times
bigger than the French and the Spanish.</p>
        <p>STT software performance can be further
observed in Table 2: for the transcription of the
whole corpus, AT scores a lower WER and finds
more correct matches than GCS. On the other
hand, deletions in GCS are more than double than
in AT, whereas substitutions and insertions are
higher in AT than in GCS. In any case, the number
of deletions and insertions between AT and GCS
are different probably because the two services
make use of different language model weights.</p>
        <p>Figure 2 shows that transcription performance
are very similar in Italian and Spanish with
“Excellent” quality, but not in French. For this reason,
we cannot impute the bad GCS performance to
audio quality. In the Italian sub-corpus, performance
are also similar with “Good” quality, but not in
the Spanish, where both services performed badly.</p>
        <p>The negative impact of audio quality is also
evident in the French sub-corpus, despite WER
values are much higher than Italian.</p>
        <p>Results in Figure 3 display the annotated
disturbing events found in the C&amp;C corpus that were
assumed to negatively affect the performance of
STT software in terms of WER. The element that
provides the minor disturbance is shift, although
the scored WER value for this tag is higher than
the one calculated on the overall evaluation. About
the other disturbing elements, they show a
major impact on the transcription of both STT
services. Overall, AT performance is better with most
disturbing elements. The only exception is
represented by code-switching events in foreign
languages for which GCS had a better performance.
5</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>Conclusion</title>
      <p>In this article we conducted a preliminary
research experiment testing the ability of STT
software to correctly transcribe digitized historical
oral sources on magnetic tape. It should be noted,
that since this preliminary work has been
conducted on a small sample of data, our results are
only indicative of which elements represent the
biggest obstacle for STT software performance.</p>
      <p>In spite of disturbing elements and the variation
of audio quality in the recordings, we demonstrate
that with our dataset and in terms of WER, AT
performed more accurate transcriptions compared to
GCS . On the other hand, GCS was better at
recognizing foreign words. Table 2 shows that AT
introduces less incorrect words but more insertions
and substitutions. This should be taken into
consideration when working with automatic
information extraction tools (e.g., Named Entity
Recognition algorithms) applied to automatic
transcriptions. Further analysis should investigate the cause
of this trend, to verify if this behavior is also due
to alignment or tokenization errors.</p>
      <p>With respect to software performance
evaluations in relation to variables characterizing analog
recordings of speech, we found evidence that
audio quality drastically impacts on the number of
mismatches. Observations about the incidence of
disturbing elements, on the other hand, cannot be
generalized since sub-corpora are in three different
languages and have three different sizes.
Throughout the analyses we noted that the most negative
impact on transcription, in terms of the increase of
WER, is caused by the presence of some specific
recurring elements, i.e.: code-switching (foreign),
overlaps and probably even the production of L2
speakers (Figure 3). Nonetheless, given the
necessity of preserving historical documents in a more
time and cost effective way, we came to the
conclusion that researchers working on the
preservation of historical recordings will benefit from the
use of the T-chain. This is because the reduction
by half of the time required for manual
transcriptions in slow motion does compensate the lack of
accuracy. This means that researchers working
on the collection and preservation of oral archives
will be able to focus on filling the gap between
human and machine output.</p>
      <p>Further contributions will be necessary for
conducting experiments on L1 and L2 data separately,
cross-language testings reducing the Italian subset
to the size of the French and Spanish sub-corpora
and evaluating the impact of incorrect
transcriptions on WER. Language identification through
code-switching is another important problem for
automatic transcription. Both services recently
provided this functionality, but while we are
writing this paper, the Google Cloud is still a preview
version. As soon as the feature will be available
the performance of automatic language
recognition algorithms should also be investigated,
especially because this feature is essential for
automatizing the transcription of entire archives.</p>
    </sec>
    <sec id="sec-4">
      <title>Acknowledgments</title>
      <p>This paper is produced under the FONTI 4.0
project, financed by resources from the Regional
Operational Program co-financed with the
European Social Fund 2014-2020 of the Veneto
Region. An important acknowledge is due to
Fondazione Giorgio Cini, Venice for making available
its precious audio material as well as for its help in
the recording analysis, and Matteo Petten o` for his
contribution in the Jupyter notebook development.</p>
      <p>Hedda Lausberg and Han Sloetjes. 2016. The
revised neuroges–elan system: An objective and
reliable interdisciplinary analysis tool for nonverbal
behavior and gesture. Behavior research methods,
48(3):973–993.</p>
      <p>Meredith Moore, Michael Saxon, Hemanth
Venkateswara, Visar Berisha, and Sethuraman
Panchanathan. 2019. Say what? A dataset for
exploring the error patterns that two ASR engines
make. Proceedings of the Annual Conference of the
International Speech Communication Association,
INTERSPEECH, 2019-Septe:2528–2532.</p>
      <p>Andrew Cameron Morris, Viktoria Maier, and Phil
Green. 2004. From wer and ril to mer and wil:
improved evaluation measures for connected speech
recognition. In Eighth International Conference on
Spoken Language Processing.</p>
      <p>Niccolo` Pretto, Alessandro Russo, Federica
Bressan, Valentina Burini, Antonio Roda`, and Sergio
Canazza. 2020. Active preservation of analogue
audio documents: A summary of the last seven years
of digitization at csc. In Proceedings of the 17th
Sound and Music Computing Conference, SMC20,
pages 394–398, Torino, Italy.</p>
      <p>Niccolo` Pretto, Nadir Dalla Pozza, Alberto Padoan,
Anthony Chmiel, Kurt James Werner, Alessandra
Micalizzi, Emery Schubert, Antonio Roda`, Simone
Milani, and Sergio Canazza. 2021a. A
worklfow and novel digital filters for compensating speed
and equalization errors on digitized audio open-reel
tapes. In Proceedings of Audio Mostly 2021, AM21,
Trento, Italy.</p>
      <p>Niccolo` Pretto, Edoardo Micheloni, Anthony Chmiel,
Nadir Dalla Pozza, Dario Marinello, Emery
Schubert, and Sergio Canazza. 2021b. Multimedia
Archives: New Digital Filters to Correct
Equalization Errors on Digitized Audio Tapes. Advances in
Multimedia, 2021:5410218.</p>
      <p>Vincent J Samar and Dale Evan Metz. 1988. Criterion
validity of speech intelligibility rating-scale
procedures for the hearing-impaired population.
Journal of Speech, Language, and Hearing Research,
31(3):307–316.</p>
      <p>Thomas Schmidt and Kai Wo¨rner. 2014.
Exmaralda. In Jacques Durand, Ulrike Gut, and Gjert
Kristoffersen, editors, The Oxford handbook of
corpus phonology. Oxford University Press.</p>
      <p>Thomas Schorrsidt. 2011. A tei-based approach to
standardising spoken language transcription.
Journal of the Text Encoding Initiative, 1.</p>
      <p>Arjan van Hessen, Silvia Calamai, Henk van den
Heuvel, Stefania Scagliola, Norah Karrouche,
Jeannine Beeken, Louise Corti, and Christoph Draxler.
2020. Speech, voice, text, and meaning: A
multidisciplinary approach to interview data through the
use of digital tools. In Companion Publication of
the 2020 International Conference on Multimodal
Interaction, ICMI ’20 Companion, page 454–455,
New York, NY, USA. Association for Computing
Machinery.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          <string-name>
            <given-names>Lou</given-names>
            <surname>Burnard</surname>
          </string-name>
          and Syd Bauman, editors,
          <year>2007</year>
          .
          <article-title>TEI P5: Guidelines for Electronic Text Encoding and Interchange, chapter A Gentle Introduction to XML. Text Encoding Initiative Consortium</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          <string-name>
            <given-names>Sergio</given-names>
            <surname>Canazza and Giovanni De Poli</surname>
          </string-name>
          .
          <year>2020</year>
          . Four Decades of Music Research, Creation, and
          <article-title>Education at Padua's Centro di Sonologia Computazionale</article-title>
          .
          <source>Computer Music Journal</source>
          ,
          <volume>43</volume>
          (
          <issue>4</issue>
          ):
          <fpage>58</fpage>
          -
          <lpage>80</lpage>
          ,
          <fpage>10</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          <string-name>
            <given-names>Rahhal</given-names>
            <surname>Errattahi</surname>
          </string-name>
          , Asmaa El Hannani, and
          <string-name>
            <given-names>Hassan</given-names>
            <surname>Ouahmane</surname>
          </string-name>
          .
          <year>2018</year>
          .
          <article-title>Automatic speech recognition errors detection and correction: A review</article-title>
          .
          <source>Procedia Computer Science</source>
          ,
          <volume>128</volume>
          :
          <fpage>32</fpage>
          -
          <lpage>37</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          <string-name>
            <given-names>Foteini</given-names>
            <surname>Filippidou</surname>
          </string-name>
          and
          <string-name>
            <given-names>Lefteris</given-names>
            <surname>Moussiades</surname>
          </string-name>
          .
          <year>2020</year>
          .
          <article-title>A benchmarking of ibm, google and wit automatic speech recognition systems</article-title>
          .
          <source>In IFIP International Conference on Artificial Intelligence Applications and Innovations</source>
          , pages
          <fpage>73</fpage>
          -
          <lpage>82</lpage>
          . Springer.
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          <string-name>
            <given-names>Daniel</given-names>
            <surname>Jurafsky</surname>
          </string-name>
          and
          <string-name>
            <given-names>James</given-names>
            <surname>Martin</surname>
          </string-name>
          .
          <year>2008</year>
          .
          <article-title>Speech and Language Processing: An Introduction to Natural Language Processing</article-title>
          ,
          <source>Computational Linguistics, and Speech Recognition</source>
          , volume
          <volume>2</volume>
          .
          <string-name>
            <surname>Pearson</surname>
          </string-name>
          , Prentice Hall,
          <volume>02</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          <string-name>
            <given-names>Evgeny</given-names>
            <surname>Kostuchenko</surname>
          </string-name>
          , Dariya Novokhrestova, Marina Tirskaya, Alexander Shelupanov, Mikhail Nemirovich-Danchenko,
          <string-name>
            <given-names>Evgeny</given-names>
            <surname>Choynzonov</surname>
          </string-name>
          , and
          <string-name>
            <given-names>Lidiya</given-names>
            <surname>Balatskaya</surname>
          </string-name>
          .
          <year>2019</year>
          .
          <article-title>The evaluation process automation of phrase and word intelligibility using speech recognition systems</article-title>
          .
          <source>In International Conference on Speech and Computer</source>
          , pages
          <fpage>237</fpage>
          -
          <lpage>246</lpage>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>