<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <journal-title-group>
        <journal-title>Dec</journal-title>
      </journal-title-group>
    </journal-meta>
    <article-meta>
      <title-group>
        <article-title>On incrementing interpretability of machine learning models from the foundations: a study on syllabic speech units</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Vincenzo Norman Vitale</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Loredana Schettino</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Francesco Cutugno</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>DIETI - University of Naples Federico II</institution>
          ,
          <addr-line>Italy</addr-line>
          ,
          <country country="IT">Italy</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Interdepartmental Research Center Urban/Eco, University of Naples Federico II</institution>
          ,
          <country country="IT">Italy</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2023</year>
      </pub-date>
      <volume>02</volume>
      <issue>2023</issue>
      <fpage>0000</fpage>
      <lpage>0002</lpage>
      <abstract>
        <p>English. Modern ASR systems generally encode information by employing representations that favour performance indicators such as Word Error Rate (WER), making the interpretation of results and the diagnosis of any error extremely dificult if not impossible. In particular, within the context of end-to-end ASR systems, studies have been devoted to investigating the degrees of explainability of such systems by considering the use of diferent sets of linguistic features. This work explores the potential of diferent machine learning algorithms by considering features extracted from syllabic units of analysis and highlights that relying on syllabic Mel-Frequency Cepstral Coeficients increases the interpretability of complex techniques. In fact, the latter currently extract basic units in ways that are highly skewed toward operational convenience. The proposed method would reduce the need for computational resources both in training and in the inference phases, which results in economical and less time-consuming processes.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>
        1. Introduction
sub-word units reducing their impact on memory and
thus allowing for the creation of bigger models with
mil1 The advent of Deep Neural Networks (DNN) enabled lions or even billions of parameters aimed at catching
modern ASR systems and, more in general, Natural Lan- a wider range of natural language nuances. On the one
guage Processing (NLP) systems to perform at their best hand, these techniques definitely improve systems’
perwhen fed with enough training data and supplied with formances and capabilities. On the other hand, they also
suficient computational resources. The recent tendency reduce models’ interpretability from their foundations,
is to focus eforts on incrementing performance indica- which not only makes them increasingly similar to black
tors like Word Error Rate (WER), making DNN models boxes but also augments their need for computational
behind the scenes increasingly complex and larger, with resources. Wav2Vec2 authors [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ] suggest that "switching
the efect of a dramatic reduction in their interpretability to a seq2seq architecture and a word piece vocabulary"
and an increase in the number of parameters consid- would result in performance gains. In line with this, the
ered and therefore in the required computation efort employment of larger and linguistically motivated units,
[
        <xref ref-type="bibr" rid="ref1 ref2">1, 2</xref>
        ]. As an example, state-of-the-art End-to-End (E2E) like syllables, could bring several advantages. Firstly, it
ASR systems [
        <xref ref-type="bibr" rid="ref3 ref4 ref5 ref6">3, 4, 5, 6</xref>
        ] employ self-supervised learning would improve performance in terms of WER and
comtechniques to determine, based on huge amounts of unla- putational resources required to train and operate these
belled data, the best representation of the speech signal systems. Secondly, it would increment the system’s
inbased on fixed-length units, which results in adaptable terpretability, allowing domain experts (i.e. linguists,
systems. In the same way, Big Language Models (Big especially phoneticians) to dive deep into error analysis,
LM) employ advanced encoding techniques, like those which means favoring interpretable rather than
computabased on Byte Pair Encoding (BPE)[
        <xref ref-type="bibr" rid="ref7 ref8 ref9">7, 8, 9</xref>
        ], to encode tionally eficient but poorly understandable inputs. The
main contributions of this study are:
• the proposal of an interpretable approach to
speech-oriented feature extraction based on
syllable;
• a comparison of various classification techniques
with diferent interpretability grades.
      </p>
    </sec>
    <sec id="sec-2">
      <title>2. Related Work</title>
      <p>
        of vocalic (V) and consonantal (C) sounds with diferent
degrees of complexity. However, in many languages, CV
2.1. Explaining Modern ASR is described as the most common structure and the most
resistant to phonetic variation and the related reduction
Among the drawbacks of modern Deep Neural Networks phenomena [25, 26, 27].
(DNN) based systems, the most frequently cited are their While disagreements have mainly concerned
determinpoor interpretability [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ], the lack of suficient training ing boundaries between diferent units, the alternation
corpora [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ] and the high demand for computational re- of diferent units can be grounded on the principles of
sources [
        <xref ref-type="bibr" rid="ref12 ref13 ref14">12, 13, 14</xref>
        ]. In recent years, studies have been sonority scale and onset maximization [28]. Thus, the
devoted to the investigation of the degrees of explain- syllable can be described as a sequence of speech sounds
ability of ASR and, more in general, of speech-related where the onset of the sequence is less intense than the
systems based on DNN. Some of these works aim to preceding coda.
interpret the internal model dynamics and overall ‘be- Based on the structural integrity of the syllable,
evhaviour’ through model-output backtracing or simula- idence has been provided that the syllable rather than
tions via explainability methods [
        <xref ref-type="bibr" rid="ref15 ref16 ref17">15, 16, 17</xref>
        ]. In some phonetic segments can represent a relevant basic unit of
other cases, ‘probing’techniques have been employed to speech production and perception [29, 30]. In fact, [26]
investigate what’s encoded in DNN layers at diferent shows that the observable variation in connected speech
‘depths’ [
        <xref ref-type="bibr" rid="ref18 ref19">18, 19</xref>
        ] by introducing probes aimed at catch- is more systematic at the level of the syllable than at one
ing intermediate internal representations to be used for of the phonetic segments.
various tasks (e.g., regression or classification). Through
classification and measurements, some probing studies
have analysed how the accent of pronunciation in dif- 3. Material and Methods
ferent English varieties influences the performance of
DeepSpeech2 [
        <xref ref-type="bibr" rid="ref20 ref21">20, 21</xref>
        ]. These studies employed linguistic- To achieve our defined goal, we considered an input
conrelated features and highlighted how the contextual pho- sisting of syllables from datasets manually annotated by
netic information contained in intermediate representa- domain experts and evaluated the performances achieved
tions influenced the classification. A further study used by diferent classification methods when relying on
disprobing to investigate the multi-temporal modelling of tinct sets of syllable-based features.
phonetic information in the Wav2Vec2 ASR [
        <xref ref-type="bibr" rid="ref22 ref3">3, 22</xref>
        ]. Some
authors have also proposed a spectrogram-like repre- 3.1. Corpus and annotation
sentation of emissions that could be used for speaker
identification and speech synthesis [
        <xref ref-type="bibr" rid="ref23">23</xref>
        ].
      </p>
      <p>However, although some studies did consider
linguistic features to explain the behaviour of existing models,
these were related to isolated fixed-length segments and,
to the best of our knowledge, did not take into account
larger linguistically meaningful units (i.e. syllables).</p>
      <sec id="sec-2-1">
        <title>This study is based on the Italian and Spanish datasets of</title>
        <p>the Nocando corpus [31] which consists of spoken
narrative texts produced by 11 Italian and 6 Spanish subjects.</p>
        <p>The audio files and their transcriptions were processed
using the WebMAUS Basic services [32] which provided
automatic phonetic transcriptions. The latter were
manually edited in Praat [33] and syllabified according to the
principles of sonority sequencing and onset
maximization [28]. Syllabic units were also annotated for their
phonetic structural pattern (CV, CCV, CCVC, CVC, VC,
V).
2.2. Syllabic unit structure
The notion of syllable is quite well-known in linguistic
studies. However, its definition has been much debated
since it consists of dynamic and complex structures that The Italian dataset consists of 940 syllables. As
excan be analysed from diferent perspectives, i.e., phono- pected, the structural patterns are not evenly distributed,
logical or phonetics, and involve various aspects, like but the following distribution is observed: CV (65%), CVC
articulatory gestures coordination, intensity modulation (16%), CCV (9%), VC (3%), CCVC (3%), V (2%).
[24]. Acoustically, a syllable is described as being char- As for the Spanish dataset, it consists of 609 syllables.
acterized by an intensity peak in the speech signal sur- The occurring patterns do not difer much from the Italian
rounded by less intense aggregated sounds. ones: CV (65%), CVC (14%), CCV (7%), VC (5%), V (4%),</p>
        <p>The essential element of a syllabic unit is the sonor- CCVC (3%).
ity peak, the nucleus, which usually consists of vowel
sounds. The nucleus can be accompanied by aggregated 3.2. Syllable-based features
consonant sounds at the beginning of the unit
preceding the nucleus, the onset, or at the end following it, the For our tasks, we assumed the hand-annotated syllables
coda. Diferent languages allow syllabic combinations as base units. These were considered in two diferent
ways. At first, we look at syllables as a single piece of
signal, which is how they have been traditionally considered
and processed. Then, we consider them as a signal
presumably made up of three components, namely the onset,
the nucleus and the coda. The following four feature sets
were considered.
• Lastly, we considered Convolutional Neural</p>
        <p>
          Network (CNN) [41, 42] as they represent
stateof-the-art in speech processing tasks [
          <xref ref-type="bibr" rid="ref3">3</xref>
          ]. The
considered CNN7 consists of a Conv2d layer with
the ReLu activation function, followed by a
FeedForward with a SoftMax. In particular, we choose
to compare two settings: the first one with a
kernel size of 3x3; The second one considers a larger
context with a size of 3x9.
• OPSM consisting of the GeMAPS [34] set from
the OpenSmile toolkit [35]. It is composed of
62 features and provides information about the
whole considered signal, namely the syllable. In the first phase, we compared K-Means, HAC and
• MFCC consists of 13 Mel Frequency Cepstrum SVM as, by their very nature, they are considered more
Coeficients, which represent the most salient interpretable than neural networks. On the one hand,
Kinformation for speech recognition tasks [36], ex- Means and HAC allowed us to explore how sample
grouptracted for each syllable part2, i.e. onset, nucleus ing is afected when the numerousness of clusters is fixed
and coda. or not, without external supervision. Then, SVM provides
• Full namely the concatenation OPSM and MFCC. a robust and interpretable way of supervisingly
evaluat• PCA consists of the Principal Component Analy- ing how samples group when a model is set to learn a
sis 3 (with an explained variance 95%) of the Full few interpretable parameters. Lastly, we evaluated the
set. performance of a CNN on the MFCC feature set to
compare it with the best-performing method among those
In order to avoid biases due to dimensionality, all the from the previous phase, allowing us to compare how
considered feature sets were normalized to achieve zero- an interpretable yet powerful method performs against
mean and unitary variance 4. one of the fundamental building blocks of modern DNN.
We compare performances through the micro averaged
3.3. The experimental protocol F1 score (Equation 1), which is particularly suitable for
multi-classification tasks.
        </p>
        <p>This study consists of a classification task that concerns
samples labelled with syllable patterns and aims at
classifying them on the basis of the considered feature sets. In
particular, we compared the following four techniques:
• The K-Means3[38] a vector-quantization method
which divides n objects in k clusters based on their
mean distance.5
• Hierarchical Agglomerative Clustering
(HAC)3 [39] is a greedy technique which aims
at grouping (or splitting) clusters based on a
similarity measure. The final output is a clusters
hierarchy which could be divided based on the
number of desired clusters.6
• The Support Vector Machine (SVM)3 [40] is a
versatile algorithm used for classification and
regression tasks, whose objective is to find a
hyperplane in a multi-dimensional space that enables
the classification of the considered data points.</p>
        <p>2MFCC components were extracted through the Librosa library
at version 0.9.2.</p>
        <p>3Defined in Scikit-learn [37] version 1.1.3.</p>
        <p>4Normalization was achieved through the Scikit-learn (version
1.1.3) StandardScaler.</p>
        <p>5K-Means parameters: k=6, tolerance to declare
convergence=1e-4, initialization through the k-means++ method,
random state=42, algorithm=LLoyd’s EM algorithm.</p>
        <p>6HAC parameters: clusters=6, metric=euclidean, linkage=ward
 1 =</p>
        <p>+ 12 * (  +   )</p>
        <p>(1)</p>
        <p>Given the fuzzy boundary between syllabic units and
the high degree of variability within each syllable
structure class, which not only concerns the presence or
absence of segments but also their phonetic specification,
the described techniques are applied considering the
pooled types of syllable samples.</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>4. Results</title>
      <p>Figure 1 reports the F1-score achieved by K-Means, HAC
and SVM over all the considered feature sets. The SVM
classifier outperforms both clustering methods on any
feature set. However, this was not our primary goal. Note
how, for any of the considered methods, the performance
diference between the MFCC set and the PCA one is
rather small.</p>
      <p>For the SVM, results reported are referred to the
optimal configuration, which has been found through a
grid search on C (between 0.5 and 10 with a step of 0.5),
gamma (within 0.01, 0.001 and 0.0001) and kernel type
(within rbf, polynomial and sigmoid). The train,
validation and dev set were respectively 60/20/20 of the original</p>
      <sec id="sec-3-1">
        <title>7Implemented with Pytorch 1.13 and Pytorch-lightning 1.8.3</title>
        <p>dataset. Data splits were balanced on the combination of
the pattern (i.e. the label) and language.</p>
      </sec>
      <sec id="sec-3-2">
        <title>The confusion matrix that reports the output of the</title>
        <p>SVM classifier operating on the basis of MFCC (Fig. 2),
highlights that better performances concern the CV
structure, which is also the more frequent in the data. As
for the other structures, misclassification cases mostly
concern their identification as CV, which reveals that
when considering syllabic units actually occurring in the
speech signal a particularly high similarity emerges
between more complex patterns, i.e. CCV and CVC, and
the CV pattern.</p>
        <p>Lastly, Fig. 3 reports the comparison between the SVM
classifier and the considered CNN configurations on the
MFCC feature set. Still, the SVM performed better than
both CNN-based configurations.</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>5. Discussion and Conclusions</title>
      <p>In this study, we evaluated the use of phonetic syllables
as basic units for speech-related tasks aimed at
preserving and, if possible, incrementing the interpretability of
diferent learning techniques. We employed four
diferent feature sets extracted upon the assumption of the
phonetic syllable as a fundamental unit, considering
different points of view: the more theoretically informed
MFCC-based that is strongly tied to the signal; the more
analytic OPSM based on Opensmile statistical analysis;
the most extended which is a combination of both; the
most computationally eficient based on the PCA
analysis. Then, upon these feature sets, we evaluated the
performance of three well-known machine learning
techniques known for being highly interpretable. Finally, we
compared the best-performing model, namely the SVM,
with a convolutional network on the MFCC set,
obtaining comparable performances. Our preliminary results
highlight that a set of features aimed at keeping things
interpretable, namely the MFCC, lets diferent methods
achieve performances that are comparable to those of
richer (Full), analytic (OPSM) or computationally
optimized (PCA) sets, which do not retain the same
interpretability grade. These findings corroborate the idea
that training speech-oriented learning models on larger
and linguistically meaningful units could increase the
capacity of domain experts and software/ml engineers
to diagnose system failures and, at the same time, help
reduce the efort and computational resources needed
for signal preprocessing. Ongoing analyses involve the
enlargement of the annotated datasets to improve the
results of further classification trials. In Appendix we
reported some preliminary results of a classification trial on
an extended dataset, about 26 minutes of hand-annotated
speech, consisting of 3589 phonetic syllables. In future
works, we plan to extend this kind of study to recent
architectures like Squeezeformer[43] or CNN-BLSTM[44].
work layer hear? analyzing hidden representa- [37] F. Pedregosa, G. Varoquaux, A. Gramfort, V. Michel,
tions of end-to-end asr through speech synthesis, B. Thirion, O. Grisel, M. Blondel, P. Prettenhofer,
in: ICASSP 2020-2020 IEEE International Confer- R. Weiss, V. Dubourg, J. Vanderplas, A. Passos,
ence on Acoustics, Speech and Signal Processing D. Cournapeau, M. Brucher, M. Perrot, E.
Duch(ICASSP), IEEE, 2020, pp. 6434–6438. esnay, Scikit-learn: Machine learning in Python,
[24] J. Laver, L. John, Principles of phonetics, Cambridge Journal of Machine Learning Research 12 (2011)
university press, 1994. 2825–2830.
[25] P. Maturi, I suoni delle lingue, i suoni dell’italiano: [38] J. A. Hartigan, M. A. Wong, Algorithm as 136: A
knuova introduzione alla fonetica, Bologna: il means clustering algorithm, Journal of the royal
staMulino, 2014. tistical society. series c (applied statistics) 28 (1979)
[26] S. Greenberg, Speaking in shorthand–a syllable- 100–108.</p>
      <p>centric perspective for understanding pronuncia- [39] F. Murtagh, A survey of recent advances in
hierartion variation, Speech Communication 29 (1999) chical clustering algorithms, The computer journal
159–176. 26 (1983) 354–359.
[27] L. Schettino, V. N. Vitale, F. Cutugno, Syllabic reduc- [40] W. S. Noble, What is a support vector machine?,
tion in Italian connected speech: towards the inte- Nature biotechnology 24 (2006) 1565–1567.
gration of linguistic and computational approaches, [41] K. O’Shea, R. Nash, An introduction to
conin: Proceedings of the 20th International Congress volutional neural networks, arXiv preprint
of Phonetic Sciences (ICPhS 2023), 2023, pp. 2015– arXiv:1511.08458 (2015).</p>
      <p>2019. [42] S. Albawi, T. A. Mohammed, S. Al-Zawi,
Under[28] M. Nespor, Fonologia, Bologna: Il Mulino, 1993. standing of a convolutional neural network, in:
[29] F. Albano Leoni, The boundaries of the syllable, in: 2017 international conference on engineering and
D. Russo (Ed.), The Notion of Syllable across His- technology (ICET), Ieee, 2017, pp. 1–6.
tory, Theories and Analysis, Cambridge Scholars [43] S. Kim, A. Gholami, A. Shaw, N. Lee, K. Mangalam,
Publishing, 2016. J. Malik, M. W. Mahoney, K. Keutzer,
Squeeze[30] F. Cangemi, O. Niebuhr, Rethinking reduction and former: An eficient transformer for automatic
canonical forms, Rethinking reduction (2018) 277– speech recognition, Advances in Neural
Informa302. tion Processing Systems 35 (2022) 9361–9373.
[31] L. Brunetti, S. Bott, J. Costa, E. Vallduví, A mul- [44] D. Wang, X. Wang, S. Lv, End-to-end mandarin
tilingual annotated corpus for the study of infor- speech recognition combining cnn and blstm,
Symmation structure 1, in: Grammatik und Korpora metry 11 (2019) 644.
2009. Dritte internationale Konferenz, Mannheim,
22.-24.09. 2009. Grammar &amp; corpora 2009, 2011.
[32] T. Kisler, U. Reichel, F. Schiel, Multilingual
processing of speech via web services, Computer Speech
&amp; Language 45 (2017) 326–347.
[33] P. Boersma, D. Weenink, Praat: doing phonetics by
computer [computer program]. version 5.3. 51,
Online: http://www. praat. org/retrieved, last viewed
on 12 (1999-2022).
[34] F. Eyben, K. R. Scherer, B. W. Schuller, J. Sundberg,</p>
      <p>E. André, C. Busso, L. Y. Devillers, J. Epps, P. Laukka,
S. S. Narayanan, et al., The geneva minimalistic
acoustic parameter set (gemaps) for voice research
and afective computing, IEEE transactions on
affective computing 7 (2015) 190–202.
[35] F. Eyben, M. Wöllmer, B. Schuller, Opensmile: the
munich versatile and fast open-source audio
feature extractor, in: Proceedings of the 18th ACM
international conference on Multimedia, 2010, pp.</p>
      <p>1459–1462.
[36] X. Huang, A. Acero, H.-W. Hon, R. Reddy, Spoken
language processing: A guide to theory, algorithm,
and system development, Prentice hall PTR, 2001.</p>
      <sec id="sec-4-1">
        <title>In Figures 5 and 4, we reported the confusion matrix of</title>
        <p>the best-performing SVM classifier on MFCC features, in
the normalized and non-normalized versions respectively.
These preliminary results were obtained on an extended
version of the dataset which is currently under further
analysis.</p>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>Y.</given-names>
            <surname>Belinkov</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Ali</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Glass</surname>
          </string-name>
          ,
          <article-title>Analyzing phonetic and graphemic representations in end-toend automatic speech recognition</article-title>
          , arXiv preprint arXiv:
          <year>1907</year>
          .
          <volume>04224</volume>
          (
          <year>2019</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>D. V.</given-names>
            <surname>Carvalho</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E. M.</given-names>
            <surname>Pereira</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. S.</given-names>
            <surname>Cardoso</surname>
          </string-name>
          ,
          <article-title>Machine learning interpretability: A survey on methods and metrics</article-title>
          ,
          <source>Electronics</source>
          <volume>8</volume>
          (
          <year>2019</year>
          )
          <fpage>832</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>A.</given-names>
            <surname>Baevski</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Zhou</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Mohamed</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Auli</surname>
          </string-name>
          , wav2vec
          <volume>2</volume>
          .
          <article-title>0: A framework for self-supervised learning of speech representations</article-title>
          ,
          <source>Advances in neural information processing systems</source>
          <volume>33</volume>
          (
          <year>2020</year>
          )
          <fpage>12449</fpage>
          -
          <lpage>12460</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>W.-N.</given-names>
            <surname>Hsu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Bolte</surname>
          </string-name>
          , Y.
          <string-name>
            <surname>-H. H. Tsai</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          <string-name>
            <surname>Lakhotia</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          <string-name>
            <surname>Salakhutdinov</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Mohamed</surname>
          </string-name>
          , Hubert:
          <article-title>Self-supervised speech representation learning by masked prediction of hidden units</article-title>
          ,
          <source>IEEE/ACM Transactions on Audio, Speech, and Language Processing</source>
          <volume>29</volume>
          (
          <year>2021</year>
          )
          <fpage>3451</fpage>
          -
          <lpage>3460</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>A.</given-names>
            <surname>Baevski</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W.-N.</given-names>
            <surname>Hsu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Q.</given-names>
            <surname>Xu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Babu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Gu</surname>
          </string-name>
          ,
          <string-name>
            <surname>M.</surname>
          </string-name>
          <article-title>Auli, Data2vec: A general framework for self-supervised learning in speech, vision and language</article-title>
          ,
          <source>in: International Conference on Machine Learning, PMLR</source>
          ,
          <year>2022</year>
          , pp.
          <fpage>1298</fpage>
          -
          <lpage>1312</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>Z.</given-names>
            <surname>Zhang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Zhou</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Ao</surname>
          </string-name>
          , S. Liu,
          <string-name>
            <given-names>L.</given-names>
            <surname>Dai</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Wei</surname>
          </string-name>
          , Speechut:
          <article-title>Bridging speech and text with hiddenunit for encoder-decoder based speech-text pretraining</article-title>
          ,
          <source>in: Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing</source>
          ,
          <year>2022</year>
          , pp.
          <fpage>1663</fpage>
          -
          <lpage>1676</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>R.</given-names>
            <surname>Sennrich</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Haddow</surname>
          </string-name>
          ,
          <string-name>
            <surname>A. Birch,</surname>
          </string-name>
          <article-title>Neural machine translation of rare words with subword units</article-title>
          ,
          <source>arXiv preprint arXiv:1508.07909</source>
          (
          <year>2015</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>A.</given-names>
            <surname>Radford</surname>
          </string-name>
          ,
          <string-name>
            <surname>J. Wu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Child</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Luan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Amodei</surname>
          </string-name>
          ,
          <string-name>
            <given-names>I.</given-names>
            <surname>Sutskever</surname>
          </string-name>
          , et al.,
          <article-title>Language models are unsupervised multitask learners</article-title>
          ,
          <source>OpenAI blog 1</source>
          (
          <year>2019</year>
          )
          <article-title>9</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>C.</given-names>
            <surname>Rafel</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Shazeer</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Roberts</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Lee</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Narang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Matena</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Zhou</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P. J.</given-names>
            <surname>Liu</surname>
          </string-name>
          ,
          <article-title>Exploring the limits of transfer learning with a unified text-totext transformer</article-title>
          ,
          <source>The Journal of Machine Learning Research</source>
          <volume>21</volume>
          (
          <year>2020</year>
          )
          <fpage>5485</fpage>
          -
          <lpage>5551</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>Y.</given-names>
            <surname>Zhang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Tiňo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Leonardis</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Tang</surname>
          </string-name>
          ,
          <article-title>A survey on neural network interpretability</article-title>
          ,
          <source>IEEE Transactions on Emerging Topics in Computational Intelligence</source>
          <volume>5</volume>
          (
          <year>2021</year>
          )
          <fpage>726</fpage>
          -
          <lpage>742</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>W. X.</given-names>
            <surname>Zhao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Zhou</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Tang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Hou</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Min</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Zhang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Zhang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Dong</surname>
          </string-name>
          , et al.,
          <article-title>A survey of large language models</article-title>
          ,
          <source>arXiv preprint arXiv:2303.18223</source>
          (
          <year>2023</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>E.</given-names>
            <surname>Strubell</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Ganesh</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>McCallum</surname>
          </string-name>
          ,
          <article-title>Energy and policy considerations for deep learning in nlp</article-title>
          , arXiv preprint arXiv:
          <year>1906</year>
          .
          <volume>02243</volume>
          (
          <year>2019</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <given-names>D.</given-names>
            <surname>Patterson</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Gonzalez</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Q.</given-names>
            <surname>Le</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Liang</surname>
          </string-name>
          , L.
          <string-name>
            <surname>-M. Munguia</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          <string-name>
            <surname>Rothchild</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          <string-name>
            <surname>So</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          <string-name>
            <surname>Texier</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          <string-name>
            <surname>Dean</surname>
          </string-name>
          ,
          <article-title>Carbon emissions and large neural network training</article-title>
          ,
          <source>arXiv preprint arXiv:2104.10350</source>
          (
          <year>2021</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <given-names>D.</given-names>
            <surname>Patterson</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Gonzalez</surname>
          </string-name>
          ,
          <string-name>
            <given-names>U.</given-names>
            <surname>Hölzle</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Q.</given-names>
            <surname>Le</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Liang</surname>
          </string-name>
          , L.
          <string-name>
            <surname>-M. Munguia</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          <string-name>
            <surname>Rothchild</surname>
            ,
            <given-names>D. R.</given-names>
          </string-name>
          <string-name>
            <surname>So</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          <string-name>
            <surname>Texier</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          <string-name>
            <surname>Dean</surname>
          </string-name>
          ,
          <article-title>The carbon footprint of machine learning training will plateau, then shrink</article-title>
          ,
          <source>Computer</source>
          <volume>55</volume>
          (
          <year>2022</year>
          )
          <fpage>18</fpage>
          -
          <lpage>28</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <given-names>A.</given-names>
            <surname>Holzinger</surname>
          </string-name>
          , G. Langs,
          <string-name>
            <given-names>H.</given-names>
            <surname>Denk</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Zatloukal</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Müller</surname>
          </string-name>
          ,
          <article-title>Causability and explainability of artificial intelligence in medicine</article-title>
          ,
          <source>Wiley Interdisciplinary Reviews: Data Mining and Knowledge Discovery</source>
          <volume>9</volume>
          (
          <year>2019</year>
          )
          <article-title>e1312</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [16]
          <string-name>
            <given-names>B.</given-names>
            <surname>Kailkhura</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Gallagher</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Kim</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Hiszpanski</surname>
          </string-name>
          , T. Han,
          <article-title>Reliable and explainable machine-learning methods for accelerated material discovery</article-title>
          ,
          <source>npj Computational Materials</source>
          <volume>5</volume>
          (
          <year>2019</year>
          )
          <fpage>1</fpage>
          -
          <lpage>9</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [17]
          <string-name>
            <given-names>P.</given-names>
            <surname>Angelov</surname>
          </string-name>
          , E. Soares,
          <article-title>Towards explainable deep neural networks (xdnn</article-title>
          ),
          <source>Neural Networks</source>
          <volume>130</volume>
          (
          <year>2020</year>
          )
          <fpage>185</fpage>
          -
          <lpage>194</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          [18]
          <string-name>
            <given-names>J.</given-names>
            <surname>Yosinski</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Clune</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Nguyen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Fuchs</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Lipson</surname>
          </string-name>
          ,
          <article-title>Understanding neural networks through deep visualization</article-title>
          ,
          <source>arXiv preprint arXiv:1506.06579</source>
          (
          <year>2015</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          [19]
          <string-name>
            <given-names>G.</given-names>
            <surname>Alain</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Bengio</surname>
          </string-name>
          ,
          <article-title>Understanding intermediate layers using linear classifier probes</article-title>
          ,
          <source>arXiv preprint arXiv:1610.01644</source>
          (
          <year>2016</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          [20]
          <string-name>
            <given-names>T.</given-names>
            <surname>Viglino</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Motlicek</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Cernak</surname>
          </string-name>
          ,
          <article-title>End-to-end accented speech recognition</article-title>
          .,
          <source>in: Interspeech</source>
          ,
          <year>2019</year>
          , pp.
          <fpage>2140</fpage>
          -
          <lpage>2144</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          [21]
          <string-name>
            <given-names>A.</given-names>
            <surname>Prasad</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Jyothi</surname>
          </string-name>
          ,
          <article-title>How accents confound: Probing for accent information in end-to-end speech recognition systems</article-title>
          ,
          <source>in: Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics</source>
          ,
          <year>2020</year>
          , pp.
          <fpage>3739</fpage>
          -
          <lpage>3753</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          [22]
          <string-name>
            <given-names>D.</given-names>
            <surname>Ma</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Ryant</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Liberman</surname>
          </string-name>
          ,
          <article-title>Probing acoustic representations for phonetic properties</article-title>
          ,
          <source>in: ICASSP 2021-2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)</source>
          , IEEE,
          <year>2021</year>
          , pp.
          <fpage>311</fpage>
          -
          <lpage>315</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          [23]
          <string-name>
            <surname>C.-Y. Li</surname>
            ,
            <given-names>P.-C.</given-names>
          </string-name>
          <string-name>
            <surname>Yuan</surname>
            , H.-
            <given-names>Y.</given-names>
          </string-name>
          <string-name>
            <surname>Lee</surname>
          </string-name>
          ,
          <article-title>What does a net-</article-title>
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>