<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <journal-title-group>
        <journal-title>These authors contributed equally.
$ taufique.ahmed@tudublin.ie (T. Ahmed); luca.longo@tudublin.ie (L. Longo)
 lucalongo.eu/about (L. Longo)</journal-title>
      </journal-title-group>
    </journal-meta>
    <article-meta>
      <title-group>
        <article-title>Latent space interpretation and visualisation for understanding the decisions of convolutional variational autoencoders trained with EEG topographic maps⋆</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Taufique Ahmed</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Luca Longo</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>School of Computer Science, Technological University Dublin</institution>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2023</year>
      </pub-date>
      <volume>000</volume>
      <fpage>0</fpage>
      <lpage>0001</lpage>
      <abstract>
        <p>Learning essential features and forming simple representations of electroencephalography (EEG) signals are dificult problems. Variational autoencoders (VAEs) can be used with EEG signals to learn the salient features of EEG data. But explainability should disclose knowledge of how the model makes its decision. The key contribution of this research is the combining of known components in a pipeline that allows us to give meaningful visualisations that help us understand which component of latent space is responsible for capturing which region of brain activation in EEG topographic maps. The results reveal that each component in the latent space contributes to capturing at least two generating factors in topographic maps. This pipeline can be used to produce EEG topographic maps of any scale. Furthermore, assist us in understanding each component of latent space responsible for activating a portion of the brain.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;Electroencephalography</kwd>
        <kwd>Convolutional variational autoencoder</kwd>
        <kwd>latent space interpretation</kwd>
        <kwd>deep learning</kwd>
        <kwd>spectral topographic maps</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>
        Electroencephalography (EEG) is a method of recording brain activity (electrical potentials)
using electrodes placed on the scalp [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. Some research, for example, has transformed EEG
signals into topographic power head maps to preserve spatial information [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]. Convolutional
neural networks are frequently employed to reduce their dimensionality and automatically learn
essential features [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]. An Autoencoder (AE) is a deep learning neural network architecture that
uses unsupervised learning to learn eficient codings without the usage of labelled input [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ]. A
Variational Autoencoder (VAE) is a form of autoencoder that creates a probabilistic model of the
input sample and then reconstructs it using that model. VAEs have shown a wide application
with electroencephalographic (EEG) signals [
        <xref ref-type="bibr" rid="ref5 ref6 ref7">5, 6, 7</xref>
        ]. However, research into interpreting the
latent space of a variational autoencoder to determine the importance of each latent space
component in capturing the generating factor in spatially preserving EEG topographic maps
is limited. In this study, the goal is to tackle the research problem to learn the importance of
each latent component of VAE, trained with spectral topographic EEG maps for capturing the
generative factors of EEG data. Therefore, the research question being addressed is: Can we
understand the reconstruction capacity of a convolutional variational autoencoder trained with
spectral topographic maps by interpreting and visualising its learnt latent space representation?
The rest of the work is organised as follows. Section 2 investigates related work, whereas
Section 3 describes an empirical study and its methodology. Section 4 presents the experimental
results and findings. Finally, Section 5 concludes the manuscript by describing the contribution
to the body of knowledge and highlighting future work directions.
      </p>
    </sec>
    <sec id="sec-2">
      <title>2. Related Work</title>
      <p>
        Traditional Autoencoders (AE) aim to learn prominent latent representations from unlabeled
input while ignoring irrelevant features. Variational Autoencoders (VAEs) was recently proposed
as an efective extension of AEs, for modeling a data’s probability distribution and learning a
latent space, usually of a lower dimension. It is ideal for unsupervised learning to understand the
impact and importance of each latent component for capturing the number of true generative
factors. VAE-based latent space analysis and decoding of EEG signals are important since they
can precisely define and determine the latent relevant features [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ]. VAE has been constructed
with two distinct encoders to map the input into  and , respectively, and then deliver the
concatenated code to the decoder to reconstruct the input[
        <xref ref-type="bibr" rid="ref9">9</xref>
        ]. Another study used VAE and
manually adjusted the latent activations, allowing the user to see the efect of diferent latent
values on the generated output [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ]. After all, if we want learned latent space representation to
be interpretable, the latent component must have clear-cut meaning [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ]. The researcher also
illustrated how a VAE model’s latent space may be made more explainable by utilizing latent
space regularisation [
        <xref ref-type="bibr" rid="ref12 ref13 ref14">12, 13, 14</xref>
        ]. The majority of time series data are mapped to prominent
representational characteristics and interpretation of its latent space results in improved
clustering performance [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ]. Because the disentanglement of latent space performs the clustering
operation, no further clustering approach is required [
        <xref ref-type="bibr" rid="ref16">16, 17</xref>
        ]. Despite broad application and
study into the interpretation of latent space, knowing VAE reconstruction decisions and the
impact of its components on capturing the number of genuine generating elements remains
inadequate.
      </p>
    </sec>
    <sec id="sec-3">
      <title>3. Research design and methodology</title>
      <p>In this study, if CNN-VAE is trained with spatially preserved EEG topographic maps and its
interpretation of the learned latent space representation provides the knowledge of how well
each component of the latent representation contributes to capturing the number of true
generative factors in spatially preserving EEG topographic maps via visual plausibility. The
detailed design of this research is illustrated in figure 1.</p>
      <p>The DEAP dataset was chosen because it contains multi-channel EEG recordings with 32
participants who watched 40 one-minute music video clips and tasks [18]. These EEG signals
were transformed into 40 × 40 interpolated topographic maps that preserve spatial information
about brain activation [19], a similar experiment was carried out in another investigation [20],
as illustrated in (figure 1, B). Following the creation of the topographic maps, a Convolutional
Variational Autoencoder (CNN-VAE) is built. The encoder network of CNN-VAE takes a 40 × 40
tensor (as seen in figure 1, C) and defines the approximate posterior distribution ( | ). The
CNN-VAE decoder is a generative network that takes a latent space  as input and returns
the reconstructed EEG topographic maps. The architecture (figure 1, C) is made up of four 2D
convolutional layers, each followed by a max pooling layer to minimize the dimension of the
feature maps. In each convolutional layer, ReLU is employed as the activation function. To avoid
overfitting, an early stopping strategy with a patience value of ten epochs is used, which indicates
that training is stopped if the validation loss does not improve for ten consecutive epochs.
To examine the number of generative factors captured from each active latent component,
the reconstructed EEG topographic map from the decoder of CNN-VAE with latent space
representation of only one active component is passed as an input to the K-means algorithm. In
terms of how well samples are clustered with other samples that are similar to each other, the
silhouette score is used to evaluate the quality of clusters generated using clustering methods
such as K-Means. The reconstruction capacity of this model is evaluated by Structural Similarity
Index (SSIM), Mean Squared Error (MSE), Mean Absolute Error (MAE), and Mean Absolute
Percentage Error (MAPE) derived for the reconstructed topographic maps.</p>
    </sec>
    <sec id="sec-4">
      <title>4. Results &amp; Discussion</title>
      <p>The reconstruction capacity of the CNN-VAE model and its explainability is described with
two diferent scenarios because it contains two decoder networks. One trained with all latent
components, whereas the other trained with only one active latent at a time. In the second
scenario, interpreting the disentangled representation of CNN-VAE, where its decoder network
is trained only with one latent component alternatively and the remaining 24 components
are set to zeros to examine the impact of each component for generating the patterns in EEG
topo maps. The results show that each component contributes diferently to capturing the
generating aspects in topo maps. The empirical experiment was carried out using test data, with
10 samples chosen at random to assess the impact of each latent component in capturing the
number of fundamental generative elements in spatially preserving EEG topographic maps. In
order to analyse the results Using visual plausibility, ten images of test data and reconstructed
images with active latent space components 0 are plotted, with findings clearly suggesting that
each component is learning two to three patterns from those EEG topographic maps. These
explanations can be used to gain the trust of stakeholders by demonstrating the visual plausibility
of each latent component in capturing the generative components in EEG topographic maps.</p>
    </sec>
    <sec id="sec-5">
      <title>5. Conclusion</title>
      <p>Research on the interpretation of disentangled representations of VAE trained with spatially
preserving EEG topographic maps is currently limited. A CNN-VAE decoder network is trained
with alternatively one active latent component, and the remaining component is set to zero
because the mean value is close to zero in the distribution learned from each latent
component. The results with visual plausibility show that each component contributes diferently to
capturing and generating aspects in topo maps. Hence, this pipeline helps us understand each
component of latent space responsible for activating a part of the brain region. Future studies
will include understanding the decision of CNN-VAE through the interpretation of its latent
space via clustering and visual plausibility, taking into account the signal-to-noise ratio and
correlation values across the input and output of the architecture.
intelligence, volume 33, 2019, pp. 4610–4617.
[17] V. Prasad, D. Das, B. Bhowmick, Variational clustering: Leveraging variational
autoencoders for image clustering, in: 2020 international joint conference on neural networks
(IJCNN), IEEE, 2020, pp. 1–10.
[18] S. Koelstra, C. Muhl, M. Soleymani, J.-S. Lee, A. Yazdani, T. Ebrahimi, T. Pun, A. Nijholt,
I. Patras, Deap: A database for emotion analysis; using physiological signals, IEEE
transactions on afective computing 3 (2011) 18–31.
[19] T. Ahmed, L. Longo, Examining the size of the latent space of convolutional variational
autoencoders trained with spectral topographic maps of eeg frequency bands, IEEE Access
10 (2022) 107575–107586. doi:10.1109/ACCESS.2022.3212777.
[20] A. V. Chikkankod, L. Longo, On the dimensionality and utility of convolutional
autoencoder’s latent space trained with topology-preserving spectral eeg head-maps, Machine
Learning and Knowledge Extraction 4 (2022) 1042–1064.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>C.</given-names>
            <surname>Binnie</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Prior</surname>
          </string-name>
          , Electroencephalography.,
          <source>Journal of Neurology, Neurosurgery &amp; Psychiatry</source>
          <volume>57</volume>
          (
          <year>1994</year>
          )
          <fpage>1308</fpage>
          -
          <lpage>1319</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>E. W.</given-names>
            <surname>Anderson</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G. A.</given-names>
            <surname>Preston</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C. T.</given-names>
            <surname>Silva</surname>
          </string-name>
          ,
          <article-title>Using python for signal processing and visualization</article-title>
          ,
          <source>Computing in science &amp; engineering 12</source>
          (
          <year>2010</year>
          )
          <fpage>90</fpage>
          -
          <lpage>95</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>M.</given-names>
            <surname>Taherisadr</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Joneidi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Rahnavard</surname>
          </string-name>
          ,
          <article-title>Eeg signal dimensionality reduction and classification using tensor decomposition and deep convolutional neural networks</article-title>
          ,
          <source>in: 2019 IEEE 29th International Workshop on Machine Learning for Signal Processing (MLSP)</source>
          , IEEE,
          <year>2019</year>
          , pp.
          <fpage>1</fpage>
          -
          <lpage>6</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>Y.</given-names>
            <surname>Bengio</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Courville</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Vincent</surname>
          </string-name>
          ,
          <article-title>Representation learning: A review and new perspectives</article-title>
          ,
          <source>IEEE transactions on pattern analysis and machine intelligence</source>
          <volume>35</volume>
          (
          <year>2013</year>
          )
          <fpage>1798</fpage>
          -
          <lpage>1828</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>S. M.</given-names>
            <surname>Abdelfattah</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G. M.</given-names>
            <surname>Abdelrahman</surname>
          </string-name>
          ,
          <string-name>
            <surname>M.</surname>
          </string-name>
          <article-title>Wang, Augmenting the size of eeg datasets using generative adversarial networks</article-title>
          ,
          <source>in: 2018 International Joint Conference on Neural Networks (IJCNN)</source>
          , IEEE,
          <year>2018</year>
          , pp.
          <fpage>1</fpage>
          -
          <lpage>6</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>J. F.</given-names>
            <surname>Hwaidi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T. M.</given-names>
            <surname>Chen</surname>
          </string-name>
          ,
          <article-title>A noise removal approach from eeg recordings based on variational autoencoders</article-title>
          ,
          <source>in: 2021 13th International Conference on Computer and Automation Engineering (ICCAE)</source>
          , IEEE,
          <year>2021</year>
          , pp.
          <fpage>19</fpage>
          -
          <lpage>23</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>K.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Yu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Zhu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Liu</surname>
          </string-name>
          , L. Wu,
          <article-title>Feature extraction and identification of alzheimer's disease based on latent factor of multi-channel eeg</article-title>
          ,
          <source>IEEE Transactions on Neural Systems and Rehabilitation Engineering</source>
          <volume>29</volume>
          (
          <year>2021</year>
          )
          <fpage>1557</fpage>
          -
          <lpage>1567</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>X.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Zhao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Song</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Zhang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Pan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Wu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Huo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Niu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <article-title>Latent factor decoding of multi-channel eeg for emotion recognition through autoencoder-like neural networks</article-title>
          ,
          <source>Frontiers in neuroscience 14</source>
          (
          <year>2020</year>
          )
          <fpage>87</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>Z.</given-names>
            <surname>Zheng</surname>
          </string-name>
          , L. Sun,
          <article-title>Disentangling latent space for vae by label relevant/irrelevant dimensions</article-title>
          ,
          <source>in: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition</source>
          ,
          <year>2019</year>
          , pp.
          <fpage>12192</fpage>
          -
          <lpage>12201</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>T.</given-names>
            <surname>Spinner</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Körner</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Görtler</surname>
          </string-name>
          ,
          <string-name>
            <given-names>O.</given-names>
            <surname>Deussen</surname>
          </string-name>
          ,
          <article-title>Towards an interpretable latent space: an intuitive comparison of autoencoders with variational autoencoders</article-title>
          ,
          <source>in: IEEE VIS</source>
          <year>2018</year>
          ,
          <year>2018</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>E.</given-names>
            <surname>Mathieu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Rainforth</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Siddharth</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y. W.</given-names>
            <surname>Teh</surname>
          </string-name>
          ,
          <article-title>Disentangling disentanglement in variational autoencoders</article-title>
          ,
          <source>in: International Conference on Machine Learning, PMLR</source>
          ,
          <year>2019</year>
          , pp.
          <fpage>4402</fpage>
          -
          <lpage>4412</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>N.</given-names>
            <surname>Bryan-Kinns</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Banar</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Ford</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Reed</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Zhang</surname>
          </string-name>
          , S. Colton,
          <string-name>
            <given-names>J.</given-names>
            <surname>Armitage</surname>
          </string-name>
          , et al.,
          <article-title>Exploring xai for the arts: Explaining latent space in generative music (</article-title>
          <year>2022</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <given-names>A.</given-names>
            <surname>Pati</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Lerch</surname>
          </string-name>
          ,
          <article-title>Attribute-based regularization of latent spaces for variational autoencoders</article-title>
          ,
          <source>Neural Computing and Applications</source>
          <volume>33</volume>
          (
          <year>2021</year>
          )
          <fpage>4429</fpage>
          -
          <lpage>4444</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <given-names>P.</given-names>
            <surname>Cristovao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Nakada</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Tanimura</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Asoh</surname>
          </string-name>
          ,
          <article-title>Generating in-between images through learned latent space representation using variational autoencoders</article-title>
          ,
          <source>IEEE Access 8</source>
          (
          <year>2020</year>
          )
          <fpage>149456</fpage>
          -
          <lpage>149467</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <given-names>F.</given-names>
            <surname>Ding</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Yang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Luo</surname>
          </string-name>
          ,
          <article-title>Clustering by directly disentangling latent space</article-title>
          ,
          <source>in: 2022 IEEE International Conference on Image Processing (ICIP)</source>
          , IEEE,
          <year>2022</year>
          , pp.
          <fpage>341</fpage>
          -
          <lpage>345</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [16]
          <string-name>
            <given-names>S.</given-names>
            <surname>Mukherjee</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Asnani</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Lin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Kannan</surname>
          </string-name>
          , Clustergan:
          <article-title>Latent space clustering in generative adversarial networks</article-title>
          ,
          <source>in: Proceedings of the AAAI conference on artificial</source>
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>