<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Convolutional neural network for cognitive task prediction from EEG's auditory steady state responses</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Daniela Montilla-Trochez</string-name>
          <email>daniela.montilla@postgrado.uv.cl</email>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Rodrigo Salas</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Alejandro Bertin</string-name>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Inga Griskova-Bulanova</string-name>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Paulo Lisboa</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Carolina Saavedra</string-name>
          <email>carolina.saavedra@uv.cl</email>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Centro de Investigación y Desarrollo en Ingeniería en Salud</institution>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Liverpool John Moores University</institution>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>Universidad de Valparaíso</institution>
        </aff>
        <aff id="aff3">
          <label>3</label>
          <institution>Vilnius University</institution>
        </aff>
      </contrib-group>
      <abstract>
        <p>The prediction of cognitive tasks from electroencephalography (EEG) signals have allowed to discriminate the cognitive states emitted by the subjects and to carry out robust monitoring of cognition; a fact that is associated with the attention and performance of an individual's behavior, allowing greater control in the experiments. The objective of this work is to perform the prediction of tasks in the function of the auditory steady-state response (ASSR). Twenty-two subjects underwent three types of tasks: counting, reading and rest, accompanied by a constant stimulus. Images were obtained from the Inter Trial phase coherence (ITPC) to train classification algorithms based on convolutional neural networks (CNN) in order to separate the tasks performed by the subjects. Performance evaluation of the classification algorithm shows very good separation between count, read and rest with an AUROC of 0.95. This is significantly better than a feedforward neural network and a pre-trained convolutional deep neural network.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Introduction</title>
      <p>
        Task detection from electroencephalography (EEG) signal allows us to discriminate between specific
cognitive states and so monitor cognition. This is associated with the attention and performance of
an individual’s behavior
        <xref ref-type="bibr" rid="ref9">Papakostas et al. (2017)</xref>
        . Parameters such as functional connectivity can be
used to distinguish between cognitive states based on the individual’s brain
        <xref ref-type="bibr" rid="ref3">Gaut et al. (2018)</xref>
        , also
considering that the last components of the evoked potentials are related to discrimination tasks
that reveal complex cognitive processes
        <xref ref-type="bibr" rid="ref11">Saavedra and Bougrain (2012)</xref>
        ;
        <xref ref-type="bibr" rid="ref10">Saavedra et al. (2019)</xref>
        .
      </p>
      <p>
        It is important to highlight that discrimination of mental tasks has the purpose of monitoring the
behavior of an individual, rather than establishing a correlation between EEG measurements and
the final result of the tasks. In particular, there are patterns that might be able to detect cognitive
states between different users.
        <xref ref-type="bibr" rid="ref8">Palaniappan and Raveendran (2001)</xref>
        .
      </p>
      <p>
        Part of the behavior is associated with the individual’s auditory quality, which can be monitored
through evoked potentials that estimate hearing sensitivity. In fact, specifically the Auditory Steady
State Responses (ASSRs) are used to measure the ability of local cortical networks to generate
activity and thus be able to differentiate individuals with normal hearing sensitivity from those with
varying degrees of auditory sensorineural loss.
        <xref ref-type="bibr" rid="ref6">Korczak et al. (2012)</xref>
        .
      </p>
      <p>
        It should be noted that ASSRs are obtained when an auditory stimulus that is periodically
presented produces an electroencephalographic response. Although there are investigations that
address the study of tasks and use artificial neural networks for the classification of waveforms of
Event-related Potential (ERP)
        <xref ref-type="bibr" rid="ref5">Gupta et al. (1995)</xref>
        , from the EEG, there are few that involve a constant
auditory stimulus.
      </p>
      <p>
        On the other hand, deep neural networks have been applied successfully in different fields
and with performances that outperforms conventional machine learning techniques. In particular,
convolutional neural networks are a special type of deep networks that have proved very effective
in classifying images because they have the ability to extract relevant characteristics, which they
use with a non-linear classifier
        <xref ref-type="bibr" rid="ref4">Goodfellow et al. (2016)</xref>
        . On the other hand, a wide variety of
deep learning models have been proposed in order to classify images in complex contexts (see for
example
        <xref ref-type="bibr" rid="ref7">Mellado et al. (2019)</xref>
        )
      </p>
      <p>The objective of the present work is to make the prediction of the mental task performed by a
subject from the EEG signal when it is under a steady state response to an auditory stimulus.</p>
    </sec>
    <sec id="sec-2">
      <title>Methods and Material</title>
    </sec>
    <sec id="sec-3">
      <title>Acquisition of EEG Signals</title>
      <p>
        We used a dataset from
        <xref ref-type="bibr" rid="ref14">Voicikas et al. (2016)</xref>
        consisting of 28 healthy young male subjects. The
auditory stimulus used was the Click trial, which consisted of 20 identical bursts of white sound
with a duration of 1.5 ms. The subjects underwent three tasks: the first was to count the number
of presentations of the stimulus. In the end, the subjects who report the number of stimuli
specified was requested to guarantee attention, in the second the subject had to ignore the
stimulus presented and try to keep his/her mind blank and, finally, in the third task he/she was
asked to make a silent reading of an easily readable text and presented on a computer screen. At
the end of the experiment, the subjects were briefly questioned about the content of the material
to control the attention.
      </p>
      <p>The channels that were selected to extract the information of the response to the stimulus were
F3, F1, Fz, F2, F4, FC3, FC1, FCz, FC2, FC4, C3, C1, Cz, C2, C4. The sampling frequency was set at
1024Hz, while the number of trials per subject varies between 100 and 120.</p>
    </sec>
    <sec id="sec-4">
      <title>Coherent Averaging</title>
      <p>The coherent averaging of N trials of EEG signals consists in obtaining the average in each instant
of time of segments of signals of equal size and that are started with the stimulus applied. The
objective of applying this technique is to reduce noise and random activations, but on the other
hand it is expected to highlight the evoked potentials in response to the stimulus. In this work,
coherent averages were applied every 15 trials of EGG signals of the same task.</p>
    </sec>
    <sec id="sec-5">
      <title>Inter Trial Phase Coherence</title>
      <p>
        The Inter Trial Phase Coherence (ITPC) is a computational technique that averages the complex
representation of a unit vector, obtained from the phase angle of a trial at a given time, represented
using Euler’s formula. The method was introduced by
        <xref ref-type="bibr" rid="ref13">Tallon-Baudry et al. (1996)</xref>
        . The ITPC is
mathematically defined by the following equation:
n
      </p>
      <p>IT P Ctf = óóóóón*1 Ér=1 eiktfr óóóóó (1)
where n represents the number of trials and eiktfr is the complex polar representation of a phase
angle k on trial r, at time-frequency point tf .</p>
      <p>
        The result of the ITPC is an image whose pixels have values that are in the range of 0 to 1,
where 0 indicates evenly distributed phase angles and 1 indicates completely identical phase angles
corresponding to complete coherence
        <xref ref-type="bibr" rid="ref2">Delorme and Makeig (2004)</xref>
        .
(a) Count
(b) Read
(c) Rest
      </p>
    </sec>
    <sec id="sec-6">
      <title>Classifiers</title>
      <p>
        In this work, 3 types of classifiers will be evaluated, which are explained below:
1. Feedforward or Fully Connected Neural Network: It is a type of artificial neural network
consisting of 1 layer of input neurons, 1 or more layers of hidden neurons and 1 output layer.
All neurons in one layer are connected to all neurons in the next layer. This type of network
has no recurrence, lateral connections, nor connections to layers farther than the consecutive
ones. The learning algorithm used is Backpropagation. These networks have the property of
being universal approximators. (More details of these networks see
        <xref ref-type="bibr" rid="ref1">Allende et al. (2001)</xref>
        )
In this work, the network architecture used consists of 2 hidden layers. The learning algorithm
used is: RMSprop.
2. Convolutional Neural Network (CNN): these are a type of artificial neural networks that
have been successfully applied in computer vision. Its name comes from the mathematical
operation that is carried out in at least one of its layers, the convolution. A CNN is composed
of at least 3 layers and these are:
• Convolution layer: The convolution operation receives the image as input and then
applies a filter or kernel on it. This layer returns a map of characteristics of the original
image and whose dimensions will decrease according to the kernel size.
• Pooling layer: The purpose of this layer is to reduce the spatial dimensions of the input
volume for the next convolutional layer without affecting depth. The reduction in size
and the loss of information is favorable due to the decrease in the size of the network
that leads to a lower overload in the calculation in the following layers and can also
reduce the overfitting.
• Fully Connected Layer: This is used as the last layer in the CNN. The neurons of the filters
are flattened and the information pass through non-linear activation functions. This layer
is responsible for classifying the images. The number of output neurons is equivalent to
the number of classes.
      </p>
      <p>It should be noted that a CNN can consist of several Convolution layers and several Pooling
layers. In this work, the CNN network architecture is composed of two stages, in the first
stage there are the 4 layers of convolution with kernel (3x3), followed by layers of average
pooling and max pooling and 4 layers of batch normalization, which they are responsible
for the extraction of features and dimensionality reduction respectively, in the second stage
there are 2 fully-connected layers, responsible for classification. The learning algorithm used
is called RMSprop.</p>
      <p>
        In this work, an ad-hoc CNN model was developed for the available data, where the resulting
architecture tries to preserve the parsimony. Because in addition to the CNN model ITPC
images are incorporated, the model will be called Hybrid-CNN
3. VGG16: The VGG16 model is a convolutional neural network with specific architecture and
has been applied in different contexts (see
        <xref ref-type="bibr" rid="ref12">Simonyan and Zisserman (2014)</xref>
        ). The architecture
of the model is composed of 16 layers, of which 13 are convolutional and 3 fully-connected.
      </p>
    </sec>
    <sec id="sec-7">
      <title>Framework of the proposed model</title>
      <p>In this paper, a framework for the classification of tasks from ASSR signals is proposed. The scheme
used is shown in figure 2. The proposed scheme consists of the following stages:
1. Data Acquisition: The EEG signals were recorded using a ANT device of 50mV/V and 64</p>
      <p>WaveGuard EEG channels.
2. Channel selection and coherent trial averaging: From the EGG signals, the 15 channels closest
to Cz were selected where the greatest response to the stimulus is visualized (see section
Acquisition of EEG Signals). A coherent averaging of 15 trials for each channel is performed.
3. ITPC: The ITPC method explained in the Inter Trial Phase Coherence section is applied to
generate the spectral images. The resulting images have dimensions 21  717. An image bank
with 7155 samples was obtained.
4. Feature Extraction: This stage consists of a series of filters that are convolved with the input
signals. Afterwards, an activation functions of max-pooling type are applied which generates
a downsampling and, at the same time, they work as detectors of relevant characteristics.
5. Classification: At this stage, fully connected non-linear activation neuron layers were used.</p>
      <p>From the activated filters obtained from the previous stage a characteristic vector is generated
that is processed by the neuronal classifier. The result is the final classification in one of the
following tags: Read, Count and Rest.</p>
    </sec>
    <sec id="sec-8">
      <title>Results</title>
      <p>This section presents the results of the different types of classifiers that were implemented to
predict the tasks from the ASSR data, the best results for each model are reported. In the figure 3 a
distribution of the confusion matrix of each of the classifiers can be visualized
(a) Feedforward Artificial Neural Network</p>
      <p>As can be seen in the ROC curves (figure 4), the VGG16 and hybrid-CNN models are above the
non-discrimination line, otherwise the FANN model is very close, tending the but performance of
the three models, the VGG16 have a good performance with respect to the classification of the
tasks but it is not an optimal model, on the contrary the hybrid-CNN algorithm shows the best
performance in all the tasks obtaining the best results of area under the curve.
(a) Feed forward</p>
    </sec>
    <sec id="sec-9">
      <title>Conclusion</title>
      <p>We have proposed the application of a convolutional neural network to analyze and classify EEG
signals of steady-state auditory responses. To improve the performance of the convolutional neural
network, coherent averaging of 15 trials was performed and images were then obtained by applying
the ITPC method. The results show that with the pipeline of the proposed model for the prediction
of cognitive tasks can be made with an an AUROC of 0.95, corresponding to a sensitivity (recall)
score of 0.85.</p>
      <p>Future work is required in order to increase the number of subjects in the study, specially
considering people with some alteration or that presents cognitive difficulties.</p>
    </sec>
    <sec id="sec-10">
      <title>Acknowledgments</title>
      <p>The authors acknowledge the support of the grant REDI170367 from CONICYT.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          <string-name>
            <surname>Allende</surname>
            <given-names>H</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Moraga</surname>
            <given-names>C</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Salas R. Artificial</surname>
          </string-name>
          <article-title>Neural Networks in Time Series Forecasting: A Comparative Analysis</article-title>
          .
          <source>Kybernetika</source>
          .
          <year>2001</year>
          ;
          <volume>38</volume>
          (
          <issue>6</issue>
          ):
          <fpage>685</fpage>
          -
          <lpage>707</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          <string-name>
            <surname>Delorme</surname>
            <given-names>A</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Makeig</surname>
            <given-names>S.</given-names>
          </string-name>
          <article-title>EEGLAB: an open source toolbox for analysis of single-trial EEG dynamics including independent component analysis</article-title>
          .
          <source>Journal of neuroscience methods</source>
          .
          <year>2004</year>
          ;
          <volume>134</volume>
          (
          <issue>1</issue>
          ):
          <fpage>9</fpage>
          -
          <lpage>21</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          <string-name>
            <surname>Gaut</surname>
            <given-names>G</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Li</surname>
            <given-names>X</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Turner</surname>
            <given-names>B</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Cunningham</surname>
            <given-names>WA</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lu</surname>
            <given-names>ZL</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Steyvers</surname>
            <given-names>M.</given-names>
          </string-name>
          <article-title>Predicting Task and Subject Differences with Functional Connectivity and BOLD Variability</article-title>
          . arXiv:
          <volume>180704745</volume>
          [q-bio].
          <year>2018</year>
          Jul; http://arxiv.org/abs/
          <year>1807</year>
          . 04745, arXiv:
          <year>1807</year>
          .04745.
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          <string-name>
            <surname>Goodfellow</surname>
            <given-names>I</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bengio</surname>
            <given-names>Y</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Courville</surname>
            <given-names>A</given-names>
          </string-name>
          .
          <article-title>Deep Learning</article-title>
          . MIT Press;
          <year>2016</year>
          . http://www.deeplearningbook.org.
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          <string-name>
            <surname>Gupta</surname>
            <given-names>L</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Molfese</surname>
            <given-names>DL</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Tammana</surname>
            <given-names>R.</given-names>
          </string-name>
          <article-title>An artificial neural-network approach to ERP classification. Brain and cognition</article-title>
          .
          <year>1995</year>
          ;
          <volume>27</volume>
          (
          <issue>3</issue>
          ):
          <fpage>311</fpage>
          -
          <lpage>330</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          <string-name>
            <surname>Korczak</surname>
            <given-names>P</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Smart</surname>
            <given-names>J</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Delgado</surname>
            <given-names>R</given-names>
          </string-name>
          ,
          <string-name>
            <surname>M Strobel</surname>
            <given-names>T</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bradford</surname>
            <given-names>C</given-names>
          </string-name>
          , Auditory
          <string-name>
            <surname>Steady-State Responses</surname>
          </string-name>
          ;
          <year>2012</year>
          . https://www. ingentaconnect.com/content/aaa/jaaa/2012/023/003/art03, doi: info:doi/10.3766/jaaa.23.
          <issue>3</issue>
          .3.
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          <string-name>
            <surname>Mellado</surname>
            <given-names>D</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Saavedra</surname>
            <given-names>C</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chabert</surname>
            <given-names>S</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Torres</surname>
            <given-names>R</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Salas</surname>
            <given-names>R</given-names>
          </string-name>
          .
          <article-title>Self-Improving Generative Artificial Neural Network for Pseudo-Rehearsal Incremental Class Learning</article-title>
          . Preprints.
          <year>2019</year>
          ;
          <volume>2019070121</volume>
          :
          <fpage>1</fpage>
          -
          <lpage>17</lpage>
          . Doi:
          <volume>10</volume>
          .20944/preprints201907.
          <fpage>0121</fpage>
          .
          <year>v1</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          <string-name>
            <surname>Palaniappan</surname>
            <given-names>R</given-names>
          </string-name>
          , Raveendran P.
          <article-title>Cognitive task prediction using parametric spectral analysis of EEG signals</article-title>
          .
          <source>Malaysian Journal of Computer Science</source>
          .
          <year>2001</year>
          ;
          <volume>14</volume>
          (
          <issue>1</issue>
          ):
          <fpage>58</fpage>
          -
          <lpage>67</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          <string-name>
            <surname>Papakostas</surname>
            <given-names>M</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Tsiakas</surname>
            <given-names>K</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Giannakopoulos</surname>
            <given-names>T</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Makedon</surname>
            <given-names>F</given-names>
          </string-name>
          .
          <article-title>Towards predicting task performance from EEG signals</article-title>
          .
          <source>In: 2017 IEEE International Conference on Big Data (Big Data)</source>
          ;
          <year>2017</year>
          . p.
          <fpage>4423</fpage>
          -
          <lpage>4425</lpage>
          . doi:
          <volume>10</volume>
          .1109/BigData.
          <year>2017</year>
          .
          <volume>8258478</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          <string-name>
            <surname>Saavedra</surname>
            <given-names>C</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Salas</surname>
            <given-names>R</given-names>
          </string-name>
          , Bougrain L.
          <article-title>Wavelet-based semblance methods to enhance single-trial ERP detection</article-title>
          . To be published in
          <source>Computational Intelligence and Neuroscience</source>
          .
          <year>2019</year>
          ; .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          <string-name>
            <surname>Saavedra</surname>
            <given-names>C</given-names>
          </string-name>
          , Bougrain L.
          <article-title>Processing stages of visual stimuli and event-related potentials</article-title>
          . In: The NeuroComp/KEOpS'12 workshop;
          <year>2012</year>
          . .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          <string-name>
            <surname>Simonyan</surname>
            <given-names>K</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zisserman</surname>
            <given-names>A</given-names>
          </string-name>
          .
          <article-title>Very deep convolutional networks for large-scale image recognition</article-title>
          .
          <source>arXiv preprint arXiv:14091556</source>
          .
          <year>2014</year>
          ; .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          <string-name>
            <surname>Tallon-Baudry</surname>
            <given-names>C</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bertrand</surname>
            <given-names>O</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Delpuech</surname>
            <given-names>C</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Pernier</surname>
            <given-names>J</given-names>
          </string-name>
          .
          <article-title>Stimulus specificity of phase-locked and non-phase-locked 40 Hz visual responses in human</article-title>
          .
          <source>Journal of Neuroscience</source>
          .
          <year>1996</year>
          ;
          <volume>16</volume>
          (
          <issue>13</issue>
          ):
          <fpage>4240</fpage>
          -
          <lpage>4249</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          <string-name>
            <surname>Voicikas</surname>
            <given-names>A</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Niciute</surname>
            <given-names>I</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ruksenas</surname>
            <given-names>O</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Griskova-Bulanova</surname>
            <given-names>I</given-names>
          </string-name>
          .
          <article-title>Effect of attention on 40Hz auditory steady-state response depends on the stimulation type: Flutter amplitude modulated tones versus clicks</article-title>
          .
          <source>Neuroscience Letters</source>
          .
          <year>2016</year>
          ;
          <volume>629</volume>
          :
          <fpage>215</fpage>
          -
          <lpage>220</lpage>
          . doi:
          <volume>10</volume>
          .1016/j.neulet.
          <year>2016</year>
          .
          <volume>07</volume>
          .019.
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>