<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <journal-title-group>
        <journal-title>New York City, USA, July</journal-title>
      </journal-title-group>
    </journal-meta>
    <article-meta>
      <title-group>
        <article-title>WISE: Web-based Interactive Speech Emotion Classification</article-title>
      </title-group>
      <contrib-group>
        <aff id="aff0">
          <label>0</label>
          <institution>Dept. of Clinical and Social Sciences in Psychology University of Rochester</institution>
          ,
          <addr-line>Rochester, NY</addr-line>
          ,
          <country country="US">USA</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Sefik Emre Eskimez</institution>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2016</year>
      </pub-date>
      <volume>10</volume>
      <issue>2016</issue>
      <fpage>2</fpage>
      <lpage>7</lpage>
      <abstract>
        <p>The ability to classify emotions from speech is beneficial in a number of domains, including the study of human relationships. However, manual classification of emotions from speech is time consuming. Current technology supports the automatic classification of emotions from speech, but these systems have some limitations. In particular, existing systems are trained with a given data set and cannot adapt to new data nor can they adapt to different users' notions of emotions. In this study, we introduce WISE, a web-based interactive speech emotion classification system. WISE has a web-based interface that allows users to upload speech data and automatically classify the emotions within this speech using pre-trained models. The user can then adjust the emotion label if the system classification of the emotion does not agree with the user's perception, and this updated label is then fed back into the system to retrain the models. In this way, WISE enables the emotion classification models to be adapted over time. We evaluate WISE by simulating the user interactions with the system using the LDC dataset, which has known, ground-truth labels. We evaluate the benefit of the user feedback enabled by WISE in situations where manually classifying emotions in a large dataset is costly, yet trained models alone will not be able to accurately classify the data.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>Accurately estimating emotions of conversational partners
plays a vital role in successful human communication. A
social-functional approach to human emotion emphasizes the
interpersonal function of emotion for the establishment and
maintenance of social relationships [Campos et al., 1989],
[Ekman, 1992], [Keltner and Kring, 1998]. According to
[Campos et al., 1989] “Emotions are not mere feelings, but
rather are processes of establishing, maintaining, or
disrupting relations between the person and the internal or external
environment, when such relations are significant to the
individual.” Thus, the expression and recognition of emotions
allows the facilitation of social bonds through the conveyance
of information about one’s internal state, disposition,
intentions, and needs.</p>
      <p>In many situations, audio is the only recorded data for a
social interaction, and estimating emotions from speech
becomes a critical task for psychological analysis. Today’s
technology allows for gathering vast amounts of emotional speech
data from the web, yet analyzing this content is impractical.
This fact prevents many interesting large-scale investigations.</p>
      <p>Given the amount of speech data that proliferates, there
have been many attempts to create automatic emotion
classification systems. However, the performance of these systems
is not as high as necessary in many situations. Many
potential applications would benefit from automated emotion
classification systems, such as call-center monitoring [Petrushin,
1999; Gupta, 2007], service robot interactions [Park et al.,
2009; Liu et al., 2013] and driver assistance systems [Jones
and Jonsson, 2005; Tawari and Trivedi, 2010]. Indeed, there
are many automated systems today that focus on speech
[Sethu et al., 2008; Busso et al., 2009; Rachuri et al., 2010;
Bitouk et al., 2010; Stuhlsatz et al., 2011; Yang, 2015].
However, emotion classification accuracy of fully automated
systems is still not satisfactory in many practical situations.</p>
      <p>In this study, we propose WISE, a web-based
interactive speech emotion classification system. This system uses
a web-based interface that allows users to easily upload a
speech file to the server for emotion analysis, without the
need for installing any additional software. Once the speech
files are uploaded, the system classifies the emotions using a
model trained on previously labeled training samples. Each
classification is also associated with a confidence value. The
user can either accept or correct the classification, to “teach”
the system the user’s specific concept of emotions. Over
time, the system adapts its emotion classification models to
the user’s concept, and can increase its classification
accuracy with respect to the user’s concept of emotions.</p>
      <p>The key contribution of our work is that we provide an
interactive speech-based emotion analysis framework. This
framework combines the machine’s computational power
with human users’ high emotion classification accuracy.
Compared to purely manual labeling, it is much more
efficient. Compared to fully automated systems, it is much more
accurate. This opens up possibilities for large-scale speech
emotion analysis with high accuracy.</p>
      <p>The proposed framework only considers offline labeling
and returns labels in three categories: emotion, arousal and
valance with time codes. To evaluate our system, we have
simulated the user-interface interactions in several settings,
by providing ground truth labels on behalf of the user. One
of the scenarios is designed to be a baseline, with which we
can compare the remaining scenarios. In another scenario,
we test if the system can adapt to the samples whose speaker
is unknown to the system. The next scenario tests how the
system’s classification confidence of a sample effects the
system’s accuracy. The full system is available for researchers to
use. 1</p>
      <p>The rest of the paper is organized as follows. Section 2
contains a review of the related work. Section 3 describes
the WISE web user-interface, while Section 4 explains the
automated speech-based emotion recognition system used in
this work. We evaluate the WISE system in Section 5, and
conclude our work in Section 6.
2</p>
    </sec>
    <sec id="sec-2">
      <title>Related Work</title>
      <p>All-in-one frameworks for automatic emotion classification
from speech, such as EmoVoice [Vogt et al., 2008] and
OpenEar [Eyben et al., 2009], are standalone software packages
with various capabilities, including audio recording, audio
file reading, feature extraction, and emotion classification.</p>
      <p>EmoVoice allows the user to create a personal
speechbased emotion recognizer, and it can track the emotional state
of the user in real-time. Each user records their own speech
emotion corpus to train the system, and the system can then
be used for real-time emotion classification for the same user.
The system outputs the x- and y-coordinates of an
arousalvalance coordinate system with time codes. It is reported in
[Vogt et al., 2008] that EmoVoice has been used in several
systems including humanoid robot-human and virtual
agenthuman interactions. EmoVoice does not consider user
feedback once the classifier is trained, whereas in our system, the
user can continually train and improve the system.</p>
      <p>OpenEar is an emotion classification multi-platform
software package that includes libraries for feature extraction
written in C++ and pre-trained models as well as scripts to
support model building. One of its main modules is named
SMILE (Speech and Music Interpretation by Large-Space
Extraction), and it can extract more than 500K features in
real-time. The other main module allows external classifiers
and libraries such as LibSVM [Chang and Lin, 2011] to be
integrated and used in classification. OpenEar also supports
popular machine learning frameworks’ data formats, such
as the Hidden Markov Model Toolkit (HTK) [Young et al.,
2006], WEKA [Hall et al., 2009], and scikit-learn for Python
[Pedregosa et al., 2011], and therefore allows easy
transition between frameworks. OpenEar’s capability of batch
processing, combined with its advantage in transitioning to other
learning frameworks, makes it appealing for large databases.</p>
      <p>ANNEMO (ANNotating EMOtions) [Ringeval et al.,
2013] is a web-based annotation tool that allows labeling
arousal, valence and social dimensions in audio-visual data.
The states are represented between -1 and 1, where the user
changes the values using a slider. The social dimension is
1http://www.ece.rochester.edu/projects/wcng
represented by categories rather than numerical values, and
those are agreement, dominance, engagement, performance
and rapport. No automatic classification/labeling modules are
included in ANNEMO.</p>
      <p>In contrast, WISE is a web-based system and can be used
easily without installing any software, unlike EmoVoice and
OpenEar. WISE is similar to ANNEMO in terms of the
webbased labeling aspect, however WISE only considers audio
data and provides automatic classification as well.
3</p>
    </sec>
    <sec id="sec-3">
      <title>Web-based Interaction</title>
      <p>Our system’s interface, shown in Figure 2, is web-based,
allowing easy, secure access and use without installing any
other software except a modern browser.</p>
      <p>When a user uploads an audio file, the waveform appears
on the main screen, allowing the user to select different parts
of the waveform. Selected parts can be played and labeled
independently. These selected parts will also be added to a
list, as shown in the bottom-left side of Figure 2. The user
can download this list by clicking on the “save” button in the
interface.</p>
      <p>The labeling scheme is restricted to three categories:
emotion, arousal and valence. Emotion category elements are
anger, disgust, fear, happy, neutral, sadness. Arousal category
elements are active, passive and neutral, and valance category
elements are positive, negative and neutral. Our future work
includes adding user defined emotion labels into the system.</p>
      <p>The user can request labels from the automated emotion
classifier by clicking on the “request label” button. The
system then shows suggested labels to the user.
There are various automated speech-based emotion
classification systems [Sethu et al., 2008; Busso et al., 2009;
Rachuri et al., 2010; Bitouk et al., 2010; Stuhlsatz et al.,
2011] that consider different features, feature selection
methods, classifiers and decision mechanisms. Our system is
based on [Yang, 2015], which provides a confidence value
along with the classification label.
4.1</p>
      <sec id="sec-3-1">
        <title>Features</title>
        <p>Speech samples are divided into overlapping frames for
feature extraction. The window and hop sizes are set to 60 ms
and 10 ms, respectively. For every frame that contains speech,
the following features are calculated: fundamental frequency
(F0), 12 mel-frequency cepstral coefficients (MFCCs),
energy, frequency and bandwidth of first four formants,
zerocrossing rate, spectral roll-off, brightness, centroid, spread,
skewness, kurtosis, flatness, entropy, roughness, and
irregularity, in addition to the derivatives of these features.
Statistical values such as minimum, maximum, mean, standard
deviation and range (i.e., max-min) are calculated from all frames
within the sample. Additionally, speaking rate is calculated
over the entire sample. Hence, the final feature vector length
is 331.</p>
      </sec>
      <sec id="sec-3-2">
        <title>4.2 Feature Selection</title>
        <p>The system employs the support vector machine (SVM)
recursive feature elimination method [Guyon et al., 2002]. This
approach takes advantage of SVM weights to detect which
features are better than others. After the SVM is trained, the
features are ranked according to the order of their weights.</p>
        <p>The last ranked feature is eliminated from the list and the
process starts again, until there are no features left. Features are
ranked in reverse order of elimination order. The top 80 best
features are chosen to be used in the classification system.
Note that in Section 5.2, the features are selected beforehand
and not updated when a new sample is added to the system.
4.3</p>
      </sec>
      <sec id="sec-3-3">
        <title>Classifier</title>
        <p>Our system uses a one-against-all (OAA) binary SVM with
radial basis function (RBF) for each emotion, arousal and
valance category element, for a total of 12 SVMs. The trained
SVMs calculate confidence scores for any sample that is
being classified. The system labels the sample with the class of
the binary classifier with maximum classification confidence
on the considered sample.</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>Evaluation</title>
      <p>To evaluate WISE and the benefit of user-assisted labeling of
the data, we have simulated user-interface interactions using
the LDC database as the source of data for training, validation
and testing.
5.1</p>
      <sec id="sec-4-1">
        <title>Dataset</title>
        <p>We use the Linguistic Data Consortium (LDC) Emotional
Prosody Speech and Transcripts [Liberman et al., 2002]
database in our simulations. The LDC database contains
samples from 15 emotion categories; however, in our evaluation,
we only use 6 of the emotions as listed in Section 3. The LDC
database contains acted speech, voiced by 7 professionals, 4
female and 3 male. The transcripts are in English and contain
semantically neutral utterances, such as dates and times.
5.2</p>
      </sec>
      <sec id="sec-4-2">
        <title>Simulations</title>
        <p>We have simulated user-interface interactions in different
scenarios for which WISE can be used to enable user feedback
to improve classification accuracy. In these simulations, there
are three data groups: training, test and validation. We
assume that validation data represents the samples where the
user provides the “correct” label. In each iteration, the
system evaluates the test data using the current models, and at
the end of each iteration, a sample from the validation data
is added to the training data to update the models. Next, we
describe the different scenarios in detail.</p>
      </sec>
      <sec id="sec-4-3">
        <title>Scenario 0 - Baseline</title>
        <p>In this scenario, the data from 1 of the 7 speakers is used
for testing, while the remaining 6 speakers’ data are used for
training and validation. Since only a limited amount of data
is available from each speaker in the next scenarios, we also
limit the amount of the validation data in this scenario. In
this way, the baseline becomes more comparable to the other
scenarios.</p>
        <p>The training data starts with N samples from each class for
each category. For the emotion classification, there are only
2 samples available in each class (emotion) for the validation
data. However, the arousal and valance categories have half
the number of classes that the emotion category has,
therefore, there are 3 samples available in each class that can be
used in validation data for these categories. After the data are
chosen randomly, the system simulates the interaction
process. This process is repeated for all speakers, and the results
are averaged over all 7 speakers and 200 trials.</p>
      </sec>
      <sec id="sec-4-4">
        <title>Scenario I</title>
        <p>This scenario has the same settings as Scenario 0, except this
time, the testing data, as well as the validation data are chosen
from a speaker, and the training data is chosen among the
remaining 6 speakers’ data.</p>
      </sec>
      <sec id="sec-4-5">
        <title>Scenario II</title>
        <p>This scenario has the same settings as Scenario I with a
single difference: in each round, the validation data has been
ordered in ascending order of the classifier’s confidence level
in classifying them. Therefore in each iteration, the sample,
on which the system has the least confidence, is added to the
training data from the validation data.</p>
      </sec>
      <sec id="sec-4-6">
        <title>Discussion</title>
        <p>Figures 3-5 show the classification accuracy versus the
number of added samples for each scenario for the emotion,
arousal and valence, respectively. Note that the error bars
represent the standard deviation of the results over the 7 speakers
and 200 trials.</p>
        <p>Scenario I shows the ability of WISE to enable adaptation
of the models. In many situations, trained models of
automatic systems have no information on the speaker to be
classified. The comparison of classification accuracy between
Scenario 0 and Scenario I shows that adaptation to unknown
data is vital for accurate emotion estimation, as the accuracy
increases greatly when data from the new user are added.</p>
        <p>For example, in Scenarios I and II, when N is 4 for the
emotion category, the system’s initial accuracy starts around
37% and increases to approximately 63%, as can be seen in
Figure 3, where on the other hand in Scenario 0, accuracy
can only increase to approximately 41%. In Scenarios I and
II , when N is 10, the classification accuracy starts higher
then the previous case, yet with the same number of added
samples, they converge to the same percentage. This enables
the possibility of using pre-trained models in our system that
are trained on available databases.</p>
        <p>The results of Scenario II suggest that adding the samples
with low classification confidence are slightly more beneficial
than adding a sample for which the system already has more
confidence. Figures 3-5 show that the classifier in Scenario II
converges to a slightly higher classification accuracy than the
one in Scenario I. This can be seen especially in the arousal
category results.
6</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>Conclusion</title>
      <p>This study introduced and evaluated the WISE system, which
is an interactive web-based emotion analysis framework to
assist in the classification of human emotion from voice data.</p>
      <p>The full system is available for the community to use. The
evaluation results show that the system can adapt to the user’s
choices and can increase the future classification accuracy
when the speaker of the sample is unknown. Hence, WISE
will enable adaptive, large scale emotion classification.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [Bitouk et al.,
          <year>2010</year>
          ]
          <string-name>
            <given-names>Dmitri</given-names>
            <surname>Bitouk</surname>
          </string-name>
          , Ragini Verma, and
          <string-name>
            <given-names>Ani</given-names>
            <surname>Nenkova</surname>
          </string-name>
          .
          <article-title>Class-level spectral features for emotion recognition</article-title>
          .
          <source>Speech Commun</source>
          .,
          <volume>52</volume>
          (
          <issue>7</issue>
          ):
          <fpage>613</fpage>
          -
          <lpage>625</lpage>
          ,
          <year>2010</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [Busso et al.,
          <year>2009</year>
          ]
          <string-name>
            <given-names>C.</given-names>
            <surname>Busso</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Lee</surname>
          </string-name>
          , and
          <string-name>
            <given-names>S.</given-names>
            <surname>Narayanan</surname>
          </string-name>
          .
          <article-title>Analysis of emotionally salient aspects of fundamental frequency for emotion detection</article-title>
          .
          <source>IEEE Transactions on Audio, Speech, and Language Processing</source>
          ,
          <volume>17</volume>
          (
          <issue>4</issue>
          ):
          <fpage>582</fpage>
          -
          <lpage>596</lpage>
          , May
          <year>2009</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [Campos et al.,
          <year>1989</year>
          ] Joseph J Campos, Rosemary G Campos,
          <article-title>and Karen C Barrett. Emergent themes in the study of emotional development and emotion regulation</article-title>
          .
          <source>Dev Psychol</source>
          .,
          <volume>25</volume>
          (
          <issue>3</issue>
          ):
          <fpage>394</fpage>
          ,
          <year>1989</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          <source>[Chang and Lin</source>
          , 2011]
          <article-title>Chih-Chung Chang and Chih-Jen Lin</article-title>
          .
          <article-title>LIBSVM: A library for support vector machines</article-title>
          .
          <source>ACM Transactions on Intelligent Systems and Technology</source>
          ,
          <volume>2</volume>
          :
          <issue>27</issue>
          :
          <fpage>1</fpage>
          -
          <lpage>27</lpage>
          :
          <fpage>27</fpage>
          ,
          <year>2011</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          <source>[Ekman</source>
          , 1992]
          <string-name>
            <given-names>Paul</given-names>
            <surname>Ekman</surname>
          </string-name>
          .
          <article-title>An argument for basic emotions</article-title>
          .
          <source>Cognition Emotion</source>
          ,
          <volume>6</volume>
          (
          <issue>3</issue>
          -4):
          <fpage>169</fpage>
          -
          <lpage>200</lpage>
          ,
          <year>1992</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [Eyben et al.,
          <year>2009</year>
          ]
          <string-name>
            <given-names>Florian</given-names>
            <surname>Eyben</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Martin</given-names>
            <surname>Wllmer</surname>
          </string-name>
          , and Bjrn Schuller. openear
          <article-title>- introducing the munich open-source emotion and affect recognition toolkit</article-title>
          . In In ACII, pages
          <fpage>576</fpage>
          -
          <lpage>581</lpage>
          ,
          <year>2009</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          <source>[Gupta</source>
          , 2007]
          <string-name>
            <given-names>Purnima</given-names>
            <surname>Gupta</surname>
          </string-name>
          .
          <article-title>Two-Stream Emotion Recognition For Call Center Monitoring</article-title>
          .
          <source>In Interspeech</source>
          <year>2007</year>
          ,
          <year>2007</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [Guyon et al.,
          <year>2002</year>
          ]
          <string-name>
            <given-names>Isabelle</given-names>
            <surname>Guyon</surname>
          </string-name>
          , Jason Weston, Stephen Barnhill, and
          <string-name>
            <given-names>Vladimir</given-names>
            <surname>Vapnik</surname>
          </string-name>
          .
          <article-title>Gene selection for cancer classification using support vector machines</article-title>
          .
          <source>Mach</source>
          . Learn.,
          <volume>46</volume>
          (
          <issue>1-3</issue>
          ):
          <fpage>389</fpage>
          -
          <lpage>422</lpage>
          ,
          <year>March 2002</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [Hall et al.,
          <year>2009</year>
          ]
          <string-name>
            <given-names>Mark</given-names>
            <surname>Hall</surname>
          </string-name>
          , Eibe Frank, Geoffrey Holmes, Bernhard Pfahringer,
          <string-name>
            <given-names>Peter</given-names>
            <surname>Reutemann</surname>
          </string-name>
          , and
          <string-name>
            <surname>Ian</surname>
            <given-names>H.</given-names>
          </string-name>
          <string-name>
            <surname>Witten</surname>
          </string-name>
          .
          <article-title>The weka data mining software: An update</article-title>
          .
          <source>SIGKDD Explor</source>
          . Newsl.,
          <volume>11</volume>
          (
          <issue>1</issue>
          ):
          <fpage>10</fpage>
          -
          <lpage>18</lpage>
          ,
          <year>November 2009</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          <source>[Jones and Jonsson</source>
          , 2005]
          <article-title>Christian Martyn Jones</article-title>
          and
          <string-name>
            <given-names>IngMarie</given-names>
            <surname>Jonsson</surname>
          </string-name>
          .
          <article-title>Automatic recognition of affective cues in the speech of car drivers to allow appropriate responses</article-title>
          .
          <source>In Proceedings of the 17th Australia Conference on Computer-Human Interaction:</source>
          Citizens Online:
          <article-title>Considerations for Today and the Future</article-title>
          ,
          <source>OZCHI '05</source>
          , pages
          <fpage>1</fpage>
          -
          <lpage>10</lpage>
          , Narrabundah, Australia, Australia,
          <year>2005</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          <source>[Keltner and Kring</source>
          , 1998]
          <string-name>
            <given-names>Dacher</given-names>
            <surname>Keltner and Ann M Kring. Emotion</surname>
          </string-name>
          ,
          <article-title>social function, and psychopathology</article-title>
          .
          <source>Rev. Gen. Psychol.</source>
          ,
          <volume>2</volume>
          (
          <issue>3</issue>
          ):
          <fpage>320</fpage>
          ,
          <year>1998</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [Liberman et al.,
          <year>2002</year>
          ]
          <string-name>
            <given-names>Mark</given-names>
            <surname>Liberman</surname>
          </string-name>
          , Kelly Davis,
          <string-name>
            <given-names>M</given-names>
            <surname>Grossman</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N</given-names>
            <surname>Martey</surname>
          </string-name>
          , and
          <string-name>
            <given-names>J</given-names>
            <surname>Bell</surname>
          </string-name>
          .
          <article-title>Emotional prosody speech and transcripts</article-title>
          .
          <source>In Proc. LDC</source>
          ,
          <year>2002</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [Liu et al.,
          <year>2013</year>
          ]
          <string-name>
            <surname>Chih-Yin</surname>
            <given-names>Liu</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Tzu-Hsin</surname>
            <given-names>Hung</given-names>
          </string-name>
          , Kai-Chung Cheng, and
          <string-name>
            <surname>Tzuu-Hseng S Li</surname>
          </string-name>
          .
          <article-title>Hmm and bpnn based speech recognition system for home service robot</article-title>
          .
          <source>In Advanced Robotics and Intelligent Systems (ARIS)</source>
          , 2013 International Conference on, pages
          <fpage>38</fpage>
          -
          <lpage>43</lpage>
          . IEEE,
          <year>2013</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [Park et al.,
          <year>2009</year>
          ]
          <string-name>
            <given-names>J. S.</given-names>
            <surname>Park</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. H.</given-names>
            <surname>Kim</surname>
          </string-name>
          , and
          <string-name>
            <given-names>Y. H.</given-names>
            <surname>Oh</surname>
          </string-name>
          .
          <article-title>Feature vector classification based speech emotion recognition for service robots</article-title>
          .
          <source>IEEE Transactions on Consumer Electronics</source>
          ,
          <volume>55</volume>
          (
          <issue>3</issue>
          ):
          <fpage>1590</fpage>
          -
          <lpage>1596</lpage>
          ,
          <year>August 2009</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [Pedregosa et al.,
          <year>2011</year>
          ]
          <string-name>
            <given-names>F.</given-names>
            <surname>Pedregosa</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Varoquaux</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Gramfort</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Michel</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Thirion</surname>
          </string-name>
          ,
          <string-name>
            <given-names>O.</given-names>
            <surname>Grisel</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Blondel</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Prettenhofer</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Weiss</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Dubourg</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Vanderplas</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Passos</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Cournapeau</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Brucher</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Perrot</surname>
          </string-name>
          , and
          <string-name>
            <given-names>E.</given-names>
            <surname>Duchesnay</surname>
          </string-name>
          .
          <article-title>Scikit-learn: Machine learning in Python</article-title>
          .
          <source>Journal of Machine Learning Research</source>
          ,
          <volume>12</volume>
          :
          <fpage>2825</fpage>
          -
          <lpage>2830</lpage>
          ,
          <year>2011</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          <source>[Petrushin</source>
          , 1999]
          <article-title>Valery A. Petrushin. Emotion in speech: Recognition and application to call centers</article-title>
          . In In Engr, pages
          <fpage>7</fpage>
          -
          <lpage>10</lpage>
          ,
          <year>1999</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [Rachuri et al.,
          <year>2010</year>
          ]
          <article-title>Kiran K Rachuri, Mirco Musolesi</article-title>
          , Cecilia Mascolo,
          <string-name>
            <surname>Peter J Rentfrow</surname>
            , Chris Longworth, and
            <given-names>Andrius</given-names>
          </string-name>
          <string-name>
            <surname>Aucinas</surname>
          </string-name>
          .
          <article-title>Emotionsense: a mobile phones based adaptive platform for experimental social psychology research</article-title>
          .
          <source>In Proc. 12th ACM Int. Conf. on Ubiquitous Computing</source>
          , pages
          <fpage>281</fpage>
          -
          <lpage>290</lpage>
          ,
          <year>2010</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          [Ringeval et al.,
          <year>2013</year>
          ]
          <string-name>
            <given-names>F.</given-names>
            <surname>Ringeval</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Sonderegger</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Sauer</surname>
          </string-name>
          , and
          <string-name>
            <given-names>D.</given-names>
            <surname>Lalanne</surname>
          </string-name>
          .
          <article-title>Introducing the recola multimodal corpus of remote collaborative and affective interactions</article-title>
          .
          <source>In 2013 10th IEEE International Conference and Workshops on Automatic Face and Gesture Recognition (FG)</source>
          , pages
          <fpage>1</fpage>
          -
          <lpage>8</lpage>
          ,
          <year>April 2013</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          [Sethu et al.,
          <year>2008</year>
          ]
          <string-name>
            <given-names>Vidhyasaharan</given-names>
            <surname>Sethu</surname>
          </string-name>
          , Eliathamby Ambikairajah, and
          <string-name>
            <given-names>Julien</given-names>
            <surname>Epps</surname>
          </string-name>
          .
          <article-title>Empirical mode decomposition based weighted frequency feature for speech-based emotion classification</article-title>
          .
          <source>In Proc. IEEE ICASSP</source>
          , pages
          <fpage>5017</fpage>
          -
          <lpage>5020</lpage>
          ,
          <year>2008</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          [Stuhlsatz et al.,
          <year>2011</year>
          ]
          <string-name>
            <given-names>A.</given-names>
            <surname>Stuhlsatz</surname>
          </string-name>
          , C. Meyer,
          <string-name>
            <given-names>F.</given-names>
            <surname>Eyben</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Zielke</surname>
          </string-name>
          , G. Meier, and
          <string-name>
            <given-names>B.</given-names>
            <surname>Schuller</surname>
          </string-name>
          .
          <article-title>Deep neural networks for acoustic emotion recognition: Raising the benchmarks</article-title>
          .
          <source>In Acoustics, Speech and Signal Processing (ICASSP)</source>
          ,
          <year>2011</year>
          IEEE International Conference on, pages
          <fpage>5688</fpage>
          -
          <lpage>5691</lpage>
          , May
          <year>2011</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          <source>[Tawari and Trivedi</source>
          , 2010]
          <string-name>
            <given-names>A.</given-names>
            <surname>Tawari</surname>
          </string-name>
          and
          <string-name>
            <given-names>M.</given-names>
            <surname>Trivedi</surname>
          </string-name>
          .
          <article-title>Speech based emotion classification framework for driver assistance system</article-title>
          .
          <source>In Intelligent Vehicles Symposium (IV)</source>
          ,
          <year>2010</year>
          IEEE, pages
          <fpage>174</fpage>
          -
          <lpage>178</lpage>
          ,
          <year>June 2010</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          [Vogt et al.,
          <year>2008</year>
          ]
          <string-name>
            <given-names>Thurid</given-names>
            <surname>Vogt</surname>
          </string-name>
          , Elisabeth Andr, and
          <string-name>
            <given-names>Nikolaus</given-names>
            <surname>Bee</surname>
          </string-name>
          .
          <article-title>Emovoice - a framework for online recognition of emotions from voice</article-title>
          .
          <source>In In Proceedings of Workshop on Perception and Interactive Technologies for Speech-Based Systems</source>
          , Springer, Kloster Irsee,
          <year>2008</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          <source>[Yang</source>
          , 2015]
          <string-name>
            <given-names>Na</given-names>
            <surname>Yang</surname>
          </string-name>
          .
          <article-title>Algorithms for affective and ubiquitous sensing systems and for protein structure prediction</article-title>
          .
          <source>PhD thesis</source>
          , University of Rochester,
          <year>2015</year>
          . http://hdl.handle.net/
          <year>1802</year>
          /29666.
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          [Young et al.,
          <year>2006</year>
          ]
          <string-name>
            <given-names>S. J.</given-names>
            <surname>Young</surname>
          </string-name>
          , G. Evermann,
          <string-name>
            <given-names>M. J. F.</given-names>
            <surname>Gales</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Hain</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Kershaw</surname>
          </string-name>
          , G. Moore,
          <string-name>
            <given-names>J.</given-names>
            <surname>Odell</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Ollason</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Povey</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Valtchev</surname>
          </string-name>
          , and
          <string-name>
            <given-names>P. C.</given-names>
            <surname>Woodland</surname>
          </string-name>
          .
          <source>The HTK Book, version 3</source>
          .4. Cambridge University Engineering Department, Cambridge, UK,
          <year>2006</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>