<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>A Preliminary Assessment of Game Event Detection in Emotional Mario Task at MediaEval 2021</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Van-Tu Ninh</string-name>
          <email>tu.ninhvan@adaptcentre.ie</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Tu-Khiem Le</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Manh-Duy Nguyen</string-name>
          <email>manh.nguyen5@mail.dcu.ie</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Sinéad Smyth</string-name>
          <email>sinead.smyth@dcu.ie</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Graham Healy</string-name>
          <email>healy@dcu.ie</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Cathal Gurrin</string-name>
          <email>cathal.gurrin@dcu.ie</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>School of Computing, Dublin City University</institution>
          ,
          <country country="IE">Ireland</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>School of Psychology, Dublin City University</institution>
          ,
          <country country="IE">Ireland</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2021</year>
      </pub-date>
      <fpage>13</fpage>
      <lpage>15</lpage>
      <abstract>
        <p>The Emotional Mario task at MediaEval 2021 presents a new challenge of analysing the gameplay of ten participants on the wellknown Super Mario Bros video game by detecting key events using facial and biometrics data. Our purpose in this work is to evaluate the application of emotion-related features in other domains of afective computing in game event detection. In this working notes paper, we present our work on in-game event detection using the conventional Random Forest model with a combination of Blood Volume Pulse and Electrodermal Activity statistical features with the facial expressions of the player as the input. In addition, we also investigate the evaluation of using the in-game visual features in another pipeline with the same Random Forest model to compare the eficiency of using in-game visual features in the model. The source code of our work can be found at https://github.com/nvtu/Emotional-Mario-Analysis.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>INTRODUCTION</title>
      <p>
        Being referred to as engines of experience, games act as a source of
external stimuli that can trigger responses in human emotion (e.g.,
a person might feel intense stress when fighting against a boss in
a game). However, the connection between games and human’s
emotions has not been comprehensively studied, which presents
an open area of research. Therefore, the Emotional Mario Task was
initiated to analyse this relationship [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ]. The task employed 10
volunteers to play various stages in the Super Mario Bros video game
and capture their reactions using a webcam and an E4 wristband.
The ultimate goal is to (1) predict five key events in the game, and
(2) summarise the gameplay by aggregating the best moments in
the game. In this work, we focus mainly on the first task. Our aim
is to analyse the contribution of facial expressions and
physiological signals recorded from wearable devices to the detection and
classification of five key events in the game.
2.1
      </p>
    </sec>
    <sec id="sec-2">
      <title>APPROACH</title>
    </sec>
    <sec id="sec-3">
      <title>Data Processing and Feature Extraction</title>
      <p>
        2.1.1 Face, Game frame, and Sensor Synchronization and
Processing: There are three types of data in the dataset captured using
diferent devices with diferent sampling rates, which are: face video,
in-game video and sensor data. Apart from the data-synchronisation
codes given by the task organisers, we also modify the source code
in the Github repository provided in [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ] to extract all relevant
frames corresponding to the actions in the game. For sensor data
recorded from Empatica E4 device, the Blood Volume Pulse (BVP)
and Accelerometer are pruned to 60 Hz from the original sampling
rate of 64 Hz and 32 Hz respectively, to match the sampling rate
of the video. For facial data, the Face Emotion Recognition (FER)
features provided by the task organisers [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ] extracted using the FER
package [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ] are inputted as a 7-dimensional vector into the model
for training. Even though the use of in-game video is not
recommended in this task, we also extract game-frame deep features from
a ResNet-50 model pre-trained on the ImageNet dataset. These deep
features are the same as the ones used in the preliminary work on
the same dataset in [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ], which is a 2048-dimensional vector.
      </p>
      <p>
        2.1.2 Blood Volume Pulse (BVP). For Blood Volume Pulse (BVP)
feature extraction, we extract statistical features commonly used
for stress detection and emotion recognition using physiological
signals. We use the Neurokit21 library, which employs the Elgandi
processing pipleline to clean the photoplethysmogram (PPG) signal
[
        <xref ref-type="bibr" rid="ref6">6</xref>
        ] and detect systolic peaks [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]. We then compute heart rate (HR),
time-domain and frequency-domain of heart rate variability (HRV)
using the extracted systolic peaks with a window-size of 60 seconds.
For frequency-domain HRV features, the same parameters of low
(LF: 0.04-0.15 Hz) and high (HF: 0.15-0.4 Hz) frequency bands as in
[
        <xref ref-type="bibr" rid="ref9">9</xref>
        ] are used. Finally, the feature vector is standardised. This feature
extraction process results in a 27-dimensional vector.
      </p>
      <p>
        2.1.3 Electrodermal Activity (EDA):. We followed previous
research [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ] in stress detection analysis to extract statistical EDA
features. Using Neurokit2 library, we extract components of the
EDA signal that comprise Skin Conductance Response (SCR), Skin
Conductance Level (SCL), SCR Peaks, SCR Onsets, and SCR
Amplitude. Then, the statistical EDA features from the combination
of four works [
        <xref ref-type="bibr" rid="ref1 ref4 ref8 ref9">1, 4, 8, 9</xref>
        ] are computed except for the slope of EDA
signal along the time-axis, which results in a 35-dimensional vector.
Finally, the feature vector is standardised.
2.2
      </p>
    </sec>
    <sec id="sec-4">
      <title>Game Event Detection Models</title>
      <p>In total, we develop two models whose names are A and B,
respectively. Model A, which detects game events based on diferent
combinations of emotion-related features, comprises of two stages.</p>
      <p>As illustrated in Figure 1 (black arrow), the first stage of the
model aims at detecting if a game event happens at a timestamp
1https://github.com/neuropsychology/NeuroKit
while the second stage concentrates on classifying the
corresponding game event (flag reached, life lost, status up, status down, new
stage). Both stages employ a Random Forest model implemented in
scikit-learn2 and incremental trees3 libraries with the same
configuration of parameters. For the first stage training, as the number of
samples of game-event/no-game-event is imbalanced which afects
the learning process of the model, we shufle the non-game-event
samples, then divide them into batches whose size is equal to the
one of game-event samples, and apply incremental training to the
Random Forest model. The non-default parameter values that we
employ in model A are shown in table 1.</p>
      <p>
        Model B, which classifies game events using deep visual features
extracted from game frames combined with BVP statistical features,
is a simple incremental training Random Forest with the same
parameter values as in Table 1 except for the number of estimators
(100), minimum samples for splitting (default value), and maximum
depth (default value).
The organisers evaluate the runs based on exact event matching
and event time-frame matching in a range of +/- one second and
+/ifve seconds using precision, recall, and f1 score. [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ]. In our paper,
we report the evaluation results of both exact event matching and
time-frame matching in range of +/- five seconds.
3.2
      </p>
    </sec>
    <sec id="sec-5">
      <title>Results</title>
      <p>In total, we submitted three runs to the task. As described in section
2, model A is used with emotion-related features as input, while
model B used additional ResNet-50 visual features of gameplay. In
our prior experiment, we also tried using model B with
emotionrelated features as input to detect the event without success
potentially due in part to the highly imbalanced nature of the dataset.
The results in table 2 and 3 both show that there is a large gap in the
precision of correct event detection between using emotion-related
features extracted from physiological signals and using visual
features from game-frame. This suggests that the game-frames contain
a lot of information about the event compared to non-visual data.
As demonstrated in table 2 and 3, the precision score of the model
A is extremely low, while the recall score is considerably higher
than other attempts in the task, which shows that the number of
false positive predictions is significantly high. This means that a
proper approach of event detection using emotion-related features
has not been constructed successfully yet and further research on
this task needs to be conducted.</p>
    </sec>
    <sec id="sec-6">
      <title>ACKNOWLEDGMENTS</title>
      <p>This publication is funded as part of Dublin City University’s
Research Committee and research grants from Science Foundation
Ireland and co-funded by the European Regional Development
Fund under grant numbers SFI/13/RC/2106, SFI/13/RC/2106_P2,
SFI/12/RC/2289_P2, and 18/CRT/6223.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>Jongyoon</given-names>
            <surname>Choi</surname>
          </string-name>
          , Beena Ahmed, and
          <string-name>
            <surname>Ricardo</surname>
          </string-name>
          Gutierrez-Osuna.
          <year>2011</year>
          .
          <article-title>Development and evaluation of an ambulatory stress monitor based on wearable sensors</article-title>
          .
          <source>IEEE transactions on information technology in biomedicine 16</source>
          ,
          <issue>2</issue>
          (
          <year>2011</year>
          ),
          <fpage>279</fpage>
          -
          <lpage>286</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>Mohamed</given-names>
            <surname>Elgendi</surname>
          </string-name>
          , Ian Norton, Matt Brearley, Derek Abbott, and
          <string-name>
            <given-names>Dale</given-names>
            <surname>Schuurmans</surname>
          </string-name>
          .
          <year>2013</year>
          .
          <article-title>Systolic peak detection in acceleration photoplethysmograms measured from emergency responders in tropical conditions</article-title>
          .
          <source>PLoS One</source>
          <volume>8</volume>
          ,
          <issue>10</issue>
          (
          <year>2013</year>
          ),
          <year>e76585</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>Justin</given-names>
            <surname>Shenk</surname>
          </string-name>
          et al.
          <year>2021</year>
          .
          <article-title>Facial Expression Recognition with a deep neural network as a PyPI package</article-title>
          . (
          <year>2021</year>
          ). https://github.com/justinshenk/ fer
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>Jennifer</given-names>
            <surname>Healey</surname>
          </string-name>
          and
          <string-name>
            <given-names>Rosalind W.</given-names>
            <surname>Picard</surname>
          </string-name>
          .
          <year>2005</year>
          .
          <article-title>Detecting stress during real-world driving tasks using physiological sensors</article-title>
          .
          <source>IEEE Transactions on Intelligent Transportation Systems</source>
          <volume>6</volume>
          (
          <year>2005</year>
          ),
          <fpage>156</fpage>
          -
          <lpage>166</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>Mathias</given-names>
            <surname>Lux</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Riegler</surname>
          </string-name>
          , Henrik Svoren,
          <string-name>
            <given-names>S.</given-names>
            <surname>Hicks</surname>
          </string-name>
          ,
          <string-name>
            <surname>Duc-Tien</surname>
            <given-names>DangNguyen</given-names>
          </string-name>
          , Kristine Jorgensen, Vajira Thambawita, and
          <string-name>
            <given-names>P.</given-names>
            <surname>Halvorsen</surname>
          </string-name>
          .
          <year>2021</year>
          .
          <article-title>Emotional Mario Task at MediaEval 2021</article-title>
          . In MediaEval.
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>Mohsen</given-names>
            <surname>Nabian</surname>
          </string-name>
          , Yu Yin, Jolie Wormwood, Karen S Quigley, Lisa F Barrett,
          <string-name>
            <given-names>and Sarah</given-names>
            <surname>Ostadabbas</surname>
          </string-name>
          .
          <year>2018</year>
          .
          <article-title>An open-source feature extraction tool for the analysis of peripheral physiological data</article-title>
          .
          <source>IEEE journal of translational engineering in health and medicine</source>
          <volume>6</volume>
          (
          <year>2018</year>
          ),
          <fpage>1</fpage>
          -
          <lpage>11</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <surname>Van-Tu</surname>
            <given-names>Ninh</given-names>
          </string-name>
          , Sinéad Smyth,
          <string-name>
            <surname>Minh-Triet Tran</surname>
            , and
            <given-names>Cathal</given-names>
          </string-name>
          <string-name>
            <surname>Gurrin</surname>
          </string-name>
          .
          <year>2021</year>
          .
          <article-title>Analysing the Performance of StressDetection Models on ConsumerGrade Wearable Devices</article-title>
          . In SoMeT.
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>Kizito</given-names>
            <surname>Nkurikiyeyezu</surname>
          </string-name>
          , Anna Yokokubo, and
          <string-name>
            <given-names>Guillaume</given-names>
            <surname>Lopez</surname>
          </string-name>
          .
          <year>2020</year>
          .
          <article-title>Efect of Person-Specific Biometrics in Improving Generic Stress Predictive Models</article-title>
          .
          <source>Sensors and Materials</source>
          <volume>32</volume>
          (02
          <year>2020</year>
          ),
          <fpage>703</fpage>
          -
          <lpage>722</lpage>
          . https://doi.org/10.18494/SAM.
          <year>2020</year>
          .2650
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>Philip</given-names>
            <surname>Schmidt</surname>
          </string-name>
          , Attila Reiss, Robert Duerichen, Claus Marberger, and Kristof Van Laerhoven.
          <year>2018</year>
          .
          <article-title>Introducing wesad, a multimodal dataset for wearable stress and afect detection</article-title>
          .
          <source>In Proceedings of the 20th ACM international conference on multimodal interaction</source>
          .
          <volume>400</volume>
          -
          <fpage>408</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <surname>Henrik</surname>
            <given-names>Svoren</given-names>
          </string-name>
          , Vajira Thambawita,
          <string-name>
            <given-names>P.</given-names>
            <surname>Halvorsen</surname>
          </string-name>
          , Petter Jakobsen, Enrique Alejandro García Ceja, Farzan Majeed Noori, Hugo Lewi Hammer, Mathias Lux,
          <string-name>
            <given-names>M.</given-names>
            <surname>Riegler</surname>
          </string-name>
          , and
          <string-name>
            <given-names>S.</given-names>
            <surname>Hicks</surname>
          </string-name>
          .
          <year>2020</year>
          .
          <article-title>Toadstool: A Dataset for Training Emotional Intelligent Machines Playing Super Mario Bros</article-title>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>