<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>A Methodological Framework for the Exploratory Analysis of Multimodal Features in Learning Activities</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Alejandro Andrade</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Marcelo Worsley</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Indiana University</institution>
          ,
          <addr-line>Blommington, IN 47405</addr-line>
          ,
          <country country="US">USA</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Northwestern University</institution>
          ,
          <addr-line>Evanston, IL 60208</addr-line>
          ,
          <country country="US">USA</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>We make a call for the formalization of a methodological framework that allows researchers to perform reliable and valid multimodal data analysis. We believe a first step is to guide the collection and interpretation of multimodal data. What we call the MMLA Exploratory Framework suggests that data collection needs to take place at least at two time points to account for within-subject variation and learning gains. We briefly describe the data set we are currently using to illustrate the application of this exploratory framework.</p>
      </abstract>
      <kwd-group>
        <kwd>Cognitive Disequilibrium</kwd>
        <kwd>Complex Systems Concepts Learning</kwd>
        <kwd>Methodology</kwd>
        <kwd>Multimodal Learning Analytics</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Introduction</title>
      <p>Collecting multimodal data from student learning activities is becoming ubiquitous in
learning sciences research. Multimodal features—such as body position, facial
expressions, and paralinguistic elements of speech—can be extracted from sensors and
audiovisual recordings. Some of the objectives for capturing these multimodal features are
(a) to better understand the learning process, (b) to develop fine-grained metrics and
assessment of student learning, and (c) to improve learning experiences by providing
feedback and support for pedagogical decision-making. In order to use multimodal
features for these purposes, however, researchers have to carefully find and validate how
such features interact with student learning. Thus, we identify two general challenges
for the collection of reliable and valid multimodal learning features: (a) the collection
(which includes gathering, integration, analysis, and visualization) of multimodal data,
and (b) the validation of multimodal features during learning activities
Our goal in this paper is to make a call for the formalization of a methodological
framework that allows researchers to identify reliable and valid multimodal data sets intended
for any of the learning analytics goals listed above. As is typical in nascent research,
we need to start from an exploratory analysis. Thus, we believe our first step is to
develop what we call an MMLA Exploratory framework. The MMLA exploratory
framework places multimodal features and learning indicators along two dimensions (see
Figure 1). Exploratory MMLA examines the changes in natural groupings within
students’ multimodal features at various stages of the learning activity (dimension 1), and
then correlates these changes with shifts in student understanding levels (dimension 2).
We propose that each participants’ multimodal features and understanding levels have
to be measured at least at two time points to account for the within-subject variation
and learning gains (measured as change in student understanding). In order to see
whether such variation is meaningful in the learning context, a correlation between
variation in multimodal features and variation in student learning gains has to be
stablished. Our working hypothesis is that such correlation represents the interplay between
multimodal features and learning. In finding evidence that such an association exists,
one can then devise ways in which to use MMLA to support learning.</p>
      <p>
        Our current work includes illustrating the MMLA Exploratory Framework with
affect data from cognitive interviews with elementary students. Specifically, we suggest
an exploratory analysis of affective states, through facial expressions and speech
prosodics, to understand how affect interacts with learning gains. We believe that
multimodal features provide a vantage point to uncover students’ affective states,
experienced during a learning activity in interaction with a tool or with another person [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ].
While using these learning tools, students experience some affective responses, and it
has been suggested that these affective responses are correlated with student learning
[
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]. We collect affective indicators from two types of indicators: (a) we capture
students’ facial action units that reveal affective states from facial expressions, and (b)
extract MFCC features that represent students’ speech prosodic elements (e.g.,
hesitation, confidence). We correlate these affective measures with student learning gains.
      </p>
    </sec>
    <sec id="sec-2">
      <title>Description of the Data Set</title>
      <sec id="sec-2-1">
        <title>Learning Content</title>
        <p>
          The learning content is complex systems concepts, dynamic equilibrium in interacting
feedback loops in particular. feedback loops are a key concept to reason about
interactions among organisms in an ecosystem [
          <xref ref-type="bibr" rid="ref3">3</xref>
          ]. Hokayem, Ma and Jin [
          <xref ref-type="bibr" rid="ref3">3</xref>
          ] describe a
learning progression for feedback loops at the elementary school level (see Figure 2). In this
progression, students move from an incipient understanding of one-way simple
causality (see Figure 2.a) to a two-way simple causality (see Figure 2.b) to a two-way cyclical
relationship that demonstrates dynamic change in both populations (see Figure 2.c).
The activity is a cognitive interview, where a set of two questions where designed to
elicit a cognitive disequilibrium. Cognitive disequilibrium is assumed to be observable
in a sudden change of a student’s affective state [
          <xref ref-type="bibr" rid="ref2">2</xref>
          ].
Audiovisual recordings of the set of questions collected, which represent about 1 min
each of the total 20 interview (see Excerpt 1 for an example). Transcripts of the
conversations were created, where a student is responsible of approx. 40% of the uttered
words in average (of about 300 total words per interview section). We believe the
repetition of words given the similar structure of the questions and answers can make these
data suitable for a cleaner analysis of the prosodic element in the students’ speech. In
observing the video data, changes in students’ facial expressions are often observable
right after the student hears the second question.
        </p>
        <p>Excerpt 1
[00:03:10] Interviewer: So, if the number of wolves goes up, what
would happen to the number of sheep?
[00:03:21] Interviewee: (raises red ball) the wolves would go up and
(lowers yellow ball) the sheep would go
down.
[00:03:23] Interviewer: Why?
[00:03:24] Interviewee: Because there would be lots of sheep for them
to eat.
[00:03:46] Interviewer: Great. And if the number of sheep goes up?
[00:03:51] Interviewee: If the number of sheep goes up? Then the
number of wolves would also go up (raises
red ball in line with yellow) because there
would be more sheep for them to eat.
[00:04:10] Interviewer: That was great, let's move on to the next
scenario.
2.5</p>
      </sec>
      <sec id="sec-2-2">
        <title>Measures and Multimodal Features</title>
        <p>
          Student understanding was measured using a coding scheme for the analysis of
students’ verbal explanations of feedback loops [
          <xref ref-type="bibr" rid="ref3">3</xref>
          ]. Facial expressions were extracted as
facial action units from the video. Prosodic elements, were extracted as MFCC features
from the audio file.
3
        </p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>Challenges thus Far</title>
      <p>
        A first challenge has been the extraction of Facial Action Units. We are using the
OpenFace software [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ] to extract facial action units (see Figure 3). The software has some
limitations such as faces have to be minimum 100 pixels long and mostly frontal. There
are also some obstructions to the face when the students move their hands to answer
the questions.
      </p>
      <p>A second challenge is the synchronization of transcript and audio files. We are using
the P2FA software, which is a Python extension that makes use of the HTK speech
recognition tool kit, for aligning the words in the transcript and the audio wave file.
This operation requires a clean transcription without annotations or commentaries, only
the uttered words.</p>
      <p>A third challenge is the abstract representation of student behavior from multimodal
features. After the features are extracted, a statistical approach is required to model the
students’ behaviors. These statistical models have to account for temporal dependencies
in the data. When using latent models or cluster analysis, as is common during
exploratory analyses, one needs to interpret the meaning of these natural groupings post facto.</p>
    </sec>
    <sec id="sec-4">
      <title>Conclusions</title>
      <p>MMLA analysis faces several challenges along the way. Two major challenges have
been identified: (a) Data pre-processing such as synchronization and identification of
features to be extracted; and (b) How to statistically represent student behavior from
fine-grained multimodal features and how to account for temporally dependent data.
We argue that identifying and explaining the usefulness of multimodal features requires
a validation approach. To help this validation process, we suggest an MMLA
Exploratory Framework to guide the collection and interpretation of data. The exploratory
framework suggests data collection at least at two time points to account for
withinsubject variation and learning gains.
5</p>
    </sec>
    <sec id="sec-5">
      <title>Acknowledgments</title>
      <p>This work has been partially funded by an NSF Data Consortium Fellowship.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Baltrušaitis</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Mahmoud</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Robinson</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <year>2015</year>
          .
          <article-title>Cross-dataset learning and personspecific normalisation for automatic Action Unit detection</article-title>
          .
          <source>In Automatic Face and Gesture Recognition (FG)</source>
          ,
          <year>2015</year>
          11th IEEE International Conference and Workshops on IEEE,
          <fpage>1</fpage>
          -
          <lpage>6</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <given-names>D</given-names>
            <surname>'mello</surname>
          </string-name>
          , S. and
          <string-name>
            <surname>Graesser</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <year>2012</year>
          .
          <article-title>Dynamics of affective states during complex learning</article-title>
          .
          <source>Learning and Instruction 22</source>
          ,
          <issue>2</issue>
          ,
          <fpage>145</fpage>
          -
          <lpage>157</lpage>
          . DOI= http://dx.doi.org/https://doi.org/10.1016/j.learninstruc.
          <year>2011</year>
          .
          <volume>10</volume>
          .001.
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Hokayem</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          , Ma, J., and
          <string-name>
            <surname>Jin</surname>
          </string-name>
          , H.,
          <year>2015</year>
          .
          <article-title>A learning progression for feedback loop reasoning at lower elementary level</article-title>
          .
          <source>Journal of Biological Education</source>
          <volume>49</volume>
          ,
          <issue>3</issue>
          ,
          <fpage>246</fpage>
          -
          <lpage>260</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Worsley</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          and
          <string-name>
            <surname>Blikstein</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <year>2015</year>
          .
          <article-title>Using learning analytics to study cognitive disequilibrium in a complex learning environment</article-title>
          .
          <source>In Proceedings of the Fifth International Conference on Learning Analytics And Knowledge - LAK'15 ACM</source>
          , New York, NY, USA,
          <fpage>426</fpage>
          -
          <lpage>427</lpage>
          . DOI= http://dx.doi.org/10.1145/2723576.2723659.
          <string-name>
            <surname>Author</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          :
          <article-title>Article title</article-title>
          .
          <source>Journal</source>
          <volume>2</volume>
          (
          <issue>5</issue>
          ),
          <fpage>99</fpage>
          -
          <lpage>110</lpage>
          (
          <year>2016</year>
          ).
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>