<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>The Multimodal Study of Blended Learning Using Mixed Sources: Dataset and Challenges of the SpeakUp Case</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Mar a Jesus Rodr guez-Triana</string-name>
          <email>maria.rodrigueztriana@epfl.ch</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Luis P. Prieto</string-name>
          <email>luis.prieto@tlu.ee</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Adrian Holzer</string-name>
          <email>adrian.holzer@epfl.ch</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Denis Gillet</string-name>
          <email>denis.gillet@epfl.ch</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Ecole Polytechnique Federale de Lausanne</institution>
          ,
          <addr-line>Lausanne</addr-line>
          ,
          <country country="CH">Switzerland</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Tallinn University</institution>
          ,
          <addr-line>Tallinn</addr-line>
          ,
          <country country="EE">Estonia</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>Social media applications have been proposed as a tool to complement students' formal learning experiences, often to increase interactivity and participation. However, evidence regarding the bene ts and challenges of such applications is still con icting. In our latest study to explore this conundrum, we have gathered a multimodal dataset that showcases the teaching and learning processes co-occurring simultaneously on a physical space (face-to-face university lectures) and a digital one (SpeakUp, a social media app). The raw data, provided by di erent sources and informants, were transformed and analyzed using mixed (quantitative and qualitative) techniques. In this contribution, we describe the multiple pieces that composed our dataset, and the steps we took in the multimodal analyses to explore the learning experience occurring in both the physical and digital spaces. This dataset and analysis pipeline illustrates not only challenges and limitations speci c to our study, but also more general ones. Several such challenges and limitations, commonplace in blended learning settings analyzed using mixed (multimodal) methods, are synthesized at the end of our paper.</p>
      </abstract>
      <kwd-group>
        <kwd>Multimodal learning analytics (MMLA)</kwd>
        <kwd>blended learning</kwd>
        <kwd>mixed methods</kwd>
        <kwd>social media</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>
        Technology-enhanced learning (TEL) is, almost invariably, blended in nature:
we create new digital spaces and channels to interact and learn { yet we still
inhabit the physical world and also learn through it. This inherently blended
nature of learning not only has prompted a methodological turn towards mixed
methods [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ], but also holds great promises for the rise of multimodal learning
analytics (MMLA), as researchers strive to understand more deeply the learning
processes and outcomes occurring in both kinds of spaces [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ].
      </p>
      <p>
        One example of such research is our ongoing project to study the usage
of social media applications to complement formal, co-located learning
experiences (e.g., face-to-face university courses). Currently there is no consensus as
to whether the bene ts that such applications provide in terms of engagement
and interaction, outweigh their potential cost as a source of distraction [
        <xref ref-type="bibr" rid="ref4 ref6 ref7">6, 4, 7</xref>
        ].
To help in clarifying these issues, we are performing a case study in an authentic
setting, one of our university courses at the Ecole Polytechnique Federale de
Lausanne (EPFL) in Switzerland.
      </p>
      <p>In this face-to-face course, composed mainly of lectures with more than a
hundred students, the usage of a social media tool (SpeakUp3) was proposed
in order to foster the (otherwise limited) interaction between students and with
the instructors. In a typical usage scenario with SpeakUp, teachers create a
chatroom that students can join. Inside the chatroom, any user can anonymously
post text messages, comment on existing messages, and up/down-vote them.
Moreover, SpeakUp had an added value from the analytics perspective: it allows
the chatroom creator to download all the traces collected inside of the room,
including both actions and posts.</p>
      <p>
        This paper provides an overview of the multimodal dataset gathered and
the analyses performed in the study. A more detailed account of the setting,
and a partial analysis of the data (including the e ects of SpeakUp usage on
student engagement, distraction, social interaction and teaching style, as well
as their relationship with learning outcomes), are described in [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]. We end our
contribution by re ecting on the limitations and challenges that we have faced,
as they pertain to those common in the multimodal study of blended learning
situations through a mix of quantitative and qualitative analyses.
2
      </p>
    </sec>
    <sec id="sec-2">
      <title>Multimodal Dataset</title>
      <p>The dataset for our study was gathered throughout a whole university course on
Communication, which involved 6 face-to-face sessions which lasted 90 minutes
each. An average of 3 lecturers and 145 students attended per session. We
combined quantitative and qualitative data coming from four types of informants (3
teachers, 145 students, 4 assistants, 1 researcher, plus the SpeakUp system itself)
using di erent data gathering techniques, namely: questionnaires, observations,
video recordings and system logs. Figure 1 o ers an overview of the dataset. In
each session, the following data were gathered:
{ The researcher video recorded the session (focusing mostly on the front of
the class) [r vid]
{ The researcher also wrote down timestamped observations about what was
happening during the session (e.g., beginning, start, breaks, activities, topics
discussed, interventions, problems with the app, number of students in the
room) [r obs]
{ In parallel, the assistants involved in the course kept track of the students
who participated face-to-face in the session (e.g., posing questions or
participating in the general discussions) [a obs]
3 http://speakup.info
{ SpeakUp logs were used to track the activity of the instructors and students
joining the chatroom4 (e.g., timestamped posted messages, number of likes
and dislikes, etc.) [sp log]</p>
      <p>In addition, there were a number of complementary data sources collected at
the beginning and at the end of the course:
{ To understand the student predisposition towards the usage of technology
and social media for learning, we conducted a questionnaire at the beginning
of the rst session (based on 7-point Likert scale questions) [s que1].
{ To gather the student and teacher perceptions on how the tool usage a ected
the engagement and attention, we conducted questionnaires at the end of the
rst session [s que2] [t que]. While the former was made of 7-point Likert
scale questions, the later combined 5-point Likert scale and open questions.
4 Since in SpeakUp, users join anonymously the chatrooms, there was no way to gure
out who was behind each user identi er. Thus, we asked the students to freely reveal
their identities just for research purposes. It is noteworthy that SpeakUp logs are
multimodal since they contain not only activity traces but also the text posted by
the users.
{ At the end of the course, students answered a test composed of
multiplechoice questions about the topics discussed in the di erent sessions. The
scores [s sco] were used for exploring whether the user's actions on SpeakUp
during the topic discussions had an impact on the student answers.</p>
      <p>It is noteworthy that, although the study was carried out in an university
course on Communication, the data sources used and the data gathering
techniques applied are not dependant to the content or the educational context.
Thus, the same strategies and techniques could be applied in other settings.
3</p>
    </sec>
    <sec id="sec-3">
      <title>Multimodal Data Analysis</title>
      <p>
        As detailed in [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ], a rst partial analysis was performed by transforming and
integrating the data from the rst session only. This rst analysis involved both
qualitative analyses (manual coding of the actions recorded in the videos, and
content analysis of the messages generated by the users), as well as quantitative
ones (descriptive statistics and exploratory computational analyses of system
logs). Figure 1 o ers an overview of the di erent analyses applied to the dataset.
      </p>
      <p>Qualitative analyses. We manually coded all the messages and comments
generated during the lesson (thus enriching [sp log]) to determine whether
they were relevant to learning, and the direction of the interaction (students
to teachers, students to students, students to all, and teachers to students).
In a similar way, and in order to understand these topics as they occurred in
the face-to-face channel of the classroom, the video recording of the lesson was
also coded (enriching [r vid]), according to several categories: which actor was
speaking; what topic to appear in the scoring test [s sco] was being discussed,
if any; what teacher action was being performed at that moment (e.g.,
presentation/lecturing, asking questions, providing answers, noting technical or other
kinds of problems); who was the target of the interaction, if any (e.g., a teacher,
students, or all the class); and nally, what supporting resources were being
used, if any (e.g., slides, videos, SpeakUp).</p>
      <p>Quantitative analyses. The user activity was measured applying descriptive
statistical analyses to the actions tracked in the SpeakUp logs (e.g., number
of posted messages, number of likes and dislikes, etc.) [sp log], to the
face-toface utterances registered by the researcher and the assistants [r obs] [a obs],
and to the results of the manually-coded video [r vid] and SpeakUp comments
[sp log]. Furthermore, a clustering analysis (using a k-means algorithm) was
performed on each student's activity features [sp log] (number of messages,
responses, likes/dislikes, etc.), in order to identify usage pro les. Finally, the
data from the users' activity [sp log] were triangulated with the teachers' and
students' perceptions from the questionnaires [t que] [s que1] [s que2] and
the scores obtained by the students in the nal test [s sco], to understand the
impact of such engagement and participation in the learning outcomes.</p>
    </sec>
    <sec id="sec-4">
      <title>Limitations and Challenges</title>
      <p>Despite the richness of the dataset and the usefulness of the analyses described
in the previous sections, it presents several limitations, especially apparent in
terms of reproducibility and scalability of our approach. While some of these
limitations are speci c to the particular implementation of our study, others
represent widespread challenges in MMLA that tries to study blended learning
settings using a mix of quantitative and qualitative techniques:</p>
      <p>Manual data gathering. The fact that several of our data sources originate
directly from manual work by human actors (e.g., observations by researchers
or assistants). The lack of tooling to easily (and consistently) timestamp, label
and export such manual data poses limitations to the scaling and di usion of
this kind of e orts.</p>
      <p>Data integration. In our study, di erent units of analysis or measure were
used. For instance, while the videos allowed us to measure the length of the
face-to-face interventions, the logs informed us only about discrete
computermediated events, without a duration. This makes merging and comparison di
cult, and illustrates a common issue when using multimodal datasets: the
heterogeneous nature of the di erent data sources. Despite the e orts put in order
to adopt interoperable standards and speci cations (like Caliper or xAPI), this
heterogeneity will be hardly avoidable, requiring multiple analysis techniques.</p>
      <p>User identi cation and anonymization. The usage of multiple data sources
entails the need to identify a certain user across data sources. Computer-mediated
user actions are often easy to trace; however, in video or audio data this can be
a challenging (if done automatically) or cumbersome task (if done manually).
In our particular study, the situation was even more complicated because both
the questionnaires and the system logs provided by SpeakUp were anonymous.
This brings up the tension between user traceability and privacy, which will
be brought to the forefront by the requirements of the recent European
General Data Protection Regulation (GDPR EU 2016/679). This kind of regulation
may lead, in the near future, to technological tools that only expose anonymous
data (hence hindering learning analytics and interventions that target speci c
students).</p>
      <p>
        Manual data analysis. As it happens with data gathering, the need for
human involvement during the analysis limits the scalability of our approach. In
our study, both comments and videos were manually analysed by teachers and/or
researchers. Although there are potential solutions that could facilitate the
manual content analyses (e.g., crowdsourcing by letting the users tag themselves the
comments), others like the video analyses remain still a challenge. We
envision that alternative MMLA techniques and approaches (such as speech or text
analyses [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]) could ameliorate the aforementioned limitations, and contribute to
avert or circumvent the MMLA challenges in the mixed-method study of blended
learning phenomena. For example, voice recognition techniques could
automatize part of the video analyses, identifying the di erent speakers participating
during the session. In addition, transcriptions could be automatically generated
by applying speech recognition to the audio recorded during the sessions. Later
on, content analyses could be applied to the transcriptions, the comments post
by the students, and the questions of the test, to explore the relations among
them. However, it should be noticed that, to put all these ideas in action, further
research would be necessary to provide speech recognition and content analyses
solutions applicable to di erent languages.
5
      </p>
    </sec>
    <sec id="sec-5">
      <title>Acknowledgements</title>
      <p>This work was partially funded by the European Union in the context of the
Go-Lab Integrated Project (FP7-ICT action, grant no. 317601), SiWay
(FP7ICT, IMAILE grant no. 619231-459.76301), the Next-Lab Innovation Action
(H2020-ICT, grant no. 731685), and CEITER (H2020-WIDESPREAD-2014-2,
CSA action, grant no.669074)</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <given-names>P.</given-names>
            <surname>Blikstein</surname>
          </string-name>
          and
          <string-name>
            <given-names>M.</given-names>
            <surname>Worsley</surname>
          </string-name>
          .
          <article-title>Multimodal Learning Analytics and Education Data Mining: Using Computational Technologies to Measure Complex Learning Tasks</article-title>
          .
          <source>Journal of Learning Analytics</source>
          ,
          <volume>3</volume>
          (
          <issue>2</issue>
          ):
          <volume>220</volume>
          {
          <fpage>238</fpage>
          ,
          <year>2016</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <given-names>X.</given-names>
            <surname>Ochoa</surname>
          </string-name>
          and
          <string-name>
            <given-names>M.</given-names>
            <surname>Worsley</surname>
          </string-name>
          . Editorial:
          <article-title>Augmenting Learning Analytics with Multimodal Sensory Data</article-title>
          .
          <source>Journal of Learning Analytics</source>
          ,
          <volume>3</volume>
          (
          <issue>2</issue>
          ):
          <volume>213</volume>
          {
          <fpage>219</fpage>
          ,
          <year>2016</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>M. J. Rodr</surname>
            guez-Triana,
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Holzer</surname>
            ,
            <given-names>L. P.</given-names>
          </string-name>
          <string-name>
            <surname>Prieto</surname>
            , and
            <given-names>D.</given-names>
          </string-name>
          <string-name>
            <surname>Gillet</surname>
          </string-name>
          .
          <article-title>Examining the e ects of social media in co-located classrooms: A case study based on speakup</article-title>
          . In K. Verbert,
          <string-name>
            <given-names>M.</given-names>
            <surname>Sharples</surname>
          </string-name>
          , and T. Klobucar, editors,
          <source>11th European Conference on Technology Enhanced Learning: Adaptive and Adaptable Learning</source>
          , volume
          <volume>9891</volume>
          , pages
          <fpage>247</fpage>
          {
          <fpage>262</fpage>
          ,
          <string-name>
            <surname>Lyon</surname>
          </string-name>
          (France),
          <year>2016</year>
          . Springer International Publishing.
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <given-names>F.</given-names>
            <surname>Sana</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Weston</surname>
          </string-name>
          , and
          <string-name>
            <given-names>N. J.</given-names>
            <surname>Cepeda</surname>
          </string-name>
          .
          <article-title>Laptop multitasking hinders classroom learning for both users and nearby peers</article-title>
          .
          <source>Computers &amp; Education</source>
          ,
          <volume>62</volume>
          :
          <fpage>24</fpage>
          {
          <fpage>31</fpage>
          ,
          <year>2013</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <given-names>J.-W.</given-names>
            <surname>Strijbos</surname>
          </string-name>
          and
          <string-name>
            <given-names>F.</given-names>
            <surname>Fischer</surname>
          </string-name>
          .
          <article-title>Methodological challenges for collaborative learning research</article-title>
          .
          <source>Learning and Instruction</source>
          ,
          <volume>17</volume>
          (
          <issue>4</issue>
          ):
          <volume>389</volume>
          {
          <fpage>393</fpage>
          ,
          <year>2007</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6. Y.-T. Sung,
          <string-name>
            <surname>K.-E. Chang</surname>
          </string-name>
          , and T.-C. Liu.
          <article-title>The e ects of integrating mobile devices with teaching and learning on students' learning performance: A meta-analysis and research synthesis</article-title>
          .
          <source>Computers &amp; Education</source>
          ,
          <volume>94</volume>
          :
          <fpage>252</fpage>
          {
          <fpage>275</fpage>
          ,
          <year>2016</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <given-names>D. L.</given-names>
            <surname>Worthington</surname>
          </string-name>
          and
          <string-name>
            <given-names>D. G.</given-names>
            <surname>Levasseur</surname>
          </string-name>
          .
          <article-title>To provide or not to provide course powerpoint slides? the impact of instructor-provided slides upon student attendance and performance</article-title>
          .
          <source>Computers &amp; Education</source>
          ,
          <volume>85</volume>
          :
          <fpage>14</fpage>
          {
          <fpage>22</fpage>
          ,
          <year>2015</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>