<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Multimodal Transcript of Face-to-Face Group-Work Activity Around Interactive Tabletops</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Xavier Ochoa</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Katherine Chiluiza</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Roger Granda</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Gabriel Falcones</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>James Castells</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>ESPOL</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>xavier</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>kchilui</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>roger.granda</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>gabriel.falcones</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>james.castells</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>bruno.guaman}@cti.espol.edu.ec</string-name>
        </contrib>
      </contrib-group>
      <abstract>
        <p>This paper describes a multimodal system around a multi-touch tabletop to collect different data sources for group-work activities. The system collects data from various cameras, microphones and the logs of the activities performed in the multi-touch tabletop. We conducted a pilot study with 27 students in an authentic classroom to explore the feasibility of capture individual and group interactions of each participant in a collaborative database design activity. From the raw data, we extracted low-level features (e.g. tabletop action, gaze interaction, verbal intervention, emotions) and generated some visualizations as annotated transcripts of what happened in a work session. We evaluated teacher's perception about how the automated multimodal transcript could potentially support the understanding of group-work activities. Results from teachers' perceptions pointed out that the multimodal transcript could become a valuable tool to understand group-work rapport and performance.</p>
      </abstract>
      <kwd-group>
        <kwd>multimodal transcripts</kwd>
        <kwd>collaboration</kwd>
        <kwd>group-work visualizations</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>INTRODUCTION</title>
      <p>
        With the evolution of multi-user tabletop devices, the opportunities to enhance collaboration in
several contexts have been extended significantly, especially in collaborative learning contexts.
Several authors have studied the effect of introducing this particular technology in collaborative
sessions, reporting positive results on enhancing communication skills between participants
        <xref ref-type="bibr" rid="ref10 ref11 ref3">(Kharrufa, et al. 2013; Heslop, 2015)</xref>
        . The data-capture capabilities of tabletops in learning contexts
present new opportunities to better understand the collaboration and learning processes during
group-work activities in the classroom. For instance, collaboration interactions gathered from a
tabletop setting could help teachers by making group-work orchestration easier
(MartinezMaldonado et al., 2011) or help students to reflect about their collaboration experience. However,
using only the data produced by the tabletops provide a narrow picture of those processes because
students interact through a variety of modes (speech, gaze, posture, gestures, etc.) and not all of the
actions are perceived or recorded by the tabletop software
        <xref ref-type="bibr" rid="ref14">(Martinez-Maldonado et al. 2017)</xref>
        . A
common approach to obtain a more holistic view of the collaboration is to complement the data
captured by the tabletops with the capture and analysis from other sources such as video recordings
        <xref ref-type="bibr" rid="ref3">(Al-Qaraghuli, 2013)</xref>
        . However, the visualization of these data, especially if several communication
modalities want to be captured, could become cluttered and confusing for teachers who seek to
provide instant feedback to students after a group-work activity.
      </p>
      <p>
        In this sense, an automatic multimodal transcript has proven to be an efficient and comprehensive
method to represent and visualize temporal information from several sources
        <xref ref-type="bibr" rid="ref4">(Bezemer &amp; Mavers,
2011)</xref>
        . Thus, combining multimodal collaboration features into a time-based visualization could help
teachers to make sense of collaboration processes through the observation of students’ actions and
emotions in the group-work activity
        <xref ref-type="bibr" rid="ref12 ref16">(Martinez-Maldonado et al., 2011; Tang et al., 2010)</xref>
        . Even though
some studies have added new dimensions of collaboration to the data obtained from tabletop, most
of the efforts to create automated transcripts from group’s interactions have been focused on one or
two modalities. For example,
        <xref ref-type="bibr" rid="ref13">Martinez-Maldonado et al. (2013)</xref>
        presented an approach to identify
common patterns of collaboration by mining student logs and detected speech. In another work,
        <xref ref-type="bibr" rid="ref2">Adachi et al. (2015)</xref>
        captured and visualized gaze and talking participation of members in a co-located
conversation to provide feedback that in turn would help balancing participation. These studies serve
as a baseline for our analysis; however, we want to explore the potential of generating an automated
transcript by combining multiple modalities to inform teachers about groups’ interactions around a
tabletop. Besides, in a classroom interaction research context, the focus has been almost exclusively
on teacher talk or teacher–student talk and not in student-student talk in group-work e.g.
studentstudent rapport building in group-work (Ädel, 2011). Ultimately, we want to know if it is possible for
teachers to determine more evidence about collaboration, such as group rapport from the transcripts
generated.
      </p>
    </sec>
    <sec id="sec-2">
      <title>MULTIMODAL INTERACTIVE TABLETOP SYSTEM</title>
      <p>
        The system is a variation of a prototype presented by
        <xref ref-type="bibr" rid="ref7">Echeverria et al. (2017)</xref>
        , that fosters the
collaborative design process. Besides the sensors used in the previous version (Kinect V1, coffee table
and three tablets), this new version adds three multimodal selfies (
        <xref ref-type="bibr" rid="ref6">Dominguez et al., 2015</xref>
        ) with lapel
microphones attached for capturing individual speech. The multimodal selfies are located around the
tabletop to capture synchronized video and audio for each participant. Additionally, a multimodal
selfie was used to capture the video of the entire group session. Figure 1 shows the components of
the system and the working prototype.
      </p>
      <p>
        The software of the system has three different applications: a tabletop application, a management
web application, and a recording application. The tabletop application was designed following the
design principles described in
        <xref ref-type="bibr" rid="ref17">Wong-Villacres et al. (2015)</xref>
        . It allows the participants to develop a
database design through the creation and modification of several interactive objects (entities,
attributes, relations). The management web application communicates with the tabletop application
and with the multimodal selfies to control the execution of the session start the recordings. It also
allows the instructor to view the solution developed by the students
        <xref ref-type="bibr" rid="ref7">(see Echeverria, et al. 2017 for
details)</xref>
        . The recording application was deployed in each multi-modal selfie. It controls the
synchronization of the recordings by the implementation of a publish/subscribe solution using the
lightweight MQTT connectivity protocol1. The recordings obtained are further processed to
automatically tag them according to the proposed audio and video features (see section 2.1). All logs
and features are stored in a relational database. In addition, raw audio and video data are saved in a
NAS Server, associated with a code for identifying each student.
2.1
      </p>
      <p>Multimodal Transcript
The multimodal transcript combines a set of automatic features (extracted from video, audio and
tabletop action logs) into a timeline where the teacher can observe the moment each interaction took
place, and how the group session developed through time. The following sections present the set of
features and details of the multimodal transcript.</p>
      <sec id="sec-2-1">
        <title>2.1.1 Audio Features</title>
        <p>From the speech recorded by the system, we used a Speech-to-Text recognition software, to obtain
an automated transcript of the conversation among participants in the group. Then, we extracted the
speech sections from the recorded individual audio of each participant, and then converted to text
using Google's Cloud Speech API2. In this way, we obtained the conversation between the participants
along with the time each verbal interaction took place. Google's API results using Spanish language
are not as accurate as results obtained using English language. In spite of that, it is still useful to
retrieve sentences with words related to the design problem.</p>
      </sec>
      <sec id="sec-2-2">
        <title>2.1.2 Video Features</title>
        <p>
          Mutual gaze and smiles has been considered as non-verbal indicators of rapport in previous work,
          <xref ref-type="bibr" rid="ref9">(Harrigan et al., 1985)</xref>
          . Thus, we believe that those features could be a valuable feature to be depicted
in the multimodal transcript. Key points from the face of each participant recorded by the Multimodal
Selfie were extracted using the OpenPose Library
          <xref ref-type="bibr" rid="ref5">(Cao et al., 2016)</xref>
          . This library retrieves the
coordinates of 20 points (e.g. eyes, nose, ears, etc.). These face key points were analyzed on every
frame of the recordings, and an algorithm was developed to automatically estimate the moments
when a participant is looking towards to another. This evaluation was carried out by counting false
positives and false negatives from all detections made by the algorithm. To evaluate the accuracy, we
considered 5 different videos of 5 minute-length from the original group sessions (see section 4 for
details). According to our evaluation, this feature has an error rate of 15.1%. In addition, we extracted
the emotions each participant demonstrated during the activity. Video frames of individual recordings
from the Multimodal Selfie were processed using the Microsoft Emotion API3. Thus, for every second,
one frame of the participant's face is sent to the API, which returns an array of scores determining
levels of happiness, anger, disgust, among others, with values from 0 to 1. To evaluate the accuracy
of the emotion recognition software, videos of three students (25 min approx.) were randomly
1 http://mqtt.org/
2 https://cloud.google.com/speech/
3 https://azure.microsoft.com/en-us/services/cognitive-services/emotion/
3
selected from all the groups that participated in a pilot study (see section 4 for details). We selected
happiness as the emotion to be evaluated because it was the most common detected emotion in the
recorded sessions. A human evaluated the videos by annotating if the student was happy or not for
each second. Since the API
returns values between 0 and 1,
we selected a threshold above
0.5 to determine if the student’s
emotion corresponded to
happiness. Our evaluation
resulted in an average error rate
of 1.77%.
        </p>
      </sec>
      <sec id="sec-2-3">
        <title>2.1.3 Interactive Tabletop</title>
      </sec>
      <sec id="sec-2-4">
        <title>Logs</title>
        <p>All the interactions with the
objects on the tabletop were
recorded in a database. Each
interaction is represented by the Figure 2: An excerpt of the multimodal transcript from a group
following features: type of
interaction (CREATE, EDIT, DELETE), student identity, timestamp, and the type of object the student
created. The solution proposed by Martínez, R. et al. (2011) was used for student differentiation while
interacting with an object on the tabletop. Figure 2 shows an excerpt of a group session captured by
the system. As we can see, different features of the group’s interaction are represented (e.g. tabletop
actions, gaze, etc.) in a vertical timeline. For instance, we can observe that in t = 1, student 1 (S1) and
student 2 (S2) were looking to the right, student 3 (S3) was looking to the left, and so on.
3</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>PILOT STUDY</title>
      <p>
        The purpose of this study was to validate with teachers the results obtained from the proposed
multimodal transcript gathered from groups’ sessions. The study is divided in two parts. In a first part
of the study, twenty-four undergraduate students from a Computer Science program (20 males, 4
females, average age: 24 years), enrolled in an introductory Database Systems course and were asked
to participate in a collaborative session using the proposed system. Eight groups were conformed
(three students each) and grouped by affinity. Each group worked approximately 30 minutes in the
design session. All of the multimodal features were recorded while the students were solving the
design problem. At the end of the session, the system scored the solution proposed by the group and
the teacher gave feedback to students about their performance. In addition, each student reported
the rapport of their group according to their Enjoyable Interaction and Personal Connection
        <xref ref-type="bibr" rid="ref8">(Frisby et
al., 2010)</xref>
        .
      </p>
      <p>In the second part of the study, four teachers (3 males, 1 female, avg. age 31) with previous experience
on teaching Database Design, were invited to participate in the evaluation of the multimodal
transcript. First, teachers watched a video showing how the multimodal tabletop system works. Then,
they observed the multimodal transcripts from the data gathered from two groups, corresponding to
the highest and lowest rapport scores from self-reported data. For purposes of simplicity, only
relevant fragments and a summary of the transcript were used in the observations. Next, the teacher
answered a set of questions about the perception of Enjoyable Interaction and Personal Connection
regarding each group. They assigned a score to each group for each variable using a three-point Likert
Scale (Low: 1, Medium: 2. High. 3). Additionally, we included open questions about how the
multimodal transcript could potentially provide support to the teacher to evaluate and recreate the
group-work performance.
4</p>
    </sec>
    <sec id="sec-4">
      <title>RESULTS AND DISCUSSION</title>
      <p>As for Enjoyable Interaction, teachers perceived a low interaction in the group with lowest rapport,
for instance they reported an average of 1.25 over 3; whereas in the group with highest rapport, they
scored an average of 2.75 over 3. As for Personal Connection, a similar pattern was observed, the
group with lowest rapport was evaluated with an average of 1.25 over 3, and the one with highest
rapport scored 2.5 over 3. From these results, it seems that the multimodal transcripts were valuable
for the teachers, since they mostly agreed with the rapport reported by the members of the groups.
Some positive comments about the support the teachers perceived from the multimodal transcript
for assessing enjoyable interaction and personal connection are presented as follows. One teacher
stated: "the combination of voice and emotions presented in the transcript give me the idea of how
the students felt about the task during the session", another teacher said that: "what I observed gives
me evidences about the interaction among students … more specifically, the mix between actions and
emotions are the evidence of such interactions". In addition, there were also some critical remarks.
For instance, one teacher indicated that the transcript: "did not present enough details about the
emotions of the participants". Another teacher said that "the emotions presented are not enough to
infer the interactions that were present, I think there's the need to evidence the interrelations between
the emotions of one participant with the others".</p>
      <p>As for the perception of teachers about how the transcript would support them to evaluate the
groupwork performance, all the interviewed participants answered positively to this question. Regarding
the recreation of student work during the session using the multimodal transcript, three teachers
answered positively and one indicated that the transcript would partially support this task. During the
interviews one teacher stated that: "I could observe whether the students were working on the task or
they were debating about the task; moreover, I can observe the actions at the level of the individual.
It is easy to identify who is the one who work the most or if the task was equally distributed". Another
teacher suggested the following: "It would be nice that the students' comment could be analyzed as
well at the level of emotions". One teacher had a slightly reluctant reaction about the recreation of
student work: "I think it is still ambiguous what the emotions reflect in the transcript; however, the
actions performed using the tabletop could help me in the recreation of the work".
The validation stage of this work points out to a promising research path. Teachers were mostly
positive about the potential of the multimodal transcript to support group-work evaluation, beyond
actions and scores. Teachers valued the fact that emotions were present in the transcript. They
thought that the mix of this feature with voice and task would support the inference of interactions
between the members of the groups. Nevertheless, this work is an on-going project that needs to
further explore how to expand the meaning of emotions in the task, as well as, the reactions between
the members of the groups after some enacted emotions.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          <string-name>
            <surname>Ädel</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          (
          <year>2011</year>
          ).
          <article-title>Rapport building in student group work</article-title>
          .
          <source>Journal Of Pragmatics</source>
          ,
          <volume>43</volume>
          (
          <issue>12</issue>
          ),
          <fpage>2932</fpage>
          -
          <lpage>2947</lpage>
          . http://dx.doi.org/10.1016/j.pragma.
          <year>2011</year>
          .
          <volume>05</volume>
          .007.
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          <string-name>
            <surname>Adachi</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Haruna</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Myojin</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>Shimada</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          (
          <year>2015</year>
          ).
          <article-title>ScoringTalk and watchingmeter: Utterance and gaze visualization for co-located collaboration</article-title>
          .
          <source>SIGGRAPH Asia 2015</source>
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          <string-name>
            <surname>Al-Qaraghuli</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zaman</surname>
            ,
            <given-names>H. B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ahmad</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>Raoof</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          (
          <year>2013</year>
          , November).
          <article-title>Interaction patterns for assessment of learners in tabletop based collaborative learning environment</article-title>
          .
          <source>In Proceedings of the 25th Australian Computer-Human Interaction Conference</source>
          . (pp.
          <fpage>447</fpage>
          -
          <lpage>450</lpage>
          ). ACM.
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          <string-name>
            <surname>Bezemer</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>Mavers</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          (
          <year>2011</year>
          ).
          <article-title>Multimodal transcription as academic practice: a social semiotic perspective</article-title>
          .
          <source>International Journal of Social Research Methodology</source>
          ,
          <volume>14</volume>
          (
          <issue>3</issue>
          ),
          <fpage>191</fpage>
          -
          <lpage>206</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          <string-name>
            <surname>Cao</surname>
            ,
            <given-names>Z.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Simon</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Wei</surname>
            ,
            <given-names>S. E.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>Sheikh</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          (
          <year>2016</year>
          ).
          <article-title>Realtime multi-person 2d pose estimation using part affinity fields</article-title>
          .
          <source>arXiv preprint arXiv:1611</source>
          .
          <fpage>08050</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          <string-name>
            <surname>Domínguez</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chiluiza</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Echeverria</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>Ochoa</surname>
            ,
            <given-names>X.</given-names>
          </string-name>
          (
          <year>2015</year>
          ).
          <article-title>Multimodal selfies: Designing a multimodal recording device for students in traditional classrooms</article-title>
          .
          <source>In Proceedings of the 2015 ACM on Intl. Conf. on Multimodal Interaction</source>
          (pp.
          <fpage>567</fpage>
          -
          <lpage>574</lpage>
          ). ACM.
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          <string-name>
            <surname>Echeverria</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Falcones</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Castells</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Martinez-Maldonado</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>Chiluiza</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          (
          <year>2017</year>
          ).
          <article-title>Exploring on-time automated assessment in a co-located collaborative system</article-title>
          .
          <source>Paper presented at the 4th Intl. Conf. on eDemocracy and eGovernment, ICEDEG</source>
          <year>2017</year>
          ,
          <volume>273</volume>
          -
          <fpage>276</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          <string-name>
            <surname>Frisby</surname>
            ,
            <given-names>B. N.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>Martin</surname>
            ,
            <given-names>M. M.</given-names>
          </string-name>
          (
          <year>2010</year>
          ).
          <article-title>Instructor-student and student-student rapport in the classroom</article-title>
          .
          <source>Communication Education</source>
          ,
          <volume>59</volume>
          (
          <issue>2</issue>
          ),
          <fpage>146</fpage>
          -
          <lpage>164</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          <string-name>
            <surname>Harrigan</surname>
            ,
            <given-names>J. A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Oxman</surname>
            ,
            <given-names>T. E.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>Rosenthal</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          (
          <year>1985</year>
          ).
          <article-title>Rapport expressed through nonverbal behavior</article-title>
          .
          <source>Journal of nonverbal behavior</source>
          ,
          <volume>9</volume>
          (
          <issue>2</issue>
          ),
          <fpage>95</fpage>
          -
          <lpage>110</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          <string-name>
            <surname>Heslop</surname>
            , Philip &amp; Preston, Anne &amp; Kharrufa, Ahmed &amp; Balaam, Madeline &amp; Leat, David &amp; Olivier,
            <given-names>Patrick.</given-names>
          </string-name>
          (
          <year>2015</year>
          ).
          <article-title>Evaluating Digital Tabletop Collaborative Writing in the Classroom</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          <string-name>
            <surname>Kharrufa</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Balaam</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Heslop</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Leat</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Dolan</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>Olivier</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          (
          <year>2013</year>
          ).
          <article-title>Tables in the wild: Lessons learned from a large-scale multi-tabletop deployment</article-title>
          .
          <source>Conf. on Human Factors in Computing Systems</source>
          ,
          <volume>1021</volume>
          -
          <fpage>1030</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          <string-name>
            <surname>Martinez-Maldonado</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          , Collins,
          <string-name>
            <given-names>A.</given-names>
            ,
            <surname>Kay</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            , and
            <surname>Yacef</surname>
          </string-name>
          ,
          <string-name>
            <surname>K.</surname>
          </string-name>
          (
          <year>2011</year>
          )
          <article-title>Who did what? who said that? Collaid: an environment for capturing traces of collaborative learning at the tabletop</article-title>
          .
          <source>ACM Intl. Conf. on Interactive Tabletops and Surfaces</source>
          ,
          <string-name>
            <surname>ITS</surname>
          </string-name>
          <year>2011</year>
          , pages
          <fpage>172</fpage>
          -
          <lpage>181</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          <string-name>
            <surname>Martinez-Maldonado</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kay</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>Yacef</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          (
          <year>2013</year>
          ).
          <article-title>An automatic approach for mining patterns of collaboration around an interactive tabletop10</article-title>
          .
          <volume>1007</volume>
          /978-3-
          <fpage>642</fpage>
          -39112-5-11
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          <string-name>
            <surname>Martinez-Maldonado</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kay</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <given-names>Buckingham</given-names>
            <surname>Shum</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S. J.</given-names>
            , &amp;
            <surname>Yacef</surname>
          </string-name>
          ,
          <string-name>
            <surname>K.</surname>
          </string-name>
          (
          <year>2017</year>
          ).
          <article-title>Collocated collaboration analytics: Principles and dilemmas for mining multimodal interaction data. Human-Computer Interaction</article-title>
          ..
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          <string-name>
            <surname>Schneider</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sharma</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Cuendet</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zufferey</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Dillenbourg</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>Pea</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          (
          <year>2016</year>
          ).
          <article-title>Detecting collaborative dynamics using mobile eye-trackers</article-title>
          .
          <source>Proceedings of International Conference of the Learning Sciences, ICLS</source>
          ,
          <volume>1</volume>
          ,
          <fpage>522</fpage>
          -
          <lpage>529</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          <string-name>
            <surname>Tang</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Pahud</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Carpendale</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>Buxton</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          (
          <year>2010</year>
          ).
          <article-title>VisTACO: visualizing tabletop collaboration</article-title>
          .
          <source>In ACM International Conference on Interactive Tabletops and Surfaces</source>
          (pp.
          <fpage>29</fpage>
          -
          <lpage>38</lpage>
          ). ACM.
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          <string-name>
            <surname>Wong-Villacres</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ortiz</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Echeverría</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>Chiluiza</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          (
          <year>2015</year>
          ).
          <article-title>A tabletop system to promote argumentation in computer science students</article-title>
          .
          <source>In Proceedings of the 2015 Intl. Conf. on Interactive Tabletops &amp; Surfaces</source>
          (pp.
          <fpage>325</fpage>
          -
          <lpage>330</lpage>
          ). ACM.
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>