<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Incorporating Emotional Intelligence into Assessment Systems</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Han-Hui Por</string-name>
          <email>HPOR@ets.org</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Aoife Cahill</string-name>
          <email>ACAHILL@ets.org</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Educational Testing Service</institution>
          ,
          <addr-line>Princeton, NJ 08541</addr-line>
          ,
          <country country="US">USA</country>
        </aff>
      </contrib-group>
      <fpage>95</fpage>
      <lpage>103</lpage>
      <abstract>
        <p>This paper proposes developing emotionally intelligent assessments to increase score validity and reliability. We summarize research on three sources of data (process data, response data, and visual and sensory data) that identify students' needs during test taking and highlight the challenges in developing caring adaptive assessments. We conclude that because of its interdisciplinary nature, the development of caring assessments requires closer collaborations among researchers from diverse fields.</p>
      </abstract>
      <kwd-group>
        <kwd>Caring assessments</kwd>
        <kwd>Adaptive testing</kwd>
        <kwd>Emotions</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Introduction</title>
      <p>
        Assessments can be stressful. Under-performance due to test anxiety can undermine the
validity of a test score by failing to recognize the true performance of the student.
Caring assessments [
        <xref ref-type="bibr" rid="ref37">37</xref>
        ] are systems developed to respond to students' needs by adjusting
content sequencing, moderating the amount and type of feedback, adding visualization
aids, etc. In addition to students' ability, caring assessment systems take into account
additional information from both traditional and non-traditional sources (e.g., student
emotions, prior knowledge and opportunities to learn). Caring assessments go beyond
traditional assessments by providing the encouragement and resources that students
might need.
      </p>
      <p>
        One way to identify a student's learning needs is via their emotions during
test-taking, such as the affective-sensitive version of AutoTutor [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ], the uncertainty (i.e.,
confusion) adaptive version of ITSpoke (UNC-ITSpoke) [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ], and emotion-sensitive
versions of Cognitive Tutors and ASSISTments [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. These systems have the “emotional
intelligence” to recognize the emotions and needs of their learners and use the
information to help the learners achieve their learning goals.
      </p>
      <p>In this paper, we argue that a fair and valid caring assessment should take multiple
interdisciplinary sources of information into account. We give an overview of current
state-of-the-art capabilities -- both technological and psychometric -- that are relevant
for designing and developing caring assessments. We outline some of the many
challenges faced and suggest some areas for future work, particularly focusing on
interdisciplinary collaborations. While we are particularly interested in the role of caring
assessments in the context of summative assessments, parts of our discussion will also
refer to elements of caring assessments that will also be helpful for formative
assessments.
2</p>
    </sec>
    <sec id="sec-2">
      <title>Adaptive Testing</title>
      <p>
        In caring assessments, students see different sets of questions and aids that are tailored
to their ability and needs. In traditional computerized adaptive testing (CAT),
pre-calibrated test items or testlets (sets of items, as in multi stage adaptive testing) are
presented to students based on the quality of responses to previous questions. Given the
response (correct/incorrect), an estimate of the student's ability is updated before further
items are selected and presented. As a result of adaptive administration, different
students experience different items. By successively fielding items selected to provide the
maximum information about a student's ability, the maximum information about the
student's ability is collected with each question, resulting in a shorter exam. Compared
to a traditional static exam, a CAT exam is an assessment that tailors itself to each
student's ability. Unfortunately, an algorithm that assigns items based on ability alone
may not necessarily support a positive testing experience [
        <xref ref-type="bibr" rid="ref27">27</xref>
        ]. However, it may be
possible to leverage psychometric approaches to incorporate additional information --
beyond ability -- into item selection in CAT.
2.1
      </p>
      <sec id="sec-2-1">
        <title>Item Selection using Item Response Theory (IRT) and Bayesian Networks</title>
      </sec>
      <sec id="sec-2-2">
        <title>Models</title>
        <p>
          IRT [
          <xref ref-type="bibr" rid="ref19">19</xref>
          ] is the current dominant statistical model used in CAT to assign students
comparable, meaningful scores, even though students see sets of items tailored to their
individual needs. Within the IRT framework, the estimated ability of a student is expected
to be the same regardless of the items. The usefulness of IRT to select hints based on
students' ability was demonstrated in FOSS (Full Option Science System), an ITS that
was part of the NSF-funded Principled Assessment Designs for Inquiry (PADI)
assessment [
          <xref ref-type="bibr" rid="ref29">29</xref>
          ].
        </p>
        <p>
          To take into account both ability as well as the emotional needs of students,
parameters of items in caring assessments can be estimated using models that integrate
person- and variable-centered information such as the mixed-measurement item response
theory (MM-IRT) [
          <xref ref-type="bibr" rid="ref23 ref24">23, 24</xref>
          ]. Such models focus on unobserved characteristics by
identifying latent classes of individuals who respond to items in unexpected but distinct
ways [
          <xref ref-type="bibr" rid="ref13 ref40">13, 40</xref>
          ]. Indeed, there is a recognition that observed scores may not correspond
with unseen differences in how individuals respond to tests, such as differences in test
strategies [
          <xref ref-type="bibr" rid="ref23">23</xref>
          ] or reactions to testing procedures [
          <xref ref-type="bibr" rid="ref3">3</xref>
          ].
        </p>
        <p>
          Bayesian graph models from the Bayesian Network framework is another promising
approach. These models allow the modeling of relevant conditional probabilities to
update the likelihood of an event in the network. A major advantage is that the assessment
of item mastery can be defined using multiple latent traits. In particular, it has been
shown that the POKS (Partial Order Knowledge Structures) Bayesian modeling
approach is computationally simpler and can outperform a 2-parameter IRT model in
some instances [
          <xref ref-type="bibr" rid="ref7">7</xref>
          ]. The Andes Tutor [
          <xref ref-type="bibr" rid="ref5">5</xref>
          ] and Hydrive [
          <xref ref-type="bibr" rid="ref22">22</xref>
          ] both incorporate Bayesian
network models to select and score items.
3
        </p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>Research on Emotions and Affective Computing</title>
      <p>
        Caring assessments have to take into account students' emotional states (e.g., anxiety,
frustration) during test-taking. Using self-reported measures, Lehman and
Zapata-Rivera [
        <xref ref-type="bibr" rid="ref18">18</xref>
        ] identified the emotions that occurred when students completed
conversationbased assessments (CBAs). They found that students experienced similar emotions
across two studies and concluded that boredom, confusion, curiosity, delight,
engagement/flow, frustration, happiness/enjoyment, hope, and pride are the prevalent
emotions in CBAs. In an adaptive assessment, the use of self-reported measures disrupts
test flow and other predictive indicators are necessary to identify emotions accurately.
In addition to correct/ incorrect responses from the typical item types such as multiple
choices items, we can enhance our prediction of students' emotions using three sources
of data: process data, response data, and visual and sensory data.
      </p>
      <p>Process Data such as response time, keystrokes, mouse clicks</p>
      <p>Response Data such as responses to multiple choices or constructed response items,
text or verbal feedback from students</p>
      <p>Visual and Sensory Data such as eye tracking, heart rates, postures, facial
expressions</p>
      <p>The data and model complexity used in an assessment to predict emotions depends
on the assessment objectives. Assessments used to identify areas of students' strengths,
and weaknesses require detailed information about each student to provide very specific
and individualized support. On the other hand, assessments conducted by teachers to
inform teaching and learning direction require aggregated information of the group, and
less precise information about individual students.</p>
      <p>The amount of information also depends on the nature of the assistance the
assessment aims to provide. A formative assessment is likely to provide more assistance in
the form of additional visual aids, redirecting students' focus to the correct cues or
paraphrasing questions when necessary. A summative assessment for determining
proficiency levels can still benefit from collecting some amount of information to promptly
identify students experiencing technical difficulties during the assessment. In the case
of multi-year assessments, information can also be collected to aid in the development
of exams in subsequent years. In the next section, we summarize research that has been
done with each of the three information sources.
3.1</p>
      <sec id="sec-3-1">
        <title>Process Data</title>
        <p>
          Process data, such as typing speed, response time, keystrokes, mouse clicks and action
sequences in problem solving tasks, trace students' progress through an assessment.
Most process data can be collected in the background with minimal incremental costs
and is unobtrusive to students taking the exam. In general, they fall into three
categories: what a student does, in what order, and how long it takes to do it. For instance, an
analysis of the patterns and pauses in students' typing in a NAEP writing test showed
that students who used the delete key more often, as a measure of their attempts to edit,
had higher scores than students who did not delete as much [
          <xref ref-type="bibr" rid="ref33">33</xref>
          ]. The findings suggest
that the latter group could benefit from encouragement to edit. Wise et al. [
          <xref ref-type="bibr" rid="ref34">34</xref>
          ] found
that monitoring learners' response time and displaying warning messages when learners
exhibited rapid guessing behavior improved scores and score validity, as indicated by
the higher correlation between the test score and the learners' GPA and SAT scores.
Other studies using timing data explored students' test taking behavior [
          <xref ref-type="bibr" rid="ref15 ref16">15,16</xref>
          ].
3.2
        </p>
      </sec>
      <sec id="sec-3-2">
        <title>Response Data</title>
        <p>
          Natural language processing (NLP) techniques are widely used to automatically
measure the quality of constructed (free) responses in educational assessments. NLP is used
in automated scoring engines to assess students' level of comprehension or writing
proficiency and subsequently drive the feedback that students receive. Beigman-Klebanov
et al. [
          <xref ref-type="bibr" rid="ref2">2</xref>
          ] show that by using NLP techniques it is possible to automatically predict a
student's utility value -- a measure of how well the student can relate what they are
writing about to themselves or other people -- from the student's writing. Flor et al. [
          <xref ref-type="bibr" rid="ref10">10</xref>
          ]
show that it is possible to automatically categorize the dialogue acts (including
expressing frustration) in a collaborative problem-solving framework using NLP techniques.
NLP techniques are used on both written and spoken text. Spoken data can also provide
a rich amount of information on both speaker emotions as well as their thought process
(disfluencies, pause structure, etc.). Studies have also examined the use of low-level
linguistic features to predict student emotions during human and computer tutoring
sessions [
          <xref ref-type="bibr" rid="ref17 ref8">17, 8</xref>
          ]. Future research can focus on using complex linguistic analysis to learn
more sophisticated relationships between the content of students’ responses and their
emotions in real time assessments.
3.3
        </p>
      </sec>
      <sec id="sec-3-3">
        <title>Visual and Sensory Data</title>
        <p>
          Visual and sensory data can also be captured to provide information on the students'
progress or emotional state. The interest in sensory information stems from the findings
that increased heart-rate and perspiration often precede our actual awareness of
emotions, and studies have shown that heart rate and respiratory frequency can distinguish
between neutral (relaxed), positive (joy) and negative (anger) emotions [
          <xref ref-type="bibr" rid="ref31 ref36">31, 36</xref>
          ].
However, while pulse rate monitors can be small, most devices would likely be obstructive
when taking tests.
        </p>
        <p>
          Although advances in facial recognition technology have vastly improved in recent
years, identifying emotions accurately in real time is still a challenging task. Facial
expressions are an integral part of emotions, but can also exist independently of emotions
[
          <xref ref-type="bibr" rid="ref25">25</xref>
          ] and vice-versa. More recent developments suggest that new facial recognition
algorithms have had some success with extracting features to classify students' emotions
[
          <xref ref-type="bibr" rid="ref35">35</xref>
          ] in real time.
        </p>
        <p>
          Eye tracking also allows us to pinpoint sections of the items that students are focused
on. Studies on the usefulness of eye tracking data have provided preliminary evidence
that they provide insights into how students respond to test items and solve problems
[
          <xref ref-type="bibr" rid="ref21 ref28 ref30">21, 28, 30</xref>
          ]. A study that used eye tracking devices found that students with a history
of performing poorly on reading tests did better when they had to write a summary of
a reading passage before answering multiple-choice questions on the content [
          <xref ref-type="bibr" rid="ref32">32</xref>
          ]. The
eye-tracking data showed that those students spent more time reading the initial text,
and less time referencing the passage, suggesting that the students had built a mental
memory model of the text. The advantage was stronger in students weaker in reading.
4
        </p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>Challenges of Developing Caring Assessments</title>
      <p>
        The development of caring assessments presents numerous challenges and research
directions. Above all, the development of caring adaptive assessments requires closer
interdisciplinary collaboration. Current research on the use of process data, NLP data,
and visual/sensory data is largely focused on how these features correlate with either
students' performance [
        <xref ref-type="bibr" rid="ref21 ref28 ref30 ref33 ref34">21, 28, 30, 33, 34</xref>
        ] or human raters in the field of automated
scoring [
        <xref ref-type="bibr" rid="ref11 ref14 ref20">11, 14, 20</xref>
        ]. On the other hand, data from multiple sources would allow us to
build accurate large-scale models of behavior from which we could then generalize
students' behavior [
        <xref ref-type="bibr" rid="ref38 ref4 ref6">4, 6, 38</xref>
        ] and adapt to their needs. One area of future research is to
focus on the predictive value of data from multiple sources in predicting students'
emotions, and the impact of responding with aids on students' learning.
      </p>
      <p>Other challenges include the costs and benefits of caring assessments over traditional
ones. For caring assessments to be adopted as an industry standard, it will be necessary
to demonstrate the effectiveness of the caring components both in terms of improving
the student experience as well as contributing to overall test reliability and validity. As
the approach to caring assessments is different in different educational context (e.g.,
summative vs formative), additional work is needed to define the elements that make
up caring assessments so that the elements and combination of elements can be studied
for their effectiveness.</p>
      <p>In addition, the widespread adoption of caring assessments will be dependent on
technology. Established assessments such as the GRE, TOEFL, LSAT depends on the
capabilities of their testing centers. Therefore, if a caring assessment requires
high-resolution cameras, all test centers would need to provide that hardware. In large-scale
international assessments with test centers in all corners of the world, this is no small
challenge.
4.1</p>
      <sec id="sec-4-1">
        <title>Psychometrics Challenges in Caring Assessments</title>
        <p>The psychometrics of caring assessments also presents some challenges. The current
challenge is to adapt assessments based on students' ability and needs. While adaptive
testing is not new, we need further research to establish if available models can
accommodate multi modal, individual, and item level characteristics.</p>
      </sec>
      <sec id="sec-4-2">
        <title>Scoring Complex Data Sequences</title>
        <p>
          Another significant challenge for interactive assessments that respond to students' needs
is that students can choose to take a large combination of actions. Should a student be
rewarded with more points for taking fewer steps to get to the correct response? Recent
research in psychometrics suggests that incorporating process data in assessments is
tenable. A transition network using weighted directed networks can capture activity
sequences, with nodes representing actions and directed links connecting two actions
only if the first action is followed by the second action in the sequence [
          <xref ref-type="bibr" rid="ref39">39</xref>
          ]. As for
scoring, Shu et al. [
          <xref ref-type="bibr" rid="ref26">26</xref>
          ] proposed a Markov-IRT model to characterize and capture the
unique features of students' individual response process during a problem-solving
activity in scenario-based tasks by laying out the model structure, its assumptions, the
parameter estimation and parameter space. The Markov-IRT model allows test
developers to determine the mapping of specific combinations to scoring rubrics.
        </p>
      </sec>
      <sec id="sec-4-3">
        <title>Implications for Summative Assessments</title>
        <p>Psychometric research can also contribute to scoring issues, particularly for high stakes
summative assessments, where assigning valid and reliable scores that reflect students'
skill mastery is a critical component. These assessments involve further issues such as
score discrimination between students, in that students who score higher have better
mastery than students with lower scores, and score comparability across cohorts of
students who take different versions of the assessments.</p>
        <p>Further, standard concerns in testing that are typically of lesser importance in
learning assessments will surface. Issues such as fairness in testing, item overexposure, the
establishing of cut scores, scaling and equating of scores, reporting and use of scores
have been extensively studied and will also need to be adapted for a caring assessment.
5</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>Conclusion</title>
      <p>We posit that caring assessments have a place in both formative and summative
assessments. To get there, we will require that researchers from diverse backgrounds, such as
computer science, engineering, natural language processing, learning, and
psychometrics, work closely together to make sure that any new caring assessment is as valid and
reliable as possible.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <surname>Baker</surname>
            ,
            <given-names>R.S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gowda</surname>
            ,
            <given-names>S.M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Wixon</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kalka</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Wagner</surname>
            ,
            <given-names>A.Z.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Salvi</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Aleven</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kusbit</surname>
            ,
            <given-names>G.W.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Ocumpaugh</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rossi</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          :
          <article-title>Towards sensor-free affect detection in cognitive tutor algebra</article-title>
          .
          <source>International Educational Data Mining Society</source>
          (
          <year>2012</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <given-names>Beigman</given-names>
            <surname>Klebanov</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            ,
            <surname>Burstein</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            ,
            <surname>Harackiewicz</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            ,
            <surname>Priniski</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            ,
            <surname>Mulholland</surname>
          </string-name>
          ,
          <string-name>
            <surname>M.</surname>
          </string-name>
          :
          <article-title>Enhancing stem motivation through personal and communal values: Nlp for assessment of utility value in student writing</article-title>
          .
          <source>In: Proceedings of the 11th Workshop on Innovative Use of NLP for Building Educational Applications</source>
          . pp.
          <fpage>199</fpage>
          -
          <lpage>205</lpage>
          . Association for Computational Linguistics, San Diego, CA (
          <year>June 2016</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <surname>Bolt</surname>
            ,
            <given-names>D.M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Cohen</surname>
            ,
            <given-names>A.S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Wollack</surname>
            ,
            <given-names>J.A.</given-names>
          </string-name>
          :
          <article-title>Item parameter estimation under conditions of test speededness: Application of a mixture Rasch model with ordinal constraints</article-title>
          .
          <source>Journal of Educational Measurement</source>
          <volume>39</volume>
          (
          <issue>4</issue>
          ),
          <fpage>331</fpage>
          -
          <lpage>348</lpage>
          (
          <year>2002</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <surname>Calvo</surname>
            ,
            <given-names>R.A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>D'Mello</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          :
          <article-title>Affect detection: An interdisciplinary review of models, methods, and their applications</article-title>
          .
          <source>IEEE Transactions on affective computing 1(1)</source>
          ,
          <fpage>18</fpage>
          -
          <lpage>37</lpage>
          (
          <year>2010</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <surname>Conati</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gertner</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Vanlehn</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          :
          <article-title>Using bayesian networks to manage uncertainty in student modeling. User modeling and user-adapted interaction 12(4</article-title>
          ),
          <fpage>371</fpage>
          -
          <lpage>417</lpage>
          (
          <year>2002</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <surname>von Davier</surname>
            ,
            <given-names>A.A.</given-names>
          </string-name>
          :
          <article-title>Computational psychometrics in support of collaborative educational assessments</article-title>
          .
          <source>Journal of Educational Measurement</source>
          <volume>54</volume>
          (
          <issue>1</issue>
          ),
          <fpage>3</fpage>
          -
          <lpage>11</lpage>
          (
          <year>2017</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          7.
          <string-name>
            <surname>Desmarais</surname>
            ,
            <given-names>M.C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Pu</surname>
            ,
            <given-names>X.:</given-names>
          </string-name>
          <article-title>A bayesian student model without hidden nodes and its comparison with item response theory</article-title>
          .
          <source>International Journal of Artificial Intelligence in Education</source>
          <volume>15</volume>
          (
          <issue>4</issue>
          ),
          <fpage>291</fpage>
          -
          <lpage>323</lpage>
          (
          <year>2005</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <given-names>D</given-names>
            <surname>'Mello</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.K.</given-names>
            ,
            <surname>Dowell</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            ,
            <surname>Graesser</surname>
          </string-name>
          ,
          <string-name>
            <surname>A.C.</surname>
          </string-name>
          :
          <article-title>Cohesion relationships in tutorial dialogue as predictors of affective states</article-title>
          .
          <source>In: AIED</source>
          . pp.
          <fpage>9</fpage>
          -
          <lpage>16</lpage>
          (
          <year>2009</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <surname>DMello</surname>
            ,
            <given-names>S.K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lehman</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Graesser</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          :
          <article-title>A motivationally supportive affect-sensitive autotutor</article-title>
          .
          <source>In: New perspectives on affect and learning technologies</source>
          , pp.
          <fpage>113</fpage>
          -
          <lpage>126</lpage>
          . Springer (
          <year>2011</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <surname>Flor</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Yoon</surname>
            ,
            <given-names>S.Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hao</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Liu</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>von Davier</surname>
          </string-name>
          , A.:
          <article-title>Automated classification of collaborative problem solving interactions in simulated science tasks</article-title>
          .
          <source>In: Proceedings of the 11th Workshop on Innovative Use of NLP for Building Educational Applications</source>
          . pp.
          <fpage>31</fpage>
          -
          <lpage>41</lpage>
          . Association for Computational Linguistics, San Diego, CA (
          <year>June 2016</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <string-name>
            <surname>Flor</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Yoon</surname>
            ,
            <given-names>S.Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hao</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Liu</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>von Davier</surname>
          </string-name>
          , A.:
          <article-title>Automated classification of collaborative problem solving interactions in simulated science tasks</article-title>
          .
          <source>In: Proceedings of the 11th Workshop on Innovative Use of NLP for Building Educational Applications</source>
          . pp.
          <fpage>31</fpage>
          -
          <lpage>41</lpage>
          . Association for Computational Linguistics, San Diego, CA (
          <year>June 2016</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          12.
          <string-name>
            <surname>Forbes-Riley</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Litman</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          :
          <article-title>Benefits and challenges of real-time uncertainty detection and adaptation in a spoken dialogue computer tutor</article-title>
          .
          <source>Speech Communication</source>
          <volume>53</volume>
          (
          <issue>9-10</issue>
          ),
          <fpage>1115</fpage>
          -
          <lpage>1136</lpage>
          (
          <year>2011</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          13.
          <string-name>
            <surname>Hernández</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Drasgow</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>González-Romá</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          :
          <article-title>Investigating the functioning of a middle category by means of a mixed-measurement model</article-title>
          .
          <source>Journal of Applied Psychology</source>
          <volume>89</volume>
          (
          <issue>4</issue>
          ),
          <volume>687</volume>
          (
          <year>2004</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          14.
          <string-name>
            <surname>Jeon</surname>
            ,
            <given-names>J.H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Yoon</surname>
          </string-name>
          , S.Y.:
          <article-title>Acoustic feature-based non-scorable response detection for an automated speaking proficiency assessment</article-title>
          .
          <source>In: Thirteenth Annual Conference of the International Speech Communication Association</source>
          (
          <year>2012</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          15.
          <string-name>
            <surname>Lee</surname>
            ,
            <given-names>Y.H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Haberman</surname>
            ,
            <given-names>S.J.:</given-names>
          </string-name>
          <article-title>Investigating test-taking behaviors using timing and process data</article-title>
          .
          <source>International Journal of Testing</source>
          <volume>16</volume>
          (
          <issue>3</issue>
          ),
          <fpage>240</fpage>
          -
          <lpage>267</lpage>
          (
          <year>2016</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          16.
          <string-name>
            <surname>Lee</surname>
            ,
            <given-names>Y.H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Jia</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          :
          <article-title>Using response time to investigate students' test-taking behaviors in a naep computer-based study</article-title>
          .
          <source>Large-scale Assessments in Education</source>
          <volume>2</volume>
          (
          <issue>1</issue>
          ),
          <volume>8</volume>
          (
          <year>2014</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          17.
          <string-name>
            <surname>Lehman</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>DMello</surname>
          </string-name>
          , S.K.:
          <article-title>Predicting student affect through textual features during expert tutoring sessions. Presented at the annual meeting of the Society for Text and Discourse (</article-title>
          <year>2010</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          18.
          <string-name>
            <surname>Lehman</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zapata-Riveria</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          :
          <article-title>Student Emotions in Conversation-Based Assessments</article-title>
          .
          <source>IEEE Transactions on Learning Technologies (In Print)</source>
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          19.
          <string-name>
            <surname>Lord</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          :
          <article-title>Application of item response theory to practical testing problems</article-title>
          . Hillsdale, NJ, Lawrence Erlbaum Ass (
          <year>1980</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          20.
          <string-name>
            <surname>Madnani</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Cahill</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Riordan</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          :
          <article-title>Automatically scoring tests of proficiency in music instruction</article-title>
          .
          <source>In: Proceedings of the 11th Workshop on Innovative Use of NLP for Building Educational Applications</source>
          . pp.
          <fpage>217</fpage>
          -
          <lpage>222</lpage>
          (
          <year>2016</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          21.
          <string-name>
            <surname>Mayer</surname>
          </string-name>
          , R.E.:
          <article-title>Unique contributions of eye-tracking research to the study of learning with graphics</article-title>
          .
          <source>Learning and instruction 20(2)</source>
          ,
          <fpage>167</fpage>
          -
          <lpage>171</lpage>
          (
          <year>2010</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          22.
          <string-name>
            <surname>Mislevy</surname>
            ,
            <given-names>R.J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gitomer</surname>
            ,
            <given-names>D.H.</given-names>
          </string-name>
          :
          <article-title>The role of probability-based inference in an intelligent tutoring system</article-title>
          .
          <source>ETS Research Report Series 1995(2)</source>
          (
          <year>1995</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          23.
          <string-name>
            <surname>Mislevy</surname>
            ,
            <given-names>R.J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Verhelst</surname>
          </string-name>
          ,N.:
          <article-title>Modeling item responses when different subjects employ different solution strategies</article-title>
          .
          <source>ETS Research Report Series 1987(2)</source>
          (
          <year>1987</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          24.
          <string-name>
            <surname>Rost</surname>
            ,
            <given-names>J.:</given-names>
          </string-name>
          <article-title>A logistic mixture distribution model for polychotomous item responses</article-title>
          .
          <source>British Journal of Mathematical and Statistical Psychology</source>
          <volume>44</volume>
          (
          <issue>1</issue>
          ),
          <fpage>75</fpage>
          -
          <lpage>92</lpage>
          (
          <year>1991</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref25">
        <mixed-citation>
          25.
          <string-name>
            <surname>Rozin</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Cohen</surname>
            ,
            <given-names>A.B.</given-names>
          </string-name>
          :
          <article-title>High frequency of facial expressions corresponding to confusion, concentration, and worry in an analysis of naturally occurring facial ex- pressions of americans</article-title>
          .
          <source>Emotion</source>
          <volume>3</volume>
          (
          <issue>1</issue>
          ),
          <volume>68</volume>
          (
          <year>2003</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref26">
        <mixed-citation>
          26.
          <string-name>
            <surname>Shu</surname>
            ,
            <given-names>Z.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bergner</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zhu</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hao</surname>
          </string-name>
          , J.,
          <string-name>
            <surname>von Davier</surname>
            ,
            <given-names>A.A.</given-names>
          </string-name>
          :
          <article-title>An Item Response Theory Analysis of Problem-Solving Processes in Scenario-Based Tasks</article-title>
          .
          <source>Psychological Test and Assessment Modeling</source>
          <volume>59</volume>
          (
          <issue>1</issue>
          ),
          <volume>109</volume>
          (
          <year>2017</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref27">
        <mixed-citation>
          27.
          <string-name>
            <surname>Shute</surname>
            ,
            <given-names>V.J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hansen</surname>
            ,
            <given-names>E.G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Almond</surname>
          </string-name>
          , R.G.:
          <article-title>You can't fatten a hog by weighing it-or can you? evaluating an assessment for learning system called aced</article-title>
          .
          <source>International Journal of Artificial Intelligence in Education</source>
          <volume>18</volume>
          (
          <issue>4</issue>
          ),
          <fpage>289</fpage>
          -
          <lpage>316</lpage>
          (
          <year>2008</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref28">
        <mixed-citation>
          28.
          <string-name>
            <surname>Tai</surname>
            ,
            <given-names>R.H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Loehr</surname>
            ,
            <given-names>J.F.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Brigham</surname>
            ,
            <given-names>F.J.:</given-names>
          </string-name>
          <article-title>An exploration of the use of eye-gaze tracking to study problem-solving on standardized science assessments</article-title>
          .
          <source>International journal of research &amp; method in education 29(2)</source>
          ,
          <fpage>185</fpage>
          -
          <lpage>208</lpage>
          (
          <year>2006</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref29">
        <mixed-citation>
          29.
          <string-name>
            <surname>Timms</surname>
            ,
            <given-names>M.J.:</given-names>
          </string-name>
          <article-title>Using item response theory (IRT) to select hints in an ITS</article-title>
          .
          <source>Frontiers in Artificial Intelligence and Applications</source>
          <volume>158</volume>
          ,
          <issue>213</issue>
          (
          <year>2007</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref30">
        <mixed-citation>
          30.
          <string-name>
            <surname>Tsai</surname>
            ,
            <given-names>M.J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hou</surname>
            ,
            <given-names>H.T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lai</surname>
            ,
            <given-names>M.L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Liu</surname>
            ,
            <given-names>W.Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Yang</surname>
            ,
            <given-names>F.Y.</given-names>
          </string-name>
          :
          <article-title>Visual attention for solving multiple-choice science problem: An eye-tracking analysis</article-title>
          .
          <source>Computers &amp; Education</source>
          <volume>58</volume>
          (
          <issue>1</issue>
          ),
          <fpage>375</fpage>
          -
          <lpage>385</lpage>
          (
          <year>2012</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref31">
        <mixed-citation>
          31.
          <string-name>
            <surname>Valderas</surname>
            ,
            <given-names>M.T.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bolea</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Laguna</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Vallverdú</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bailón</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          :
          <article-title>Human emotion recognition using heart rate variability analysis with spectral bands based on respiration</article-title>
          .
          <source>In: Engineering in Medicine and Biology Society (EMBC)</source>
          ,
          <year>2015</year>
          37th Annual International Conference of the IEEE. pp.
          <fpage>6134</fpage>
          -
          <lpage>6137</lpage>
          . IEEE (
          <year>2015</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref32">
        <mixed-citation>
          32.
          <string-name>
            <surname>Wang</surname>
            ,
            <given-names>Z.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Sabatini</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          , OReilly, T.,
          <string-name>
            <surname>Feng</surname>
          </string-name>
          , G.:
          <article-title>How individual differences interact with task demands in text processing</article-title>
          .
          <source>Scientific Studies of Reading</source>
          <volume>21</volume>
          (
          <issue>2</issue>
          ),
          <fpage>165</fpage>
          -
          <lpage>178</lpage>
          (
          <year>2017</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref33">
        <mixed-citation>
          33.
          <string-name>
            <surname>White</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Kim</surname>
            ,
            <given-names>Y.Y.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chen</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Liu</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          :
          <article-title>Performance of Fourth-Grade Students in the 2012 NAEP Computer-Based Writing Pilot Assessment: Scores, Text Length, and Use of Editing Tools</article-title>
          . Working Paper Series.
          <source>NCES</source>
          <year>2015</year>
          -
          <volume>119</volume>
          . National Center for Education Statistics (
          <year>2015</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref34">
        <mixed-citation>
          34.
          <string-name>
            <surname>Wise</surname>
            ,
            <given-names>S.L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Bhola</surname>
            ,
            <given-names>D.S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Yang</surname>
          </string-name>
          , S.T.:
          <article-title>Taking the Time to Improve the Validity of Low-Stakes Tests: The Effort-Monitoring CBT</article-title>
          .
          <source>Educational Measurement: Issues and Practice</source>
          <volume>25</volume>
          (
          <issue>2</issue>
          ),
          <fpage>21</fpage>
          -
          <lpage>30</lpage>
          (
          <year>2006</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref35">
        <mixed-citation>
          35.
          <string-name>
            <surname>Yang</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Alsadoon</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Prasad</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Singh</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Elchouemi</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          :
          <article-title>An Emotion Recognition Model Based on Facial Recognition in Virtual Learning Environment</article-title>
          .
          <source>Procedia Computer Science</source>
          <volume>125</volume>
          ,
          <fpage>2</fpage>
          -
          <lpage>10</lpage>
          (
          <year>2018</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref36">
        <mixed-citation>
          36.
          <string-name>
            <surname>Yu</surname>
            ,
            <given-names>S.N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chen</surname>
            ,
            <given-names>S.F.</given-names>
          </string-name>
          :
          <article-title>Emotion state identification based on heart rate variability and genetic algorithm</article-title>
          .
          <source>In: Engineering in Medicine and Biology Society (EMBC)</source>
          ,
          <year>2015</year>
          37th Annual International Conference of the IEEE. pp.
          <fpage>538</fpage>
          -
          <lpage>541</lpage>
          . IEEE (
          <year>2015</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref37">
        <mixed-citation>
          37.
          <string-name>
            <surname>Zapata-Rivera</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          :
          <article-title>Toward caring assessment systems</article-title>
          .
          <source>In: Adjunct Publication of the 25th Conference on User Modeling, Adaptation and Personalization</source>
          . pp.
          <fpage>97</fpage>
          -
          <lpage>100</lpage>
          . ACM (
          <year>2017</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref38">
        <mixed-citation>
          38.
          <string-name>
            <surname>Zapata-Rivera</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Liu</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Chen</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hao</surname>
          </string-name>
          , J.,
          <string-name>
            <surname>von Davier</surname>
            ,
            <given-names>A.A.</given-names>
          </string-name>
          :
          <article-title>Assessing science inquiry skills in an immersive, conversation-based scenario</article-title>
          . In:
          <article-title>Big data and learning analytics in higher education</article-title>
          , pp.
          <fpage>237</fpage>
          -
          <lpage>252</lpage>
          . Springer (
          <year>2017</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref39">
        <mixed-citation>
          39.
          <string-name>
            <surname>Zhu</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Shu</surname>
            ,
            <given-names>Z.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Davier</surname>
            ,
            <given-names>A.A.</given-names>
          </string-name>
          :
          <article-title>Using Networks to Visualize and Analyze Process Data for Educational Assessment</article-title>
          .
          <source>Journal of Educational Measurement</source>
          <volume>53</volume>
          (
          <issue>2</issue>
          ),
          <fpage>190</fpage>
          -
          <lpage>211</lpage>
          (
          <year>2016</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref40">
        <mixed-citation>
          40.
          <string-name>
            <surname>Zickar</surname>
            ,
            <given-names>M.J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gibby</surname>
            ,
            <given-names>R.E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Robie</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          :
          <article-title>Uncovering faking samples in applicant, incumbent, and experimental data sets: An application of mixed-model item response theory</article-title>
          .
          <source>Organizational Research Methods</source>
          <volume>7</volume>
          (
          <issue>2</issue>
          ),
          <fpage>168</fpage>
          -
          <lpage>190</lpage>
          (
          <year>2004</year>
          )
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>