<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Understanding Students' Problem-Solving Processes via Action Sequence Analyses</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Ruhan Circi</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Juanita Hicks</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>American Institutes for Research</institution>
        </aff>
      </contrib-group>
      <abstract>
        <p>The transition of the National Assessment of Educational Progress (NAEP) to digitally based assessments (DBAs) allowed for the collection of data that can provide insights into students' problemsolving processes. When students interact with a NAEP DBA item, their recorded timestamped events in the process data form sequences. We refer to action sequences as the series of clicks or other actions a student makes within an item. Using data from one released block of the NAEP 2017 mathematics assessment for grade 4, this study aims to provide an understanding of the relationships among action sequence characteristics, item characteristics and student performance. First, we extract individual actions sequences across items. Second, we categorize each individual action into one of four activities: Browsing, Passive investigation, Active investigation, or Decision. This categorization enables us to investigate sequence patterns within and across different items. Sequence characteristics are summarized from two perspectives: a) the time spent on each activity is calculated for each student across items and b) the within-sequence entropy (Shannon, 1948) and turbulence (Elzinga, 2006) of the sequences are calculated to quantify students' action mobility. We found that the time students spend on “Decision” and “Passive investigation” activities can be used to predict student performance.</p>
      </abstract>
      <kwd-group>
        <kwd>NAEP</kwd>
        <kwd>Process data</kwd>
        <kwd>Digitally based assessments</kwd>
        <kwd>sequence mining</kwd>
        <kwd>action sequences</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. BACKGROUND</title>
      <p>
        In 2017 the National Assessment of Educational Progress (NAEP)
transitioned from paper-based assessments (PBAs) to digitally
based assessments (DBAs). DBAs allow us to capture student
interactions with the test screen that are recorded as timestamped
events. These records form data known as process data.
It has become commonplace to include response time (RT) in
addition to responses in the psychometric models to account for
speed and accuracy
        <xref ref-type="bibr" rid="ref3">(e.g., Goldhammer, 2015)</xref>
        , and to examine the
relationship between response time and item- and person-level
factors
        <xref ref-type="bibr" rid="ref9">(e.g., Masters, Schnipke, &amp; Connor, 2005)</xref>
        . Response time
is used to examine psychometric quality of items and students’
test-taking behaviors and it is concluded to be promising for
various assessment elements. Yet, the process data contains richer
information such as actions that students use during their
problem-solving processes and the allocation of the time students
spend on particular activities within a single response time
remained unexplored.
      </p>
      <p>When students interact with a NAEP DBA item, their recorded
timestamped events in the process data form sequences. These
sequences contain information about the order, mobility, and
duration of the tasks students take throughout the problem-solving
process and may shed light on the processes underlying the
students’ test-taking behaviors. In this study, we divide response
times into subcategories using the action definitions to provide a
more meaningful understanding of student test taking behavior
and examine the differences across item types and student
performance.</p>
    </sec>
    <sec id="sec-2">
      <title>1.1 Literature</title>
      <p>
        Process data is most commonly used to calculate response time
(RT), defined as the time an examinee takes to complete an item
or assessment. Due to the association of RT with psychological
and cognitive processes
        <xref ref-type="bibr" rid="ref4">(e.g., Huff &amp; Sireci, 2001)</xref>
        , RT is often
used to make decisions such as setting assessment time limits
        <xref ref-type="bibr" rid="ref11">(e.g., van der Linden, 2011)</xref>
        and capturing aberrant test-taking
behaviors
        <xref ref-type="bibr" rid="ref8">(e.g., Marianti, Fox, Avetisyan, Veldkamp, &amp; Tijmstra,
2014)</xref>
        .
      </p>
      <p>
        However, RT alone may not provide sufficient information to
draw inferences about the processes underlying students’ test
taking behaviors
        <xref ref-type="bibr" rid="ref5">(Lee &amp; Haberman, 2016)</xref>
        . In fact, RT could
consist of the time for various components in the problem-solving
process such as preparation (e.g., forming a response plan) and
writing down/typing the response. The decomposition of RT can
differ depending on item types
        <xref ref-type="bibr" rid="ref7">(e.g., Li, Banerjee, &amp; Zumbo,
2017)</xref>
        . Thus, to ensure the validity of inferences drawn from RTs,
it is necessary to understand what students actually do throughout
the RT.
      </p>
    </sec>
    <sec id="sec-3">
      <title>2. CURRENT STUDY</title>
      <p>
        In the assessment setting, RT could consist of the times for
various components in the problem-solving process such as
preparation (e.g., forming a response plan) and writing
down/typing the response. The decomposition of RT can be
different for different item types
        <xref ref-type="bibr" rid="ref7">(e.g., Li, Banerjee, &amp; Zumbo,
2017)</xref>
        . Since the NAEP mathematics assessment consists of items
with a mix of item types (e.g., multiple choice, constructed
response), using the decomposition of RT for different tasks (e.g.,
investigation, decision) rather than total RT could be helpful when
different decisions (e.g., setting assessment time limits, capturing
aberrant test-taking behaviors) are to be made based on the time.
A more fine-grained understanding of the relationships among
RTs and students’ problem-solving behaviors can be gained by
analyzing students’ action sequences, which can further improve
the usefulness of RT in psychometric research (e.g., determining
non-response categories such as omit and not reach).
      </p>
      <p>
        The goals of this study are: a) identify and describe the action
sequences of students in a meaningful way, b) examine mobility
across actions, c) differentiate profiles of action sequences, and d)
explore students’ performance in connection to sequence clusters.
Steps taken for current project can be presented as follows: First,
individual actions are extracted. Second, students’ response
processes are represented as sequences consisting of four tasks,
i.e., Browsing, Passive investigation, Active investigation, and
Decision (See definitions in Table 1). Since the variation across
time for individual actions can be very large, we decided to use a
set cut off point (2 seconds) for defining each action. In the end,
we recoded the sequence of student actions in these groups for
further analyses (See Figure 1 for an example). Then, the
characteristics of the sequences are summarized from two
perspectives: a) the time spent on each task is calculated for each
student, which allows the decomposition of the RT, and b) the
within-sequence entropy
        <xref ref-type="bibr" rid="ref10">(Shannon, 1948)</xref>
        and turbulence
        <xref ref-type="bibr" rid="ref2">(Elzinga, 2006)</xref>
        of the sequences are calculated to quantify
students’ action mobility.
In addition to summarizing sequence characteristics in a
descriptive manner, this study examines the relationships among
the sequence characteristics, item characteristics and students’
item responses. Specifically, to examine the relationship between
sequence characteristics and item characteristics, the RT
decomposition and students’ action mobility are compared across
different items. Furthermore, representative sequence(s) are
identified for each item with the use of a sequence dissimilarity
measure and a clustering algorithm. The representative
sequence(s) can inform the typical response process of an item.
Finally, to examine the relationship between sequence
characteristics and student performance, sequence characteristics,
such as the time duration of each task, within-sequence entropy
and turbulence, are used as features to predict students’ item
scores. The results could inform which feature(s) of the sequences
best contribute to correct/incorrect item responses or the
presence/absence of the responses. Moreover, the score
distributions are compared across sequence clusters.
      </p>
      <p>In sum, this study, by decomposing RT and examining the
relationships among the sequence characteristics, item
characteristics and student performance, aims to inform more
meaningful ways of calculating RT (e.g., different ways of RT
calculation for different items) and the validity of score categories
such as “omit” and “not reach”. For instance, if the sequences of
students who were scored as “not reach” were found to contain
some actions that are related to making responses (i.e., the
“Decision” actions), the scores of these students may be
considered as “omit” as opposed to “not reach”.</p>
    </sec>
    <sec id="sec-4">
      <title>2.1 Research Questions</title>
      <p>Specifically, the following research questions are examined in the
current study:
RQ1. What actions do students take and what are the
characteristics of the action sequences (mobility, time
distribution) throughout the RTs of the NEAP math items?
RQ2. How do students’ action sequences differ across different
item types (e.g., multiple-choice item, constructed-response
item)?
RQ3. Which action sequence characteristic(s) best predict the
item scores?</p>
    </sec>
    <sec id="sec-5">
      <title>3. DATA</title>
      <p>We used data from one of the released blocks from NAEP 2017
Grade 4 Mathematics assessment. One of the released blocks
includes 29,100 4th graders in both public and private schools and
consists of 14 cognitive items. The sample was collected using the
conventional NAEP sampling procedures, i.e., a two-stage
stratified random sampling design with schools selected in the
first stage and students in the second stage. In the data cleaning
procedure, students with accommodation or interruptions were
excluded. Comparisons of the demographic composition of the
two samples, full sample and analytical sample, are presented in
Table 2.
proficiency or in individualized education program, the results for
these variables are not included. Detail may not sum to totals
because of rounding.</p>
      <p>A small non-significant difference in the proportion of White
(50.5 % in analytical, and 48.9% in full sample) and Hispanic
students (24.4% in analytical and 25.9% in full sample) are
observed. A significant difference in term of NSLP non-eligible
category is found (46% vs. 43.8%).</p>
    </sec>
    <sec id="sec-6">
      <title>4. ANALYSIS</title>
      <p>To construct sequences and decompose RT from the process data,
we followed two steps (See Figure 1 for a demonstration of the
procedure): a) Recoding the actions into four task categories (i.e.,
Browsing, Passive investigation, Active investigation, Decision;
See definitions in Table 1); and b) Calculating the time duration
of each task. Thus, students’ item response processes were
represented as sequences whose lengths are proportional to the
time durations. Since the variation of time students spend on an
item can be large (i.e., range from 0.01 second to 30 minutes),
using a small-time unit (e.g., 0.01 second) could result in
extremely long sequences that exceed the computer memory
capacity. Therefore, we decided to use 2 seconds as the time unit
while constructing the sequences. Only actions in students’ initial
item visit (i.e., actions between the first pair of “Enter Item” and
“Exit Item” actions) were included in the sequence. Students
whose initial item visit lasts longer than 8 minutes (480 seconds)
were excluded from the analyses to avoid extremely long
sequences. For all the items in the MA block, the percentages of
students with initial item visit longer than 8 minutes are lower
than 1%.</p>
      <p>
        The mean time spent on each action as well as the action mobility
were summarized as the sequence characteristics. The number of
task transitions, Shannon entropy
        <xref ref-type="bibr" rid="ref10">(Shannon, 1948)</xref>
        and turbulence
        <xref ref-type="bibr" rid="ref2">(Elzinga, 2006)</xref>
        measures were used to quantify the action
mobility.
      </p>
      <p>
        To examine how students’ action sequences differ across different
item types (e.g., multiple-choice item, constructed-response item),
the characteristics of sequences were summarized and compared
across different items. To identify the typical response process for
an item, the hierarchical agglomerative clustering algorithm was
applied to all the students’ sequences based on the optimal
matching edit distance
        <xref ref-type="bibr" rid="ref6">(Levenshtein, 1966)</xref>
        matrix. The medoids
of the clusters (i.e., the sequence that is the nearest to the virtual
center of the cluster) were treated as the representative sequences
that represent the typical response processes for an item. As no
study to our knowledge has been done to determine the optimal
number of clusters when the clustering is based on the edit
distance matrix. Ward’s algorithm was used to form clusters by
maximizing within cluster homogeneity. We chose the number of
clusters by visually inspecting the dendrogram and assessing the
interpretability of the clusters. Specifically, for each item, we
examined the cluster medoids when the number of clusters ranged
from 2 to 4 and chose the number of clusters that resulted in
interpretable clusters from practical perspectives. All sequence
analyses were performed using the TraMineR R package.
To examine the relationship between the sequence characteristics
and student performance, the sequence characteristics were used
as features to predict the item scores using the regression tree
        <xref ref-type="bibr" rid="ref1">(Breiman, 2017)</xref>
        . In addition, the score distributions were
compared across the sequence clusters identified based on the
hierarchical clustering algorithm and edit distance.
      </p>
    </sec>
    <sec id="sec-7">
      <title>5. RESULTS</title>
      <p>For the purposes of this paper, we present the results for two
selected items1 listed in Table 3. The items are different item
types (Item A is multiple-choice item while Item B is constructed
response item) and are close in the presentation order. Thus, the
two items were chosen to demonstrate the difference in sequence
characteristics between items of different types (with minimal
confounding of the presentation order).</p>
    </sec>
    <sec id="sec-8">
      <title>5.1 Response Time Decomposition</title>
      <p>The average time students spent on each recoded behavior actions,
i.e., browsing, passive investigation, active investigation, and
1 https://nces.ed.gov/NationsReportCard/nqt/Search
decision are shown in Figure 2. For Item A, the “decision” task
had the highest average time among the four tasks; however, for
Item B, the “passive investigation” task had the highest time. On
average, students spent 10 seconds browsing Item A by executing
scroll on the screen, while students hardly spent any time
browsing Item B. Such difference in the browsing time could be
associated with the content of the items: Item A needs to be
solved by inspecting and comparing the heights of the trees which
may result in browsing actions, while Item B is a straightforward
computational item which may not require much browsing.
Table 4 lists the summary statistics of three mobility measures,
i.e., the number of task transitions, within-sequence entropy, and
turbulence. Task transition refers to switching among the four
tasks (i.e., browsing, passive investigation, active investigation,
and decision) in the sequence. The average task transitions for
item A and B are 2.28 and 2.13, respectively. As for the
withinsequence entropy and turbulence measures, higher values indicate
larger mobility. On average, item A is found to have higher
within-sequence entropy and turbulence.</p>
    </sec>
    <sec id="sec-9">
      <title>5.2 Sequence Characteristics</title>
      <p>Figure 3 presents the state distributions at each time unit for the
two selected items. Each unit of the x-axis represents 2 seconds.
For instance, in the first 2 seconds, students who conducted
“passive investigation” make up the largest proportion in both
items. When responding to Item A, more than 10% of the students
were browsing the item in the first 2 seconds; when interacting
with Item B, nearly no students browsed the item in this time unit.</p>
    </sec>
    <sec id="sec-10">
      <title>5.3 Typical Response Process</title>
      <p>Figure 4 shows the representative sequences for Item A and Item
B. A representative sequence refers to the sequence with the
smallest sum of edit distance to the rest of sequences; the
representative sequence is considered to be representative of the
typical response process of an item. As the sequence length is
proportional to the time duration, the overall time duration of the
typical response process is shorter for Item A than Item B. We
observe that, in the typical response process;
• for Item A, the student conducts passive investigation,
browses the item, and makes response decisions,
sequentially.
• for Item B, the student conducts passive and active
investigations and makes response decisions.
While identifying a single typical response process for an item is
desirable for the purpose of interpretation, a single sequence may
not be enough to represent all the sequences. It is possible that
there are multiple response process archetypes for an item. Thus,
we conducted hierarchical agglomerative clustering based on the
edit distance matrix. In the clustering process, each unit is a
student. After examining the dendrogram and the interpretability
of the clusters, we chose to retain three clusters (labeled as Type
1, Type 2, and Type 3). The weighted cluster sizes and students’
demographic characteristics by sequence clusters found in Item B
are presented in Table 5.</p>
    </sec>
    <sec id="sec-11">
      <title>5.4 Relationship Between the Sequence</title>
    </sec>
    <sec id="sec-12">
      <title>Characteristics and Student Performance</title>
      <p>Figure 6 shows the regression tree learned from the process data
of Item B. Time durations of browsing, passive investigation,
active investigation and decision, number of task transitions,
within-sequence entropy and turbulence are used to predict item
scores. Item B has five score categories, i.e., incorrect, correct, off
task, omitted, and not reached. Each box in Figure 6 is called a
“node” and the five decimals in each box are the predicted
proportions of students having the five score categories in that
node. The name (and color) of the node is determined by the score
category that has the highest proportion among the five categories.
For example, as the first split was performed with the decision
time, for students with decision time longer than 14 seconds (25%
of the students in the sample have decision time longer than 14
seconds), the predicted proportions of getting “incorrect” and
“correct” scores are 0.70 and 0.28, respectively. In addition, all
the splits in this regression tree are performed with either decision
or passive investigation time durations.</p>
    </sec>
    <sec id="sec-13">
      <title>5.5 Relationship between the Sequence</title>
    </sec>
    <sec id="sec-14">
      <title>Cluster and Student Performance</title>
      <p>Table 6 lists the score distribution of scores within each sequence
cluster found in Item B. Item B is a fill-in-blank item, which is the
fourth item in the block with a difficulty level of 0.29. In all three
clusters, the proportion of students getting “incorrect” score was
the highest among the five score categories. The similarity in the
score distributions across sequence clusters implies that no clear
pattern on the performance difference has been found among
students with different response process archetypes.</p>
      <p>Off
task
0.2
0.1
0.2</p>
    </sec>
    <sec id="sec-15">
      <title>6. DISCUSSION</title>
    </sec>
    <sec id="sec-16">
      <title>6.1 Summary</title>
      <p>In summary, this study provided insights into the decomposition
of RT by constructing action sequences from students’ process
data. In particular, the action sequences contained information of
the time duration, order and mobility of the tasks students
executed to solve the NAEP mathematics items. By presenting the
sequences of two selected NAEP released items as examples, this
paper demonstrated the differences in RT decomposition and
typical response processes between items of different types (i.e., a
multiple-choice item vs a fill-in-blank item). This methodology
and set of results suggest that examining action sequences and RT
decomposition can be a useful way to mine process data and
uncover educational processes. Also, action sequence mining can
be useful to analyze high variance data such as process data.
Response process archetypes were found by conducting a
hierarchical clustering algorithm using the edit distance matrix of
students’ action sequences. As for the relationship between
student performance and sequence characteristics, the time
students spent on “Decision” and “Passive investigation” were
incorporated in the learned regression tree of the example
fill-inblank item, meaning that these components of RT can be used to
predict the scores of this item. Further, among the 10,000 students
who correctly responded to Item B, 48.6% had their action
sequences clustered into Type 1, 26.3% into Type 2 and 25.2%
into Type 3, which implied that students who responded to the
item correctly may have different response processes.</p>
    </sec>
    <sec id="sec-17">
      <title>6.2 Limitations and Future Research</title>
      <p>As an initial exploration of action sequences in the NAEP
mathematics items, this study has limitations and opens up
opportunities for future research. First, the actions were
categorized into four tasks (browsing, passive investigation, active
investigation and decision) in this study. However, this may not
be the only way to categorize the actions. For instance, in a
multiple-choice item, the actions could be recoded based on
students’ selected options. Thus, sequences that reflect students’
trajectory of answer changes can be constructed.</p>
      <p>Second, the number of clusters was determined only based on the
dendrogram and the interpretability of the clusters in this study.
To better justify the choice of the number of clusters, future
studies could develop quantitative measures to determine the
optimal number of clusters based on the edit distance matrix.
Finally, this study only included a limited number of sequence
characteristics as features to learn the regression tree. Other
features such as the frequencies of subsequences (e.g., the
frequency of a student switching from passive investigation to
active investigation and then to decision), together with feature
selection algorithms, could be incorporated in future studies.</p>
    </sec>
    <sec id="sec-18">
      <title>7. REFERENCES</title>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <surname>Breiman</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          (
          <year>2017</year>
          ).
          <article-title>Classification and regression trees</article-title>
          .
          <source>Routledge.</source>
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <surname>Elzinga</surname>
            ,
            <given-names>C. H.</given-names>
          </string-name>
          (
          <year>2006</year>
          ).
          <article-title>Turbulence in categorical time series</article-title>
          .
          <source>Mathematical Population Studies.</source>
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <surname>Goldhammer</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          (
          <year>2015</year>
          ).
          <article-title>Measuring ability, speed, or both? Challenges, psychometric solutions, and what can be gained from experimental control</article-title>
          .
          <source>Measurement: Interdisciplinary Research and Perspectives</source>
          ,
          <volume>13</volume>
          (
          <issue>3-4</issue>
          ),
          <fpage>133</fpage>
          -
          <lpage>164</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <surname>Huff</surname>
            ,
            <given-names>K. L.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>Sireci</surname>
            ,
            <given-names>S. G.</given-names>
          </string-name>
          (
          <year>2001</year>
          ).
          <article-title>Validity issues in computer-based testing</article-title>
          .
          <source>Educational Measurement: Issues and Practice</source>
          ,
          <volume>20</volume>
          (
          <issue>3</issue>
          ),
          <fpage>16</fpage>
          -
          <lpage>25</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <surname>Lee</surname>
            ,
            <given-names>Y.-H.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>Haberman</surname>
            ,
            <given-names>S. J.</given-names>
          </string-name>
          (
          <year>2016</year>
          ).
          <article-title>Investigating testtaking behaviors using timing and process data</article-title>
          .
          <source>International Journal of Testing</source>
          ,
          <volume>16</volume>
          (
          <issue>3</issue>
          ),
          <fpage>240</fpage>
          -
          <lpage>267</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <surname>Levenshtein</surname>
            ,
            <given-names>V. I.</given-names>
          </string-name>
          (
          <year>1966</year>
          ).
          <article-title>Binary codes capable of correcting deletions, insertions, and reversals</article-title>
          .
          <source>Soviet Physics Doklady</source>
          ,
          <volume>10</volume>
          (
          <issue>8</issue>
          ),
          <fpage>707</fpage>
          -
          <lpage>710</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <surname>Li</surname>
            ,
            <given-names>Z.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Banerjee</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>Zumbo</surname>
            ,
            <given-names>B. D.</given-names>
          </string-name>
          (
          <year>2017</year>
          ).
          <article-title>Response time data as validity evidence: Has it lived up to its promise and, if not, what would it take to do so</article-title>
          . In B.D.
          <string-name>
            <surname>Zumbo</surname>
          </string-name>
          &amp;
          <string-name>
            <surname>A.M. Hubley</surname>
          </string-name>
          (Eds.),
          <source>Understanding and Investigating Response Processes in Validation Research</source>
          (pp.
          <fpage>159</fpage>
          -
          <lpage>177</lpage>
          ). Springer.
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <surname>Marianti</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Fox</surname>
            ,
            <given-names>J.-P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Avetisyan</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Veldkamp</surname>
            ,
            <given-names>B. P.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>Tijmstra</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          (
          <year>2014</year>
          ).
          <article-title>Testing for aberrant behavior in response time modeling</article-title>
          .
          <source>Journal of Educational and Behavioral Statistics</source>
          ,
          <volume>39</volume>
          (
          <issue>6</issue>
          ),
          <fpage>426</fpage>
          -
          <lpage>451</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <surname>Masters</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Schnipke</surname>
            ,
            <given-names>D. L.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>Connor</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          (
          <year>2005</year>
          ).
          <article-title>Comparing item response times and difficulty for calculation items</article-title>
          .
          <source>In annual meeting of the American Educational Research Association</source>
          , Montreal, Canada.
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <surname>Shannon</surname>
            ,
            <given-names>C. E.</given-names>
          </string-name>
          (
          <year>1948</year>
          ).
          <article-title>A mathematical theory of communication</article-title>
          .
          <source>Bell System Technical Journal</source>
          ,
          <volume>27</volume>
          (
          <issue>3</issue>
          ),
          <fpage>379</fpage>
          -
          <lpage>423</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <surname>van der Linden</surname>
          </string-name>
          , W. J. (
          <year>2011</year>
          ).
          <article-title>Setting time limits on tests</article-title>
          .
          <source>Applied Psychological Measurement</source>
          ,
          <volume>35</volume>
          (
          <issue>3</issue>
          ),
          <fpage>183</fpage>
          -
          <lpage>199</lpage>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>