<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Using POMDPs to Forecast Kindergarten Students Reading Comprehension</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Russell G. Almond</string-name>
          <email>ralmond@fsu.edu</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Umit Tokac</string-name>
          <email>ut08@my.fsu.edu</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Stephanie Al Otaiba</string-name>
          <email>salotaiba@smu.edu</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Department of Teaching and Learning, Southern Methodist University</institution>
          ,
          <addr-line>Dallas, TX 75275</addr-line>
          ,
          <country country="US">USA</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Educational Psychology and</institution>
          ,
          <addr-line>Lerning Systems</addr-line>
          ,
          <institution>Flordia State University</institution>
          ,
          <addr-line>Tallahassee, FL 32306</addr-line>
          ,
          <country country="US">USA</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>Summative assessment of student abilities typically comes at the end of the instructional period, too late for educators to use the information for planning instruction. This paper explores the possibility of using Hierarchical Linear Models to forecast students end of year performance. Because these models are closely related to partially observed Markov decision processes (POMDPs), these should support extensions to instructional planning to meet educational goals. Despite the new notation, the POMDP models are subject to a familiar problem from the educational context: scale identi ability. This paper describes how this problem manifests itself and looks at one potential solution.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>INTRODUCTION</title>
      <p>
        There is a long tradition in education of separating
instruction and assessment: summative assessment
of what a student learns comes at the end of the
unit/semester/year. As limited time is allocated for
assessment, such assessments are typically limited in
their reliability (accuracy of measurement) and
content validity (coverage of the targeted knowledge, skills
and ability). Because summative assessment comes
at the end of instructions, instructors are not able to
make changes to their instructions to maximize
student learning
        <xref ref-type="bibr" rid="ref3">(Almond, 2010)</xref>
        .
      </p>
      <p>
        <xref ref-type="bibr" rid="ref6">Bennett (2007)</xref>
        suggested breaking the summative
assessment into four or six periodic assessments. First,
spreading the cost (student time taken away from
direct instruction) over multiple measurement occasions
allows for longer testing providing both greater content
      </p>
      <p>Some of the work took place while she was at the
Florida Center for Reading Research, Tallahassee, FL</p>
      <p>
        Almond (2007) noted that the forecasting could be
done using a partially observed Markov decision
process
        <xref ref-type="bibr" rid="ref8">(POMDP; Boutilier, Dean, &amp; Hanks, 1999)</xref>
        : the
latent variables describing student pro ciency form an
unobserved Markov process, and the periodic
assessments provide observable evidence about the state of
those latent variables. The instructional activities
chosen between time points are the measurement space,
and in fact, the students response to instruction often
provides important clues about their pro ciency and
speci c learning problems
        <xref ref-type="bibr" rid="ref12">(Marcotte &amp; Hintze, 2009)</xref>
        .
        <xref ref-type="bibr" rid="ref2">Almond (2009)</xref>
        notes the similarity between POMDPs
and other frameworks more commonly used in
education, such as latent growth modeling
        <xref ref-type="bibr" rid="ref19">(Singer &amp; Willett,
2003)</xref>
        and hierarchical linear modeling
        <xref ref-type="bibr" rid="ref18">(HLM;
Raudenbush &amp; Byrk, 2002)</xref>
        . The principle di erence is one of
emphasis: in the POMDP framework, the emphasis is
usually on estimating the individuals latent state for
the purpose of planning. In the HLM and multilevel
growth model, the emphasis is usually on estimating
the e ectiveness of various activities. This paper looks
at the problem of forecasting using HLM models both
directly and through conversion to POMDP
parameterizations.
      </p>
      <p>The purpose of our study is to try to t a
POMDPbased latent growth model using Bayesian methods to
a set of data documenting the development of Reading
skills in a number of Kindergarten students. Once the
model is successfully t, we will use it to predict the
end-of-year status of the students.
given special instructions nor a prescribed curriculum,
although most of them used the same curriculum.
2</p>
    </sec>
    <sec id="sec-2">
      <title>THE DATA</title>
      <p>
        This study uses longitudinal data about reading
development originally collected by the Florida Center
for Reading Research
        <xref ref-type="bibr" rid="ref4">(Al Otaiba et al., 2011)</xref>
        . The
reading skills for this initial cohort of students was
measured three times (Fall, Winter and Spring)
during Kindergarten, and follow-up measurements were
taken at the end of 1st, 2nd and 3rd grade. There
were 247 students in the initial sample, but only 224
were still in the area at the end of the rst year.
During Kindergarten, children rapidly develop in
Reading and pre-Reading skills (e.g., oral
vocabulary and letter identi cation). Consequently, not all
measures are appropriate for all time points.
Consequently, di erent measures were collected at di erent
time points. Table 1 shows the measures that were
collected during Kindergarten:
      </p>
      <p>Additionally, teacher and school identi ers are
available for each child. For this cohort teachers were not
3</p>
    </sec>
    <sec id="sec-3">
      <title>THE POMDP FRAMEWORK</title>
      <p>
        <xref ref-type="bibr" rid="ref1">Almond (2007)</xref>
        provides a generalized model for how
a POMDP can represent measurement of a developing
pro ciency across multiple time points (Figure 1).
      </p>
      <p>Activity</p>
      <p>Activity
S
O
t=1
Assessment</p>
      <p>Growth</p>
      <p>S
O
t=2</p>
      <p>
        S
O
t=3
In this gure, the nodes marked S represent the
latent student pro ciency as it evolves over time. At
each time slice, there is generally some kind of
measurement of student progress represented by the
observable outcomes O. Note that these may be di
erent for di erent time slices (c.f., Table 1). Following
the terminology of evidence-centered assessment
design
        <xref ref-type="bibr" rid="ref13">(ECD; Mislevy, Steinberg, &amp; Almond, 2003)</xref>
        we
call this an evidence model. In general, both the pro
ciency variables at Measurement Occasion m, Sm, and
the observable outcome variables on that occasion, Om
are multivariate.
      </p>
      <p>
        Extending the ECD terminology,
        <xref ref-type="bibr" rid="ref1">Almond (2007)</xref>
        calls
the model for the Sm's, the pro ciency growth model.
Following the normal logic of POMDPs this is
expressed with two parts: the rst is the initial pro
ciency model, which gives the population distribution
for pro ciency at the rst measurement occasion. The
second is an action model, which gives a probability
distribution for change in pro ciency over time that
depends on the instructional activity chosen between
measurement occasions.
3.1
      </p>
      <sec id="sec-3-1">
        <title>PROFICIENCY GROWTH MODEL</title>
        <p>For the data from the Al Otaiba et al. (2011) study, the
latent pro ciency is obviously Reading. The question
immediately arises as to how many dimensions to use
to represent the reading construct. As the students are
entering the study in Kindergarten, components of the
reading skill, such as oral vocabulary and phonemic
awareness are less tightly correlated than they are with
older children. (In the fall of the Kindergarten year the
correlation between the LW and PV scores in the Al
Otaiba et al. study was r = :46, n = 247, while in
the spring it had increased to r = :56, n = 224.) As
and 6 years 4 months at the time of the rst testing
(with a few students 7 years or older). This represents
a considerable variation in maturity, and potentially
in initial ability.</p>
        <p>We de ne the following model for Measurement
Occasion 1:</p>
        <p>Rn1</p>
        <p>N ( s(n); s(n))
3.2</p>
      </sec>
      <sec id="sec-3-2">
        <title>EVIDENCE MODELS</title>
        <p>Because we are assuming that Reading pro ciency is
unidimensional, we do not need to specify which of the
measures in Table 1 are relevant to which pro ciencies.
Thus, the evidence model is a collection of simple
regressions, for each observation Ynmi for Individual n
at Measurement Occasion m on Instrument i, we have:
Ynmi = ai + biRnm + nmi
nmi</p>
        <p>N (0; !i)
a starting point, we will t a unidimensional model
of Reading, representing it with a single continuous
variable: Rnm the reading ability of Individual n on
Measurement Occasion m.
3.1.1</p>
      </sec>
      <sec id="sec-3-3">
        <title>Model for Growth</title>
        <p>In the rst cohort of the Al Otaiba et al. (2011) study,
teachers were not given speci c instructions about
curriculum or activity between the time points. We
therefore do not have a dependency on activity to
measure here. However we do expect there to be some
classroom-to-classroom di erences, so we will make
the growth parameters dependent on the classroom
(The teacher e ect is part of the classroom e ect,
however aspects of the peer group and environment
are captured as well). Let c(n) be the classroom to
which Student n belongs. Note also that classrooms
are nested within schools, so school e ects are
considered part of the general classroom e ect.</p>
        <p>Following this logic, for Measurement Occasion m &gt; 1,
de ne:</p>
        <p>Rnm = Rn(m 1) + ( c(n)m + 0m) Tnm + nm
(1)
nm</p>
        <p>N (0; c(n)mp</p>
        <p>Tnm)
Here Tnm is the time between Measurement
Occasions m and m 1 for individual n. Here 0m is an
average growth rate, and cm is a classroom speci c
growth rate. Note that the residual standard
deviation depends on both a classroom speci c rate, cm,
and the time elapsed between measurements. This is
consistent with the model that student ability is
growing according to a nonstationairy Wiener (Brownian
motion) process.
3.1.2</p>
      </sec>
      <sec id="sec-3-4">
        <title>Model For Initial Pro ciency</title>
        <p>Children entering Kindergarten have very diverse
language and early literacy backgrounds. There are
considerable di erences in the amount of experience with
print material the child experiences at home, breadth
and depth of vocabulary used with the child, as well
as a wide variety of preschool experiences. As a child's
preschool and early home experiences are at least
partially dependent on their parents' social and economic
status, and within-school socio-economic status tends
to be more homogeneous than across school status, we
model the initial status as dependent on the school.
Let s(n) be the school attended (during Kindergarten)
for Student n.</p>
        <p>
          There is also a considerable variation in the age at
entry. In the Al Otaiba et al. (2011) study, 95% of the
children were between the ages of 5 years 2 months
(2)
(3)
(4)
(5)
Note that the slope parameter bi actually encodes a
relative importance for the various measures.
One advantage of this structure is that we we do not
need to explicitly specify the data collection structure
(Table 1). Instead, we can simply set the values of
measures not recorded in each wave to missing values.
Because each of the instruments are well established
          <xref ref-type="bibr" rid="ref22 ref9">(Woodcock et al., 2001; Good &amp; Kaminkski, 2002)</xref>
          ,
we know some of their critical psychometric
properties. In particular, the reliability of Instrument i, i
is documented in the handbooks for the measures. In
classical test theory, the reliability is the squared
correlation between the true score of an examinee and the
observed score. With a bit of algebra, this de nition
is is equivalent to:
Here the notation Varn( ) indicates that the variance
is taken over individuals (with measurement occasion
and instrument held constant). Solving Equation 5
for Varn( nmi) yields an estimate for !i2 for each
measurement occasion. We took the median of the three
estimates as our base estimate for !i2, !~i.
        </p>
        <p>One drawback of the classical test theory concept of
reliability is that it is dependent on the population being
measured. Thus, as the sample in the Al Otaiba et al.
(2011) is slightly di erent from the norming samples
used in the development of the WJ-III and DIBELS
measures, we expect our observed reliability will
differ slightly from the published values. What we do is
set up priors for !i using !~i as the prior mean. In
particular,
where Gamma( ; ) is a gamma distribution with
shape parameter and rate parameter . We note
that any gamma distribution with = !~i2 will have
the proper mean. The shape parameter is then e
ectively a tuning parameter giving the strength the prior
distribution, or equivalently the relative weight of the
published reliabilities and the observed error
distribution. We initially chose a value of = 100 weights the
prior knowledge as equivalent to 100 observations, but
later increased it to 1000 when we were experiencing
convergence problems.
A problem that frequently arises in educational models
using latent variables is the identi ability of the scale.
In particular, suppose we replaced Rnmi with Rn0mi =
Rnmi + c for an arbitrary constant c, and replaced ai
with a0i = ai bic. The likelihood of the observed data
Ynmi (implicit in Equation 3) would be identical. A
similar problem arises if we replace Rnmi with Rn00mi =
cRnmi and bi with b0i0 = bi=c. Additional constraints
must be added to the model to identify the scale and
location of the latent variable R.</p>
        <p>A frequently used convention in psychometrics is to
identify the scale and location of the latent variable
by assuming that the population mean and variance
for the latent variable is 0 and 1 (i.e., that the latent
variable has an approximately unit normal
distribution). In this case we can identify the scale for Rn1 by
constraining Ps s = 0 and S1 Ps s = 1, where S is
the total number of schools in the study.</p>
        <p>Because this is a temporal model, there exists another
complication. We need to identify the scale of Rnm for
m &gt; 1. In particular, the mean and variance of the
innovations 0m and tm can cause similar identi
ability to the scale and location for Rnm that the initial
mean and variance caused for Rnm. In this case we
apply a di erent solution. We assume that the
properties of the instruments, and their relationships to the
latent reading pro ciency do not vary across time (at
least for the time points they are in use). Note that
in Equation 3, the slope, bi and intercept, ai do not
vary across time. This establishes a common scale for
all time points.</p>
        <p>Our initial thinking was that this would be enough
to identify the model. Unfortunately, because of
the structural missing data additional constraints are
needed. These are described below.</p>
        <p>
          <xref ref-type="bibr" rid="ref5">Bafumi, Gelman, Park, and Kaplan (2005</xref>
          ) present a
di erent approach to enforcing identi ability. They let
the model be unidenti ed while tting the data, but
then transform the estimates when evaluating the data
          <xref ref-type="bibr" rid="ref1 ref16 ref17 ref6">(i.e., they enforce the constraint by manipulating the
samples in R and coda R Development Core Team,
2007; Plummer, Best, Cowles, &amp; Vines, 2006 rather
than in BUGS or JAGS)</xref>
          . For example, rather than
constraining P s = 0, they would estimate s freely,
but post hoc would adjust the sample from the rth
cycle, (sr)0 = (sr) P (sr), making appropriate
adjustments to the other parameters. They claim that the
resulting model mixes better, however, there is some
di culty in guring out how the post hoc adjustments
will a ect other parameters in the model.
4
        </p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>PROBLEMS WITH</title>
    </sec>
    <sec id="sec-5">
      <title>FITTING</title>
    </sec>
    <sec id="sec-6">
      <title>MODEL</title>
      <p>
        We attempted to t the model described in the
previous section with MCMC using JAGS
        <xref ref-type="bibr" rid="ref15">(Plummer,
2012)</xref>
        .1 After some initial di culties we removed the
teacher and school e ects (intending to add them again
after we t the simpler model). This also allowed us
to restrict the prior distribution for Rn1 to be a unit
normal distribution (zero mean, variance one). This
is a common identi ability constraint imposed in
psychometric models.
4.1
      </p>
      <sec id="sec-6-1">
        <title>FIVE MEASURE MODEL</title>
        <p>Our initial experiments involved ve of the six
measures (the PSF measure was left out due to a
mistake in the model setup). We ran three Markov chains
using random starting positions and found that the
models did not converge. Or more properly, the
evidence model parameters (ai, bi, and !i) for the
DIBELS NWF (nonsense word uency) measure did not
converge. Table 2 shows the posterior mean of the
evidence model parameters for the ve measures (because
the MCMC chain did not reach the stationary state,
this may not be the true posterior).</p>
        <p>
          Note in Table 2 that the estimated residual variance
is extremely low, indicating a nearly perfect
correlation between the latent Reading variable and the NWF
measure. In this case, the MCMC chain looks like it is
somehow using that measure to identify the scale of the
latent variable. Furthermore, the slope for that
variables is twice as high as the slope for other variables in
1Actually, we did some of our early model tting using
WinBUGS
          <xref ref-type="bibr" rid="ref11">(D. J. Lunn, Thomas, Best, &amp; Spiegelhalter,
2000)</xref>
          . Some of the identi cation problems we were having
in WinBUGS we are not having in JAGS. JAGS may be
using slightly better samplers which may take care of issues
that occur when the predictor variables in regressions are
not centered
          <xref ref-type="bibr" rid="ref15">(Plummer, 2012)</xref>
          . Similar improvements may
have been made in OpenBUGS
          <xref ref-type="bibr" rid="ref10">(D. Lunn, Spiegelhalter,
Thomas, &amp; Best, 2009)</xref>
          , the successor to WinBUGS, but
we have not tested this model using OpenBUGS.
Trace plots of the evidence models show the problem.
Figure 2 shows an example of extremely slow mixing,
that is characteristic of identi ability problems.
Depending on the values of the other variables in the
system (particularly the latent reading variables) higher
or lower slopes may be sensible. Looking at the trace
plots of Rnm for several students show similar poor
mixing for m &gt; 1. We would expect similar problems
with the trace plots for 0m, but the mixing looks good
on those chains.
        </p>
        <p>It is likely that the problem is some complex
interaction between using 0m and the b's to identify mean
growth, or the a's which de ne the starting point for
growth. Note that the problematic measure, NWF,
was not measured at the rst time point. Thus the
constraint on the distribution of Rn0 will not de ne its
scale in the second or third measurement occasions.
As the problematic measure may be the ones which
were not recored at all three time points, we ran the
model again, dropping the ISF and NWF measures
(the ones not observed at the rst or third
measurement occasion). The new model also did not converge,
although the focus of the problem has now moved from
the NWF measure to the LNF measure.</p>
        <p>Table 4 shows the new estimates from the unconverged
posterior. Again, the variance for the measure that did
not converge is substantially smaller than that of the
other measures, and the slope is substantially higher.
Again the trace plots (Figure 3) show poor mixing,
as do similar plots for the Rnm measures for m &gt; 1.
There is also an indication of a trend that indicates
that the chains have not covered the whole of the
posterior distribution.</p>
        <p>Evidence Model Parameters, 3 Measure</p>
        <p>What is required is a method for xing the value of
ai for measures that were not collected at the initial
time point. One possible way to do this would be to
simply set ai = 0. This is not unreasonable, if all of
22 23 24
N = 5000 Bandwidth = 0.1265
25
8
ty .0
ienD .04
s
00 −1.0 −0.8 −0.6 −0.4 −0.2 0.0
.</p>
        <p>y
the variables are on a standardized scale: it implies
that the average trajectory of the average student will
pass through the average of the scores.</p>
        <p>This required that the scores all be on the same scale
(especially problematic with the WJ-III and DIBELS
scores based on di erent development and norming
sample. Fortunately, for these data all six measures
were collected in the winter time period.
Subtracting the mean of the Winter scores and dividing by the
standard deviation for each measure produced
standardized scores. This standardization together with
the constraint ai = 0 caused the models to converge.
4.4</p>
      </sec>
      <sec id="sec-6-2">
        <title>SIX MEASURE MODEL</title>
        <p>Using the standardized data and the additional
constraint of a1 = 0, we again t the model using MCMC.
This time, we got convergence on all of the evidence
model parameters (Figure 4).</p>
        <p>Table 5 shows the mean of the latent Reading variable
for the rst ve students in the sample. This appears
to be well behaved with all of the students showing
growth across the three time points.</p>
      </sec>
    </sec>
    <sec id="sec-7">
      <title>FUTURE DIRECTIONS AND</title>
    </sec>
    <sec id="sec-8">
      <title>CHALLENGES</title>
      <p>The key to getting this model to converge was the
standardization of the measure scales. Fortunately,
this data set had a time period where all six measures
were applied to the same population. Consequently,
standardizing the scales at this time point put the
measures on a comparable scale, which then made the xed
intercept constraint meaningful.</p>
      <p>
        It is di cult to see how this generalizes to cases in
which there is not a single time point in which all
measures are collected. This is a problem with the cohort
examined in this study when we look at the data
gathered in rst and second grades. As the students
reading abilities develop, new and more di cult measures
of reading become appropriate. Linking these back
to the old scale is a di cult problem. This problem
is well known in the educational literature under the
name \vertical scaling"
        <xref ref-type="bibr" rid="ref20">(von Davier, Carstensen and
von Davier, 2006, provide a review of the literature)</xref>
        .
Now that the model without teacher or school e ects
converges, the next step is to add those back into the
model. Also, we should use cross-validation to
evaluate how well the model predicts students scores. The
Al Otaiba et al. (2011) data set has long term follow-up
for a substantial portion of the students, so we can see
how well the model can predict First and Second grade
reading scores as well. Finally, we can look at the
rules for classi cation in to special instruction, to see
whether integrating the data across multiple measures
provides a better picture of the student than looking
at one measure alone.
      </p>
      <sec id="sec-8-1">
        <title>Acknowledgments</title>
        <p>We would like to thank the Florida Center for
Reading Research for allowing us access to the data used in
this paper. The data were originally collected as part
of a larger National Institute of Child Health and
Human Development Early Child Care Research Network
study.</p>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          <string-name>
            <surname>Almond</surname>
            ,
            <given-names>R. G.</given-names>
          </string-name>
          (
          <year>2007</year>
          ).
          <article-title>Cognitive modeling to represent growth (learning) using Markov decision processes</article-title>
          .
          <source>Technology, Instruction, Cognition and Learning (TICL)</source>
          ,
          <volume>5</volume>
          , 313{
          <fpage>324</fpage>
          . Available from http://www .oldcitypublishing.com/TICL/TICL.html
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          <string-name>
            <surname>Almond</surname>
            ,
            <given-names>R. G.</given-names>
          </string-name>
          (
          <year>2009</year>
          ).
          <article-title>Estimating parameters of periodic assessment models (Research Report No</article-title>
          . To appear).
          <source>Educational Testing Service.</source>
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          <string-name>
            <surname>Almond</surname>
            ,
            <given-names>R. G.</given-names>
          </string-name>
          (
          <year>2010</year>
          ).
          <article-title>Using evidence centered design to think about assessments</article-title>
          . In V. J.
          <string-name>
            <surname>Shute</surname>
            &amp;
            <given-names>B. J.</given-names>
          </string-name>
          <string-name>
            <surname>Becker</surname>
          </string-name>
          (Eds.),
          <article-title>Innovative assessment for the 21st century: Supporting educational needs</article-title>
          . (pp.
          <volume>75</volume>
          {
          <fpage>100</fpage>
          ). Springer.
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          <string-name>
            <given-names>Al</given-names>
            <surname>Otaiba</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            ,
            <surname>Folsom</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. S.</given-names>
            ,
            <surname>Schatschnneider</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            ,
            <surname>Wanzek</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            ,
            <surname>Greulich</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            ,
            <surname>Meadows</surname>
          </string-name>
          ,
          <string-name>
            <surname>J.</surname>
          </string-name>
          , et al. (
          <year>2011</year>
          ).
          <article-title>Predicting rst-grade reading performance from kindergarten response to tier 1 instruction</article-title>
          .
          <source>Exceptional Children</source>
          ,
          <volume>77</volume>
          (
          <issue>4</issue>
          ),
          <volume>453</volume>
          {
          <fpage>470</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          <string-name>
            <surname>Bafumi</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gelman</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Park</surname>
            ,
            <given-names>D. K.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>Kaplan</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          (
          <year>2005</year>
          ).
          <article-title>Practial issues in implementing and understanding bayesian ideal point estimation</article-title>
          .
          <source>Political Analysis</source>
          ,
          <volume>13</volume>
          ,
          <fpage>171</fpage>
          -
          <lpage>187</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          <string-name>
            <surname>Bennett</surname>
            ,
            <given-names>R. E.</given-names>
          </string-name>
          (
          <year>2007</year>
          , May).
          <article-title>Assessment of, for, and as learning: Can we have all three? Paper presented at the Institute of Educational Assessors National Conference</article-title>
          , London, England. Available from http://www.ioea.org.uk/ Home/news and events/annual conference/ day1/randy bennett.aspx
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          <string-name>
            <surname>Black</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>Wiliam</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          (
          <year>1998</year>
          ).
          <article-title>Assessment and classroom learning</article-title>
          .
          <source>Assessment in Education: Principles, Policy, and Practice</source>
          ,
          <volume>5</volume>
          (
          <issue>1</issue>
          ),
          <volume>7</volume>
          {
          <fpage>74</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          <string-name>
            <surname>Boutilier</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Dean</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>Hanks</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          (
          <year>1999</year>
          ).
          <article-title>Decision-theoretic planning: Structural assumptions and computational leverage</article-title>
          .
          <source>Journal of Arti cial Intelligence Research</source>
          ,
          <volume>11</volume>
          ,
          <fpage>1</fpage>
          -
          <lpage>94</lpage>
          . Available from citeseer.ist.psu.edu/ boutilier99decisiontheoretic.html
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          <string-name>
            <surname>Good</surname>
            ,
            <given-names>R. H.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>Kaminkski</surname>
            ,
            <given-names>R. A</given-names>
          </string-name>
          . (Eds.). (
          <year>2002</year>
          ).
          <article-title>Dynamic indicators of basic early literacy skills (6th ed</article-title>
          .) [Computer software manual]. Available from https://dibels.uoregon.edu/
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          <string-name>
            <surname>Lunn</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Spiegelhalter</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Thomas</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>Best</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          (
          <year>2009</year>
          ).
          <article-title>The BUGS project: Evolution, critique and future directions (with discussion)</article-title>
          . Statistics in Medicine,
          <volume>28</volume>
          , 3049{
          <fpage>3082</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          <string-name>
            <surname>Lunn</surname>
            ,
            <given-names>D. J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Thomas</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Best</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>Spiegelhalter</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          (
          <year>2000</year>
          ).
          <article-title>WinBUGS { a Bayesian modeling framework: concepts, structure, and extensibility</article-title>
          .
          <source>Statistics and Computing</source>
          ,
          <volume>10</volume>
          , 325{
          <fpage>337</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          <string-name>
            <surname>Marcotte</surname>
            ,
            <given-names>A. M.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>Hintze</surname>
            ,
            <given-names>J. M.</given-names>
          </string-name>
          (
          <year>2009</year>
          ).
          <article-title>Incremental and predictive utility of formative assessment methods of reading comprehension</article-title>
          .
          <source>Journal of School Psychology</source>
          ,
          <volume>47</volume>
          ,
          <fpage>315</fpage>
          -
          <lpage>335</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          <string-name>
            <surname>Mislevy</surname>
            ,
            <given-names>R. J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Steinberg</surname>
            ,
            <given-names>L. S.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>Almond</surname>
            ,
            <given-names>R. G.</given-names>
          </string-name>
          (
          <year>2003</year>
          ).
          <article-title>On the structure of educational assessment (with discussion)</article-title>
          .
          <source>Measurement: Interdisciplinary Research and Perspective</source>
          ,
          <volume>1</volume>
          (
          <issue>1</issue>
          ),
          <fpage>3</fpage>
          -
          <lpage>62</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          <string-name>
            <surname>Pelligrino</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Glaser</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>Chudowsky</surname>
          </string-name>
          , N. (Eds.). (
          <year>2001</year>
          ).
          <article-title>Knowing what students know: The science and design of educational assessment</article-title>
          .
          <source>National Research Council.</source>
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          <string-name>
            <surname>Plummer</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          (
          <year>2012</year>
          , May).
          <source>JAGS version 3.2.0 user manual (3.2</source>
          .0 ed.) [Computer software manual]. Available from http://mcmc-jags .sourceforge.net/
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          <string-name>
            <surname>Plummer</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Best</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Cowles</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>Vines</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          (
          <year>2006</year>
          ).
          <article-title>coda: Output analysis and diagnostics for MCMC [Computer software manual]</article-title>
          .
          <source>(R package version 0</source>
          .
          <fpage>10</fpage>
          -7)
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          <string-name>
            <given-names>R</given-names>
            <surname>Development Core</surname>
          </string-name>
          <string-name>
            <surname>Team.</surname>
          </string-name>
          (
          <year>2007</year>
          ).
          <article-title>R: A language and environment for statistical computing [Computer software manual]</article-title>
          . Vienna, Austria. Available from http://www.R-project.org
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          <string-name>
            <surname>Raudenbush</surname>
            ,
            <given-names>S. W.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>Byrk</surname>
            ,
            <given-names>A. S.</given-names>
          </string-name>
          (
          <year>2002</year>
          ).
          <article-title>Hierarchical linear models (second edition ed</article-title>
          .).
          <source>Sage Publications.</source>
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          <string-name>
            <surname>Singer</surname>
            ,
            <given-names>J. D.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>Willett</surname>
            ,
            <given-names>J. B.</given-names>
          </string-name>
          (
          <year>2003</year>
          ).
          <article-title>Applied longitudinal data analysis: Modeling change and event occurrence (1st ed</article-title>
          .). Oxford University Press, USA.
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          <string-name>
            <surname>von Davier</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Carstensen</surname>
            ,
            <given-names>C. H.</given-names>
          </string-name>
          , &amp; von
          <string-name>
            <surname>Davier</surname>
            ,
            <given-names>A. A.</given-names>
          </string-name>
          (
          <year>2006</year>
          ).
          <article-title>Linking competencies in educational settings and measuring growth (Research Report No</article-title>
          . RR-
          <volume>06</volume>
          -12). ETS.
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          <string-name>
            <surname>Wiggins</surname>
            ,
            <given-names>G. P.</given-names>
          </string-name>
          (
          <year>1998</year>
          ).
          <article-title>Educative assessment: Designing assessments to inform and improve student performance. Jossey-Bass.</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          <string-name>
            <surname>Woodcock</surname>
            ,
            <given-names>R. W.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>McGrew</surname>
            ,
            <given-names>K. S.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>Mather</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          (
          <year>2001</year>
          ).
          <article-title>Wj-iii tests of cognitive abilities and achievement</article-title>
          [Computer software manual].
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>