<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Bayesian Network Models for Adaptive Testing</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Martin Plajner</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Jirˇ´ı Vomlel</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Institute of Information Theory and Automation, Academy of Sciences of the Czech Republic</institution>
          ,
          <addr-line>Pod voda ́renskou veˇzˇ ́ı 4, Prague 8, CZ-182 08</addr-line>
          ,
          <country country="CZ">Czech Republic</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Institute of Information Theory and Automation, Academy of Sciences of the Czech Republic</institution>
          ,
          <addr-line>Pod voda ́renskou veˇzˇ ́ı 4, Prague 8, CZ-182 08</addr-line>
          ,
          <country country="CZ">Czech Republic</country>
        </aff>
      </contrib-group>
      <fpage>24</fpage>
      <lpage>33</lpage>
      <abstract>
        <p>Computerized adaptive testing (CAT) is an interesting and promising approach to testing human abilities. In our research we use Bayesian networks to create a model of tested humans. We collected data from paper tests performed with grammar school students. In this article we first provide the summary of data used for our experiments. We propose several different Bayesian networks, which we tested and compared by cross-validation. Interesting results were obtained and are discussed in the paper. The analysis has brought a clearer view on the model selection problem. Future research is outlined in the concluding part of the paper.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>INTRODUCTION</title>
      <p>
        The testing of human knowledge is a very large field of
human effort. We are in touch with different ability and skill
checks almost daily. The computerized form of testing is
also getting an increased attention with the growing spread
of computers, smart phones and other devices which allow
easy impact on the target groups. In this paper we focus on
the Computerized Adaptive Testing (CAT)
        <xref ref-type="bibr" rid="ref1 ref8">(van der Linden
and Glas, 2000; Almond and Mislevy, 1999)</xref>
        .
      </p>
      <p>
        CAT aims at creating shorter tests and thus it takes less time
without sacrificing its reliability. This type of test is
computer administered. The test has an accompanied model
which models a student (a student model). This model is
constructed based on samples of previous students. During
the testing the model is updated to reflect abilities of one
particular student who is in the process of testing. At the
same time we use the model to adaptively select next
questions to be asked in order to ask the most appropriate one.
This leads to collection of significant information in shorter
time and allows to ask less questions. We provide an
additional description of the testing process in the Section 4 and
more information can be found also in
        <xref ref-type="bibr" rid="ref6">(Milla´n et al., 2000)</xref>
        .
It seems that there is a large possibility of applications of
CAT in the domain of educational testing
        <xref ref-type="bibr" rid="ref10 ref11 ref9">(Vomlel, 2004a;
Weiss and Kingsbury, 1984)</xref>
        .
      </p>
      <p>
        In this paper we look into the problem of using Bayesian
network models
        <xref ref-type="bibr" rid="ref3">(Kjaerulff and Madsen, 2008)</xref>
        for adaptive
testing
        <xref ref-type="bibr" rid="ref5">(Milla´n et al., 2010)</xref>
        . Bayesian network is a
conditional independence structure and its usage for CAT can
be understood as an expansion of the Item Response
Theory (IRT)
        <xref ref-type="bibr" rid="ref1">(Almond and Mislevy, 1999)</xref>
        . IRT has been
successfully used in testing for many years already and
experiments using Bayesian networks in CAT are also being
made
        <xref ref-type="bibr" rid="ref10 ref7 ref9">(Mislevy, 1994; Vomlel, 2004b)</xref>
        .
      </p>
      <p>We discuss the construction of Bayesian network
models for data collected in paper tests organized at grammar
schools. We propose and experimentally compare different
Bayesian network models. To evaluate models we simulate
tests using parts of collected data. Results of all proposed
models are discussed and further research is outlined in the
last section of this paper.
2</p>
    </sec>
    <sec id="sec-2">
      <title>DATA COLLECTION</title>
      <p>We designed a paper test of mathematical knowledge
of grammar school students focused on simple
functions (mostly polynomial, trigonometric, and
exponential/logarithmic). Students were asked to solve different
mathematical problems1 including graph drawing and
reading, calculation of points on the graph, root finding,
description of function shape and other function properties.
The test design went through two rounds. First, we
prepared an initial version of the test. This version was carried
out by a small group of students. We evaluated the first
version of the test and based on this evaluation we made
changes before the main test cycle. Problems were updated
and changed to be better understood by students. Few
prob1In this case we use the term mathematical “problem” due to
its nature. In general tests, terms “question” or “item” are often
used. In this article all of these terms are interchangeable.
lems were removed completely from the test, mainly
because the information benefit of the problem was too low
due to its high or low difficulty. Moreover we divided
problems into subproblems in the way that:
(a) it is possible to separate the subproblem from the main
problem and solve it independently or
(b) it is not possible to separate the subproblem, but it
represents a subroutine of the main problem solution.
Note that each subproblem of the first type can be viewed
as a completely separate problem. On the other hand,
subproblems of the second type are inseparable pieces of a
problem.</p>
      <p>Next we present an example of a problem that appeared in
the test.</p>
      <p>Example 2.1. Decide which of the following functions
f (x)
g(x)
= x2
=
2x</p>
      <p>8</p>
      <p>The final version of test contains 29 mathematical
problems. Each one of them is graded with 0–4 points. These
problems have been further divided into 53 subproblems.
Subproblems are graded so that the sum of their grades is
the grade of the parent problem, i.e., it falls into the set
f0; : : : ; 4g. Usually a question is divided into two parts
each graded by at most two points2. The granularity of
subproblems is not the same for all of them and is a subset
of the set f0; : : : ; 4g. All together, the maximal possible
score to obtain in the test is 120 points. In an alternative
evaluation approach, each subproblem is evaluated using
the Boolean values (correct/wrong). The answer is
evaluated as correct only if the solution of the subproblem and
the solution method is correct unless there is an obvious
numerical mistake.</p>
      <p>We organized tests at four grammar schools. In total 281
students participated in the testing. In addition to
problem solutions, we also collected basic personal data from
students including age, gender, name, and their grades in
mathematics, physics, and chemistry from previous three
school terms. The primal goal of the tests was not the
student evaluation. The goal was to provide them
valuable information about their weak and strong points. They
could view their result (the scores obtained in each
individual problem) as well as a comparison with the rest of the
test group. The comparisons were provided in the form of
quantiles in their class, school and all participants.</p>
      <p>2There is one exception from this rule: The first problem is
very simple and it is divided into 8 parts, each graded by zero or
one point (summing to the total maximum of 8).</p>
      <p>The Table 1 shows the average scores of the grammar
schools (the higher the score the better the results). We
also computed correlations between the score and average
grades from Mathematics, Physics, and Chemistry from
previous three school terms. The grades are from the set
f1; 2; 3; 4; 5g with the best grade being 1 and the worst
being 5. These correlations are shown in the Table 2.
Negative numbers mean that a better grade is correlated with a
better result, which confirms our expectation.
In this section we discuss different Bayesian network
models we used to model relations between students’ math
skills and students’ results when solving mathematical
problems. All models discussed in this paper consists of
the following:</p>
      <p>A set of n variables we want to estimate fS1; : : : ; Sng.
We will call them skills or skill variables. We will
use symbol S to denote the multivariable (S1; : : : ; Sn)
taking states s = (s1; : : : sn).</p>
      <p>A set of m questions (math problems) fX1; : : : ; Xmg.
We will use the symbol X to denote the multivariable
(X1; : : : ; Xm) taking states x = (x1; : : : ; xm).
A set of arcs between variables that define relations
between skills and questions and, eventually, also
inbetween skills and inbetween questions.</p>
      <p>The ultimate goal is to estimate the values of skills, i.e., the
probabilities of states of variables S1; : : : ; Sn.
3.1</p>
      <sec id="sec-2-1">
        <title>QUESTIONS</title>
        <p>The solution of math problems were either evaluated using
a numeric scale or using a Boolean scale as explained in the
previous section. Although the numeric scale carries more
information and thus it seems to be a better alternative,
there are other aspects discouraging such a choice. The
main problem is the model learning. The more the states
the higher the number of model parameters to be learned.
With a limited training data it may be difficult to reliably
estimate the model parameters.</p>
        <p>We consider two alternatives in our models. Variables
corresponding to problems’ solutions (questions) can either be
Boolean, i.e. they have two states only 0 and 1 or
integer, i.e. each Xi takes mi states f1; : : : ; mig,
mi 2 N, where mi is the maximal number points for
the corresponding math problem.</p>
        <p>In Section 5 we present results of experiments with both
options.
3.2</p>
      </sec>
      <sec id="sec-2-2">
        <title>SKILL NODES</title>
        <p>We assume the student responses can be explained by skill
nodes that are parents of questions. Skill nodes model the
student abilities and, generally, they are not directly
observable. Several decisions are to be made during the model
creation.</p>
        <p>The first decision is the number of skill nodes itself. Should
we expect one common skill or should it rather be several
different skills each related to a subset of questions only?
In the later case it is necessary to specify which skills are
required to solve each particular question (i.e. a math
problem). Skills required for the successful solution of a
question become parents of the considered question.
Most networks proposed in this paper have only one skill
node. This node is connected to all questions. The student
is thus modelled by a single variable. Ordinarily, it is not
possible to give a precise interpretation to this variable.
We created two models with more than one skill node. One
of them is with the Boolean scale of question nodes and
the other is with the numeric scale. We used our expert
knowledge of the field of secondary school mathematics
and our experiences gained during the evaluation of paper
tests. In these model we included 7 skill nodes with arcs
connecting each of them to 1 – 4 problems.</p>
        <p>Another issue is the state space of the skill nodes. As an
unobserved variable, it is hard to decide how many states it
should have. Another alternative is to use a continuous skill
variable instead of a discrete one but we did not elaborate
more on this option. In our models we have used skill nodes
with either 2 or 3 states (si 2 f1; 2g or si 2 f1; 2; 3g).
We tried also the possibility of replacing the unobserved
skill variable by a variable representing a total score of the
test. To do this we had to use a coarse discretization. We
divided the scores into three equally sized groups and thus we
obtained an observed variable having three possible states.
The states represent a group of students with “bad”,
“average”, and “good” scores achieved. The state of this variable
is known if all questions were included in the test. Thus,
during the learning phase the variable is observed and the
information is used for learning. On the other hand, during
the testing the resulting score is not known – we are trying
to estimate the group into which would this test subject fall.
In the testing phase the variable is hidden (unobserved).
3.3</p>
      </sec>
      <sec id="sec-2-3">
        <title>ADDITIONAL INFORMATION</title>
        <p>As mentioned above, we have collected not only solutions
to problems but also additional personal information about
students. This additional information may improve the
quality of the student model. On the other hand it makes
the model more complex (more parameters need to be
estimated). It may mislead the reasoning based solely on
question answers (especially later when sufficient
information about a student is collected from his/her answers).
The additional variables are Y1; : : : ; Y` and they take states
y1; : : : ; y`. We tested both versions of most of the models,
i.e. models with or without the additional information.
3.4</p>
      </sec>
      <sec id="sec-2-4">
        <title>PROPOSED MODELS</title>
        <p>In total we have created 14 different models that differ in
factors discussed above. The combinations of parameters’
settings are displayed in the Table 3. One model type is
shown in the Figure 1. It is the case of ”tf plus” which is
a network with one hidden skill node and with the
additional information3. Models that differ only by number of
states of variables have the same structure. Models with
the “obs” infix in the name and “o” in the ID have the skill
variable modified to represent score groups rather than skill
(as explained earlier in the part 3.2). Models without
additional information do not contain the part of variables on
the right hand side of the skill variable S1. Figure 2 shows
the structure of the expert models with 7 skill variables in
the middle part of the figure.
4</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>ADAPTIVE TESTS</title>
      <p>All proposed models are supposed to serve for adaptive
testing. In this section we describe the process of adaptive
testing with the help of these models.</p>
      <p>At first, we select the model which we want to use. If this
model contains additional information variables it is
necessary to insert observed states of these variables before we
start selecting and asking questions. Next, following steps
are repeated:</p>
      <p>The next question to be asked is selected.</p>
      <p>The question is asked and a result is obtained.</p>
      <p>The result is inserted into the network as evidence.
3Please note that the missing problems and problem numbers
are due to the two-cycled test creation and problems removal.</p>
      <p>The network is updated with this evidence.</p>
      <p>(optional) Subsequent answers are estimated.</p>
      <p>
        This procedure is repeated as long as necessary. It means
until we reach a termination criterion which can be either
a time restriction, the number of questions, or a confidence
interval of the estimated variables. Each of these criterion
would lead to a different learning strategy
        <xref ref-type="bibr" rid="ref10 ref9">(Vomlel, 2004b)</xref>
        ,
but because such strategy would be NP-Hard
        <xref ref-type="bibr" rid="ref4">(L´ın, 2005)</xref>
        .
We have chosen an heuristic approach based on greedy
entropy minimization.
4.1
      </p>
      <sec id="sec-3-1">
        <title>SELECTING NEXT QUESTION</title>
        <p>One task to solve during the procedure is the selection of
the next question. It is repeated in every step of the testing
and it is described below.</p>
        <sec id="sec-3-1-1">
          <title>Let the test be in the state after s</title>
        </sec>
        <sec id="sec-3-1-2">
          <title>1 steps where</title>
          <p>Xs
=</p>
          <p>fXi1 : : : Xin j i1; : : : ; in 2 f1; : : : ; mgg
are unobserved (unanswered) variables and
e =</p>
          <p>fXk1 = xk1 ; : : : ; Xko = xko jk1; : : : ; ko 2 f1; : : : ; mgg
is evidence of observed variables – questions which were
already answered and, possibly, the initial information. The
goal is to select a variable from Xs to be asked as the next
question. We select a question with the largest expected
information gain.</p>
          <p>We compute the cumulative Shannon entropy over all skill
variables of S given evidence e. It is given by the following
Assume we decide to ask a question X0 2 Xs with possible
outcomes x01; : : : ; x0p. After inserting the observed outcome
the entropy over all skills changes. We can compute the
value of new entropy for evidence extended by X0 = x0j ,
j 2 f1; : : : ; pg as:</p>
          <p>H(e; X0 = x0j ) =
n
X X
i=1 si</p>
          <p>P (Si = sije; X0 = x0j )
log P (Si = sije; X0 = x0j )
:
This entropy H(e; X0 = x0j ) is the sum of individual
entropies over all skill nodes. Another option would be to
compute the entropy of the joint probability distribution
of all skill nodes. This would take into account
correlations between these nodes. In our task we want to estimate
marginal probabilities of all skill nodes. In the case of high
correlations between two (or more) skills the second
criterion would assign them a lower significance in the model.
This is the behavior we wanted to avoid. The first
criterion assigns the same significance to all skill nodes which
seems to us as a better solution. Given the objective of the
question selection, the greedy strategy based on the sum of
entropies provides good results. Moreover, the
computational time required for the proposed method is lower.
Now, we can compute the expected entropy after answering
question X0:</p>
          <p>EH(X0; e)
=</p>
          <p>p
X P (X0 = x0j je) H(e; X0 = x0j ) :
j=1
Finally, we choose a question X that maximizes the
information gain IG(X0; e)</p>
          <p>X
IG(X0; e)
=
=
arg max IG(X0; e) ; where
X02Xs</p>
          <p>H(e) EH(X0; e) :</p>
        </sec>
      </sec>
      <sec id="sec-3-2">
        <title>4.2 INSERTION OF THE SELECTED QUESTION</title>
        <p>The selected question X is given to the student and his/her
answer is obtained. This answer changes the state of
variable X from unobserved to an observed state x . Next, the
question together with its answer is inserted into the
vector of evidence e. We update the probability distributions
P (Sije) of skill variables with the updated evidence e. We
also recompute the value of entropy H(e). The question
X is also removed from Xs forming a set of unobserved
variables Xs+1 for the next step s and selection process can
be repeated.
4.3</p>
      </sec>
      <sec id="sec-3-3">
        <title>ESTIMATING SUBSEQUENT ANSWERS</title>
        <p>In experiments presented in the next section we will use
individual models to estimate answers for all subsequent
questions in Xs+1. This is easy since we enter evidence
e and perform inference to compute P (X0 = x0je) for all
states of X0 2 Xs+1 by invoking the distribute and collect
evidence procedures in the BN model.
5</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>MODEL EVALUATION</title>
      <p>
        In this section we report results of tests performed with
networks proposed in Section 3 of this paper. The
testing was done by 10-fold cross-validation. For each model
we learned the corresponding Bayesian network from 190 of
randomly divided data. The model parameters were learned
using Hugin’s
        <xref ref-type="bibr" rid="ref2">(Hugin, 2014)</xref>
        implementation of the EM
algorithm. The remaining 110 of the dataset served as a
testing set. This procedure was repeated 10 times to obtain 10
networks for each model type.
      </p>
      <p>The testing was done as described in Section 4. For every
model and for each student from the testing data we
simulated a test run. Collected initial evidence and answers were
inserted into the model. During testing we estimated
answers of the current student based on evidence collected so
far. At the end of the step s we computed probability
distributions P (Xije) for all unobserved questions Xi 2 Xs+1.
Then we selected the most probable state of Xi:
xi
=
arg max P (Xi = xlje) :</p>
      <p>xl
By comparing this value to the real answer x0i we obtained
a success ratio of the response estimation for all questions
Xi 2 Xs+1 of test (student) t in step s</p>
      <p>SRts
I(expr)
=
=</p>
      <p>PXi2Xs+1 I(xi = x0 )</p>
      <p>i ; where
jXs+1j
1 if expr is true
0 otherwise.</p>
      <p>The total success ratio of one model in the step s for all test
data (N = 281) is defined as</p>
      <p>SRs
=</p>
      <p>PN
t=1 SRts :</p>
      <p>N
We will refer to the success rate in the step s as to elements
of sr = (SR0; SR1; : : :), where SR0 is the success rate of
the prediction before asking any question.</p>
      <p>0
0.714
0.749
0.714
0.746
0.714
0.747
0.715
0.684
0.717
0.684
0.684
0.686
0.716
0.684
Table 4 shows success rates of proposed networks for
selected steps s = 0; 1; 5; 15; 25; 30. The network ID
corresponds to the ID from the Table 3. The most important part
of the tests are the first few steps, which is because of the
nature of CAT. We prefer shorter tests therefore we are
interested in the early progression of the model (in this case
approximately up to the step 20). During the final stages
of testing we estimate results of only a couple of questions
which in some cases may cause rapid changes of success
rates. Questions which are left to the end of the test do
not carry a large amount of information (because of the
entropy selection strategy). This may be caused by two
possible reasons. The first one is that the state of the
question is almost certain and knowing it does not bring any
additional information. The second possibility is that the
question connection with the rest of the model is weak and
because of that it does not change much the entropy of skill
variables. In the latter case it is also hard to predict the
state of such question because its probability distribution
also does not change much with additional evidence.
From an analysis of success rates we have identified
clusters of models with similar behavior. For models with
integer valued questions and also for models with Boolean
questions three clusters of models with similar success
ratio emerged:
models with skill variable of 3 states,
models with skill variable of 2 states, and
the expert model.</p>
      <p>We selected the best model from each cluster to display
success ratios SRs in steps s in Figure 3 for Boolean
questions and in Figure 4 for integer valued questions. We made
the following observations:</p>
      <p>Models with the skill variable with 3 states were more
successful.
Models with skill variable with 2 states were better at
the very end of tests, but this test stage is is not very
important for CAT since the tests usually terminates at
early stages as explained above.</p>
      <p>The expert model achieved medium quality prediction
in the middle stage but its prediction ability decreases
in the second half of the tests.</p>
      <p>We would like to point out that the distinction between
models is basically only by differences of skill variables
used in the models. The influence of additional
information is visible only at the very beginning of testing. As can
be seen in the Table 4 “+” models are scoring better in the
initial estimation and then in the first one. After that both
models follow almost the same track. In the late stages of
the test, models with additional information are estimating
worse than their counterparts without information. It
suggests that models without additional information are able
to derive the same information by getting answers to few
questions (in the order of a couple of steps).</p>
      <p>It is easy to observe that the expert model does not provide
as good results as other models especially during the
second half of the testing. As was stated above the second part
of the testing is not as important, nevertheless we have
investigated causes for these inaccuracies. The main possible
reason for this behavior may be the complexity of this type
of model. With seven skill nodes and various connections
to question nodes this model contains a significantly higher
number of parameters to be fitted. It is possible that our
limited learning sample leads to over-fitting. We have
explored the conditional probability tables (CPTs) of models
used during cross-validation procedure to see how sparse
they are. Our observation is shown in the Table 5. The
number AZT is the average of the total number of zeros in
cross-validation models for the specific configuration and
AS is the average sparsity of CPTs rows in these models.
We can see that in the same type of scales (Boolean or
numeric) the sparsity of expert models is significantly higher.
This can be improved by increasing data volume or
decreasing the model’s complexity. This finding is consistent
with the above explained possible cause for inaccuracies.
In addition we can observe that there is also an increase in
sparsity when more skill variables states are introduced. It
seems to us as a good idea to further explore the space
between one skill variable and seven skill variables as well as
the number of their states to provide a better insight into
this problem and to draw out more general conclusions.
In Figures 5 and 6 we compare which questions were often
selected by the tested models at different stages of the tests.
Figure 5 is for Boolean questions and Figure 6 for integer
valued questions. Only three models (the same as for
success ratio plots) were selected because other models share
common behavior with others from the same cluster. On
the horizontal axis there is the step when the question was
asked, on the vertical axis are questions by their ID. The
darker the cell in the graph the more tests used the
corresponding variable in the corresponding time. Even though
it provides only a rough presentation it is possible to notice
different patterns of behavior. Especially, we would like to
point out the clouded area of the expert model where it is
clear that the individual tests were very different. Expert
models are apparently less sure about the selection of the
next question. This may be caused by a large set of skill
variables which divide the effort of the model into many
directions. This behavior is not necessarily unwanted
because it provides very different test for every test subject
which may be considered positive, but it is necessary to
maintain the prediction success rates.
6</p>
    </sec>
    <sec id="sec-5">
      <title>CONCLUSION AND FUTURE</title>
    </sec>
    <sec id="sec-6">
      <title>RESEARCH</title>
      <p>In this paper we presented several Bayesian network
models designed for adaptive testing. We evaluated their
performance using data from paper tests organized at grammar
schools. In the experiments we observed that:</p>
      <p>Larger state space of skill variables is beneficial.
Clearly, models with 3 states of the hidden skill
variable behave better during the most important stages of
the tests. Test with hidden variables with more than 3
states are still to be done.</p>
      <p>Expert model did not score as good as simpler models
but it showed a potential for its improvements. The
proposed expert model is much more complex than
other models in this paper and probably it can improve
its performance with more data collected.</p>
      <p>Additional information provided improves results
only during the initial stage. This fact is positive
because obtaining such additional information may be
hard in practice. Additionally, it can be considered
politically incorrect to make assumption about student
skills using this type of information.</p>
      <p>In the future we plan to explore models with one or two
hidden variables having more than three states, expert models
with skill nodes of more than 2 states, and try to add
relations between skills into the expert model to improve its
performance. We would also like to compare our current
results with standard models used in adaptive testing like
the Rash and IRT models.
b2+
b3
b2e
n2+
n3
n2e
0
5
10
15
25
30
35
40
20
step
step
25
30
35
40
0
3
0
3
0
1
0
4
0
3
0
1
s
r
a
v 0
2
s
r
a
v 0
2
0
0
0
15
steps
0
5
10
20
25
0
5
10</p>
      <p>15
stSetpesp
20
25
0
5
10</p>
      <p>15
steps
20
25
The work on this paper has been supported from GACR
project n. 13-20012S.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          <string-name>
            <surname>Almond</surname>
            ,
            <given-names>R. G.</given-names>
          </string-name>
          and
          <string-name>
            <surname>Mislevy</surname>
            ,
            <given-names>R. J.</given-names>
          </string-name>
          (
          <year>1999</year>
          ).
          <source>Graphical Models and Computerized Adaptive Testing. Applied Psychological Measurement</source>
          ,
          <volume>23</volume>
          (
          <issue>3</issue>
          ):
          <fpage>223</fpage>
          -
          <lpage>237</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          <string-name>
            <surname>Hugin</surname>
          </string-name>
          (
          <year>2014</year>
          ).
          <article-title>Explorer, ver. 8.0, comput</article-title>
          .
          <source>software</source>
          <year>2014</year>
          , http://www.hugin.com.
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          <string-name>
            <surname>Kjaerulff</surname>
          </string-name>
          , U. B. and
          <string-name>
            <surname>Madsen</surname>
            ,
            <given-names>A. L.</given-names>
          </string-name>
          (
          <year>2008</year>
          ).
          <source>Bayesian Networks and Influence Diagrams</source>
          . Springer.
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          <string-name>
            <surname>L´ın</surname>
            ,
            <given-names>V.</given-names>
          </string-name>
          (
          <year>2005</year>
          ).
          <article-title>Complexity of finding optimal observation strategies for bayesian network models</article-title>
          .
          <source>In Proceedings of the conference Znalosti</source>
          , Vysoke´ Tatry.
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          <string-name>
            <surname>Milla´n</surname>
          </string-name>
          , E.,
          <string-name>
            <surname>Loboda</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          , and Pe´
          <string-name>
            <surname>rez-de-la Cruz</surname>
            ,
            <given-names>J. L.</given-names>
          </string-name>
          (
          <year>2010</year>
          ).
          <article-title>Bayesian networks for student model engineering</article-title>
          .
          <source>Computers &amp; Education</source>
          ,
          <volume>55</volume>
          (
          <issue>4</issue>
          ):
          <fpage>1663</fpage>
          -
          <lpage>1683</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          <string-name>
            <surname>Milla´n</surname>
          </string-name>
          , E.,
          <string-name>
            <surname>Trella</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Prez-de-la Cruz</surname>
          </string-name>
          , J., and
          <string-name>
            <surname>Conejo</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          (
          <year>2000</year>
          ).
          <article-title>Using bayesian networks in computerized adaptive tests</article-title>
          . In Ortega, M. and
          <string-name>
            <surname>Bravo</surname>
          </string-name>
          , J., editors,
          <source>Computers and Education in the 21st Century</source>
          , pages
          <fpage>217</fpage>
          -
          <lpage>228</lpage>
          . Springer.
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          <string-name>
            <surname>Mislevy</surname>
            ,
            <given-names>R. J.</given-names>
          </string-name>
          (
          <year>1994</year>
          ).
          <article-title>Evidence and inference in educational assessment</article-title>
          .
          <source>Psychometrika</source>
          ,
          <volume>59</volume>
          (
          <issue>4</issue>
          ):
          <fpage>439</fpage>
          -
          <lpage>483</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          <string-name>
            <surname>van der Linden</surname>
          </string-name>
          , W. J. and
          <string-name>
            <surname>Glas</surname>
            ,
            <given-names>C. A.</given-names>
          </string-name>
          (
          <year>2000</year>
          ).
          <source>Computerized Adaptive Testing: Theory and Practice</source>
          . Springer.
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          <string-name>
            <surname>Vomlel</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          (
          <year>2004a</year>
          ).
          <article-title>Bayesian networks in educational testing</article-title>
          .
          <source>International Journal of Uncertainty, Fuzziness and Knowledge-Based Systems</source>
          ,
          <volume>12</volume>
          (supp01):
          <fpage>83</fpage>
          -
          <lpage>100</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          <string-name>
            <surname>Vomlel</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          (
          <year>2004b</year>
          ).
          <article-title>Buliding Adaptive Test Using Bayesian Networks</article-title>
          .
          <source>Kybernetika</source>
          ,
          <volume>40</volume>
          (
          <issue>3</issue>
          ):
          <fpage>333</fpage>
          -
          <lpage>348</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          <string-name>
            <surname>Weiss</surname>
            ,
            <given-names>D. J.</given-names>
          </string-name>
          and
          <string-name>
            <surname>Kingsbury</surname>
            ,
            <given-names>G. G.</given-names>
          </string-name>
          (
          <year>1984</year>
          ).
          <article-title>Application of Computerized Adaptive Testing to Educational Problems</article-title>
          .
          <source>Journal of Educational Measurement</source>
          ,
          <volume>21</volume>
          :
          <fpage>361</fpage>
          -
          <lpage>375</lpage>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>