<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Computer Adaptive Testing Using the Same-Decision Probability</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Suming Chen</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Arthur Choi</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Adnan Darwiche</string-name>
          <email>darwicheg@cs.ucla.edu</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Computer Science Department University of California</institution>
          ,
          <addr-line>Los Angeles</addr-line>
          ,
          <country country="US">USA</country>
        </aff>
      </contrib-group>
      <fpage>34</fpage>
      <lpage>43</lpage>
      <abstract>
        <p>Computer Adaptive Tests dynamically allocate questions to students based on their previous responses. This involves several challenges, such as determining when the test should terminate, as well as which questions should be asked. In this paper, we introduce a Computer Adaptive Test that uses a Bayesian network as the underlying model. Additionally, we show how the notion of the Same-Decision Probability can be used as an information gathering criterion in this context - to determine which further questions are needed and if so, which further questions should be asked. We show empirically that utilizing the Same-Decision Probability is a viable and intuitive approach for determining question selection in Bayesian-based Computer Adaptive Tests, as its usage allows us to ask fewer questions while still maintaining the same level of precision and recall in terms of classifying competent students.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>INTRODUCTION</title>
      <p>
        Computer Adaptive Tests have recently become
increasingly popular, as they have the ability to adapt to each
individual unique user
        <xref ref-type="bibr" rid="ref1 ref2 ref34">(Vomlel, 2004; Almond, DiBello,
Moulder, &amp; Zapata-Rivera, 2007; Almond, Mislevy, Steinberg,
Yan, &amp; Williamson, 2015)</xref>
        . This in turn allows for the test
to tailor itself specifically based on user responses and its
current estimation of the user’s knowledge level. For
example, if the user answers a series of questions correctly, the
test can adjust and curate some more difficult questions.
On the other hand, if the user answers a series of questions
incorrectly, the test can adjust and present easier questions.
Bayesian networks have been used as the principal base
model for many Computer Adaptive Tests
        <xref ref-type="bibr" rid="ref2 ref26 ref27 ref34">(Milla´n &amp;
Pe´rezDe-La-Cruz, 2002; Vomlel, 2004; Munie &amp; Shoham, 2008;
Almond et al., 2015)</xref>
        , as they offer powerful approaches to
      </p>
      <p>
        A key question in this domain is when the test should
terminate. Some students may perform so well or so poorly that
the system can recognize that further testing is unnecessary,
as asking further questions would have very little
probability of reversing the initial diagnosis.
        <xref ref-type="bibr" rid="ref26">(Milla´n &amp;
Pe´rez-DeLa-Cruz, 2002)</xref>
        discusses some stopping criteria that are
used to determine when further questioning is needed, as
well as some selection criteria that are used to determine
which questions are actually asked.
      </p>
      <p>
        In this paper, we discuss the creation of a Computer
Adaptive Test, and then take a recently introduced notion, called
the Same-Decision Probability (SDP)
        <xref ref-type="bibr" rid="ref15">(Darwiche &amp; Choi,
2010)</xref>
        , and show its usefulness as an information gathering
criteria in this domain, in contrast to standard criteria. The
SDP quantifies the stability of threshold-based decisions in
Bayesian networks and is defined as the probability that a
current decision would stay the same, had we observed
further information.
      </p>
      <p>The paper is structured as follows: We first present some
motivation for the constructed Computer Adaptive Test.
We then discuss some related work in the educational
diagnosis field. Following that, we then show how the
SameDecision Probability can be used as a stopping and
selection criterion for our Computer Adaptive Test. Finally, we
discuss experiment setup, present empirical results
demonstrating the usefulness of the Same-Decision Probability,
and then conclude the paper.
2</p>
    </sec>
    <sec id="sec-2">
      <title>MOTIVATION</title>
      <p>The Berkeley Free Clinic1 is a clinic staffed entirely by
volunteers and offers medical and dental treatment for the
sur1http://www.berkeleyfreeclinic.org/
rounding community. In particular, the dental section of the
clinic offers a wide variety of services ranging from
cleanings, x-ray services, fillings, and extractions. Due to the
expertise required for these services, the volunteers need
to be highly trained to apply their knowledge in a practical
setting. The volunteers are thus required to undergo a
training period in order to prepare them to serve at the clinic.
Recently, it was decided that in order to better evaluate
the knowledge of the volunteers, the volunteers would go
through a competency exam to determine whether or not
they had sufficient knowledge. The idea of this test is that
if a volunteer was found to be inadequate, that he/she would
have to undergo further instruction. The coordinators of the
clinic worked in conjunction with the volunteer dentists in
order to create a written test that thoroughly tested the
different concepts and skills necessary for a volunteer. The
written test consists of 63 questions that includes questions
ranging from showing a portrait of a certain instrument
(e.g. a perioprobe) and asking “Name this instrument”, to
“What is in a bridge and crown set up tray?”. A portion of
this test can be seen in Figures 1, 2, 3, and 4.</p>
      <p>This test was given to 22 subjects. Sixteen of the
subjects were evaluated prior to the test-taking to be competent
volunteers, whereas the other 6 subjects were evaluated to
be non-competent volunteers. The clinic coordinators set
the pass threshold to be at 60%, meaning that volunteers
needed to correctly answer 60% of the questions in order to
be considered competent. The test proved to be effective in
that of the 16 competent volunteers, 15 of them passed, and
of the 6 non-competent volunteers, none of them passed.
The test taking duration ranged from 30 minutes to 1 hour.
Feedback from the participants indicated that they felt this
test was fair and covered the necessary bases to be a
volunteer, but that the test was too time-consuming.</p>
      <p>
        The complaints about the duration of the test is in line with
the discoveries of
        <xref ref-type="bibr" rid="ref16">(Garc´ıa, Amandi, Schiaffino, &amp; Campo,
2007)</xref>
        , who discover that a significant problem with
webbased tests is that they are too long, which may hinder
students’ ability to perform as they simply become bored and
careless. This serves as the chief motivation for our work,
as we want to turn this test into a Computer Adaptive Test
(CAT), so that we can present students a test with fewer
questions, but just as many relevant questions — this way
we can decrease overall test duration without
compromising the effectiveness of the test.
3
      </p>
    </sec>
    <sec id="sec-3">
      <title>RELATED WORK</title>
      <p>
        There has been a surge of interest in applying artificial
intelligence techniques in educational diagnosis.
Educational diagnosis a broad field that covers Interactive
Tutoring Systems (ITS)
        <xref ref-type="bibr" rid="ref14 ref19 ref31">(Conati et al., 2002; Suebnukarn &amp;
Haddawy, 2006; Gujarathi &amp; Sonawane, 2012)</xref>
        , as well
as Computer Adaptive Tests (CAT)
        <xref ref-type="bibr" rid="ref25 ref27">(Munie &amp; Shoham,
2008; Milla´n, Descalco, Castillo, Oliveira, &amp; Diogo, 2013)</xref>
        .
When compared to traditional techniques, these
applications have been shown to be highly effective in increasing
the efficiency of student learning
        <xref ref-type="bibr" rid="ref25 ref26 ref30 ref34 ref4 ref5">(Milla´n &amp;
Pe´rez-De-LaCruz, 2002; Vomlel, 2004; Sinharay, 2006; Brusilovsky &amp;
Milla´n, 2007; Beal, Arroyo, Cohen, Woolf, &amp; Beal, 2010;
Milla´n et al., 2013)</xref>
        .
      </p>
      <p>
        In the work of
        <xref ref-type="bibr" rid="ref26">(Milla´n &amp; Pe´rez-De-La-Cruz, 2002)</xref>
        , they
use Bayesian networks as the framework for constructing
Computer Adaptive Tests (CATs). They introduce a model
for representing student knowledge, where knowledge is
modeled as different interrelated concepts. They noted that
student knowledge can be diagnosed (inferred) by treating
student answers as evidence. In addition to this model,
they introduced an adaptive testing algorithm. To
evaluate their model, they used 180 simulated students. They
found that the introduction of adaptive question selection
improves both accuracy and efficiency. Their model
incorporates notions such as “slipping”, where a student might
answer a question incorrectly even if the concept is known.
They generate students randomly by sampling from the
network, where each student is assumed to hold knowledge of
different concepts.
      </p>
      <p>
        Similar results were shown in
        <xref ref-type="bibr" rid="ref34">(Vomlel, 2004)</xref>
        , who discuss
applications of Bayesian networks in educational
diagnosis, especially in skill diagnosis. They show that modeling
dependence between various skills allows for higher
quality of diagnosis. Additionally, they show that Computer
Adaptive Tests allow them to best select a fitting question
to test for competency in certain skills. They found that
computer adaptive testing substantially reduces the number
of questions that may need to be asked.
      </p>
      <p>
        Just like
        <xref ref-type="bibr" rid="ref26">(Milla´n &amp; Pe´rez-De-La-Cruz, 2002)</xref>
        ,
        <xref ref-type="bibr" rid="ref34">(Vomlel,
2004)</xref>
        and
        <xref ref-type="bibr" rid="ref2">(Almond et al., 2015)</xref>
        found that Bayesian
networks to be useful to model the links between various
related proficiencies, and modeled unseen questions as
unobserved variables. They use diagnostic Bayesian networks in
order to model uncertainty, where experts build a Bayesian
network and calibrate it from data.
      </p>
      <p>
        Another application of Bayesian networks to model
student knowledge is discussed in
        <xref ref-type="bibr" rid="ref27">(Munie &amp; Shoham, 2008)</xref>
        .
They use a Bayesian network to model qualifying exams
for graduate students at Stanford, where their goal is to
select the questions (observations) that can best measure the
knowledge in order to determine whether or not the student
should pass. They also use a threshold function and stress
that reducing the number of questions asked is important.
They also prove the NP-hardness of deciding an optimal
set of questions to assign a student.
      </p>
      <p>
        Additionally, a variety of work has been done on using
Bayesian networks in the educational diagnosis field for
threshold-based decision making under uncertainty, such
as 1) in
        <xref ref-type="bibr" rid="ref35">(Xenos, 2004)</xref>
        , where the authors model the
educational experience of various students and use
thresholdbased decisions to determine whether or not a student is
likely to fail, 2) in
        <xref ref-type="bibr" rid="ref3">(Arroyo &amp; Woolf, 2005)</xref>
        , where the goal
is to determine the type of learner a student was (in terms of
attitude) 3) in
        <xref ref-type="bibr" rid="ref17">(Gertner, Conati, &amp; VanLehn, 1998)</xref>
        , where
the goal is to infer what part of a problem a student is
having trouble with, in order to provide hint-based
remediation, 4) in
        <xref ref-type="bibr" rid="ref6">(Butz, Hua, &amp; Maguire, 2004)</xref>
        , where the
constructed ITS can assess knowledge using a BN as well as
recommending certain remediation, and a threshold-based
decision is used to determine if a student is competent in
an area or not.
      </p>
      <p>
        Some other work in the field of educational diagnosis is
found in
        <xref ref-type="bibr" rid="ref5">(Brusilovsky &amp; Milla´n, 2007)</xref>
        , which describes a
model that details the relationship between the knowledge
of certain concepts and how the student will perform on
certain questions. They note that if a user demonstrates
lack of knowledge, the model can be used to locate the
most likely concepts that will remedy the situation. They
found that using a Bayesian model allows them to compute
the probability of a student answering a question correctly
given competency in some domain.
      </p>
      <p>
        More recently, the work on educational diagnosis
modeling has focused on even more intricate modeling and
remediation methods. For instance,
        <xref ref-type="bibr" rid="ref29">(Rajendran, Iyer, Murthy,
Wilson, &amp; Sheard, 2013)</xref>
        has focused on modeling user
emotional states, in order to detect when users are
frustrated as well as the cause of the frustration. By finding the
root cause of the frustration, special measures may be taken
to remediate. Similarly,
        <xref ref-type="bibr" rid="ref12">(Chrysafiadi &amp; Virvou, 2013)</xref>
        has
done significant work in modeling a student’s performance,
progress, and behavior. The developed e-learning system
has been deployed into production.
        <xref ref-type="bibr" rid="ref13">(Chrysafiadi &amp; Virvou,
2014)</xref>
        finds that modeling a user’s detailed characteristics,
such as knowledge, errors, and motivation can improve the
quality of a student’s learning process as it allows for
better remediation. The need for personalized remediation is
also further stressed as
        <xref ref-type="bibr" rid="ref32">(Vandewaetere &amp; Clarebout, 2014)</xref>
        studies the importance of giving learners instruction and
support directly. They stress that learner models need to be
adjusted and updated with new information about the
learners knowledge, effective states, and behavior in order to be
maximally effective.
4
4.1
      </p>
    </sec>
    <sec id="sec-4">
      <title>COMPUTER ADAPTIVE TESTING</title>
      <sec id="sec-4-1">
        <title>USING A BAYESIAN NETWORK</title>
        <p>
          To address the problem of having such a time-consuming
test, we believe that using a Computerized Adaptive Test
(CAT) could differentiate the competent students from the
non-competent students with fewer questions than a
standard test. Asking fewer questions while still accurately
measuring a student’s ability is a canonical problem in
educational diagnosis, and is discussed thoroughly in
          <xref ref-type="bibr" rid="ref16 ref25 ref26 ref27">(Milla´n
&amp; Pe´rez-De-La-Cruz, 2002; Garc´ıa et al., 2007; Munie &amp;
Shoham, 2008; Milla´n et al., 2013)</xref>
          . Our main goal is to
best measure whether or not a student is competent with
a limited subset of questions. This means that we have to
determine when enough questions have been asked, as well
as which additional questions should be asked, so as to
ensure that overall the Computer Adaptive Test is comprised
of fewer questions.
        </p>
        <p>
          Bayesian networks have been found to be especially useful
for decision making under uncertainty
          <xref ref-type="bibr" rid="ref28">(Nielsen &amp; Jensen,
2009)</xref>
          . We believe that they provide us a natural
mechanism to model both 1) student knowledge and 2) questions
that may be potentially asked to measure student
knowledge. In these models, Bayesian networks can model the
relationships between various proficiencies. These models
detail the relationship between knowledge of certain
concepts and how the student will perform on certain
questions.
        </p>
        <p>
          Using a Bayesian network thus allows us to predict how
likely a student is to have knowledge in a certain concept
based on the student’s answers. Hence, we worked with
the clinic coordinators and experts to create a Bayesian
network structure and elicit the parameters for this exam.
Our developed model is similar to the models developed
by
          <xref ref-type="bibr" rid="ref1 ref25 ref26 ref27 ref34 ref5">(Milla´n &amp; Pe´rez-De-La-Cruz, 2002; Vomlel, 2004;
Almond et al., 2007; Munie &amp; Shoham, 2008; Brusilovsky
&amp; Milla´n, 2007; Milla´n et al., 2013)</xref>
          , where Bayesian
networks are used to evaluate and in a sense, “diagnose” a
student’s degree of knowledge or competency.
        </p>
        <p>In our network, we have a main variable of interest that
is representative of the student’s total level of knowledge.
We refer to this variable as the decision variable (D), and
it serves as an overall measure of our belief that a student
is competent. General competency is determined by
competency in a collection of specialized fields of knowledge,
which are represented by a variety of latent concept
variables. The final type of variable is a question variable that
represents a question that may be asked to the student. An
answered question, whether correct or incorrect, will act as
evidence that can influence our belief on the student’s
competency. The graphical structure of the model can be seen
in Figure 5.</p>
        <p>
          Similarly to
          <xref ref-type="bibr" rid="ref17 ref33 ref7">(Gertner et al., 1998; VanLehn &amp; Niu, 2001;
Cantarel, Weaver, McNeill, Zhang, Mackey, &amp; Reese,
2014)</xref>
          , the clinic coordinators/experts determined that the
“pass threshold” should be set at 0:8, meaning that if given
some evidence (questions answered), the posterior
probability of the decision variable was found to be over 0:8,
then the student would be considered as competent.2
According to this pass threshold, we found once again that
of the 16 competent volunteers, 15 of them passed, and of
the 6 non-competent volunteers, none of them passed. The
volunteers’ scores are shown in Table 1. From this table
we can see that our constructed Bayesian model allows us
to accurately predict a student’s competency.
Although the test has a total of 63 questions, there may be
a point during the adaptive test where we can determine
that it is unnecessary to ask any further questions. For
educational diagnosis in Bayesian networks,
          <xref ref-type="bibr" rid="ref25 ref26">(Milla´n &amp;
Pe´rezDe-La-Cruz, 2002; Milla´n et al., 2013)</xref>
          shows there are two
standard ways to determine if more information gathering
(additional questions posed) is necessary. The first involves
a fixed test that asks a set number of questions, and then
determines if the student is competent after all questions have
been answered. This static method is clearly ineffective, as
the student is only evaluated at the conclusion of the test.
The second way involves terminating the test once the
posterior probability of the decision variable is above or
below some threshold. This stopping criteria is also seen
in
          <xref ref-type="bibr" rid="ref20 ref21 ref23 ref24">(Hamscher, Console, &amp; de Kleer, 1992; Heckerman,
Breese, &amp; Rommelse, 1995; Kruegel, Mutz, Robertson, &amp;
Valeur, 2003; Lu &amp; Przytula, 2006)</xref>
          . Note that
          <xref ref-type="bibr" rid="ref26">(Milla´n &amp;
Pe´rez-De-La-Cruz, 2002)</xref>
          compares an approach that
utilizes a set number of questions with a threshold-based
approach — they found that with a set number of questions
they were able to diagnose students correctly 90.27% of the
time, whereas by using an adaptive criterion they were able
to diagnose students correctly 94.54% of the time, while
requiring fewer number of questions to be asked. This is
a clear indication that using a less trivial stopping criterion
can be ultimately beneficial.
        </p>
        <p>
          Another possibility of a stopping criterion involves
computing the value of information of some observations, and
if that value is not high enough, to then stop information
gathering. However, this involves either computing the
myopic value of information and just computing the
usefulness of making one observation, or computing the
nonmyopic value of information and computing the usefulness
of making several observations. The former is fairly easy to
compute, but can prove to be very limited as often the
combined usefulness of some observations is greater than its
parts. For example, if a student answers a question
incorrectly, it may not be very telling. However, if that student
also answers a question incorrectly for an entirely
orthogonal field, these two mistakes combined may be indicative
of a student’s lack of competency. Note that computing the
non-myopic value of information is useful in determining if
a set of observations has significant value, but is intractable
to compute if the set of observations is large
          <xref ref-type="bibr" rid="ref22">(Krause &amp;
Guestrin, 2009)</xref>
          .
        </p>
        <p>
          To decide whether or not enough information has been
gathered, we use a non-myopic stopping criterion and
compute the Same-Decision Probability (SDP)
          <xref ref-type="bibr" rid="ref15 ref9">(Darwiche &amp;
Choi, 2010; Chen, Choi, &amp; Darwiche, 2014)</xref>
          .
        </p>
        <p>Definition 1 Suppose we are making a decision based on
whether Pr (D= d j e) T for some evidence e and
threshold T . If H is a set of variables that are available
to observe, then the SDP is:</p>
        <p>SDP(d; H; e; T ) =</p>
        <p>X[Pr (d j h; e)
h</p>
        <p>T ]Pr (h j e): (1)
Here, [ ] is an indicator function which is 1 if
and 0 otherwise.
is true,
In short, the SDP is a measure of how robust a decision
is with respect to some unobserved variables and will help
determine how much more information is necessary.
In this case, based on the student’s responses, at any point
in time, we can compute the posterior probability that they
are competent and thus determine a decision of whether or
not they are competent. Keep in mind that this decision is
temporary and can be reversed if more questions are
answered — we can compute the SDP over the remaining
unanswered questions to determine how likely it is that the
current decision (deciding whether a student is competent
or non-competent) would remain the same even if the
remaining questions were answered. If the SDP is high, that
is an indication that we can terminate the test early. We
found that using the SDP allowed us to cut the test duration
significantly while maintaining diagnosis accuracy. Details
of our experiment setup and results can be found in
Section 5.1.
4.3</p>
      </sec>
      <sec id="sec-4-2">
        <title>SELECTION CRITERION FOR</title>
      </sec>
      <sec id="sec-4-3">
        <title>INFORMATION GATHERING</title>
        <p>
          We have now motivated the usefulness of the SDP as a
stopping criterion for Computer Adaptive Tests, as it tells us
when we can terminate the test. However, the question
remains: if computing the SDP indicates that more questions
are necessary, which questions should we ask?
In a standard, non-adaptive test, the ordering of the
questions cannot be controlled. In the adaptive setting we have
more control — based on a test taker’s answers on the test
so far, we can select questions that have the most potential
to give us further insight on the test taker. Our goal here is
to select questions such that the expected SDP after
observing those questions is maximal. The expected SDP
          <xref ref-type="bibr" rid="ref10">(Chen,
Choi, &amp; Darwiche, 2015)</xref>
          is defined as the following:
Definition 2 Let G be a subset of the available features H
and let D(ge) be the decision made after observing
features G, i.e., D(ge) = d if Pr (d j g; e) T . The E-SDP
is then:
        </p>
        <p>E-SDP(D; G; H; e; T )
= X SDP(D(ge); H n G; ge; T ) Pr (g j e):
g
(2)
The expected SDP is thus a measure of how similar our
decision, after observing G, is to the decision made after
observing H. We want to find a question that will maximize
the expected SDP. In other words, we want to ask
questions such that no matter how they are answered, will, on
average, minimize the usefulness of any remaining
questions. By maximizing this objective, we reduce the overall
number of questions that need to be answered before test
termination.</p>
        <p>
          As stated before, computing the value of information (VOI)
is essential to the process of information gathering. Since
for our model it is intractable to compute the non-myopic
value of information, that leaves only computing the
myopic VOI as an available possibility. The VOI of
observing a variable may depend on various objective functions,
for instance, how much the observation reduces the
entropy of the decision variable. There is a comprehensive
overview of these different objective functions in
          <xref ref-type="bibr" rid="ref22">(Krause
&amp; Guestrin, 2009)</xref>
          .3
          <xref ref-type="bibr" rid="ref10 ref9">(Chen et al., 2014, 2015)</xref>
          compares the
approach maximizing these standard objective functions to
maximizing the expected SDP and finds that maximizing
the expected SDP is more effective in reducing the number
of questions that need to be answered, particularly in cases
when the decision threshold is extreme.
        </p>
        <p>Our approach involves selecting the question that leads to
the highest expected SDP. We found that selecting
variables to optimize the expected SDP allowed us to
substantially decrease the number of questions selected compared
to other selection criteria. Details of experiment setup and
experimental results can be found in Section 5.2.
5</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>EXPERIMENTS</title>
      <p>In this section, we empirically evaluate the SDP as a
stopping and selection criterion for our adaptive test. We
compare the SDP against the standard criteria used by other
Bayesian Computer Adaptive Tests. First, we introduce
some notation used throughout the experimental section.
We use standard notation for variables and their
instantiations, where variables are denoted by upper case letters
(e.g. X) and their instantiations by lower case letters (e.g.
x). Sets of variables are then denoted by bold upper case
letters (e.g. X) and their instantiations by bold lower case
letters (e.g. x). The primary decision variable, measuring
a student’s overall competency, is denoted by D; with two
states + (competent) and (non-competent).</p>
      <p>
        We have the completed tests for 22 subjects, where for each
subject, we have the responses of the n = 63 test questions
3Additionally,
        <xref ref-type="bibr" rid="ref18">(Golovin &amp; Krause, 2011)</xref>
        studies the usage of
adaptive submodularity for selection criteria and shows that
objective functions that satisfy this notion can be readily
approximated. For our problem, our objective function does not satisfy
the notion of adaptive submodularity.
(to be more precise, we know whether each question was
answered correctly or incorrectly). We denote this dataset
by T, where:
      </p>
      <p>T = fe1; : : : ; e22g;
and where each ei 2 T is an instantiation of the n test
responses for student i. Hence, the probability</p>
      <p>Pr (D = + j ei)
denotes the posterior probability that student i is competent
given their test results. Using our Bayesian network, we
decide that student i is competent if Pr (D = + j ei) 0:80
(otherwise, we decide that they are not sufficiently
competent). We report the quantity Pr (D = + j ei) for each
student, in Table 1.</p>
      <p>To evaluate the SDP as a stopping and selection
criterion, we simulate partially-completed tests from the
fullycompleted tests T. In particular, we take the
fullycompleted test results ei, for each student i, and generate a
set of partially-completed tests Qi = fqi;1; : : : ; qi;ng. In
particular, we randomly permute the questions, and take the
first j questions of the permutation as a partially-completed
test qi;j . Hence, a partially-completed test qi;j+1 adds one
additional test question to test qi;j , and test qi;n
corresponds to the fully-completed test ei.</p>
      <p>In our experiments, we use those partially-completed tests
qi;j that have at least 10 questions, i.e., where j 10 (we
assume 10 questions to be the minimum number of
questions, where we can begin to evaluate the competency of a
student). Moreover, for each of the 22 students i, we
simulated 50 sets of partially-completed tests Qi based on 50
random permutations, giving us a total of 22 50 = 1; 100
sets of partially-completed tests.
5.1</p>
      <sec id="sec-5-1">
        <title>STOPPING CRITERION EXPERIMENTS</title>
        <p>Using our partially-completed tests, we evaluate the SDP
as a stopping criterion, against more traditional methods.
In particular, we take each set of partially-completed tests
Qi, and going from test qi;10 up to test qi;n 1, we check
whether each stopping criterion is satisfied, i.e., a decision
is made to stop asking questions. Note that given test qi;n,
the only decision is to stop, since there are no more test
questions to ask. For the SDP, when we evaluate the test
qi;j , we compute the SDP with respect to the remaining
n j unanswered questions (i.e., we treat them as the set
of available observables H).</p>
        <p>As for other, more traditional, stopping criteria, we
consider (1) stopping after a fixed number of questions have
been answered (after which a student’s competence is
determined), or (2) stopping early once the posterior
probability Pr (D = + j qi;j ) surpasses a given threshold T ,
after which we deem a student to be competent.</p>
        <p>
          T
0.750
0.775
0.800
0.825
0.850
0.875
0.900
0.925
0.950
In Table 2, we report the results where our stopping
criterion is based on asking a fixed number of questions. In
Table 3, we highlight the results where we use instead
a posterior probability threshold
          <xref ref-type="bibr" rid="ref26 ref27 ref3 ref35">(Milla´n &amp;
Pe´rez-De-LaCruz, 2002; Xenos, 2004; Arroyo &amp; Woolf, 2005;
Munie &amp; Shoham, 2008)</xref>
          . These results are based on
averages over our 1; 100 sets of partially-completed tests Qi,
which we simulated from our original dataset. In both
tables, we see that there is a clear trade-off between the
accuracy of the test, and the number of questions that we
ask. (Note again that students were already evaluated to
be competent/non-competent prior to the test, and that
precision/recall is based on this prior evaluation).
a lower recall. In fact, we see that once the threshold is
set high enough, the recall actually drops: Some students
who should be diagnosed as competent are in fact being
diagnosed incorrectly as non-competent once the threshold
is set too high.
        </p>
        <p>Consider now the case where we set a very large threshold
on the SDP (0.999). In this case, the precision and recall
are equivalent to the case where our criterion is to ask a set
number of n 1 = 62 questions (as in Table 2). In contrast,
using the SDP criterion, we ask only an average of 42.05
questions, meaning that the SDP as a stopping criterion has
the same robustness as the ”set number of questions”
stopping criterion — while asking nearly 20 fewer questions.</p>
        <p>It is clear from our results that if our goal is to ask fewer
questions, while maintaining competitive precision and
recall rates, using the SDP as a stopping criterion is a
compelling alternative to the traditional stopping criteria of 1)
asking a set number of questions and 2) checking to see if
the posterior probability of competency surpasses a
threshold — using the SDP as a stopping criterion allows us to
reduce the number of questions asked while still
maintaining the same precision and recall.
5.2</p>
      </sec>
      <sec id="sec-5-2">
        <title>SELECTION CRITERION EXPERIMENTS</title>
        <p>
          We next consider experiments that compare various
question selection criteria such as 1) random selection, where
we select the next question randomly (as in non-adaptive
or linear testing), 2) information gain (mutual information)
          <xref ref-type="bibr" rid="ref26 ref34">(Milla´n &amp; Pe´rez-De-La-Cruz, 2002; Vomlel, 2004)</xref>
          , and 3)
margins of confidence
          <xref ref-type="bibr" rid="ref22">(Krause &amp; Guestrin, 2009)</xref>
          . Note
that as this model does not fit the decision-theoretic setting
(there are no assigned utilities), we do not consider utility
maximization as a selection criterion. In addition to the
above, we evaluate the SDP as a selection criterion.
Since our adaptive test involves selecting one question at a
time, our goal is to select a variable H 2 H
(corresponding to a question not yet presented) that leads to the greatest
gain in the SDP, i.e., the SDP gain. Our selection criteria
is based on asking the question which yields the highest
SDP gain, using information gain as a tie-breaker, when
multiple questions have the same SDP gain. A tie-breaker
We next consider the SDP as a stopping criterion. In
particular, we compute the SDP with respect to all unanswered
questions, and then make a stopping decision when the
computed SDP surpasses a given threshold T . Note that
when we use a threshold T = 1:0, then we commit to a
stopping decision only when no further observations will
change the decision (i.e., the probability of making the
same decision is 1.0). In Table 4, we report the results of
using the SDP as a stopping criterion.
        </p>
        <p>
          We see that even when our threshold is set to a relatively
small value (0.850), we still attain precision and recall rates
that are comparable to those obtained by asking nearly all
questions (as in Table 2). For the stopping criterion based
on posterior probability thresholds (in Table 3), we can see
that we can attain a higher precision, but at the expense of
is needed as the SDP gain of observing a single variable
can be zero. In this case, observing a single variable is not
enough to change our decision (multiple observations may
be needed). In our experiments, we compute the SDP gain
using a branch-and-bound algorithm, as in
          <xref ref-type="bibr" rid="ref8">(Chen, Choi, &amp;
Darwiche, 2013)</xref>
          . The advantages of this algorithm are (1)
it can prune the corresponding search space, and (2) it can
cache and reuse intermediate calculations, resulting in
significant computational savings.
        </p>
        <p>We compare how quickly these selection criteria allow us
to stop information gathering, each using the SDP as a
stopping criterion; see Table 5. Here, we took
partiallycompleted tests (all tests where at least 30 questions were
answered), and used each selection criterion to select
additional questions to ask, until the SDP dictates that we stop
asking questions. On average, we find that (1) the random
selection criterion asks an additional 8.18 questions, (2)
information gain asks an additional 3.67 questions, (3)
margins of confidence ask an additional 3.57 questions, and (4)
our approach based on the SDP gain requires the fewest
additional questions: 2.77.</p>
        <p>
          Note that while the random question selection criterion
clearly has the poorest performance, there is a relatively
modest improvement based on our approach, using the
SDP gain, compared to the more common approaches
based on information gain
          <xref ref-type="bibr" rid="ref26 ref34">(Milla´n &amp; Pe´rez-De-La-Cruz,
2002; Vomlel, 2004)</xref>
          , and margins of confidence
          <xref ref-type="bibr" rid="ref22">(Krause
&amp; Guestrin, 2009)</xref>
          . Nevertheless, the benefits of our
approach, as a selection criteria, are still evident. Further,
the results suggest the potential of using the SDP gain in a
less greedy way, where we ask multiple questions at a time,
which we consider a promising direction for future work.
The SDP has been shown to be highly intractable, being
PPPP-complete
          <xref ref-type="bibr" rid="ref11">(Choi, Xue, &amp; Darwiche, 2012)</xref>
          .
Therefore, the computational requirements of the SDP, as a
stopping and selection criterion, may be higher than other, more
common approaches. This complexity depends largely on
103
) 102
s
(
e
iTgm101
n
i
n
n
euR100
g
a
r
e
v
A10-1
10-2
the number of unasked questions that must be considered.
Figure 6 shows a plot of the running times for the SDP
stopping and selection criteria, in comparison to other
measures. Note that the SDP selection criterion is on
average more efficient than the SDP stopping criterion, even
though the selection criterion includes an SDP gain
computation, as well as an information gain computation (as a
tie breaker). Here, there are cases that can be detected that
allow us to skip the SDP gain computation
          <xref ref-type="bibr" rid="ref8">(Chen et al.,
2013)</xref>
          , leaving just the relatively efficient information gain
computation. In general, we note that there is a trade-off:
the SDP, as stopping and selection critera, provides
valuable information (leading to fewer questions asked), but the
SDP is also more computationally demanding.
        </p>
        <p>Algorithm Runtime Comparison
SDP-STOP
SDP-SEL
THRESH</p>
        <p>IG
5
10 15 20
Number of Unasked Questions
25
30
We created a Computer Adaptive Test using a Bayesian
network as the underlying model and showed how the notion
of the SDP can be used as an information gathering
criterion in this context. We showed that it can act as a stopping
criterion for determining if further questions are needed,
and as a selection criterion for determining which questions
should be asked. Finally, we have shown empirically that
the SDP is a valuable information gathering tool, as its
usage allows us to ask fewer questions while still maintaining
the same level of precision and recall for diagnosis.</p>
      </sec>
      <sec id="sec-5-3">
        <title>Acknowledgments</title>
        <p>This work has been partially supported by ONR grant
#N00014-12-1-0423, NSF grants #IIS-1514253 and
#IIS1118122, and a Google Research Award. We also thank the
coordinators and volunteers at the Berkeley Free Clinic.</p>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          <string-name>
            <surname>Almond</surname>
            ,
            <given-names>R. G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>DiBello</surname>
            ,
            <given-names>L. V.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Moulder</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>Zapata-Rivera</surname>
          </string-name>
          , J.
          <string-name>
            <surname>-D.</surname>
          </string-name>
          (
          <year>2007</year>
          ).
          <article-title>Modeling diagnostic assessments with Bayesian networks</article-title>
          .
          <source>Journal of Educational Measurement</source>
          ,
          <volume>44</volume>
          (
          <issue>4</issue>
          ),
          <fpage>341</fpage>
          -
          <lpage>359</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          <string-name>
            <surname>Almond</surname>
            ,
            <given-names>R. G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Mislevy</surname>
            ,
            <given-names>R. J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Steinberg</surname>
            ,
            <given-names>L. S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Yan</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>Williamson</surname>
            ,
            <given-names>D. M.</given-names>
          </string-name>
          (
          <year>2015</year>
          ).
          <source>Bayesian Networks in Educational Assessment</source>
          . Springer.
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          <string-name>
            <surname>Arroyo</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>Woolf</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          (
          <year>2005</year>
          ).
          <article-title>Inferring learning and attitudes from a Bayesian network of log file data</article-title>
          .
          <source>In Proceedings of the 12th International Conference on Artificial Intelligence in Education</source>
          , pp.
          <fpage>33</fpage>
          -
          <lpage>40</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          <string-name>
            <surname>Beal</surname>
            ,
            <given-names>C. R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Arroyo</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          , Cohen,
          <string-name>
            <given-names>P. R.</given-names>
            ,
            <surname>Woolf</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B. P.</given-names>
            , &amp;
            <surname>Beal</surname>
          </string-name>
          ,
          <string-name>
            <surname>C. R.</surname>
          </string-name>
          (
          <year>2010</year>
          ).
          <article-title>Evaluation of Animalwatch: An Intelligent Tutoring System for arithmetic and fractions</article-title>
          .
          <source>Journal of Interactive Online Learning</source>
          ,
          <volume>9</volume>
          (
          <issue>1</issue>
          ),
          <fpage>64</fpage>
          -
          <lpage>77</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          <string-name>
            <surname>Brusilovsky</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          , &amp; Milla´n,
          <string-name>
            <surname>E.</surname>
          </string-name>
          (
          <year>2007</year>
          ).
          <article-title>User models for adaptive hypermedia and adaptive educational systems</article-title>
          .
          <source>In The adaptive web</source>
          , pp.
          <fpage>3</fpage>
          -
          <lpage>53</lpage>
          . Springer-Verlag.
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          <string-name>
            <surname>Butz</surname>
            ,
            <given-names>C. J.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Hua</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>Maguire</surname>
            ,
            <given-names>R. B.</given-names>
          </string-name>
          (
          <year>2004</year>
          ).
          <article-title>A web-based Intelligent Tutoring System for computer programming</article-title>
          .
          <source>In Web Intelligence</source>
          , pp.
          <fpage>159</fpage>
          -
          <lpage>165</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          <string-name>
            <surname>Cantarel</surname>
            ,
            <given-names>B. L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Weaver</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>McNeill</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Zhang</surname>
          </string-name>
          , J.,
          <string-name>
            <surname>Mackey</surname>
            ,
            <given-names>A. J.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>Reese</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          (
          <year>2014</year>
          ).
          <article-title>Baysic: a Bayesian method for combining sets of genome variants with improved specificity and sensitivity</article-title>
          .
          <source>BMC bioinformatics</source>
          ,
          <volume>15</volume>
          (
          <issue>1</issue>
          ),
          <fpage>104</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          <string-name>
            <surname>Chen</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Choi</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>Darwiche</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          (
          <year>2013</year>
          ).
          <article-title>An exact algorithm for computing the Same-Decision Probability</article-title>
          .
          <source>In Proceedings of the 23rd International Joint Conference on Artificial Intelligence</source>
          , pp.
          <fpage>2525</fpage>
          -
          <lpage>2531</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          <string-name>
            <surname>Chen</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Choi</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>Darwiche</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          (
          <year>2014</year>
          ).
          <article-title>Algorithms and applications for the Same-Decision Probability</article-title>
          .
          <source>Journal of Artificial Intelligence Research (JAIR)</source>
          ,
          <volume>49</volume>
          ,
          <fpage>601</fpage>
          -
          <lpage>633</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          <string-name>
            <surname>Chen</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Choi</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>Darwiche</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          (
          <year>2015</year>
          ).
          <article-title>Value of information based on Decision Robustness</article-title>
          .
          <source>In Proceedings of the 29th Conference on Artificial Intelligence (AAAI).</source>
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          <string-name>
            <surname>Choi</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Xue</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>Darwiche</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          (
          <year>2012</year>
          ).
          <article-title>Same-Decision Probability: A confidence measure for threshold-based decisions</article-title>
          .
          <source>International Journal of Approximate Reasoning (IJAR)</source>
          ,
          <volume>2</volume>
          ,
          <fpage>1415</fpage>
          -
          <lpage>1428</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          <string-name>
            <surname>Chrysafiadi</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>Virvou</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          (
          <year>2013</year>
          ).
          <article-title>Persiva: An empirical evaluation method of a student model of an intelligent e-learning environment for computer programming</article-title>
          .
          <source>Computers &amp; Education</source>
          ,
          <volume>68</volume>
          ,
          <fpage>322</fpage>
          -
          <lpage>333</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          <string-name>
            <surname>Chrysafiadi</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>Virvou</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          (
          <year>2014</year>
          ).
          <article-title>Kem cs: A set of student's characteristics for modeling in adaptive programming tutoring systems</article-title>
          .
          <source>In Information, Intelligence, Systems and Applications</source>
          ,
          <string-name>
            <surname>IISA</surname>
          </string-name>
          <year>2014</year>
          , The 5th International Conference on, pp.
          <fpage>106</fpage>
          -
          <lpage>110</lpage>
          . IEEE.
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          <string-name>
            <surname>Conati</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gertner</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>VanLehn</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          (
          <year>2002</year>
          ).
          <article-title>Using Bayesian networks to manage uncertainty in student modeling. User Modeling</article-title>
          and
          <string-name>
            <surname>User-Adapted</surname>
            <given-names>Interaction</given-names>
          </string-name>
          ,
          <volume>12</volume>
          (
          <issue>4</issue>
          ),
          <fpage>371</fpage>
          -
          <lpage>417</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          <string-name>
            <surname>Darwiche</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>Choi</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          (
          <year>2010</year>
          ).
          <article-title>Same-Decision Probability: A confidence measure for threshold-based decisions under noisy sensors</article-title>
          .
          <source>In 5th PGM</source>
          , pp.
          <fpage>113</fpage>
          -
          <lpage>120</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          <string-name>
            <surname>Garc</surname>
            ´ıa,
            <given-names>P.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Amandi</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Schiaffino</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>Campo</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          (
          <year>2007</year>
          ).
          <article-title>Evaluating Bayesian networks precision for detecting students learning styles</article-title>
          .
          <source>Computers &amp; Education</source>
          ,
          <volume>49</volume>
          (
          <issue>3</issue>
          ),
          <fpage>794</fpage>
          -
          <lpage>808</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          <string-name>
            <surname>Gertner</surname>
            ,
            <given-names>A. S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Conati</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>VanLehn</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          (
          <year>1998</year>
          ).
          <article-title>Procedural help in Andes: Generating hints using a Bayesian network student model</article-title>
          .
          <source>In Proceedings of the 13th National Conference on Artificial intelligence</source>
          , pp.
          <fpage>106</fpage>
          -
          <lpage>111</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          <string-name>
            <surname>Golovin</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>Krause</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          (
          <year>2011</year>
          ).
          <article-title>Adaptive submodularity: Theory and applications in active learning and stochastic optimization</article-title>
          .
          <source>JAIR</source>
          ,
          <volume>42</volume>
          ,
          <fpage>427</fpage>
          -
          <lpage>486</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          <string-name>
            <surname>Gujarathi</surname>
            ,
            <given-names>M. V.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>Sonawane</surname>
            ,
            <given-names>M. S.</given-names>
          </string-name>
          (
          <year>2012</year>
          ).
          <article-title>Intelligent Tutoring System: A case study of mobile mentoring for diabetes</article-title>
          .
          <source>IJAIS</source>
          ,
          <volume>3</volume>
          (
          <issue>8</issue>
          ),
          <fpage>41</fpage>
          -
          <lpage>43</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          <string-name>
            <surname>Hamscher</surname>
            ,
            <given-names>W.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Console</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          , &amp; de Kleer, J. (Eds.). (
          <year>1992</year>
          ).
          <article-title>Readings in Model-Based Diagnosis</article-title>
          . Morgan Kaufmann Publishers Inc.
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          <string-name>
            <surname>Heckerman</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Breese</surname>
            ,
            <given-names>J. S.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>Rommelse</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          (
          <year>1995</year>
          ).
          <article-title>Decisiontheoretic troubleshooting</article-title>
          .
          <source>Communications of the ACM</source>
          ,
          <volume>38</volume>
          (
          <issue>3</issue>
          ),
          <fpage>49</fpage>
          -
          <lpage>57</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          <string-name>
            <surname>Krause</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>Guestrin</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          (
          <year>2009</year>
          ).
          <article-title>Optimal value of information in graphical models</article-title>
          .
          <source>Journal of Artificial Intelligence Research (JAIR)</source>
          ,
          <volume>35</volume>
          ,
          <fpage>557</fpage>
          -
          <lpage>591</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          <string-name>
            <surname>Kruegel</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Mutz</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Robertson</surname>
            ,
            <given-names>W.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>Valeur</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          (
          <year>2003</year>
          ).
          <article-title>Bayesian event classification for intrusion detection</article-title>
          .
          <source>In Proceedings of the Annual Computer Security Applications Conference (ACSAC).</source>
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          <string-name>
            <surname>Lu</surname>
          </string-name>
          , T.-C., &amp;
          <string-name>
            <surname>Przytula</surname>
            ,
            <given-names>K. W.</given-names>
          </string-name>
          (
          <year>2006</year>
          ).
          <article-title>Focusing strategies for multiple fault diagnosis</article-title>
          .
          <source>In Proceedings of the 19th International FLAIRS Conference</source>
          , pp.
          <fpage>842</fpage>
          -
          <lpage>847</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref25">
        <mixed-citation>
          <string-name>
            <surname>Milla´n</surname>
          </string-name>
          , E.,
          <string-name>
            <surname>Descalco</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Castillo</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Oliveira</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>Diogo</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          (
          <year>2013</year>
          ).
          <article-title>Using Bayesian networks to improve knowledge assessment</article-title>
          .
          <source>Computers &amp; Education</source>
          ,
          <volume>60</volume>
          (
          <issue>1</issue>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref26">
        <mixed-citation>
          <string-name>
            <surname>Milla´n</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          , &amp; Pe´
          <string-name>
            <surname>rez-De-La-Cruz</surname>
            ,
            <given-names>J. L.</given-names>
          </string-name>
          (
          <year>2002</year>
          ).
          <article-title>A Bayesian diagnostic algorithm for student modeling and its evaluation. User Modeling</article-title>
          and
          <string-name>
            <surname>User-Adapted</surname>
            <given-names>Interaction</given-names>
          </string-name>
          ,
          <volume>12</volume>
          (
          <issue>2-3</issue>
          ),
          <fpage>281</fpage>
          -
          <lpage>330</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref27">
        <mixed-citation>
          <string-name>
            <surname>Munie</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>Shoham</surname>
            ,
            <given-names>Y.</given-names>
          </string-name>
          (
          <year>2008</year>
          ).
          <article-title>Optimal testing of structured knowledge</article-title>
          .
          <source>In Proceedings of the 23rd National Conference on Artificial intelligence</source>
          , pp.
          <fpage>1069</fpage>
          -
          <lpage>1074</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref28">
        <mixed-citation>
          <string-name>
            <surname>Nielsen</surname>
            ,
            <given-names>T. D.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>Jensen</surname>
            ,
            <given-names>F. V.</given-names>
          </string-name>
          (
          <year>2009</year>
          ).
          <article-title>Bayesian networks and decision graphs</article-title>
          . Springer Science &amp; Business Media.
        </mixed-citation>
      </ref>
      <ref id="ref29">
        <mixed-citation>
          <string-name>
            <surname>Rajendran</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Iyer</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Murthy</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          , Wilson,
          <string-name>
            <given-names>C.</given-names>
            , &amp;
            <surname>Sheard</surname>
          </string-name>
          ,
          <string-name>
            <surname>J.</surname>
          </string-name>
          (
          <year>2013</year>
          ).
          <article-title>A theory-driven approach to predict frustration in an its. Learning Technologies</article-title>
          , IEEE Transactions on,
          <volume>6</volume>
          (
          <issue>4</issue>
          ),
          <fpage>378</fpage>
          -
          <lpage>388</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref30">
        <mixed-citation>
          <string-name>
            <surname>Sinharay</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          (
          <year>2006</year>
          ).
          <article-title>Model diagnostics for Bayesian networks</article-title>
          .
          <source>Journal of Educational and Behavioral Statistics</source>
          ,
          <volume>31</volume>
          (
          <issue>1</issue>
          ),
          <fpage>1</fpage>
          -
          <lpage>33</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref31">
        <mixed-citation>
          <string-name>
            <surname>Suebnukarn</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>Haddawy</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          (
          <year>2006</year>
          ).
          <article-title>A Bayesian approach to generating tutorial hints in a collaborative medical problembased learning system</article-title>
          .
          <source>Artificial intelligence in Medicine</source>
          ,
          <volume>38</volume>
          (
          <issue>1</issue>
          ),
          <fpage>5</fpage>
          -
          <lpage>24</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref32">
        <mixed-citation>
          <string-name>
            <surname>Vandewaetere</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>Clarebout</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          (
          <year>2014</year>
          ).
          <article-title>Advanced technologies for personalized learning, instruction, and performance</article-title>
          .
          <source>In Handbook of Research on Educational Communications and Technology</source>
          , pp.
          <fpage>425</fpage>
          -
          <lpage>437</lpage>
          . Springer.
        </mixed-citation>
      </ref>
      <ref id="ref33">
        <mixed-citation>
          <string-name>
            <surname>VanLehn</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>Niu</surname>
            ,
            <given-names>Z.</given-names>
          </string-name>
          (
          <year>2001</year>
          ).
          <article-title>Bayesian student modeling, user interfaces and feedback: A sensitivity analysis</article-title>
          .
          <source>International Journal of Artificial Intelligence in Education</source>
          ,
          <volume>12</volume>
          (
          <issue>2</issue>
          ),
          <fpage>154</fpage>
          -
          <lpage>184</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref34">
        <mixed-citation>
          <string-name>
            <surname>Vomlel</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          (
          <year>2004</year>
          ).
          <article-title>Bayesian networks in educational testing</article-title>
          .
          <source>International Journal of Uncertainty, Fuzziness and KnowledgeBased Systems</source>
          ,
          <volume>12</volume>
          (
          <issue>supp01</issue>
          ),
          <fpage>83</fpage>
          -
          <lpage>100</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref35">
        <mixed-citation>
          <string-name>
            <surname>Xenos</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          (
          <year>2004</year>
          ).
          <article-title>Prediction and assessment of student behaviour in open and distance education in computers using Bayesian networks</article-title>
          .
          <source>Computers &amp; Education</source>
          ,
          <volume>43</volume>
          (
          <issue>4</issue>
          ),
          <fpage>345</fpage>
          -
          <lpage>359</lpage>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>