<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Adaeze Nwaigwe</string-name>
          <email>adaeze.nwaigwe@faculty.umuc</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>University of Maryland University College 3501</institution>
          <addr-line>Unversity Blvd East Adelphi, MD 207831 412 608 8747 .edu</addr-line>
          ,
          <country country="US">USA</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>We use the Additive Factors Model to drive the evaluation of the student model of an Intelligent Tutoring System. Using data from the Andes Physics Tutor, applying the simple location heuristic and implementing the Additive Factors Model tool in the Pittsburgh's Science of Learning Center's DataShop, we discover possible ways to improve the student model of the Andes Intelligent Tutor.</p>
      </abstract>
      <kwd-group>
        <kwd>Student modeling</kwd>
        <kwd>learning curves</kwd>
        <kwd>additive factors model</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1.1 BACKGROUND</title>
      <p>
        Learning curves derived from student models drive evaluation,
revision and improvement of the Intelligent Tutor. The AFM is a
statistical algorithm which models learning and performance by
using logistical regression performed over the “error rate”
learning curve data [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. If a student is learning the knowledge
component (KC) or skill being measured, the learning curve is
expected to follow a so-called “power law of practice” [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]. If such
a curve exists, it presents evidence that the student is learning the
skill being measured or conversely, that the skill represents what
the student is learning.
      </p>
      <p>
        While use of learning curves is now a standard technique for
assessing the cognitive models of Intelligent Tutors, the technique
requires that a method is instated for attributing blame to skills or
KCs. This simply means that each error a student makes must be
blamed on a skill or set of skills. Four different heuristics for error
attribution have been proposed and tested. These heuristics are
guided by whether the method is driven by location – the simple
location heuristic (LH), the model-based location heuristic
(MLH); or by the temporal order of events – the temporal
heuristic (TH), the model-based temporal heuristic (MTH); and
whether the choice of the student model is leveraged (MLH,
MTH) [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ].
2 EVALUATING THE STUDENT MODEL
2.1Adapting the Andes Log data for the AFM
Algorithm
The log data used for this work was obtained from the Andes
Intelligent Tutor [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ] and encompassed four problems in the area
of electric field, across 102 students. The data was collected in
Spring 2005 at the US Naval Academy during its regular physics
class and as part of the PSLC’s LearnLab facility that provides
researchers, access to run experiments in or perform secondary
analyzes of data collected from one of seven available
technologyenhanced courses running at multiple high school and college
sites (see http://learnlab.org).
      </p>
      <p>Prior to using the AFM tool on the dataset, the simple location
heuristic (LH) was applied to error transactions in the Andes log
data which had missing KCs. That is, when the Andes failed to
assign blame to a KC on an error transaction, the LH will select
the first correctly implanted KC in the same location as the error.
The LH was applied to about 44% of the original data. Table 1
depicts a summary of the LH data.
2.2 Generating Model Values using AFM
The Datashop’s AFM algorithm was used to compute statistical
measures of goodness of fit for the model - Akaike Information
Criterion (AIC) and Bayesian Information criterion (BIC), as well
as to generate learning curves for the Andes log data.
3 RESULTS AND DISCUSSION
We found that there were 5 groups of KCs – “Low and Flat”, “No
learning”, “Still high”, “Too Little data” and “Good”. The “Low
and Flat” group indicated KCs where students likely received too
much practice. It appears that although students mastered the KCs
they continued to receive tasks for them. It may be better to
reduce the required number of tasks or change Andes’ knowledge
tracing parameters so that students get fewer opportunities with
these KCs. The “Still high” group suggests KCs, which students
continued to struggle with. Increasing opportunities for practice
for these KCs might improve the student model. The “No
learning” group indicated KCs where the slope of the predicted
learning curve showed no apparent learning. A step towards
improving the student model could be to explore whether each of
these KCs can be split into multiple KCs. The new KCs may
better reflect the variation in difficulty and transfer of learning
that may be happening across problem steps, which are currently
labeled by each KC. The KCs in the “Too Little data” group seem
to be KCs for which students were exposed to insufficient practice
opportunities for the data to be meaningful. For these KCs, adding
more tasks or merging similar KCs might provide data that is
interpretable. The KCs that appeared “Good” may reflect those in
which there was substantial student learning. Table 2 shows the
different group of KCs, their frequencies and AIC and BIC scores.
Figures 1a – 1d show the different groups of KCs. Intercept (logit)
and intercept (probability) both indicate KC difficulty. Higher
intercept values indicate more difficult KCs. The slope parameter
indicates the KC learning rate. Higher values suggest students will
learn such KCs faster.</p>
      <sec id="sec-1-1">
        <title>KC Name</title>
      </sec>
      <sec id="sec-1-2">
        <title>Intercept (logit) 0.06</title>
      </sec>
      <sec id="sec-1-3">
        <title>Intercept (probability) 0.52</title>
      </sec>
      <sec id="sec-1-4">
        <title>Slope 0.000</title>
      </sec>
      <sec id="sec-1-5">
        <title>Number of Students</title>
      </sec>
      <sec id="sec-1-6">
        <title>Number of Unique Steps</title>
      </sec>
      <sec id="sec-1-7">
        <title>Total Number of Steps</title>
      </sec>
      <sec id="sec-1-8">
        <title>Total Number of Transactions</title>
      </sec>
      <sec id="sec-1-9">
        <title>Total Student Hours # of Knowledge Component Model Low and</title>
        <p>Flat
2
AIC
BIC
# of Knowledge Components
2
5,857
71,300
107.02
34</p>
      </sec>
      <sec id="sec-1-10">
        <title>Intercept (logit) 1.77 0.63</title>
        <p>Intercept
(probability)
0.85
0.65
0.120
0.037</p>
      </sec>
      <sec id="sec-1-11">
        <title>Actual data predicted</title>
      </sec>
      <sec id="sec-1-12">
        <title>Intercept (logit) -0.28 -0.01</title>
      </sec>
      <sec id="sec-1-13">
        <title>Intercept (probability) 0.43 0.50</title>
      </sec>
      <sec id="sec-1-14">
        <title>Slope 0.000 0.000</title>
        <p>
          4 CONCLUSION AND FUTURE WORK
This paper presented how the AFM can be used to evaluate the
student model of the Andes Physics Tutor. Refining four of the
five groups of KCs identified, might improve the Andes student
model. A further approach would to use Learning Factors
Analysis [
          <xref ref-type="bibr" rid="ref1">1</xref>
          ] algorithm to automatically find better student models
by searching through a space of KC models. The next step is to
explore these options and measure their effect.
        </p>
      </sec>
    </sec>
    <sec id="sec-2">
      <title>5 ACKNOWLEDGMENTS</title>
      <p>Our thanks to the Pittsburgh Science of Learning Center for
providing the analysis tool for this work, to Bob Hausmann and
Kurt VanLehn for dataset access.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <surname>Koedinger</surname>
            ,
            <given-names>K.R.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>McLaughlin</surname>
            ,
            <given-names>E.A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Stamper</surname>
            ,
            <given-names>J.C.</given-names>
          </string-name>
          2012
          <source>Automated Student Model Improvement. Proceedings of the 5th International Conference on Educational Data Mining</source>
          , Chania, Greece, pp.
          <fpage>17</fpage>
          -
          <lpage>24</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <surname>Mathan</surname>
            <given-names>S.</given-names>
          </string-name>
          &amp;
          <string-name>
            <surname>Koedinger</surname>
            <given-names>K.</given-names>
          </string-name>
          <year>2005</year>
          .
          <article-title>Fostering the Intelligent Novice: Learning From Errors With Metacognitive Tutoring</article-title>
          .
          <source>Educational Psychologist</source>
          .
          <volume>40</volume>
          (
          <issue>4</issue>
          ), pps.
          <fpage>257</fpage>
          -
          <lpage>265</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <surname>Nwaigwe</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          &amp;
          <string-name>
            <surname>Koedinger</surname>
            ,
            <given-names>K.R.</given-names>
          </string-name>
          <year>2011</year>
          .
          <article-title>The Simple Location Heuristic is Better at Predicting Students' Changes in Error Rate Over Time Compared to the Simple Temporal Heuristic</article-title>
          .
          <source>Proceedings of the 4th International Conference on Educational Data Mining. Eindhoven</source>
          , Netherlands.
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <surname>VanLehn</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Lynch</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Schultz</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Shapiro</surname>
            ,
            <given-names>J. A.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Shelby</surname>
            ,
            <given-names>R. H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Taylor</surname>
          </string-name>
          , L., et al.
          <year>2005</year>
          .
          <article-title>The Andes physics tutoring system: Lessons learned</article-title>
          .
          <source>International Journal of Artificial Intelligence and Education</source>
          ,
          <volume>15</volume>
          (
          <issue>3</issue>
          ),
          <fpage>147</fpage>
          -
          <lpage>204</lpage>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>