<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Regularizing Student Parameters of Individualized Bayesian Knowledge Tracing Model</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Tracing</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Michael V. Yudelson Carnegie Mellon University 5000</institution>
          <addr-line>Forbes Ave. Pittsburgh, PA 15213 1(412) 268-5595</addr-line>
        </aff>
      </contrib-group>
      <abstract>
        <p>We propose a workflow to be implemented in a new workflows architecture of the LearnSphere We present the motivation and initial validation of the value of regularizing student-level parameters in an individualized Bayesian Knowledge Tracing model. Theoretically grounded, regularization of iBKT models, as we show, leads to increased model accuracy. Bayesian Knowledge Tracing (BKT) model individualization has been an active research topic recently. Individualization is accounting for student population variance that may come in different forms. One of the ways to individualize BKT is to introduce student-specific parameters. The new individualized BKT (iBKT) has been shown to fit the data better and even some promise, if deployed, to be able to save students time by better accounting for personal learning paths [7]. One of the toolkits that support iBKT models and is available via the LearnSphere workflow is hmm-scalable1. iBKT via hmmscalable, as compared to regular BKT [5], splits every BKT parameter into per-skill and per-student components and allows a subset or all BKT parameters to be individualized. Both studentand skill-level parameter are fit with the help one of the gradientbased approaches including stochastic or conjugate gradient descent, or a Lagrange updates method [6].</p>
      </abstract>
      <kwd-group>
        <kwd>Bayesian Knowledge regularization</kwd>
        <kwd>model-fitting</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. INTRODUCTION</title>
      <p>
        The solvers implemented in the core hmm-scalable utility that
the LearnSphere iBKT wrapper workflow uses are fitting both
skill and student parameters in fixed-factor fashion. However, it is
often advised to treat student-related variables as random factors.
The logic behind this is that students at hand are sampled from a
larger population of students [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]. Treating student factors as fixed
effects could lead to fixed fallacy [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ] and deflated generalizability
of the model.
      </p>
      <p>In this paper, we propose to take advantage of the regularization
feature in hmm-scalable. This feature adds a penalty to the
objective log-likelihood function in the form of the L2 norm. That
is an approximation of treating the regularized parameter as a
random factor. The influence of the L2 term is weighted by the λ
coefficient. The larger the λ, the stronger the effect of the
regularization. In simpler terms, regularization penalizes deviation
of parameter values from the centroid value. In regularized</p>
      <p>hmm-scalable toolkit,
https://iedms.github.io/standard-bkt
the
public
version
regression, the centroid is usually 0, and regularization takes a
form of scarcity-inducing mechanism. In the case of iBKT, as it is
implemented in hmm-scalable, an appropriate centroid is 0.5.
If student-level parameters have a value of 0.5, this effectively
means no student-level effect on the relevant still parameter.
Making the right choice of the λ that would lead to an improved
fit (and, potentially, the generalizability) is a problem in an of
itself. The hmm-scalable does not support the search for an
optimal λ directly. We propose to perform this search as part of
the wrapping LearnSphere workflow.
2.
2.1</p>
    </sec>
    <sec id="sec-2">
      <title>WORKFLOW</title>
    </sec>
    <sec id="sec-3">
      <title>Data Inputs</title>
    </sec>
    <sec id="sec-4">
      <title>METHOD</title>
      <p>The student-parameter-regularizing workflow takes the same data
inputs with regular BKT workflow. These are integer correctness
variable (correct encoded as 1, incorrect as 2), student as a
character factor, the problem as a character factor, and delimited
skill(s) as character factor.
2.2</p>
    </sec>
    <sec id="sec-5">
      <title>Workflow Model</title>
      <p>The workflow performs a sequence of runs of a chosen iBKT
model with a set of λ weights for the regularizing L2 parameter
penalty, where the first value is 0 (the default), and the rest could
be defined by the user. For example, from 0.1 to 4, with a 0.1
increment. In this model, the skill parameters are not regularized.
2.3</p>
    </sec>
    <sec id="sec-6">
      <title>Workflow Outputs</title>
      <p>The output of the workflow is built on the standard output of the
BKT workflow. An array of statistical fit metrics (Log-likelihood,
AIC, BIC, RMSE, Accuracy) is presented along with the λ weight
used for the regularization. The main outcome is whether the
regularization improves the fit of the model, as compared to the
non-regularized version, and at what λ value. A graphical
webfriendly plot of a selected via drop-down fit metric across all
tested λ values could help better visualize the effect of the
regularization. Skill and student level parameters could also be
printed for the winning model.</p>
    </sec>
    <sec id="sec-7">
      <title>3. DISCUSSION</title>
      <p>As part of pilot-testing the method we propose to be implemented
as a LearnSphere workflow, we have run a series of iBKT models
with student-level parameter regularizations for a set of λ weights.
Initially, we used values from 0 to 6 with a step of 1. Later, we
added values 0.1 to 0.9 with a step of 0.1, 1.5, 2.5, and 3.5.
We used the 2010 KDD Cup Challenge Set B (available via PSLC
DataShop2), the largest available dataset of student learning data.
2 2010 KDD Cup https://pslcdatashop.web.cmu.edu/KDDCup
The overall number of student-step transactions in that dataset is
over 20 million. The data in that dataset was collected by the
application called Cognitive Tutor. It was donated by Carnegie
Learning Inc. and contained data of students working on a Bridge
to Algebra course.</p>
      <p>
        For each setting, we performed 5 runs of 2-fold student-stratified
cross-validation. As identified in [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ], this setup of the
crossvalidation allows for the most potent way to reliably rank
alternative models when using cross-validation. One of the
guarantees of this cross-validation approach is appropriately sized
confidence interval around the accuracy metrics.
      </p>
      <p>
        For the sake of selecting the best λ, we used simple average across
5x2=10 statistical goodness of fit values. We picked RMSE and
Accuracy as our metrics. Figures 1 and 2 show the plots of the
RMSE and Accuracy (respectively) for the set of tested λ weights.
Although RMSE and Accuracy plots are not entirely in sync: the
best (lowest) RMSE value corresponds to λ=2.0, and the best
(highest) Accuracy value corresponds to λ=4.0, it is, arguably,
appropriate to choose λ=2.0, thus trusting RMSE metric.
To underline the usefulness of the regularizing student parameters
in iBKT models, we have performed the model ranking F-test
described in [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ] by making pairwise comparisons between
multiple student-stratified 2-fold cross-validated models. We have
used a Majority Class model that always predicted correct
outcome, a shipped model that used parameters from the deployed
Cognitive Tutor, a standard BKT model fit by using
ExpectationMaximization (EM) method implemented in hmm-scalable,
and iBKT non-regularized, and regularized models fit using
gradient-based Lagrangian updates solver. The results of the
pairwise comparisons are given in Table 1.
p-value of the difference
Here we can see that BKT models reliably (with a p-value less
than 0.001) beat Majority Class and Shipped models both in terms
of Accuracy and RMSE. In terms of accuracy, iBKT models beat
standard BKT, while the difference between the best-regularized
iBKR and non-regularized iBKT is not significant.
      </p>
      <p>If we look at RMSE, BKT models, again, prevail. There is now
the statistical difference between iBKT models: regularized iBKT
model has an edge. Standard BKT model fit using EM is not
statistically different from either iBKT model in terms of RMSE.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <surname>Alpaydin</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          (
          <year>1999</year>
          )
          <article-title>Combined 5 × 2 cv F test for comparing supervised classification learning algorithms</article-title>
          .
          <source>Neural Computation</source>
          ,
          <volume>11</volume>
          (
          <issue>8</issue>
          ),
          <fpage>1885</fpage>
          -
          <lpage>1892</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <surname>Baayen</surname>
            ,
            <given-names>R. H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Davidson</surname>
            ,
            <given-names>D. J.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>Bates</surname>
            <given-names>D. M.</given-names>
          </string-name>
          (
          <year>2008</year>
          )
          <article-title>Mixed-effects modeling with crossed random effects for subjects and items</article-title>
          .
          <source>Journal of Memory and Language</source>
          ,
          <volume>59</volume>
          (
          <issue>4</issue>
          ),
          <fpage>390</fpage>
          -
          <lpage>412</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <surname>Cen</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Koedinger</surname>
            ,
            <given-names>K. R.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>Junker</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          (
          <year>2008</year>
          )
          <article-title>Comparing Two IRT Models for Conjunctive Skills</article-title>
          . In Woolf,
          <string-name>
            <given-names>B.P</given-names>
            , A ̈ımeur, E.,
            <surname>Nkambou</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            , and
            <surname>Lajoie</surname>
          </string-name>
          , S. (Eds.),
          <source>Proceedings of the 9th international conference on Intelligent Tutoring Systems (ITS '08)</source>
          , Springer-Verlag, Berlin/Heidelberg, (pp.
          <fpage>796</fpage>
          -
          <lpage>798</lpage>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <surname>Clark</surname>
            ,
            <given-names>H. H.</given-names>
          </string-name>
          (
          <year>1973</year>
          ).
          <article-title>The language-as-fixed-effect fallacy: A critique of language statistics in psychological research</article-title>
          .
          <source>Journal of Verbal Learning and Verbal Behavior</source>
          ,
          <volume>12</volume>
          ,
          <fpage>335</fpage>
          -
          <lpage>359</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <surname>Corbett</surname>
            ,
            <given-names>A. T.</given-names>
          </string-name>
          and
          <string-name>
            <surname>Anderson</surname>
            ,
            <given-names>J. R.</given-names>
          </string-name>
          (
          <year>1995</year>
          )
          <article-title>Knowledge tracing: Modeling the acquisition of procedural knowledge. User Modeling</article-title>
          and
          <string-name>
            <surname>User-Adapted Interaction</surname>
          </string-name>
          ,
          <volume>4</volume>
          (
          <issue>4</issue>
          ),
          <fpage>253</fpage>
          -
          <lpage>278</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <surname>Levinson</surname>
            ,
            <given-names>S. E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Rabiner</surname>
            ,
            <given-names>L. R.</given-names>
          </string-name>
          , and
          <string-name>
            <surname>Sondhi</surname>
            ,
            <given-names>M. M.:</given-names>
          </string-name>
          <article-title>An Introduction to the Appli- cation of the Theory of Probabilistic Functions of a Markov Process to Automatic Speech Recognition</article-title>
          .
          <source>Bell System Technical Journal</source>
          ,
          <volume>62</volume>
          (
          <issue>4</issue>
          ):
          <fpage>1035</fpage>
          -
          <lpage>1074</lpage>
          . (
          <year>1983</year>
          )
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <surname>Yudelson</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Koedinger</surname>
            ,
            <given-names>K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Gordon</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          (
          <year>2013</year>
          )
          <article-title>Individualized Bayesian Knowledge Tracing Models</article-title>
          . In: Lane,
          <string-name>
            <given-names>H.C.</given-names>
            , and
            <surname>Yacef</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            ,
            <surname>Mostow</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            ,
            <surname>Pavlik</surname>
          </string-name>
          , P.I. (eds.)
          <source>Proceedings of 16th International Conference on Artificial Intelligence in Education (AIED</source>
          <year>2013</year>
          ), Memphis, TN. LNCS vol.
          <volume>7926</volume>
          , (pp.
          <fpage>171</fpage>
          -
          <lpage>180</lpage>
          ).
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>