<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Statistical Approaches to the Model Comparison Task in Learning Analytics</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Josh Gardner</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Christopher Brooks</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>School of Information, University of Michigan</institution>
          ,
          <country country="US">USA</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>Comparing the performance of predictive models of student success has become a central task in the eld of learning analytics. In this paper, we argue that research seeking to compare two predictive models requires a sound statistical approach for drawing valid inferences about comparisons between the performance of such models. We present an overview of work from the statistics and machine learning communities and evaluate several methodological approaches, highlighting four approaches that are suitable for the model comparison task. We apply two of these methods to a learning analytics dataset from a MOOC conducted at the University of Michigan, providing open-source code in R for reproducing this analysis on other datasets. We o er several practical considerations for future implementations, as well as suggestions for future research in both the learning analytics community and the broader eld of machine learning.</p>
      </abstract>
      <kwd-group>
        <kwd>Learning analytics</kwd>
        <kwd>predictive modeling</kwd>
        <kwd>machine learning</kwd>
        <kwd>model evaluation</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>Identifying e ective models for the prediction of student success is one of the cornerstone tasks in the
eld of learning analytics. For a eld concerned with understanding and improving the educational process,
identifying students for whom those improvements need to be made is a critical challenge. Additionally, as
the elds of learning analytics and educational data mining have matured, so too have the underlying data
mining and machine learning techniques, as well as toolkits and hardware which are used to apply these
techniques broadly. As a result, it is now possible for even non-experts to leverage machine learning rapidly,
and to build predictive models without having to understand details of techniques or implementations thereof
{ a development that has had an undoubtedly positive overall impact on the study and practice of learning
analytics.</p>
      <p>Prior learning analytics research comparing the performance of supervised learning models to predict
student success has been extensive, particularly in the past four years, but the use of statistical techniques
to compare these models has been mixed. In a review of recent research produced by the International
Educational Data Mining Society (2013-2016) and the International Conference on Learning Analytics and
Knowledge (2015-2016), we uncovered 12 papers speci cally focused on using student behavioral data to
predict either pass/fail or dropout outcomes. However, of these 12 papers, seven reported simple matrices of
model performance metrics using estimates obtained from cross-validation on the training set{an approach
which we will argue below is not appropriate for model comparison{with most also using inferential language
not appropriate for the methodology of their comparison. This nding was, unfortunately, not surprising, and
we nd it to be generally representative of work across the eld. Many otherwise excellent learning analytics
research projects simply present a matrix of accuracies or model performance statistics and conclude that the
best or most preferable modeling method is the one with the highest accuracy, AUC, sensitivity/speci city,
F-score (or whatever the outcome of interest may be) on their training dataset(s)-entirely passing over the
task of evaluating these di erences in performance. Correcting this problem by adopting statistically sound
methods for the comparison of predictive models will allow the learning analytics community to build an
enduring, reproducible base of knowledge about methods for predicting student success.</p>
      <p>This paper aims to begin this methodological discussion with the broader community by focusing on
the two-model comparison task: determining whether two models di er in performance on a given dataset.
In Section 2, we situate the challenge of comparing predictive models of student success within the larger
strands of research in the learning analytics community and in the larger machine learning and statistics
communities, presenting an overview of research on the two-model comparison task. In Section 3, we dive
deeper into this task, presenting a framework for understanding the task of comparing supervised machine
learning models. We provide an overview of research on methods for statistically valid, high-reproducibility
evaluation of performance di erences between classi ers, and present methods which are not valid for the
two-model task. In Section 4, we continue this discussion, detailing statistically valid methods for the
two-model task. In Section 5, we produce a sample application of two of these valid tests to a real-world
MOOC dataset, including sample code written in R. In Section 6, we provide practical recommendations
and additional considerations for learning analytics researchers implementing these methods in practice.
Finally, we conclude and present directions for future research in Section 7.</p>
      <p>
        This paper speci cally addresses problems that emerge when researchers have a single dataset, and need
to use this dataset to both build models and obtain reliable estimates and comparisons of model performance.
While we can easily obtain an unbiased estimate of the mean and variance of the di erence if there is a
su cient supply of data and a holdout set is used [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ], in practice this approach is often neither practical nor
desirable in the context of learning analytics, giving rise to the problems addressed in this paper. Learning
analytics researchers often work with datasets which are limited in size, and face a tension between (a)
using as much data as possible for model training to avoid starving statistical models of the data necessary
for optimal performance, and (b) reserving data for model evaluation to obtain reliable estimates of this
performance.
      </p>
      <p>These situations, both characterized by limited data, are particularly prevalent in learning analytics and
educational data mining. The privacy restrictions surrounding educational data limit learning analytics and
educational data mining researchers more than machine learning researchers in many other elds.
Additionally, researchers are often speci cally interested in building models for narrow foci - such as admissions
to a given university - and not interested in the models' general performance across institutions, leading
to datasets which are limited in size with no option for collecting additional data. As such, adopting and
consistently implementing a set of statistically robust techniques to evaluate classi ers in the context of
learning analytics is especially critical.</p>
    </sec>
    <sec id="sec-2">
      <title>2. Prior Research</title>
      <p>
        The broader eld of machine learning research, having confronted the task of model comparison and
evaluation earlier than the eld of learning analytics, has a relatively robust literature of both experimental
and theoretical results on methods for obtaining reliable estimates of model performance. Resampling was
the accepted method for generating a sample of estimates for applying a t-test throughout much of the 1990s
[
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]. However, several researchers began to investigate the soundness of such an approach in the 1990s [
        <xref ref-type="bibr" rid="ref3 ref4 ref5">3, 4, 5</xref>
        ].
[
        <xref ref-type="bibr" rid="ref6">6</xref>
        ] provided a seminal early analysis of ve statistical tests for comparing supervised classi cation learning
algorithms, presenting experimental results from simulated data to estimate the empirical performance
of several common statistical tests. This work formed the foundation for future analyses and presented
several results that endure today, including the problems with unadjusted resampled t-tests discussed below.
[
        <xref ref-type="bibr" rid="ref7">7</xref>
        ] extended this analysis, proposing a modi cation to the t-statistic for resampling (shown below) and
demonstrating both a theoretical justi cation and empirical evidence of its e ectiveness. Another important
dimension to statistical tests for model comparisons is introduced in [
        <xref ref-type="bibr" rid="ref1 ref2 ref8">1, 2, 8</xref>
        ] by formalizing measurements of
replicability (in addition to power and Type I error rate). These papers present further empirical evidence
about the performance of several sampling schemes and test statistics on both simulated and real datasets.
Several key ndings from this literature are discussed in detail in Section 3 below; for additional reading,
see references.
      </p>
      <p>
        Further research extends these conversations to the multi-model comparison task, including [
        <xref ref-type="bibr" rid="ref10 ref11 ref12 ref9">9, 10, 11, 12</xref>
        ].
This work provides insight into the issue of comparing multiple models, and even doing so across multiple
datasets or domains. Extending the analysis in this paper to multiple models ts the most realistic cases
in which researchers encounter these types of comparison tasks; however, the fundamental issues are the
same as the two-model case on which we focus herein. While the multiple comparisons case is analogous, it
utilizes tests which speci cally correct for multiple comparisons (ANOVA, Newman-Keuls, etc.) { and once a
statistically rigorous approach to model comparison has been adopted, taking the step from the 2-model case
to the n-model case requires only minor methodological changes. The multiple-model-multiple-dataset case,
by contrast, faces an entirely di erent set of considerations than the task evaluated here [
        <xref ref-type="bibr" rid="ref11 ref12 ref13 ref14 ref9">11, 9, 13, 12, 14</xref>
        ].
A thorough discussion is beyond the scope of this paper, although the authors intend to address it in future
work.
      </p>
    </sec>
    <sec id="sec-3">
      <title>3. Statistical Comparisons of Classi ers</title>
      <sec id="sec-3-1">
        <title>3.1. Basic Terminology</title>
        <p>In this paper, we address the problem of comparing the performance of two learning algorithms1, LA
and LB, on a single dataset D. Typically, the di erence in model performance is estimated using a sample
of di erences in performance of LA and LB, PA and PB, where the observed di erences x1; :::; xn = PA;1
PB;1; :::; PA;n PB;n across several iterations of model training and testing, formed on multiple training and
testing partitions Dt;1:::Dt;n and D=Dt;1:::D=Dt;n respectively (each training partition Dt;n might be a single
random subsample or a set of training folds using k-fold cross-validation; the corresponding test partition
D=Dt;n is the test dataset). Performance can be measured using any metric and loss function which can be
calculated from trained models using the testing data, including accuracy, area under the Receiver Operating
Characteristic curve (AUC), sensitivity, speci city, etc. Often x1; :::; xn is collected using multiple runs, R,
of a partitioning method (such as k-fold cross-validation), where the same randomized partitioning method
is applied in each of k folds or iterations for each run. This leads to a total of R k = n samples in x1; :::; xn.</p>
        <p>
          In a hypothesis test, we use this sample to draw conclusions about the underlying population from which
it was drawn. To do so, we calculate a test statistic, referred to in the literature as t or z, and, utilizing
assumptions about the population from which our sample was drawn, make a determination as to whether
the observed value of the test statistic provides su cient evidence to reject a null hypothesis, H0 (which in
this case holds that the true average of the di erences in performance is zero; that is H0 : = 0). Speci cally,
we calculate the probability of observing the test statistic under H0, and reject H0 if this probability is less
than the desired level of signi cance . A Type I error occurs when a hypothesis test rejects H0 when
it is in fact true; a Type II error occurs when a hypothesis test accepts H0 when it is in fact false. The
expected frequency of a Type I error is ; this is also referred to as the \acceptable" level of Type I error
throughout this paper and the broader literature [
          <xref ref-type="bibr" rid="ref6">6</xref>
          ]. The power of a test is related to its Type II error rate,
and represents its ability to correctly lead us to reject H0 when it is truly false (power measures how well
we identify true e ects when they do indeed exist).
        </p>
      </sec>
      <sec id="sec-3-2">
        <title>3.2. Model Evaluation as Experiment</title>
        <p>When comparing and evaluating supervised machine learning methods, it is useful to think about model
evaluation as an experiment. Indeed, much of the research discussed above is grounded in this language,
and it provides a familiar framework for understanding the challenges of classi er performance comparison
and the methods used to address them. This experiment has two parts:
(1) a sampling component, where we collect a set of samples of di erences in the performance of
two learners, and
(2) a hypothesis testing component where these samples are used to evaluate hypotheses about
comparisons between the performance of learning algorithms LA and LB on D.</p>
        <p>1These learning algorithms are also referred to as \learners" throughout as is common in machine learning literature, but
this is not to be confused with the use of \learners" to refer to students in a course as is common in educational literature. In
the context of this paper we always refer to a \learner" as a machine learning algorithm.</p>
        <p>The method adopted in (2) typically entails a set of assumptions about the underlying distribution from
which the sample in (1) is drawn; often this includes an assumption that the observations are independent
and identically distributed draws from a normal distribution.</p>
        <p>The challenge of model comparison in the current context primarily occurs at (2), the hypothesis testing
stage: in many (but not all) cases, the way model performance data is collected { using a series of overlapping
training and testing datasets, such as in k-fold cross validation { renders the analogy to simple random
sampling invalid. In particular, the underlying assumption of independence between the observations in
the sample relied upon by common testing methods (such as the Student's t-test) is violated. In this
case, one must adjust either the approach to (1) so that the sampling more closely meets the assumptions
of the hypothesis test used, or adjust (2) to more accurately re ect the distribution of the experimental
sample. Below, we discuss a series of techniques for improving both (1) and (2), reviewing the literature
to highlight tests which are not appropriate for the model comparison task as well as those which are
well-suited to it. While there is no single best approach to this challenge and further research is needed,
adopting these methods { and encouraging further conversations within the learning analytics community {
will represent tremendous progress as the eld attempts to move toward consensus around e ective student
success modeling techniques.</p>
      </sec>
      <sec id="sec-3-3">
        <title>3.3. The Two-Model Comparison Task; Hypothesis Testing for Model Comparison</title>
        <p>
          While the statistical community has developed a robust set of tools for statistical inference and hypothesis
testing, these tools often do not extend directly to the task of model comparison (or, perhaps more accurately,
the model comparison task does not conform to these tests). In practice, we often implement hypothesis tests
whose assumptions are not entirely met, but are closely approximated by a sample. When the properties of
this sample deviate too far from these assumptions, it can produce undesirable results, including elevated
probability of rejecting the null hypothesis when it is true (Type I error) or decreased ability to discern
a true e ect when it exists (low power). Empirical analyses have helped to determine exactly how much
various hypothesis tests are impacted by samples of predictive model performance which do not fully match
their assumptions, as measured by their Type I and II errors, power, and reproducibility [
          <xref ref-type="bibr" rid="ref1 ref2 ref6 ref8">6, 8, 2, 1</xref>
          ].
        </p>
        <p>Throughout this paper, we apply statistical testing in the context of what we will refer to as the two-model
comparison task, or simply the model comparison task. This is the task that this paper seeks to provide
e ective tools for evaluating:</p>
        <p>The two-model comparison task: Given dataset D and two statistical learners LA and LB,
identify whether the two learners di er in their performance in predicting the class labels of D.</p>
        <p>This task is distinct from, but related to, the determination of whether one model performs \better" than
another (the details of this distinction are discussed with respect to one-sided and two-sided tests below).</p>
        <p>The challenge in this task in the situations described above is that we would like to use as much of our
limited dataset D as possible in each iteration to obtain a realistic estimate of model performance (because
withholding available data would either deprive L of useful training data in that would improve its test
performance P , or testing data that could be used to improve the estimate of this performance). As a
result, we devise methods to form overlapping training sets Dt;1:::Dt;nsuch that each estimate PA;1; :::; PA;n
and PB;1; :::; PB;n is not independent from the others (methods for building these training sets include
resampling and k-fold cross-validation, discussed below). If we conduct hypothesis tests under the incorrect
assumption that our data are independent, we will make errors more often than expected due to chance,
and will not have a reliable estimate of the frequency of making such errors. To correct our probability
estimates, we can adjust how we build our sample (how each successive Dt is formed, or how we calculate
Pm;nfrom it), or adjust how we evaluate the sample (how t is calculated and tested). Below, we rst discuss
a series of methods that, in various ways, fail to accurately draw inferences about populations when applied
to the model comparison task, and then discuss statistically valid approaches to this task. Speci cally,
in the following subsection, \Invalid Methods for Model Comparison," we catalogue several problematic
approaches to the model comparison test. In the next section, \Valid Methods for Model Comparison," we
highlight successful attempts to correct some of the methods considered previously. Finally, we present an
application of these valid methods to a MOOC dataset in our case study.</p>
      </sec>
      <sec id="sec-3-4">
        <title>3.4. Invalid Methods for Model Comparison</title>
        <p>A model evaluation approach consists of both a sampling method and a hypothesis testing method;
these together determine whether the approach is statistically valid for the model comparison task. The
approaches evaluated below, and indeed any such approach, might be statistically invalid for a number of
reasons, but regardless of the underlying mechanism they will typically display either (a) an in ated Type I
error rate, (b) low statistical power, or (c) low reproducibility.2 Methods are typically evaluated using these
performance rates and by some combination of theoretical and experimental evaluation on real or simulated
data.</p>
        <p>Perhaps the most common approach to the model comparison task, both in the eld of learning analytics
and in the broader eld of applied machine learning research, is using no statistical test. Often researchers,
having taken the appropriate steps to collect their data, train a set of learners LA and LB, estimate their
performance on D, and present the estimates of performance (and their di erences) to speak for themselves.
Little needs to be said about this approach, which amounts to simply displaying sample averages, other
than that it provides insu cient evidence about whether there exist true di erences in model performance.
This approach leaves entirely unaddressed the question of whether such results may be spurious and caused
simply by randomness inherent in the sampling process.</p>
        <p>Another common approach to model comparison involves using repeated random sampling to build a
successive series of training and testing sets from D and then applying a hypothesis test to the results of
model training and testing on these samples. The intuition behind such approaches is that if the data is
shu ed su ciently between runs and an adequate number of model performance estimates are built on these
successive shu ings, the observed di erences in model performance x1; :::; xn = PA;1 PB;1; :::; PA;n PB;n
will conform to (or closely approximate) the assumptions of the test statistic used to evaluate its distribution.
For several randomization methods and the test statistics they are generally evaluated with, however, this
is not the case.</p>
        <p>
          One such method is the resampling method. Resampling consists of repeatedly (n times) drawing a
speci ed number of observations (typically two thirds of the training data [
          <xref ref-type="bibr" rid="ref6">6</xref>
          ]) without replacement to form
training sets Dt;1:::Dt;n with the remaining data in each iteration used for model evaluation. Both learners
LA and LB are trained and evaluated on the respective training and testing sets, and their di erences
in performance x1; :::; xn = PA;1 PB;1; :::; PA;n PB;n are used as estimates of the di erence in model
performance. The sample population x1; :::; xn is used to compute a test statistic t with a Student's
tdistribution with n 1 degrees of freedom under the assumption that the individual di erences x1; :::; xn
were drawn independently from a normal distribution.
        </p>
        <p>
          However, these di erences will not conform to a normal distribution, for two reasons. First, PA;1and PB;1
are not independent: they both use identical test/training sets and are correlated with one another. Second,
the training and test sets overlap across trials, to varying degrees (depending on the size of each resample).
While this approach can be adjusted to more closely match the assumptions of the test statistic, this is not
the case with the unadjusted t-test often used with resampled data. This approach is shown to have an
extremely high Type I error rate (indeed, the highest of all methods evaluated in [
          <xref ref-type="bibr" rid="ref6">6</xref>
          ]). An additional concern
with this method is that the di erence between the samples can always be made statistically signi cant
by increasing the number of resamples (an undesirable and problematic property for a statistical test). A
similar method is the bootstrap method. Bootstrapping di ers from the resampling approach only in
that the observations drawn from D to form Dt are drawn with replacement. Similar to the resampling
method, then, the bootstrap method produces training datasets with a high degree of overlap, and it has
been observed to produce in ated Type I error rates [
          <xref ref-type="bibr" rid="ref13">13</xref>
          ].
        </p>
        <p>
          10-fold cross validation with an unadjusted t-test is another randomization approach that has proven
to be problematic for comparing model performance. In 10-fold cross validation, the dataset is partitioned
into k = 10 random folds, with a series of learners LA;1:::LA;10 and LB;1:::LB;10 trained on a set Dt;k
consisting of a subset of k 1 folds and evaluated on the kth fold. Due to the fact that each training set Dt;k
2Replicability is de ned as how well the results of an experiment can be reproduced, or the degree to which the outcome of
a test depends on the particular random partitioning of the data used to perform it [
          <xref ref-type="bibr" rid="ref1 ref2">1, 2</xref>
          ].
shares k 2 folds (and therefore 80% of the available training data) with each other training set in a given
run, this method also produces high rates of Type I error [
          <xref ref-type="bibr" rid="ref2 ref6 ref8">8, 2, 6</xref>
          ]. This is not due to an inherent aw in
the cross-validation approach itself; several authors have o ered modi cations to both the cross-validation
process and its test statistic which will be discussed below.
        </p>
        <p>In contrast to randomization-based approaches, which prescribe an approach to sampling the model
performance data, other methods utilize di erent test statistics in an e ort to produce more reliable evaluations
of a given sample. One such method is the di erence-in-proportions test. In this test, a test statistic,
z, is calculated based on a simple comparison between the overall error rates PA and PB, of the algorithms
LAand LB, with p = PA 2 PB being the average of the two error rates:
z =</p>
        <p>PA
p2p(1</p>
        <p>PB
p)=n</p>
        <p>
          However, this test faces similar (and additional) problems as those discussed previously. First, the
unadjusted z-statistic still assumes independence between the measured error rates PA and PB, an assumption
which is violated by the identical training sets used to estimate PA and PB. [
          <xref ref-type="bibr" rid="ref6">6</xref>
          ] identi es potential corrections
to this statistic, but observes that these corrections are rarely used in practice and does not empirically test
the corrected performance. Additionally, even if corrections for non-independence are applied, the di erence
in proportions test is unable to account for variation resulting from the choice of training sets: it only makes
a single comparison between the model error rates PA and PB, and thus is structurally unable to measure
variation that occurs with multiple successive runs of model training and testing.
        </p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>4. Valid Methods for Model Comparison</title>
      <p>The main goal of this paper is to describe and implement statistically valid methods for the two-model
comparison task. While several such methods are described in this section, we select two for detailed
description and implementation in the next section of this paper, with a speci c emphasis on methods
which match the practical conditions under which learning analytics researchers might apply these tests.</p>
      <p>
        The corrected resampled t-test is a method that has seen wide adoption in the machine learning
community. [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ], noting that the in ated Type I Error rate observed by [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ] was due to an underestimation
of the variance of resamples, proposed a correction to the t-test (shown in Table 1 below). They propose an
adjustment to the estimated variance of a traditional t-statistic, from
to
where the number of samples J = k r, S2j is the sample variance of the estimates, n1 is the number of
samples used for training, and n2 is the number of samples used for testing. This modi cation is intended
to account for the correlation between resamples, which the resampled t-test \boldly" assumes to be zero,
by instead estimating that correlation as = n1n+2n2 [14; speci cally see sections 3 and 4 for discussion
and theoretical justi cation of this modi cation]. This modi ed test statistic is used with the Student's
t-distribution and n 1 degrees of freedom, in otherwise exactly the same way as an unadjusted Student's
t-statistic. This adjustment has been shown to have acceptable levels of Type I error and high replicability,
particularly when using many (R = 100) runs [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. Additionally, the correction avoids another critical aw
with the unadjusted resampled t-test, mentioned above, that the signi cance of the test statistic can be
increased without bound simply by increasing k (the number of resamples).
      </p>
      <p>
        This same corrected test statistic can be used for cross-validation in a general approach we refer to as
corrected cross-validation [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ]. Cross-validation di ers from resampling in at least two ways that are
relevant to statistical testing. First, the test and training sets are guaranteed not to overlap within each
(1)
(2)
(3)
fold. Second, while training sets do overlap across folds (each training set shares k 2 folds of data with
the others), they do so in a consistent way. Together, due at least in part to these two factors, it has been
empirically demonstrated that a corrected cross-validation approach achieves superior performance relative
to the corrected resampled t-test [
        <xref ref-type="bibr" rid="ref1 ref2 ref6">2, 1, 6</xref>
        ].
      </p>
      <p>
        Several proposals have o ered slightly modi ed versions of the corrected cross-validation approach with
di erent corrections to t, and speci c proposed values for R and k. A 5x2 cross-validation approach (R =
5; k = 2) was proposed in [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ], seeking to avoid correlation between the accuracy measured across folds
by using only two folds, which results in nonoverlapping training sets within each run. When this 5x2cv
paired t-test is used in combination with a modi ed test statistic t~ with 5 degrees of freedom (shown in
table 1), it achieves an acceptable level of Type I error [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ]. However, it has been widely recognized that this
test achieves low power [
        <xref ref-type="bibr" rid="ref15 ref6 ref8">15, 6, 8</xref>
        ]. Additionally, the test has been shown to have low replicability due to the
large random variation inherent in using only 2 folds per run and the fact that the modi ed test statistic t~
depends only on the di erence x1; 1 and not on the full set of di erences x1;1; :::; xr;k [
        <xref ref-type="bibr" rid="ref1 ref6 ref8">8, 1, 6</xref>
        ]. Additionally,
despite the improved performance, the choice of R = 5 is somewhat arbitrary, and has been criticized as ad
hoc and lacking in theoretical justi cation [
        <xref ref-type="bibr" rid="ref15 ref16">15, 16</xref>
        ].
      </p>
      <p>
        Another proposal that has been shown to have both better power and higher replicability than the
5x2cv approach is 10x10cv with a variance correction, known as the corrected repeated k-fold cv test.
This test utilizes a variance correction similar to the corrected resampled t-test, but pairs the modi ed test
statistic with a cross-validation approach. Together, these modi cations achieve an acceptable Type I error
rate (equal to , as expected) and better statistical power than 5x2cv [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. Notably, in a direct comparison
between corrected resampling with R = 100 and 10x10cv (both of which require 100 iterations of model
training and thus the same computational e ort), 10x10cv with the corrected repeated k-fold cv test achieves
higher replicability [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ].
      </p>
      <p>
        A novel sampling-based approach to the model comparison task, evaluated in [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ], shows acceptable Type I
Error, high power, and high replicability: the sorted runs sampling scheme. This is a modi ed approach
to cross-validation, in which the results PR;1; :::; PR;k for each run are rst sorted, and then the averages
of the ordered folds are taken according to their respective rank across runs. This results in a sample of k
estimates for each learner, where P1 = R1 PR
      </p>
      <p>i=1 min(P1:::k) is the average of the lowest scores for a given
learner in each of the R runs, P2 is the average of the second-lowest score for a learner in each run, etc. An
illustration of this scheme is shown in Figure 1.</p>
      <p>
        The sorted runs sampling scheme has been shown to achieve a better Type I Error rate than 10x10cv,
but at the cost of reduced power [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ]. The sorted runs sampling scheme also has the appealing result that it
results in a sample for which the independence assumption is not heavily violated, so that no correction in
variance or degrees of freedom is required, allowing it to be used with a normal, uncorrected t-statistic and
df = n 1.
      </p>
      <p>Below, we present an implementation of the 10x10cv with the corrected repeated k-fold cv test
and the sorted runs sampling scheme with a standard t-test. We select the 10x10cv as an example
of an approach that can be easily applied to the cross-validation that is already common in most predictive
models of student performance, favoring 10x10cv over 5x2cv because of the former's greater statistical power,
and over the corrected resampling approach because of its greater reproducibility. Additionally, using 10
folds (instead of 2) substantially reduces the pessimistic bias of the model performance estimates Pm;k on
each fold, because the models are trained on datasets closer in size to that of the actual dataset we are
estimating their performance on, jDj. We utilize the sorted runs sampling scheme both because it has been
demonstrated to perform at least as well as 10x10cv with the adjusted t-statistic, and because it shows
even better reproducibility { one of the key components of an e ective model comparison test discussed
above. However, there are situations in which one might cautiously utilize an alternative approach (such
as the adjusted resampled t-test or 5x2cv) instead of these approaches, particularly if training two learners
a total of 100 times each (R k = 100) is not practical. Additional e ective methods for evaluating model
performance are mentioned in sections 5 and 6 below.</p>
    </sec>
    <sec id="sec-5">
      <title>5. Case Study - Evaluating Classi ers on MOOC Data</title>
      <p>In this section, we present a case study in model comparison to illustrate the application of the two
approaches identi ed above that are both statistically and practically appropriate for comparing classi cation
algorithms in this case. These are the corrected repeated k-fold cross-validation test with 10x10cv, and the
sorted runs sampling scheme with a paired t-test.</p>
      <p>
        Our dataset is compiled from a nancial literacy MOOC run on the Coursera platform through the
University of Michigan between February and May 2014. From the raw Coursera clickstream logs, a set of
features similar to that of [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ] was assembled. In particular, an appended featureset was used, where the
same set of p features is collected for each week w of the course, leading to a \wider" dataset each successive
week. The intuition behind this approach is that it expands the feature space for temporal prediction
problems, allowing the learner to exploit more information about each student and their activity patterns
as the course proceeds. This particular analysis utilized an appended feature set representing the learner
      </p>
      <sec id="sec-5-1">
        <title>Quiz</title>
        <p>tempts
Quiz Exams
Quizzes
Human Graded</p>
      </sec>
      <sec id="sec-5-2">
        <title>Forum Posts</title>
      </sec>
      <sec id="sec-5-3">
        <title>Direct Nodes</title>
      </sec>
      <sec id="sec-5-4">
        <title>Thread Nodes</title>
      </sec>
      <sec id="sec-5-5">
        <title>Count of attempted exam questions. Count of attempted human-graded quiz questions (0 for all learners; no human-graded quizzes in this course).</title>
        <p>Count of forum posts (this includes both
forum posts and comments, which Coursera
tracks separately).</p>
        <p>Number of distinct users a given user
responded to on the forums (direct-reply).</p>
        <p>
          Number of distinct uses a given user posted in
the same forum with (thread-reply).
activity for the rst two complete weeks of the course. The original dataset consisted of 51,088 observations
of 16 features (8 per week) based on those of [
          <xref ref-type="bibr" rid="ref10">10</xref>
          ]. Students who had shown no activity in the rst two weeks
of the course were removed as inactive students, because including them makes the prediction problem too
easy (this is a common approach; as an example, see [
          <xref ref-type="bibr" rid="ref17">17</xref>
          ]). 75% of these remaining observations were used for
the model training and evaluation (to retain a holdout set for a robustness check comparing hypothesis test
results to true test performance or for future research) resulting in a dataset with jDj = 23; 537 observations.
The feature set was used to predict a binary variable for whether a learner would register any activity in
the third week of the course. 4,074 observations were positive (i.e., dropouts), resulting in a dataset with a
positive:negative outcome class balance of approximately 1:5.
        </p>
        <p>The third week of activity was chosen as the target for prediction because it appears to be a turning
point in the course: many learners register no activity after the rst or second week of the course, so the
dropout rate is much higher than in subsequent weeks (and we might expect many of those learners to be
\explorers" not interested in course completion; see gure 2). Unlike the early dropouts, a learner who shows
activity in the rst two weeks, persisting to the third week, and then drops out is more likely to be making
an honest e ort at completion. Additionally, the large decrease in dropouts after week 3 suggests that a
higher proportion learners who persist past the third week are able and motivated to successfully complete
the course, and that e ective predictions (leading to e ective interventions) here could substantially impact
course completion.</p>
        <p>
          Two models { a multilayer perceptron (also known as a feedforward arti cial neural network) and a
classi cation tree { were t to the data3. Hyperparameters for the models are tuned internally by the caret
package in R, and the only hyperparameter setting that was adjusted was setting the number of layers for
the neural network (L = 3). These models were selected because both are used in student success models
(the former is gaining widespread use, and the latter is common in previous research) and both make few
3The multilayer perceptron used was from the implementation in the RSNNS package in R [
          <xref ref-type="bibr" rid="ref18">18</xref>
          ]; the classi cation tree was
an implementation from the rpart package which itself is based on Breiman's ori nal formulation [
          <xref ref-type="bibr" rid="ref19">19</xref>
          ]; the caret package was
used as a wrapper for the training/testing procedure [
          <xref ref-type="bibr" rid="ref7">7</xref>
          ]. We refer the interested reader to these citations for details on the
implementations of these models.
and weak (if any) assumptions about the underlying data and require no pre-training feature selection.
Both models were t and evaluated using the same subsamples, allowing for valid comparisons between the
models (see [
          <xref ref-type="bibr" rid="ref7">7</xref>
          ] for details on implementation in R). Caret provides statistics for each model's performance
across each individual resample, and our implementation uses identical resamples for both models. Because
both of the evaluation approaches identi ed above could utilize 10x10cv, this was used (the sorted runs
sampling scheme can be used with any number of folds, but the adjusted cross-validation approach achieved
the performance discussed above speci cally with 10x10cv).4
        </p>
        <p>The results in Table 3 show the sample averages for three di erent performance metrics-accuracy, area
under the Receiver Operating Characteristic, and speci city { and the resulting comparisons between them.
These results are intended to provide a multidimensional view of model performance and to give three
points of comparison for the di erent test statistics used. Additionally, the three performance metrics are
used to highlight di erent approaches to the challenge of the model comparison task in the context of
highly-unbalanced student performance datasets where a speci c class (such as a dropout or failure) is of
particular interest. Note that, in practice, conducting three di erent comparisons would require a correction
for multiple comparisons, such as a Bonferroni correction, and would e ectively be moving beyond the
two-model comparison task. Since these results are for demonstration purposes only, we will not provide a
correction.</p>
        <p>A detailed review of the results shows several important practical features of the two-model comparison
task and the two test statistics used to evaluate this comparison. In general, the two test statistics (corrected
10x10cv and sorted runs) provide similar p-value estimates. The corrected 10x10cv provides slightly more
conservative p-value estimates in each case, which re ects its more conservative variance correction. The
increased estimated variance for the 10x10cv approach may be more conservative than necessary (variance
too large, shrinking the test statistic), or the uncorrected variance of the sorted runs test statistic may
not be conservative enough (variance too small, in ating the test statistic). Interestingly, in this particular
application, the two test statistics yielded nearly identical numerators for each of the three comparisons
(accuracy, AUC, speci city), despite quite di erent methods for calculating those sample means (sorted
mean across each run, vs. mean across all folds within run). The di erence in p-value for the sorted runs
test statistic, then, is entirely due to the di erence in estimated variance and the di erence in degrees of
freedom used to calculate p-values from the test statistics. Note that both test statistics use the normal
calculation for a Student's t-distribution (with degrees of freedom calculated as shown in Table 2).</p>
        <p>4Sample code for implementing this model comparison is available in the following repository: https://bitbucket.org/
jpgard/lak-mla-2017/src</p>
        <p>The two test statistics track each other closely, with both yielding their lowest and highest p-values for
the same comparisons (AUC and speci city, respectively). Additionally, the di erent values of these test
statistics across metrics clearly re ect the fact that we are not always warranted in concluding that an
observed di erence is likely to re ect a true di erence in model performance. This is a desirable property
of null hypothesis statistical testing, which allows us to distinguish between true and potentially spurious
results. Particularly in the case of speci city (the proportion of true positives, i.e. dropouts, which are
correctly identi ed as positives), which is a metric of great interest to practitioners given the highly unbalanced
nature of many learning analytics datasets and the high cost of correctly identifying the minority class (in
this case, dropouts), both test statistics show that a conclusion that the two models had true di erences
in performance would be weak at best (depending on the chosen ), while a reported result of the simple
di erences in average might lead readers to conclude otherwise.</p>
        <p>
          The di erences observed here almost certainly re ect the peculiarities of this particular dataset, but this
dataset is peculiar in one way that is quite similar to most other student performance datasets { it is highly
unbalanced, with a dropout:retention ratio of approximately 1:5, as noted previously. While there is little
research on how these test statistics perform on unbalanced datasets, [
          <xref ref-type="bibr" rid="ref2">2</xref>
          ] notes that the replicability of the
sorted runs sampling scheme actually increases as the data becomes more unbalanced, because learners tend
to predict the majority class more as that class becomes more prevalent (they do not conduct this evaluation
for any other sampling scheme).
        </p>
      </sec>
    </sec>
    <sec id="sec-6">
      <title>6. Using Tests in Practice - Cautions and Caveats</title>
      <p>The above application demonstrates an approach to the two-model comparison task in an e ort to begin
a larger discussion about e ective methods for model comparison and evaluation. However, grounding
research in statistically valid methodology is not simply about implementing a new test statistic. Several
other considerations need to be accounted for in the research and modeling process.</p>
      <p>One critical distinction, made at the outset of this paper, is between the two-model comparisons
discussed in this work and multiple-model comparisons (or comparisons across multiple datasets). In practice,
researchers are likely to encounter the multiple comparison scenario, and methods beyond those described
in this paper should be used to avoid invalid inferences. We intend to address such methods in the context
of predictive models of student success in future work.</p>
      <p>
        Another potential pitfall that is not new but has received renewed focus is that of obtaining overly
optimistic results from about model comparison through multiple, repeated rounds of viewing model
comparisons [
        <xref ref-type="bibr" rid="ref20 ref21">20, 21</xref>
        ]. By repeatedly building di erent models or feature sets, collecting model performance
samples, and testing these various iterations, researchers increase the likelihood of observing spurious
correlation and mistaking this for a true di erence in model performance. The broad array of approaches to both
feature extraction and modeling make this a uniquely challenging element of implementing these procedures
in practice, and suggest that a deepened emphasis on pre-registration (specifying hypotheses and methods
prior to experimentation) would help diminish such concerns in future evaluation of student performance
predictors.
      </p>
      <p>Another consideration { but one which works in the opposite direction of the optimistic bias above { is
that approaches which use cross-validation are subject to a pessimistic bias due to smaller datasets being
used for training than the actual size of D a model trained on k 1 folds only contains k k 1 of the data (i.e.,
models trained in a 10-fold cross-validation approach only use 90% of the data). The e ective size of this
bias depends both on the size of the dataset and the sensitivity of the modeling approach(es) to this size.
The multilayer perceptron used above, for example, is likely more sensitive to the training dataset size than
the classi cation tree, and the performance achieved with the perceptron on this relatively small dataset
should be seen as a conservative estimate of its performance on datasets of size jDj. Further research on this
issue is recommended below, and it re ects a larger set of open questions about whether researchers ought
to use di erent corrections for di erent types of models in the comparison task.</p>
      <p>
        A nal consideration, again familiar to statisticians, is the distinction between two-sided and one-sided
signi cance test, and its implications for conclusions about model performance [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ]. The tests discussed
in this paper, and in almost all of the literature described above, are generally two-sided tests, although a
similar methodology can (cautiously) be used in conjunction with one-sided tests, if such a test is appropriate.
Researchers should keep in mind the di erence between these two tests: in particular, the two-sided test
allows us to evaluate HA : 6= 0 while a one-sided test allows us to evaluate HA : &gt; 0 or HA : &lt; 0.
      </p>
      <p>While the ability to draw a speci c conclusion about whether PA PB 0 (that is, whether model
A performs better than model B) is attractive, using a two-sided test is typically only warranted if the
researcher has some a priori reason to believe that a given classi er performs better than another. If this
is not the case, and a researcher views model performance data before making a determination about the
use or direction of a one-sided test, then the results are e ectively biased. Using a one-sided test entirely
disregards the possibility of a performance comparison going in the opposite direction-but in the general
case such an assumption seems unwarranted. Further discussion of this issue in the context of the model
comparison task is also recommended below.</p>
      <p>
        As mentioned above, the two tests implemented here { the corrected repeated k-fold cross-validation test
with 10x10cv and the sorted runs sampling scheme with a Student's t-test { are not the only acceptable
model comparison methods. While we highlight our reasons for favoring these approaches in Table 2 and the
analysis above, there are situations in which implementing these two speci c tests may not be practical-such
as when conducting 100 iterations of model training is not feasible. In these cases, utilizing one of the other
approaches, such as adjusted resampling or adjusted 5x2cv, would still be far preferable to not conducting a
statistical test. The researcher would need to keep in mind the caveats to these methods mentioned in Table
2 and above (for instance, about the lower replicability or power of certain tests). Other approaches not
discussed here may also be useful to researchers, such as the ROC Convex Hull (ROCCH) method proposed
by [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ]. This method is particularly useful when the learners being compared will be used to predict on
new datasets with target conditions (class balances, misclassi cation cost) that are di erent and unknown
in advance; it is not, however, an approach based on hypothesis testing. [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ] also highlights McNemar's
test as a potential alternative, but this test has the same inability to measure (and therefore account for)
sample variance as the di erence-of-proportions test, and showed no substantial performance bene ts in that
analysis. Additionally, while the non-parametric McNemar's test may be a more promising alternative to
t-testing due to its limited assumptions about the data, it may also su er when tested on non-independent
overlapping data sets [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ], though the degree to which this a ects its performance is unclear.
      </p>
    </sec>
    <sec id="sec-7">
      <title>7. Conclusion</title>
      <p>In this paper, we present an overview of the two-model comparison task, with an emphasis on its
application to predictive models of student success. We provide an overview of predictive research in learning
analytics, and of statistical and machine learning research directed at adjusting traditional sampling and
hypothesis testing approaches to the model comparison task. We present a detailed explication of several
of these methods, including invalid methods that researchers should avoid using, and methods that have
demonstrated empirical and theoretical support. Finally, we demonstrate an application of two tests, the
corrected 10x10 cross-validation and the sorted runs sampling scheme with a t-test, to an actual comparison
of two classi ers in the context of MOOC dropout prediction. This application was used to compare and
contrast the two methods, both of which are acceptable for use in future research.</p>
      <sec id="sec-7-1">
        <title>7.1. Future Research</title>
        <p>
          As model comparison and evaluation rises as a priority for applied data science in general, and
learning analytics in particular, further research is needed on several fronts. Overall, additional studies should
evaluate the several di erent approaches for sampling and hypothesis testing discussed in this paper. A
robust foundation of theoretical and applied research, with the latter using both simulated and real data,
is an essential prerequisite to e ective evaluation of student success models. Future research should
revisit the ndings of previous predictive modeling experiments and seek to evaluate their results through
more statistically rigorous frameworks.5 Additionally, this research should include discussions of end-to-end
methodologies for e ective machine learning comparison and evaluation, including developing a consensus
around unbiased approaches to feature extraction, model selection, and model evaluation. Such analyses
will need to extend beyond the two-model case discussed here, and should be applicable to n &gt; 2
classi ers. Given the variety of modeling approaches that continue to emerge, future research also needs to
evaluate whether all models should be treated equally in the case of hypothesis testing { as is the current
standard { or whether adjustments should be made for di erent levels of variation across di erent types of
models ([
          <xref ref-type="bibr" rid="ref23">23</xref>
          ] presented an early analysis of this challenge). Similarly, researchers should begin more active
discussions on how other statistical methods, such as one-sided vs. two-sided hypothesis testing, should be
implemented in experiments and under what conditions. Researchers should continue to discuss methods
for evaluating classi er performance on highly skewed data, moving the conversation beyond mere accuracy
and developing speci c methodologies to evaluate classi er performance on unbalanced datasets. While this
paper highlights a two-class classi cation problem, there is a particular need for simple, e ective measures
of multiclass performance metrics, as many student success prediction problems are multiclass (i.e., grade
prediction). Finally, null hypothesis statistical testing is not the only reasonable approach to this
problem: Bayesian methods are also attracting increased interest in the machine learning community and these
methods have valuable additional insights to add to the model comparison task. Although a comparison ogf
null hypothesis statistical testing and Bayesian testing is beyond the scope of this paper, t is our hope that
future research further explores Bayesian testing methods and utilizes these insights, such as the concept of
practical signi cance, to evaluate predictive models of student success[
          <xref ref-type="bibr" rid="ref14">14</xref>
          ].
5For a promising e ort in this vein, see [
          <xref ref-type="bibr" rid="ref22">22</xref>
          ].
        </p>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>R. R.</given-names>
            <surname>Bouckaert</surname>
          </string-name>
          , E. Frank,
          <article-title>Evaluating the replicability of signi cance tests for comparing learning algorithms</article-title>
          ,
          <source>in: Advances in Knowledge Discovery and Data Mining</source>
          , Springer, Berlin, Heidelberg,
          <year>2004</year>
          , pp.
          <volume>3</volume>
          {
          <fpage>12</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>R. R.</given-names>
            <surname>Bouckaert</surname>
          </string-name>
          ,
          <article-title>Estimating replicability of classi er learning experiments</article-title>
          ,
          <source>in: Proceedings of the Twenty- rst International Conference on Machine Learning, ICML '04</source>
          ,
          <string-name>
            <surname>ACM</surname>
          </string-name>
          , New York, NY, USA,
          <year>2004</year>
          , pp.
          <volume>15</volume>
          {.
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>D. D.</given-names>
            <surname>Jensen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P. R.</given-names>
            <surname>Cohen</surname>
          </string-name>
          ,
          <article-title>Multiple comparisons in induction algorithms</article-title>
          ,
          <source>Mach. Learn</source>
          .
          <volume>38</volume>
          (
          <year>2000</year>
          )
          <volume>309</volume>
          {
          <fpage>338</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>A.</given-names>
            <surname>Feelders</surname>
          </string-name>
          , W. Verkooijen,
          <article-title>On the statistical comparison of inductive learning methods</article-title>
          , in: D. Fisher, H.-J. Lenz (Eds.),
          <source>Learning from Data, Lecture Notes in Statistics</source>
          , Springer New York,
          <year>1996</year>
          , pp.
          <volume>271</volume>
          {
          <fpage>279</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>F.</given-names>
            <surname>Provost</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Fawcett</surname>
          </string-name>
          ,
          <article-title>Robust classi cation for imprecise environments</article-title>
          ,
          <source>Mach. Learn</source>
          .
          <volume>42</volume>
          (
          <issue>3</issue>
          ) (
          <year>2001</year>
          )
          <volume>203</volume>
          {
          <fpage>231</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>T. G.</given-names>
            <surname>Dietterich</surname>
          </string-name>
          ,
          <article-title>Approximate statistical tests for comparing supervised classi cation learning algorithms</article-title>
          ,
          <source>Neural Comput</source>
          .
          <volume>10</volume>
          (
          <issue>7</issue>
          ) (
          <year>1998</year>
          )
          <year>1895</year>
          {
          <year>1923</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>M.</given-names>
            <surname>Kuhn</surname>
          </string-name>
          ,
          <article-title>A short introduction to the caret package</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>R. R.</given-names>
            <surname>Bouckaert</surname>
          </string-name>
          ,
          <article-title>Choosing between two learning algorithms based on calibrated tests.</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>N.</given-names>
            <surname>Japkowicz</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Shah</surname>
          </string-name>
          ,
          <source>Evaluating Learning Algorithms: A Classi cation Perspective</source>
          , Cambridge University Press,
          <year>2011</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>W.</given-names>
            <surname>Xing</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Chen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Stein</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Marcinkowski</surname>
          </string-name>
          ,
          <article-title>Temporal predication of dropouts in MOOCs: Reaching the low hanging fruit through stacking generalization</article-title>
          ,
          <source>Comput. Human Behav</source>
          .
          <volume>58</volume>
          (
          <year>2016</year>
          )
          <volume>119</volume>
          {
          <fpage>129</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>J.</given-names>
            <surname>Demsar</surname>
          </string-name>
          ,
          <article-title>Statistical comparisons of classi ers over multiple data sets</article-title>
          ,
          <source>J. Mach. Learn. Res</source>
          .
          <volume>7</volume>
          (
          <issue>Jan</issue>
          ) (
          <year>2006</year>
          )
          <volume>1</volume>
          {
          <fpage>30</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>S.</given-names>
            <surname>Garcia</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Herrera</surname>
          </string-name>
          ,
          <article-title>An extension on\statistical comparisons of classi ers over multiple data sets"for all pairwise comparisons</article-title>
          ,
          <source>J. Mach. Learn. Res</source>
          .
          <volume>9</volume>
          (
          <issue>Dec</issue>
          ) (
          <year>2008</year>
          )
          <volume>2677</volume>
          {
          <fpage>2694</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <given-names>O. T.</given-names>
            <surname>Yildiz</surname>
          </string-name>
          , E. Alpaydin, Senior Member,
          <article-title>Ordering and nding the best of k&gt;2 supervised learning algorithms</article-title>
          ,
          <source>IEEE Trans. Pattern Anal. Mach. Intell</source>
          .
          <volume>28</volume>
          (
          <issue>3</issue>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <given-names>G.</given-names>
            <surname>Corani</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Benavoli</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Demsar</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Mangili</surname>
          </string-name>
          ,
          <string-name>
            <surname>M.</surname>
          </string-name>
          <article-title>Za alon, Statistical comparison of classi ers through bayesian hierarchical</article-title>
          modellingarXiv:
          <volume>1609</volume>
          .
          <fpage>08905</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <given-names>C.</given-names>
            <surname>Nadeau</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Bengio</surname>
          </string-name>
          ,
          <article-title>Inference for the generalization error</article-title>
          ,
          <source>Mach. Learn</source>
          .
          <volume>52</volume>
          (
          <issue>3</issue>
          ) (
          <year>2003</year>
          )
          <volume>239</volume>
          {
          <fpage>281</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [16]
          <string-name>
            <given-names>I. H.</given-names>
            <surname>Witten</surname>
          </string-name>
          , E. Frank,
          <string-name>
            <given-names>M. A.</given-names>
            <surname>Hall</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C. J.</given-names>
            <surname>Pal</surname>
          </string-name>
          ,
          <source>Data Mining: Practical Machine Learning Tools and Techniques</source>
          , Morgan Kaufmann,
          <year>2016</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [17]
          <string-name>
            <given-names>K.</given-names>
            <surname>Veeramachaneni</surname>
          </string-name>
          , U.
          <string-name>
            <surname>-M. O'Reilly</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          <string-name>
            <surname>Taylor</surname>
          </string-name>
          ,
          <article-title>Towards feature engineering at scale for data from massive open</article-title>
          online coursesarXiv:
          <volume>1407</volume>
          .
          <fpage>5238</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          [18]
          <string-name>
            <given-names>C.</given-names>
            <surname>Bergmeir</surname>
          </string-name>
          ,
          <string-name>
            <surname>J.</surname>
          </string-name>
          <article-title>Ben tez, Neural networks in R using the stuttgart neural network simulator: RSNNS</article-title>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Stat</surname>
          </string-name>
          . Softw.
          <volume>46</volume>
          (
          <issue>1</issue>
          ) (
          <year>2012</year>
          )
          <volume>1</volume>
          {
          <fpage>26</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          [19]
          <string-name>
            <given-names>L.</given-names>
            <surname>Breiman</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Friedman</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C. J.</given-names>
            <surname>Stone</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R. A.</given-names>
            <surname>Olshen</surname>
          </string-name>
          ,
          <article-title>Classi cation and regression trees</article-title>
          , CRC press,
          <year>1984</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          [20]
          <string-name>
            <given-names>T.</given-names>
            <surname>Hastie</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Tibshirani</surname>
          </string-name>
          ,
          <string-name>
            <surname>J. Friedman,</surname>
          </string-name>
          <article-title>The elements of statistical learning</article-title>
          .
          <year>2001</year>
          , NY Springer.
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          [21]
          <string-name>
            <given-names>G.</given-names>
            <surname>Vanwinckelen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Blockeel</surname>
          </string-name>
          ,
          <article-title>On estimating model accuracy with repeated cross-validation</article-title>
          ,
          <source>in: BeneLearn 2012: Proceedings of the 21st Belgian-Dutch Conference on Machine Learning</source>
          ,
          <year>2012</year>
          , pp.
          <volume>39</volume>
          {
          <fpage>44</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          [22]
          <string-name>
            <surname>J. M. L. Andres</surname>
            ,
            <given-names>R. S.</given-names>
          </string-name>
          <string-name>
            <surname>Baker</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          <string-name>
            <surname>Siemens</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          <string-name>
            <surname>GASEVIC</surname>
          </string-name>
          ,
          <string-name>
            <surname>C.</surname>
          </string-name>
          <article-title>A. Spann, Replicating 21 ndings on student success in online learning</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          [23]
          <string-name>
            <given-names>S. L.</given-names>
            <surname>Salzberg</surname>
          </string-name>
          ,
          <article-title>On comparing classi ers: Pitfalls to avoid and a recommended approach, Data Min</article-title>
          .
          <source>Knowl. Discov</source>
          .
          <volume>1</volume>
          (
          <issue>3</issue>
          ) (
          <year>1997</year>
          )
          <volume>317</volume>
          {
          <fpage>328</fpage>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>