<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Using Course Evaluations and Student Data to Predict Computer Science Student Success</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Anlan Du∗</string-name>
          <email>amd5wf@virginia.edu</email>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Alexandra Plukis∗</string-name>
          <email>aplukis@asu.edu</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Huzefa Rangwala</string-name>
          <email>hrangwal@gmu.edu</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Arizona State University</institution>
          ,
          <addr-line>Tempe, AZ</addr-line>
          ,
          <country country="US">United States</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>George Mason University</institution>
          ,
          <addr-line>Fairfax, VA</addr-line>
          ,
          <country country="US">United States</country>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>University of Virginia</institution>
          ,
          <addr-line>Charlottesville, VA</addr-line>
          ,
          <country country="US">United States</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2017</year>
      </pub-date>
      <abstract>
        <p>As the field of computer science has grown, the question of how to improve retention in computer science, especially for females and minorities, has grown increasingly important. Previous research has looked into attitudes among those who leave CS, as well as the impact of taking specific courses; we build on this body of research using large-scale analysis of course evaluations and students' academic history. Our goal is to understand their potential connection to a student's performance and retention within the CS major. We process course-specific data, faculty evaluations, and student demographic data through various machine learning-based classifiers to understand the predictive power of each feature. We find our algorithm performs significantly better for higher-performing students than lower-performing ones, but do not find that evaluations significantly improve predictions of students doing well in courses and staying in the major.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>INTRODUCTION</title>
      <p>Among the most important aspects of a college education are the
classes a student takes. Often, college students use introductory
courses to decide what they would like to study and pursue. Bad
experiences in an introductory course might detract from a student’s
ifrst impression of a field, while a good experience in a course might
improve his or her opinion, even boosting retention and improving
skills upon graduation [13]. Therefore, it is key that administrators
and professors alike understand which course characteristics
maintain interest and improve student outcomes. Such information can
impact administrative decisions, such as who is assigned to teach
particular courses and the recommended sequence of courses.</p>
      <p>The digitization of student records and course evaluations ofers
a unique opportunity to apply big data modeling techniques to study
retention. George Mason University, the data source for this work,
keeps anonymized records on students’ academic records in high
school, demographic data, and their course loads and grades at the
university. They also administer standardized course evaluations
across all courses. Various data mining and modeling techniques,
such as decision trees and support vector machines, can be applied
to these datasets and their results compared. Using this data, one
∗Both authors contributed equally to this research.</p>
      <p>Copyright © 2021 for this paper by its authors. Use permitted under Creative Commons
License Attribution 4.0 International (CC BY 4.0).
can more easily find patterns that reveal how diferent traits afect
student retention.</p>
      <p>George Mason also ofers a unique opportunity to analyze the
impact of professor gender on student success. George Mason’s
engineering faculty is 26.8% female, more than 1.5 times higher
than the national average of 15.7% [10][16]. A larger female faculty
means that analyses of the impact of instructor gender are less likely
to be swayed by a single professor and therefore more statistically
significant.
2</p>
    </sec>
    <sec id="sec-2">
      <title>RELATED WORK</title>
      <p>Our work builds upon previous research regarding student both
college retention and achievement in courses, both generally and
between demographic groups [7] . Demographic disparities are
particularly evident in the number of degrees awarded. For instance,
during George Mason University’s 2017-2018 school year only 15.8%
of the total 196 computer science (CS) degrees were awarded to
females. This lack of representation is even more pronounced for
minority students—only 6 CS degrees were awarded to African
American students and 16 awarded to their Hispanic counterparts
[9]. These disparities have led to a large body of research into
retention for minorities in STEM and specifically [ 8][1][15]. Bettinger
and Long researched the impact of female faculty on female
retention in majors or repeated interest in classes and found mixed
results: some disciplines such as statistics and mathematics
beneifted from an early female professor introduction, while others saw
a decrease in female retention. The authors pointed out that it was
dificult to gauge the exact impact of female professors in fields
that had low levels of females in faculty, such as engineering and
physics. We hope to improve upon on this because George Mason’s
School of Engineering female full time academic faculty make up
26.8%, far surpassing the national average of 15.7% [16] [10].</p>
      <p>The issue of student performance and retention extends beyond
under-represented minorities. Cucuringu et al. used fifteen years
of student data to find classes that optimized a student’s likelihood
of successfully completing a course of study with high grades [5].
They also took the step of segmenting a student population into
sub-groups based on various characteristics, so as to understand the
nuances that diferent types of students might experience. Morsy
and Karypis used a similarly broad, qualitative approach to predict
student performance based on previous classes taken [14].</p>
      <p>Research specific to CS retention has also been conducted:
Biggers et al. incorporated interviews of students who left CS, seeking
to find the qualitative sentiments that afected both female and
male students’ decisions [2]. We combine these two approaches
by using data on students’ individual demographics, grades, and
course history to understand how each factor may contribute to
both student performance and choice of major. Additionally, we
incorporate student evaluations for the courses they take to
understand the role that these qualitative elements may play in these
outcomes, as suggested by Biggers et al [2].</p>
      <p>Research on course evaluations suggests they may prove
informative with regards to a student’s academic experience. Much
research has studied the relationship between the ease of a course,
often represented by the grade a student receives, and the rating of
the faculty. One well-known meta-analysis by Cohen argued that
students are fairly accurate in their assessments of instructional
eficacy [ 6]. Centra’s study built upon this notion, and further
emphasized that students do not give higher evaluations to professors
in a quid pro quo for higher grades: both extremely easy and
dififcult courses sufered in student evaluations, while courses with
appropriate dificulty received the best evaluations [ 3]. Feldman
analyzed the contributory power of various teacher
characteristics to a teacher’s overall rating and student achievement, finding
that preparation, organization, clarity, and students’ feelings of
engagement contributed most strongly to overall performance [11].
He also highlighted some myths about student evaluations, citing
research that suggests that they can, in fact, be informative. We
incorporate evaluations in order to expand on these questions of
student evaluation eficacy, and understand what they say about
students’ experiences and choices.
3</p>
    </sec>
    <sec id="sec-3">
      <title>PROBLEM DESCRIPTION</title>
      <p>
        The objective of this study is to investigate a few questions relating
course quality—defined using faculty traits such as gender and
instructional evaluations—to student retention in computer science.
Specifically, we will address the following inquiries:
(
        <xref ref-type="bibr" rid="ref1">1</xref>
        ) Which course features, if any, in lower division CS courses
improve graduation retention for students?
(
        <xref ref-type="bibr" rid="ref2">2</xref>
        ) Which course features, if any, of instructors in introductory
      </p>
      <p>
        CS courses can predict student success in future CS courses?
(
        <xref ref-type="bibr" rid="ref3">3</xref>
        ) Do non-CS courses that are required by CS majors, like
calculus, have an impact on major retention for students? If
so, which courses and features have the largest impact?
4
4.1
      </p>
    </sec>
    <sec id="sec-4">
      <title>MATERIALS</title>
    </sec>
    <sec id="sec-5">
      <title>Dataset</title>
      <p>Our dataset consisted of records containing first time freshman
student enrollment and course evaluation data for 20,825 George
Mason students over the span of eight years, from Summer 2009
to Fall 2018. All student data were collected and anonymized in
accordance with GMU’s Institutional Review Board policies. The
student data contained demographics data such as age, sex, and
race; admissions data such as high school, SAT score, and high
school GPA; and course data such as declared major, graduation
year, courses taken, and grades received. Students who transferred
into GMU were not included in this dataset because they likely
had completed introductory courses at their previous institutions,
rendering that first-year data inaccessible to us. We also collected
course evaluation data on 87,629 GMU courses from Summer 2009
to Spring 2019, 8,243 of which were computer science, or computer
science-adjacent courses. The evaluations are averages of all of the
student evaluations for that specific course and section, so there is
1 evaluation available for each unique course GMU ofers. This data
was collected from the GMU evaluation site 1, which is publicly
available while on campus. As these are publicly available
documents on campus and the identifying features were anonymized,
they are exempt research under GMU’s IRB policy 2. To collect data
on professor gender, we reviewed pronoun usage in departmental
documents and consulted faculty members when documentation
was insuficient.</p>
      <p>The courses we describe as CS-adjacent are courses taught by or
in conjunction with the Department of Computer Science at GMU.
These CS-adjacent courses include Information Technology,
Computer Game Design, Software Engineering, Electrical and Computer
Engineering, and Information Systems. After discarding course data
with no grades or grades not translating to the A-F scale and
applying our course filters, we had records for 57,627 student-course
enrollments.
4.2</p>
    </sec>
    <sec id="sec-6">
      <title>Definitions</title>
      <p>We frequently discuss “student success” within the computer
science major. In this paper, our definition of “success” is divided into
three categories:</p>
      <p>Completion of a computer science degree: A student is
deifned as graduating with a computer science degree if he or she
graduated with a major in either computer science or applied
computer science. A student is defined as not graduating with a CS
degree if he or she graduated, but not with a CS major. Because
we are focused on retention, not graduation, we only included in
our data students who had had enough time to graduate. By not
including students who transfer or drop out of GMU, or who simply
have not graduated yet, we reduced the number of confounding
variables that are not directly related to students’ experiences in
CS.</p>
      <p>Fulfilment of a student’s potential in a course: A student’s
“potential” in CS211 is defined as the term GPA of the semester
in which CS112—the direct pre-requisite—was taken. Our interest
in this stems from its potential in combination with predictions
of passing a course. Students who perform below their “potential”
within CS211, despite passing and receiving credit for the course,
might still benefit from administrator involvement. Alternatively,
the characteristics of students performing above their potential may
highlight positive factors that should continue to be proliferated
on an institutional level.</p>
      <p>Passing a course for credit: A student is defined as passing
a course for credit if he or she receives a C grade or above. At
GMU, computer science BS students “must earn a C or better in
any course intended to satisfy a prerequisite for a computer science
course ... [s]tudents may attempt an undergraduate course taught
by the Volgenau School of Engineering twice.” 3. In our research,
we specifically target student success in CS112 and CS211 because
they are required courses for CS/ACS majors and pre-requisites for
all other programming courses. Figure 1 visualizes the contrast in
pass rates for first and second attempts in CS211: within our dataset,</p>
      <sec id="sec-6-1">
        <title>1https://irr2.gmu.edu/</title>
        <p>
          2https://rdia.gmu.edu/topics-of-interest/human-or-animal-subjects/humansubjects/exempt-research/
3https://catalog.gmu.edu/colleges-schools/engineering/computer-science/computerscience-bs/#admissionspoliciestext
only 19.8% of students attempting CS211 for the first time did not
receive credit, versus 63.3% of students on their second attempt.
For this work, we compared performance of predictive models that
were trained on three diferent sets of data, which are fully described
in Appendix 9.2:
(
          <xref ref-type="bibr" rid="ref1">1</xref>
          ) Baseline predictions based on high school performance and
student demographic data, as well as basic course
information such as the term in which a course was taken and a
student’s GPA in that term.
(
          <xref ref-type="bibr" rid="ref2">2</xref>
          ) Baseline features in addition to instructor gender and course
evaluations for the classes, either CS-only or math and CS,
taken by each student.
(
          <xref ref-type="bibr" rid="ref3">3</xref>
          ) Baseline features, plus the course numbers as unique
identifiers that were distinct for each section and semester of
a class, but common to all of the students who took that
section.
        </p>
        <p>We chose to use machine learning classifiers because they can often
pick up on more intricate patterns and correlations than linear and
other basic statistical models can would. We decided to test these
distinct data sets because they each highlight a component of
student courses that may be significant to students’ performance and
ultimate retention. The full list of features used in each experiment
are described in Appendix 9.</p>
        <p>We used seven classifiers from the Python sci-kit learn library:
Random Forest, Gradient Boosting, AdaBoost, SVC, Decision Tree,
Neural Net, and Naive Bayes. For each of these models, we
performed 5-fold cross-validation, recording the resulting the averages
and standard deviations. In order to account for imbalances in our
dataset, we decided upon area under an ROC curve (ROC AUC)
and F1 score as our main metrics, because they take into account
precision and recall in addition to overall accuracy.
5.1</p>
      </sec>
    </sec>
    <sec id="sec-7">
      <title>Pre-Processing</title>
      <p>We consolidated student data for all students who took at least one
CS class, of whom there were 15,552. To better incorporate summer
student data, we moved summer courses to the proceeding fall term.
Then, we calculated percentile values for students’ SAT scores and
high school GPAs, enabling us to compare these metrics along
a standard scale of 0 to 1. Next, for models predicting retention,
we removed all students who had not yet graduated, leaving us
with 7,602 students. Lastly, we dropped all students with empty
values for any of the columns used in training. This left a dataset
of 1,476 students who took at least one CS or CS-adjacent class
before graduating. Of those, 330 graduated with a CS or ACS major,
or 22.35%. This left us with an imbalanced dataset, leading to our
decision to use F1 score and ROC AUC to characterize our models.</p>
      <p>For the grade prediction portion, students who received no
grade—meaning they audited or did not complete the class—were
not included in the data. This left 1,728 students who took both
CS112 and CS211 at GMU at least once. In the cases where students
took these courses multiple times, only the initial course attempt
was used so as to only capture their original experience in the class.
Predicting grades for only first attempts of CS211 ofers an earlier
lfagging system for at-risk students.</p>
      <p>We wanted to understand the impact that not only general
instructor qualities, but also “exemplary” instructors, had on student
grades. To that end, each grade prediction model was run with
the course evaluations processed in one of two ways: percentiles
or flags. Percentiles, which capture the general quality of an
instructor, had each evaluation entry into a percentile relative to the
other courses. Flags, which served to identify exemplary
instructors, transformed each entry into a binary feature based on whether
it was in the top 10% of evaluation scores in that category.</p>
      <p>Although evaluations ofer more data than can usually be gleaned
from student records, we tried to capture the elements in a course
that cannot be captured in evaluations or records. We did so by
creating unique course IDs for each course, so as to highlight especially
good courses, good times of day for students, and good
connections between students in courses—all of which are not explicitly
quantified in our data.
5.2</p>
    </sec>
    <sec id="sec-8">
      <title>Experiments</title>
      <p>As mentioned previously, we had three main groups of datasets.
The second group, which includes the course evaluation data, was
then run on three diferent subsets: first, it was trained with just
the "overall teaching" and "overall course" evaluation scores for
the first CS and math courses, then the overall evaluations from
the first two courses, then all available course evaluation metrics
for the first two courses in each area. For graduation prediction,
both math and CS courses were included in the evaluation data in
order to capture a full snapshot of introductory courses. For grade
prediction, only CS courses are included so as to not diminish the
dataset of non-CS or non-STEM students, who often do not have
the same rigorous math requirements.</p>
      <p>Our rationale in deploying some tests with just two course
evaluation features per course was that the added dimensionality of
running the models on all of the features (many of which were
positively correlated) might hinder performance. The baseline was
meant to be the control for the predictive capabilities of only
basic course features and student demographic information, so that
subsequent tests might reveal how much predictive power the
additional data might have added. The full list of features used for each
of these experiments are listed in Appendix 9.1.</p>
      <p>All of our experiments deal with binary classification, and as such
require binary flagging for the classes of interest. In grade prediction
experiments, those who are at risk—of either not receiving credit
or not fulfilling their potential in a course—are flagged with a 1. In
the graduation predictions, students who graduate with a computer
science or applied computer science degree are flagged with a 1.</p>
      <p>For each experiment, we ran 5-fold cross validation on our
models, using a deterministic seed to generate our training-testing splits
so that we could directly compare splits before and after the
models were trained. We performed Student’s t-tests on our results to
understand the significance of any diferences in performance.
5.3</p>
    </sec>
    <sec id="sec-9">
      <title>Fairness</title>
      <p>In order to check that the predictions were not favoring certain
students already predisposed to graduating with a CS degree or
passing their courses, we decided to separate the students into
groups based on their academic abilities coming into college. We
consider a prediction algorithm to be fair if its F1 score remains
statistically similar regardless of the student’s quartile standing.
We used high school GPA (HS GPA) and total SAT scores to have
one metric of school success and one metric of testing success to
create a fuller understanding of student academic ability. These
two scores were transformed into percentiles, averaged together,
then transformed into a percentile once more. This final percentile
calculation divided the students into evenly sized groups.</p>
      <p>The students were then separated into 4 groups based on their
percentile standings, as pictured in Figures 3 and 4. To test the
fairness implications, 5-fold splits were trained on all students and
then tested only on certain quartiles. This way, we could clearly
see any disparity in performance for all students versus those in
separate groups of students.</p>
      <p>We used these quartiles to test for fairness by training each of
our models on the full datasets, splitting up the testing sets based
on the quartiles, and calculating the metrics based on these results.
We then compared these quartile results with the results for all
students to determine if there was a significant diference between
them, and therefore a disparity in fairness for difering groups.
6</p>
    </sec>
    <sec id="sec-10">
      <title>RESULTS</title>
      <sec id="sec-10-1">
        <title>Our results are divided into three sections:</title>
        <p>
          (
          <xref ref-type="bibr" rid="ref1">1</xref>
          ) Performance metrics (F1 Score, ROC AUC, Accuracy) for our
baseline models;
(
          <xref ref-type="bibr" rid="ref2">2</xref>
          ) Comparison between baseline models and models that
include course evaluation and other instructor data;
(
          <xref ref-type="bibr" rid="ref3">3</xref>
          ) Fairness: Comparison between prediction of each academic
quartile versus prediction of all students
6.1
        </p>
      </sec>
    </sec>
    <sec id="sec-11">
      <title>Baseline Performance</title>
      <p>The Percentiles column indicates evaluation scores were
converted to percentiles; the Flags column indicates binary flags of
the top 10% of scores were used.4
4Note that for models using Discrete IDs, we do not use numerical evaluation data, so
there is no distinction between the two categories’ results.</p>
      <sec id="sec-11-1">
        <title>1 Overall Eval</title>
        <p>ll 2 Overall Evals</p>
      </sec>
      <sec id="sec-11-2">
        <title>A 2 Full Evals</title>
      </sec>
      <sec id="sec-11-3">
        <title>Discrete IDs</title>
      </sec>
      <sec id="sec-11-4">
        <title>1 Overall Eval</title>
      </sec>
      <sec id="sec-11-5">
        <title>1 2 Overall Evals</title>
      </sec>
      <sec id="sec-11-6">
        <title>Q 2 Full Evals</title>
      </sec>
      <sec id="sec-11-7">
        <title>Discrete IDs</title>
      </sec>
      <sec id="sec-11-8">
        <title>1 Overall Eval</title>
      </sec>
      <sec id="sec-11-9">
        <title>2 2 Overall Evals</title>
      </sec>
      <sec id="sec-11-10">
        <title>Q 2 Full Evals</title>
      </sec>
      <sec id="sec-11-11">
        <title>Discrete IDs</title>
      </sec>
      <sec id="sec-11-12">
        <title>1 Overall Eval</title>
      </sec>
      <sec id="sec-11-13">
        <title>3 2 Overall Evals</title>
      </sec>
      <sec id="sec-11-14">
        <title>Q 2 Full Evals</title>
      </sec>
      <sec id="sec-11-15">
        <title>Discrete IDs</title>
      </sec>
      <sec id="sec-11-16">
        <title>1 Overall Eval</title>
      </sec>
      <sec id="sec-11-17">
        <title>4 2 Overall Evals</title>
      </sec>
      <sec id="sec-11-18">
        <title>Q 2 Full Evals</title>
      </sec>
      <sec id="sec-11-19">
        <title>Discrete IDs</title>
      </sec>
      <sec id="sec-11-20">
        <title>1 Overall Eval</title>
        <p>ll 2 Overall Evals</p>
      </sec>
      <sec id="sec-11-21">
        <title>A 2 Full Evals</title>
      </sec>
      <sec id="sec-11-22">
        <title>Discrete IDs</title>
      </sec>
      <sec id="sec-11-23">
        <title>1 Overall Eval</title>
      </sec>
      <sec id="sec-11-24">
        <title>1 2 Overall Evals</title>
      </sec>
      <sec id="sec-11-25">
        <title>Q 2 Full Evals</title>
      </sec>
      <sec id="sec-11-26">
        <title>Discrete IDs</title>
      </sec>
      <sec id="sec-11-27">
        <title>1 Overall Eval</title>
      </sec>
      <sec id="sec-11-28">
        <title>2 2 Overall Evals</title>
      </sec>
      <sec id="sec-11-29">
        <title>Q 2 Full Evals</title>
      </sec>
      <sec id="sec-11-30">
        <title>Discrete IDs</title>
      </sec>
      <sec id="sec-11-31">
        <title>1 Overall Eval</title>
      </sec>
      <sec id="sec-11-32">
        <title>3 2 Overall Evals</title>
      </sec>
      <sec id="sec-11-33">
        <title>Q 2 Full Evals</title>
      </sec>
      <sec id="sec-11-34">
        <title>Discrete IDs</title>
      </sec>
      <sec id="sec-11-35">
        <title>1 Overall Eval</title>
      </sec>
      <sec id="sec-11-36">
        <title>4 2 Overall Evals</title>
      </sec>
      <sec id="sec-11-37">
        <title>Q 2 Full Evals</title>
      </sec>
      <sec id="sec-11-38">
        <title>Discrete IDs</title>
      </sec>
      <sec id="sec-11-39">
        <title>Experiment</title>
        <p>In all of these t-tests, our null hypothesis was that evaluations
and specific courses taken by a student do not improve student
success predictions. If this were true, results from the baseline set of
data would be the same as results that included course information
because the course information would add no predictive power.
None of our experiments proved to have a significant improvement
over our baseline, so we fail to reject our null hypothesis and do
not find that evaluations improve predictions of student success.
6.3</p>
      </sec>
    </sec>
    <sec id="sec-12">
      <title>Fairness Across Student Quartiles</title>
      <p>Tables 6, 7, and 8 show fairness t-tests. These are tests of whether
the performance of each experimental model is better or worse at
predicting results for a specific quartile, versus predicting results
for all students. They capture the statistical significance of
discrepancies in performance when run on diferent groups of students.</p>
      <p>The null hypothesis in these tests is that there is no diference
between the F1 scores for all students and those of each quartile.
In other words: the null hypothesis is that the predictions are fair.
The lowest p-scores we found are in bold or, if they are statistically
significant, are highlighted.</p>
      <p>Table 6 shows the models’ fairness in predicting whether
students passed CS211.</p>
      <p>Table 7 shows fairness in predicting whether students achieved
their potential grades. This table difers much from Table 6 in that
many of the p-values listed here are significant at the 0.05 level. All
of the significant results are clustered within the first and second
quartiles, which are the bottom two quartiles in our groupings.</p>
      <p>Table 8 shows models’ fairness in predicting whether students
graduate with a CS major. While the significant t statistics in Table
7 were positive—indicating that the models perform best on the
ifrst and second quartiles—we see that performance for the lower
two quartiles is negative. Additionally, t statistics are significantly
better for students in the top quartile. This suggests significant
fairness disparities in these prediction models: the F1 scores of
prediction across the entire student body overlap strongly with the
F1 scores across the third quartile, but vary widely from those of
both stronger and poorer overall performers. This is in spite of the
quartiles being represented in the overall dataset in equal numbers
of data points.</p>
      <sec id="sec-12-1">
        <title>Percentiles</title>
        <p>t Statistic p-value</p>
      </sec>
      <sec id="sec-12-2">
        <title>Flags</title>
        <p>t Statistic
type
Overall, our results show that adding evaluations to predictions
does not significantly improve predictions over the baseline of
student demographics and basic course features. The addition of
instructor gender, too, was not significant. Even when gender was
included for all courses used in predictions, the results did not
improve drastically. However, there were many semesters for which
there were no female instructors available at all to teach a course, so
there could not be any direct comparisons between students with
male instructors and those with female instructors. Although there
are slight improvements for some experiments beyond the baseline,
notably those involving the discrete and continuous unique IDs,
they do not reach a significance level of 0.05. This suggests that the
impact of student evaluation data and instructor gender on student
performance is not immediately visible.</p>
        <p>Generally, our experiments performed better when predicting
whether students would achieve their "potential" grades than whether
they would receive credit. When comparing p-values between
passing and potential for all experiments, as in Tables 4, 5, and 6,
predicting student potential seems to improve upon predicting passing
even when compared to their respective baselines. We attribute
this to the fluid nature of a student’s forecast grade: if a students
forecast grade is a C, then there are 3 possible grades that this
student could get and still fulfill at or above his or her potential.
Similarly, for students whose predicted grades are A’s, there is only
one possible grade, an A, with which they can achieve at or above
their potential. This imbalance on both sides of the forecast grades
means the models can make an easier prediction of achieving below
a potential grade because there are generally more options on the
lower end of the grade scale than on the higher end.</p>
        <p>In addition, because the evaluations we have access to are only
averages for all students in a course and do not reflect each
student’s personal evaluation of the professor, each student in a section
of a course would have the same evaluations. This large amount
of overlap between students, who then experienced diferent
outcomes with their success, seemed to negatively impact the models
in experiments where full sets of evaluations were used. This
issue was slightly assuaged with the use of unique IDs, but not at a
significant level for most quartiles, see Tables 5, 6, and 7. For this
reason, evaluation sets where evaluations are unique to each student
would provide and interesting contrast to this work—individualized
evaluations might provide high quality features for prediction.</p>
      </sec>
    </sec>
    <sec id="sec-13">
      <title>7.2 Fairness</title>
      <p>Our fairness results reach significance levels especially often in
the third and fourth quartiles—see Tables 8, 9, and 10—which are
the lower academic quartiles. These quartile models underperform
against the models for all students, frequently significantly. This
is a cause for concern–the students for which our models predict
well on, quartiles 1 and 2, are the quartiles in which students often
already perform well. There are a few reasons that might contribute
to this underperformance on the lower quartile of students. One is
that our quartiles are artificially created—although high school GPA
and SAT scores are indicators of academic success in high school,
they do not necessarily represent the same success in college. In
addition, we split students into quartile depending on the percentile
of their averaged HS GPA and SAT scores, not on any visible clusters
within the data. These artificial clusters might not represent true
student groups.</p>
    </sec>
    <sec id="sec-14">
      <title>8 CONCLUSION</title>
      <p>Our data suggests there is a pressing need to understand how
students of diferent academic calibers experience the same curricula,
given the disparities in their ultimate outcomes. It also suggests that
evaluations of courses, at least as they are structured in our data,
do not ofer significant insights as to how a student will perform
or whether he or she will remain in computer science. Lastly, we
ifnd that the starting course number or code may have some
predictive power, suggesting that diferent courses may significantly
impact the outcomes of students. The question now becomes one of
identifying how we measure the diferent features of these courses.</p>
      <p>There are several possible expansions on our methodology. As
previously mentioned, our data is imbalanced, and using techniques
such as cost-aware training, oversampling, undersampling, or
Synthetic Minority Oversampling Technique [4] might enable a more
balanced weighting of results and greater accuracy in identifying
points in the minority classes. Another proposed fairness-oriented
metric is the Absolute Between-ROC Area metric, which measures
the absolute area between two ROC curves. In doing so, it measures
disparities in prediction across every possible decision threshold,
as opposed to just one[12]. Lastly, we would like to grid-search for
hyperparameters that optimize F1 score and ROC AUC, rather than
accuracy.</p>
      <p>In addition to course evaluations for computer science classes,
we also scraped course evaluations for other classes. In the future,
we hope to use this dataset to apply such retention analysis to all
majors. Given that GMU has unique student body, with many
transfer students and non-traditional graduates, we would like to also
include these students in a future analysis to track diferences in
their progressions through their majors. This also begs the question
of whether our results would be diferent at a school with more
four-year students. Despite the fact that our data does not indicate
that evaluations can improve predictions of student success, we are
interested in the outcomes of research into this avenue at schools
with difering evaluation styles to see if these results can be
improved upon. In addition, the fairness concerns raised in this paper
around difering performances for students with varying academic
statuses are of concern. We would like to see the improvement
of grade prediction techniques both for all students and for each
quartile or minority demographic.</p>
    </sec>
    <sec id="sec-15">
      <title>ACKNOWLEDGMENTS</title>
      <p>This work was funded by the National Science Foundation as a
part of GMU’s Computer Science Research Experience for
Undergraduates, and graciously supported by Dr. Karen Lee and GMU’s
OSCAR ofice. We would also like to thank Drs. Mark Snyder and
Huzefa Rangwala for their guidance and knowledge. Finally, we
thank GMU for the use of their facilities and this opportunity.</p>
      <sec id="sec-15-1">
        <title>Features</title>
        <sec id="sec-15-1-1">
          <title>Race</title>
          <p>Sex</p>
        </sec>
        <sec id="sec-15-1-2">
          <title>High School GPA</title>
        </sec>
        <sec id="sec-15-1-3">
          <title>SAT Total Score</title>
        </sec>
        <sec id="sec-15-1-4">
          <title>SAT Verbal Score</title>
        </sec>
        <sec id="sec-15-1-5">
          <title>SAT Math Score</title>
        </sec>
        <sec id="sec-15-1-6">
          <title>Average Percentile</title>
        </sec>
        <sec id="sec-15-1-7">
          <title>Class Term Taken</title>
        </sec>
        <sec id="sec-15-1-8">
          <title>Term GPA</title>
        </sec>
        <sec id="sec-15-1-9">
          <title>Instructor Gender</title>
        </sec>
        <sec id="sec-15-1-10">
          <title>Grade Points</title>
        </sec>
        <sec id="sec-15-1-11">
          <title>Overall Evaluations</title>
        </sec>
        <sec id="sec-15-1-12">
          <title>All Evaluations</title>
        </sec>
        <sec id="sec-15-1-13">
          <title>Course ID</title>
        </sec>
      </sec>
      <sec id="sec-15-2">
        <title>Meaning</title>
        <p>Categorical variable, includes option for no race
listed.</p>
        <p>Categorical variable: male, female, and no gender
listed.</p>
        <p>Continuous variable.</p>
        <p>Continuous variable, out of 1600. Empty cells are
iflled in with the median of the total SAT scores.
Continuous variable, out of 800. Empty cells are
iflled in with the median of the total SAT verbal
scores.</p>
        <p>Continuous variable, out of 800. Empty cells are
iflled in with the median of the total SAT math
scores
Continuous variable, between 0 and 1. Average of
the HS GPA and SAT Total percentiles for each
student.</p>
        <p>Continuous variable, indicates the term in which
the student took the course used for prediction and
the course being predicted.</p>
        <p>Continuous variable, the non-cumulative GPA for
the term in which the student took the course used
for prediction and the course being predicted.
Binary variable, split between male and female.
Continuous variable, the grade received in the
course used for predicting the second course.
Continuous or binary, depending on the treatment
of the specific test—flagging or percentiles. These
are defined as SET (as seen in Appendix 9.2)
questions 15 and 16.</p>
        <p>Continuous or binary, depending on the treatment
of the specific test—flagging or percentiles. These
are defined as SET (as seen in Appendix 9.2)
questions 1 through 14.</p>
        <p>Binary, represents the unique course taken by a
student: ID is discipline, course number, section
number, term taken, and binary digit indicating
a summer term. Students in the same course and
section will all have a 1.
✕
✕
✕
✕
✕
✕
✕
✕
✕
✕
✕
✕
✕
✕
✕
✕
✕
✕
✕
✕
✕
✕
✕
✕
✕
✕
✕
✕
✕
✕
✕
✕
✕
✕
✕
✕
✕
✕
✕
✕
✕
✕
✕
✕
✕
✕
✕</p>
      </sec>
    </sec>
    <sec id="sec-16">
      <title>A PREDICTION FEATURES</title>
      <p>Table 9 shows the features used for each of the predictions, categorized by the type of experiment being described.</p>
      <sec id="sec-16-1">
        <title>Baseline</title>
      </sec>
      <sec id="sec-16-2">
        <title>Overall SET All SET IDs</title>
        <p>Each of these sections were rated on a scale of 1 to 5, with a NA option available. Questions 15 and 16 are the “overall” evaluations used in
certain experiments.</p>
        <p>
          (
          <xref ref-type="bibr" rid="ref1">1</xref>
          ) Course requirements and expectations were clear.
(
          <xref ref-type="bibr" rid="ref2">2</xref>
          ) The course was well organized.
(
          <xref ref-type="bibr" rid="ref3">3</xref>
          ) The instructor helped me to better understand the course material.
(
          <xref ref-type="bibr" rid="ref4">4</xref>
          ) Feedback (written comments and suggestions on papers, solutions provided, class discussion, etc.) was helpful.
(
          <xref ref-type="bibr" rid="ref5">5</xref>
          ) The instructor showed respect for the students.
(
          <xref ref-type="bibr" rid="ref6">6</xref>
          ) The instructor was accessible either in person or electronically.
(
          <xref ref-type="bibr" rid="ref7">7</xref>
          ) The course grading policy was clear.
(
          <xref ref-type="bibr" rid="ref8">8</xref>
          ) Graded work reflected what was covered in the course.
(
          <xref ref-type="bibr" rid="ref9">9</xref>
          ) The assignments (projects, papers, presentations, etc.) helped me learn the material.
        </p>
        <p>
          (
          <xref ref-type="bibr" rid="ref10">10</xref>
          ) The textbook and/or assigned readings helped me understand the material.
(
          <xref ref-type="bibr" rid="ref11">11</xref>
          ) Assignments and exams were returned in a reasonable amount of time.
(
          <xref ref-type="bibr" rid="ref12">12</xref>
          ) The instructor covered the important aspects of the course as outlined in the syllabus.
(
          <xref ref-type="bibr" rid="ref13">13</xref>
          ) The instructor made the class intellectually stimulating.
(
          <xref ref-type="bibr" rid="ref14">14</xref>
          ) The instructor encouraged the students to be actively involved in the material through discussion, assignments, and other activities.
(
          <xref ref-type="bibr" rid="ref15">15</xref>
          ) My overall rating of the teaching.
        </p>
        <p>
          (
          <xref ref-type="bibr" rid="ref16">16</xref>
          ) My overall rating of this course.
        </p>
        <p>C</p>
      </sec>
    </sec>
    <sec id="sec-17">
      <title>EXTENDED GRADE PREDICTION RESULTS</title>
      <p>type
model
F1</p>
      <sec id="sec-17-1">
        <title>Percentiles AUC Acc F1</title>
      </sec>
      <sec id="sec-17-2">
        <title>Top 10% Flags AUC Acc</title>
        <p>1 overall Gradient 0.665 0.872 0.809 0.689 0.875 0.826
issagPn 2 2ovfeurlalll BBGGoorraaooddssttiieeiinnnnggtt ±±0000.. 66..0078340696 ±±0000.. 88..0078336024 ±±0000.. 88..0012222229 ±±0000.. 66..0078552433 ±±0000.. 88..0077335432 ±±0000.. 88..0012335102</p>
        <p>Boosting ±0.056 ±0.034 ±0.027 ±0.054 ±0.032 ±0.029
1 overall Random 0.799 0.770 0.728 0.803 0.773 0.734
iltteaoPn 2 2ovfeurlalll RRFFaaoonnrrddeeoossmmtt ±±0000.. 88..0000114800 ±±0000.. 77..0088118788 ±±0000.. 77..0034119133 ±±0000.. 77..0099106606 ±±0000.. 77..0077111377 ±±0000.. 77..0022102548</p>
        <p>Forest ±0.016 ±0.030 ±0.019 ±0.008 ±0.013 ±0.010
Table 10: Highest performing models from each of the experiments, evaluation treatments, and grade prediction styles. The
best performers in each grade prediction style block are highlighted.</p>
        <p>The experiment that improved upon the baseline power of prediction most utilizes unique course IDs to represent individual courses taken.
The results of this type of experiment are displayed in Table 2. Table 2 contains the results of the ID experiments predicting student grades.
The top performing models in Table 2 outperform the baseline predictive powers in F1, AUC, and accuracy measures, and the significance of
these experiments is explored in Table 4.</p>
        <p>F1</p>
      </sec>
      <sec id="sec-17-3">
        <title>Potential AUC Acc</title>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>E. P.</given-names>
            <surname>Bettinger</surname>
          </string-name>
          and
          <string-name>
            <given-names>B. T.</given-names>
            <surname>Long</surname>
          </string-name>
          .
          <year>2005</year>
          .
          <article-title>Do Faculty Serve as Role Models? The Impact of Instructor Gender on Female Students</article-title>
          .
          <source>The American Economic Review</source>
          <volume>95</volume>
          ,
          <issue>2</issue>
          (
          <year>2005</year>
          ),
          <fpage>152</fpage>
          -
          <lpage>157</lpage>
          . http://www.jstor.org/stable/4132808
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>Maureen</given-names>
            <surname>Biggers</surname>
          </string-name>
          , Anne Brauer, and
          <string-name>
            <given-names>Tuba</given-names>
            <surname>Yilmaz</surname>
          </string-name>
          .
          <year>2008</year>
          .
          <article-title>Student Perceptions of Computer Science: A Retention Study Comparing Graduating Seniors with Cs Leavers</article-title>
          .
          <source>SIGCSE Bull</source>
          .
          <volume>40</volume>
          ,
          <issue>1</issue>
          (mar
          <year>2008</year>
          ),
          <fpage>402</fpage>
          -
          <lpage>406</lpage>
          . https://doi.org/10.1145/1352322. 1352274
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <surname>John</surname>
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Centra</surname>
          </string-name>
          .
          <year>2003</year>
          .
          <article-title>Will Teachers Receive Higher Student Evaluations by Giving Higher Grades</article-title>
          and
          <source>Less Course Work? Research in Higher Education</source>
          <volume>44</volume>
          ,
          <issue>5</issue>
          (
          <year>2003</year>
          ),
          <fpage>495</fpage>
          -
          <lpage>518</lpage>
          . https://doi.org/10.1023/A:1025492407752
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <surname>Nitesh</surname>
            <given-names>V.</given-names>
          </string-name>
          <string-name>
            <surname>Chawla</surname>
            , Kevin W. Bowyer, Lawrence O. Hall, and
            <given-names>W. Philip</given-names>
          </string-name>
          <string-name>
            <surname>Kegelmeyer</surname>
          </string-name>
          .
          <year>2002</year>
          .
          <article-title>SMOTE: synthetic minority over-sampling technique</article-title>
          .
          <source>Journal of artificial intelligence research 16</source>
          (
          <year>2002</year>
          ),
          <fpage>321</fpage>
          -
          <lpage>357</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>Chantal</given-names>
            <surname>Cherifi</surname>
          </string-name>
          , Hocine Cherifi, MÃąrton Karsai, and Mirco Musolesi (Eds.).
          <year>2018</year>
          .
          <article-title>Complex Networks &amp; Their Applications VI: Proceedings of Complex Networks 2017 (The Sixth International Conference on Complex Networks and Their Applications)</article-title>
          .
          <source>Studies in Computational Intelligence</source>
          , Vol.
          <volume>689</volume>
          . Springer International Publishing, Cham. https://doi.org/10.1007/978-3-
          <fpage>319</fpage>
          -72150-7
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <surname>Peter</surname>
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Cohen</surname>
          </string-name>
          .
          <year>1981</year>
          .
          <article-title>Student Ratings of Instruction and Student Achievement: A Meta-Analysis of Multisection Validity Studies</article-title>
          .
          <source>Review of Educational Research</source>
          <volume>51</volume>
          ,
          <issue>3</issue>
          (
          <year>1981</year>
          ),
          <fpage>281</fpage>
          -
          <lpage>309</lpage>
          . https://doi.org/10.2307/1170209
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>Paulo</given-names>
            <surname>Cortez</surname>
          </string-name>
          and
          <string-name>
            <given-names>Alice</given-names>
            <surname>Silva</surname>
          </string-name>
          .
          <year>2008</year>
          .
          <article-title>Using data mining to predict secondary school student performance</article-title>
          .
          <source>EUROSIS</source>
          (01
          <year>2008</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <surname>Benjamin</surname>
            <given-names>J.</given-names>
          </string-name>
          <string-name>
            <surname>Drury</surname>
            , John Oliver Siy, and
            <given-names>Sapna</given-names>
          </string-name>
          <string-name>
            <surname>Cheryan</surname>
          </string-name>
          .
          <year>2011</year>
          .
          <article-title>When Do Female Role Models Benefit Women? The Importance of Diferentiating Recruitment From Retention in STEM</article-title>
          .
          <source>Psychological Inquiry</source>
          <volume>22</volume>
          ,
          <issue>4</issue>
          (
          <year>2011</year>
          ),
          <fpage>265</fpage>
          -
          <lpage>269</lpage>
          . https: //doi.org/10.1080/1047840X.
          <year>2011</year>
          .620935
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          <article-title>[9] Ofice of Institutional Efectiveness and</article-title>
          <string-name>
            <given-names>Planning. 2019. Degrees</given-names>
            <surname>Conferred By Degree And Demographic - Year</surname>
          </string-name>
          2017-
          <volume>18</volume>
          ,
          <string-name>
            <given-names>All</given-names>
            <surname>Terms</surname>
          </string-name>
          . http://irr2.gmu.edu/New/N_ Degree/DegDegreeDetail.cfm
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          <source>[10] Ofice of Institutional Efectiveness and Planning</source>
          .
          <year>2019</year>
          .
          <article-title>Full-Time Academic Faculty Demographic Profiles Two-Year Comparisons</article-title>
          . http://irr2.gmu.edu/New/N_ Faculty/FullTimeFacComp.cfm
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <surname>Kenneth</surname>
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Feldman</surname>
          </string-name>
          .
          <year>1996</year>
          .
          <article-title>Identifying exemplary teaching: Using data from course and teacher evaluations</article-title>
          .
          <source>New Directions for Teaching and Learning</source>
          <year>1996</year>
          ,
          <volume>65</volume>
          (Mar
          <year>1996</year>
          ),
          <fpage>41</fpage>
          -
          <lpage>50</lpage>
          . https://doi.org/10.1002/tl.37219966509
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <surname>Josh</surname>
            <given-names>Gardner</given-names>
          </string-name>
          , Christopher Brooks, and
          <string-name>
            <given-names>Ryan</given-names>
            <surname>Baker</surname>
          </string-name>
          .
          <year>2019</year>
          .
          <article-title>Evaluating the Fairness of Predictive Student Models Through Slicing Analysis</article-title>
          .
          <source>In Proceedings of the 9th International Conference on Learning Analytics &amp; Knowledge (LAK19)</source>
          . ACM, New York, NY, USA,
          <fpage>225</fpage>
          -
          <lpage>234</lpage>
          . https://doi.org/10.1145/3303772.3303791
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <given-names>Jef</given-names>
            <surname>Kastner Gregory Warren Bucks</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Kathleen A.</given-names>
            <surname>Ossman</surname>
          </string-name>
          and
          <string-name>
            <given-names>F James</given-names>
            <surname>Boerio</surname>
          </string-name>
          .
          <year>2015</year>
          .
          <article-title>First-year Engineering Courses' Efect on Retention and Workplace Performance</article-title>
          .
          <source>In 2015 ASEE Annual Conference &amp; Exposition. ASEE Conferences</source>
          , Seattle, Washington. https://peer.asee.org/24114.
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <given-names>Sara</given-names>
            <surname>Morsy</surname>
          </string-name>
          and
          <string-name>
            <given-names>George</given-names>
            <surname>Karypis</surname>
          </string-name>
          .
          <year>2017</year>
          .
          <article-title>Cumulative knowledge-based regression models for next-term grade prediction</article-title>
          .
          <source>Proceedings of the 17th SIAM International Conference on Data Mining</source>
          ,
          <string-name>
            <surname>SDM</surname>
          </string-name>
          <year>2017</year>
          (jan
          <year>2017</year>
          ),
          <fpage>552</fpage>
          -
          <lpage>560</lpage>
          . http://www.scopus. com/inward/record.url?scp=
          <volume>85027876583</volume>
          &amp;partnerID=8YFLogxK
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <given-names>Laird</given-names>
            <surname>Townsend</surname>
          </string-name>
          .
          <year>1994</year>
          .
          <article-title>How Universities Successfully Retain and Graduate Black Students</article-title>
          .
          <source>The Journal of Blacks in Higher Education</source>
          <volume>4</volume>
          (
          <year>1994</year>
          ),
          <fpage>85</fpage>
          -
          <lpage>89</lpage>
          . http: //www.jstor.org/stable/2963380
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [16]
          <string-name>
            <surname>Brian</surname>
            <given-names>L.</given-names>
          </string-name>
          <string-name>
            <surname>Yoder</surname>
          </string-name>
          .
          <year>2017</year>
          .
          <article-title>Engineering by the Numbers</article-title>
          . Engineering College Proifles &amp; Statistics
          <string-name>
            <surname>Book</surname>
          </string-name>
          (
          <year>2017</year>
          ). http://www.asee.org/papers-and-publications/ publications/college-profiles/15EngineeringbytheNumbersPart1.pdf
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>