<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Evaluations as Research Tools: Gender Differences in Academic Self-Perception and Care Work in Undergraduate Course Reviews</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>David Lang</string-name>
          <email>dnlang86@stanford.edu</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Youjie Chen</string-name>
          <email>minachen@stanford.edu</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Andreas Paepcke</string-name>
          <email>paepcke@stanford.edu</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Mitchell L. Stevens</string-name>
          <email>stevens4@stanford.edu</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Stanford University Stanford</institution>
          ,
          <addr-line>CA</addr-line>
          ,
          <country country="US">USA</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>Student course reviews are rarely considered as research instruments, yet their ubiquity makes them promising tools for education data science. To illustrate this potential, we use a corpus of student reviews to observe gender differences in how students appraise their own learning and in the advice they give to to future students. We find systematic differences in who submits course reviews, with female and academically high-achieving students more likely to submit. Among submitters, we find (a) females understate their achievement of learning goals relative to males earning the same grades; (b) females offer lengthier written advice to future students than males; (c) advice written by females exhibits more positive tone, even after accounting for grades and course selections.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;care work</kwd>
        <kwd>course evaluations</kwd>
        <kwd>gender</kwd>
        <kwd>higher education</kwd>
        <kwd>topic models</kwd>
        <kwd>survey design 1</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>
        While a variety of student characteristics of are of interest to
education data scientists, we focus on students’ gender for reasons both
practical and theoretical. For privacy purposes, our case university
has currently granted researcher access to only a few variables
describing review submitters; we utilize those available data here. Yet
we also have two theoretical motivations for focusing on gender.
First, borrowing from social psychology, we recognize that women
tend to under-estimate their own abilities, while men to to
overestimate, conditional on measured accomplishment [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ]. Second,
borrowing from feminist social science, we posit that submission
of a course review is a form of care work – a voluntary investment
in the well-being of others – and thus implicated differently in
feminine and masculine gender roles and identities [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ]. Our findings
comport with the contours of these larger literatures in ways that
are both important in their own right, and instructive for any future
deployments of course reviews for education data science.
We pursue three sets of analyses below. In the first set, we
observe variation in rates of review submission by gender and earned
grades. These analyses illustrate how researchers might test the
representativeness of corpora of reviews. Second, we observe how
submitters respond to multiple close-ended review prompts targeting
self-assessments of learning, but that are phrased differently. These
analyses illustrate how review design may interact with student
characteristics to produce patterned variation in reported learning
progress. Third, we conduct computational text analyses of
submissions to an open-ended review prompt. These analyses illustrate
how qualitative reviews can be efficiently leveraged for scientific
insight.
2.
2.1
      </p>
    </sec>
    <sec id="sec-2">
      <title>RELATED WORK</title>
    </sec>
    <sec id="sec-3">
      <title>Course Reviews</title>
      <p>
        Research on course reviews typically has focused on questions of
their value as an instruments for evaluating the quality of instruction.
Analyses conducted at scale typically focus on whether measures of
learning are correlated with instructional quality [
        <xref ref-type="bibr" rid="ref27">27</xref>
        ].
      </p>
      <p>
        One potential concern about course reviews from the
pscyhometrics liaterature is the potential for differential item function, a
phenomenon in which respondents of equal ability will exhibit different
responses to a given survey item or question. Studies of differential
item function in course reviews have focused on the quantitative
difficulty of a class, or characteristics of the instructor [
        <xref ref-type="bibr" rid="ref17">17</xref>
        ] [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ].
Relatively little work has focused on how student characteristics may
be associated with differential item function on course reviews.
If findings from copious research on product reviews translate to
academic course reviews, we would expect that students with
highvalence opinions about a course are more likely to respond, resulting
in a bimodal or j-shaped distribution [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ]. In practice, these findings
may not translate. There are often other incentives for filling out
reviews, for example, giving students earlier access to their final
grades as an incentive to respond. Those who submit reviews may
not be representative of the larger population of students who
enrolled in a particular course, or of the overall campus population.
This problem is exacerbated when analysts to not have access of
reviewer characteristics such as gender or grades [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. Past work that
has tried to adjust for non-response bias in review has suggested
that non-response bias tends to favor positive reviews [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ].
There is varying theory around why students may opt to submit
reviews. Studies conclude that students are more likely to respond
to course evaluations if they are majoring in the subject of the course
[
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. Other work suggests that female students are generally more
likely to respond to course evaluations than males [
        <xref ref-type="bibr" rid="ref23">23</xref>
        ]. However,
little work has focused specifically on this topic in Computer Science
courses.
      </p>
      <p>
        Under experimental conditions where researchers manipulated the
information content and valence of course reviews, researchers
found that these factors had material effects on course enrollment
decisions. Students were more likely to enroll in courses if course
evaluations had positive valence, particularly if there was a large
number of such evaluations [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ]. Similar work found that exposure
to positive or negative course reviews had modest to large effects on
students’ expected performance within a course, and their likelihood
of recommending the course in the future[
        <xref ref-type="bibr" rid="ref13">13</xref>
        ]. These findings are
particularly relevant for CS courses as CS courses tend to have
relatively enrollments compared to other subjects.
      </p>
    </sec>
    <sec id="sec-4">
      <title>2.2 Text Analysis</title>
      <p>
        There is a burgeoning literature on using computational text analysis
and methods to quantify differences in corpora based on
characteristics of the author and the text. These techniques have been quickly
adopted to educational applications but we have seen relatively few
instances of text analysis of course reviews. When text analysis of
course evaluations are done, they are typically focused on keyword
extraction and on predicting Likert item responses as a function of
the text [
        <xref ref-type="bibr" rid="ref24">24</xref>
        ].
      </p>
      <p>We group text analyses methods into the following three categories:</p>
      <sec id="sec-4-1">
        <title>2.2.1 Dictionary and rule-based approaches</title>
        <p>
          Dictionary-based approaches characterize the words of documents
into groups of predefined categories such as sentiment. The most
popular of these dictionaries is the Linguistic Inquiry and Word
Count (LIWC) dictionary [
          <xref ref-type="bibr" rid="ref18">18</xref>
          ]. In addition to grouping words into
75 distinct categories and themes (e.g. family, power, death, etc), the
dictionary generates four psycho-social variables that were validated
on college application essays through a rating process. Each of these
variables is scored on a 1 to 99 interval where 1 is a complete lack of
the construct or and 99 is highly pronounced form of the construct.
These constructs are:
1. Tone- This is a summary variable describing the emotional
quality of the text. A score of 99 reflects a positive tone and a
score of 1 reflects a negative tone. A score of 50 represents
neutral valence.
2. Analytic- This is a measure of how much formal logic is used
in the text. A score of 1 indicates little use of formal logic and
a score of 99 exhibits statement with a great deal of formal
logic.
3. Authenticity- This is a measure of the sincerity/honest of
a text. A score of 1 indicates insincerity and a score of 99
indicates high sincerity.
4. Clout- This is a measure of the text’s authority, relative
position, and confidence. A score of 1 suggests relatively little
authority and a score of 99 suggests high authority.
        </p>
        <p>Researchers can also create their own custom measures of text.
We take advantage of this affordance by capturing mentions of
instructors’ names.</p>
      </sec>
      <sec id="sec-4-2">
        <title>2.2.2 Token-based approaches</title>
        <p>
          Token-based approaches treat every word in a text as input into a
model. These approaches often result in the loss of syntactic
meaning but are often very effective at classifying documents.
Tokenbased approaches have proven effective at detecting socioeconomic
features of authors such as race, gender, and income in college
application essays [
          <xref ref-type="bibr" rid="ref4">4</xref>
          ]. Other applications have generated algorithms with
high predictive validity on classroom observation and evaluation
rubrics [
          <xref ref-type="bibr" rid="ref15">15</xref>
          ].
        </p>
      </sec>
      <sec id="sec-4-3">
        <title>2.2.3 Unsupervised approaches</title>
        <p>
          The basic premise behind unsupervised approaches is that texts
include multiple topics, and topics comprise words. Using
unsupervised methods such as Latent Dirichelet Allocation (LDA) , we
can group texts categorically. These same methods have been
augmented recently to allow the distribution of topics to co-vary with
other relevant metadata, a technique known as structural topic
modeling [
          <xref ref-type="bibr" rid="ref22">22</xref>
          ]. In this case, we can examine the concentration of topics
by features such as student gender or grades. This method further
allows us to perform statistical inference to see if topic preponderance
varies systematically by characteristics of authors.
        </p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>2.3 Gender Differences in Academic Experiences, Skill Perception and Care Work</title>
      <p>
        Our work has three motivations from prior social-science literature
on higher education and gender. The first is that male and female
students may have different experiences when taking the same courses.
For example, women are less comfortable asking questions and have
less confidence in CS courses than their male peers [
        <xref ref-type="bibr" rid="ref19">19</xref>
        ]. This
"gender confidence gap" grows as students take more advanced courses
[
        <xref ref-type="bibr" rid="ref3">3</xref>
        ]. Analyses of communal academic resources in CS programs
find substantial differences in how contributions by male and
female users are acknowledged Github and StackOverflow [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ] [
        <xref ref-type="bibr" rid="ref26">26</xref>
        ].
Consequences of these phenomena may extend beyond college,as
women with degrees in STEM fields are less likely than men to enter
STEM occupations [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ]. While course reviews cannot capture
empirical variation in experience per se, they can capture how submitters
make sense of those experiences.
      </p>
      <p>
        Second, gendered differences in skill perception may influence
how students report their experiences and learning gains in reviews.
While women tend to approach STEM fields with less confidence,
men tend to over-estimate their abilities. Experimental work by
Correll [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ] found that men expressed inflated perceptions of their
own skill at completing quantitative tasks compared to women
performing at the same level of measured accomplishment. Together
these inquiries suggest that course reviews may bear traces of
gendered patterns of academic self-perceptions. Our third motivation is
the gendered character of care work. Social scientists define care
work as work that attends to the well-being of others. It comprises
activities and services intended to help other people develop their
capabilities and pursue their goals [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ]. Care work is consistently
associated with femininity and female role expectations, and often
is unpaid or poorly compensated [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ]. To the extent that submitting
course reviews is an act of assistance – to improve classes and to
inform future students – it is appropriately theorized as a form of
care work. Thus we might expect that female and male students will
approach the task of course reviews with different dispositions, such
that the number, extensiveness, and content of course evaluations
may vary by gender of submitters.
      </p>
    </sec>
    <sec id="sec-6">
      <title>3. RESEARCH QUESTIONS</title>
      <p>Our working hypothesis is that course reviews will exhibit gendered
patterns of academic experience, self-perceptions and advice-giving.
Specifically: (1) reviews from male students will exhibit stronger
professed strong learning gains (2) reviews from female students
will exhibit characteristics of care work.</p>
      <p>We group our analyses into two parts. The first part examines
variation by gender and earned grades on review submission rates and
on Likert-scale items on course reviews. The second part
examines variation in male and female responses to a qualitative review
prompt eliciting advice for future students considering the same
courses.</p>
    </sec>
    <sec id="sec-7">
      <title>3.1 Review Submission Rates and Likert Items</title>
      <p>H1: Female students will respond to course evaluations more
often than males.</p>
      <p>Our care work hypothesis is that female students will be more
responsive to institutional requests for reviews. We investigate this
hypotheses using an exact binomial-two-sample test. We examine
results by gender and grade.</p>
      <p>H2: There are systematic differences in response rate by grade.
There are many competing theories of how grades might influence
response rates to course evaluations. If students have a poor grade,
they may be more inclined to view the evaluation as an opportunity
to retaliate against the grader. Alternatively, students who receive
a low grade may opt to avoid opportunities to reflect on negative
experiences. We will investigate this hypothesis utilizing a simple
c2 test of response-rates by grade.</p>
      <p>H3: Female students will understate their achievements relative
to their male counterparts. We hypothesize that men are more
likely to see course reviews as a form of positive self-reflection and
promotion, and that females are more likely reviews as a form of
care work. We believe these differences will have stronger valence
in items that focus on a students’ accomplishments rather than other
constructs such as student learning. We will model these analyses
as a fixed-effect regression model with the following specification:
Yi j = b1Malei + Gradesi j + G j + ei j
(1)
The subscripts i and j correspond to indices for student and course.
The Y variable corresponds to our focal outcome variable, in this
case, responses to a Likert item. Male corresponds to a student’s
self-reported indicator variable of whether the student identifies
as male and b1 corresponds to the associated coefficient with this
variable. We represent course effects with Gi to control for factors
like the difficulty of the course or instructional quality. We also
control for grades with an additional fixed-effect for each possible
grade a student could receive 2. The error is represented by e. Errors
are clustered at the course level.</p>
    </sec>
    <sec id="sec-8">
      <title>3.2 Open-Response Questions</title>
      <p>We pay particular interest to open-response items in course reviews.
We suspect that such items are may be the most valuable and least
explored element of course reviews. As such, we may be able to
detect subtle differences in qualitative responses.</p>
      <sec id="sec-8-1">
        <title>3.2.1 Psychosocial variables</title>
        <p>H4: Course evaluations written by females will express more
positive and sincere sentiment.</p>
        <p>Given our care work hypothesis, we believe that female students
will express more positive sentiment in open-response items. We
use the same analytical strategy as an equation 1 using LIWC’s
tone variable. Specifically, we examine gender differences in these
psycho-social variables after controlling for variation that can can
be attributed to the course, or to student grade. We report outcomes
in standardized effect sizes to facilitate interpretability.
We also hypothesize that a corollary to the care work hypothesis
is that female students will use more "I" statements and tentative
language. This tendency would manifest as reviews written by
female students exhibiting more authentic language.</p>
        <p>H5: Course evaluations written by male students will express
more clout. Based on prior literature pertaining to a confidence
gap in CS by gender, we hypothesize this trend should manifest
with less expressions of clout and authority in course evaluations by
female authors.</p>
      </sec>
      <sec id="sec-8-2">
        <title>3.2.2 Hand-crafted rules</title>
        <p>H6: Female students will write more on course evaluations and
mention the instructor more often.</p>
        <p>We hypothesize that care work will manifest in other ways beyond
psycho-social variables. Specifically: female submitters will put
more effort into reviews by writing more; and they will take a more
individualized approach by mentioning the instructor explicitly.
2in our analyses, there are over twenty grade types, including + and
- variants as well as credit and nocredit courses. We report A,B,C,D,
and not passing grades for simplicity
We have crafted two simple measures to facilitate investigation of
this hypothesis: the length of each response in number of words,
and a capture of each instance of an instructor name.</p>
      </sec>
      <sec id="sec-8-3">
        <title>3.2.3 Topic models</title>
        <p>H7: There will by systematic variation in topics depending on the
author’s gender.</p>
        <p>Our final analysis is exploratory using structural topic models to
identify whether qualitative components of the corpora
systematically vary with gender of submitter The goals of this analysis are
to develop efficient means of sorting and categorizing qualitative
components of course reviews.</p>
      </sec>
    </sec>
    <sec id="sec-9">
      <title>4. DATA</title>
      <p>Data comprise information describing enrollments in courses offered
through the Computer Science (CS) Department of a private research
university during the 2015-16 and 2016-17 academic years, and the
entire population of formal reviews submitted by students enrolled
in those courses. Reviews were administered near the end of the
academic term but before the beginning of the term’s official final
exam period. As an incentive for submitting reviews, students were
given the ability to see their final course grades a bit earlier than
non-submitters.</p>
      <p>In total these data yield 11,255 student responses from 251 courses.
Courses range in character from very large introductory
lecture-andlab formats to small advanced seminars. Institutional data made
available to us for analysis include each student’s grade, gender,
GPA, declared major (if known), and academic year. We combine
these data with the corpus of reviews submitted for CS courses
during the study period specified above. Approximately one-third of
submitted reviews from female students, and approximately half are
from undergraduates. We cannot track or identify students enrolling
in multiple CS courses during the study period, however we can
compute and generate response rates by grade and gender.
We limit our analysis to responses in which submitters offered a
response to the review’s only open-ended question. That question
reads:
"What would you like to say about this course to a student who is
considering taking it in the future?"
The prompt is very well aligned with our care work hypotheses, in
that it specifically asks submitters to give advice to a hypothetical
future student. Individual responses vary substantially in length:
from a single character to over 5,964 characters (the latter equivalent
to 1004 words). The mean response length is 132 characters –
approximately the length of a tweet. The entire corpus of responses
to this question is 300,000 words.</p>
      <p>Additionally we analyze responses to two review prompts with
five-point Likert responses: 3</p>
      <p>How much did you learn from this course?</p>
      <p>How well did you achieve the learning goals of this course?
3we will limit this analysis to complete cases due to the fact that one
item was not consistently administered across courses. We ignore
questions pertaining to quality of instruction and focus on student
learning goals.
Aggregated responses to these two prompts appear in Figure 1.</p>
    </sec>
    <sec id="sec-10">
      <title>5. ANALYSES</title>
    </sec>
    <sec id="sec-11">
      <title>5.1 H1: Response Rates By Gender</title>
      <p>Rates of review submission by student gender and earned grade are
reported in Figure 2. Two features are notable. First, females are
more likely to submit overall. On average, females submit to 78.0%
of opportunities to do so; males, 74.5% (p&lt;.001).</p>
      <p>Second, those receiving higher grades in a course are more likely
to submit reviews Females receiving a grade of "A" are 3.7% more
likely to respond to submit than their male counterparts. The
gender submission gap is greatest among students receiving a grade of
"B," with female "B" recipients 6.5% (p&lt;.001)more likely to
submit than males. We do not observe statistically significant gender
differences in submission rates for those receiving grades below
"B," however such grades represent fewer than 5% of grades in the
research sample.</p>
    </sec>
    <sec id="sec-12">
      <title>5.2 H2: Variation in Submission by Grade</title>
      <p>We also examined whether review submission varied systematically
grades. Figure 2 indicates a strong positive correlation between
grade and likelihood of submission. A student with a grade of A or
higher has an 80% chance of responding to the evaluation, while
students who do not pass the course or receive credit without a
grade responded approximately 50 percent of the time. Effectively,
this means that students who fail courses are represented by review
submissions about half as often as students who excel in a courses.
Differences are statistically significant with a c2 statistic of 395.37
and a p-value of less than 0.001.</p>
    </sec>
    <sec id="sec-13">
      <title>5.3 H3: Reports of Learning and Goal-Meeting</title>
      <p>We observe the proportion of students reporting having achieved the
learning goals of a course Extremely Well by grade in Figure 3. Not
surprisingly, we find a strong direct correlation with course grade,
such that reported goal achievement declines with grade. What is
striking is that at every grade level, there is a clear gap in reported
goal achievement, with males more likely to report achievement
than females earning the same grade.</p>
      <p>We extend this analysis to see if this same pattern occurs with the
question of how much students learn. Using the same specification
as described in equation 1 in table 1. We see that after controlling
for grades and course, males and females exhibit no differences</p>
      <p>Male
in self-reported measures of ‘how much they learned’ in a class.
However, when we look at a similar question about ‘achievement
of learning goals’, we see a stark difference. Male students are 8%
points more likely than females to state they mastered the learning
goals of a course. This finding suggests two concerns. First, given
the similarity of these questions, we see that subtle differences in
phrasing yield substantial differences in student responses. Second,
female students report lower-level of mastery even after controlling
for grades. Notably, these surveys are collected before students
know their final grades. These perceptions may change after this
information is revealed to them.
5.4
5.4.1
graphicx</p>
    </sec>
    <sec id="sec-14">
      <title>Open Text Responses</title>
      <sec id="sec-14-1">
        <title>Psycho-social variables</title>
        <p>We report our analyses for H4 and H5 in table 2. We find modest
variation by gender in how submitters describe their experience in
the same course, conditional on grades. On average, submissions
from males evince slightly more negative and slightly less authentic
language. While these gender differences are highly significant, their
magnitude is modest: on the order of a tenth of a standard deviation.
Nevertheless, they are consistent with our care work hypotheses. To
wit, men are somewhat more critical and less honest in their reviews
than women, suggesting greater empathy and investment among
female submitters.</p>
        <p>With respect to our hypothesis around clout, we find little evidence
that qualitative open-responses exhibit any significant differences in</p>
      </sec>
      <sec id="sec-14-2">
        <title>5.4.2 Handcrafted features</title>
        <p>We report the standardized results of our analysis in table 3. We
observe marginally significant differences in the frequency with which
submissions from males and females mention instructor name, with
women approximately two percent more likely to mention.
Submissions from women are also lengthier – about .15 of a standard
deviation. While modest in magnitude, these statistically significant
findings comport with our care work hypotheses that female
submitters approach the task of submitting reviews with more attention to
specificity and investment.</p>
      </sec>
    </sec>
    <sec id="sec-15">
      <title>5.5 Topic Models</title>
      <p>
        We ran structural topic models while allowing submitter gender
to vary with topic prevalence. We tuned the optimal number of
topics from 2 to 50 using an exclusivity measure called FREX
[
        <xref ref-type="bibr" rid="ref2">2</xref>
        ] [
        <xref ref-type="bibr" rid="ref20">20</xref>
        ]. FREX (See Equation 2) is a harmonic weighting of the
frequency (F ) with which a word occurs in a topic; and exclusivity
(E), how frequently the word occurs in a given topic relative to
others. The parameter w corresponds to a tuning parameter of the
relative importance of these features. We used the default parameter
of w = :7 to favor topics that had more exclusivity.
      </p>
      <p>F REX = (</p>
      <p>+
w
E
1 w ) 1</p>
      <p>F
(2)
Using this criterion, We found the locally optimal parameter to be
32 distinct topics. We then hand labelled each topic, observing the
ten (10) statements that had the highest probability labels of that
topic see Figure 4 (Top). The most common topics were comments
about a course being a tutorial, a suggestion to take the course, or
a positive review. The least common topics were highly specific
suggestions and issues pertaining to course prerequisites. While we
did not see substantial gender variation in topics overall, there are
exceptions of note. First, submissions from males are more likely to
talk about math prerequisites and claims that instruction was poorly
organized or of poor quality. They were also more likely to discuss
course organization and instruction. Submissions from females were
more likely to bear topics pertaining to workload, study practices,
and attendance. These patterns provide at least modest evidence
that women are offering relatively more specific advice that may be
relevant to larger numbers of future students.</p>
      <p>We note an important caveat to this analysis, however. In contrast
with the above studies results of the topic models presented in this
section do not control for grades or course selections, thus reported
gender differences in topic prevalence may be an artifact of these
other factors. We attempted to model the data with all of these
parameters but found the models to be degenerate.</p>
    </sec>
    <sec id="sec-16">
      <title>6. DISCUSSION</title>
      <p>Even while they are controversial for evaluating instructors and
instruction, course reviews are ubiquitous features of the US higher
education landscape and potentially powerful tools for education
data science. In the work presented here we have sought to
demonstrate the promise of course reviews as a window into students’
perceptions of their academic experiences and their orientation to
the task of submitting evaluations. Taking advantage of archival data
that included 11,255 submitted to 251 computer science courses at
a single university between 2015-2017 that was linked to
administrative information describing submitters’ gender (M/F) and grades,
we found patterned variation in who submits course reviews, and
how.</p>
      <p>In three observational studies we found that (a) women and those
earning high grades were disproportionately likely to submit
reviews (b) the phrasing of close-ended review prompts influenced
patterns of response by gender (c) responses to qualitative review
prompts differed subtly but significantly by gender, with women
writing somewhat more positive, individualized, and lengthier
reviews. These empirical findings comport with theoretical insights
from educational social psychology and feminist social science,
which suggest gender variation in how men and women perceive
their own academic accomplishments and their obligations for the
well-being of others.</p>
      <p>While the empirical findings presented here are modest, they suggest
the promise of leveraging course reviews for cumulative science in
at least two ways.</p>
      <p>First, we note that the inquiries presented here are based entirely on
the premise that course reviews and submitter demographic
information are "found" data. To the extent that virtually every US college
and university possesses data such as these, we can only imagine the
number and variety of insights that might be gained from parallel
investigations at other schools. To a nascent field whose promise
lies substantially in observing phenomena at scale, course reviews
provide exceptionally promising sources of data for education data
science.</p>
      <p>Second, there is every reason to imagine that education data
scientists might collaborate with school administrators to more explicitly
and conscientiously instrument reviews for systematic experimental
and quasi-experimental research. The basic conditions for such
inquiries are already in place and sustained by established
administrative rhythms: schools have offices conducting the reviews, students
anticipate receiving them, and they take place multiple times a year.
It is possible to imagine substantial scientific insight through the
linkage review subsmissions with with administrative data
describing characteristics of submitters. The initial efforts presented here
provide an inkling of this promise.</p>
      <p>
        As with any novel research strategy, pursuing education data science
through course reviews comes with important ethical considerations
regarding participant consent and responsible use. We are grateful
that such discussions are already well underway nationwide [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ]
and we hope that our own illustrative work here might helpfully
contribute to them. Indeed, addressing questions of responsible use
of student data in the context of course reviews may have the
additional benefit of improving the collective value of an institutional
practice currently regarded with ambivalence and suspicion but that,
in whatever form, will likely be part of the academic landscape for
a long time.
      </p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <surname>Meredith</surname>
            <given-names>J. D.</given-names>
          </string-name>
          <string-name>
            <surname>Adams</surname>
            and
            <given-names>Paul D.</given-names>
          </string-name>
          <string-name>
            <surname>Umbach</surname>
          </string-name>
          .
          <year>2012</year>
          .
          <article-title>Nonresponse and Online Student Evaluations of Teaching: Understanding the Influence of Salience, Fatigue</article-title>
          , and Academic Environments. Research in Higher Education
          <volume>53</volume>
          ,
          <issue>5</issue>
          (8
          <year>2012</year>
          ),
          <fpage>576</fpage>
          -
          <lpage>591</lpage>
          . DOI: http://dx.doi.org/10.1007/s11162-011-9240-5
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <surname>Edoardo</surname>
            <given-names>M</given-names>
          </string-name>
          <string-name>
            <surname>Airoldi and Jonathan M Bischof</surname>
          </string-name>
          .
          <year>2012</year>
          .
          <article-title>A Poisson convolution model for characterizing topical content with word frequency and exclusivity</article-title>
          . arxiv.org (
          <year>2012</year>
          ). https://arxiv.org/abs/1206.4631http: //arxiv.org/abs/1206.4631
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>Christine</given-names>
            <surname>Alvarado</surname>
          </string-name>
          , Yingjun Cao, and
          <string-name>
            <given-names>Mia</given-names>
            <surname>Minnes</surname>
          </string-name>
          .
          <year>2017</year>
          .
          <article-title>Gender Differences in Students' Behaviors in CS Classes throughout the CS Major</article-title>
          .
          <source>In Proceedings of the 2017 ACM SIGCSE Technical Symposium on Computer Science Education. ACM</source>
          , New York, NY, USA,
          <fpage>27</fpage>
          -
          <lpage>32</lpage>
          . DOI: http://dx.doi.org/10.1145/3017680.3017771
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>A.J.</given-names>
            <surname>Alvero</surname>
          </string-name>
          , Noah Arthurs, anthony lising antonio, Benjamin W. Domingue, Ben Gebre-Medhin,
          <string-name>
            <given-names>Sonia</given-names>
            <surname>Giebel</surname>
          </string-name>
          , and Mitchell L.
          <string-name>
            <surname>Stevens</surname>
          </string-name>
          .
          <year>2020</year>
          .
          <article-title>AI and Holistic Review</article-title>
          .
          <source>In Proceedings of the AAAI/ACM Conference on AI, Ethics, and Society</source>
          . ACM, New York, NY, USA,
          <fpage>200</fpage>
          -
          <lpage>206</lpage>
          . DOI: http://dx.doi.org/10.1145/3375627.3375871
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <surname>David</surname>
            <given-names>N.</given-names>
          </string-name>
          <string-name>
            <surname>Beede</surname>
          </string-name>
          , Tiffany A.
          <string-name>
            <surname>Julian</surname>
          </string-name>
          , David Langdon,
          <string-name>
            <surname>George</surname>
            <given-names>McKittrick</given-names>
          </string-name>
          ,
          <string-name>
            <given-names>Beethika</given-names>
            <surname>Khan</surname>
          </string-name>
          , and
          <string-name>
            <given-names>Mark E.</given-names>
            <surname>Doms</surname>
          </string-name>
          .
          <year>2011</year>
          .
          <article-title>Women in STEM: A Gender Gap to Innovation</article-title>
          .
          <source>SSRN Electronic Journal (8</source>
          <year>2011</year>
          ). DOI: http://dx.doi.org/10.2139/ssrn.1964782
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>Michael</given-names>
            <surname>Brown</surname>
          </string-name>
          and
          <string-name>
            <given-names>Carrie</given-names>
            <surname>Klein</surname>
          </string-name>
          .
          <year>2020</year>
          .
          <article-title>Whose Data? Which Rights? Whose Power? A Policy Discourse Analysis of Student Privacy Policy Documents</article-title>
          .
          <source>The Journal of Higher Education</source>
          (
          <year>2020</year>
          ),
          <fpage>1</fpage>
          -
          <lpage>30</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <surname>Shelley</surname>
            <given-names>J.</given-names>
          </string-name>
          <string-name>
            <surname>Correll</surname>
          </string-name>
          .
          <year>2004</year>
          .
          <article-title>Constraints into Preferences: Gender, Status, and Emerging Career Aspirations</article-title>
          .
          <source>American Sociological Review</source>
          <volume>69</volume>
          ,
          <issue>1</issue>
          (2
          <year>2004</year>
          ),
          <fpage>93</fpage>
          -
          <lpage>113</lpage>
          . DOI: http://dx.doi.org/10.1177/000312240406900106
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>Erica</given-names>
            <surname>DeFrain and Erica DeFrain</surname>
          </string-name>
          .
          <year>2016</year>
          .
          <article-title>An Analysis of Differences in Non-Instructional Factors Affecting Teacher-Course Evaluations over Time and Across Disciplines</article-title>
          . (
          <year>2016</year>
          ). https://repository.arizona.edu/handle/10150/621018
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>Paula</given-names>
            <surname>England</surname>
          </string-name>
          .
          <year>2005</year>
          .
          <article-title>Emerging Theories of Care Work</article-title>
          .
          <source>Annual Review of Sociology 31</source>
          ,
          <issue>1</issue>
          (8
          <year>2005</year>
          ),
          <fpage>381</fpage>
          -
          <lpage>399</lpage>
          . DOI: http: //dx.doi.org/10.1146/annurev.soc.
          <volume>31</volume>
          .041304.122317
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>Nancy</given-names>
            <surname>Folbre</surname>
          </string-name>
          .
          <year>1995</year>
          . “
          <article-title>Holding hands at midnight”: The paradox of caring labor</article-title>
          .
          <source>Feminist Economics</source>
          <volume>1</volume>
          ,
          <issue>1</issue>
          (3
          <year>1995</year>
          ),
          <fpage>73</fpage>
          -
          <lpage>92</lpage>
          . DOI:http://dx.doi.org/10.1080/714042215
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>Maarten</given-names>
            <surname>Goos</surname>
          </string-name>
          and
          <string-name>
            <given-names>Anna</given-names>
            <surname>Salomons</surname>
          </string-name>
          .
          <year>2017</year>
          .
          <article-title>Measuring teaching quality in higher education: assessing selection bias in course evaluations</article-title>
          .
          <source>Research in Higher Education</source>
          <volume>58</volume>
          ,
          <issue>4</issue>
          (6
          <year>2017</year>
          ),
          <fpage>341</fpage>
          -
          <lpage>364</lpage>
          . DOI: http://dx.doi.org/10.1007/s11162-016-9429-8
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <surname>Nan</surname>
            <given-names>Hu</given-names>
          </string-name>
          , Jie Zhang, and
          <string-name>
            <given-names>Paul A.</given-names>
            <surname>Pavlou</surname>
          </string-name>
          .
          <year>2009</year>
          .
          <article-title>Overcoming the J-shaped distribution of product reviews</article-title>
          .
          <source>(10</source>
          <year>2009</year>
          ). DOI: http://dx.doi.org/10.1145/1562764.1562800
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <given-names>Neneh</given-names>
            <surname>Kowai-Bell</surname>
          </string-name>
          , Rosanna E. Guadagno, Tannah Little, Najean Preiss, and
          <string-name>
            <given-names>Rachel</given-names>
            <surname>Hensley</surname>
          </string-name>
          .
          <year>2011</year>
          .
          <article-title>Rate My Expectations: How online evaluations of professors impact students' perceived control</article-title>
          .
          <source>Computers in Human Behavior</source>
          <volume>27</volume>
          ,
          <issue>5</issue>
          (9
          <year>2011</year>
          ),
          <fpage>1862</fpage>
          -
          <lpage>1867</lpage>
          . DOI: http://dx.doi.org/10.1016/J.CHB.
          <year>2011</year>
          .
          <volume>04</volume>
          .009
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <given-names>Cong</given-names>
            <surname>Li</surname>
          </string-name>
          and
          <string-name>
            <given-names>Xiuli</given-names>
            <surname>Wang</surname>
          </string-name>
          .
          <year>2013</year>
          .
          <article-title>The power of eWOM: A re-examination of online student evaluations of their professors</article-title>
          .
          <source>Computers in Human Behavior</source>
          <volume>29</volume>
          ,
          <issue>4</issue>
          (7
          <year>2013</year>
          ),
          <fpage>1350</fpage>
          -
          <lpage>1357</lpage>
          . DOI: http://dx.doi.org/10.1016/J.CHB.
          <year>2013</year>
          .
          <volume>01</volume>
          .007
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <given-names>Jin</given-names>
            <surname>Liu</surname>
          </string-name>
          and
          <string-name>
            <given-names>Julie</given-names>
            <surname>Cohen</surname>
          </string-name>
          .
          <year>2020</year>
          .
          <article-title>Measuring Teaching Practices at Scale: A Novel Application of Text-as-Data Methods | EdWorkingPapers</article-title>
          . (
          <year>2020</year>
          ). https://www.edworkingpapers.com/ai20-
          <fpage>239</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [16]
          <string-name>
            <surname>Anna</surname>
            <given-names>May</given-names>
          </string-name>
          , Johannes Wachs, and
          <string-name>
            <given-names>Anikó</given-names>
            <surname>Hannák</surname>
          </string-name>
          .
          <year>2019</year>
          .
          <article-title>Gender differences in participation and reward on Stack Overflow</article-title>
          .
          <source>Empirical Software Engineering</source>
          <volume>24</volume>
          ,
          <issue>4</issue>
          (8
          <year>2019</year>
          ),
          <fpage>1997</fpage>
          -
          <lpage>2019</lpage>
          . DOI: http://dx.doi.org/10.1007/s10664-019-09685-x
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [17]
          <string-name>
            <given-names>Arunachalam</given-names>
            <surname>Narayanan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>William J</given-names>
            .
            <surname>Sawaya</surname>
          </string-name>
          , and
          <string-name>
            <given-names>Michael D.</given-names>
            <surname>Johnson</surname>
          </string-name>
          .
          <year>2014</year>
          .
          <article-title>Analysis of Differences in Nonteaching Factors Influencing Student Evaluation of Teaching between Engineering and Business Classrooms</article-title>
          .
          <source>Decision Sciences Journal of Innovative Education</source>
          <volume>12</volume>
          ,
          <issue>3</issue>
          (7
          <year>2014</year>
          ),
          <fpage>233</fpage>
          -
          <lpage>265</lpage>
          . DOI:http://dx.doi.org/10.1111/dsji.12035
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          [18]
          <string-name>
            <given-names>JW</given-names>
            <surname>Pennebaker</surname>
          </string-name>
          , RL Boyd,
          <string-name>
            <given-names>K</given-names>
            <surname>Jordan</surname>
          </string-name>
          , and
          <string-name>
            <given-names>K</given-names>
            <surname>Blackburn</surname>
          </string-name>
          .
          <year>2015</year>
          .
          <article-title>The development and psychometric properties of LIWC2015</article-title>
          . (
          <year>2015</year>
          ). https: //repositories.lib.utexas.edu/handle/2152/31333
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          [19]
          <string-name>
            <surname>Katie</surname>
            <given-names>Redmond</given-names>
          </string-name>
          ,
          <string-name>
            <given-names>Sarah</given-names>
            <surname>Evans</surname>
          </string-name>
          , and
          <string-name>
            <given-names>Mehran</given-names>
            <surname>Sahami</surname>
          </string-name>
          .
          <year>2013</year>
          .
          <article-title>A large-scale quantitative study of women in computer science</article-title>
          at Stanford University.
          <source>In Proceeding of the 44th ACM technical symposium on Computer science education - SIGCSE '13</source>
          . ACM Press, New York, New York, USA, 439. DOI:http://dx.doi.org/10.1145/2445196.2445326
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          [20]
          <string-name>
            <given-names>J</given-names>
            <surname>Reich</surname>
          </string-name>
          , DH Tingley,
          <string-name>
            <given-names>J</given-names>
            <surname>Leder-Luis</surname>
          </string-name>
          , and
          <string-name>
            <given-names>M</given-names>
            <surname>Roberts</surname>
          </string-name>
          .
          <year>2014</year>
          .
          <article-title>Computer-assisted reading and discovery for student generated text in massive open online courses</article-title>
          . (
          <year>2014</year>
          ). https://papers.ssrn.com/sol3/papers.cfm?abstract_id=
          <fpage>2499725</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          [21]
          <string-name>
            <surname>Lauren</surname>
            <given-names>A Rivera</given-names>
          </string-name>
          and
          <string-name>
            <given-names>András</given-names>
            <surname>Tilcsik</surname>
          </string-name>
          .
          <year>2019</year>
          .
          <article-title>Scaling down inequality: Rating scales, gender bias, and the architecture of evaluation</article-title>
          .
          <source>American Sociological Review</source>
          <volume>84</volume>
          ,
          <issue>2</issue>
          (
          <year>2019</year>
          ),
          <fpage>248</fpage>
          -
          <lpage>274</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          [22]
          <string-name>
            <surname>Margaret</surname>
            <given-names>E.</given-names>
          </string-name>
          <string-name>
            <surname>Roberts</surname>
          </string-name>
          ,
          <string-name>
            <surname>Brandon M. Stewart</surname>
            , Dustin Tingley, Christopher Lucas, Jetson Leder-Luis, Shana Kushner Gadarian, Bethany Albertson, and
            <given-names>David G.</given-names>
          </string-name>
          <string-name>
            <surname>Rand</surname>
          </string-name>
          .
          <year>2014</year>
          .
          <article-title>Structural Topic Models for Open-Ended Survey Responses</article-title>
          .
          <source>American Journal of Political Science</source>
          <volume>58</volume>
          ,
          <issue>4</issue>
          (
          <issue>10</issue>
          <year>2014</year>
          ),
          <fpage>1064</fpage>
          -
          <lpage>1082</lpage>
          . DOI:http://dx.doi.org/10.1111/ajps.12103
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          [23]
          <string-name>
            <surname>Linda</surname>
            <given-names>J.</given-names>
          </string-name>
          <string-name>
            <surname>Sax</surname>
          </string-name>
          ,
          <string-name>
            <surname>Shannon K. Gilmartin</surname>
          </string-name>
          , and
          <string-name>
            <surname>Alyssa</surname>
            <given-names>N.</given-names>
          </string-name>
          <string-name>
            <surname>Bryant</surname>
          </string-name>
          .
          <year>2003</year>
          .
          <article-title>Assessing response rates and nonresponse bias in web and paper surveys</article-title>
          . (
          <year>2003</year>
          ). DOI: http://dx.doi.org/10.1023/A:1024232915870
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          [24]
          <string-name>
            <given-names>T</given-names>
            <surname>Sliusarenko</surname>
          </string-name>
          , LH Clemmensen International . . . ,
          <source>and Undefined</source>
          <year>2013</year>
          .
          <year>2013</year>
          .
          <article-title>Text Mining in Students' Course Evaluations</article-title>
          . pdfs.semanticscholar.org (
          <year>2013</year>
          ). https://pdfs.semanticscholar.org/cb02/ b880ef86371461b3ebe46d2f8c293b43c7a2.pdf
        </mixed-citation>
      </ref>
      <ref id="ref25">
        <mixed-citation>
          [25]
          <string-name>
            <surname>Philip</surname>
            <given-names>Stark</given-names>
          </string-name>
          , Kellie Ottoboni, and
          <string-name>
            <given-names>Anne</given-names>
            <surname>Boring</surname>
          </string-name>
          .
          <year>2016</year>
          .
          <article-title>Student Evaluations of Teaching (Mostly) Do Not Measure Teaching Effectiveness</article-title>
          .
          <source>ScienceOpen Research</source>
          (
          <year>2016</year>
          ). DOI:http: //dx.doi.org/10.14293/s2199-
          <fpage>1006</fpage>
          .1.sor-edu.
          <source>aetbzc.v1</source>
        </mixed-citation>
      </ref>
      <ref id="ref26">
        <mixed-citation>
          [26]
          <string-name>
            <surname>Josh</surname>
            <given-names>Terrell</given-names>
          </string-name>
          , Andrew Kofink, Justin Middleton, Clarissa Rainear, Emerson Murphy-Hill,
          <string-name>
            <given-names>Chris</given-names>
            <surname>Parnin</surname>
          </string-name>
          , and
          <string-name>
            <given-names>Jon</given-names>
            <surname>Stallings</surname>
          </string-name>
          .
          <year>2017</year>
          .
          <article-title>Gender differences and bias in open source: pull request acceptance of women versus men</article-title>
          .
          <source>PeerJ Computer Science</source>
          <volume>3</volume>
          (
          <issue>5</issue>
          2017),
          <year>e111</year>
          . DOI: http://dx.doi.org/10.7717/peerj-cs.
          <fpage>111</fpage>
        </mixed-citation>
      </ref>
      <ref id="ref27">
        <mixed-citation>
          [27]
          <string-name>
            <surname>Bob</surname>
            <given-names>Uttl</given-names>
          </string-name>
          ,
          <string-name>
            <given-names>Carmela A.</given-names>
            <surname>White</surname>
          </string-name>
          , and Daniela Wong Gonzalez.
          <year>2017</year>
          .
          <article-title>Meta-analysis of faculty's teaching effectiveness: Student evaluation of teaching ratings and student learning are not related</article-title>
          .
          <source>Studies in Educational Evaluation</source>
          <volume>54</volume>
          (9
          <year>2017</year>
          ),
          <fpage>22</fpage>
          -
          <lpage>42</lpage>
          . DOI: http://dx.doi.org/10.1016/J.STUEDUC.
          <year>2016</year>
          .
          <volume>08</volume>
          .007
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>