<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Comparing Feature Engineering Approaches to Predict Complex Programming Behaviors</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Wengran Wang</string-name>
          <email>wwang33@ncsu.edu</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Yudong Rao</string-name>
          <email>yrao3@ncsu.edu</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Yang Shi</string-name>
          <email>yshi26@ncsu.edu</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Alexandra Milliken</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Chris Martens</string-name>
          <email>crmartens@ncsu.edu</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Tiffany Barnes</string-name>
          <email>tmbarnes@ncsu.edu</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Thomas W. Price</string-name>
          <email>twprice@ncsu.edu</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>North Carolina State University</institution>
          ,
          <country country="US">USA</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>syntax tree (AST), and a variety of approaches have been proposed to extract features from these ASTs to use in learning algorithms, but no work has directly compared their e ectiveness. In this paper, we do so by comparing three di erent feature engineering approaches for classifying the behavior of novices' open-ended programming projects according to expert labels. In order to evaluate the e ectiveness of these feature engineering approaches, we handlabeled a dataset of novice programs from the Scratch repository to indicate the presence of ve complex, game-related programming behaviors. We compared these feature engineering approaches by evaluating their classi cation e ectiveness. Our results show that the three approaches perform similarly across di erent target labels. However, we also nd evidence that all approaches led to over tting, suggesting the need for future research to select and reduce code features, which may reveal advantages in more complex feature engineering approaches.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. INTRODUCTION</title>
      <p>
        Automatically classifying student code using machine
learning has many applications in computer science education,
such as to automatically grade students' code [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ], to predict
when students are unlikely to succeed at a task and may
bene t from feedback [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ], and to propagate feedback on
particular misconceptions to students who need it [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ].
However, a fundamental challenge in applying machine
learning to source code is how to represent that code in a way
the learner can understand. The structure of programming
code is traditionally represented as an abstract syntax tree
Copyright c 2020 for this paper by its authors. Use permitted under
Creative Commons License Attribution 4.0 International (CC BY 4.0).
(AST), where nodes and their children correspond to
speci c code elements (e.g., an if statement), and the tree can
be arbitrarily large. However, the vast majority of machine
learning models take xed-length vectors as input.
In many domains, researchers have addressed this challenge
of code representation by extracting a set of features from
source code, which can then represent the code in the model.
For example, a simple Bag-of-Words (BoW) approach
represents code as a binary vector, where each element indicates
the presence or absence of a speci c AST node anywhere in
the code (e.g., [
        <xref ref-type="bibr" rid="ref11 ref18 ref6">6, 18, 11</xref>
        ]). However, simple feature
extraction approaches like BoW do not capture the complex
structural relationships among AST nodes in student code, which
may be important for many classi cation tasks. A number
of other feature extraction approaches have been proposed
(e.g., [
        <xref ref-type="bibr" rid="ref2 ref28 ref4">2, 28, 4</xref>
        ]), but no work has directly compared their
e ectiveness for classifying student code. Further, feature
extraction can be especially di cult in the domain of
computer science education, where students' code may cover a
large and sparse solution space, with little overlap among
solutions paths [
        <xref ref-type="bibr" rid="ref13 ref21 ref27">21, 27, 13</xref>
        ]. This suggests the need to
develop new feature extraction approaches that address this
challenge.
      </p>
      <p>In this paper, we compared the e ectiveness of three code
feature extraction approaches, on a challenging and
generalizable classi cation task. This task classi es programming
game design projects to identify the presence and absence of
complex game behaviors. We found that the three code
feature extraction approaches had similar performance across
all behaviors, and the performance of speci c feature
extraction approaches is dependent on factors such as the
properties of the target label, the size of training data, and the
prevalence of positive labels. Our work contributes to
educational data mining for CS education by comparing the
a ordance of di erent feature engineering approaches and
evaluating their e ectiveness in predicting the presence of
complex game behaviors.</p>
    </sec>
    <sec id="sec-2">
      <title>2. RELATED WORK</title>
      <p>In this section, rst, we discuss the relevance and importance
of automatic code classi cation for improving computing
education through personalization and scalability. We then
summarize state-of-the-art feature engineering approaches
that related work has used for various code analysis
purposes.</p>
    </sec>
    <sec id="sec-3">
      <title>2.1 Applications of Classifying Student Code</title>
      <p>
        Manual labeling of student code is a frequent practice in
computing education. It is often done by instructors and
researchers, for example, to grade student program
submissions, identify misconceptions [
        <xref ref-type="bibr" rid="ref15 ref23">15, 23</xref>
        ], or to pro le a
programming dataset to identify when particular code
features are used [
        <xref ref-type="bibr" rid="ref12">12</xref>
        ]. However, labeling tasks are quite
timeconsuming and hard to scale, leading many researchers to
investigate methods for automatically labeling code.
Researchers have used di erent approaches to automatically
classify and analyze students' code, such as using correct
program submissions to generate rubric-based auto-graders
[
        <xref ref-type="bibr" rid="ref8">8</xref>
        ], or using programming homework grades to infer
students' knowledge [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]. Elmadami et al., for example, built a
data-driven misconception classi er in EER-Tutor, an
Intelligent Tutoring System that provides tutorials for database
design. Using association rule mining, EER-Tutor
categorizes frequent failing patterns as indicators of misconceptions
[
        <xref ref-type="bibr" rid="ref9">9</xref>
        ]. Similarly, Mao et al. developed a classi er that predicts
a student's success in completing a programming task based
on their programming code trajectory with just one minute
of data from a student [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ]. This classi er, if implemented in
programming education, could help instructors or learning
systems to prompt students with suggestions or feedback
when they most need it. In addition, our prior work has
shown the e cacy of adding an automatic code classi er to a
learning system. We developed an unsupervised classi er to
identify completions of 11 sub-goals in a Snap! block-based
programming task [
        <xref ref-type="bibr" rid="ref27">27</xref>
        ]. We then integrated this classi er in
a programming environment to detect sub-goal completions
and provide timely positive feedback to students, which
signi cantly increased the time students spent engaged with
the programming task [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ]. These results suggest that, in
programming learning environments, automatic code
classication can help provide adaptive and scalable feedback to
support students.
      </p>
    </sec>
    <sec id="sec-4">
      <title>2.2 AST Structural Feature Extraction</title>
      <p>
        Researchers have used various approaches to extract a xed
set of features from code to use in machine learning models.
Structural feature extraction looks for patterns in an AST
and creates a binary input vector, indicating the presence
or absence of these patterns, or counts of the frequency of
their occurrence. For example, Bag-of-Words (BoW) is a
common feature extraction approach, adapted from
natural language processing, where each possible type of AST
node becomes a binary feature. Figure 1 shows how BoW
features transform a piece of code into an input vector by
indicating the presence or absence of each feature in the
programming code AST. The BoW approach has been used in
various code classi cation tasks, such as to predict students'
success in completing a program, or to summarize functions
of code snippets [
        <xref ref-type="bibr" rid="ref11 ref18 ref6">6, 18, 11</xref>
        ]. For example, Azcona et al. used
BoW to represent students' code, and found that after using
BoW feature extraction to convert program code into
vectors, a simple Naive Bayes model predicted correctness of
short pieces of Python code submissions with 59.4%
accuracy [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ]. This suggests that even a relatively simple feature
extraction approach, such as BoW, extracts useful
information that can predict meaningful labels for student code with
some success.
      </p>
      <p>
        BoW features represent a single AST node, regardless of
its neighbors. However, meaningful programming patterns
usually include nodes that are structurally connected with
each other in the AST. For example, in Figure 2(b), the
piece of programming code completes a behavior that, when
a sprite 1 touches a bullet, the game ends. In order to
accomplish this behavior, the \Stop" block must be inside of
the \if" block - otherwise, the behavior would be di erent.
In order to extract structural information - such as the
requirement that the \stop" block be inside the \if" for the
behavior (shown in Figure 2), more complex features can be
extracted. For example, researchers have extracted features
corresponding to paths within the AST. For example, the
root path for a given node consists of the path from that
node to the root node [
        <xref ref-type="bibr" rid="ref21">21</xref>
        ], and this has been used for
datadriven hint generation [
        <xref ref-type="bibr" rid="ref21">21</xref>
        ]. The Code2Vec algorithm [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ]
also decomposes the AST into a collection of paths for use
in a deep neural network (discussed further in Section 2.3),
which was used to predict method names of code from Java
GitHub repositories (not student code). Others have used
n-Grams, which are n-length sequences of nodes extracted
from a attened representation of the AST. For example, a
vertical n-Gram is created by a depth- rst iteration of its
nodes, and a horizontal n-Gram is created by a
breadthrst iteration of all children in an AST subtree (shown in
Figure 2(c)). Akram et al. used code n-Grams as features
to predict the rubric-based grades of students' block-based
programs with a Gaussian Process model that achieved an
R-squared of 0.94, higher than the 0.88 achieved by the
baseline BoW approach [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ].
      </p>
      <p>
        Many of these structural features can be represented more
generally as a type of pq-Grams. A pq-Gram is a subtree
that includes a target node, along with its (p 1) ancestor
nodes, as well as q of its child nodes. For example,
Figure 2(b) shows how a pq-Gram can be extracted for the
target node \script (2)", with p ancestor nodes and q child
nodes. If a node has fewer than (p 1) ancestors or q
children, the pq-Gram includes this information by noting these
missing nodes as \null". pq-Grams were introduced as part
of a method to calculate di erences between tree-structured
data [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ], such as between a student's AST and a correct
solution's AST, in order to provide automated hints [
        <xref ref-type="bibr" rid="ref28">28</xref>
        ].
Using this notation, we can also consider the features
extracted by the BoW and root path approaches to be
pqGrams: p(1)q(0)-Grams and p(1)q(0)-Grams, respectively.
Most horizontal n-Grams can be represented as
p(0)q(n)Grams, and vertical n-Gram can be viewed as sub-arrays
of root paths, and can be represented as p(n)q(0)-Grams.
These AST structural feature extractions have been shown
to be e ective for representing and analyzing student code
[
        <xref ref-type="bibr" rid="ref27">27</xref>
        ]. Despite the variety of feature extraction approaches,
no work has compared the e cacy of these approaches,
especially pq-Gram feature extraction, which has not been used
previously for code classi cation tasks. This comparison is
important, since having features with too little expressivity
(e.g., BoW) will not capture important AST structural
information (causing the resulting model to under t), but having
features with too much expressivity (e.g., treating an entire
AST as a feature) will lead to overly-speci c features that
1A sprite in Scratch is similar to an object-oriented class.
In Scratch game design projects, an actor of the game is
usually represented by a sprite.
do not generalize to new, unseen instances, causing the
resulting model to over t. Empirical evaluation is needed to
nd an appropriate balance.
      </p>
      <p>
        Researchers have also explored clustering these simple
features to represent more complex structural relationships.
For example, Zhi et al. automatically clustered pq-Grams
into what they called \features" [
        <xref ref-type="bibr" rid="ref27">27</xref>
        ], which clusters pq-Grams
that performs a meaningful programming sub-goal.
However, this clustering method was not applied to feature
engineering for supervised classi cation. Mao et al. used
Recent Temporal Patterns (RTPs) to transform a highly
condensed feature set into a Multivariate State Sequence,
including information such as feature co-occurrence and
precedence. They used this feature engineering approach to
predict whether a student is unlikely to succeed in a given
task. Using RTPs, their model performs better than
simple feature extraction approaches, with an 18.8% increase in
classi cation accuracy, showing that creating feature
combinations o ers more information on students' programming
status [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ]. However, these existing approaches were only
applied in short programming tasks with speci c goals.
      </p>
    </sec>
    <sec id="sec-5">
      <title>2.3 Distributed Code Embeddings</title>
      <p>
        So far, we have discussed ways to represent code as
onehot or count vector. Neural Network models, in addition,
commonly contain an embedding layer that learns a
multidimensional representation on top of the feature extraction
approaches we have discussed. In addition, many embedded
approaches use sequential models, so that the embedding
is trained based on relative locations of each code element
[
        <xref ref-type="bibr" rid="ref11 ref18">11, 18</xref>
        ]. These embeddings can be learned in an
unsupervised manner [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ], similar to how word embeddings such as
Word2Vec [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ] are learned. They can also be learned in a
supervised manner through back-propagation during model
training [
        <xref ref-type="bibr" rid="ref18 ref4">4, 18</xref>
        ]. Depending on the architecture of Neural
Network, these embedded approaches also vary in ways to
represent code as vectors to feed into the models. For
example, Alon et al. used a leaf-to-leaf approach, connecting
the shortest path from each two leaf nodes [
        <xref ref-type="bibr" rid="ref3 ref4">4, 3</xref>
        ]. Iyer et al.
used embedded Bag-of-Words code representations to feed
into an LSTM model for conducting code retrieval tasks,
and have shown that their model can answer programming
questions by nding highly-relevant code snippets [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ].
Although distributed code representation is a powerful way
to interpret programs, they typically require a large amount
of training data - Code2Vec, for example, was trained on
more than 14M pieces of programming code. This
requirement on the number of training samples is not suitable for
our task.
      </p>
    </sec>
    <sec id="sec-6">
      <title>3. EXPERIMENT AND RESULTS</title>
      <p>Our goal in this study was to compare the e ectiveness of the
existing code feature extraction approaches outlined in
Section 2.2. To do so, we compared the e ectiveness of BoW,
n-Gram, and pq-Gram approaches in classifying game
behaviors.</p>
    </sec>
    <sec id="sec-7">
      <title>3.1 Dataset and Classification Task: Labeling</title>
    </sec>
    <sec id="sec-8">
      <title>Open-Ended Scratch Projects</title>
      <p>
        The student code used in our evaluation comes from the
Scratch community [
        <xref ref-type="bibr" rid="ref22">22</xref>
        ], an online, novice-friendly,
blockbased programming website, where users create and remix
interactive programming projects, such as games and
animations. We chose the Scratch repository because it
includes diverse, open-ended programs from learners around
the world, which are not constrained to a single assignment
or goal. This might be analogous to submissions to an
openended nal programming course project. On Feb 18th, 2020,
we scraped the 6247 most trendy projects, from the Game
genre in the Scratch community. Among them, we selected
the rst 457 projects based on the creation date. Among
the 457 projects, we excluded 44 projects that had over 50%
broken or unused code 2. Our dataset for the classi cation
task includes 413 projects, with an average of 1201 AST
nodes in each project 3.
      </p>
      <p>An important classi cation task for Scratch game projects is
to identify whether a given project includes a speci c game
behavior (i.e., a game mechanic), for example, whether the
player can jump (like in the classic Mario game).
Algorithm 1 shows one pseudocode example of how a platformer
jump can be implemented in Scratch 4. In this example, this
behavior is implemented by two threads, to ensure that the
actor jumps using gravity, and stops when landing on the top
of a platform. Based on our observations of students' code,
PlatformerJump is the most complex behavior in the ve
behaviors, usually including a large amount of code, spread
across di erent sprites and scripts. However, even with a
relatively less complex behavior (e.g., CollisionChangeVar ),
students can still implement the behavior in a wide variety
of ways.</p>
      <p>
        The ability to detect these game behaviors automatically
would allow researchers to better understand novice
programming behavior by pro ling the whole Scratch
repository, including millions of projects, to nd popular
combinations of behaviors (e.g., in [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]). It would also enable
researchers to instantly identify what type of game a
student is currently working on, in order to o er them highly
customized feedback or examples. This task also represents
a di cult challenge for code classi cation, since these game
behaviors are comprised of many code elements, which may
be dispersed throughout a student's code (e.g., the jump
behavior, shown in Algorithm 1), and which may be
implemented in diverse ways, creating a large and sparse
programming state-space [
        <xref ref-type="bibr" rid="ref27">27</xref>
        ]. These challenging properties are
shared by many other programming code classi cation tasks,
such as identifying misconceptions [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ] and predicting learner
performance [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ].
      </p>
      <p>
        In order to create meaningful categories of game behaviors,
the rst author investigated 13 student game design project
submissions, from an undergraduate programming course in
a large, public research university. After thoroughly
examining the submissions, the rst author decomposed each game
into a set of discrete game behaviors, under the criteria that
these behaviors are general enough to be reused in other
games. We identi ed 24 game behaviors. From these, we
selected ve that represented a diverse range of complexity
and frequency of use. One author developed a de nition for
each behavior label (see Table 1) and trained another author
on how to label projects, during which they jointly labeled
20 projects with the presence or absence of these ve game
behaviors. The two authors then individually labeled the
2Many Scratch game design projects come from remixing
and reusing existing projects [
        <xref ref-type="bibr" rid="ref24">24</xref>
        ], and novices do not always
have the ability to e ectively modify these projects [
        <xref ref-type="bibr" rid="ref19 ref20">19, 20</xref>
        ],
leading to abandoned or broken designs [
        <xref ref-type="bibr" rid="ref20">20</xref>
        ].
3In block-based languages such as Scratch, an AST node
generally has a corresponding block that matches the node.
So this also means that we have an average of about 1201
blocks in each project.
4Algorithm 1 is an example implementation of the
PlatformerJump behavior. The exact game logic is not
important, but it illustrates the complexity of Scratch game
behaviors.
1. These behaviors are implemented by a variety of blocks,
sometimes more than 30 blocks across di erent sprites
(i.e., similar to object-oriented classes) or code scripts
(i.e., threads).
2. Students implemented these behaviors in a variety of
di erent ways, using varying types and numbers of
blocks. This makes expert-authored rule-based static
analysis (e.g., [
        <xref ref-type="bibr" rid="ref25 ref7">25, 7</xref>
        ]) ine ective at detecting the
presence of these behaviors.
3. Although the selected game behaviors are typical within
certain game genres, the prevalence of individual
behaviors is often quite infrequent, as is shown by the
counts of projects in Table 1, with a range of 5.3% to
46%. This creates imbalanced datasets, which pose a
challenge in training classi ers that can lead a model
to be biased towards the majority class.
      </p>
      <p>While the above characteristics make our classi cation task
challenging, they are also common characteristics in many
important code classi cation tasks. Code indicating
misconceptions, low performers, or notable strategies may also
be complex, diverse, or rare. The results of our evaluation,
therefore, may be able to generalize to these tasks as well.</p>
    </sec>
    <sec id="sec-9">
      <title>3.2 Experiment Setup</title>
      <p>In order to understand how well these feature extraction
approaches capture meaningful information for predicting
the presence of complex game behaviors, we compared the
BoW, n-Gram, and pq-Gram approaches. Here we present
our experiment setup.</p>
      <sec id="sec-9-1">
        <title>KeyboardMove</title>
      </sec>
      <sec id="sec-9-2">
        <title>CollisionChangeVar</title>
      </sec>
      <sec id="sec-9-3">
        <title>PlatformerJump</title>
      </sec>
      <sec id="sec-9-4">
        <title>MoveWithMouse</title>
      </sec>
      <sec id="sec-9-5">
        <title>CollisionStopGame</title>
        <p>An actor 5 moves in the direction indicated by the
player on the keyboard
When one actor touches another, a variable
changes (e.g. score)
An actor can jump and then falls down with grav- 81/413
ity
An actor moves when the user moves or clicks the 49/413
mouse
The game ends when an actor touches another 25/413
4 - 6
20 - 50
2 - 4
3 - 4
Feature extraction. We rst extracted BoW, n-Gram,
and pq-Gram features from the training dataset. To reduce
the number of irrelevant features, we extracted features that
have more than 5% support (i.e., the percentage of projects
where this feature exists). When extracting n-Grams of a
speci c n value, we consider both horizontal and vertical
ones, as introduced in Section 2.2, for they each extract
different AST structural information. We extracted n-Grams
with n 2 f1; 2; : : : ; 10g, and also extracted pq-Grams, with
p 2 f1; 2; 3g, and q 2 f1; 2; 3; 4g. The exact subset of these
n-grams or pq-Grams used was determined by
hyperparameters, as discussed below. At each increase of n in n-Grams,
we kept features that were extracted by smaller ns, but
removed duplicated smaller features based on the rule that
when two features always co-appear, and one is a subset of
another. We extracted pq-Grams using the same approach.
Training and evaluation. To train our model, we used a
Support Vector Machine (SVM) model with a linear kernel,
and used the regularization parameter as a
hyperparameter, with values in f0.01, 0.1, 1, 10, 100g. We employed
ve-fold cross-validation to evaluate our feature set. Within
each round of cross-validation, we used 41 of the training set
as the validation set to tune the hyper-parameters. When
we extracted features for n-Grams, we used the maximum
n as a hyperparameter, with n 2 f1; 2; : : : ; 10g. Similarly,
we also used the maximum ps and qs as hyperparameters
when extracting pq-Grams, with p 2 f1; 2; 3g, and q 2
f1; 2; 3; 4g. The values of the hyperparameters were
determined by their F1 scores on the validation set at each
round of cross-validation. Since many of our target labels
are highly imbalanced, the accuracy score o ers little
information on how well our model performs in predicting target
labels. We therefore use F1 scores to tune hyperparameters.</p>
      </sec>
    </sec>
    <sec id="sec-10">
      <title>3.3 Results</title>
      <p>Figure 3 shows how the feature sets perform across the ve
target game behaviors. We present the F1 scores of
prediction in each target behavior, with its precision (P) and
recall (R), shown in the brackets. In this rst experiment, our
results show that all classi ers perform similarly. One
exception to this is the PlatformerJump behavior, where BoW
(F1 = 0.58) does notably worse than n-Gram (F1 = 0.72),
or pq Gram (F1 = 0.68). We note that PlatformerJump is
easily the most complex behavior (shown in Algorithm 1),
which can be completed in many ways. This suggests that
there may be some advantage to more expressive feature
representations for identifying more complex program
prop</p>
      <p>Figure 3 also shows that F1 scores of all approaches
decrease as the prevalence of positive samples decreases (i.e.,
with more class imbalance). For example, on the y-axis of
Figure 3, we have marked each label with the prevalence of
its positive samples. As the prevalence of features decreases
from 197/413 (48%) in KeyboardMove to 25/413 (6%) in
CollisionStopGame, The feature extraction methods
perform increasingly worse. For more common behaviors such
as KeyboardMove, all approaches had su cient data to
accurately identify the behavior (F1 = 0.83-0.87), and even
simpler approaches such as BoW were expressive enough
to nd discriminating features, e.g., the \WhenKeyPressed"
block. However, when less positive training data are
available, none of the approaches perform well, suggesting the
possibility that the models are over tting.</p>
      <p>We therefore investigated the training F1 scores for each
model, shown in Figure 4. The results con rm that the
models are likely over tting with all three feature
extraction approaches, especially when the prevalence of positive
samples is small, such as in MoveWithMouse and
CollisionStopGame. This is unsurprising, given that all feature
extraction approaches produced hundreds (BoW) to over a
thousand (pq-Gram) features, and the training data never
exceeded 200 positive instances. Because even our simplest
feature extraction approach (BoW) was clearly over tting,
it is unclear whether more expressive feature representations
(n-Grams, pq-Grams) would hold an advantage under other
circumstances (e.g., more training data, with additional
feature selection6).</p>
    </sec>
    <sec id="sec-11">
      <title>4. DISCUSSION AND CONCLUSION</title>
      <p>
        We have compared di erent approaches for extracting
predictive features from the program code. We have also
presented the rst step towards automated classi cation of
openended block-based program behaviors, going beyond existing
rule-based analysis [
        <xref ref-type="bibr" rid="ref1 ref17">17, 1</xref>
        ]. Here we discuss the insights we
have derived from our results, and we discuss our future task
to improve our current classi cation approach and address
underlying challenges.
      </p>
      <p>Overall, we found no evidence that di erent feature
extraction approaches led to better or worse code classi cation
results. However, we note that in all cases, this was due to
over tting, due to a large number of features and a relatively
small amount of training data. This is supported by the fact
that all approaches did worse for behaviors with more class
imbalance. We are therefore unable to conclude whether
our results would generalize to a larger dataset, where the
richer feature of n-Gram or pq-Grams may be better
leveraged. However, we also note that in the domain of computer
science education, courses often only have a relatively small
number of students, causing the size of training data to be
also small. This means that in programming code feature
extraction approaches, we should explore ways to reduce the
number of features, through feature selection and
dimen6We did attempt basic feature selection approaches,
reducing the number of features to less than 100, but this
either failed to improve or worsened performance, suggesting
a need for future work exploring how to address the over
tting.
sionality reduction. It may also be helpful to develop ways
to more e ciently label data - which serves as the building
block for many intelligent algorithms.</p>
      <p>
        Our classi cation performance also varied considerably across
tasks, in some cases quite well (e.g. KeyboardMove), but in
others quite poorly (e.g. CollisionStopGame). To
understand why, we note that, unlike prior work, our code
classication task used a relatively small dataset (n = 413),
consisting of large programming projects. The Scratch projects
we analyzed had an average of 1201 AST nodes. By
contrast, code classi cation tasks in prior work have
generally been applied to smaller code input, with a much larger
amount of training data. For example, Code2Vec is
implemented in a code classi cation task for predicting method
names from programs with an average length of 7 lines [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ],
although the training sample is of size 14M. Iyer et al.
implemented an LSTM-based code summarization model, but
only on programming code with an average of 38 tokens
(i.e., the number of elements in the program). Speci cally,
in the computer science education domain, many
programming code analysis approaches are evaluated on short
programming tasks, such as drawing a geometric shape using
nested loops [
        <xref ref-type="bibr" rid="ref13 ref26 ref27">26, 13, 27</xref>
        ], or implementing a short algorithm
(e.g., a bubble sort algorithm [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]). Mou et al. evaluated
their LSTM-based code classi er's performance on
classifying function methods, and concluded that longer programs
(i.e., with longer length of code) had relatively lower classi
cation performance compared to shorter programming tasks
[
        <xref ref-type="bibr" rid="ref18">18</xref>
        ]. This may explain why our results were not as strong,
and more prone to over tting.
      </p>
      <p>In conclusion, in this work, we compared features with
different levels of expressivity (i.e., Bag-of-Words, n-Grams,
and pq-Grams), in a challenging task to classify meaningful
game design behaviors in open-ended Scratch projects. Our
results show that our model may be over tting with all three
di erent feature extraction approaches, and that we need to
explore ways to reduce feature dimensions and increase data
size to improve performance.</p>
    </sec>
    <sec id="sec-12">
      <title>5. ACKNOWLEDGEMENTS</title>
      <p>This material is based upon work supported by the National
Science Foundation under Grant No. 1917885.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>E.</given-names>
            <surname>Aivaloglou</surname>
          </string-name>
          and
          <string-name>
            <given-names>F.</given-names>
            <surname>Hermans</surname>
          </string-name>
          .
          <article-title>How kids code and how we know: An exploratory study on the scratch repository</article-title>
          .
          <source>In Proceedings of the 2016 ACM Conference on International Computing Education Research</source>
          , pages
          <volume>53</volume>
          {
          <fpage>61</fpage>
          ,
          <year>2016</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>B.</given-names>
            <surname>Akram</surname>
          </string-name>
          et al.
          <article-title>Assessment of students' computer science focal knowledge, skills, and abilities in game-based learning environments</article-title>
          .
          <year>2019</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>U.</given-names>
            <surname>Alon</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Zilberstein</surname>
          </string-name>
          ,
          <string-name>
            <surname>O. Levy</surname>
          </string-name>
          , and
          <string-name>
            <given-names>E.</given-names>
            <surname>Yahav</surname>
          </string-name>
          .
          <article-title>A general path-based representation for predicting program properties</article-title>
          .
          <source>In ACM SIGPLAN Notices</source>
          , volume
          <volume>53</volume>
          , pages
          <fpage>404</fpage>
          {
          <fpage>419</fpage>
          . ACM,
          <year>2018</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>U.</given-names>
            <surname>Alon</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Zilberstein</surname>
          </string-name>
          ,
          <string-name>
            <surname>O. Levy</surname>
          </string-name>
          , and
          <string-name>
            <surname>E. Yahav.</surname>
          </string-name>
          <article-title>code2vec: Learning distributed representations of code</article-title>
          .
          <source>Proceedings of the ACM on Programming Languages</source>
          ,
          <volume>3</volume>
          (POPL):
          <volume>40</volume>
          ,
          <year>2019</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>N.</given-names>
            <surname>Augsten</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. H.</given-names>
            <surname>Bo</surname>
          </string-name>
          <article-title>hlen, and</article-title>
          <string-name>
            <given-names>J.</given-names>
            <surname>Gamper</surname>
          </string-name>
          .
          <article-title>Approximate matching of hierarchical data using pq-grams</article-title>
          .
          <source>In VLDB</source>
          , volume
          <volume>5</volume>
          , pages
          <fpage>301</fpage>
          {
          <fpage>312</fpage>
          ,
          <year>2005</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>D.</given-names>
            <surname>Azcona</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Arora</surname>
          </string-name>
          ,
          <string-name>
            <given-names>I.-H.</given-names>
            <surname>Hsiao</surname>
          </string-name>
          ,
          <article-title>and</article-title>
          <string-name>
            <given-names>A.</given-names>
            <surname>Smeaton</surname>
          </string-name>
          . user2code2vec:
          <article-title>Embeddings for pro ling students based on distributional representations of source code</article-title>
          .
          <source>In Proceedings of the 9th International Conference on Learning Analytics &amp; Knowledge</source>
          , pages
          <volume>86</volume>
          {
          <fpage>95</fpage>
          . ACM,
          <year>2019</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>M.</given-names>
            <surname>Ball</surname>
          </string-name>
          . Lambda:
          <article-title>An autograder for snap</article-title>
          .
          <source>Technical report, Technical Report</source>
          . Electrical Engineering and Computer Sciences University of California at Berkeley,
          <year>2018</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>N.</given-names>
            <surname>Diana</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Eagle</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Stamper</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Grover</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Bienkowski</surname>
          </string-name>
          , and
          <string-name>
            <given-names>S.</given-names>
            <surname>Basu</surname>
          </string-name>
          .
          <article-title>Data-driven generation of rubric criteria from an educational programming environment</article-title>
          .
          <source>In Proceedings of the 8th International Conference on Learning Analytics and Knowledge</source>
          , LAK '
          <volume>18</volume>
          , page
          <volume>16</volume>
          {
          <fpage>20</fpage>
          , New York, NY, USA,
          <year>2018</year>
          .
          <article-title>Association for Computing Machinery</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>M.</given-names>
            <surname>Elmadani</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Mathews</surname>
          </string-name>
          ,
          <article-title>and</article-title>
          <string-name>
            <given-names>A.</given-names>
            <surname>Mitrovic</surname>
          </string-name>
          .
          <article-title>Data-driven misconception discovery in constraint-based intelligent tutoring systems</article-title>
          .
          <year>2012</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>L.</given-names>
            <surname>Gusukuma</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A. C.</given-names>
            <surname>Bart</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Kafura</surname>
          </string-name>
          , and
          <string-name>
            <given-names>J.</given-names>
            <surname>Ernst</surname>
          </string-name>
          .
          <article-title>Misconception-driven feedback: Results from an experimental study</article-title>
          .
          <source>In Proceedings of the 2018 ACM Conference on International Computing Education Research</source>
          , pages
          <volume>160</volume>
          {
          <fpage>168</fpage>
          . ACM,
          <year>2018</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>S.</given-names>
            <surname>Iyer</surname>
          </string-name>
          ,
          <string-name>
            <surname>I. Konstas</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Cheung</surname>
          </string-name>
          , and
          <string-name>
            <given-names>L.</given-names>
            <surname>Zettlemoyer</surname>
          </string-name>
          .
          <article-title>Summarizing source code using a neural attention model</article-title>
          .
          <source>In Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)</source>
          , pages
          <year>2073</year>
          {
          <year>2083</year>
          ,
          <year>2016</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>C. M.</given-names>
            <surname>Lewis</surname>
          </string-name>
          .
          <article-title>The importance of students' attention to program state: a case study of debugging behavior</article-title>
          .
          <source>In Proceedings of the ninth annual international conference on International computing education research</source>
          , pages
          <volume>127</volume>
          {
          <fpage>134</fpage>
          ,
          <year>2012</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <given-names>Y.</given-names>
            <surname>Mao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Zhi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Khoshnevisan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T. W.</given-names>
            <surname>Price</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Barnes</surname>
          </string-name>
          , and
          <string-name>
            <given-names>M.</given-names>
            <surname>Chi</surname>
          </string-name>
          .
          <article-title>One minute is enough: Early prediction of student success and event-level di culty during novice programming tasks</article-title>
          .
          <source>In EDM</source>
          ,
          <year>2019</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <given-names>S.</given-names>
            <surname>Marwan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T. W.</given-names>
            <surname>Price</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Chi</surname>
          </string-name>
          , and
          <string-name>
            <given-names>T.</given-names>
            <surname>Barnes</surname>
          </string-name>
          .
          <article-title>Immediate data-driven positive feedback increases engagement on programming homework for novices</article-title>
          .
          <year>2020</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <given-names>O.</given-names>
            <surname>Meerbaum-Salant</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Armoni</surname>
          </string-name>
          , and
          <string-name>
            <given-names>M.</given-names>
            <surname>Ben-Ari</surname>
          </string-name>
          .
          <article-title>Habits of programming in scratch</article-title>
          .
          <source>ITiCSE'11 - Proceedings of the 16th Annual Conference on Innovation and Technology in Computer Science</source>
          , pages
          <volume>168</volume>
          {
          <fpage>172</fpage>
          ,
          <year>2011</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [16]
          <string-name>
            <given-names>T.</given-names>
            <surname>Mikolov</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Chen</surname>
          </string-name>
          , G. Corrado,
          <string-name>
            <given-names>J.</given-names>
            <surname>Dean</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Sutskever</surname>
          </string-name>
          , and
          <string-name>
            <surname>G. Zweig.</surname>
          </string-name>
          word2vec. URL https://code. google. com/p/word2vec,
          <year>2013</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [17]
          <string-name>
            <given-names>J.</given-names>
            <surname>Moreno-Leon</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Robles</surname>
          </string-name>
          , and
          <string-name>
            <given-names>M.</given-names>
            <surname>Roman-Gonzalez</surname>
          </string-name>
          .
          <article-title>Dr. scratch: Automatic analysis of scratch projects to assess and foster computational thinking</article-title>
          .
          <source>RED. Revista de Educacion a Distancia</source>
          , (
          <volume>46</volume>
          ):
          <volume>1</volume>
          {
          <fpage>23</fpage>
          ,
          <year>2015</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          [18]
          <string-name>
            <given-names>L.</given-names>
            <surname>Mou</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Zhang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Wang</surname>
          </string-name>
          , and
          <string-name>
            <given-names>Z.</given-names>
            <surname>Jin</surname>
          </string-name>
          .
          <article-title>Convolutional neural networks over tree structures for programming language processing</article-title>
          .
          <source>In Thirtieth AAAI Conference on Arti cial Intelligence</source>
          ,
          <year>2016</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          [19]
          <string-name>
            <given-names>G. L.</given-names>
            <surname>Nelson</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Xie</surname>
          </string-name>
          ,
          <article-title>and</article-title>
          <string-name>
            <given-names>A. J.</given-names>
            <surname>Ko</surname>
          </string-name>
          .
          <article-title>Comprehension rst: evaluating a novel pedagogy and tutoring system for program tracing in cs1</article-title>
          .
          <source>In Proceedings of the 2017 ACM Conference on International Computing Education Research</source>
          , pages
          <volume>2</volume>
          {
          <fpage>11</fpage>
          . ACM,
          <year>2017</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          [20]
          <string-name>
            <given-names>S.</given-names>
            <surname>Papert</surname>
          </string-name>
          . Children, computers and powerful ideas,
          <year>1990</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          [21]
          <string-name>
            <given-names>T. W.</given-names>
            <surname>Price</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Dong</surname>
          </string-name>
          , and
          <string-name>
            <given-names>T.</given-names>
            <surname>Barnes</surname>
          </string-name>
          .
          <article-title>Generating data-driven hints for open-ended programming</article-title>
          .
          <source>International Educational Data Mining Society</source>
          ,
          <year>2016</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          [22]
          <string-name>
            <given-names>M.</given-names>
            <surname>Resnick</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Maloney</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Monroy-Hernandez</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Rusk</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Eastmond</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Brennan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Millner</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Rosenbaum</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Silver</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Silverman</surname>
          </string-name>
          , et al.
          <article-title>Scratch: programming for all</article-title>
          .
          <source>Communications of the ACM</source>
          ,
          <volume>52</volume>
          (
          <issue>11</issue>
          ):
          <volume>60</volume>
          {
          <fpage>67</fpage>
          ,
          <year>2009</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          [23]
          <string-name>
            <given-names>T.</given-names>
            <surname>Sirkia</surname>
          </string-name>
           and
          <string-name>
            <given-names>J.</given-names>
            <surname>Sorva</surname>
          </string-name>
          .
          <article-title>Exploring programming misconceptions: an analysis of student mistakes in visual program simulation exercises</article-title>
          .
          <source>In Proceedings of the 12th Koli Calling International Conference on Computing Education Research</source>
          , pages
          <volume>19</volume>
          {
          <fpage>28</fpage>
          ,
          <year>2012</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          [24]
          <string-name>
            <given-names>V.</given-names>
            <surname>Vasudevan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Kafai</surname>
          </string-name>
          , and
          <string-name>
            <given-names>L.</given-names>
            <surname>Yang</surname>
          </string-name>
          .
          <article-title>Make, wear, play: remix designs of wearable controllers for scratch games by middle school youth</article-title>
          .
          <source>In Proceedings of the 14th international conference on interaction design and children</source>
          , pages
          <volume>339</volume>
          {
          <fpage>342</fpage>
          ,
          <year>2015</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref25">
        <mixed-citation>
          [25]
          <string-name>
            <given-names>W.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Zhi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Milliken</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Lytle</surname>
          </string-name>
          , and
          <string-name>
            <given-names>T.</given-names>
            <surname>Price</surname>
          </string-name>
          . Crescendo:
          <article-title>Engaging students to self-paced programming practices</article-title>
          . In To be published in
          <source>the 51st ACM Technical Symposium on Computer Science Education (SIGCSE '20)</source>
          ,
          <year>2020</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref26">
        <mixed-citation>
          [26]
          <string-name>
            <surname>M. Wu</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          <string-name>
            <surname>Mosse</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          <string-name>
            <surname>Goodman</surname>
            , and
            <given-names>C.</given-names>
          </string-name>
          <string-name>
            <surname>Piech</surname>
          </string-name>
          .
          <article-title>Zero shot learning for code education: Rubric sampling with deep learning inference</article-title>
          .
          <source>In Proceedings of the AAAI Conference on Arti cial Intelligence</source>
          , volume
          <volume>33</volume>
          , pages
          <fpage>782</fpage>
          {
          <fpage>790</fpage>
          ,
          <year>2019</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref27">
        <mixed-citation>
          [27]
          <string-name>
            <given-names>R.</given-names>
            <surname>Zhi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T. W.</given-names>
            <surname>Price</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Lytle</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Dong</surname>
          </string-name>
          , and
          <string-name>
            <given-names>T.</given-names>
            <surname>Barnes</surname>
          </string-name>
          .
          <article-title>Reducing the state space of programming problems through data-driven feature detection</article-title>
          .
          <source>In Educational Data Mining in Computer Science Education (CSEDM) Workshop@ EDM</source>
          ,
          <year>2018</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref28">
        <mixed-citation>
          [28]
          <string-name>
            <given-names>K.</given-names>
            <surname>Zimmerman</surname>
          </string-name>
          and
          <string-name>
            <given-names>C. R.</given-names>
            <surname>Rupakheti</surname>
          </string-name>
          .
          <article-title>An automated framework for recommending program elements to novices (n)</article-title>
          .
          <source>In 2015 30th IEEE/ACM International Conference on Automated Software Engineering (ASE)</source>
          , pages
          <fpage>283</fpage>
          {
          <fpage>288</fpage>
          . IEEE,
          <year>2015</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>