<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Predicting student performance over time. A case study for a blended-learning engineering course</article-title>
      </title-group>
      <contrib-group>
        <aff id="aff0">
          <label>0</label>
          <institution>Universitat Autonoma de Barcelona</institution>
          ,
          <addr-line>Edifici D, Campus UAB, 08193 Bellaterra</addr-line>
          ,
          <country country="ES">Spain</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Universitat Oberta de Catalunya, Rambla del Poblenou</institution>
          ,
          <addr-line>156, 08018 Barcelona</addr-line>
          ,
          <country country="ES">Spain</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2019</year>
      </pub-date>
      <fpage>0000</fpage>
      <lpage>0002</lpage>
      <abstract>
        <p>In recent years, different studies have focused in analyzing whether it is possible to explain and predict performance of students based on information we know about them, and in particular, on that obtained from Learning Management Systems (LMSs). A review of existing literature shows we can still raise no conclusion, and in particular when dealing with face to face (F2F) studies. In this article, we analyze the performance of a first-year engineering course, offered in a higher education institution (a public university). The course under analysis lasts for 12 weeks and is offered with flipped classroom methodology. Activities that students should follow out of class are scheduled in advance, and communicated to students during the learning period. In addition, there has been a previous effort to align learning activities and learning outcomes. The goal is to determine if prediction models fed with data gathered during the learning process can provide an accurate estimator of students at risk. This risk evaluation will be done considering as core data those reflecting activity, being of particular relevance, traces stored in LMS as part of the learning process. Our study demonstrates performance can be estimated based on this data, with increasing accuracy over time. Activity performed by the student is linked to academic result, and this relation is verified even when not taking into account any graded results obtained during the learning process.</p>
      </abstract>
      <kwd-group>
        <kwd>learning analytics</kwd>
        <kwd>performance prediction</kwd>
        <kwd>student modelling</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>Introduction</title>
      <p>
        General adoption of LMSs has motivated a growing interest for Educational Data
Mining and Learning Analytics in general and academic performance prediction in
particular. Prediction is one of the most explored areas in both fields, increasing its relative
weight over time. Research papers related to prediction were around 30% in 1995-2005
[
        <xref ref-type="bibr" rid="ref1">1</xref>
        ], while this number increases to over 40% in more recent studies [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ].
      </p>
      <p>This rising interest for prediction is more than justified, due to its potential impact
at all educational levels. At the macro level, it can help institutional managers to
implement educational policies addressed to reduce failure and increase overall quality.
Predicting and understanding the reasons behind prediction results can be a powerful tool
to increase global academic performance.</p>
      <p>While this macro level is undoubtedly of great interest, we will focus on the teaching
impact. We consider prediction a potential lever with different potential applications.
If we can provide early prediction, we can redress risk behaviors. This change in
behavior can be implemented in a passive way – i.e. just informing students of the
potential risk – or in a more active fashion – i.e. implementing specific teaching measures -.
If we consider late predictions, they can help to understand causes of failure and
redesign pedagogical approaches for forthcoming editions of the same course.</p>
      <p>Our study will be carried out in a first course engineering subject in a public
oncampus university. Being an on-campus university is relevant, as attending on-campus
classes mixes with on-line activities. The subject under analysis has been selected due
to its blended-learning methodology, where class attendance mixes with on-line
lectures and activities. We aim to estimate probability of success, focusing on behavior of
the student regarding the different activities. In particular, considering behavior of the
students both in on-campus and on-line activities.</p>
      <p>The goal of this research is to evaluate to what extent risk of failure in a flipped face
to face (F2F) course can be predicted based on the analysis of the student’s behavior.
The underlying hypothesis is that student performance (in terms of pass-fail) can be
predicted, in good measure, by his/her activity during the course even not taking into
account grades gathered in evaluative activities performed during the course. In order
to validate the hypothesis we raise two research questions (RQ):
 RQ1: Can student’s final performance in terms of “pass-fail” in a course be
anticipated by analyzing his/her behavior regarding the fulfillment of the programmed
learning activities without taking into account grades obtained in evaluative
assessments?
 RQ2: How is this prediction influenced when limiting data to those gathered in the
early stages of the course?</p>
      <p>Affirmative answer to the first question would suggest activity performance is linked
to academic result. This would open a path to better understand the learning process for
this particular subject and to suggest potential improvements in pedagogical design.
Regarding the second question, early prediction can be useful to redress individual
student behavior and reduce overall failure. Answers to both questions would help to
improve teaching quality.
2</p>
    </sec>
    <sec id="sec-2">
      <title>Theoretical framework</title>
      <p>
        Learning analytics (LA), defined formally by Siemens [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ] and Ferguson [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ], cover a
full set of studies dealing with the extraction of meaningful information from data
retrieved in the learning process. Ferguson [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ] focuses on “measurement, collection,
analysis and reporting of data”. Closely related, but with different goals, we find
Educational Data Mining (EDM), more commonly accepted definition by Baker and Yacef
[
        <xref ref-type="bibr" rid="ref2">2</xref>
        ] which tends to focus on techniques.
      </p>
      <p>A core concept in both fields is data, being of particular interest those gathered
during the learning process. While the process of data gathering related to the learning
process is intrinsic in online universities, most on-campus universities were not born
with this idea in mind. In recent years, and with the general adoption of LMS systems,
there has been a shift towards data gathering and analysis.</p>
      <p>
        The extraction of information from LMS data is a topic on its own. Agudo-Peregrina,
Iglesias-Pradas, Conde-González, and Hernández-García [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ] suggest to begin by
classifying information around two main axes: interactions based on agent and interactions
based on frequency of use. Each of these axes will include a number of specific
variables, which depend on the study and the expected output [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ].
      </p>
      <p>
        Conijn, Snijders, Kleingeld, and Matzat [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ] compiled pre-existing work and
summarized variables considered of potential interest in the literature. The use of variables
which are linked to the learning process is common. In particular, we can find number
of resources viewed, quizzes started, sessions or total clicks. It is likely to remark that
the different studies analyzed provide different impact and influence of variables
depending on the course under consideration.
      </p>
      <p>
        According to [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ] there is also no consensus on the influence of a given variable. This
fact is also reflected in [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ], concluding that in order to get better results “we need to get
a better insight into what the LMS data represents”.
      </p>
      <p>
        Both compilations ( [
        <xref ref-type="bibr" rid="ref5 ref6">5, 6</xref>
        ]) show the huge number of different variables that are
present in different studies. This is also present in LMS related web sites who focus on
information gathered from LMS systems [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ]. Whichever the initial variable set is, a
selection process will be mandatory, in particular if the number of variables is high in
relation to the number of samples.
      </p>
      <p>
        Before entering the prediction process itself, the nature of the problem must be
focused. Failure analysis can be approached as a regression problem (i.e. estimating final
graded performance of student) or as a classification problem (i.e. analyzing whether
the student will pass or fail). The classification approach is common in the literature,
with studies suggesting better performance and potential detection of meaningful
patterns [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ].
      </p>
      <p>
        Once variables are selected, and considering we face a classification problem,
different studies use different methods for prediction. To have some examples, the range
goes from simple decision trees [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ], to behavioral clustering [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ]. Different research
compilations regarding techniques ( [
        <xref ref-type="bibr" rid="ref11 ref12">11, 12</xref>
        ]) show there is no universal method that
provides suitable results for all situations. In our case, and considering our study is not
focused on the techniques themselves, we will evaluate results with most common
methods, without being tied to a particular one, putting the focus on the interpretation
of results.
      </p>
      <p>
        Whichever the technique, evaluating model goodness is the next required step. Due
to the classification nature of the problem, area under curve (AUC) is a potential
indicator of model goodness. AUC “is a one-number measure of a model’s discrimination
performance, i.e., the extent to which a model successfully separates the positive and
the negative observations” [
        <xref ref-type="bibr" rid="ref13">13</xref>
        ]. AUC can also help in cases where we are not dealing
with large datasets [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ].
      </p>
      <p>
        Some studies raise concerns about use of AUC as performance indicator [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ]. In
particular, AUC makes no difference between errors, although when thinking of
failure/success classification this can be of potential interest. Depending on the potential
application of the prediction, false negative and false positive errors could have
different impact, and this information is not contained in AUC.
      </p>
      <p>
        LA and EDM reseearch does not have unique indicators for evaluating model
performance. Due to this fact, comparison of published research works is not
straightforward ( [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ]). Different studies apply different metrics regarding model
validation. Bowers, Sprott and Taff ( [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ]) suggest a framework for comparison. This
framework considers global accuracy, but at the same time, includes a graphical view
with information related to sensitivity and specificity.
      </p>
      <p>Considering the potential drawback of AUC as a unique indicator, and also the need
to compare to previous research, keeping accuracy, sensitivity and specificity besides
AUC can help to effectively compare with pre-existing works. AUC can be a general
indicator of the overall quality of the classification, while rest of parameters provide
additional information and allow to compare with previous research.</p>
      <p>This same compilation includes articles with different time scenarios. Best
performing models are fed with long-time data (math achievement trajectories from grades
712, non-cumulative GPA (Grade Point Average) from grades 9-12, and student
engagement trajectories from grades 8-12). All of them include evaluative data as input
variables to the model.</p>
      <p>The impact of graded activities gathered during the course is common in the
literature. At the same time, some studies compiled look for pre-existing variables that could
be also of potential interest to performance models. This variables can be external to
the learning process, and can include social or economic aspects.</p>
      <p>
        As a final consideration, [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ] also concludes that “the predictive utility of many
variables is dependent upon course site design and pedagogical goals”. It seems clear that
while there has been a technical approach to data mining, there has not been such an
evolution on seeking the interpretation or generation of relevant information from the
data stored in information systems in general and LMS in particular.
3
      </p>
    </sec>
    <sec id="sec-3">
      <title>Methodology</title>
      <p>We face a classification task, without being restricted to a particular data mining
technique. Models will be fed with activity data. This activity data will also include a time
scope in order to evaluate three kind of models according to the time when the data are
gathered: early (4 weeks), medium (8 weeks) and late (12 weeks).
3.1</p>
      <sec id="sec-3-1">
        <title>Suitable techniques</title>
        <p>
          While it is not the goal of the paper to discuss about data mining algorithms, we did not
want to restrict our study to a particular technique. We selected those present in
literature compilation regarding student performance prediction [
          <xref ref-type="bibr" rid="ref12">12</xref>
          ]. Selected techniques
were naive Bayes, neural networks, decision trees (including both gradient boosted
trees – GBT- and random forest –RF-) and Support Vector Machines (SVM). For all
of these techniques, we will keep classification error, sensitivity and specificity as
parameters to compare with existing literature, and AUC as an additional check.
        </p>
        <p>
          Regarding models, we define the true positive class as those students likely to fail
who actually fail. Those students marked as failing who really pass will be considered
False Positives (i.e. Type I errors), while students marked as passing who really fail
will be the false negative class (i.e. Type II errors). This approach will permit direct
comparison with results in [
          <xref ref-type="bibr" rid="ref16">16</xref>
          ].
3.2
        </p>
      </sec>
      <sec id="sec-3-2">
        <title>Variables</title>
        <p>
          Due to the different and high number of variables present in the literature, and to the
fact that they normally include graded activities, we decided to begin from scratch, but
keeping in mind lessons learnt from previous compilations ( [
          <xref ref-type="bibr" rid="ref5 ref6">5, 6</xref>
          ]). In particular, we
look for meaningful variables linked to the learning process.
        </p>
        <p>We reviewed our course design, and looked for knowledge derived from our
teaching experience. We considered three core concepts as fundamental to explain academic
results: class attendance, continuous working and flipped behavior. Not all this piece
of information was kept as structured data before performing this study.</p>
        <p>In particular, we had no information regarding class attendance. Class attendance is
not mandatory, and there is no specific control, as students can decide – without
academic impact – whether they attend classes or not. The introduction of the flipped
classroom methodology made us think about potential non-intrusive techniques to estimate
it.</p>
        <p>This estimation was performed through the use of a learning engagement tool
(Socrative). This tool was introduced as part of the course design to help the detection
of areas that need reinforcement. Questions are performed to students during class to
evaluate contents that are clear and those that need reinforcement. Questions have no
impact in grades. They help instructors to focus on specific areas depending on the
answers students provide. Information in the logs allow us to provide an estimation of
student attendance to class. We summarize attendance in each of the periods (early,
medium and late attendance). It is an estimation – and not an exact value – as the tool
is not used in every class.</p>
        <p>Continuous working is complex to evaluate and measure. In order to keep simple
and at the same time meaningful variables, we opted to keep the volume of information
collected in the LMS log file per user and week. We kept one variable for each week
of the course that reflects the amount of log lines the LMS. For each of the periods
(early, medium, late) we consolidate work in the whole period into a single variable.</p>
        <p>We raised concerns regarding activities performed offline. To capture this offline
activity, we reinforced the need to use the LMS as part of the pedagogical design. Users
can obviously work offline, but video lessons and problem solving require access to the
platform. In this way, we can assume users with greater activity levels are those with
greater number of log entries.</p>
        <p>The above group of variables can reflect continuous work but does not directly link
to flipped behavior. The flipped methodology would make advisable to review certain
topics before attending class. The list of required activities and due dates is part of the
course design. These activities and dates are communicated to students in advance on
a per-week basis. So, we included a new set of variables, reflecting for each week the
amount of work that was assigned to that week and was effectively performed on time.</p>
        <p>The need to get this information requires that all instructors share a common set of
activities instructed to students. Each of the activities will have a unique indicator. Once
this indicator is located in the log files for a given user, date can be compared to due
date for that activity. This approach makes it possible to compute on-time performance
of activities for every student, provided that all instructors set the same dates for activity
performance.</p>
        <p>So far, we have variables reflecting class attendance for each of the periods. We also
have a per-week estimation of workload performed based on the log data, and finally
the amount of work assigned to each of the weeks performed on time. For this last two
datasets we also keep the total work performed in the period.</p>
        <p>
          Due to the high number of variables, a forward selection process will be necessary.
This is done to keep the recommended ratio between number of variables and number
of samples avoiding potential overfit [
          <xref ref-type="bibr" rid="ref17">17</xref>
          ]. This operation will be done for each of the
time scopes (early, medium, late) under analysis.
3.3
        </p>
      </sec>
      <sec id="sec-3-3">
        <title>A word on pre-existing data</title>
        <p>
          Different studies have analyzed pre-existing variables which can condition students’
outcomes [
          <xref ref-type="bibr" rid="ref16">16</xref>
          ]. We discarded variables without direct link to the learning process. In
our case, and after discussion, we kept the grade you get when entering the university,
and the fact of being new or repeating student. Variables such as city of residence or
family income were not considered due to our focus on activity.
        </p>
        <p>Regarding the grade the student enters the university with, we thought that under the
same conditions, students with higher entering grades should be more likely to pass.
Regarding the fact of being new or repeating student, our experience shows that
repeating students show different behavior than those being enrolled in the subject for the
first time.
3.4</p>
      </sec>
      <sec id="sec-3-4">
        <title>Summary</title>
        <p>
          We have compared results with the compilation in [
          <xref ref-type="bibr" rid="ref16">16</xref>
          ], where impact of different
variables in published models is shown. For clarity, we have added just our Bayesian
models results, as they offer best performance in terms of AUC for early and late
activity, and for mid-term is close to maximum. Results are shown in figure 1, marked as
“Early prediction”, “Mid-term prediction” and “Late prediction”. The compilation
includes 110 indicators (depicted as numbers in Figure 1) from 36 different prediction
works:
We also analyzed which were the variables that show relevant in the different time
scenarios for the Bayesian method analyzed. Table 3 summarizes the results for the
different time scopes:
        </p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>Discussion</title>
      <p>Before answering RQ1 and RQ2, we need to validate to what extent the models outlined
in this paper are of potential interest. This question can be answered analyzing Figure
1 and results in Table 2.</p>
      <p>Figure 1 shows that final performance of student can be predicted based on the
activity data analyzed. We can create models that anticipate success based on this activity
data, being those models more accurate the longer the period under consideration. AUC
values confirm also models are better the longer the period we analyze.</p>
      <p>Prior to comparing with other LA/EDM works, we want to remark our focus has
been set on activity. We could obtain potential better models by analyzing partial grades
obtained by students, but this could mean losing the focus on the impact of student
behavior regarding the subject. This is particularly relevant for us, even more
considering we are facing changes in methodology, such as the flipped behavior.</p>
      <p>
        Limiting analysis to early stages provides average results. While we believe a longer
time period would be advisable, in particular if the goal is to take actions which can
derive costs, performance is similar to other published results. In particular, if we look
into the works compiled in [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ], we find studies with similar performance, such as [
        <xref ref-type="bibr" rid="ref18 ref19">18,
19</xref>
        ] – with indicators depicted as 3,72 and 73 in Figure 1–. In the first case, indicators
included in the study are socio-economic, while in the second they are related to
extracurricular activities. As a noticeable point, none of them includes grades as predictor
variables.
      </p>
      <p>
        Within our constraints, to get better results it is necessary to broaden the time scope.
Doing so – medium and late models – we get results similar to [
        <xref ref-type="bibr" rid="ref20 ref21">20, 21</xref>
        ] (whose variables
are depicted as points 1,15 and 36 in Figure 1). Those studies have also broader time
scopes (minimum 1 year) and do not constrain to limit grading data. Best predicting
scenarios in [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ] include always graded data [
        <xref ref-type="bibr" rid="ref22 ref23">22, 23</xref>
        ].
      </p>
      <p>Being able to predict results considering whether activities have been performed or
not is relevant from a pedagogical point of view and opens potential future lines. While
it is not the main goal of this research we also have analyzed results in Table 3 regarding
individual variable impact. If we deep into variable details, attending to class, or
performing homework makes a difference. Looking this fact from another angle, we can
tell students that coming to class and doing what they are instructed to will help them
to pass the subject.</p>
      <p>Class attendance is present in all cases, independently of time scope. Regarding
homework, during the first period, it is relevant your attitude in the first weeks, and just
before first partial test. When the period is longer, it becomes more relevant the amount
of work performed in the whole period. We believe there is even room for improvement
with this same dataset trying to look for other pieces of information that can remain
unnoticed inside the huge data volume.</p>
      <p>
        We have compared our findings with results in [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ] . In that case, authors consider
“there is a relation between some type of interactions and academic performance in
online courses, whereas this relation is non-significant in the case of VLE-supported
F2F courses”. We believe this relation can be found also in F2F or VLE supported
studies as long as they include a pedagogical design that requires the use of VLEs. If
this is done so, evidences will be gathered in the LMSs and can show differences in
behavior.
      </p>
      <p>
        Consistently with findings by Bowers et al. [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ] for dropout flags, and in particular
for early prediction, models still lack accuracy (i.e. classification error is high). The
potential impact of this error will depend on the purpose of the prediction. If we use
them to just raise early alarms, it would not be critical. If deeper pedagogical actions
are taken to redress behaviors that can anticipate failure, there would be non-optimal
use of resources
      </p>
      <p>Although the models exposed can set the basics for targeted actions, the design of
specific policies or pedagogical interventions to reduce failure should take into account
not only global accuracy but the impact of false positives and false negatives. For
instance, we could consider small group tutoring actions for students classified as likely
to fail. The specific design of the action should be made taking into account the false
positive rate – i.e. it will affect students who would potentially pass without the action
– and at the same time the false negative rate – i.e. there is a group of students marked
as passing who will not potentially pass -. We believe this analysis opens a really
interesting future line in the field of pedagogical design.</p>
      <p>The longer the period, the higher the values for accuracy and sensitivity. In other
words, with longer periods we are more certain about final results regarding true
positive class. In our case, that means we are more certain behavior of the user could lead
to failure.</p>
      <p>While this has been a model for an individual subject, we would like to make a
reflection about robustness and portability. Computation of data in Table 2 has been done
through cross-folding validation, and shows high values of variance in some cases (i.e.
sensitivity in early models). To solve this issue it would be advisable to have a higher
number of samples (i.e. more students to analyze).</p>
      <p>Regarding portability, the process we followed to extract information shows there is
a great dependence on course design. Portability of the resulting model itself is not
straightforward, but we believe the methodological approach is. The analysis followed
can help to obtain models for any flipped classroom course. A pedagogical design that
helps gathering evidences from LMS, combined with meaningful variables and, for F2F
universities, class attendance should generate models that anticipate potential success
based on pure activity data. We believe it will be difficult to generate portable results
among different subjects even in same university unless they share a common course
design.</p>
      <p>Although it was not the goal to establish a comparison among algorithms, Bayesian
models have shown good performance related to computational cost. SVM performs
also well, but at a higher computational cost. Decision-tree family algorithms can be of
interest but would need a higher number on samples to avoid deviations. Finally, deep
learning techniques have not provided considerable gain and have higher computation
requirements.</p>
      <p>To sum up, and going back to the research questions introduced in this paper, final
performance of individual students can be anticipated considering only activity data.
Relevant aspects for success, considering the course design in our study, include class
attendance and different aspects related to homework. Regarding the influence of time,
early periods lack accuracy, and would not be optimal if the goal is to set-up actions
which involve high costs. As we consider longer periods – medium and late – the
models get better. Results for these medium and late models can be of potential help both
to redress behavior – in the case of the medium prediction – or – once course is finished
– to analyze results and improve course design for future course sessions.
6</p>
    </sec>
    <sec id="sec-5">
      <title>Open lines</title>
      <p>This paper wants to set the basics for defining specific actions to reduce failure in
engineering studies in higher education. Lines of activity include:</p>
      <p>Improve models, in particular in early periods, potentially including new data.
Deepen into the meaning of the variables selected as more relevant.</p>
      <p>Apply same methodology to other subjects in order to validate and compare
results.</p>
      <p>Define actions to reduce failure based on early and medium prediction analysis
and to improve pedagogical design based on early, medium and late predictions.</p>
      <p>
        Authors are open to collaboration in previous lines – or to carry out similar research
in other environments –. For those interested in carrying out similar research on their
own, data processing was done through Python scripts, using Scikit-learn libraries
(https://scikit-learn.org/) for modelling algorithms. In particular, sklearn.naive_bayes,
sklearn.neural_network, sklearn.tree, sklearn.ensemble (for GBT) and sklearn.svm
implementations were relevant among those used [
        <xref ref-type="bibr" rid="ref24">24</xref>
        ]. The method does not rely in any
particular LMS, but our study was carried out on Moodle platform
(https://moodle.org/). Final models were also tested on RapidMiner software
(https://rapidminer.com/) to validate results.
in
analytics,"
[Online].
      </p>
      <p>Available:</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          1.
          <string-name>
            <given-names>C.</given-names>
            <surname>Romero</surname>
          </string-name>
          and
          <string-name>
            <given-names>S.</given-names>
            <surname>Ventura</surname>
          </string-name>
          ,
          <article-title>"Educational data mining: A survey from 1995 to 2005," Expert Systems with Applications</article-title>
          , vol.
          <volume>33</volume>
          , no.
          <issue>1</issue>
          , pp.
          <fpage>135</fpage>
          -
          <issue>146</issue>
          , 7
          <year>2007</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          2.
          <string-name>
            <given-names>R. S. J. D.</given-names>
            <surname>Baker</surname>
          </string-name>
          and
          <string-name>
            <given-names>K.</given-names>
            <surname>Yacef</surname>
          </string-name>
          ,
          <article-title>"The State of Educational Data Mining in 2009: A Review and Future Visions"</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          3.
          <string-name>
            <given-names>G.</given-names>
            <surname>Siemens</surname>
          </string-name>
          ,
          <article-title>"What are Learning Analytics?</article-title>
          ,"
          <year>2010</year>
          . [Online].
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          4.
          <string-name>
            <given-names>R.</given-names>
            <surname>Ferguson</surname>
          </string-name>
          ,
          <article-title>"Learning analytics: drivers, developments and challenges Journal Item Learning analytics: drivers, developments and challenges,"</article-title>
          <source>International Journal of Technology Enhanced Learning</source>
          , vol.
          <volume>4</volume>
          , no.
          <issue>5</issue>
          , pp.
          <fpage>304</fpage>
          -
          <lpage>317</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          5.
          <string-name>
            <given-names>Á. F.</given-names>
            <surname>Agudo-Peregrina</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Iglesias-Pradas</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Á</surname>
          </string-name>
          .
          <article-title>Conde-González and Á. HernándezGarcía, "Can we predict success from log data in VLEs? Classification of interactions for learning analytics and their relation with performance in VLE-supported F2F and online learning," Computers in Human Behavior</article-title>
          , vol.
          <volume>31</volume>
          , pp.
          <fpage>542</fpage>
          -
          <issue>550</issue>
          , 2
          <year>2014</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          6.
          <string-name>
            <given-names>R.</given-names>
            <surname>Conijn</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Snijders</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Kleingeld</surname>
          </string-name>
          and
          <string-name>
            <given-names>U.</given-names>
            <surname>Matzat</surname>
          </string-name>
          ,
          <article-title>"Predicting student performance from LMS data: A comparison of 17 blended courses using moodle LMS,"</article-title>
          <source>IEEE Transactions on Learning Technologies</source>
          ,
          <year>2017</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>7. "The Indicators Project - Dabbling https://indicatorsproject.wordpress.com/.</mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          8.
          <string-name>
            <given-names>P.</given-names>
            <surname>Strecht</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Cruz</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Soares</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Mendes-Moreira</surname>
          </string-name>
          and
          <string-name>
            <given-names>R.</given-names>
            <surname>Abreu</surname>
          </string-name>
          ,
          <article-title>"A Comparative Study of Classification and Regression Algorithms for Modelling Students' Academic Performance"</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          9.
          <string-name>
            <given-names>J.</given-names>
            <surname>Yoo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Yoo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Lance</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Hankins</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Yoo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Yoo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Lance</surname>
          </string-name>
          and
          <string-name>
            <given-names>J.</given-names>
            <surname>Hankins</surname>
          </string-name>
          ,
          <article-title>"Student progress monitoring tool using treeview,"</article-title>
          <source>in Proceedings of the 37th SIGCSE technical symposium on Computer science education - SIGCSE '06</source>
          , New York, New York, USA,
          <year>2006</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          10.
          <string-name>
            <surname>P. J. H. Y. I.-H. J. Yeonjeong</surname>
          </string-name>
          ,
          <article-title>"Clustering blended learning courses by online behavior data: A case study in a Korean higher education institute,"</article-title>
          <source>The Internet and Higher Education</source>
          , pp.
          <fpage>1</fpage>
          -
          <lpage>11</lpage>
          ,
          <year>2016</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          11.
          <string-name>
            <given-names>F.</given-names>
            <surname>Castro</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Vellido</surname>
          </string-name>
          , À. Nebot and
          <string-name>
            <given-names>F.</given-names>
            <surname>Mugica</surname>
          </string-name>
          ,
          <article-title>"Applying Data Mining Techniques to eLearning Problems"</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          12.
          <string-name>
            <surname>A. M. Shahiri</surname>
            ,
            <given-names>W.</given-names>
          </string-name>
          <string-name>
            <surname>Husain</surname>
            ,
            <given-names>N. '.</given-names>
          </string-name>
          <string-name>
            <surname>Aini</surname>
            and
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Rashid</surname>
          </string-name>
          ,
          <article-title>"ScienceDirect The Third Information Systems International Conference A Review on Predicting Student's Performance using Data Mining Techniques,"</article-title>
          <source>Procedia - Procedia Computer Science</source>
          , vol.
          <volume>72</volume>
          , pp.
          <fpage>414</fpage>
          -
          <lpage>422</lpage>
          ,
          <year>2015</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          13.
          <string-name>
            <given-names>S.</given-names>
            <surname>Rosset</surname>
          </string-name>
          ,
          <article-title>"Model selection via the AUC</article-title>
          .,
          <article-title>" in Machine Learning</article-title>
          ,
          <source>Proceedings of the 21st International Conference</source>
          ,
          <year>2004</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          14.
          <string-name>
            <surname>C. Ferri</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          <string-name>
            <surname>Hernández-Orallo</surname>
            and
            <given-names>R.</given-names>
          </string-name>
          <string-name>
            <surname>Modroiu</surname>
          </string-name>
          ,
          <article-title>"An experimental comparison of performance measures for classification," Pattern Recognition Letters</article-title>
          , vol.
          <volume>30</volume>
          , pp.
          <fpage>27</fpage>
          -
          <lpage>38</lpage>
          ,
          <year>2008</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          15.
          <string-name>
            <surname>J. M. Lobo</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Jiménez-Valverde</surname>
            and
            <given-names>R.</given-names>
          </string-name>
          <string-name>
            <surname>Real</surname>
          </string-name>
          ,
          <article-title>"AUC: a misleading measure of the performance of predictive distribution models,"</article-title>
          <source>Global Ecology and Biogeography</source>
          , vol.
          <volume>17</volume>
          , no.
          <issue>2</issue>
          , pp.
          <fpage>145</fpage>
          -
          <issue>151</issue>
          ,
          <fpage>1</fpage>
          <string-name>
            <surname>3</surname>
          </string-name>
          <fpage>2008</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          16.
          <string-name>
            <surname>A. J. Bowers</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          <string-name>
            <surname>Sprott</surname>
            and
            <given-names>S.</given-names>
          </string-name>
          <string-name>
            <surname>Taff</surname>
          </string-name>
          ,
          <article-title>"Do We Know Who Will Drop Out? A Review of the Predictors of Dropping out of High School: Precision, Sensitivity and Specificity,"</article-title>
          vol.
          <volume>96</volume>
          , no.
          <issue>2</issue>
          , pp.
          <fpage>77</fpage>
          -
          <lpage>100</lpage>
          ,
          <year>2013</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          17.
          <string-name>
            <given-names>F. E. J.</given-names>
            <surname>Harrell</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K. L.</given-names>
            <surname>Lee</surname>
          </string-name>
          and
          <string-name>
            <given-names>D. B.</given-names>
            <surname>Mark</surname>
          </string-name>
          ,
          <article-title>" "Multivariable prognostic models: issues in developing models, evaluating assumptions and adequacy, and measuring and reducing errors,"</article-title>
          <source>Stat Med</source>
          , pp.
          <fpage>361</fpage>
          -
          <lpage>387</lpage>
          ,
          <year>1996</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          18.
          <string-name>
            <surname>A. I.</surname>
          </string-name>
          <article-title>""Mother got tired of taking care of my baby." A study of dropouts," Austin</article-title>
          , Texas: Austin Independent School District.,
          <year>1982</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          19.
          <string-name>
            <given-names>J. L.</given-names>
            <surname>Mahoney and R. B. Cairns</surname>
          </string-name>
          ,
          <article-title>"Do extracurricular activities protect against early school dropout?,"</article-title>
          <source>Developmental psychology</source>
          , vol.
          <volume>33</volume>
          , no.
          <issue>2</issue>
          , pp.
          <fpage>241</fpage>
          -
          <lpage>253</lpage>
          ,
          <year>1997</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          20. E. Allensworth and
          <string-name>
            <given-names>J.</given-names>
            <surname>Easton</surname>
          </string-name>
          ,
          <article-title>"What matters for staying on-track and graduating in chicago public high schools: A close look at course grades, failures, and attendance in the freshman year</article-title>
          ,"
          <year>2007</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          21.
          <string-name>
            <given-names>D.</given-names>
            <surname>Doss</surname>
          </string-name>
          ,
          <article-title>"Ninth grade course enrollment and dropping out," in Annual Meeting of the American Educational Research Association</article-title>
          , San Francisco,
          <year>1986</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          22.
          <string-name>
            <given-names>A. J.</given-names>
            <surname>Bowers</surname>
          </string-name>
          and
          <string-name>
            <given-names>R.</given-names>
            <surname>Sprott</surname>
          </string-name>
          ,
          <article-title>"Examining the multiple trajectories associated with dropping out of high school: A growth mixture model analysis,"</article-title>
          <source>Journal of Educational Research</source>
          , vol.
          <volume>105</volume>
          , no.
          <issue>3</issue>
          , pp.
          <fpage>176</fpage>
          -
          <issue>195</issue>
          ,
          <fpage>1</fpage>
          <string-name>
            <surname>4</surname>
          </string-name>
          <fpage>2012</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          23.
          <string-name>
            <given-names>B.</given-names>
            <surname>Muthén</surname>
          </string-name>
          ,
          <article-title>"Latent Variable Analysis: Growth Mixture Modeling and Related Techniques for Longitudinal Data," in Handbook of quantitative methodology for the social sciences, D</article-title>
          . Kaplan, Ed.,
          <string-name>
            <surname>Sage</surname>
            <given-names>Publications</given-names>
          </string-name>
          ,
          <year>2004</year>
          , pp.
          <fpage>345</fpage>
          -
          <lpage>368</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          24.
          <string-name>
            <surname>«API Reference -</surname>
          </string-name>
          scikit
          <source>-learn 0.21</source>
          .2 documentation» [Online]. Available: https://scikitlearn.org/stable/modules/classes.html.
        </mixed-citation>
      </ref>
      <ref id="ref25">
        <mixed-citation>
          25.
          <string-name>
            <given-names>L. P.</given-names>
            <surname>Macfadyen</surname>
          </string-name>
          and
          <string-name>
            <given-names>S.</given-names>
            <surname>Dawson</surname>
          </string-name>
          ,
          <article-title>"Mining LMS data to develop an “early warning system” for educators: A proof of concept,"</article-title>
          <source>Computers &amp; Education</source>
          , vol.
          <volume>54</volume>
          , no.
          <issue>2</issue>
          , pp.
          <fpage>588</fpage>
          -
          <issue>599</issue>
          , 2
          <year>2010</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>