<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>A predictive model for identifying students with dropout profiles in online courses</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Marcelo A. Santana</string-name>
          <email>marcelo.almeida@nti.ufal.br</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Evandro B. Costa</string-name>
          <email>evandro@ic.ufal.br</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Baldoino F. S. Neto</string-name>
          <email>baldoino@ic.ufal.br</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Italo C. L. Silva</string-name>
          <email>italocarlo@nti.ufal.br</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Joilson B. A. Rego</string-name>
          <email>jotarego@gmail.com</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Institute of Computing, Federal University of Alagoas</institution>
        </aff>
      </contrib-group>
      <abstract>
        <p>Online education often deals with the problem related to the high students' dropout rate during a course in many areas. There is huge amount of historical data about students in online courses. Hence, a relevant problem on this context is to examine those data, aiming at nding e ective mechanisms to understand student pro les, identifying those students with characteristics to drop out at early stage in the course. In this paper, we address this problem by proposing predictive models to provide educational managers with the duty to identify students whom are in the dropout bound. Four classi cation algorithms with di erent classi cation methods were used during the evaluation, in order to nd the model with the highest accuracy in prediction the pro le of dropouts students. Data for model generation were obtained from two data sources available from University. The results showed the model generated by using SVM algorithm as the most accurate among those selected, with 92.03% of accuracy.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;Dropout</kwd>
        <kwd>Distance Learning</kwd>
        <kwd>Educational Data Mining</kwd>
        <kwd>Learning Management Systems</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. INTRODUCTION</title>
      <p>
        Every year, the registration marks in E-learning modality
has increased considerably, in 2013, 15.733 courses were
offered, in E-learning or semi-presence modality.
Furthermore, the institutions are very optimistic, 82% of researched
places, believe that the amount of registration marks will
have a considerable expansion in 2015 [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ], showing the
Elearning evolution and its importance as a tool for citizen's
formation. The Learning Management Systems (LMS) [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ]
can be considered one of factors that has had an important
role for popularization of this learning modality [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ].
Despite the rapid growth of online courses, there has also
been rising concern over a number of problems. One issue
in particular that is di cult to ignore is that these online
courses also have high dropout rates. Speci cally, in Brazil,
in 2013, according with the latest Censo, published by the
Elearning Brazilian Association (ABED), the dropout average
was about 19,06% [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ].
      </p>
      <p>Beyond the hard task on identifying the students who can
have possible risk of dropping out, the same dropout also
brings a huge damage to current nancial and social
resources. Thus, the society also loses when they are poorly
managed, once the student lls the vacancy but he gives up
the course before the end.</p>
      <p>
        Online education often deals with the problem related to the
high students' dropout rate during a course in many areas.
There is huge amount of historical data about students in
online courses. Hence, a relevant problem on this context is to
examine those data, aiming at nding e ective mechanisms
to understand student pro les, identifying those students
with characteristics to drop out at early stage in the course.
In this paper, we address this problem by proposing
predictive models to provide educational managers with the duty
of identifying students who are in the dropout bound. This
predictive model took in consideration academic elements
related with their performance at the initial disciplines of
the course. Data from System Information course at
Federal University of Alagoas (UFAL) were used to build this
model, which uses a very known LMS, called Moodle.
A tool to support the pre-processing phase was used in order
to prepare data for application of Data Mining algorithms.
The Pentaho Data Integration [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ] tool covers the extraction
areas, transformation and data load (ETL), making easier
the archive generation in the compatible format with the
data mining software adopted, called WEKA[
        <xref ref-type="bibr" rid="ref5">5</xref>
        ].
Therefore, for what was exposed above, it justi es the
needing of an investment to develop e cient prediction methods,
assessment and follow up of the students with dropout risk,
allowing a future scheduling and adoption of proactive
measures aiming the decrease of the stated condition.
The rest of the paper is organized as follows. Section 2
presents some related work. Section 3 Environment for
Construction of predictive model. Afterwards, we present the
experiment settings in Section 4, and in Section 5 we
discuss the results of the experiment. Section 6 presents some
concluding remarks and directions of future work.
      </p>
    </sec>
    <sec id="sec-2">
      <title>2. RELATED WORK</title>
      <p>
        Several studies have been conducted in order to nd out the
reasons of high dropout indices in online courses. Among
them, Xenos [
        <xref ref-type="bibr" rid="ref18">18</xref>
        ] makes a review of the Open University
students enrolled in a computing course. In this studies, ve
acceptable reasons, that might have caused the dropout, were
identi ed: Professional (62,1%), Academic (46%), Family
(17,8%), Health Issues (9,5%), Personal Issues (8,9%).
According to Barroso and Falca~o (2004) [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ] the motivational
conditions to the dropout are classi ed in three groups: i)
Economic - Impossibility of remaining in the course because
of socio-economics issues; ii) Vocational - The student is not
identi ed with the chosen course. iii) Institutional -
Failure on initial disciplines, previous shortcomings of earlier
contents, inadequacy with the learning methods.
Manha~es et al.[
        <xref ref-type="bibr" rid="ref14">14</xref>
        ] present a novel architecture that uses
EDM techniques to predict and identify those who are at
dropout risk. The paper shows initial experimental results
using real world data about of three undergraduate
engineering courses of one the largest Brazilian public university.
According to the experiments, the classi er Naive Bayes
presented the highest true positive rate for all datasets used in
the experiments.
      </p>
      <p>
        A model for predicting students' performance levels is
proposed by Erkan Er [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ]. Three machine learning algorithms
were employed: instance-based learning Classi er, Decision
Tree and Naive Bayes. The overall goal of the study is to
propose a method for accurate prediction of at-risk students
in an online course. Speci cally, data logs of LMS, called
METU-Online, were used to identify at-risk students and
successful students at various stages during the course. The
experiment were realized in two phases: testing and
training. These phases were conducted at three steps which
correspond to di erent stages in a semester. At each step, the
number of attributes in the dataset had been increased and
all attributes were included at nal stage. The important
characteristic of the dataset was that it only contained
timevarying attributes rather than time-invariant attributes such
as gender or age. According to the author, these data did
not have signi cant impact on overall results.
      </p>
      <p>
        Dekker [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ] in your paper presents a data mining case study
demonstrating the e ectiveness of several classi cation
techniques and the cost-sensitive learning approach on the dataset
from the Electrical Engineering department of Eindhoven
University of Technology. Was compared two decision tree
algorithms, a Bayesian classi er, a logistic model, a
rulebased learner and the Random Forest. Was also considered
the OneR classi er as a baseline and as an indicator of the
predictive power of particular attributes. The experimental
results show that rather simple classi ers give a useful result
with accuracies between 75 and 80% that is hard to beat
with other more sophisticated models. We demonstrated
that cost-sensitive learning does help to bias classi cation
errors towards preferring false positives to false negatives.
We believe that the authors could get better results by
making some adjustments to the parameters of the algorithms.
Jaroslav [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ], aims to research to develop a method to
classify students at risk of dropout throughout the course. Using
personal data of students enriched with data related to
social behaviours, Jaroklav uses dimensionality reduction
techniques and various algorithms in order to nd which of the
best results managing to get the accuracy rates of up to
93:51%, however the best rates are presented at the end of
the course. Whereas the goal is to identify early on dropout,
the study would be more relevant if the best results were
obtained results at the beginning of the course.
      </p>
      <p>In summary, several studies investigating the application of
EDM techniques to predict and identify students who are
at risk dropout. However, those works share similarities:
(i) identify and compare algorithm performance in order to
nd the most relevant EDM techniques to solve the
problem or (ii) identify the relevant attributes associated with
the problem. Some works use past time-invariant student
records (demographic and pre-university student data). In
this study, contribution to those presented in this section,
makes the junction between two di erent systems, gathering
a larger number of attributes, variables and time invariant.
Besides being concerned with the identi cation and
comparison of algorithms, identify the attributes of great relevance
and solve the problem the predict in more antecedence the
likely to dropout students.</p>
    </sec>
    <sec id="sec-3">
      <title>3. ENVIRONMENT FOR CONSTRUCTION</title>
    </sec>
    <sec id="sec-4">
      <title>OF PREDICTIVE MODEL</title>
      <p>This subsection presents an environment for construction
for a predictive model for supporting educators in the task
of identifying prospective students with dropout pro les in
online courses. The environment is depicted in Figure 1.
The proposed environment in this work is composed by three
layers: Data source, Model development and Model. The
data sources are located in the rst layer. Data about all
students enrolled at the University are stored in two data
sources: The rst one contains students' personal data, for
example: age, gender, income, marital status and grades
from the academic control system used by the University.
Information related with frequency of access, participation,
use of the tools available, and grades of students related
the activities proposed within the environment are kept in
second data source.</p>
      <p>
        In the second layer, the pre-processing [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ] activity over the
data is initiated. Sequential steps are executed in this layer
in order to prepare them to data mining process. In the
original data some information can not be properly
represented in a expected format by data mining algorithm, data
redundancy or even data with some kind of noise. These
problems can produce misleading results or make the
algorithm execution becomes computationally more expensive.
This layer is divided into the following stages: data
extraction, data cleaning, data transformation, data selection and
the choice of algorithm that best ts the model. Just below,
will be displayed brie y each step of this layer.
      </p>
      <p>Data extraction: The extraction phase establishes the
connection with the data source and performs the extraction of
the data.</p>
      <p>Data cleaning: This routine tries to ll missing values, smooth
out noise while identifying outliers, and correct data
inconsistencies.</p>
      <p>Data transformation: In this step, data are transformed and
consolidated into appropriate forms for mining by
performing summary or aggregation operations. Sometimes, data
transformation and consolidation are performed before the
data selection process, particularly in the case of data
warehousing. Data reduction may also be performed to obtain a
smaller representation of the original data without sacri
cing its integrity.</p>
      <p>Data selection: In this step, relevant data to the analysis
task are retrieved from the database.</p>
      <p>Choice of algorithm: An algorithm to respond with quality
in terms of accuracy, which has students elusive pro le, was
considered the algorithm that best applies to the model.
Finally, the last layer is the presentation of the model. This
layer is able to post-processing the result obtained in the
lower layer and presenting it to the end-user of a most
understandable way.</p>
    </sec>
    <sec id="sec-5">
      <title>4. EXPERIMENT SETTINGS</title>
      <p>
        The main objective of this present research is to build a
predictive model for supporting educators in the hard task
of identifying prospective students with dropout pro les in
online courses, using Educational Data Mining (EDM)
techniques [
        <xref ref-type="bibr" rid="ref16">16</xref>
        ]. This section is organized as follows: Section 4.1
describes the issue which drives our assessment. Section 4.2
shows which data were selected for to the data group
utilized in the experiment and which algorithms were chosen for
data mining execution. Section 4.3 indicates the employed
tools during the execution of experiment. Finally, Section
4.4 shows every step in experiment execution, including data
consolidation, data preprocessing and algorithms execution.
      </p>
    </sec>
    <sec id="sec-6">
      <title>4.1 Planning</title>
      <p>The research question that we would like to answer is:
RQ.Is our predictive model able to early identify the
students with dropout risk?
In order to answer this question, EDM techniques with four
di erent classi cation methods were used, aiming to get a
predictive model which answers us with quality in precise
ways which students have a dropout pro le, taking in
consideration only data about the initial disciplines of a speci ed
course.</p>
    </sec>
    <sec id="sec-7">
      <title>4.2 Subject Selection</title>
      <sec id="sec-7-1">
        <title>4.2.1 Data Selection</title>
        <p>
          The Federal University of Alagoas o ers graduation courses,
postgraduate courses and E-learning courses. In the on line
courses, there are more than 1800 registered students[
          <xref ref-type="bibr" rid="ref4">4</xref>
          ].
An E-learning course is usually partitioned in semesters,
where di erent disciplines are taught along these semesters.
Each semester usually has ve disciplines per semester, and
each discipline has a duration between ve to seven weeks.
Anonymous data, from the Information Systems E-learning
course, were selected from this environment, relative to rst
semester in 2013. Data of one discipline (Algorithm and
Data Structure I), chosen based on its relevance, were
analysed. Such discipline has about 162 students enrolled.
        </p>
      </sec>
      <sec id="sec-7-2">
        <title>4.2.2 Machine Learning Algorithms Selection</title>
        <p>
          In this work to predict student dropouts, four machine
learning algorithms were used, using di erent classi cation
methods. The methods used were: simple probabilistic classi er
based on the application of Bayes' theorem, decision tree,
support vector's machine and multilayer neural network.
These techniques have been successfully applied to solve
various classi cation problems and function in two phases: (i)
training and (ii) testing phase. During the training phase
each technique is presented with a set of example data pairs
(X, Y), where X represents the input and Y the respective
output of each pair [
          <xref ref-type="bibr" rid="ref13">13</xref>
          ]. In this study, Y can receive one
of the following values, \approved" or \reproved", that
corresponds the student situation in discipline.
        </p>
      </sec>
    </sec>
    <sec id="sec-8">
      <title>4.3 Instrumentation</title>
      <p>
        The Pentaho Data Integration [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ] tool was chosen to realize
all preprocessing steps on selected data. Pentaho is a
opensource software, developed in Java, which covers extraction
areas, transform and load of the data [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ], making easier the
creation of an model able to : (i) extract information from
data sources, (ii) attributes selection, (iii) data discretization
and (iv) le generation in a compatible format with the data
mining software.
      </p>
      <p>
        For execution of selected classi cation algorithms (see
Section 4.2.2), the data mining tool Weka was selected. Such
algorithms are implemented on Weka software as NaiveBayes
(NB), J48 (AD), SMO (SVM), MultilayerPerceptron (RN)
[
        <xref ref-type="bibr" rid="ref17">17</xref>
        ] respectively. Weka is a software of open code which
contains a machine learning algorithms group to data's mining
task [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ].
      </p>
      <p>
        Some features were taken in consideration for Weka [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ]
adoption, such as: ease of acquisition, facility and
availability to directly download from the developer page with no
operation cost; Attendance of several algorithms versions
set in data mining and availability of statistical resources to
compare results among algorithms.
      </p>
    </sec>
    <sec id="sec-9">
      <title>4.4 Operation</title>
      <p>The evaluation of experiment was executed on HP Probook
2.6 GHz Core-I5 with 8Gb of memory, running Windows
8.1.</p>
      <sec id="sec-9-1">
        <title>4.4.1 Data’s Preprocessing</title>
        <p>
          Real-world data tend to be dirty, incomplete, and
inconsistent. Data preprocessing techniques can improve data
quality, thereby helping to improve the accuracy and e ciency
of the subsequent mining process [
          <xref ref-type="bibr" rid="ref11">11</xref>
          ].
        </p>
        <p>
          Currently, the data is spread in two main data sources:
LMS Moodle, utilized by the University as assistance on
Elearning teaching, including data which show the access
frequency, student's participation using the available tools, as
well as the student's success level related to proposed
activities. Meanwhile, student's personal les as age, sex, marital
status, salary and disciplines grades are kept in the
Academic Control System (ACS), which is a Software designed
to keep the academic control of the whole University [
          <xref ref-type="bibr" rid="ref4">4</xref>
          ].
Aiming to reunite a major data group and work only with
relevant data to the research question that we want to
answer, we decided to perform consolidation of these two data
source in a unique major data source, keeping their integrity
and ensuring that only relevant information will be used
during data mining algorithms execution.
        </p>
        <p>
          Careful integration can help reduce and avoid redundancies
and inconsistencies in the resulting data set. This can help
improve the accuracy and speed of the data mining
process [
          <xref ref-type="bibr" rid="ref11">11</xref>
          ].
        </p>
        <p>To maintain the integrity and reliability between data, a
mandatory attribute, with unique value and present between
in both data sources, was chosen. Thus, the CPF attribute
was chosen to make data uni cation between the two
selected data sources, once it permits the unique identi cation
among selected students.</p>
        <p>In order to facilitate algorithms execution and
comprehension of results,predicting the dropout in an early stage of
the study. In order to achieve a high rate of accuracy and
minimum of false negatives, i.e. students that have not been
recognized to be in danger of dropout. Some attributes were
transformed, as we can seen below:</p>
        <p>The corresponding attributes related with discipline
grades were discretized in a ve-group-value (A,B,C,D
e E), depending on the discipline's achieved grades.
The student with a grade higher or equal 9, was
allocated for \A" group. Those ones who had their grades
between 8,99 and 7 were allocated for \B" group. the
\C" students are those that had a grade between 6,99
and 5, and those who had grades under 5,99 stayed at
\D" group and nally those that doesn't have a grade
associated were allocated in "E" group.</p>
        <p>Every student was labelled as approved or reproved
based on the situation informed by the academics
registers. The nal score of each discipline is composed by
two tests, if the student did not succeed in obtaining
the minimum average, he will be leaded to the nal
reassessment and nal test.</p>
        <p>In the \City" attribute, some inconsistencies were found,
where di erent data about the same city were
registered in database. For instance, the instances of Ouro
Branco and Ouro Branco/AL are related to same
city. This problem was totally solved, with application
of techniques for grouping attributes.</p>
        <p>The attribute \age" had to be calculated. For this, the
student's birth date, registered in database, was taken
in consideration.</p>
        <p>When all the attributes were used the accuracy was low.
That is why we utilized feature selection methods to
reduce the dimensionality of the student data extracted from
dataset. We improved the pre-processing method the data.
In order to preserve reliability of attributes for classi cation
after the reduction. We use InfoGainAttributeEval
algorithm that builds a rank of the best attributes considering
the extent of information gain based on the concept of
entropy.</p>
        <p>After this procedure, we reduced the set of attributes from
17 to 13 most relevant. The list of the re ned set of
attributes in relevance ordercan be found in Table 1.
Taking in consideration that the main objective is to predict
student's nal situation with the earlier advance as possible
inside the given discipline, to this study we will only use
data until the moment of the rst test.</p>
        <p>The Figure 2 presents all the executed stages, during the
preprocessing phase, in order to generate a compatible le
with the mining software.</p>
      </sec>
      <sec id="sec-9-2">
        <title>4.4.2 Algorithms Execution</title>
        <p>
          The k-fold method was applied to make a assessment the
model generalization capacity, with k=10 (10-fold cross
validation). The cross validation method, consists in splitting
of the model in k subgroups mutually exclusive and with the
same size, from these subgroups, one subgroup is selected for
test and the remaining k-1's are utilized for training. The
average error rate of each training subgroup can be used as
an estimate of the classi er's error rate. When Weka
implements the cross validation, it trains the classi er k times to
calculate the average error rate and nally, leads the build
classi er back utilizing the model as a training group. Thus,
the average error rate provides a better solution in terms of
classi er's error accuracy reliability [
          <xref ref-type="bibr" rid="ref12">12</xref>
          ].
        </p>
        <p>In order to get the best results of the algorithms without
losing generalization, some parameters of SVM algorithms
were adjusted.</p>
        <p>
          The rst parameter was set the parameter \C". This
parameter is for the soft margin cost function, which controls
the in uence of each individual support vector; this process
involves trading error penalty for stability [
          <xref ref-type="bibr" rid="ref3">3</xref>
          ].
        </p>
        <p>
          The default kernel used by Weka tool is the polynomial we
changed to the Gaussian setting the parameters Gamma.
Gamma is the free parameter of the Gaussian radial basis
function [
          <xref ref-type="bibr" rid="ref3">3</xref>
          ].
        </p>
        <p>After several adjustments to the values of the two
parameters mentioned above, which showed the best results in term
of accuracy and lower false positive rate, was C = 9.0 and
Gamma = 0.06 parameter.</p>
        <p>
          For comparison of results related to selected algorithms, we
used Weka Experiment Environment (WEE). The WEE
allows the selection of one or more algorithms available in the
tool as well as analyse the results, in order to identify, if a
classi er is, statistically, better than the other. In this
experiment, the cross validation method, with the parameter
\k=10" [
          <xref ref-type="bibr" rid="ref5">5</xref>
          ], is used in order to calculate the di erence on
the results in each one of the algorithms related to a chosen
standard algorithm (baseline).
        </p>
      </sec>
    </sec>
    <sec id="sec-10">
      <title>5. RESULTS AND DISCUSSIONS</title>
      <p>In this section, the results of the experiment, described in
Section 4, are analyzed.</p>
      <p>
        The WEE tool calculated the average accuracy of each
classi er. Table 2 shows the result of each algorithms
execution. The accuracy represents the percentage of the test
group instance which are correctly classi ed by the model
built during training phases. If the built model has a high
accuracy, the classi er is treated as e cient and can be put
into production [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ].
RN
90.86
0.85
0.15
0.93
0.07
positives is not suitable to our solution. In this case, we have
considered the algorithm which has the lower false positive
rates.
      </p>
      <p>As we can see on table 2 the algorithm SVM presented a low
false positive rate and better accuracy. Therefore, only the
best algorithm was considered to our solution. The Naive
Bayes classi er had the worst result in terms of accuracy
and a high false positive rate. The other ones had an error
average of 8%, and then, we end up with 8% of the students
with dropout risk not so correctly classi ed.</p>
    </sec>
    <sec id="sec-11">
      <title>5.1 Research Question</title>
      <p>
        As can be seen in table 2, in our experiment, the SVM
algorithm obtained 92% of accuracy. According to Han J. et
al. [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ] if the accuracy of the classi er is considered
acceptable, the classi er can be used to classify future data tuples
for which the class label is not known. Thus, the results
are pointing to the viability of model able to early identify
a possible student's dropout, based on their failures in the
initial disciplines.
      </p>
    </sec>
    <sec id="sec-12">
      <title>5.2 Statistical Significance Comparision</title>
      <p>
        We often need compare di erent learning schemes on the
same problem to see which is the better one to use. This
is a job for a statistical device known as the t-test, or
Student's t-test. A more sensitive version of the t-test known
as a paired t-test it was used. [
        <xref ref-type="bibr" rid="ref17">17</xref>
        ]. Using this value and
desired signi cance level (5%), consequently one can say that
these classi ers with a certain degree of con dence (100
signi cance level) are signi cantly di erent or not. By using
the t-test paired in the four algorithms, performed via Weka
analysis tool, observed that the SVM algorithm is signi
cantly respectful of others.
      </p>
    </sec>
    <sec id="sec-13">
      <title>5.3 Threats to validity</title>
      <p>The experiment has taken in consideration data from the
Information System course and the Data Structure Algorithm
discipline. However, the aforementioned discipline was
chosen, based on its importance in the context of Information
System course.</p>
    </sec>
    <sec id="sec-14">
      <title>6. CONCLUSION AND FUTURE WORK</title>
      <p>Understand the reasons behind the dropout in E-learning
education and identify in which aspects can be improved is
a challenge to the E-learning. One factor, which has been
pointed as in uencer of students' dropout, is the academic
element related with their performance at the initial
disciplines of the course.</p>
      <p>Comparing the results among the four algorithms, we can
verify that the accuracy oscillates around 85.5 to 92.03%.
Furthermore, a classi er which has a high error rate to false
This research has addressed dropout problem by proposing
predictive models to provide educational managers with the
duty to identify students whom are in the dropout bound.
The adopted approach allowed us to perform predictions
at an initial discipline phase. The preliminaries results has
shown that prediction model to identify students with dropout
pro les is feasible. These predictions can be very useful to
educators, supporting them in developing special activities
for these potential students, during the teaching-learning
process.</p>
      <p>As an immediate future work, some outstanding points still
should be regarded to the study's improvement, as apply the
same model in di erent institution databases with di erent
teaching methods and courses, including new factors related
to dropout as: professional, vocational and family data,
execute some settings in algorithms' parameters in order to
have the best achievements. Furthermore, a integrated
software to LMS, to provide this feedback to educators, will be
developed using this built model.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <surname>Abed -</surname>
          </string-name>
          E-learning Brazilian Association. http://www.abed.org.br/.
          <source>Accessed December</source>
          <year>2014</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <surname>Pentaho - Pentaho Data</surname>
          </string-name>
          Integration. http://www.pentaho.com/.
          <source>Accessed January</source>
          <year>2015</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          <article-title>[3] SVM - support vector machines (svms)</article-title>
          . http://www.svms.org/parameters/.
          <source>Accessed December</source>
          <year>2014</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4] UFAL - Federal University of Alagoas. http://www.ufal.edu.br/.
          <source>Accessed January</source>
          <year>2015</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5] Weka - the University of Waikato. http://www.cs.waikato.ac.nz/ml/weka/.
          <source>Accessed January</source>
          <year>2015</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>M. F.</given-names>
            <surname>Barroso</surname>
          </string-name>
          and
          <string-name>
            <given-names>E. B.</given-names>
            <surname>Falcao</surname>
          </string-name>
          . University dropout:
          <source>the case of ufrj physics institute. IX National Meeting of Research in Physics Teaching</source>
          ,
          <year>2004</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>J.</given-names>
            <surname>Bayer</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Bydzovska</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Geryk</surname>
          </string-name>
          ,
          <string-name>
            <surname>T.</surname>
          </string-name>
          <article-title>Obs vac, and L. Popel nsky. Predicting drop-out from social behaviour of students</article-title>
          . In
          <string-name>
            <surname>A. H. M. Y. Kalina Yacef</surname>
            ,
            <given-names>Osmar</given-names>
          </string-name>
          <string-name>
            <surname>Zaiane</surname>
          </string-name>
          and J. Stamper, editors,
          <source>Proceedings of the 5th International Conference on Educational Data Mining - EDM</source>
          <year>2012</year>
          , pages
          <fpage>103</fpage>
          {
          <fpage>109</fpage>
          ,
          <string-name>
            <surname>Greece</surname>
          </string-name>
          ,
          <year>2012</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>G.</given-names>
            <surname>Dekker</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Pechenizkiy</surname>
          </string-name>
          , and
          <string-name>
            <given-names>J.</given-names>
            <surname>Vleeshouwers</surname>
          </string-name>
          .
          <article-title>Predicting students drop out: A case study</article-title>
          . In T. Barnes,
          <string-name>
            <given-names>M. C.</given-names>
            <surname>Desmarais</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Romero</surname>
          </string-name>
          , and S. Ventura, editors,
          <source>EDM</source>
          , pages
          <volume>41</volume>
          {
          <fpage>50</fpage>
          ,
          <year>2009</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>E.</given-names>
            <surname>Er</surname>
          </string-name>
          .
          <article-title>Identifying at-risk students using machine learning techniques: A case study with is 100</article-title>
          . In
          <source>International Journal of Machine Learning and Computing</source>
          , pages
          <volume>476</volume>
          {
          <fpage>481</fpage>
          ,
          <string-name>
            <surname>Singapore</surname>
          </string-name>
          ,
          <year>2012</year>
          . IACSIT Press.
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>M.</given-names>
            <surname>Hall</surname>
          </string-name>
          , E. Frank,
          <string-name>
            <given-names>G.</given-names>
            <surname>Holmes</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Pfahringer</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Reutemann</surname>
          </string-name>
          ,
          <string-name>
            <given-names>and I. H.</given-names>
            <surname>Witten</surname>
          </string-name>
          .
          <article-title>The weka data mining software: An update</article-title>
          .
          <source>SIGKDD Explor</source>
          . Newsl.,
          <volume>11</volume>
          (
          <issue>1</issue>
          ):
          <volume>10</volume>
          {
          <fpage>18</fpage>
          ,
          <string-name>
            <surname>Nov</surname>
          </string-name>
          .
          <year>2009</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>J.</given-names>
            <surname>Han</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Kamber</surname>
          </string-name>
          , and
          <string-name>
            <given-names>J.</given-names>
            <surname>Pei</surname>
          </string-name>
          .
          <article-title>Data Mining: Concepts and Techniques</article-title>
          . Morgan Kaufmann Publishers Inc., San Francisco, CA, USA, 3rd edition,
          <year>2011</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>S. B.</given-names>
            <surname>Kotsiantis</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Pierrakeas</surname>
          </string-name>
          , and
          <string-name>
            <given-names>P. E.</given-names>
            <surname>Pintelas</surname>
          </string-name>
          .
          <article-title>Preventing student dropout in distance learning using machine learning techniques</article-title>
          . In V. Palade,
          <string-name>
            <given-names>R. J.</given-names>
            <surname>Howlett</surname>
          </string-name>
          , and L. C. Jain, editors,
          <source>KES</source>
          , volume
          <volume>2774</volume>
          of Lecture Notes in Computer Science, pages
          <volume>267</volume>
          {
          <fpage>274</fpage>
          . Springer,
          <year>2003</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <given-names>I.</given-names>
            <surname>Lykourentzou</surname>
          </string-name>
          ,
          <string-name>
            <given-names>I.</given-names>
            <surname>Giannoukos</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Nikolopoulos</surname>
          </string-name>
          , G. Mpardis, and
          <string-name>
            <given-names>V.</given-names>
            <surname>Loumos</surname>
          </string-name>
          .
          <article-title>Dropout prediction in e-learning courses through the combination of machine learning techniques</article-title>
          .
          <source>Comput. Educ.</source>
          ,
          <volume>53</volume>
          (
          <issue>3</issue>
          ):
          <volume>950</volume>
          {
          <fpage>965</fpage>
          ,
          <string-name>
            <surname>Nov</surname>
          </string-name>
          .
          <year>2009</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <surname>L. M. B. Manh</surname>
          </string-name>
          <article-title>~aes, S. M. S. da</article-title>
          <string-name>
            <surname>Cruz</surname>
            , and
            <given-names>G.</given-names>
          </string-name>
          <article-title>Zimbra~o. Wave: An architecture for predicting dropout in undergraduate courses using edm</article-title>
          .
          <source>In Proceedings of the 29th Annual ACM Symposium on Applied Computing, SAC '14</source>
          , pages
          <fpage>243</fpage>
          {
          <fpage>247</fpage>
          , New York, NY, USA,
          <year>2014</year>
          . ACM.
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <given-names>M.</given-names>
            <surname>Pretorius</surname>
          </string-name>
          and
          <string-name>
            <surname>J. van Biljon</surname>
          </string-name>
          .
          <article-title>Learning management systems: Ict skills, usability and learnability</article-title>
          .
          <source>Interactive Technology and Smart Education</source>
          ,
          <volume>7</volume>
          (
          <issue>1</issue>
          ):
          <volume>30</volume>
          {
          <fpage>43</fpage>
          ,
          <year>2010</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [16]
          <string-name>
            <given-names>C.</given-names>
            <surname>Romero</surname>
          </string-name>
          and
          <string-name>
            <given-names>S.</given-names>
            <surname>Ventura</surname>
          </string-name>
          .
          <article-title>Educational data mining: A review of the state of the art</article-title>
          .
          <source>Systems, Man, and Cybernetics</source>
          , Part C:
          <article-title>Applications</article-title>
          and Reviews, IEEE Transactions on,
          <volume>40</volume>
          (
          <issue>6</issue>
          ):
          <volume>601</volume>
          {
          <fpage>618</fpage>
          ,
          <string-name>
            <surname>Nov</surname>
          </string-name>
          <year>2010</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [17]
          <string-name>
            <given-names>I. H.</given-names>
            <surname>Witten</surname>
          </string-name>
          , E. Frank, and
          <string-name>
            <given-names>M. A.</given-names>
            <surname>Hall</surname>
          </string-name>
          .
          <source>Data Mining: Practical Machine Learning Tools and Techniques</source>
          . Morgan Kaufmann Publishers Inc., San Francisco, CA, USA, 3rd edition,
          <year>2011</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          [18]
          <string-name>
            <given-names>M.</given-names>
            <surname>Xenos</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Pierrakeas</surname>
          </string-name>
          , and
          <string-name>
            <given-names>P.</given-names>
            <surname>Pintelas</surname>
          </string-name>
          .
          <article-title>A survey on student dropout rates and dropout causes concerning the students in the course of informatics of the Hellenic Open University</article-title>
          . Computers Education,
          <volume>39</volume>
          (
          <issue>4</issue>
          ):
          <volume>361</volume>
          {
          <fpage>377</fpage>
          ,
          <year>2002</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>