<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <journal-title-group>
        <journal-title>F. Coscia);</journal-title>
      </journal-title-group>
    </journal-meta>
    <article-meta>
      <title-group>
        <article-title>University student dropout prediction for female students in the Computer Engineering career</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Ellen L. Méndez Xavier</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Fabrizio Coscia</string-name>
          <email>fabricoscia@fpuna.edu.py</email>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Christian von Lücken</string-name>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Lorenzo</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Paraguay</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Lorenzo</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Paraguay</string-name>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Department of Computer Science Education, Facultad Politécnica - Universidad Nacional de Asunción)</institution>
          ,
          <addr-line>San Lorenzo</addr-line>
          ,
          <country country="PY">Paraguay</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Department of Computer Science Education, Facultad Politécnica - Universidad Nacional de Asunción</institution>
          ,
          <addr-line>San</addr-line>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>Department of Computer Science Education, Facultad Politécnica - Universidad Nacional de Asunción</institution>
          ,
          <addr-line>San</addr-line>
        </aff>
      </contrib-group>
      <volume>000</volume>
      <fpage>0</fpage>
      <lpage>0002</lpage>
      <abstract>
        <p>University student dropout is a complex phenomenon raising concern both academically and socially. This issue affects students from different areas and contexts; it appears as the premature interruption of higher education, and it has meaningful consequences both on the people involved and on society as a whole. This paper uses prediction models to analyze academic and socioeconomic data about students of the career of Computer Engineering of the Facultad Politécnica of the Universidad Nacional de Asunción (FP-UNA) focusing on women. These models show their effectiveness to predict dropout, and therefore can be useful tools for educational management and to develop preventive actions.</p>
      </abstract>
      <kwd-group>
        <kwd>higher education</kwd>
        <kwd>women</kwd>
        <kwd>gender</kwd>
        <kwd>computing</kwd>
        <kwd>student dropout 1</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>University student dropout is a complex issue that affects students, educational institutions
and national social development as a whole. It is defined as a definite or temporary interruption
of university studies by a student, who will not get the corresponding academic degree or
diploma.</p>
      <p>
        In 2017, the Worl Bank published a study called “Turning Point: Higher Education in Latin
America”. It mentions that only 50% of higher education students get to finish their career and
graduate. It is estimated that students in Latin America and the Caribbean take 36% more time
on average to finish their career in comparison to the rest of the world [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ].
      </p>
      <p>
        University dropout in Paraguay is associated to economic and family problems, poor
academic performance, low motivation and incorrect career choice, among other factors. Also,
most of the students in the country have to combine their studies with some type of labor [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ].
      </p>
      <p>
        Data in [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ], show that in average, less women than men enter university, but once they are
registered, their rate of career completion is higher than that of men. The same behavior has
been seen in a specific study in the context of computer careers in the Facultad Politécnica of
the Universidad Nacional de Asunción (FP-UNA), which indicates that the number of women
in comparison to that of men is lower at the time of entering the university, they are more
effective if we standardize the number to evaluate graduation [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ]. Also, in [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ], we can see that
women are disproportionally underrepresented in these careers.
      </p>
      <p>
        Even though some common challenges have been identified for both men and women, such
as curricular and financial barriers to achieve success, women have their own difficulties due
to cultural elements and gender stereotypes that can affect their academic performance. These
questions may also include home responsibilities and the lack of a family support system while
pursuing their engineering studies [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ] [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ].
      </p>
      <p>This paper takes into consideration academic performance data, as well as socioeconomic
data, in order to evaluate dropout possibilities for women and provide a list of students with
high dropout risk. We have used a number of techniques from Data Mining (DM) to evaluate
these data, to allow us to automatically find the relationships in the data group analyzed. This
process has the potential to improve academic management and, as a consequence, improve
student retention. In Section II, we explain the analysis methods applied. In Section III we
present the results, and then the main conclusions of this paper.</p>
    </sec>
    <sec id="sec-2">
      <title>2. University Dropout.</title>
      <p>From an individual point of view, dropping out means failing to complete a determined
course of action or failing to reach a wanted goal, in pursuit of which the person entered a
higher education institution in particular. Therefore, dropping out depends not only on
individual intentions but also on social and intellectual processes through which people create
the goals they want to achieve in a given university [11]. In other words, Tinto defines dropping
out as a situation a student faces when they aspire to achieve something but cannot manage to
finish their educational project.</p>
      <sec id="sec-2-1">
        <title>2.1. Theoretical Dropout Model.</title>
        <p>According to Tinto, students act according to the exchange theory for building their social
and academic integration, expressed in terms of goals and institutional commitment levels. The
researcher defines the dropout theoretical model based on two main components:
•
•</p>
        <p>Intention and goals: students’ willingness to finish their career.</p>
        <p>Institutional commitment: degree of attachment to or identification with the university
they belong to.</p>
        <p>The interaction with the academic system and the social system of the university modifies
a student’s perception in terms of the two components defined previously, and they influence
the decision of dropping out of or remaining as part of the institution.</p>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>3. Information Processing.</title>
      <p>The Knowledge Discovery in Databases (KDD) process is a methodology used to discover
knowledge from large data groups. This is no trivial process since it allows identifying valid,
original, potentially useful and understandable data patterns.</p>
      <p>
        Figure 1 presents the KDD process. It can be seen that Data Mining is a relevant element in
this process. Likewise, data preparation, selection and cleaning are no less important, together
with the incorporation of previous knowledge and result interpretation. [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ]. These steps, when
applied in an iterative and interactive manner, facilitate extracting useful knowledge from
analyzed data.
      </p>
      <p>This paper has considered the KDD process for data preparation, pre-processing and
transformation from an academic database, and the socioeconomic characterization of students
of the Computer Engineering career in FP-UNA. The data was used as an entry point to apply
different data mining algorithms with the KNIME 2 software. Later, we worked on result
fragmentation and comparisons.</p>
      <sec id="sec-3-1">
        <title>3.1. Algorithms.</title>
        <p>For the proposed data analysis, the selected algorithms are Decision trees, Naïve Bayes and
Random Forest [8]. These algorithms have proven to be useful to make predictions for data
groups that require identifying patterns from different variables such as the university dropout
problem requires with academic and socioeconomic data.</p>
        <p>They were chosen based on the fact that they are widely known, they have low complexity
but are very efficient and also because this is the first time they have been used in our context.
Decision trees and Naive Bayes are often the first algorithms taught to students due to their
simplicity, interpretability and strength in a variety of applications [9].</p>
      </sec>
      <sec id="sec-3-2">
        <title>3.1.1. Decision Tree.</title>
        <p>Tree structure, where every internal node indicates a test on an attribute, every branch
represents a test result, and every leave node (or terminal node) has a class label [12].
2 KNIME is a data mining tool that allows developing models in a visual manner.
https://www.knime.com/knimeanalytics-platform</p>
        <p>Simulating a path for a data group instance is as follows: initially we bear in mind we have
already generated a tree from training data, also the classification instance has the following
characteristics:
•
•
•
•</p>
        <sec id="sec-3-2-1">
          <title>Average entering score: 80</title>
        </sec>
        <sec id="sec-3-2-2">
          <title>Has children: No</title>
        </sec>
        <sec id="sec-3-2-3">
          <title>School modality: Technical</title>
        </sec>
        <sec id="sec-3-2-4">
          <title>Civil status: Single</title>
          <p>According to the tree branches, the nodes are evaluated, and the corresponding label
instance is located.</p>
        </sec>
      </sec>
      <sec id="sec-3-3">
        <title>3.1.2. Naïve Bayes.</title>
        <p>The algorithm is based on a conditional probability concept and aims at giving higher
importance to those events that are really relevant to the data group [10].</p>
        <p>Conditional probability P(A|B) is the probability of event A happening, knowing that
another event B is also happening, and it is expressed as the following equation:
 ( | ) =
 ( ∩  )
 ( )
=
 ( | ) ( )
 ( )</p>
        <p>The assumption of attribute independence is generally a poor assumption and is often
breached for true data groups [13].</p>
        <p>For the simulation of the functioning of this algorithm considering the same instance used
for the decision tree, the two main ones are defined first:</p>
        <p>Then, through training data, the different probabilities for every event are calculated with
the characteristics of the evaluated instance:
•
•
•
•
•
•
•
•
•
•
•
•</p>
        <p>A: The student drops out of the career
B: The student does not drop out of the career
P(A) = 0.4
P(B) = 0.6
P(Average score = 80 | A) = 0.3
P(Average score = 80 | B) = 0.4
P(Has children = No | A) = 0.6
P(Has children = No | B) = 0.8
P(School modality = Technical | A) = 0.5
P(School modality = Technical | B) = 0.6
P(Civil status = Single | A) = 0.7</p>
        <p>P(Civil status = Single | B) = 0.5</p>
        <p>Finally, in order to determine whether the evaluated student is dropping out of the career or
not, the product of the probabilities involved in every case is calculated:
 ( ) = 0.4 ∗ 0.3 ∗ 0.6 ∗ 0.5 ∗ 0.7 = 0.0252
 ( ) = 0.6 ∗ 0.4 ∗ 0.8 ∗ 0.6 ∗ 0.5 = 0.0576</p>
        <p>According to the results, it can be seen that the evaluated student is not dropping out of the
career.</p>
      </sec>
      <sec id="sec-3-4">
        <title>3.1.3. Random Forest.</title>
        <p>In [10] it is mentioned that there are two approximations in the evolution of data mining:
•
•
•</p>
        <p>Developing new algorithms (which happens very occasionally), or adjusting the
parameters of a well-known algorithm.</p>
        <p>Combining classifiers that are kind of simple in order to create a more complex one.</p>
        <p>When the classifiers are decision trees, the resulting algorithm is called Random Forest.</p>
        <p>In order to simulate the functioning of this algorithm, considering the same instance as in
the previous issue, first of all it is assumed that 3 trees are used to form the algorithm and by
combining training data, each of them takes a main attribute as root node. Then the instance is
evaluated by every tree and each one defines if the career is abandoned or not. Finally, the
different decisions are counted and the one with the highest number is the final decision (see
Figure 2).</p>
        <sec id="sec-3-4-1">
          <title>Tree 1:</title>
          <p>"Average entering score" as the main characteristic.
Average score &gt; 70 -&gt; No dropping out</p>
        </sec>
        <sec id="sec-3-4-2">
          <title>Tree 3:</title>
          <p>"Has children" as a main attribute.</p>
          <p>Students without children have lower dropout rates -&gt; No dropping out.</p>
          <p>Taking into consideration the three trees, the student does not drop out of the career.</p>
        </sec>
      </sec>
      <sec id="sec-3-5">
        <title>3.2. Data Classification.</title>
        <p>The data used are divided into three major groups:</p>
        <p>Admission data: Provided by the department of Information Technology and
Communication (ITCs) of FP-UNA, including subject scores required for admission to a
career in the Polytechnical University. Also, there are other data included, such as
gender, home address, telephone number, civil status, city, date and place of birth,
among others.</p>
        <p>Socio-economic data: Provided by the department of Information Technology and
Communication (ITCs) of FP-UNA, the data are gathered through surveys at different
moments of the school year. They include data related to the student´s social and
economic situation. These data include if the student has a job or not, information in
relation to the kind of high school training received, family data, among others.
3. Subject data during the career: It includes the grades of the 52 subjects of the students
of the career.</p>
      </sec>
      <sec id="sec-3-6">
        <title>3.3. Data Characterization.</title>
        <p>In the sample there are 1,676 students who entered the university from 2000 until 2023, for
the Computer Engineering career, with the gender distribution seen in Figure 2. Since the focus
of this paper is analyzing female student dropout, we will be looking at the corresponding data
for women (373 students).</p>
        <p>Men
80%</p>
        <p>Women</p>
        <p>Men</p>
        <p>Women
20%</p>
        <p>The career admission process requires passing a number of examinations of basic knowledge
for the career and getting the highest scores until taking the number of seats available. The data
from these evaluations are used to obtain the admission average. A standardization is made
considering the subjects evaluated for admission present yearly variations, thus applying the
Admission Average.</p>
        <p>For the group of data referring to student socioeconomic information, the following
attributes were selected considering the possible values indicated below.</p>
        <p>Finally, for academic data, we looked at all the Computer Engineering career subjects except
the Final Dissertation Paper (FDP). Together with the Department of Academic Statistics of
FPUNA, it was decided that the average of all of a subject’s attempts should be taken into account
in order to reflect the student’s whole academic process.</p>
        <p>In order to identify the characteristics of those students who have graduated from the career,
the Academic Department provided a subgroup of academic data from 265 Computer
Engineering students who have already finished their studies. This way, the GRADUATION
attribute was created, and the dropout concept was applied as explained in the following
paragraph.</p>
      </sec>
      <sec id="sec-3-7">
        <title>3.4. Data Pre-processing.</title>
        <p>Before introducing data in the KNIME tool, they went through preprocessing for
information cleaning and optimization, so that it can be used better in the mentioned software.</p>
        <p>Initially, there were 6 files with the following characteristics:
•
•
•
•
•
•
21042023_nota_IIN.csv: this includes students’ grades and other data about those
grades.
21042023_nota_retroactiva_IIN.csv: the same as the previous file but it has
students’ older grades.
28042023_nota_IIN_convalidaciones.csv: it has validation grades of students who
have changed their career.
puntaje ingreso.csv: it has data related to student admission, besides their grades, it
also includes sone socioeconomic data.
socioeconomica.csv: it has the answers to the socioeconomic survey carried out by
FP-UNA.
03042023_egresados_IIN.csv: it is a list of Computer Engineering students who
have finished their career successfully.</p>
        <p>The data treatment process was the following:
1. We took the file 21042023_nota_IIN.csv and for every record, we grouped the grades
per student, bearing in mind that a student may have more than one grade per subject
in case such student failed such subject. Then we repeated the process taking files
21042023_nota_retroactiva_IIN.csv and 28042023_nota_IIN_convalidaciones.csv.
2. Then we processed the file puntaje_ingreso.csv by assigning admission scores to the
corresponding students. Next, we carried out a similar process with the file
socioeconomica.csv.
3. The last file, 03042023_egresados_IIN.csv, was used to add the GRADUATION
column and identify those students who have already finished their career.
4. Once all the files were processed, we started to apply the dropping out definition set in
order to identify students who are presumed career dropouts.
5. The last step was grouping all the data in a final file that has socioeconomic data,
admission average grades, and average grades per subject.</p>
      </sec>
      <sec id="sec-3-8">
        <title>3.5. Dropout Definition at FP-UNA.</title>
        <p>For the Universidad Nacional de Asunción there is no formal definition for school dropout.
In fact, the bylaws do not mention dropping out at any moment as a matter of discussion or
analysis. In order to come closer to a definition of dropping out in FP-UNA, the Coordination
of Academic Statistics considers dropping out as a student’s lack of registration for four
consecutive periods (two years).</p>
      </sec>
      <sec id="sec-3-9">
        <title>3.6. Other considerations.</title>
        <p>Data volume in relation to academic information is a lot larger than that of socioeconomic
data, also considering that such data are gathered through surveys which some students decide
not to answer since they are not mandatory. This makes it necessary to evaluate the impact of
combining both groups. Therefore, we applied a strategy of evaluating academic data alone
independently and later, academic and socioeconomic data, in order to compare their impact on
the application of data mining techniques.</p>
      </sec>
      <sec id="sec-3-10">
        <title>3.7. Model Validation.</title>
        <p>The confusion matrix presents a graphic vision of errors made by the classification model in
a table. It is a graphic model to visualize the level of correctness of a prediction model. In
literature, it is also known as a contingency table or error matrix [10].</p>
        <p>P
N</p>
        <sec id="sec-3-10-1">
          <title>True Class</title>
          <p>Predicted Class</p>
          <p>P N
TP FN
FP</p>
          <p>TN</p>
          <p>In essence, this matrix indicates the number of instances that have been correctly and
incorrectly classified. The parameters indicated are:
•
•
•
•</p>
          <p>True Positive (TP): number of correct classifications of the positive kind (P).
True Negative (TN): number of correct classifications of the negative kind (N).
False Negative (FN): number of incorrect classifications of the positive kind classified
as negative.</p>
          <p>False Positive (FP): number of incorrect classifications of the negative kind classified
as positive.</p>
          <p>From the confusion matrix, we have defined a group of metrics that allow quantifying the
goodness of a classification model. An error (ERR) is the sum of incorrect predictions on the
total number of predictions. On the contrary, accuracy (ACC) is the number of correct
predictions on the total number of predictions, as the following equation shows:
ERR =</p>
          <p>FP + FN</p>
          <p>FP + FN + TP + TN
ACC =</p>
          <p>TP + TN
FP + FN + TP + TN
= 1 − ERR</p>
        </sec>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>4. Results Obtained</title>
      <p>For the application of the different algorithms, the data were divided into two subsamples
at random. The first one with 50% of records which were used as training data, and the second
one with the other 50% of records which were used to test the prediction ability of the model.</p>
      <p>Next, the results of the application of the 3 algorithms with their validation data through a
confusion matrix technique are shown, both for academic data processing and for academic and
socioeconomic data processing focused on female population.</p>
      <sec id="sec-4-1">
        <title>4.1. Decision Tree</title>
        <p>With this algorithm, taking only academic data, we classified 81 instances correctly, while
42 instances were classified incorrectly. This implies a 65.32% correctness rate. In relation to
academic and socioeconomic data, the correctness rate is 70%.</p>
        <p>The results of the confusion table including the results of academic data analysis and
academic and socioeconomic data analysis can be seen in Table 2.</p>
      </sec>
      <sec id="sec-4-2">
        <title>4.2. Naïve Bayes</title>
        <p>We used this algorithm to classify 96 instances correctly, while 28 instances were classified
incorrectly. This implies a 77.41% correctness rate for academic data, and a 78.35% rate for
socioeconomic and academic data. The results of the confusion table can be seen in Table 3.</p>
      </sec>
      <sec id="sec-4-3">
        <title>4.3. Random Forest</title>
        <p>With this algorithm, we classified 107 instances correctly, while 17 instances were classified
incorrectly. This implies an 89.29% correctness rate for academic data and 84.67% with the
inclusion of socioeconomic data. The results of the confusion data can be seen in Table 4.</p>
        <p>Based on the results obtained in the correctness rate of the algorithms, Figure 3 shows a
comparison of the 3 algorithms both for analyzing only academic data and for analyzing
academic and socioeconomic data.</p>
        <p>seRandom Forest
u
q
i
n
h
c
tge Naive Bayes
n
i
s
s
e
c
rPo Decision Trees
0
20
40 60</p>
        <p>Correctness rate
Academic + Socioeconomic</p>
        <p>Academic
80
100</p>
        <p>By analyzing the data carefully, the algorithms have identified 123 students with dropout
probabilities and have confirmed the validation data. Table 5 shows numbers identified by
algorithm in relation to data group.</p>
        <sec id="sec-4-3-1">
          <title>Academic</title>
          <p>67
69
69</p>
        </sec>
        <sec id="sec-4-3-2">
          <title>Academic +</title>
        </sec>
        <sec id="sec-4-3-3">
          <title>Socioeconomic</title>
          <p>77
57
70
In the lists provided by the algorithms, there are 19 students who appear in all of them.</p>
          <p>The maximum number of coincidences is the appearance of 19 students in all the algorithms
executed. In the results of algorithms applied to the academic data group, we see that 52 out of
80 students show up in the 3 lists. In relation to academic data including socioeconomic data,
42 out of 84 students appear in all 3 lists.</p>
        </sec>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>5. Conclusions.</title>
      <p>In this analysis, we have studied the application of data mining algorithms to predict
university dropout, taking into consideration two different stages: one based exclusively on
academic data, and another one that includes socioeconomic variables, applied to a data group
associated to women exclusively.</p>
      <p>Even though all the algorithms were able to identify students who might drop out of the
career properly, none of them got the whole group. Also, the group of algorithms was not able
to identify all real dropout cases. The existence of a number of students who have been
identified by all the algorithms as having dropout probabilities is relevant, other students were
identified by 2 of them only. This suggests that students in the result lists could be classified
using the number of algorithms as criteria to spot them as students with dropout possibilities,
in order to prioritize assistance by the academic manager.</p>
      <p>We have seen that we were able to get useful results, but necessary improvements could be
achieved by using other algorithms, a higher number of data, different data preprocessing
strategies, among others. Even though it is easy to think that a higher number of variables to
be analyzed could lead to better predictive capacity, the existing socioeconomic data did not
improve results.</p>
      <p>For future work, we expect to be able to carry out a detailed characterization analysis of the
results obtained in order to identify common patterns and higher repetition of noticeable
characteristics. Also, it would be important to carry out an institutional analysis on the current
situation of students that have confirmed their dropping out and make a validation of their
reasons for doing that.</p>
    </sec>
    <sec id="sec-6">
      <title>Special Thanks</title>
      <p>To the Coordination of Statistics of the Academic Direction of FP-UNA, for their support
during the preparation of this analysis and the approximation to definitions of aspects that have
no institutional definition.
prediction in a Chilean public university through classification based on Decision Trees with
optimized parameters, University Training, issue 11, number 3, 2018.)
[7] A. Lourens and D. Bleazard. Applying predictive analytics in identifying students at
risk: A case study. South African Journal of Higher Education, vol. 30, no. 2, pp. 129–142,
2016.
[8] Pranckevičius, T., &amp; Marcinkevičius, V. (2017). Comparison of Naive Bayes, Random</p>
      <sec id="sec-6-1">
        <title>Forest, Decision Tree, Support Vector Machines, and Logistic Regression</title>
      </sec>
      <sec id="sec-6-2">
        <title>Classifiers for Text Reviews Classification. Balt. J. Mod. Comput., 5.</title>
        <p>[9] Witten I., Frank E., &amp; Hall M. (2011). Data Mining: Practical Machine Learning Tools
and Techniques. Morgan Kaufmann.
[10] Roma, J. C., Quiles, R. C., Roig, J. G., y Alfonso, J. M. (2017). Minería de datos: Modelos
y algoritmos. Editorial UOC, S.L. (Data mining: Models and algorithms.)
[11] Tinto, V. (1989). Definir la deserción: Una cuestión de perspectiva. Revista de
Educación Superior. 18(71). (Defining dropout: A matter of Perspective. Higher Education
Magazine.)
[12] Maimon, O. y Rokach, L. (2005). Data Mining and Knowledge Discovery Handbook.</p>
        <p>Springer.
[13] Mosquera, R., Castrillón, O., y Parra, L. (2018). Máquinas de soporte vectorial,
clasificador naive bayes y algoritmos genéticos para la predicción de riesgos
psicosociales en docentes de colegios públicos colombianos. Información
tecnológica, 29(6):153–162 (Vectoral support machines, naive bayes classifier and genetic
algorithms for psychosocial risk prediction in teachers from Colombian public schools.
Technologic information.)</p>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <surname>Mundial</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          (
          <year>2017</year>
          ).
          <article-title>Momento decisivo: La educación superior en América Latina y el Caribe</article-title>
          .
          <source>Direcciones en Desarrollo</source>
          , Washington. (
          <article-title>Decisive moment: Higher education in Latin America and the Caribbean</article-title>
          .
          <source>Developing Guidelinnes</source>
          , Washington.)
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <surname>I. Acuña</surname>
          </string-name>
          , “Acceso y permanencia de estudiantes de educación superior en la ciudad de Pilar”
          <article-title>Ciencia Latina Revista Científica Multidisciplinar</article-title>
          , vol.
          <volume>5</volume>
          , no.
          <issue>6</issue>
          , pp.
          <volume>12</volume>
          <fpage>372</fpage>
          -
          <lpage>12</lpage>
          384,
          <year>2021</year>
          .
          <article-title>(”Access and stay of higher education students in Pilar City” Multidisciplinary Scientific Magazine Latin Science</article-title>
          )
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <surname>Madar</surname>
            ,
            <given-names>N.K.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>Danoch</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>Morera</surname>
            ,
            <given-names>L.S.</given-names>
          </string-name>
          (
          <year>2022</year>
          ).
          <article-title>Dropouts of women in engineering studies: A comparable case</article-title>
          .
          <source>Journal of Entrepreneurship Education</source>
          ,
          <volume>25</volume>
          (
          <issue>1</issue>
          ),
          <fpage>1</fpage>
          -
          <lpage>10</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <surname>Méndez</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          ,
          <string-name>
            <surname>von Lücken</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          , &amp;
          <string-name>
            <surname>Cantero</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          (
          <year>2022</year>
          ).
          <article-title>Applications, admissions and graduations of women in computer science careers for the universidad Nacional de Asunción</article-title>
          . Proceedings http://ceur-ws.
          <source>org ISSN</source>
          ,
          <volume>1613</volume>
          ,
          <fpage>0073</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>O.</given-names>
            <surname>Maimon</surname>
          </string-name>
          and
          <string-name>
            <given-names>L.</given-names>
            <surname>Rokach</surname>
          </string-name>
          ,
          <source>Data Mining and Knowledge Discovery Handbook</source>
          . Springer,
          <year>2005</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>P.</given-names>
            <surname>Ramírez</surname>
          </string-name>
          and
          <string-name>
            <given-names>E.</given-names>
            <surname>Grandón</surname>
          </string-name>
          .
          <article-title>Predicción de la deserción académica en una universidad pública chilena a través de la clasificación basada en Árboles de decisión con parámetros optimizados, Formación universitaria</article-title>
          , vol.
          <volume>11</volume>
          , no.
          <issue>3</issue>
          ,
          <year>2018</year>
          . (Academic droput
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>