<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>mining over open data for a longitudinal assessment of municipal public education in Brazil</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Arthur Scanoni</string-name>
          <email>arthur.scanoni@ufrpe.br</email>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Rogério Silva Filho</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Paulo Adeodato</string-name>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Kellyton Brito</string-name>
          <email>kellyton.brito@ufrpe.br</email>
        </contrib>
      </contrib-group>
      <pub-date>
        <year>2022</year>
      </pub-date>
      <abstract>
        <p>The popularization of large-scale assessments in education, such as PISA in Europe and ENEM in Brazil, and the availability of their related data on contemporaneous government open data repositories, have fostered the creation of public value. Much of the current research aims to analyze education at a student level and focuses on high school education, and few studies on earlier educational stages at the municipal level can be found. This paper presents a study focused on analyzing Brazilian educational open data for assessing public education on a municipal level for elementary and middle school education to understand the correlations between contextual indexes and expenditures and educational achievement. For this, we have applied a data mining-based approach and statistical methods to correlate the features and student performance from 2013 to 2019. The main educational results indicate that the highest positive correlations with students' performance are related to teacher training, followed by financial investments. With regard to the Brazilian open data scenario, certain improvements were identified, such as a large amount of available, updated data. However, historical challenges such as a lack of standards and available machine-readable data are still present.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>
        Since the 1950s [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ], governments and society have agreed that transparency, ”the right to know,”
and Open Government Data (OGD), which is verified and used by the general population, may
bring about many benefits, such as increased accountability and citizen participation. At the
beginning of the 2010s, the movement resurfaced with the possibility of using Web 2.0 to publish
and consume data, and many initiatives for publishing open data portals were launched. Thus,
new benefits such as delivering better public services and increasing government eficiency
and efectiveness have been indicated, especially due to the possibility of society analyzing the
publicized data for value generation [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]. These benefits are strongly related to the educational
context. Since the 1960s [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ], research using educational data aimed at the efects of policies
has become standard in educational assessment, and the popularization of OGD has provided a
boost to this kind of research, thereby promoting the potential of the findings.
(K. Brito)
      </p>
      <p>
        Within this context, one important data source is related to large-scale educational assessment
(LSA). The databases of these exams are a relevant data source for scientific studies, and enable
governments to plan, define and validate educational policies. Although exams at higher
educational stages are mostly used to analyze educational performance in LSA [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ], elementary
and secondary stages, when children learn to read and write, and develop their critical and
logical thinking, play an essential role in this process. Moreover, most studies either analyze
data on the student level, aiming to analyze and predict student performance [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ]. Despite being
of value, this analysis nonetheless has limits regarding the use of results for understanding and
creating policies and actions at higher levels, such as a municipal or state level. Additionally,
being able to characterize the relationship between student performances at municipal/state
levels and government policies and expenditures at these levels could identify any gaps and
thereby help to support the creation of new policies and validations of those that already exist
[
        <xref ref-type="bibr" rid="ref5">5</xref>
        ].
      </p>
      <p>Within this context, this paper aims to use the Brazilian open educational data from
elementary and secondary education on a municipal level in order to understand the relationship
between contextual indexes and expenditures, and Brazilian educational achievements from a
longitudinal perspective. We have specifically considered the period from 2013 to 2019, and
have applied a data mining process and classical statistical methods to seek correlations that
could provide a national perspective on how the characteristics and investments of
municipalities are correlated with the performance of early-stage students within the complex Brazilian
educational system. The analysis includes educational results and a discussion regarding the
challenges of the current Brazilian open data scenario.</p>
      <p>The remainder of this paper is organized as follows. Section 2 presents the background to the
subject. Section 3 presents the methodology, including the research questions and hypotheses,
and an overview of the data mining process carried out to answer the questions. Section 4
presents the experiment, followed by Section 5, which contains the results and evaluation.
Lastly, Section 6 presents the concluding remarks.</p>
    </sec>
    <sec id="sec-2">
      <title>2. Background and Related Works</title>
      <p>
        Understanding the power of educational policies and the performance of educational systems is
a permanent research topic [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ]. The research arose as a reaction against the pioneering Coleman
study [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ], which was the first to identify the predominance of student backgrounds in their
outcomes and how factors related to schools and policies have a lesser influence. The essential
characteristic that diferentiates the current from earlier research is the vast amount of available
open data on the educational process [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ]. One important branch of research is studying the role
played by budgets in the educational system. Although there is a sense that more investment
implies better outcomes, some studies, such as [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ], have shown that, at least for the higher
economies, the manner in which resources are used is more important than the amount spent.
In fact, some educational systems, as in the case of Brazil, have increased the budget without,
however, changing the outcomes [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ].
      </p>
      <p>By considering the Brazilian scenario, some studies have attempted to understand the
relationship between contextual variables and educational achievement from Brazilian educational
open data, such as [9, 10]. They have often relied upon statistical predictive models through
the use of regression coeficients and feature importance in order to unravel how variables
are correlated with school achievement. They have also focused on predicting the academic
performance of individual students in a specific exam. However, these studies often fail to
discuss their model assumptions and how this could verify them in practical scenarios. Also,
there is no consensus on an appropriate set of variables to be included in their models. Choices
are often arbitrary and may sufer certain influences, mainly due to the lack of standards and
the decentralized Brazilian open data scenario, as discussed in Subsection 5.3. Moreover, very
few studies regarding the Brazilian scenario have explored the expenditure variables, mainly
because this information is not easily gathered. In this direction, [11] and [12] found that
expenditure has no or little contribution to increasing test scores, but [13] found that ”financial
resources are paramount in producing performance.”</p>
      <p>Analyzing the studies, it is clear that there are two types of research. One group used the
student and school characteristics to analyze student and school performances, and the other
group used financial and expenditure data to analyze the municipal performance. However,
some school characteristics may be aggregated to a municipal level, such as teacher data. Hence,
these characteristics may be studied jointly with financial data to provide a new viewpoint,
capable of supporting policies at diferent levels.</p>
    </sec>
    <sec id="sec-3">
      <title>3. Methodology</title>
      <p>The main objective of this research is to use open and publically available data to analyze the
characteristics of municipalities and school management, seeking to find correlations between
these characteristics and the performance of public-school students. For this, we defined three
research questions: RQ1: Is it possible to define and apply a data mining methodology over open
data to find correlations between the main characteristics and investments of municipalities and
student performance in large-scale assessments in Brazil?; RQ2: What are the main correlations
between the characteristics and investments of municipalities and student performance in
largescale assessments?; and RQ3: How have the correlations between the characteristics and investments
of municipalities and student performance varied over time?</p>
      <p>To answer RQ2 and RQ3, we defined a DM process inspired by the well-known CRISP-DM
method [14]. Thus, if these questions can be answered, it implies that RQ1 is also answered.
The defined process contains five phases: (i) domain understanding; (ii) data collection and
understanding; (iii) data preparation; (iv) modeling and analysis; and (v) results reporting
and evaluation. Section 4 presents the experiments, with details and implementation of the
methodology, from domain understanding to modeling and analysis. The last step is presented
in Section 5.</p>
    </sec>
    <sec id="sec-4">
      <title>4. Experiments</title>
      <p>Domain Understanding: In Brazil, when considering the teaching stage, they may be
elementary school (1st to 9th grades) and high school (10th to 12th grades). The elementary school is
divided in two stages: initial years (1st to 5th grades) and final years (6th to 9th grades). For a
clearer comprehension, in this paper we have named the stages as elementary school, middle
school, and high school. With a few minor exceptions, it is mainly the municipalities that are
responsible for elementary and middle schools, while state governments are responsible for
high schools, and the federal government is responsible for universities. The private sector may
be present at any level, but it is not the subject of this study. The most studied LSA in Brazil
is the ENEM, which is used for entry to most universities. However, the assessment of early
education is still underexplored. For this, there is the IDEB, an education development index. It
is based on two main metrics: each school’s pass and abandonment rates and the performance
of students in a specific exam, called ”Prova Brasil”. The IDEB has three distinct indexes, one for
elementary school, one for middle school, and one for high school. Thus, in order to evaluate
municipal education, which is responsible for elementary and middle schools, we have used the
respective IDEB index for these two stages.</p>
      <p>Data Collection and Understanding: We have focused data collection on the granularity of
municipalities, rather than students, and relevant data were divided into two groups, educational
data and economic data. The educational data, including school data and IDEB indicators, were
collected from the INEP portal [15]. Economic data were collected from the National Treasury
website and the SIOPE website [16]. Figure 1 presents a list of the collected data used as features.
In addition, the performance indicator, IDEB, was also collected from the INEP portal. Most
features presented in Figure 1 have two values, one for the elementary stage and another for
the middle stage. The corresponding value will be used in the analysis of each stage. Also,
some of the features are continuous and may be directly understood, such as the number of
students per class (SPC). However, some of the features are divided into subcategories and
present discrete values for each of the subcategories. These categories are defined according to
metrics defined by the Brazilian government. The ATT classifies schools into 5 groups, ranging
from G1 (percentual of teachers who have a teaching degree in the same subject they teach)
to G5 (teachers without a college degree). The SMC ranges from G1 (low complexity) to G6
(high complexity)., TEI ranges from G1 (low efort) to G6 (high efort), and TRI ranges from G1
(low regularity) to G6 (high regularity). As the IDEB is calculated every two years, data were
collected for the years 2013, 2015, 2017 and 2019 for each of the 5,568 Brazilian municipalities.
Each feature presented in Figure 1 is published as a diferent dataset for each year. The Brazilian
government provides all of the indicators over its open data portals already listed.</p>
      <p>Data Preparation: INEP data is in spreadsheet format (XLSX and ODS) and, despite being
complete, contains unnecessary data, such as a header with the INEP logo and merged lines and
columns. Hence, we manually cleaned spreadsheets before they could be processed. INEP data
is also usually presented in diferent granularities, such as aggregations of the
country/regions/states, aggregations of municipalities, and individual schools. It also contains data for each
separate stage, such as elementary, middle, and high school. For the objectives of this study, we
ifltered and used the aggregations by municipalities, and data regarding the elementary and
middle stages. The National Treasure data are presented with two values, the value initially
transferred plus a correction that may be available later. In the study, we considered the sum of
both as the absolute value of spending. For SIOPE data, it was not necessary to edit the data
since each one has only the value of each municipality. Fortunately, municipalities are identified
by a unique and equal ID across the diferent datasets. Thus, after cleaning and filtering the
data from the three diferent sources, it was possible to create a unique database joining all of
them together.</p>
      <p>Modeling and Analysis: All features were analyzed individually and correlated with the
two performance features: IDEB elementary school (IDEB-E) and IDEB middle school (IDEB-M).
We used the Spearman’s correlation, which evaluates the monotonic relationships between
the variables. It is most suitable in this context since it is not afected by global changes, for
example, an increase in IDEB results due to changes in methodology or an overall improvement
of education in the country [17]. It is well known that results in exams are the result of several
factors, including a diversity of topics such as financial status, parents’ level of scholarity and
personal efort, as presented in Section 2. Thus, it not expected to find that municipality indexes,
characteristics or investments in education present high correlation values. There are also
several types of interpretations for the correlation values, and we decided to follow the same
interpretations as the Mukaka study [18], i.e., to be a considerable correlation, the absolute
value must be greater than 0.3. Moreover, due to the expected low values, we also present and
analyze the variable with at least the 5 highest positive and negative correlations.</p>
      <p>In order to answer RQ2, we calculated the correlations of every feature for every year to
analyze which features were relevant in each year. We then averaged the results for each
feature to have a consistent value among the years. Next, to answer RQ3, we calculated a linear
regression on the values over the four years, and identified the percentage of the slope compared
to the mean of the values. The result is the assessment of the strength, or ”the velocity”, in
which the correlation increases or decreases during the period.</p>
    </sec>
    <sec id="sec-5">
      <title>5. Results and Evaluation</title>
      <sec id="sec-5-1">
        <title>5.1. Correlation results and discussion</title>
        <p>As planned, we calculated the Spearman correlations coeficient for each feature in relation to
the performance indicators, IDEB-E and IDEB-M, and averaged the results obtained over the
years. Unsurprisingly, most of the features with high correlations, both positive and negative,
are the same for both elementary and middle schools. Figure 2 presents the data.</p>
        <p>When considering the positive correlations, the data indicated that student performances in
a municipality were better when the teachers were more highly qualified: the percentage of
teachers with higher education (THE) and the adequacy of teacher training (ATT) at a higher
level (G1), presented the higher correlations. It is followed by investment: first the investment
per student (INVS), then the investment per teacher (INVT). On the other hand, the feature
presenting a high negative correlation in both scenarios was the age-grade distortion (AGD), i.e.,
in a municipality where a high number of students was out of the expected class, the educational
performance was worse. It is important to note that, as correlation is not causality, it may not
be assured whether the age-grade distortion is the cause or, most probably, the consequence of a
bad student performance. The second higher negative correlation is the opposite of the positive
correlation, the number of teachers with no college degree (ATT_G5), followed by the number
of schools with a high management complexity (SMC_G4 and SMC_G5). Lastly, also associated
with the teacher, (TRI) at the lower level, representing municipalities where the turnover is
high, also presented a negative correlation with student performance.</p>
        <p>As planned, we also analyzed the behaviour of features over time. Taking the correlation
values for each year, considering both the IDEB-E and IDEB-M, we calculated the slope of their
linear regression and analyzed the trend of each feature over time. With this, we calculated the
percentage of variation in relation to the mean, thereby obtaining the tendency if the correlation
is consistently increasing or decreasing. To be able to get more insights, we broadened the
analysis all variables with p &gt; 0.10. Figures 3 and 4 present the data.</p>
        <p>By analyzing data in Figures 3 and 4, most of the features present a non-monotonic variation,
i.e., they neither strictly increase nor decrease, and do not strictly present a trend. Also, none
of the five features with the highest correlations (positive or negative) presented a percentage
slope variation higher than 10%, demonstrating that the most correlated features remained
stable. However, some trends may be found, allowing additional analysis. The percentage of
schools in a city (SMC) with a low management complexity (level 2 of 6) presented a monotonic
increase in relation to IDEB-E, suggesting that having more schools with less students (50-300)
is increasing its importance. Also, the TRI_G1 presented an absolute monotonic decrease. This
signifies that, if the trend continues, the negative correlation of the teacher regularity indicator
(TRI) will soon no longer exist. Additionally, the result of a positive correlation of TEI_G4,
a high efort group, consistently increasing, in opposition to the negative correlations of the
TEI at lower levels, was unexpected. This result must be carefully assessed in further studies,
although an initial reason may be put forward. The lower level groups include teachers working
only one daily period, which may indicate that teaching may not be their primary activity.</p>
        <p>Presented data answers RQ2 and RQ3 by using the defined data mining methodology over
open data, thereby answering RQ1 afirmatively.</p>
      </sec>
      <sec id="sec-5-2">
        <title>5.2. Open data discussion</title>
        <p>Four years after being one of the founders of the Open Government Partnership (OGP), [ 19]
indicated some of the challenges of the Brazilian scenario. We may summarize these challenges
at that time as: (i) a lack of available data, because of the small quantity of datasets published in
the repositories; (ii) multiple and decentralized data sources, because although a national
opendata portal exists, it is incomplete and there are many other state and municipal repositories
not integrated within it; (iii) zombie data, without an update policy; (iv) a lack of standards for
publishing, because each variety of publisher chooses what and how to publish, as well as the
data format; and (v) one-way data, not allowing citizens to return data to the government.</p>
        <p>In this study we have verified that, considering the current state of educational open data,
some of these challenges have been solved, and others still remain. As improvements, we
highlight that there is no lack of available educational data, since we were able to collect the
data easily with regard to all municipalities in Brazil. Also, we cannot consider the data as
zombie data, because despite a certain amount of delay in publishing, since they are released
annually, all the expected and updated data were found. As challenges, we indicate that there
are still multiple and decentralized data sources and a lack of standards for publishing. Even
being published by the Federal government, the gathered data was not found in the national
open data repository, but on three diferent, independent sources, each with a diferent format
and standards. In particular, the data gathered from SIOP is not even found on an open or
transparency portal: it was gathered from a system that must be operated by an individual
who has to access the ”management reports” option, filter the desired data, and then performs
the download. Moreover, most of the collected data is ”almost” machine readable, one of the
premises of open data. INEP data is in spreadsheet format (XLSX and ODS) and, despite being
complete, it contains unnecessary data, such as a header with INEP logo and a variety of merged
lines and columns. Hence, it is necessary to manually clean the spreadsheets before they may be
processed. Lastly, almost all the collected data can not be found or downloaded without human
intervention, and was not available to be accessed from an API (Application Programming
Interface).</p>
        <p>Considering these challenges, we argue that many improvements still need to be implemented
in order to enhance the Brazilian open data scenario and foster the potential of using government
data.</p>
      </sec>
    </sec>
    <sec id="sec-6">
      <title>6. Concluding Remarks</title>
      <p>This paper has presented a data mining approach over educational open data for a longitudinal
assessment of municipal public education in Brazil, by finding correlations among the
characteristics and investments of municipalities and student performances in large-scale assessments.
For this, we defined a methodology which consisted of five steps: (i) domain understanding; (ii)
data collection and understanding; (iii) data preparation; (iv) modeling and analysis; and (v)
results reporting and evaluation. Data was collected from Brazilian open data portals, including
educational and financial data, and data from the performance of students at a municipal level,
including the 5,568 municipalities for the period from 2013 to 2019. Spearman correlations were
calculated in order to find correlations, and the regression slope over the period was calculated
and analyzed in order to verify the trends.</p>
      <p>The main results indicate that the highest positive correlations are linked to teacher training,
followed by financial investments in students and teachers. On the other hand, the age-grade
distortion presents a high negative correlation. All of the highest correlations remained stable
over time, but correlations regarding the lower level of school management complexity and the
number of full-time teachers present a positive trend.</p>
      <p>Some improvements were also identified in the Brazilian open data scenario, such as a
large amount of available, updated data. However, historical challenges, such as multiple,
decentralized data sources, the lack of standards for publishing, and the existence of datasets
that are not machine readable, and thus require human intervention, are still challenges.</p>
      <p>We argue that correlation is not causality. Also, we recognize that many other features may
be correlated and influence student performance, not only those studied. Thus, the results of
this study must be analyzed with caution. Future work may focus on identifying why some
features present a higher correlation with student performance, and how public policy may
be modeled to boost this performance. Furthermore, Brazil is a large country, well known for
having regional diferences, but data from all municipalities were analyzed together. Thus,
future studies focused on identifying the regional diferences may be suitable, as well as studies
comparing Brazilian results with results of other countries, especially in Latin America. Lastly,
future studies using multivariate correlations and machine learning models for simulating the
prediction of student performance, based on the municipalities features, may be promising.
[9] C. Gomes, E. Jelihovschi, Presenting the regression tree method and its application in a
large-scale educational dataset, Int. J. Res. Method Educ 43 (2020) 201–221. doi:10.1080/
1743727X.2019.1654992.
[10] R. Filho, P. Adeodato, K. Santos Brito, Interpreting classification models using feature
importance based on marginal local efects, 2021.
[11] M. Haddad, R. Freguglia, C. Gomes, Public spending and quality of education in brazil, J.</p>
      <p>Dev. Stud 53 (2017-10) 1679–1696. doi:10.1080/00220388.2016.1241387.
[12] A. Santos, F. Medeiros, Relationship of federal funding to ideb results in a state in brazil:
an approach based on educational data mining, in: 2020 15th Iberian Conference on
Information Systems and Technologies (CISTI, 2020, p. 1–4. doi:10.23919/CISTI49556.
2020.9140924.
[13] A. R. M. Valle, R. Gomes, Analyzing the importance of financial resources for
educational efectiveness, Int. J. Product. Perform. Manag 63 (2014-01) 4–21. doi: 10.1108/
IJPPM-08-2012-0085.
[14] C. Shearer, The crisp-dm model: the new blueprint for data mining, J. data Warehous 5
(2000) 13–22.
[15] Instituto nacional de estudos e pesquisas educacionais anísio teixeira, “dados abertos
inep, 2022. URL: https://www.gov.br/inep/pt-br/acesso-a-informacao/dados-abertos.
[16] Fundo Nacional Educação, Sistema de informações sobre orçamentos públicos em educação,
2022. URL: https://www.fnde.gov.br/siope/relatorio-gerencial/dist/indicador.
[17] P. Schober, C. Boer, L. Schwarte, Correlation coeficients, Anesth. Analg 126 (2018-05)
1763–1768. doi:10.1213/ANE.0000000000002864.
[18] M. Mukaka, A guide to appropriate use of correlation coeficient in medical research,</p>
      <p>Malawi Med. J 24 (2012).
[19] K. S. Brito, M. S. Costa, V. Garcia, S. L. Meira, Is brazilian open government data actually
open data? an analysis of the current scenario, Int. J. E-Planning Res 4 (2015-04) 57–73.
doi:10.4018/ijepr.2015040104.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>W.</given-names>
            <surname>Parks</surname>
          </string-name>
          ,
          <article-title>Open government principle: Applying the right to know under the constitution</article-title>
          ,
          <source>Georg. Wawhingt. Law Rev</source>
          <volume>26</volume>
          (
          <year>1957</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>M.</given-names>
            <surname>Janssen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Charalabidis</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Zuiderwijk</surname>
          </string-name>
          ,
          <article-title>Benefits, adoption barriers and myths of open data and open government</article-title>
          ,
          <source>Information Systems Management</source>
          <volume>29</volume>
          (
          <year>2012</year>
          )
          <fpage>258</fpage>
          -
          <lpage>268</lpage>
          . doi:
          <volume>10</volume>
          .1080/10580530.
          <year>2012</year>
          .
          <volume>716740</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>J.</given-names>
            <surname>Coleman</surname>
          </string-name>
          , Equality of educational opportunity,
          <source>Equity Excell. Educ</source>
          <volume>6</volume>
          (
          <year>1968</year>
          )
          <fpage>19</fpage>
          -
          <lpage>28</lpage>
          . doi:
          <volume>10</volume>
          .1080/0020486680060504.
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>S.</given-names>
            <surname>Salloum</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Alshurideh</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Elnagar</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Shaalan</surname>
          </string-name>
          ,
          <article-title>Mining in educational data: Review and future directions</article-title>
          ,
          <source>in: Proceedings of the International Conference on Artificial Intelligence and Computer Vision</source>
          , volume
          <volume>AICV2020</volume>
          ,
          <year>2020</year>
          , p.
          <fpage>92</fpage>
          -
          <lpage>102</lpage>
          . doi:
          <volume>10</volume>
          .1007/ 978- 3-
          <fpage>030</fpage>
          - 44289-
          <issue>7</issue>
          _
          <fpage>9</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <surname>O.E.C.D.</surname>
          </string-name>
          ,
          <article-title>Does money buy strong performance in pisa?</article-title>
          ,
          <source>PISA Focus 13</source>
          (
          <year>2012</year>
          )
          <article-title>4</article-title>
          ,.
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>D.</given-names>
            <surname>Hernández-Torrano</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Courtney</surname>
          </string-name>
          ,
          <article-title>Modern international large-scale assessment in education: an integrative review and mapping of the literature</article-title>
          ,
          <source>Large-scale Assessments Educ</source>
          <volume>9</volume>
          (
          <issue>2021</issue>
          -12)
          <fpage>17</fpage>
          ,. doi:
          <volume>10</volume>
          .1186/s40536- 021- 00109- 1.
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>C.</given-names>
            <surname>Fischer</surname>
          </string-name>
          ,
          <article-title>Mining big data in education: Afordances and challenges</article-title>
          ,
          <source>Rev. Res. Educ</source>
          <volume>44</volume>
          (
          <year>2020</year>
          )
          <fpage>130</fpage>
          -
          <lpage>160</lpage>
          . doi:
          <volume>10</volume>
          .3102/0091732X20903304.
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>P.</given-names>
            <surname>Adeodato</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Filho</surname>
          </string-name>
          ,
          <article-title>Where to aim? factors that influence the performance of brazilian secondary schools</article-title>
          ,
          <source>in: Proceedings of The 13th International Conference on Educational Data Mining</source>
          ,
          <year>2020</year>
          , p.
          <fpage>5</fpage>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>