<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>Tracking learners' knowledge and skills development</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Andrea Zanellati</string-name>
          <email>andrea.zanellati2@unibo.it</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Maurizio Gabbrielli</string-name>
          <email>maurizio.gabbrielli@unibo.it</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Olivia Levrini</string-name>
          <email>olivia.levrini2@unibo.it</email>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Department of Computer Science and Engineering, University of Bologna</institution>
          ,
          <addr-line>Mura Anteo Zamboni 7, Bologna</addr-line>
          ,
          <country country="IT">Italy</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Department of Physics and Astronomy, University of Bologna</institution>
          ,
          <addr-line>Viale Berti Pichat 6/2, Bologna</addr-line>
          ,
          <country country="IT">Italy</country>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>Workshop Proce dings</institution>
        </aff>
      </contrib-group>
      <abstract>
        <p>This Ph.D. research proposal aims to investigate how exploiting educational data to track the learners' development of knowledge and skills, thus embedding this information in automated tools designed to enhance teaching and learning. The encoding of learners' knowledge and skills is a crucial issue which can be exploited in addressing several tasks, such as underachievement prediction and personalized learning. However, some challenges characterized how to design the encoding and include it in automated tools: dealing with several formats of data (among which also text, video, images, and audio recording), tackling the strong dependence of educational data from the context where they are collected, and consider ethical issues related to explainability and fairness. With this position paper, we introduce the research questions which lead the project, a brief state of the art about techniques used to tackle the students' knowledge and skills encoding, the methodology and the expected results. Specifically, we aim to investigate which data can be used to fulfill our main purpose, test our encoding solutions in two case studies (underachievement prediction and knowledge tracing), and assess the contribution of our encoding to tackle them. As for the methodology, we want to explore strategies of Informed Machine Learning, that is to say incorporating an external knowledge source in the machine learning pipeline, which can improve the explainability and fairness of the models and handle the influence of the external context on the educational data.</p>
      </abstract>
      <kwd-group>
        <kwd>knowledge tracing</kwd>
        <kwd>skill development</kwd>
        <kwd>educational computing</kwd>
        <kwd>informed machine learning</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <sec id="sec-1-1">
        <title>This position paper presents the research proposal for a</title>
        <p>Consortium, with a specific focus on the problem of how
to track the development of students’ knowledge and
skills. The paper was presented a year and a half after
the start of the Ph.D. and contains the conceptual and
motivational framework, the expected development steps,
and a summary presentation of the preliminary results
of the work done so far.</p>
        <p>The paper is organized as follows. In the next section,
we describe the background for the research, focusing
on some challenges, motivating the research questions
for the proposal, and the rational for our methodological
choices. The third section is dedicated to the
methodology. We introduce Informed Machine Learning (IML)
as a reference methodological approach and we outline</p>
        <p>https://www.unibo.it/sitoweb/andrea.zanellati2/en (A. Zanellati);
(O. Levrini)</p>
        <p>0000-0003-1632-3800 (A. Zanellati); 0000-0003-0609-8662</p>
      </sec>
    </sec>
    <sec id="sec-2">
      <title>2. Background</title>
      <sec id="sec-2-1">
        <title>2.1. Datafication in the Educational field</title>
        <p>
          In the last decade, the process of datafication in
society has become increasingly pervasive, also afecting the
educational field [
          <xref ref-type="bibr" rid="ref1">1</xref>
          ]. We assisted in a growing and
varied interest in the application of artificial intelligence
and data science techniques in this sector, with the rise
of new research fields such as Learning Analytics (LA)
and Educational Data Mining (EDM) [
          <xref ref-type="bibr" rid="ref2">2</xref>
          ]. Despite some
diferences, especially in the analysis techniques most
commonly used by the two research communities, LA
and EDM share the goal of extracting knowledge of
interest for educational stakeholders –policy-makers, didactic
coordinators, teachers, parents, and students– and using
the extracted knowledge to improve the learning process
in some way.
        </p>
        <p>In this Ph.D. research proposal, in a broad perspective,
we consider a key issue which is transversal to many
edof knowledge and skills, thus embedding this
informa(M. Gabbrielli); 0000-0003-1632-3800 (O. Levrini)</p>
        <p>
          © 2022 Copyright for this paper by its authors. Use permitted under Creative Commons License ucational situations: tracking the learners’ development
tion in automated tools designed to enhance teaching ing availability of educational data promotes the
applicaand learning. There are several tasks which can benefit tion of data science and machine learning techniques, to
from the encoding of students’ knowledge and skills de- exploit data potential in enhancing learning and teaching.
velopment, e.g. low achievement prediction models [
          <xref ref-type="bibr" rid="ref3 ref4">3, 4</xref>
          ] On the other hand, we have to consider that educational
or automated feedback system for personalized learning data are often highly context-dependent. There are few
[
          <xref ref-type="bibr" rid="ref5">5</xref>
          ]. As main objective, we aim to tackle the problem standardized large-scale educational datasets, i.e. data
of how encoding the students’ knowledge and skills de- are very heterogeneous for diferent class groups.
Furvelopment, identifying valuable data resources for its thermore, data labeling is not a common practice in
classrepresentation, and testing our solutions efectiveness in room settings, because it is not one of the main objectives
addressing the tasks listed above. There are three main of teachers or other stakeholders traditionally involved
starting considerations which motivate our proposal and in training processes. Moreover, educational data are
lead to design our research questions. often indicators of a learning competence or behavior,
whose evaluation depends to some extent also on the
2.2. Three challenges to address for evaluator. Let us consider, for example, the evaluation of
task on creative thinking skills: diferent evaluators can
automated tracking of knowledge
        </p>
        <p>result in diferent evaluation labels. Therefore, we are
and skills development afirming that the collection of data on students cannot
Firstly, the datafication process in the educational field ignore the context in which this occurs and can hardly
is characterized by several types of data, namely product be considered free from pedagogical, psychological, and
data, process data, and background data. Product data cognitive science theories, consolidated over years of
reare related to what students produced, and how they search and assumed more or less explicitly by teachers,
show their learning. They can be collected while stu- by those who design the context of learning or by who
dents are learning, e.g. personal notes during classes, carries out the data collection. The research on
Intelliproduction of diagrams, and concept maps, questions gent Tutoring Systems or Adaptive Educational Systems
answering, resolutions, and formative and summative already considers the domain model and the pedagogical
assessments. Process data deals with how students are model together with the learner model [10, 11]. However,
learning a specific content or how they behave during in these systems, they are often separate components,
their performance assessment. The possibility to gain while we are suggesting that the domain model and the
process data increased in the last years due to a spread pedagogical model directly influence the learner model.
in the use of digital technology in education, e.g Learn- This assumption relies on the issue already stated in the
ing Management Systems (LMS), Massive Open Online literature of the theory-ladenness in data-intensive
apCourses (MOOCs), and computer-based tests (CBt), both proaches [12]. We can name this as the theory-ladennes
at school and university. This was furtherly accelerated challenge.
by the recent COVID-19 pandemic. These technologies As a consequence, it is not easy to have robust and
allow to track individual students’ learning processes and fair datasets on which to apply automatic data mining
collect data such as their mouse clicks, scrolling behav- techniques, pointing out the third issue on ethics. An
unior, or time spent on diferent tasks or content resources. balanced or unrepresentative dataset may disadvantage
Also face-to-face classes allow the collection of process students not suficiently represented by the sample. The
data, although it is often challenging and more time- model –here intended as an automated detector for a
comconsuming. An example is the data collected through the monly seen outcome or measure in LA and EDM, such as
Think-Aloud protocol [6]. As for background data, they dropout, underachievement, afects, learning strategies,
usually contain student demographics (parents’ educa- and disengaged behaviors– may be prone to overfitting
tion, family income, household registration), curriculum the profile of well-represented students, resulting
inflexiplans, teachers’ quality and style, and student perfor- ble to new cases or changes that may occur in the school
mance evaluation. The previous list shows the variety of population. According to Baker [13], this is not just a
formats for educational data referable to a single learning technical challenge but it is a challenge for inclusion. In
activity: numerical, categorical and boolean variables are fact, a lot of the populations that we want to focus on,
enriched by other formats e.g. texts, images, videos, and including historically underserved and underrepresented
voice recording. This leads to the problem of how multi- populations, are the ones it is harder to collect data for.
modal data fusion can be conducted in learning analytics This can be seen as a generalizability challenge for the
[7, 8]. We refer to this challenge as the multimodal data models developed in LA and EDM. We can refer to this
challenge last point as the ethical challenge.</p>
        <p>A second issue concerns how these data are collected,
organized, and labeled [9]. On the one hand, the
increas2.3. Research Questions tion about relations between entities in certain contexts”.
There are three types of knowledge, several possible
repTo sum up, the datafication process, which afected the resentations, and diferent forms of integration, as shown
educational field, is an opportunity to promote data- in Table 1. When dealing with the approach of informed
informed decisions for revising the learning designs and machine learning in the educational field, the main source
avoiding behaviors that lead to poor learning. In partic- of prior knowledge to consider is the expert knowledge,
ular, one of the central problems is how to use data for often informal and validated through a group of
expethe design of a learner model, an essential component for rienced specialists. Also world knowledge could be a
data-informed pedagogies and educational actions. In source of information to take into account, referring to
the development of this model, there are some challenges facts from everyday life that are known to almost
everyto be taken into consideration: the multimodal data chal- one, subsuming also linguistics.
lenge, the theory-ladennes challenge, and the problems Some forms of knowledge integration in LA models
of inclusiveness, fairness, and generalizability, summa- already exist; it almost occurs with the search for synergy
rized in the ethical challenge. The considerations in the with learning design, oriented to data-informed
learnprevious subsection lead us to formulate the following ing and teaching practice that preserve the agency of
research questions. students and teachers [15, 16], overcoming purely
data</p>
        <p>Most studies in this area have a purely or highly data- driven approach. This way can be seen as integrating
driven approach, which does not consider how context prior knowledge into the final step of the machine
learnand several pedagogical assumptions can afect and be ing pipeline when its output is validated or benchmarked
integrated into the machine-learning pipeline. This leads against existing knowledge through human mediation.
us to formulate the following research questions. However, there are other forms of knowledge
integraRQ1 How diferent educational data can tion in the machine learning pipeline –Training Data,
be used for a reliable representation of Hypothesis Set, and Learning Algorithm– that could be
learners’ development of knowledge and investigated to face the challenge of the reconstruction of
skills? students’ learning trajectories and students’ competence
development. In this research proposal we want to
adRQ2 Is there any prior knowledge which dress the problem of developing knowledge and skills by
can be integrated into AI tools used for investigating which supplementary knowledge sources
tracking learners’ development of knowl- can be used, how they can be represented and where
edge and skills to improve their perfor- they can be integrated into the machine learning pipeline
mance or their explainability? (training data, hypothesis set, learning algorithm, and
The first research question is motivated by both mul- ifnal hypothesis). To do this we consider two case
studtimodal data and ethical challenges, and also wants to ies, i.e two situations in which the problem of tracking
suggest the need to reflect on what information is actually the development of knowledge or skills is relevant and
collected and expressed in the data. The second question which we propose to approach from the perspective of
inemphasizes the need to consider other sources of infor- formed machine learning. The first case study concerns
mation. The term prior knowledge here is intended in a predictive model of underachievement and represents
the perspective of Informed Machine Learning, chosen a study already started for which there are some
prelimas methodological paradigm, that we describe in the next inary results. In this first case, we present an example
section. In our discussion, we can assume the domain of feature engineering strongly driven by an explicit
inand the pedagogical models as integrative knowledge tegration of a theoretical framework. The second case
source to data. study concerns the problem of knowledge tracing. It
represents a work direction still to be developed which also
requires an in-depth analysis of what already exists in
3. Methodology: Informed the literature as attempts at hybrid approaches in which
Machine Learning a theory-laden is present. Therefore, we propose to
investigate the RQs through two case studies that allow to
use of diferent data (in the first case it is a static dataset
and in the second dynamic) for learner modeling and to
test prior knowledge integration strategies.</p>
        <sec id="sec-2-1-1">
          <title>According to von Rueden et al. [14] Informed Machine</title>
          <p>Learning describes “learning from a hybrid information
source that consists of data and prior knowledge”. It is
not a purely data-driven approach due to the integration
of an external and independent knowledge source into
the machine learning pipeline.</p>
          <p>With the term knowledge they assumed a computer
science perspective, defining it as “validated
informa4. Case studies
learning models: random forest and two neural networks
(categorical embedding neural network and feature
tok4.1. Low achievement prediction enizer transformer). Finally, we presented a
knowledgeexploiting longitudinal large-scale based methods to encode students learning. Specifically,
assessment tests in the design of the learner model, we exploit features
already present in the dataset regarding demographic
4.1.1. Problem definition and State of the Art information and the socio-cultural-economic context of
Firstly, we examine data collected through national large- the student, together with other features more related
scale assessment tests. These tests are often used to sup- with the student’s learning. This second set of features
port educational policy decisions [17] or in studies aiming is obtained through engineering the boolean features
to determine the relationship between socio-economic that record the correctness of the student’s responses to
factors and school performances. Nevertheless, they the individual items of the test. The new features are
are designed to measure students’ knowledge and skills defined based on the framework used by INVALSI for
and often to track longitudinally the students’ learning classifying the items, in terms of areas, processes, and
path [18]. These test design features enable the collection macro-processes. The rational for this choice is
twoof data that can be useful for tracking the development folds: firstly, this allows application to students from
difof knowledge and skills and building predictive models ferent cohorts who have taken diferent tests; secondly,
they are directly related to students learning in terms of
Ifnor[1th9e], rfiosrkeoxfalmonpgle-,tethrmeauunthdoerrascrheifeevretmoednattaorcodlrloepcoteudt. knowledge and skills, that are very important to design
through the PISA international large-scale assessment educational interventions to counteract the phenomenon
tests to predict math proficiency. of underachievement. The classification framework is</p>
          <p>
            Several machine learning techniques have been ex- shown in Table 2 This framework represents the source
ploited to build predictive models for students’ perfor- of integrative prior knowledge. Its representation is in
mance [
            <xref ref-type="bibr" rid="ref4">4</xref>
            ], including supervised learning, e.g., random the form of algebraic equations, with which we define the
forests, support vector machine and Bayesian network, new features, i.e. for each student a correctness rate is
unsupervised learning, e.g., k-means and hierarchical computed for each area, process, or macro-process. The
clustering, and recommender systems, e.g., collaborative integration takes place into the train set.
ifltering. Our results are summarized in table 3, which are
promising. We aim to improve the research in three
main directions. Firstly, we want to test the
transfer4.1.2. Specific objectives and outcomes ability to other disciplines such as Italian and English,
In [20] we present some preliminary results about maths which are tested by INVALSI, by using a similar
represenlow achievement prediction exploiting a very large ital- tation or encoding for students learning. Secondly, we
ian dataset (more than 700000 students). Specifically, we aim to improve the data quality by training and testing
exploit data collected through the INVALSI1 large-scale the model with students from diferent cohorts. This is
assessment test to predict at grade 5 low achievement in possible by using at least four cohorts of students and
math at the end of compulsory school at grade 10. We may improve the transferability of the models to new
used three AI tools based on state-of-the-art machine cohorts. In fact, training the model on students’ data
from diferent school years could help in avoiding
over
          </p>
        </sec>
        <sec id="sec-2-1-2">
          <title>1Italian National Institute for the Evaluation of the School System</title>
          <p>(P1) Know and master the specific contents of mathematics
(P2) Know and use algorithms and procedures
(P3) Know diferent forms of representation and move from one to the other
(P4) Solve problems using strategies in diferent fields
(P5) Recognize the measurable nature of objects and phenomena in diferent
contexts and measure quantities
(P6) Progressively acquire typical forms of mathematical thought
(P7) Use tools, models and representations in quantitative treatment
information in the scientific, technological, economic and social fields
(P8) Recognize shapes in space and use them for problem solving
Macro-process
(MP1) Formulating
(MP2) Interpreting
(MP3) Employing
iftting patterns to a specific test. Furthermore, we can try 4.2. Knowledge tracing for personalized
diferent student modeling approach, which is not driven learning
by the Invalsi theoretical framework but which take into
account other contextual information, e.g the items difi- 4.2.1. Problem definition and State of the Art
culty or the items embedding based on their texts. A last As a second case study, let us consider an instructional
point of development concerns the interpretability which unit provided through a learning management system
can be improved by comparing the feature importance (LMS). This is usual for MOOCs courses, it has also been
analysis of the random forest model with the weights the case for many students and teachers during
COVIDwhich define our neural networks. 19 pandemic [21] and potentially it may also be exploited</p>
          <p>To sum up, With this case study we want to investi- in face-to-face classes, as a tool to organize teaching
magate the potential of educational data collected through terials and manage diferent activities. As students work
longitudinal large-scale assessment tests for the repre- with the LMS they produce a wealth of data including
sentation of the development of knowledge and skills, product data (e.g. an explanation written in an electronic
and look for other prior knowledge resources that can journal, or a video recorded through a mobile app) and
improve the performance of the model. process data (e.g. the number of edits made in the
writing of this explanation, or log data). This data can be
Table 3 exploited for the well-known problem of knowledge
tracPerformance on test set ing [22], which can be described as monitoring students’
Models Accuracy Precision Recall changing knowledge states during the learning process
and accurately predicting their performance in future
Random Forest 0.77 0.62 0.67 exercises. This information can be further applied to
purCE neural network 0.76 0.76 0.76 sue personalized learning in order to maximize students’
FTT neural network 0.78 0.77 0.78 learning eficiency.</p>
          <p>The most common machine learning techniques to
handle knowledge tracing are Bayesian Network [22]
and Dynamic Bayesian Network [23], to build
probabilistic models. Another frequent approach is that of
logistic models, such as learning factor analysis [24],
performance factor analysis [25] and knowledge tracing
machines [26]. In recent years, it has been explored also
the use of deep neural networks [27], which outperform
more traditional techniques, named Deep Knowledge
Tracing (DKT).</p>
        </sec>
        <sec id="sec-2-1-3">
          <title>On the one hand, the brief background presentation in the</title>
          <p>previous section demonstrates how crucial and
transversal the proposed RQs are in the educational field. On
4.2.2. Dataset and goals the other, it highlights their complexity and the need to
address them focusing on case studies, which may be
For this case study we will consider data collected by very diferent from each other, although they share the
the ALICE project (Learning Progression Analytics - An- need to identify suitable data to represent the learners’
alyzing Learning for Individualized Competence devel- development of knowledge and skills. The comparative
opment in mathematics and science Education), led by analysis of the results obtained in diferent case studies
IPN Kiel, with the cooperation of DIPF Frankfurt and can bring out good practices or scalable solutions.
Ruhr-University Bochum. ALICE aims to exploit data Therefore, in this research project, we aim to focus
from students’ interactions with digital technologies in on diferent educational issues, referable to those
preSTEM –Science, Technology, Engineer, and Mathemat- sented in the previous section: underachievement and
ics– classroom learning, both to predict the productivity knowledge tracing. For each case study, we are going to
of students’ learning trajectories for their competence identify or build a dataset useful for defining a learner
development and to identify underlying causes of unpro- model, intended as a representation of learners’
develductivity. The data is collected through the implemen- opment of knowledge and skills, thus contributing to
tation of some instructional units in face-to-face classes RQ1. To handle RQ2, the proposed representations will
using an LMS as a teaching aid. In this context, we want be used to test several state-of-the-art solutions of
mato investigate which useful prior knowledge related to chine learning, which can tackle the educational problem
ALICE educational context can be modeled and how. Fur- that motivates each case study. Furthermore, we aim to
thermore, we aim to explore where they can be integrated identify significant sources of prior knowledge (domain
in the ML pipeline to improve the learning trajectories model and pedagogical model) and investigate how to
inanalysis. tegrate them into the machine learning tools. Hence we</p>
          <p>Our hypothesis is that the analysis of log data for the will evaluate their efectiveness, with respect to
convenknowledge tracing can benefit from information on the tional machine learning solutions, by considering models’
face-to-face context, such as the choices of the teacher in performances and their explainability, trying to come up
the exposition of the unit contents and teaching times, re- with the main goal of RQ2.
lationships peer-to-peer, or the didactic model on which
the unit itself is designed. We want to investigate the
possibility of representing one or more of these sources of References
prior knowledge through a graph or a bayesian network,
that can be used as input for a DKT network together
with the log data collected on the student’s interactions
with the learning management system.</p>
        </sec>
      </sec>
      <sec id="sec-2-2">
        <title>4.3. Remarks</title>
        <p>In both cases, we refer to data collected about students’
learning to build a learner model. However, we have to
consider that the learning dynamics are strongly
influenced by the domain model – understood as physical
space, social-relational space, disciplinary space–, as well
as by tutors/teachers and the pedagogical model they
assumed. Domain model and pedagogical model may be
considered as prior knowledge, here intended as a
separate source with respect to data about students’ learning
behaviors or performances, which can be integrated into
the machine learning pipeline, following the paradigm
of Informed Machine Learning[14].
[6] R. Jääskeläinen, Think-aloud protocol, Handbook ing model, in: 2021 IEEE 19th International
Sympoof translation studies 1 (2010) 371–374. sium on Intelligent Systems and Informatics (SISY),
[7] S. Mu, M. Cui, X. Huang, Multimodal data fusion IEEE, 2021, pp. 49–54.</p>
        <p>in learning analytics: A systematic review, Sensors [20] A. Zanellati, S. P. Zingaro, M. Gabbrielli, Student
20 (2020) 6856. low achievement prediction, Lecture Notes in
Com[8] T. Baltrušaitis, C. Ahuja, L.-P. Morency, Multimodal puter Science 13355 (2022).</p>
        <p>machine learning: A survey and taxonomy, IEEE [21] Q. Liu, S. Shen, Z. Huang, E. Chen, Y. Zheng,
transactions on pattern analysis and machine intel- A survey of knowledge tracing, arXiv preprint
ligence 41 (2018) 423–443. arXiv:2105.15106 (2021).
[9] W. Wang, G. Xu, W. Ding, Y. Huang, G. Li, J. Tang, [22] A. T. Corbett, J. R. Anderson, Knowledge tracing:
Z. Liu, Representation learning from limited educa- Modeling the acquisition of procedural knowledge,
tional data with crowdsourced labels, IEEE Transac- User modeling and user-adapted interaction 4 (1994)
tions on Knowledge and Data Engineering (2020). 253–278.
[10] R. Nkambou, R. Mizoguchi, J. Bourdeau, Ad- [23] T. Käser, S. Klingler, A. G. Schwing, M. Gross,
Dyvances in intelligent tutoring systems, volume 308, namic bayesian networks for student modeling,
Springer Science &amp; Business Media, 2010. IEEE Transactions on Learning Technologies 10
[11] V. J. Shute, D. Zapata-Rivera, Adaptive educational (2017) 450–462.</p>
        <p>systems, Adaptive technologies for training and [24] H. Cen, K. Koedinger, B. Junker, Learning factors
education 7 (2012) 1–35. analysis–a general method for cognitive model
eval[12] W. Pietsch, Aspects of theory-ladenness in data- uation and improvement, in: International
conferintensive science, Philosophy of Science 82 (2015) ence on intelligent tutoring systems, Springer, 2006,
905–916. pp. 164–175.
[13] R. Baker, Challenges for the future of educational [25] P. I. Pavlik Jr, H. Cen, K. R. Koedinger, Performance
data mining: The baker learning analytics prizes factors analysis–a new alternative to knowledge
(2019). tracing., Online Submission (2009).
[14] L. von Rueden, S. Mayer, K. Beckh, B. Georgiev, [26] J.-J. Vie, H. Kashima, Knowledge tracing machines:
S. Giesselbach, R. Heese, B. Kirsch, M. Walczak, Factorization machines for knowledge tracing, in:
J. Pfrommer, A. Pick, R. Ramamurthy, J. Garcke, Proceedings of the AAAI Conference on Artificial
C. Bauckhage, J. Schuecker, Informed machine Intelligence, volume 33, 2019, pp. 750–757.
learning - a taxonomy and survey of integrating [27] C. Piech, J. Bassen, J. Huang, S. Ganguli, M. Sahami,
prior knowledge into learning systems, IEEE Trans- L. J. Guibas, J. Sohl-Dickstein, Deep knowledge
actions on Knowledge and Data Engineering (2021) tracing, Advances in neural information processing
1–1. doi:1 0 . 1 1 0 9 / T K D E . 2 0 2 1 . 3 0 7 9 8 3 6 . systems 28 (2015).
[15] M. Blumenstein, Synergies of learning analytics
and learning design: A systematic review of student
outcomes., Journal of Learning Analytics 7 (2020)
13–32.
[16] H. Shen, L. Liang, N. Law, E. Hemberg, U.-M.</p>
        <p>O’Reilly, Understanding learner behavior through
learning design informed learning analytics, in:
Proceedings of the Seventh ACM Conference on</p>
        <p>Learning@ Scale, 2020, pp. 135–145.
[17] G. E. Fischman, A. M. Topper, I. Silova, J. Goebel, J. L.</p>
        <p>Holloway, Examining the influence of international
large-scale assessments on national education
policies, Journal of education policy 34 (2019) 470–499.
[18] L. Branchetti, F. Ferretti, A. Lemmo, A. Mafia,</p>
        <p>F. Martignone, M. Matteucci, S. Mignani, A
longitudinal analysis of the italian national standardized
mathematics tests, in: CERME 9-Ninth Congress of
the European Society for Research in Mathematics</p>
        <p>Education, 2015, pp. 1695–1701.
[19] A. Pejić, P. S. Molcer, K. Gulači, Math proficiency
prediction in computer-based international
largescale assessments using a multi-class machine
learn</p>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>A.</given-names>
            <surname>Klašnja-Milićević</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Ivanović</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Budimac</surname>
          </string-name>
          ,
          <article-title>Data science in education: Big data and learning analytics</article-title>
          ,
          <source>Computer Applications in Engineering Education</source>
          <volume>25</volume>
          (
          <year>2017</year>
          )
          <fpage>1066</fpage>
          -
          <lpage>1078</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>C.</given-names>
            <surname>Romero</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Ventura</surname>
          </string-name>
          ,
          <article-title>Educational data mining and learning analytics: An updated survey</article-title>
          ,
          <source>Wiley Interdisciplinary Reviews: Data Mining and Knowledge Discovery</source>
          <volume>10</volume>
          (
          <year>2020</year>
          )
          <article-title>e1355</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>B.</given-names>
            <surname>Albreiki</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Zaki</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Alashwal</surname>
          </string-name>
          ,
          <article-title>A systematic literature review of student performance prediction using machine learning techniques</article-title>
          ,
          <source>Education Sciences</source>
          <volume>11</volume>
          (
          <year>2021</year>
          )
          <fpage>552</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>J. L.</given-names>
            <surname>Rastrollo-Guerrero</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. A.</given-names>
            <surname>Gomez-Pulido</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Durán-Domínguez</surname>
          </string-name>
          ,
          <article-title>Analyzing and predicting students' performance by means of machine learning: A review</article-title>
          ,
          <source>Applied sciences 10</source>
          (
          <year>2020</year>
          )
          <fpage>1042</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>N. S.</given-names>
            <surname>Raj</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Renumol</surname>
          </string-name>
          ,
          <article-title>A systematic literature review on adaptive content recommenders in personalized learning environments from 2015 to 2020</article-title>
          ,
          <article-title>Journal of Computers in Education (</article-title>
          <year>2021</year>
          )
          <fpage>1</fpage>
          -
          <lpage>36</lpage>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>