<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <journal-title-group>
        <journal-title>AIxEDU</journal-title>
      </journal-title-group>
    </journal-meta>
    <article-meta>
      <title-group>
        <article-title>Large Language Models for the Assessment of Students' Authentic Tasks. A Replication Study in Higher Education</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Daniele Agostini</string-name>
          <email>daniele.agostini@unitn.it</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Federica Picasso</string-name>
          <email>federica.picasso@unitn.it</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Helga Ballardini</string-name>
          <email>helga.ballardini@unitn.it</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>. The Context: AI Assessment in Higher Education</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>University of Trento, Department of Psychology and Cognitive Sciences</institution>
          ,
          <addr-line>Corso Bettini, 84, 38068 Rovereto</addr-line>
          ,
          <country country="IT">Italy</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2024</year>
      </pub-date>
      <volume>2</volume>
      <fpage>25</fpage>
      <lpage>28</lpage>
      <abstract>
        <p>After the public release of ChatGPT (November 30th, 2022) and consequently, that of all its competitors, the use of Large Language Models (LLMs) has become widespread among the public. The most significant impact was perceived from the very beginning in the field of Education and Instruction [ 1, 2, 3, 4, 5, 6, 7]. Of particular interest for this paper is its use both by teachers and students in particular in the context of higher education [8, 4, 9]. The immediacy with which Large Language Models (LLMs) have been integrated into higher education practices, both by teachers and students, leads to questions of fundamental importance relating to their efectiveness and reliability. In this field, LLMs become the means through which teachers have the opportunity to revolutionise the interaction with students, the management of workload and the personalisation of each learning experience [2]. Although these technologies are recognised as having advantages and potential for improving learning in terms of accessibility and personalisation [7], a crucial question concerns their application in assessment practices, especially the ability to objectively and impartially evaluate students' performance. The possibilities of using these tools in the field of learning evaluation is relatively little known, which implies the need to delve deeper into the topic for its application both in pedagogical theory and in educational practice. A previous study has been already published [10] which explored the use of the main LLM in the specific context of assessing students' papers, and this is a replication study based on it. The purpose of the current study is to explore the possible use of the main LLMs in the specific context of evaluating students' written productions, with a focus on the aspects of accuracy that are evaluated with the help of a rubric proposed by the teacher. This article is part of a series of contributions that focus on this topic, in light of the principles and application of the AI-Mediated Assessment for Academics and Students (AI-MAAS) model [11].</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;Large Language Models (LLMs)</kwd>
        <kwd>AI-Assisted Assessment</kwd>
        <kwd>Rubrics</kwd>
        <kwd>Authentic Tasks</kwd>
        <kwd>Academic Assessment</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>
        and exhaustive answers. For example, LLMs allow users to avoid various typical inconvenient steps
that characterise the standard use of search engines, such as the selection of long lists of websites, the
acceptance of cookies and the appearance of advertising banners. As a result, educational institutions
and agencies have begun to incorporate LLMs and generative AI into their curricula at various levels,
developing courses to harness the potential of these innovative technologies. There is currently a
strong emphasis on AI Literacy [16, 17], which allows professionals from diferent sectors, including
educational institutions, to deepen their understanding of the fundamental elements of AI generative,
the availability of tools, the functionalities and methods of use that make LLMs efective tools in all
ifelds [ 18, 19, 20, 21, 22, 23]. However, one of the most critical issues concerns information management:
LLMs possess enormous potential due to their ability to analyse and generate data; this raises numerous
questions about accuracy, privacy and ethics in information management and ownership of output. The
challenge in this continuously and rapidly evolving field becomes the ability to pay constant attention
and critically evaluate so that end users always use LLMs responsibly [24, 25, 26, 27]. In relation to
this issue, higher education institutions have reacted by placing themselves on the defensive, so much
so that some universities, in order to counteract the possible use of LLMs by students during exam
tests, they have reintroduced the obligation to write by hand and also take oral tests [
        <xref ref-type="bibr" rid="ref8">8, 28</xref>
        ]. At the
same time, pieces of software created specifically to detect the productions generated by LLMs were
introduced on the market. However, these turned out to be inefective, causing management and legal
problems for institutions because students could be unfairly accused of sending texts generated by
artificial intelligence [ 29, 21, 30]. To avoid such inconveniences, national and international institutions
and universities promptly provided themselves with guidelines that promote ethical behaviour towards
the use of LLMs while maintaining a certain caution, allowing students and teachers to use them
efectively to carry out tasks and benefit institutions. Important international bodies and universities
moved in this direction, such as UNESCO [31, 32], the JISC National Center for AI [33], the Russell
Group [34], the French National Ministry of Education [35], the US Department of Education [36]
and University College London [37]. Assessment tasks have proven to be arguably the ones that can
profit most from the AI technology, especially in terms of sustainability. However, caution is needed as
LLMs without specific task adaptations have proven incapable and unreliable in managing students’
assessments independently [38, 33], while LLMs supported by assessment tools have been shown
to produce satisfactory results [39]. Above all, the use of artificial intelligence by students requires
that teachers know how to take ethical aspects into consideration and act with responsibility when
evaluating tasks and tests which results could have a great impact on students’ careers (for example,
motivation, grades, scholarships, acceptance into master’s or doctoral programs).
      </p>
    </sec>
    <sec id="sec-2">
      <title>2. Theoretical Framework</title>
      <p>Since the 1980s the idea of being able to use computerised systems (and now also artificial intelligence)
to assist educators in their assessment tasks and to be able to make precise, impartial and informed
decisions has been present in much literature [40, 41]. The possibility of using LLMs for learning
assessment had already been explored in the period immediately preceding the release of ChatGPT,
while the use of transformer models, including OpenAI’s GPT-3, were already well established. Tamkin
et al. [42] emphasised their educational application, which included:
• Summary: LLMs are able to summarise even very long texts. This use can help students submit
concise summaries. Furthermore, various parameters can be considered for the synthesis, and
this supports educators in providing precise information on the elements of the text that will be
evaluated.
• Questions and Answers: LLMs can "understand" various portions of text, answer questions, and
ask questions when required. These features are useful for providing interactive feedback and
learning experiences.
• Classification: LLMs can classify the text into predefined categories: this allows you to introduce
assisted assessment or classify students’ feedback.
• Plagiarism detection: by comparing the similarity between diferent texts, LLMs are very useful to
detect potential cases of plagiarism among students or to identify the misuse of original materials
by students.
• Assessment of knowledge: LLMs can assess students’ understanding of a topic based on their
written productions, especially if the information is generated from correct homework and with
the help of an assessment rubric to refer to.</p>
      <p>These five applications are fundamental to using LLMs in learning assessment. Following the
introduction of ChatGPT and other universally accessible LLMs, UNESCO published the guidelines “AI
and education: Guidance for policy-makers” [31], which suggest the following recommendations for
learning assessment:
1. Test and implement AI technologies to support the assessment of various dimensions of skills
and outcomes.
2. Use caution when using an automated assessment with closed-ended, rule-based questions.
3. Use AI-generated formative assessment as an integrated feature of learning management systems
(LMS) to analyse learning outcomes more accurately and eficiently and reduce the risk of human
bias.
4. Use the ability to provide AI-powered progressive assessments to regularly update students and
parents.
5. Examine and evaluate the use of facial recognition and other artificial intelligence capabilities for
users’ recognition and their tracking in remote online assessments.</p>
      <p>
        Based on these diferent theoretical approaches, recommendations and guidelines, the AI-Mediated
Assessment for Academics and Students (AI-MAAS) model was developed which is currently under
validation; it proposes two potential implementations of LLMs for the assessment of learning: the
ifrst one for formative evaluation and the second one for summative evaluation [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ]. In both cases,
the selected LLM must be able to evaluate using an assessment rubric provided by the teachers or
by the students. Given the novelty of the tool, so far, there are not many experiments in this field.
Martin et al. [39] worked on this opportunity, starting from the need to be able to assign students even
complex tasks that involve a certain degree of reasoning, abstract conceptualisation and reprocessing
of information; while the correction of these types of tasks (with a large quantity of long open answers)
often proves to be an unsustainable task for teachers. Some researchers, working on this aspect, have
demonstrated that in the evaluation of a chemistry task, for example, it is possible to use LLMs: in this
case, an almost perfect match was obtained between the scores assigned by human raters and the scores
generated by the LLMs. It should be highlighted, however, that Martin and colleagues did not simply
use an LLM to achieve this result. The researchers used a complex procedure that involved, among
other operations, the unsupervised machine learning technique HDBSCAN (Hierarchical Density-Based
Spatial Clustering of Applications with Noise), a cluster mapping and training of a deep neural network
classifier. The aim of this study was to test an operational model and demonstrate its feasibility. This
excellent solution represents the result of models trained on specific tasks and populations; therefore,
it cannot be assumed that the procedure applied can be replicated by any teacher not specialised in
Machine Learning. Other studies have instead used LLMs for assessment purposes without comparing
the performance of the AI with that of a teacher. These studies applied for example in the evaluation
of L2 English tasks [43] and in supporting self-assessment came to satisfactory results [44]. Machine
learning has also been applied in the evaluation of tasks related to STEM disciplines, but without using
LLM [45].
      </p>
      <p>
        Finally, a previous study [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ] explored the use of the main LLMs in the specific context of assessing
students’ papers, with a focus on their accuracy in assessing according to a rubric developed by the
teacher. The idea was that employing LLMs for assessment in higher education can enable the adoption
of teaching and assessment approaches that were previously unsustainable and unscalable. This
should help to ensure constructive alignment [46] and thereby improve the quality and efectiveness
of university teaching. The study, aimed at selecting the most human-like evaluation amongst LLMs,
highlighted that while some AI models, like ChatGPT-4 and Claude 2, performed well in most of the
assessment criteria, others, such as Microsoft Copilot and Google Bard, were far from human-like
assessment. The article recommends further research on ChatGPT and Claude, with potential inclusion
of open-source models as well as involving multi-shot prompting, expanding the student sample,
involving more evaluators, and refining and redesigning the rubrics.
      </p>
    </sec>
    <sec id="sec-3">
      <title>3. Methodology and Tools</title>
      <p>
        This is a replication study of “Are Large Language Models Capable of Assessing Students’ Written
Products? A Pilot Study in Higher Education” published in “Research Trends in Humanities Education &amp;
Philosophy, 11” [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ] that follows the same methodology with updated LLMs employed, a greater sample
of students and of human evaluators. It explores the use of leading Large Language Models (LLMs) in
the specific context of assessing student written products, focusing on their precision and ability to
evaluate according to a grading rubric developed by the teacher. The goal is to understand whether and
which models can be used by university and non-university educators who are not experts in Machine
Learning to assess students’ written products, even in the presence of open-ended tasks and questions,
thanks to grading rubrics. The pilot study was conducted at the University of Trento within the
context of a university habilitation course for secondary school teachers, during the module concerning
learning methodologies. One-hundred-fourty-two students participated anonymously, divided into
35 groups, along with 3 evaluating teachers, experts in experimental pedagogy and assessment. No
data regarding the students’ demographics was collected. The groups were tasked with carrying out an
authentic task, namely to re-designing a past educational intervention that proved to be unsuccessful,
targeted at a specific class (which could range from 1st grade of lower secondary school to 5th grade
of higher secondary school, depending on the group composition). They were instructed to identify
the past teaching approaches and strategies and to now think of diferent ones more suited to reach
the intended learning outcomes. Furthermore, students’ reflection and redesign ability was evaluated
through the rubric of reference. To complete this task, groups were given two hours and thirty minutes,
and a template for the educational design consisting of the following sections was provided: Involved
Disciplines, Class and Grade Level, Intervention Title, Teacher, Programme and Learning Objectives,
Context and Environment (formal, informal, type of setting, etc.). Moreover, a description of the
reflection process applied for renewing the formative design is required, and it is considered under
evaluation. In the part of the schedule with details, they were asked to explain the programming
with concise descriptions of the various educational activities, the teacher’s tasks, and those of the
students. Within this framework, groups had the freedom to propose their original programming. The
ifnal product of each group is thus an MS Word file containing the programming of the educational
intervention according to the described template. For the evaluation of the products, the following
grading rubric (Table 1) was prepared, consisting of five evaluation criteria with four levels for each
criterion.
      </p>
      <p>Three expert human evaluators and seven LLMs (plus one that merged the feedback of all 7 models
in a single one) evaluated all the student groups’ products. The LLMs selected for this study were the
most popular competing models at the time, and they are applied in the assessment process through
the use of big-AGI (https://big-agi.com/). Big-AGI is an AI suite created to make advanced artificial
intelligence accessible and was chosen for ease of adding several models through API, the possibility of
imparting system prompts and the function (called “beam”) for sending the same prompt to several
LLMs at the same time. Human results were then compared with the results produced by the LLMs
with various statistical analysis (see Method of Analysis section). The models used are:
1. GPT-4o: released in May 2024, GPT-4o is a multilingual, multimodal generative pretrained
transformer developed by OpenAI. The model is capable of processing and generating text,
images, and audio, making it a versatile tool for a wide range of tasks. Its multimodal capabilities
Demonstrates a lim- Shows a basic
ited understanding of understanding of
edueducational architec- cational architectures,
tures, with applica- applying them
genertions not always ap- ally correctly but with
propriate or consis- some uncertainties.
tent.</p>
      <sec id="sec-3-1">
        <title>Applies educational ar</title>
        <p>chitectures correctly,
with a good
understanding of their use in
the specific context.</p>
      </sec>
      <sec id="sec-3-2">
        <title>Selects and imple</title>
        <p>ments appropriate
teaching
strategies with a good
correlation to the
intervention goals.</p>
      </sec>
      <sec id="sec-3-3">
        <title>The objectives are well</title>
        <p>defined and generally
aligned with the
chosen teaching
architectures and strategies.</p>
      </sec>
      <sec id="sec-3-4">
        <title>The scan is clear</title>
        <p>and generally
wellstructured, with a
good progression of
activities.</p>
      </sec>
      <sec id="sec-3-5">
        <title>Provides good reflection on the changes, with clear links to the learning objectives.</title>
        <sec id="sec-3-5-1">
          <title>Excellent Level (4 points)</title>
          <p>Demonstrates a
thorough understanding of
educational
architectures, applying them
in an innovative and
contextually relevant
manner. Clearly
justifies the choices made.
Selects and
implements highly efective
and diversified
teaching strategies,
perfectly adapted
to the objectives
and context of the
intervention.</p>
          <p>The objectives are
clear, specific,
measurable and perfectly
aligned with the
chosen teaching
architectures and
strategies.</p>
        </sec>
      </sec>
      <sec id="sec-3-6">
        <title>The scan is detailed,</title>
        <p>logical and
wellstructured, with a
clear progression of
activities and realistic
timeframes.</p>
      </sec>
      <sec id="sec-3-7">
        <title>Provides deep and crit</title>
        <p>ical reflection on the
changes made, clearly
justifying each choice
in relation to the
learning objectives.</p>
        <p>Selection The teaching strate- Uses some relevant
and imple- gies chosen are limited teaching strategies,
mentation or not always appro- but their
implemenof teaching priate for the objec- tation could be more
and learning tives of the interven- targeted or diversified
strategies tion. in relation to the
intervention goals.</p>
      </sec>
      <sec id="sec-3-8">
        <title>Definition of The objectives are</title>
        <p>the intended vague, not measurable
learning out- or not aligned with
comes the chosen teaching
architectures and
strategies.</p>
      </sec>
      <sec id="sec-3-9">
        <title>The objectives are</title>
        <p>present but could be
more specific or better
aligned with the
teaching architectures
and strategies.</p>
      </sec>
      <sec id="sec-3-10">
        <title>Detailed scan- The scan is incomplete, The scan is present but</title>
        <p>ning of the in- unclear, or lacks a log- could be more detailed
tervention ical progression of ac- or better structured in
tivities. some parts.</p>
        <sec id="sec-3-10-1">
          <title>Critical reflection on the redesign process</title>
        </sec>
      </sec>
      <sec id="sec-3-11">
        <title>There is a lack of crit- Includes some reflec</title>
        <p>ical reflection on the tion on the changes,
changes made or the but the analysis could
justifications are su- be more thorough.</p>
        <p>perficial.
enable a deeper integration of diferent data formats, enhancing its utility in complex applications.</p>
        <p>Link: https://openai.com/index/hello-gpt-4o/
2. Gemini 1.5 Pro Latest: a large language model developed by DeepMind (Google), is natively
multimodal and supports an extended context window of up to two million tokens, which is
currently the longest of any large-scale foundation model. This expansion in token capacity allows
for processing more extensive sequences of data, thereby increasing its utility in tasks that require
long-term contextual understanding. Link: https://deepmind.google/technologies/gemini/pro/.
3. Claude 3.5 Sonnet: developed by Anthropic, excels in the ability to understand nuanced language,
humour, and complex instructions. It is designed to generate high-quality content in a
relatable, natural tone, showing marked improvements in areas such as writing and human-centric
communication. Link: https://www.anthropic.com/news/claude-3-5-sonnet.
4. Mistral Large (2402): is designed to excel in complex reasoning tasks, particularly in multilingual
contexts. The model is highly efective in text understanding, transformation, and code generation.
It demonstrates top-tier performance in handling sophisticated reasoning challenges, making it
a robust tool for both natural language processing and technical tasks. Link: https://mistral.ai/
news/mistral-large/.
5. Open Mixtral 8x22B (2404): is one of the latest model developed by Mistral, featuring a sparse
Mixture-of-Experts (SMoE) architecture. Despite its large size, with 141 billion parameters, only
39 billion parameters are actively engaged during processing, optimising both performance and
cost eficiency. This approach sets new standards in the AI community for balancing model
complexity with computational resource usage. Link: https://mistral.ai/news/mixtral-8x22b/.
6. Llama 3.1 70B Instruct Turbo: developed by Meta, is a 70-billion parameter language model
designed for instruction-following tasks. The model is optimised to improve interactions where clear
guidance or step-by-step reasoning is required, positioning it as an efective tool for applications
in both academic and practical domains. Link: https://ai.meta.com/blog/meta-llama-3-1/
7. Qwen2 72B Instruct: developed by Alibaba Cloud, is a 72-billion parameter language model
optimised for instruction-based tasks. It integrates the latest advancements in generative AI,
ofering improved eficiency in tasks ranging from conversational AI to complex text generation
and reasoning. Its design caters specifically to high-performance needs in both commercial and
research applications. Link:https://www.alibabacloud.com/en/solutions/generative-ai/qwen?_p_
lc=1</p>
        <p>All these LLMs can "understand" and write in Italian, but it cannot be ruled out that performance
in English may be diferent (presumably better, since most of the training is done in that language).
Mistral’s models were added for their specific training with European languages that renders them
“natively fluent in English, French, Spanish, German, and Italian, with a nuanced understanding of
grammar and cultural context”. Privacy shouldn’t be a concern since there is no data saved to LLM
provider servers due to our use of API on a local instance of big-AGI. All the conversations are saved
only locally.</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>4. Prompting</title>
      <p>This study aimed to understand which models could be used by university educators (and, potentially,
other educators) to assess students’ products. For this reason, overly sophisticated prompting techniques
were not used; instead, what an educator might do by providing clear instructions and giving the
necessary context data for evaluation was employed. The LLMs systems were promoted through the
following instruction (originally written in Italian). The first one is the System Prompt:
You are an experienced and impartial university lecturer. Your job is
to assess the quality of student assignments according to a specific
assessment rubric.
**How to respond to requests:**
* Do not express personal opinions or subjective judgements. *
Focus exclusively on the criteria provided in the rubric. * Provide
a fair and impartial assessment based on the task’s adherence to
the criteria. * Carefully review the student’s entire paper before
beginning the assessment. * Offer constructive suggestions as to
how the student might improve. * Uses clear and concise language.
* Justify the marks awarded with specific references to the paper
and the rubric. * In your assessment, take into account that the
students only had 2 hours for planning.
**Request format
Each request will include:
* **The student’s assignment:** The text of the assignment you are
to assess. * **The grading rubric:** A list of criteria with
descriptions for each grade level.
**Response format:**
Your answer should follow this format:
**Title of the paper (also called title of the paper) as it appears
in the document: [insert title here]**.
**Total score:** [Insert total score here].
**Scoring breakdown:**
| Criterion | Score | Comments |—|—|—| | [Criterion 1] | [Score]
| [Comments with specific examples from the task] | | [Criterion
2] | [Score] | [Comments with specific examples from the task] | |
[Criterion 3] | [Score] | [Comments with specific examples from the
task] | | ... | ... | ... |
**Suggestions for improvement
* [Suggestion 1] * [Suggestion 2] * ... **Answer following the answer
format provided above.**</p>
      <p>The second is the prompt that were given to the LLMs to assess the products (originally written in
Italian):</p>
      <p>Evaluate the attached teaching design (student task) that was
created by a group of students from the secondary school teaching
qualification course. The key competence of this assignment lay in
being able to design a teaching intervention that makes effective
use of teaching architectures and strategies. In particular, the
group’s competence in terms of redesign and depth of reflection is
taken into account with respect to previous instructional design. At
the same time, the instructional design had to prove effective in
achieving the goals they set themselves. Take into account that the
students only had 2 hours to design. Use the evaluation rubric below
to assess:
&lt;Starting teaching design evaluation rubric&gt;</p>
      <p>Understanding
and
application
of
teaching
Evaluation rubric:
Criterion 1
architectures:
- Insufficient (award 1 point): Limited understanding and
applications not always appropriate. - Sufficient (award 2 points):
Basic understanding with some uncertainties in application. - Good
(awarded 3 points): Correct application and good understanding.
Excellent (awarded 4 points): Thorough understanding and innovative
and relevant application.</p>
      <p>Criterion 2 - Selection and implementation of teaching strategies:
- Insufficient (award 1 point): Limited or not always adequate
strategies. - Sufficient (award 2 points): Relevant strategies but
implementation can be improved. - Good (award 3 points): Strategies
appropriate and related to the objectives. - Excellent (award 4
points): Highly effective, diverse and well adapted strategies.
Criterion 3 - Definition of learning objectives:
- Insufficient (award 1 point): Vague or non-measurable objectives.
- Sufficient (award 2 points): Objectives present but not very
specific. - Good (award 3 points): Well-defined and generally
aligned objectives. - Excellent (award 4 points): Clear, specific,
measurable and perfectly aligned objectives. Criterion 4 - Detailed
scanning of the intervention
- Insufficient (award 1 point): Incomplete or unclear scan.
Sufficient (award 2 points): Scan present but can be improved in
structure. - Good (awarded 3 points): Clear and well-structured scan.
- Excellent (awarded 4 points): Detailed, logical and well-structured
scanning.</p>
      <p>Criterion 5 - Critical reflection on redesign:
- Insufficient (award 1 point): Lack of critical reflection
or superficial justifications. - Sufficient (award 2 points):
Reflection present but not very thorough. - Good (awarded 3 points):
Good reflection with clear connections. - Excellent (award 4 points):
Deep and critical reflection, clear justifications.
&lt;end of assessment rubric&gt;
**Total score:**
**Scoring distribution:**</p>
      <p>The prompt was sent simultaneously to all the LLMs involved. Through big-AGI the authentic
task document was attached in PDF format. A zero-shot prompting procedure was used for all LLMs,
meaning that no examples of human task assessments were given to the models. It is possible for a
university instructor to provide an example that can enhance the quality of LLM assessments, however,
the goal in this instance was to choose the most suitable models for this type of evaluation, not to find
methods for optimising the results. Finally, an “eighth” LLM evaluator has been added, which uses
the “Beam” function of big.AGI software. All the nuanced answers from the 7 LLM have been sent for
consideration and synthesis to GPT-4o, resulting in an eighth assessment that considers the feedback
from all seven LLMs.</p>
    </sec>
    <sec id="sec-5">
      <title>5. Models’ Settings</title>
      <p>In order to set a common limit of length for every model’s answers, all of them have been set through
big.AGI API controls to 8128 tokens maximum. Also, the temperature was set to 0.2, that should ensure
quite strict adherence to the instructions yet leave some room for creativity in answers.
5.1. Attention to Tokens and Context
Understanding tokens and context is crucial when using a Large Language Model (LLM). Tokens can be
simplified as units of text that might consist of a word, part of a word, or even a single character. The
characteristics of tokens can vary between models. However, it is generally safe to assume that, on
average, English might require one to one and a half tokens per word, and Italian might need one and a
half to two tokens per word. The context window, another essential concept, represents the number
of tokens a language model can consider simultaneously when generating responses. This context
depends on the model used and the available memory. Exceeding a model’s context window could
cause errors if it happens in a single prompt or, in a more extended conversation, the model might start
ignoring the earlier parts of the dialogue to make room for more recent inputs. Therefore, preserving
context is vital for generating coherent and relevant responses. It is important to note that not only
GPT-4o</p>
      <sec id="sec-5-1">
        <title>Claude 3.5 Sonnet</title>
      </sec>
      <sec id="sec-5-2">
        <title>Gemini 1.5 Pro Latest</title>
      </sec>
      <sec id="sec-5-3">
        <title>Mistral Large (2402)</title>
        <p>Mixtral 8x22B (2404)</p>
      </sec>
      <sec id="sec-5-4">
        <title>Meta Llama 3.1 70B Instruct Turbo Qwen2 72B Instruct</title>
        <p>6. Method of Analysis
the user’s prompts consume context, but the model’s responses do as well. To preserve the context
window, some LLMs platforms impose a character limit on the prompts that can be sent and on the
length of the generated responses, which are shorter than the maximum context window. Contrasting
this replication study with the original one, it can be noted that context windows are decidedly wider
than the ones that were found in LLMs one year ago, posing less of threat to the coherence of the
assessment. Below Table 2 illustrates the maximum context window size for each of the models used:</p>
        <sec id="sec-5-4-1">
          <title>Large Language Model (versions available in Italy, September 2024)</title>
        </sec>
        <sec id="sec-5-4-2">
          <title>Context Window (in tokens)</title>
          <p>The analysis method for evaluating the data involved examines the levels assigned by each evaluator
(both LLMs and humans) to the various criteria of the rubric for each of the 35 group products. Each of
the seven evaluators assigned a level to each of the five criteria for every product, resulting in each
evaluator assigning a level to a total of 175 criteria. Several statistical techniques were employed to
extract insights from the data, including Principal Component Analysis (PCA), analysis of standard
deviation, and the creation of a disagreement index among evaluators. Microsoft Excel and JASP (based
on R) were used for the statistical analyses.</p>
        </sec>
      </sec>
    </sec>
    <sec id="sec-6">
      <title>7. Results</title>
      <p>The consistency of the assessment for diferent models has been tested over the evaluation of three
random tasks from the sample for three times each by each one of the models. For this test, each LLM
assessed a total of 45 criteria.</p>
      <p>From these tests, the following behaviours were observed:
• GPT-4o, Gemini 1.5 Pro Latest and Claude 3.5 Sonnet were extremely consistent, with only one
instance of a diferent assessment for one criterion, by just one point.
• Mistral Large 2402 was perfectly consistent with zero instances of diferent assessments.
• Open Mixtral 8x22B (2404) and Qwen2 72B Instruct were quite consistent, with five instances of
diferent assessments of single criterion by one point.
• Llama 3.1 70B Instruct Turbo: was the fairly inconsistent, with 19 instances of diferent criteria
assessment by one point.
7.1. Principal Component Analysis
The first analysis conducted, in addition to descriptive data, was the PCA, a dimensionality reduction
technique that allows the identification of latent variables within the data and that can represent a
general model of the data. Three principal components were identified from the PCA conducted on the
assessment data (Table 3).</p>
      <p>The first component (RC1) is formed by evaluators e1, e3, e4, e5, e6, e7, e8 loadings, which correspond
respectively to the LLMs GPT-4o, Claude 3.5 Sonnet, Mistral Large, Mixtral 8x22B, Llama 3.1 70B, Qwen2
72B and the merge of LLMs opinion. The second component (RC2) comprises those of e2, e9, e10 and
e11 corresponding to Gemini 1.5 Pro and human evaluators 1, 2 and 3. As can be appreciated in Figure 1,
Both GPT-4o and Claude 3.5 Sonnet contributes mainly to RC1 component but also to RC2. Gemini Pro
1.5 on the other hand, contributes only to RC2 component (the tiny loading to RC1 is negative). Trying
to name the identified components, RC1 could be called "LLM Evaluation Pattern" and RC2 "Human
Evaluation Pattern".
7.2. Analysis of Standard Deviation of Grades by Product and Assessment Criterion
To understand how assessments difered from criterion to criterion and from evaluator to evaluator, an
analysis was conducted on the standard deviation (SD) of the diferent variables of the study. The criteria,
numbered or abbreviated in some of the graphs, are those listed in Table 4. Firstly, an efort was made
to identify which assessment criteria had the slightest and the most SD (Table 4) to understand which
were assessed more consistently by all evaluators. The criteria with the minimum SD across all products
is Criterion 4 and 1 (“Detailed scanning of the intervention” and “Understanding and application of
teaching architectures”), with an average of about 0.5. This suggests a high level of agreement among
evaluators in assessing the quality and details of the detailed activities envisaged in the educational
design and the understanding and correct application of the teaching architectures at their bases. On the
other hand, the criterion with the maximum SD among all activities is Criterion 5 (Critical reflection on
redesign), with an average of about 0.8. This indicates a higher level of disagreement or inconsistency
in how evaluators assessed the quality of teacher’s critical reflection about their past activities and the
way in which they tried to improve them.
7.3. Agreement Index
An "Agreement Index" (AIdx) was developed to obtain a more robust metric and better understand
which evaluators assigned more similar scores for the various criteria. This index combines the average
diference between the scores assigned to a criterion and the variability of this diference. It was
calculated to understand which evaluators are most similar to the human ones for each criterion. While
LLMs evaluators are treated individually, the human benchmark is an average of the human evaluators’
(e9, e10 and e11) assessments. It is constructed as follows:
Therefore:</p>
      <sec id="sec-6-1">
        <title>Critical reflection on redesign</title>
      </sec>
      <sec id="sec-6-2">
        <title>Selection and implementation of teaching strategies</title>
      </sec>
      <sec id="sec-6-3">
        <title>Definition of learning objectives</title>
      </sec>
      <sec id="sec-6-4">
        <title>Understanding and application of teaching architectures Detailed scanning of the intervention</title>
        <p>• The "Average Diference" is the absolute average diference in scores assigned between the
evaluator in question and the average of human evaluators across all tasks and criteria.
• The "Variability of the Diference" is the standard deviation of the diference scores between the
tested evaluator and the reference evaluator, reflecting how consistent these diferences are across
diferent tasks and criteria.</p>
        <p>AIdx is calculated individually for each evaluator. It provides a single measure that encapsulates the
average magnitude of evaluation diferences relative to the reference evaluator and the consistency of
such diferences. A lower value indicates a more significant overall agreement in evaluation relative to
the human evaluator. The highest possible value for the index for an evaluator would be achieved if
they constantly evaluated at the maximum diference from the human evaluators (3 points).</p>
        <p>The LLM evaluator who provided assessments most similar to the average of human evaluators
(calculated through the AIdx) is GPT-4o, followed at a negligible distance by Claude 3.5 Sonnet. On the
other hand, the LLM evaluator with the worst AIdx is Qwen2 72B (Table 5).</p>
        <p>Focusing on the single criterion (Table 6), it can be noted how AIdx with other evaluators vary from
criterion to criterion. Unexpectedly, Qwen2 72B, the worst on the general AIdx with the “average”
human evaluator, is the single model that is most human-like in three criteria out of five. Its main
problem is that it assessed in a very diferent way from humans the most dificult criterion: criterion
number 5 “Critical reflection on redesign” (Table 6). It also did not fare optimally in criterion number
3 “Definition of learning objectives”. Other LLMs like GPT-4o and Claude 3.5 Sonnet, as well as the
Merge of the diferent LLMs feedback, keep a good Agreement Index across the board.
7.4. Assessment correlations among LLM and Human evaluators
As reported in Table 7 the model with higher correlation with human evaluation is by far Gemini 1.5
Pro (r = 0.84), followed at a distance by Claude Sonnet 3.5 (r = 0.66), then by GPT-4o (r = 0.59), the merge
of the LLMs feedback by GPT-4o (r = 0.58) and Llama 3.1 70B (r = 0.45). This suggests that Gemini 1.5
Pro’s pattern of scores across the criteria is the most similar to that of the human evaluators.</p>
        <p>But what happens excluding single criteria from the correlation analysis? That could help in
understanding what criteria makes the assessment “human” and what LLMs struggle with:
• Excluding Criterion 3 (Definition of learning objectives) : When Criterion 3 is excluded
all LLMs’ correlation indexes significantly improve. Noticeably, for Mistral Large and Qwen2
72B the jump is from being hardly correlated, or not at all (r = 0.27 and -0.06 respectively), to
being significantly correlated (r = 0.86 and 0.88). Excluding Criterion 3 also significantly reduces
the correlation of Gemini 1.5 Pro suggesting that this was the Criterion that it got right and
mostly contributed to its excellent general correlation to humans’ assessment. This suggests that
Criterion 3 may be peculiarly human-like in its application, which these models struggle to mimic
accurately. The high increase implies that Criterion 3 might involve a complex judgment that
those models are incapable to handle or contextual information that is not being passed to the
model.
• Excluding Criterion 2 (Selection and implementation of teaching strategies): Excluding
Criterion 2 doesn’t change LLMs correlation with human assessment, except for Gemini 1.5 Pro
Latest. Gemini shows an almost perfect correlation of r = 0.99 when Criterion 2 is excluded,
which is remarkable, but, even in this case, this criterion doesn’t seem to be crucial.
• Excluding Criterion 5 (Critical reflection on redesign) : The exclusion leads to a substantial
increase in correlation for Mistral Large (from r = 0.27 to 0.88) and a notable improvement for
several other models. This criterion, similarly to Criterion 3, may also represent aspects of human
judgement that are challenging for models to replicate accurately.</p>
      </sec>
    </sec>
    <sec id="sec-7">
      <title>8. Discussion</title>
      <p>
        Regarding the goal of understanding whether educators without expertise in machine learning can
employ current Large Language Models (LLMs) to assess students’ written authentic tasks using
assessment rubrics, the analyses have revealed several interesting elements:
• Diferently from a previous iteration of the study [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ], all the models have enough context window
to perform this task.
• From the PCA, it appears that human evaluators generally have a diferent pattern of evaluation
compared to LLMs.
• In contrast with the evaluation pattern, the Agreement Index (AIdx) measures both the magnitude
of the score diferences and their consistency. A high Agreement Index value suggests significant
discrepancies between the model’s scores and the average human scores, despite possibly similar
trends in the pattern. Transforming in percentage the AIdx of each model referred to the
average human an accuracy metric has been achieved. This helps to better visualise each model’s
performance (Fig. 2 and Fig. 3)
• Only Llama 3.1 70B was inconsistent in the repeated assessment of the same task.
• Gemini 1.5 Pro is the LLM model with the evaluation pattern more similar (with by far the higher
correlation) to the human’s (see Table 7). It is the only model that in the PCA results only in the
component of human assessment (Fig. 1). On the other hand, its AIdx was the second worst, just
before Qwen2 72B (Table 5, Fig. 2).
• GPT-4o and Claude 3.5 Sonnet have evaluation patterns not too dissimilar from the human’s (Fig
1, Table 7) and on average attribute marks more similar to humans than any other model (Table
5).
• Llama 3.1 70B Instruct was the best of the open models, and the fourth in total (Table 5), after
the already mentioned three proprietary models. It behaved quite well in the correlation index
with the humans’ assessment pattern with a moderate correlation (Table 7) and has a good AIdx.
The problem with this model is the inconsistency of the assessment of the same task, where it
“changed its mind” 19 times out of 45. It would be interesting to understand if that inconsistency
has to do with the quantisation applied by Together AI, the API provider used.
• Mixtral 8x22B and Mistral Large fared similarly with patterns quite dissimilar to the human’s
and AIdx which are pretty decent (similar to Llama’s). The correlation of Mistral Large with the
human pattern of evaluation, when Criterion 3 is removed, is the second highest, thus giving
reasons to follow it closely and keep it in the test pool.
• Qwen2 72B, an open LLM, would have been by far the best LLM overall (and Mistral Large would
have been the second) if it weren’t for Criterion 3. Criterion 3 posed a grave problem for Qwen2
both from the assessment pattern and from the AIdx point of view (Fig. 2 and Fig. 3).
• Criterion 3 (Definition of learning objectives), in a larger part, and Criterion 5 (Critical reflection
on redesign), in a smaller part, appear to be the most discriminative criteria in terms of capturing
what makes the human evaluation pattern unique for this assessment task (Fig. 2 and Fig. 3). These
criteria likely involve nuances and complexities in judgment that are particularly human-like
and challenging for LLMs to capture accurately, or the authors might have failed to provide all
the relevant contextual information regarding these criteria to LLMs. This last hypothesis seems
relevant because, in the previous iteration of the study, this same criteria was the easiest one for
LLM to assess in a human-like manner [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ].
      </p>
      <p>Based on the available data, it appears that the more suitable LLMs for the assessing students’
authentic tasks using an assessment rubric are Claude 3.5 Sonnet and GPT-4o. That is because they fared
well both on the assessment pattern (PCA and correlations) and in the agreement index (magnitude
and consistencies of scores). On the other hand, Gemini 1.5 Pro is the one that had by far the most
human-like assessment pattern, but fell short on the AIdx, attributing marks that were very diferent
from the humans’.</p>
      <p>Qwen2 7B and Llama 3.1 70B deserve a mention as they are open models, and if not for some flaw
would have been at the level (or better than) the aforementioned proprietary models. Llama has a problem
of inconsistency of the marks assigned for each criterion, while Qwen2 really just got one criterion very
wrong. It might be useful to know that for both of them, Together AI (https://www.together.ai/) was
used as an API provider. It applies quantisation of Floating Point 8-bit (FP8) for Llama 3.1 70B Instruct
Turbo, while Qwen2 72B Instruct is run at full-precision Floating Point 16-bit (FP16).</p>
      <p>Human evaluators have a pattern of evaluation (see the PCA) that can be usually distinguished from
the LLMs’ one, but Gemini 1.5 Pro, if not for its very diferent score attribution, has very similar patterns.
It is interesting to note that human evaluators among themselves have diferent score attributions (see
Table 5 and Figure 3), but as for critical criteria they assess similarly.</p>
      <p>
        All considered, presently, none of the LLMs can be used for autonomous evaluation for all criteria,
especially regarding the more complex and the less contextualised ones. This confirms what Webb [ 33]
highlighted. However, Claude 3.5 Sonnet, GPT-4o and, with some caution, Qwen2 72B Instruct have the
potential to be used as solid support for evaluation for the summative evaluation level as described in
the AI-MAAS (AI-Mediated Assessment for Academics and Students) model [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ].
      </p>
    </sec>
    <sec id="sec-8">
      <title>9. Conclusions</title>
      <p>
        The fundamental question of this study was whether and which current Large Language Models (LLMs)
can be used by university educators (but it applies to other educators and instructors, too), even those
without technical experience, to assess student-written authentic products in the presence of open tasks
and questions using assessment rubrics. Indeed, using these technologies could make assessment more
sustainable and scalable, allowing for more consistent alignment with declared learning objectives. This
study has allowed us to determine that Claude 3.5 Sonnet, GPT-4o and, with some caution, Qwen2 72B
Instruct have the potential to be used as solid support for summative evaluation. According to this study,
the use of LLMs can be beneficial, but only if they are used under proper supervision. They should
be seen as assistance for university educators and not as a substitute for assessments. The available
data does not indicate that they are reliable enough to perform assessments independently, even if they
are getting close to it. In fact, some criteria that is too complex or needs additional information about
the context or specific subject can be evaluated in a way that is not in line with human assessment.
This finding confirms the guidelines as stated by Miao et al. [ 31] and Webb [33]. The limitations of the
present study lie in the sample size of student products that need to be significantly increased, as well
as the number of human expert evaluators and the disciplines involved in the tests. The assessment
rubric can also be optimised and, especially for the most critical criteria (such as Criterion 3), it would
be important to experiment on its formulation to understand if it could have been a human error in
defining the criteria that made it dificult to interpret by the LLMs. The idea behind this study is that it
should be expanded and updated on a rolling basis to adjust the discussion and bring useful novelties
into the assessment practice. Future evolutions of the study might include multi-shot prompting and
the evaluation of textual feedback and assessment to tasks. Feedback that could be provided during the
assessment for each of the criteria provided in a rubric deserve particular exploration [
        <xref ref-type="bibr" rid="ref11 ref9">11, 32, 9, 42</xref>
        ].
      </p>
    </sec>
    <sec id="sec-9">
      <title>Acknowledgments</title>
      <p>The authors would like to thank Elena Benini, PhD Student, for her contribution on the assessment of
students’ authentic tasks. Thanks also to Prof. Massimo Stella for the fruitful discussion about statistical
methods. Both of them work at the Department of Psychology and Cognitive Sciences of the University
of Trento.
search ifnds, 2023. URL: https://www.bloomberg.com/company/press/
generative-ai-to-become-a-1-3-trillion-market-by-2032-research-finds/.
[14] G. Hammond, Big tech outspends venture capital firms in AI investment frenzy, 2023. URL:
https://www.ft.com/content/c6b47d24-b435-4f41-b197-2d826cce9532.
[15] Y. S. Lee, T. Kim, S. Choi, W. Kim, When does AI pay of? AI-adoption intensity, complementary
investments, and R&amp;D strategy, Technovation 118 (2022) 102590. doi:10.1016/j.technovation.
2022.102590.
[16] D. Long, B. Magerko, What is AI literacy? Competencies and design considerations, in: Proceedings
of the 2020 CHI Conference on Human Factors in Computing Systems, 2020, pp. 1–16.
[17] D. T. K. Ng, J. K. L. Leung, S. K. W. Chu, M. S. Qiao, Conceptualizing AI literacy: An exploratory
review, Computers and Education: Artificial Intelligence 2 (2021) 100041.
[18] G. Biagini, S. Cuomo, M. Ranieri, Developing and validating a multidimensional AI literacy
questionnaire: Operationalizing AI literacy for higher education, in: Proceedings of the First
International Workshop on High-Performance Artificial Intelligence Systems in Education, AIxEDU
2023, Aachen, 2023. URL: https://ceur-ws.org/Vol-3605/.
[19] D. Cetindamar, K. Kitto, M. Wu, Y. Zhang, B. Abedin, S. Knight, Explicating AI literacy of
employees at digital workplaces, IEEE Transactions on Engineering Management 71 (2024)
810–823. doi:10.1109/TEM.2021.3138503.
[20] S.-C. Kong, W. M.-Y. Cheung, G. Zhang, Evaluating an Artificial Intelligence Literacy Programme
for Developing University Students’ Conceptual Understanding, Literacy, Empowerment and
Ethical Awareness, Educational Technology &amp; Society 26 (2023) 16–30.
[21] B. Wang, P.-L. P. Rau, T. Yuan, Measuring user competence in using artificial intelligence: Validity
and reliability of artificial intelligence literacy scale, Behaviour &amp; Information Technology 42
(2023) 1324–1337. doi:10.1080/0144929X.2022.2072768.
[22] P. Weber, M. Pinski, L. Baum, Toward an Objective Measurement of AI Literacy, in: PACIS 2023</p>
      <p>Proceedings, 2023. URL: https://aisel.aisnet.org/pacis2023/60.
[23] UNESCO, Report “Guidance for generative AI in education and research”, 2023. URL: https:
//www.unesco.org/en/articles/guidance-generative-ai-education-and-research.
[24] A. Gerdes, A participatory data-centric approach to AI ethics by design, Applied Artificial</p>
      <p>Intelligence 36 (2022). doi:10.1080/08839514.2021.2009222.
[25] C. Jang, Coping with vulnerability: The efect of trust in ai and privacy-protective behaviour on the
use of ai-based services, Behaviour &amp; Information Technology (2023). doi:10.1080/0144929X.
2023.2246590.
[26] A. Majeed, S. O. Hwang, When AI Meets Information Privacy: The Adversarial Role of AI in Data</p>
      <p>Sharing Scenario, IEEE Access 11 (2023) 76177–76195. doi:10.1109/ACCESS.2023.3297646.
[27] P. Samuelson, Generative AI meets copyright, Science 381 (2023) 158–161. doi:10.1126/science.</p>
      <p>adi0656.
[28] M. A. Yeo, Academic integrity in the age of Artificial Intelligence (AI) authoring apps, TESOL</p>
      <p>Journal 14 (2023) e716. doi:10.1002/tesj.716.
[29] V. van Oijen, AI-generated text detectors: Do they work?, 2023. URL: https://communities.surf.nl/
en/ai-in-education/article/ai-generated-text-detectors-do-they-work.
[30] D. Weber-Wulf, A. Anohina-Naumeca, S. Bjelobaba, T. Foltýnek, J. Guerrero-Dib, O. Popoola,
P. Šigut, L. Waddington, Testing of detection tools for AI-generated text, International Journal for
Educational Integrity 19 (2023) 26. doi:10.1007/s40979-023-00146-z.
[31] F. Miao, W. Holmes, H. Ronghuai, Z. Hui, AI and education: Guidance for policy-makers, Technical</p>
      <p>Report, UNESCO, 2023. URL: https://unesdoc.unesco.org/ark:/48223/pf0000376709.
[32] E. Sabzalieva, A. Valentini, ChatGPT and artificial intelligence in higher education: Quick start
guide, Technical Report, UNESCO, 2023. URL: https://unesdoc.unesco.org/ark:/48223/pf0000385146.
[33] M. Webb, A Generative AI Primer, Technical Report, National Centre for AI, 2023. URL: https:
//nationalcentreforai.jiscinvolve.org/wp/2024/01/02/generative-ai-primer/.
[34] Russell Group, New principles on use of AI in education, 2023. URL: https://russellgroup.ac.uk/
news/new-principles-on-use-of-ai-in-education/.
[35] GTnum, Intelligence artificielle et éducation: Apports de la recherche et enjeux pour les politiques
publiques, 2023. URL: https://edunumrech.hypotheses.org/8726.
[36] M. A. Cardona, R. J. Rodríguez, K. Ishmael, Artificial Intelligence and the Future of Teaching and
Learning: Insights and Recommendations, Technical Report, 2023. URL: https://policycommons.
net/artifacts/3854312/ai-report/4660267/.
[37] UCL, Using generative AI (GenAI) in learning and teaching, 2023. URL: https://www.ucl.ac.uk/
teaching-learning/publications/2023/sep/using-generative-ai-genai-learning-and-teaching.
[38] Z. Swiecki, H. Khosravi, G. Chen, R. Martinez-Maldonado, J. M. Lodge, S. Milligan, N. Selwyn,
D. Gašević, Assessment in the age of artificial intelligence, Computers and Education: Artificial
Intelligence 3 (2022) 100075. doi:10.1016/j.caeai.2022.100075.
[39] P. P. Martin, D. Kranz, P. Wulf, N. Graulich, Exploring new depths: Applying machine learning
for the analysis of student argumentation in chemistry, Journal of Research in Science Teaching
(2023). doi:10.1002/tea.21903.
[40] E. Kasneci, K. Seßler, S. Küchemann, M. Bannert, D. Dementieva, F. Fischer, et al., ChatGPT for
good? On opportunities and challenges of large language models for education, Learning and
Individual Diferences 103 (2023) 102274.
[41] A. Lepage, N. Roy, A review of the literature from 1970 to 2022 on the roles of teachers and
artificial intelligence in the field of AI in education, Médiations et Médiatisations 16 (2023) 30–50.
doi:10.52358/mm.vi16.304.
[42] A. Tamkin, M. Brundage, J. Clark, D. Ganguli, Understanding the Capabilities, Limitations,
and Societal Impact of Large Language Models, arXiv preprint arXiv:2102.02503 (2021). URL:
http://arxiv.org/abs/2102.02503.
[43] O. Koraishi, Teaching English in the Age of AI: Embracing ChatGPT to Optimize EFL Materials
and Assessment, Language Education and Technology 3 (2023) Article 1.
[44] F. Ali, D. Choy, S. Divaharan, H. Y. Tay, W. Chen, Supporting self-directed learning and
selfassessment using TeacherGAIA, a generative AI chatbot application: Learning approaches and
prompt engineering, Learning: Research and Practice 9 (2023) 135–147. doi:10.1080/23735082.
2023.2258886.
[45] F. Ouyang, T. A. Dinh, W. Xu, A Systematic Review of AI-Driven Educational Assessment
in STEM Education, Journal for STEM Education Research 6 (2023) 408–426. doi:10.1007/
s41979-023-00112-x.
[46] J. Biggs, Enhancing teaching through constructive alignment, Higher Education 32 (1996) 347–364.
doi:10.1007/BF00138871.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>A.</given-names>
            <surname>Baytak</surname>
          </string-name>
          ,
          <article-title>The acceptance and difusion of generative artificial intelligence in education: A literature review</article-title>
          ,
          <source>Current Perspectives in Educational Research</source>
          <volume>6</volume>
          (
          <year>2023</year>
          )
          <article-title>Article 1</article-title>
          . doi:
          <volume>10</volume>
          .46303/ cuper.
          <year>2023</year>
          .
          <volume>2</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>S.</given-names>
            <surname>Elbanna</surname>
          </string-name>
          , L. Armstrong,
          <article-title>Exploring the integration of ChatGPT in education: Adapting for the future</article-title>
          ,
          <source>Management &amp; Sustainability: An Arab Review</source>
          <volume>3</volume>
          (
          <year>2023</year>
          )
          <fpage>16</fpage>
          -
          <lpage>29</lpage>
          . doi:
          <volume>10</volume>
          .1108/ MSAR-03-2023-0016.
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>A.</given-names>
            <surname>Extance</surname>
          </string-name>
          ,
          <article-title>ChatGPT has entered the classroom: How LLMs could transform education</article-title>
          ,
          <source>Nature</source>
          <volume>623</volume>
          (
          <year>2023</year>
          )
          <fpage>474</fpage>
          -
          <lpage>477</lpage>
          . doi:
          <volume>10</volume>
          .1038/d41586-023-03507-3.
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>S.</given-names>
            <surname>Roy</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Gupta</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Ray</surname>
          </string-name>
          ,
          <article-title>Adoption of AI ChatBot like Chat GPT in Higher Education in India: A SEM Analysis Approach, Economic Environment 4 (</article-title>
          <year>2023</year>
          )
          <fpage>130</fpage>
          -
          <lpage>149</lpage>
          . doi:
          <volume>10</volume>
          .36683/
          <fpage>2306</fpage>
          -
          <lpage>1758</lpage>
          / 2023-4-46/
          <fpage>130</fpage>
          -
          <lpage>149</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>N.</given-names>
            <surname>Saif</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S. U.</given-names>
            <surname>Khan</surname>
          </string-name>
          ,
          <string-name>
            <surname>I. Shaheen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Alotaibi</surname>
          </string-name>
          ,
          <string-name>
            <surname>M. M. Alnfiai</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          <string-name>
            <surname>Arif</surname>
          </string-name>
          ,
          <article-title>Chat-GPT; validating Technology Acceptance Model (TAM) in education sector via ubiquitous learning mechanism, Computers in Human Behavior (</article-title>
          <year>2023</year>
          )
          <article-title>108097</article-title>
          . doi:
          <volume>10</volume>
          .1016/j.chb.
          <year>2023</year>
          .
          <volume>108097</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>C. K.</given-names>
            <surname>Tiwari</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. A.</given-names>
            <surname>Bhat</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S. T.</given-names>
            <surname>Khan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Subramaniam</surname>
          </string-name>
          ,
          <string-name>
            <surname>M. A. I. Khan</surname>
          </string-name>
          ,
          <article-title>What drives students toward ChatGPT? An investigation of the factors influencing adoption and usage of ChatGPT, Interactive Technology</article-title>
          and
          <string-name>
            <given-names>Smart</given-names>
            <surname>Education</surname>
          </string-name>
          (
          <year>2023</year>
          ). doi:
          <volume>10</volume>
          .1108/ITSE-04-2023-0061.
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>F.</given-names>
            <surname>Kamalov</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D. Santandreu</given-names>
            <surname>Calonge</surname>
          </string-name>
          ,
          <string-name>
            <surname>I. Gurrib</surname>
          </string-name>
          ,
          <article-title>New era of artificial intelligence in education: Towards a sustainable multifaceted revolution</article-title>
          ,
          <source>Sustainability</source>
          <volume>15</volume>
          (
          <year>2023</year>
          )
          <fpage>12451</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>M.</given-names>
            <surname>Perkins</surname>
          </string-name>
          ,
          <article-title>Academic Integrity Considerations of AI Large Language Models in the Post-Pandemic Era: ChatGPT and Beyond</article-title>
          ,
          <source>Journal of University Teaching and Learning Practice</source>
          <volume>20</volume>
          (
          <year>2023</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>M.</given-names>
            <surname>Sullivan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Kelly</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Mclaughlan</surname>
          </string-name>
          ,
          <article-title>ChatGPT in higher education: Considerations for academic integrity and student learning</article-title>
          ,
          <source>Journal of Applied Learning &amp; Teaching</source>
          (
          <year>2023</year>
          ). doi:
          <volume>10</volume>
          .37074/ jalt.
          <year>2023</year>
          .
          <volume>6</volume>
          .1.17.
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>D.</given-names>
            <surname>Agostini</surname>
          </string-name>
          ,
          <article-title>Are large language models capable of assessing students' written products? A pilot study in higher education</article-title>
          ,
          <source>Research Trends in Humanities Education &amp; Philosophy</source>
          <volume>11</volume>
          (
          <year>2024</year>
          )
          <fpage>38</fpage>
          -
          <lpage>60</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>D.</given-names>
            <surname>Agostini</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Picasso</surname>
          </string-name>
          ,
          <article-title>Large language models for sustainable assessment and feedback in higher education</article-title>
          ,
          <source>Intelligenza Artificiale</source>
          <volume>18</volume>
          (
          <year>2024</year>
          )
          <fpage>121</fpage>
          -
          <lpage>138</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>T.</given-names>
            <surname>Babina</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Fedyk</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A. X.</given-names>
            <surname>He</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Hodson</surname>
          </string-name>
          ,
          <source>Firm Investments in Artificial Intelligence Technologies and Changes in Workforce Composition, Working Paper 31325, National Bureau of Economic Research</source>
          ,
          <year>2023</year>
          . doi:
          <volume>10</volume>
          .3386/w31325.
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <surname>Generative</surname>
            <given-names>AI</given-names>
          </string-name>
          <article-title>to become a $1.3 trillion market by 2032</article-title>
          , re-
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>