<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <issn pub-type="ppub">1613-0073</issn>
    </journal-meta>
    <article-meta>
      <title-group>
        <article-title>Analysis of AI-Generated Responses in Online Learning Discussions</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Zifeng Liu</string-name>
          <email>liuzifeng@ufl.edu</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Wanli Xing</string-name>
          <email>wanli.xing@coe.ufl.edu</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
          <xref ref-type="aff" rid="aff3">3</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Chenglu Li</string-name>
          <email>chenglu.li@utah.edu</email>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
          <xref ref-type="aff" rid="aff2">2</xref>
          <xref ref-type="aff" rid="aff4">4</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Educational Data Mining (EDM 2024)</institution>
          ,
          <addr-line>Atlanta, Georgia</addr-line>
          ,
          <country country="US">USA</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>Leveraging Large Language Models for Next-Generation Educational</institution>
        </aff>
        <aff id="aff2">
          <label>2</label>
          <institution>Text generation</institution>
          ,
          <addr-line>Explainable analysis, Llama3, Online Learning</addr-line>
        </aff>
        <aff id="aff3">
          <label>3</label>
          <institution>University of Florida (UF)</institution>
          ,
          <addr-line>Gainesville, FL 32611</addr-line>
          ,
          <country country="US">United States</country>
        </aff>
        <aff id="aff4">
          <label>4</label>
          <institution>University of Utah</institution>
          ,
          <addr-line>201 Presidents' Cir, Salt Lake City, UT 84112</addr-line>
          ,
          <country country="US">United States</country>
        </aff>
      </contrib-group>
      <abstract>
        <p>Large Language Models (LLMs) have demonstrated significant potential in enhancing online learning through features like automated question-answering systems. These systems can identify and respond to prevalent learner queries, thereby personalizing and enriching the online educational experience. However, there remains a notable gap in research regarding the performance of diferent models in educational settings, particularly in evaluating AI-generated content using explainable metrics. This study evaluates two distinct language models, Llama3-8B and GPT-2 small, to determine which better supports educational objectives in online environments. We used a t-test to statistically assess diferences in the Flesch-Kincaid readability scores between the two models and the results indicate that both models perform well in providing support for Massive Open Online Courses (MOOCs) learners regarding their readability. We further conducted an explainable analysis of the content generated by both models, the results show that although both models can generate certain support, there is still much improvement for the comprehension and accuracy of these generated contents. Our findings recommend that the selection of an AI model for educational use should be tailored to the specific learning goals and needs of the audience. Moreover, this study underscores the importance of applying explainable and transparent metrics for assessing AI-generated content to ensure its educational eficacy and ethical integrity.</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction and Prior Work</title>
      <p>
        Online discussion forums play a important role as both
pedagogical and social platforms in online learning
environments. Educational studies have consistently shown that
these forums support student learning by enhancing
engagement, improving critical thinking, and ofering increased
opportunities for reflection and collaborative knowledge
construction [
        <xref ref-type="bibr" rid="ref1 ref2">1, 2</xref>
        ]. Furthermore, the cognitive and
socioemotional support inherent in student interactions within
these forums has been found to boost both engagement and
academic achievement [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. Despite the recognized
importance of interactions within online learning communities
[
        <xref ref-type="bibr" rid="ref3 ref4 ref5">3, 4, 5</xref>
        ], online forums frequently experience low student
engagement. This lack of participation is primarily attributed
to anticipated non-responsiveness and the perceived
irrelevance of the discussed topics, which reduces students’
motivation to engage [
        <xref ref-type="bibr" rid="ref6 ref7">6, 7</xref>
        ]. Such sparse interactions can create
a vicious cycle of disengagement, where students may feel
isolated and less inclined to share. This low engagement
level in discussion forums not only deprives students of the
benefits of these crucial social settings but also contributes
to higher dropout rates [
        <xref ref-type="bibr" rid="ref10 ref11 ref8 ref9">8, 9, 10, 11</xref>
        ].
      </p>
      <p>
        To address the issue of low student participation,
researchers have developed numerous methods to enhance
interaction and engagement in online learning
communities. Some eforts have focused on creating key
learning indicators through collaboration with teachers,
using learning design frameworks, or through the iterative
refinement and empirical testing of educational systems
[
        <xref ref-type="bibr" rid="ref12 ref13 ref14 ref15">12, 13, 14, 15</xref>
        ]. These indicators yield automated, actionable
(HEXED) Workshop, co-located with the 17th International Conference on
0009-0005-5833-2141 (Z. Liu)
© 2024 Copyright for this paper by its authors. Use permitted under Creative Commons License
insights, such as the optimal timing for learning activities
and the sequencing of educational materials, supporting
tailored classroom management and enhancing students’
self-regulation. Furthermore, machine learning and
learning analytics have been employed to monitor student
behaviors, identify engagement patterns, and provide
personalized feedback [
        <xref ref-type="bibr" rid="ref16 ref17 ref18">16, 17, 18</xref>
        ]. These methods have been
applied to cluster and analyze texts posted by students online,
improving communication between teachers and students
and facilitating the tracking of individual and community
learning progress [
        <xref ref-type="bibr" rid="ref19 ref20">19, 20</xref>
        ], thereby allowing teachers to
customize interventions to meet specific needs and provide
timely support, creating a more engaging and interactive
learning environment.
      </p>
      <p>
        Recent advancements in large pre-trained language
models (LLMs) like GPT and Llama have significantly expanded
their use in educational applications [
        <xref ref-type="bibr" rid="ref21 ref22">21, 22</xref>
        ]. These models
ofer potential benefits for enhancing online learning
discussions. For instance, automated question-answering systems
can detect prevalent questions and concerns among learners
and provide automated responses, thereby improving the
online learning experience by delivering more personalized,
meaningful, and engaging educational and instructional
supports [
        <xref ref-type="bibr" rid="ref23">23</xref>
        ]. Research utilizing social support theory in
conjunction with LLMs to support online learners indicates
that AI-generated texts can provide a level of emotional and
community support comparable to that provided by human
interactions [
        <xref ref-type="bibr" rid="ref24 ref25">24, 25</xref>
        ]. Although there are other ways that
use LLMs to support online learning discussion, such as
extracting and visualizing key concepts and their
relationships from discussion threads [
        <xref ref-type="bibr" rid="ref26">26</xref>
        ], generating summaries
about long discussion threads [
        <xref ref-type="bibr" rid="ref27 ref28">27, 28</xref>
        ], and analyzing the
sentiment of posts to help identify students who might be
struggling or feeling disengaged [
        <xref ref-type="bibr" rid="ref29">29</xref>
        ]. However, automatic
text generation ofers unique advantages in supporting
online learning discussions by providing timely and
personalized responses and ofering immediate emotional support.
Despite the availability of these methods, the potential for
LLMs to provide emotional support in online discussion
forums remains an area requiring further exploration.
CEUR
      </p>
      <p>ceur-ws.org</p>
      <p>
        Despite their advanced capabilities, LLMs also
demonstrate limitations, such as generating inaccurate
information, producing ofensive outputs, and exhibiting biases [
        <xref ref-type="bibr" rid="ref30">30</xref>
        ].
These issues render LLMs unsuitable for universal
application without considerable modifications and transparent
explanations of their generated content [
        <xref ref-type="bibr" rid="ref31 ref32">31, 32</xref>
        ],
particularly in educational contexts. The extensive use of new
textgenerative models across various sectors has thus prompted
the need for robust evaluation metrics. The safety and
supportiveness perceived in the responses generated by these
models are heavily influenced by individual, contextual, and
cultural factors [
        <xref ref-type="bibr" rid="ref33">33</xref>
        ]. For example, responses that align with
certain biases might be viewed as supportive by individuals
holding those biases, potentially detracting from the
engagement and motivation of those committed to more widely
accepted values [
        <xref ref-type="bibr" rid="ref34">34</xref>
        ]. Moreover, maintaining the safety of
online discourse presents considerable challenges for both
technical and educational researchers. With estimates
suggesting that 5–30% of online discourse displays bias, varying
by domain, such biases can substantially afect the
behavior of data-driven LLMs [
        <xref ref-type="bibr" rid="ref35">35</xref>
        ]. This situation highlights the
imperative for implementing explainable metrics that can
accurately assess the trustworthiness and eficacy of
content produced by these models, especially in educational
settings.
      </p>
      <p>
        In this study, we investigate the application of
state-ofthe-art deep learning algorithms for text generation, aimed
at providing automated support for massive online
learning communities, specifically MOOCs. We assessed the
efectiveness of GPT-2 and Llama3 in generating text using
MOOC posts data1. GPT-2 has been recognized in previous
research as a leading model in text generation, noted for
its potential to provide emotional and community support
within large-scale online learning environments [
        <xref ref-type="bibr" rid="ref24 ref25">24, 25</xref>
        ].
Llama3, a robust deep-learning-based language model, was
released by Meta AI in April 2024 and is considered a
significant advancement in the field [
        <xref ref-type="bibr" rid="ref36">36</xref>
        ].
      </p>
      <p>
        One gap identified in prior research is the lack of
explainable evaluations of AI-generated texts. Consequently, the
primary objectives of this research are: (1) to determine the
extent to which deep learning-based text generation can
ofer eficient and meaningful textual support to learners in
massive online communities, and (2) to apply an
explainable metric to evaluate the generated texts and compare the
state-of-art models from a educational background. In this
context, we fine-tuned GPT-2 and Llama3-8B using 29,604
MOOC posts data and proposed a framework for explainable
and reference-free evaluation of AI-generated responses
using TIGERSCORE, a new metric developed by [
        <xref ref-type="bibr" rid="ref32">32</xref>
        ]. The
results indicate that there is potential for improvement in
utilizing these models for online learning forum support.
The main contributions of this study include:
• Applying new LLMs to provide online discussion
support for MOOCs and comparing the performance
of various popular LLMs in automating text
generation for online learning;
• Employing explainable and reference-free metrics
to evaluate the automated AI-generated support;
• Enhancing the understanding of the practical
limitations and capabilities of LLMs in educational
settings.
      </p>
      <sec id="sec-1-1">
        <title>1https://datastage.stanford.edu/StanfordMoocPosts/</title>
      </sec>
    </sec>
    <sec id="sec-2">
      <title>2. Method</title>
      <sec id="sec-2-1">
        <title>2.1. Data Source Description</title>
        <p>The Stanford MOOCPosts dataset comprises 29,604
anonymized learner forum posts from 11 public online
classes ofered by Stanford University, covering diverse
subjects such as Humanities, Medicine, and Education. This
dataset categorizes posts into questions, answers, and
opinions but also includes detailed annotations such as
sentiment ratings (1-7), levels of confusion (1-7), and urgency
ratings indicating the need for instructor attention (1-7).
We chose this dataset for its diverse contexts and extensive
representation of various academic disciplines, providing
a robust foundation for analyzing learner interactions and
engagement within online learning environments.
3060 GPU and 32 GB of RAM, utilizing Python 3. We
randomly selected 90% of the data from the dataset as training
data and used the remaining 10% for model evaluation.
During fine-tuning, the instruction provided to the model was
consistently: ”You are an online forum discussion support
assistant.” The process utilized 500 steps.</p>
        <p>
          Figure 2 depicts the architecture blocks of GPT-2 small
and Llama3-8B. Both GPT-2 small and Llama3-8B are
language models that leverage the Transformer architecture, a
framework based on an attention mechanism that does not
rely on Recurrent Networks to process sequences [
          <xref ref-type="bibr" rid="ref37">37</xref>
          ]. The
Transformer architecture utilizes Encoders to positionally
encode input sequences and Decoders to decode these
sequences, eficiently transforming one sequence into another.
2.3.1. GPT-2 small
GPT-2 is a transformer-based language model developed
by OpenAI and released in 2019. It is trained on a dataset
comprising 40GB of Internet texts, culminating in a model
with 1.5 billion parameters. Due to its exceptional
performance in text generation, OpenAI initially decided against
releasing the fully trained model, citing concerns over
potential malicious uses such as the generation of fake news
or automated email composition. In fact, GPT-2 achieved
state-of-the-art results in 7 out of 8 tested languages [
          <xref ref-type="bibr" rid="ref38">38</xref>
          ].
OpenAI released smaller versions of the model, including a
version with 124 million parameters and a medium version
with 345 million parameters, to the public for research and
experimentation. The small version includes 12 layers, and
the medium version contains 24 layers, as depicted in Figure
2. In this study, we utilize a small dataset of MOOC posts to
train both a GPT-2 small model. Using automatic evaluation
methods, we will select one language model to train with
the entire processed dataset. We employed code from 3 to
ifne-tune the GPT-2 small model.
        </p>
        <sec id="sec-2-1-1">
          <title>3https://github.com/minimaxir/gpt-2-simple</title>
          <p>
            2.3.2. Llamma3-8B
Llama is a decoder-only language model that processes
input sentences as ordered tokens and predicts subsequent
tokens. The Llama 3 model, released by Meta on April 18,
2024, was pretrained on over 15 trillion tokens sourced from
publicly accessible datasets. This corpus includes not only
publicly available instructional datasets but also over 10
million human-annotated examples. Meta has developed
the Meta Llama 3 series, a family of large language models
(LLMs) available in configurations of 8 billion and 70
billion parameters. These models, pretrained and
instructiontuned, are specifically optimized for dialogue applications
and have demonstrated superior performance over many
existing open-source chat models on standard industry
benchmarks. Llama 3 operates as an auto-regressive language
model. The instruction-tuned versions employ Supervised
Fine-Tuning (SFT)[
            <xref ref-type="bibr" rid="ref39">39</xref>
            ] and Reinforcement Learning with
Human Feedback (RLHF)[
            <xref ref-type="bibr" rid="ref40">40</xref>
            ] to enhance alignment with
human preferences concerning helpfulness and safety. We
ifne-tuned the Llama3-8B model using open resources 4.
          </p>
        </sec>
      </sec>
      <sec id="sec-2-2">
        <title>2.4. AI Generated Text Evaluation</title>
        <p>
          AI-generated text evaluation methodologies are traditionally
categorized into two primary types: intrinsic and extrinsic
methods. Intrinsic methods involve participants reading
and rating the texts based on aspects such as output quality
and user satisfaction. Extrinsic methods assess the impact
of the generated text on the success of user or system tasks
[
          <xref ref-type="bibr" rid="ref41">41</xref>
          ].
2.4.1. Readability: F-K Grade Level
The Flesch-Kincaid (F-K) Grade Level is an established tool
initially developed to assess the readability of texts for the
US Navy (Kincaid et al., 1975) and subsequently adopted as
a military standard. It has also gained widespread adoption
in academic research, utilized to evaluate the readability of
documents within medical and educational fields [
          <xref ref-type="bibr" rid="ref42 ref43">42, 43</xref>
          ].
Unlike some readability assessments, the F-K Grade Level
does not stipulate a minimum text length for evaluation.
   = 0.39 ∗    + 11.8 ∗    − 15.59 (1)
        </p>
        <p>In Equation 1, TWs refers to the total number of words
in the generated text.TSs refers to the total number of
sentences in the text. TSYs refers to the total number of syllables
in the text.</p>
        <sec id="sec-2-2-1">
          <title>4https://colab.research.google.com/drive/135ced7oHyt\</title>
          <p>
            dxu3N2DNe1Z0kqjyYIkDXp?usp=sharing
2.4.2. Explainable Error Analysis
In this study, we employed TIGERScore [
            <xref ref-type="bibr" rid="ref32">32</xref>
            ], a metric
trained to follow instructional guidance for explainable
and reference-free evaluation across a diverse range of
text generation tasks. Traditional automatic metrics
often face challenges such as dependency on reference texts,
domain specificity, and lack of transparent attribution. In
contrast, TIGERScore overcomes these limitations by being
instruction-driven and providing comprehensive error
analyses to precisely identify faults in generated texts. Unlike
other evaluation methods that yield obscure scores,
TIGERScore utilizes natural language instructions to conduct
detailed error analysis, thereby enhancing the interpretability
of its assessments.
          </p>
          <p>TIGERScore is constructed around three principal design
criteria: (1) It operates under instruction-driven protocols,
which enhances its flexibility and applicability to various
text generation challenges. For instance, the instructions
used in this study are consistent with those utilized for
ifne-tuning the GPT-2 small and Llama3-8B models. (2) It
dispenses with the need for references or exemplary
comparisons, facilitating an unbiased evaluation. (3) The model’s
outputs are highly interpretable; it not only identifies errors
but also provides a detailed analysis of each error, including
its location, nature, and the associated penalty.</p>
          <p>More specifically, TIGERScore 5 takes an instruction, an
associated input context, and a hypothesis output (in this
case, the AI-generated texts), which it evaluates for errors.
The evaluation identifies various mistakes, detailing their
specific locations, aspects, explanations, and the penalty
scores incurred. The aggregate of these deducted scores
constitutes the overall assessment of the output. Currently,
we employ TIGERScore for an explainable analysis of the
generated texts, with plans to incorporate human
evaluations in future research endeavors.</p>
        </sec>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>3. Results</title>
      <sec id="sec-3-1">
        <title>3.1. Finetune results</title>
        <p>The finetune training loss is shown in Figure 3. The loss of
GPT-2 small is relatively stable with minor fluctuations, in
contrast, Llama3-8B’s loss curve is more volatile, which may
suggest that the model is more sensitive to the training data
or encountered more optimization challenges during
training. Overall, the loss of GPT-2 small gradually decreases and
stabilizes, indicating that the model is progressively
converging throughout the training process. Although
Llama38B shows significant fluctuations, it also exhibits a general
downward trend in loss, especially in the first 100 steps. As
GPT-2 small is a relatively smaller model, it adapt or overfit
a smaller dataset more quickly, resulting in a smoother
decrease in loss. Llama3-8B, with its higher complexity, might
require more data or more sophisticated tuning strategies
to optimize, hence the larger initial fluctuations.</p>
        <p>Table 2 displays examples of texts generated by two
diferent models. Both models demonstrate a robust capacity to
produce contextually appropriate and engaging responses.
Llama3-8B’s responses generally exhibit greater creativity
and a deeper engagement with the topics, likely attributed
to its more sophisticated tuning and larger model size. In
contrast, GPT-2 small, although slightly more restrained in</p>
        <sec id="sec-3-1-1">
          <title>5https://huggingface.co/spaces/TIGER-Lab/TIGERScore</title>
          <p>its creative outputs, still efectively addresses the queries
with logically coherent and pertinent responses.</p>
        </sec>
      </sec>
      <sec id="sec-3-2">
        <title>3.2. F-K Grade Level Evaluation</title>
        <p>The Figure 4 presents a scatter plot of the F-K Grade Levels
for 200 texts generated by two diferent models, with blue
dots representing texts generated by GPT-2 small and red
dots representing texts from Llama3. The horizontal axis
denotes the text number, and the vertical axis represents
the corresponding F-K Grade Level value. We removed a
few extreme outliers for clarity. From the plot, it is evident
that the texts generated by GPT-2 small typically exhibit
lower readability scores, ranging from 4 to 10 (with an
average value of 6.67, shown in Table 3). In contrast, the texts
produced by Llama3-8B show a broader distribution of
readability scores spanning from 4 to 18, with an average value
of 11.13 which is more suitable for MOOC learners. Overall,
both models are capable of generating texts with relatively
high readability.</p>
        <p>We performed a t-test to examine the F-K readability
levels of texts generated by the two models, as shown in Table
3. Given the extremely small p-value (p&lt;0.001), we can
confidently reject the null hypothesis and accept the
alternative hypothesis that the F-K values of texts generated by
Llama3-8B are significantly higher than those generated by
GPT-2 small. This finding indicates that the texts produced
by GPT-2 small may exhibit easier readability.</p>
      </sec>
      <sec id="sec-3-3">
        <title>3.3. Explainable analysis</title>
        <p>We conducted an explainable analysis of the output of two
models using TIGERSCORE. Table 4 provides an example of
tayTYmhnoohaauudetth’cbnnsaoeucwneyomdshudblalotdeotorvabysken!eodnyIutotmadhwlueoianagitntrghhno.bt,Wtebwdtfehihohfecarieranteenuknagsytoetglohyirtnue.ehgaTaestthrtrueoewedyaifaseoyjtnufuhrttrseisotrtehdrstdeheggoaeorrmlaiaumzdd,geeaeh.ttrth.hat IIItcTwhhhbhhiealeeaalllvvalfiseeeeoctnvtucawiegudvsoesiesttdrhinyogk,atnneaastdenntatdhhddwreeetttihthhhgbheieieevgynssaehatndhsemwmosaateavnanyelgleoulrtvttsmohoehcuelhbtampaeeovcaroeotwcfilohavrtfsiihmtmtteuyaeaaadtsaatteywhnbhndoaotpiusssnrttohtbttoobieurtal.ihese.lc.attmakucvtdhotetaehbaasnacketlttsosaiiwvstcwiha-ttiasydolltt,eorbbmtehehusiasetcaryasctdhkllclaeefyasodnssr.ttdutohodde.emont.s</p>
        <p>I think that all students can learn from one another. I really enjoyed
Math shouldn’t be about right vs. wrong. Cathy’s approach to problem solving. She could see the value in making
Math can become fun when it is conceptual mistakes, but she also had a great student-friendly way of being able to
and you can adapt it to diferent situations. get students to think in a diferent way. She did not feel like she was asking
for a person or a topic, she just saw it in a diferent way.
the analysis results from Table 2. Here, we present specific
examples only. In the future, based on the scoring, we will
aggregate all AI-generated results and employ statistical
methods to evaluate the diferences between the two models.</p>
        <p>From the table, it is clear that the results from both
models might be lacking in terms of accuracy or understanding
context. The detailed analysis shows specific areas where
each model’s generated text is insuficient. For instance, in
Example 1, the text produced by GPT-2 small faces issues
related to comprehensiveness and accuracy. From a
comprehensiveness standpoint, the AI-generated text misinterprets
the initial post. It should have more efectively addressed
the themes of working hard and making mistakes that were
mentioned originally. Regarding accuracy, the evaluation
also recommends maintaining a focus on the main topic.</p>
        <p>In comparison, the results generated by Llama3-8B
performed better. In Example 1, there was only one issue related
to misunderstanding the context. In Example 2, according
to the TIGERSCORE, the results produced by Llama3-8B
had no issues. Further manual analysis would be valuable
in the future.</p>
        <p>From the explainable analysis results, we can conclude
that: (1) Using explainable metrics to evaluate large-scale
AI-generated texts is feasible; (2) Overall, the text quality
generated by Llama3-8B is superior to that of GPT-2 small
in the explainable analysis; (3) Although current LLMs ofer
many opportunities for online learning support, the quality
of this support still needs further improvement.</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>4. Discussion</title>
      <p>
        Large Language Models (LLMs) have significantly advanced
the field of Education Data Mining (EDM), presenting
innovative methodologies for the analysis of educational data
and the enhancement of learning experiences [
        <xref ref-type="bibr" rid="ref44">44</xref>
        ].
Particularly in online learning environments, AI-generated texts
derived from LLMs furnish not only a substantive level of
emotional and communal support but also contribute
critically to elevating student engagement and academic
outcomes [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. This research extends the application of LLMs
to support online discussion forums, addressing the gap
in utilizing explainable evaluations for AI-generated texts
within educational frameworks. The objectives of this study
are two aspects. The first is to ascertain the extent to which
deep learning-powered text generation can provide efective
and substantive textual support to learners within
expansive online communities, and the second is to implement an
explainable metric to assess these generated texts, thereby
facilitating a comparative analysis of cutting-edge models
against established educational benchmarks. In our
discussion, we reflect on the implications of our results concerning
the performance of two diferent language models, their
applicability in educational settings, and the broader impact
of large language models on online learning environments.
      </p>
      <p>
        Firstly, the analysis demonstrates distinct strengths
between the two models GPT-2 small and Llama3-8B. Texts
generated by Llama3-8B align well with the demands of
educational content that benefits from depth and
innovative thinking. This characteristic can enhance discussions
in online forums, where engaging and profound content
can stimulate deeper interaction among students [
        <xref ref-type="bibr" rid="ref45">45</xref>
        ].
Conversely, the simplicity and coherence in the responses from
GPT-2 small cater to scenarios where straightforward
communication is required, possibly aiding learners who benefit
from clear and concise explanations. Compared with
previous studies, with an average F-K Grade Level of 10.10 [
        <xref ref-type="bibr" rid="ref25">25</xref>
        ]
and 4.02 [
        <xref ref-type="bibr" rid="ref24">24</xref>
        ] of the GPT-2 small model, our results show
that the generated text has middle-level readability.
Furthermore, the significant diference in the F-K readability
levels between the texts generated by the GPT-2 small and
Llama3-8B, as indicated by the t-test results 3, suggests a
tailored application approach where each model’s output is
matched to specific educational needs or student groups.
      </p>
      <p>
        Secondly, the introduction of explainable metrics, such
as TIGERSCORE [
        <xref ref-type="bibr" rid="ref32">32</xref>
        ], is pivotal in assessing and
understanding the utility of large-scale responses generated by
LLMs in educational settings. Previous studies evaluated
the AI-generated text using only quantitative methods like
F-K Grade Level [
        <xref ref-type="bibr" rid="ref24">24</xref>
        ] and word perplexity [
        <xref ref-type="bibr" rid="ref25">25</xref>
        ]; for human
evaluation, they only incorporate a small amount of text,
causing manual scores to be time-consuming. The
increasing focus on model interpretability, which has led to a surge
in research dedicated to explainable metrics [
        <xref ref-type="bibr" rid="ref46 ref47">46, 47</xref>
        ]. This
study adds explainable analysis to AI-generated content
by LLMs and shows great potential for large-scale content
explainable evaluation. These metrics help refine the AI’s
output, ensuring that the generated content is engaging and
pedagogically valuable. For instance, the ability to dissect
and explain model decisions and output can facilitate the
integration of AI tools into learning environments where
transparency and trust are paramount. Educators can
leverage these insights to better scafold learning, providing
interventions that are responsive to the unique dynamics of
student interactions in online forums.
      </p>
      <p>
        Lastly, despite the potential shown by these technologies,
there are significant limitations in the texts generated by
current AI models, including issues related to accuracy,
content misunderstanding, and the production of inappropriate
content. These challenges are particularly critical in
educational contexts where the accuracy and appropriateness of
content are paramount. In this study, we identified diferent
error levels made by GPT-2 small and Llama3-8B, with
results indicating that Llama3-8B generated better outcomes
with fewer errors. This may be due to two reasons: firstly,
Llama3-8B is a more complex model than GPT-2 small,
enabling it to learn better and perform more efectively with
MOOC posts data [
        <xref ref-type="bibr" rid="ref22">22</xref>
        ]; secondly, the TIGERSCORE itself
is based on the Llama series model, which may cause the
model to favor evaluations of models similar to itself. Future
research should focus on this point and use a diverse set of
explainable and automatic evaluation metrics for analysis.
Overall, previous research has shown that the performance
of GPT-2 small surpasses that of other traditional neural
network models like RNNs. Building on this, we employed
the latest LLM models to generate automated responses to
students’ online posts, achieving better results than with
the GPT-2 model.
      </p>
      <p>In conclusion, while the advanced capabilities of
models like Llama3-8B and GPT-2 ofer exciting opportunities
for enhancing interactive learning, their integration into
educational frameworks must be handled with a keen
awareness of their limitations and a strong emphasis on ethical
implications and educational validity. This application of
explainable metrics for AI-generated content will maximize
their potential benefits while safeguarding the learning
environment against possible negative impacts of AI technology.</p>
    </sec>
    <sec id="sec-5">
      <title>5. Limitation and Future Work</title>
      <p>Considering the limitations and future objectives identified
in the project, we face several challenges and directions for
future research. First, the dataset used for fine-tuning the
model is much smaller than typical training sets, which may
limit the model’s ability to generalize efectively. Secondly,
we rely on artificially generated responses to supplement
missing posts, which could introduce biases or inaccuracies
not present in the original data. Additionally, we used only
one explainable metric for analysis; the TIGERSCORE metric
is based on Llama models, which may lead to a better score
for Llama3 than GPT-2. As not all AI-generated texts are
thoroughly reviewed by humans, errors and biases may
remain undetected and uncorrected, potentially afecting the
quality of the output. In the future, we plan to incorporate
more manual analysis as indicated in Table 5.</p>
    </sec>
    <sec id="sec-6">
      <title>6. Acknowledgments</title>
      <p>Acknowledgements This work is supported by the National
Science Foundation (NSF) of the United States under grant
numbers 1503196 and 2105695. Any opinions, findings, and
conclusions or recommendations expressed in this paper,
however, are those of the authors and do not necessarily
reflect the views of the NSF.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>R. L.</given-names>
            <surname>Moore</surname>
          </string-name>
          ,
          <string-name>
            <surname>K. M. Oliver</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          <string-name>
            <surname>Wang</surname>
          </string-name>
          ,
          <article-title>Setting the pace: Examining cognitive processing in mooc discussion forums with automatic text analysis</article-title>
          ,
          <source>Interactive Learning Environments</source>
          <volume>27</volume>
          (
          <year>2019</year>
          )
          <fpage>655</fpage>
          -
          <lpage>669</lpage>
          . doi:
          <volume>10</volume>
          .1080/10494820.
          <year>2019</year>
          .
          <volume>1610453</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>C.</given-names>
            <surname>Coman</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L. G.</given-names>
            <surname>Ţîru</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Meseşan-Schmitz</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Stanciu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. C.</given-names>
            <surname>Bularca</surname>
          </string-name>
          ,
          <article-title>Online teaching and learning in higher education during the coronavirus pandemic: Students' perspective</article-title>
          ,
          <source>Sustainability</source>
          <volume>12</volume>
          (
          <year>2020</year>
          )
          <fpage>10367</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>L.</given-names>
            <surname>Song</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S. W.</given-names>
            <surname>McNary</surname>
          </string-name>
          ,
          <article-title>Understanding students' online interaction: Analysis of discussion board postings</article-title>
          .,
          <source>Journal of Interactive Online Learning</source>
          <volume>10</volume>
          (
          <year>2011</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>P. C.</given-names>
            <surname>Abrami</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R. M.</given-names>
            <surname>Bernard</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E. M.</given-names>
            <surname>Bures</surname>
          </string-name>
          , E. Borokhovski,
          <string-name>
            <given-names>R. M.</given-names>
            <surname>Tamim</surname>
          </string-name>
          ,
          <article-title>Interaction in distance education and online learning: Using evidence and theory to improve practice</article-title>
          ,
          <source>Journal of computing in higher education 23</source>
          (
          <year>2011</year>
          )
          <fpage>82</fpage>
          -
          <lpage>103</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>R. M.</given-names>
            <surname>Wallace</surname>
          </string-name>
          ,
          <article-title>Online learning in higher education: A review of research on interactions among teachers and students</article-title>
          , Education,
          <source>Communication &amp; Information</source>
          <volume>3</volume>
          (
          <year>2003</year>
          )
          <fpage>241</fpage>
          -
          <lpage>280</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>J. C.</given-names>
            <surname>Richardson</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Maeda</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Lv</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Caskurlu</surname>
          </string-name>
          ,
          <article-title>Social presence in relation to students' satisfaction and learning in the online environment: A meta-analysis</article-title>
          ,
          <source>Computers in Human Behavior</source>
          <volume>71</volume>
          (
          <year>2017</year>
          )
          <fpage>402</fpage>
          -
          <lpage>417</lpage>
          . doi:
          <volume>10</volume>
          .1016/j.chb.
          <year>2017</year>
          .
          <volume>02</volume>
          .001.
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>T. K.</given-names>
            <surname>Chiu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T. K.</given-names>
            <surname>Hew</surname>
          </string-name>
          ,
          <article-title>Factors influencing peer learning and performance in mooc asynchronous online discussion forum</article-title>
          ,
          <source>Australasian Journal of Educational Technology</source>
          <volume>34</volume>
          (
          <year>2018</year>
          )
          <fpage>16</fpage>
          -
          <lpage>28</lpage>
          . doi:
          <volume>10</volume>
          .14742/ajet.3240.
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>H.</given-names>
            <surname>Tang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W.</given-names>
            <surname>Xing</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Pei</surname>
          </string-name>
          ,
          <article-title>Exploring the temporal dimension of forum participation in moocs</article-title>
          ,
          <source>Distance Education</source>
          <volume>39</volume>
          (
          <year>2018</year>
          )
          <fpage>353</fpage>
          -
          <lpage>372</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>M.</given-names>
            <surname>Cleveland-Innes</surname>
          </string-name>
          , P. Campbell,
          <article-title>Emotional presence, learning, and the online learning environment</article-title>
          ,
          <source>The International Review of Research in Open and Distributed Learning</source>
          <volume>13</volume>
          (
          <year>2012</year>
          )
          <fpage>269</fpage>
          -
          <lpage>292</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>M.</given-names>
            <surname>Fei</surname>
          </string-name>
          , D.-Y. Yeung,
          <article-title>Temporal models for predicting student dropout in massive open online courses</article-title>
          ,
          <source>in: 2015 IEEE international conference on data mining workshop (ICDMW)</source>
          , IEEE,
          <year>2015</year>
          , pp.
          <fpage>256</fpage>
          -
          <lpage>263</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>W.</given-names>
            <surname>Xing</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Du</surname>
          </string-name>
          ,
          <article-title>Dropout prediction in moocs: Using deep learning for personalized intervention</article-title>
          ,
          <source>Journal of Educational Computing Research</source>
          <volume>57</volume>
          (
          <year>2019</year>
          )
          <fpage>547</fpage>
          -
          <lpage>570</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>E.</given-names>
            <surname>Er</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Gómez-Sánchez</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Dimitriadis</surname>
          </string-name>
          ,
          <string-name>
            <surname>M. L. BoteLorenzo</surname>
            ,
            <given-names>J. I.</given-names>
          </string-name>
          <string-name>
            <surname>Asensio-Pérez</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          <string-name>
            <surname>Álvarez-Álvarez</surname>
          </string-name>
          ,
          <article-title>Aligning learning design and learning analytics through instructor involvement: A mooc case study</article-title>
          ,
          <source>Interactive Learning Environments</source>
          <volume>27</volume>
          (
          <year>2019</year>
          )
          <fpage>685</fpage>
          -
          <lpage>698</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <given-names>W.</given-names>
            <surname>Holmes</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Q.</given-names>
            <surname>Nguyen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Zhang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Mavrikis</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Rienties</surname>
          </string-name>
          ,
          <article-title>Learning analytics for learning design in online distance learning</article-title>
          ,
          <source>Distance Education</source>
          <volume>40</volume>
          (
          <year>2019</year>
          )
          <fpage>309</fpage>
          -
          <lpage>329</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <given-names>R.</given-names>
            <surname>Martinez-Maldonado</surname>
          </string-name>
          ,
          <article-title>A handheld classroom dashboard: Teachers' perspectives on the use of real-time collaborative learning analytics</article-title>
          ,
          <source>International Journal of Computer-Supported Collaborative Learning</source>
          <volume>14</volume>
          (
          <year>2019</year>
          )
          <fpage>383</fpage>
          -
          <lpage>411</lpage>
          . doi:
          <volume>10</volume>
          .1007/s11412- 019- 09308- z.
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <given-names>Y.</given-names>
            <surname>Zhang</surname>
          </string-name>
          , S. Sun,
          <string-name>
            <given-names>M.</given-names>
            <surname>Galley</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y. C.</given-names>
            <surname>Chen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Brockett</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Gao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Gao</surname>
          </string-name>
          , J. Liu,
          <string-name>
            <given-names>W. B.</given-names>
            <surname>Dolan</surname>
          </string-name>
          , Dialogpt:
          <article-title>Largescale generative pre-training for conversational response generation</article-title>
          ,
          <source>in: Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics: System Demonstrations</source>
          ,
          <year>2020</year>
          , pp.
          <fpage>270</fpage>
          -
          <lpage>278</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [16]
          <string-name>
            <given-names>F.</given-names>
            <surname>Yilmaz</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Yılmaz</surname>
          </string-name>
          ,
          <article-title>Learning analytics intervention improves students' engagement in online learning</article-title>
          ,
          <source>Technology, Knowledge and Learning</source>
          <volume>27</volume>
          (
          <year>2021</year>
          )
          <fpage>449</fpage>
          -
          <lpage>460</lpage>
          . doi:
          <volume>10</volume>
          .1007/s10758- 021- 09547- w.
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [17]
          <string-name>
            <given-names>L.</given-names>
            <surname>Zheng</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Zhong</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Niu</surname>
          </string-name>
          ,
          <article-title>Efects of personalised feedback approach on knowledge building, emotions, co-regulated behavioural patterns and cognitive load in online collaborative learning</article-title>
          ,
          <source>Assessment &amp; Evaluation in Higher Education</source>
          <volume>47</volume>
          (
          <year>2021</year>
          )
          <fpage>109</fpage>
          -
          <lpage>125</lpage>
          . doi:
          <volume>10</volume>
          . 1080/02602938.
          <year>2021</year>
          .
          <volume>1883549</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref18">
        <mixed-citation>
          [18]
          <string-name>
            <given-names>L.-A.</given-names>
            <surname>Lim</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Gentili</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Pardo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Kovanović</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Whitelock-Wainwright</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Gašević</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Dawson</surname>
          </string-name>
          ,
          <article-title>What changes, and for whom? a study of the impact of learning analytics-based process feedback in a large course, Learning and Instruction (</article-title>
          <year>2019</year>
          )
          <article-title>101202</article-title>
          . doi:
          <volume>10</volume>
          .1016/J.LEARNINSTRUC.
          <year>2019</year>
          .
          <volume>04</volume>
          .003.
        </mixed-citation>
      </ref>
      <ref id="ref19">
        <mixed-citation>
          [19]
          <string-name>
            <given-names>D.</given-names>
            <surname>Parmar</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Ali</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Dewan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Wen</surname>
          </string-name>
          ,
          <article-title>Automatic analysis of online course discussion forum: A short review</article-title>
          ,
          <source>2023 IEEE Canadian Conference on Electrical and Computer</source>
          Engineering (CCECE) (
          <year>2023</year>
          )
          <fpage>210</fpage>
          -
          <lpage>215</lpage>
          . doi:
          <volume>10</volume>
          .1109/CCECE58730.
          <year>2023</year>
          .
          <volume>10289065</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref20">
        <mixed-citation>
          [20]
          <string-name>
            <given-names>S.</given-names>
            <surname>Goggins</surname>
          </string-name>
          , W. Xing,
          <article-title>Building models explaining student participation behavior in asynchronous online discussion</article-title>
          ,
          <source>Computers &amp; Education</source>
          <volume>94</volume>
          (
          <year>2016</year>
          )
          <fpage>241</fpage>
          -
          <lpage>251</lpage>
          . doi:
          <volume>10</volume>
          .1016/j.compedu.
          <year>2015</year>
          .
          <volume>11</volume>
          .002.
        </mixed-citation>
      </ref>
      <ref id="ref21">
        <mixed-citation>
          [21]
          <string-name>
            <surname>OpenAI</surname>
          </string-name>
          , Gpt-4
          <source>technical report</source>
          ,
          <year>2023</year>
          . URL: https://doi. org/10.48550/arXiv.2303.08774. doi:
          <volume>10</volume>
          .48550/arXiv. 2303.08774.
        </mixed-citation>
      </ref>
      <ref id="ref22">
        <mixed-citation>
          [22]
          <string-name>
            <given-names>H.</given-names>
            <surname>Touvron</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Lavril</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Izacard</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Martinet</surname>
          </string-name>
          , M.
          <article-title>-</article-title>
          <string-name>
            <surname>A. Lachaux</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          <string-name>
            <surname>Lacroix</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          <string-name>
            <surname>Roziere</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          <string-name>
            <surname>Goyal</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          <string-name>
            <surname>Hambro</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          <string-name>
            <surname>Azhar</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Rodriguez</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Joulin</surname>
          </string-name>
          , E. Grave, G. Lample,
          <article-title>Llama: Open and eficient foundation language models</article-title>
          ,
          <source>arXiv preprint arXiv:2302.13971</source>
          (
          <year>2023</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref23">
        <mixed-citation>
          [23]
          <string-name>
            <given-names>D.</given-names>
            <surname>Parmar</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. A. A.</given-names>
            <surname>Dewan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Wen</surname>
          </string-name>
          ,
          <article-title>Automatic analysis of online course discussion forum: A short review</article-title>
          ,
          <source>in: 2023 IEEE Canadian Conference on Electrical and Computer Engineering (CCECE)</source>
          ,
          <year>2023</year>
          , pp.
          <fpage>210</fpage>
          -
          <lpage>215</lpage>
          . doi:
          <volume>10</volume>
          .1109/CCECE58730.
          <year>2023</year>
          .
          <volume>10289065</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref24">
        <mixed-citation>
          [24]
          <string-name>
            <given-names>H.</given-names>
            <surname>Du</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W.</given-names>
            <surname>Xing</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Pei</surname>
          </string-name>
          ,
          <article-title>Automatic text generation using deep learning: providing large-scale support for online learning communities</article-title>
          ,
          <source>Interactive Learning Environments</source>
          <volume>31</volume>
          (
          <year>2023</year>
          )
          <fpage>5021</fpage>
          -
          <lpage>5036</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref25">
        <mixed-citation>
          [25]
          <string-name>
            <given-names>C.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W.</given-names>
            <surname>Xing</surname>
          </string-name>
          ,
          <article-title>Natural language generation using deep learning to support mooc learners</article-title>
          ,
          <source>International Journal of Artificial Intelligence in Education</source>
          <volume>31</volume>
          (
          <year>2021</year>
          )
          <fpage>186</fpage>
          -
          <lpage>214</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref26">
        <mixed-citation>
          [26]
          <string-name>
            <given-names>G. K.</given-names>
            <surname>Wong</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y. K.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Lai</surname>
          </string-name>
          ,
          <article-title>Visualizing the learning patterns of topic-based social interaction in online discussion forums: an exploratory study</article-title>
          ,
          <source>Educational Technology Research and Development</source>
          <volume>69</volume>
          (
          <year>2021</year>
          )
          <fpage>2813</fpage>
          -
          <lpage>2843</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref27">
        <mixed-citation>
          [27]
          <string-name>
            <given-names>S.</given-names>
            <surname>Gottipati</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Shankararaman</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Ramesh</surname>
          </string-name>
          ,
          <article-title>Topicsummary: A tool for analyzing class discussion forums using topic based summarizations</article-title>
          ,
          <source>in: 2019 IEEE Frontiers in Education Conference (FIE)</source>
          , IEEE,
          <year>2019</year>
          , pp.
          <fpage>1</fpage>
          -
          <lpage>9</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref28">
        <mixed-citation>
          [28]
          <string-name>
            <given-names>O.</given-names>
            <surname>Almatrafi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Johri</surname>
          </string-name>
          ,
          <article-title>Improving moocs using information from discussion forums: An opinion summarization and suggestion mining approach</article-title>
          ,
          <source>IEEE Access 10</source>
          (
          <year>2022</year>
          )
          <fpage>15565</fpage>
          -
          <lpage>15573</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref29">
        <mixed-citation>
          [29]
          <string-name>
            <given-names>L.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Johnson</surname>
          </string-name>
          , W. Aarhus,
          <string-name>
            <given-names>D.</given-names>
            <surname>Shah</surname>
          </string-name>
          ,
          <article-title>Key factors in mooc pedagogy based on nlp sentiment analysis of learner reviews: What makes a hit</article-title>
          ,
          <source>Computers &amp; Education</source>
          <volume>176</volume>
          (
          <year>2022</year>
          )
          <fpage>104354</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref30">
        <mixed-citation>
          [30]
          <string-name>
            <given-names>C.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W.</given-names>
            <surname>Xing</surname>
          </string-name>
          , W. Leite,
          <article-title>Building socially responsible conversational agents using big data to support online learning: A case with algebra nation</article-title>
          ,
          <source>British Journal of Educational Technology</source>
          <volume>53</volume>
          (
          <year>2022</year>
          )
          <fpage>776</fpage>
          -
          <lpage>803</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref31">
        <mixed-citation>
          [31]
          <string-name>
            <given-names>W.</given-names>
            <surname>Xu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Pan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Song</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Freitag</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W. Y.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <article-title>Instructscore: Towards explainable text generation evaluation with automatic feedback</article-title>
          ,
          <source>arXiv preprint arXiv:2305.14282</source>
          (
          <year>2023</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref32">
        <mixed-citation>
          [32]
          <string-name>
            <given-names>D.</given-names>
            <surname>Jiang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Zhang</surname>
          </string-name>
          , W. Huang,
          <string-name>
            <given-names>B. Y.</given-names>
            <surname>Lin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W.</given-names>
            <surname>Chen</surname>
          </string-name>
          , Tigerscore:
          <article-title>Towards building explainable metric for all text generation tasks</article-title>
          ,
          <source>arXiv preprint arXiv:2310.00752</source>
          (
          <year>2023</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref33">
        <mixed-citation>
          [33]
          <string-name>
            <surname>I. Van De Poel</surname>
          </string-name>
          ,
          <article-title>Design for value change</article-title>
          ,
          <source>Ethics and Information Technology</source>
          <volume>23</volume>
          (
          <year>2021</year>
          )
          <fpage>27</fpage>
          -
          <lpage>31</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref34">
        <mixed-citation>
          [34]
          <string-name>
            <given-names>S.</given-names>
            <surname>Cruz</surname>
          </string-name>
          ,
          <article-title>Cognitive and afective outcomes among targets and non-targets of racist hate speech in the college setting, Doctoral dissertation</article-title>
          , Arizona State University, Arizona,
          <year>2021</year>
          . Publication No.
          <volume>2564152566</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref35">
        <mixed-citation>
          [35]
          <string-name>
            <given-names>A. C.</given-names>
            <surname>Curry</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Rieser</surname>
          </string-name>
          , #
          <article-title>metoo alexa: How conversational systems respond to sexual harassment</article-title>
          ,
          <source>in: Proceedings of the Second ACL Workshop on Ethics in Natural Language Processing</source>
          ,
          <year>2018</year>
          , pp.
          <fpage>7</fpage>
          -
          <lpage>14</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref36">
        <mixed-citation>
          [36]
          <article-title>AI@Meta, Llama 3 model card (</article-title>
          <year>2024</year>
          ). URL: https://github.com/meta-llama/llama3/blob/main/ MODEL_CARD.md.
        </mixed-citation>
      </ref>
      <ref id="ref37">
        <mixed-citation>
          [37]
          <string-name>
            <given-names>A.</given-names>
            <surname>Vaswani</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Shazeer</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Parmar</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Uszkoreit</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Jones</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A. N.</given-names>
            <surname>Gomez</surname>
          </string-name>
          , Ł. Kaiser,
          <string-name>
            <surname>I. Polosukhin</surname>
          </string-name>
          ,
          <article-title>Attention is all you need</article-title>
          ,
          <source>Advances in neural information processing systems</source>
          <volume>30</volume>
          (
          <year>2017</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref38">
        <mixed-citation>
          [38]
          <string-name>
            <given-names>A.</given-names>
            <surname>Radford</surname>
          </string-name>
          ,
          <string-name>
            <surname>J. Wu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Child</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Luan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Amodei</surname>
          </string-name>
          ,
          <string-name>
            <given-names>I.</given-names>
            <surname>Sutskever</surname>
          </string-name>
          , et al.,
          <article-title>Language models are unsupervised multitask learners</article-title>
          ,
          <source>OpenAI blog 1</source>
          (
          <year>2019</year>
          )
          <article-title>9</article-title>
          .
        </mixed-citation>
      </ref>
      <ref id="ref39">
        <mixed-citation>
          [39]
          <string-name>
            <given-names>S. J.</given-names>
            <surname>Pan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Q.</given-names>
            <surname>Yang</surname>
          </string-name>
          ,
          <article-title>A survey on transfer learning</article-title>
          ,
          <source>IEEE Transactions on knowledge and data engineering 22</source>
          (
          <year>2009</year>
          )
          <fpage>1345</fpage>
          -
          <lpage>1359</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref40">
        <mixed-citation>
          [40]
          <string-name>
            <given-names>R. S.</given-names>
            <surname>Sutton</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A. G.</given-names>
            <surname>Barto</surname>
          </string-name>
          ,
          <article-title>Reinforcement learning: An introduction</article-title>
          , MIT press,
          <year>2018</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref41">
        <mixed-citation>
          [41]
          <string-name>
            <given-names>A.</given-names>
            <surname>Belz</surname>
          </string-name>
          , E. Reiter,
          <article-title>Comparing automatic and human evaluation of nlg systems</article-title>
          ,
          <source>in: Proceedings of the 11th Conference of the European Chapter of the Association for Computational Linguistics</source>
          ,
          <year>2006</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref42">
        <mixed-citation>
          [42]
          <string-name>
            <surname>J. M. L. Williamson</surname>
            ,
            <given-names>A. G.</given-names>
          </string-name>
          <string-name>
            <surname>Martin</surname>
          </string-name>
          ,
          <article-title>Analysis of patient information leaflets provided by a district general hospital by the flesch and flesch-kincaid method</article-title>
          ,
          <source>International Journal of Clinical Practice</source>
          <volume>64</volume>
          (
          <year>2010</year>
          )
          <fpage>1824</fpage>
          -
          <lpage>1831</lpage>
          . URL: https://doi.org/10.1111/j.1742-
          <fpage>1241</fpage>
          .
          <year>2010</year>
          .
          <volume>02408</volume>
          .x. doi:
          <volume>10</volume>
          .1111/j.1742-
          <fpage>1241</fpage>
          .
          <year>2010</year>
          .
          <volume>02408</volume>
          .x.
        </mixed-citation>
      </ref>
      <ref id="ref43">
        <mixed-citation>
          [43]
          <string-name>
            <given-names>S.</given-names>
            <surname>Sabharwal</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Badarudeen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S. Unes</given-names>
            <surname>Kunju</surname>
          </string-name>
          ,
          <article-title>Readability of online patient education materials from the aaos web site</article-title>
          ,
          <source>Clinical Orthopaedics and related research 466</source>
          (
          <year>2008</year>
          )
          <fpage>1245</fpage>
          -
          <lpage>1250</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref44">
        <mixed-citation>
          [44]
          <string-name>
            <given-names>P.</given-names>
            <surname>Denny</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Gulwani</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N. T.</given-names>
            <surname>Hefernan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Käser</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Moore</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A. N.</given-names>
            <surname>Raferty</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Singla</surname>
          </string-name>
          ,
          <article-title>Generative ai for education (gaied): Advances, opportunities, and challenges</article-title>
          ,
          <source>CoRR abs/2402</source>
          .01580 (
          <year>2024</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref45">
        <mixed-citation>
          [45]
          <string-name>
            <given-names>E. M.</given-names>
            <surname>Onyema</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E. C.</given-names>
            <surname>Deborah</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A. O.</given-names>
            <surname>Alsayed</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Q.</given-names>
            <surname>Noorulhasan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Sanober</surname>
          </string-name>
          ,
          <article-title>Online discussion forum as a tool for interactive learning and communication</article-title>
          ,
          <source>International Journal of Recent Technology and Engineering</source>
          <volume>8</volume>
          (
          <year>2019</year>
          )
          <fpage>4852</fpage>
          -
          <lpage>4859</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref46">
        <mixed-citation>
          [46]
          <string-name>
            <given-names>M.</given-names>
            <surname>Zhong</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Liu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Yin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Mao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Jiao</surname>
          </string-name>
          , P. Liu,
          <string-name>
            <given-names>C.</given-names>
            <surname>Zhu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Ji</surname>
          </string-name>
          , J. Han,
          <article-title>Towards a unified multi-dimensional evaluator for text generation</article-title>
          ,
          <year>2022</year>
          . 2022b.
        </mixed-citation>
      </ref>
      <ref id="ref47">
        <mixed-citation>
          [47]
          <string-name>
            <given-names>J.</given-names>
            <surname>Fu</surname>
          </string-name>
          , S.
          <article-title>-</article-title>
          <string-name>
            <surname>K. Ng</surname>
            ,
            <given-names>Z.</given-names>
          </string-name>
          <string-name>
            <surname>Jiang</surname>
          </string-name>
          , P. Liu, Gptscore:
          <article-title>Evaluate as you desire</article-title>
          ,
          <source>arXiv preprint arXiv:2302.04166</source>
          (
          <year>2023</year>
          ).
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>