<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>with Large Language Models</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Kole Norberg</string-name>
          <email>knorberg1@carnegielearning.com</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Husni Almoubayyed</string-name>
          <email>halmoubayyed@carnegielearning.com</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Stephen E. Fancsali</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Logan De Ley</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Kyle Weldon</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>April Murphy</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Steve Ritter</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Carnegie Learning, Inc.</institution>
          ,
          <addr-line>501 Grant St, Pittsburgh, PA 15219</addr-line>
          ,
          <country country="US">USA</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2023</year>
      </pub-date>
      <abstract>
        <p>Large Language Models have recently achieved high performance on many writing tasks. In a recent study, math word problems in Carnegie Learning's MATHia adaptive learning software were rewritten by human authors to improve their clarity and specificity. The randomized experiment found that emerging readers who received the rewritten word problems spent less time completing the problems and also achieved higher mastery compared to emerging readers who received the original content. We used GPT-4 to rewrite the same set of math word problems, prompting it to follow the same guidelines that the human authors followed. We lay out our prompt engineering process, comparing several prompting strategies: zero-shot, few-shot, and chain-of-thought prompting. Additionally, we overview how we leveraged GPT's ability to write python code in order to encode mathematical components of word problems. We report text analysis of the original, human-rewritten, and GPT-rewritten problems. GPT rewrites had the most optimal readability, lexical diversity, and cohesion scores but used more low frequency words. We present our plan to test the GPT outputs in upcoming randomized field trials in</p>
      </abstract>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>-</title>
      <p>MATHia.</p>
    </sec>
    <sec id="sec-2">
      <title>1. Introduction</title>
      <p>
        Recent studies have found that building more comprehensive learner models results in better
student learning outcomes and experience. In particular, research has pointed towards strong
connections between mathematics and reading comprehension (e.g., [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]). Almoubayyed et al.
[
        <xref ref-type="bibr" rid="ref2">2</xref>
        ] built a highly-accurate machine learning model to predict students’ reading ability based on
their interactions with an introductory activity in a mathematics adaptive learning software,
MATHia.1 This model enables researchers to target emerging readers as a sub-population when
building out reading supports in MATHia (600,000+ user base) without the need for intrusive
reading tests or collecting state exam scores.
      </p>
      <p>Supporting the connection between reading skill and math outcomes, a recent randomized
ifeld trial with 12,000+ students demonstrated that improving the readability of math word
problems in Carnegie Learning’s MATHia improved student outcomes. This was particularly
true among emerging readers who were identified using the the predictive model developed in
nEvelop-O
CEUR
Workshop
Proceedings
htp:/ceur-ws.org
ISN1613-073</p>
      <p>
        CEUR Workshop Proceedings (CEUR-WS.org)
1The predictive model generalized well to diferent states and demographics in [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ].
[
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]. Specifically, predicted emerging readers who received the rewritten word problems mastered
more skills, made fewer errors, and required fewer problem opportunities to reach mastery,
saving them a significant amount of time (over a third) compared to emerging readers who
received the original problems. Despite this success, it is unclear whether manually rewriting
all word problems in MATHia is a scalable solution or indeed if the results will replicate across
hundreds of lessons in Carnegie Learning’s content portfolio.
      </p>
      <p>
        Recent developments in Large Language Models (LLMs) have allowed many language tasks
to scale up in less time and with lower cost. Given the major improvements shown in [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ], we
looked towards LLMs as a potential method to allow us to eficiently scale up the rewriting
efort. This work introduces our process of prompt engineering ChatGPT [ 5] to rewrite word
problems according to the same style guidelines that were used by human authors in [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ]. We
also include a comprehensive text analysis, using the Automatic Readability Tool for English
(ARTE, [6]), of the outputs of GPT and compare it to that of the original problems and the
human-rewrites from [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ]. Finally, we lay out our plan to carry out a randomized field trial to
test student outcomes from the GPT-4 rewrites according to a set of defined metrics.
      </p>
    </sec>
    <sec id="sec-3">
      <title>2. MATHia &amp; UpGrade</title>
      <p>
        MATHia (formerly Cognitive Tutor, [7]) is an intelligent tutoring system (ITS) developed by
Carnegie Learning and currently used by 600,000+ students across the United States. MATHia is
typically used as part of a blended curriculum in the classroom. Content in MATHia is delivered
in lessons known as ‘workspaces,’ with workspaces being classified as either ‘Concept Builders’
or ‘Mastery’ workspaces. Students work through a set of pre-defined steps in Concept Builders
to learn new concepts in mathematics by interacting with adaptive multimedia tools. In Mastery
workspaces, problems are adaptively selected, according to a Bayesian Knowledge Tracing
(BKT; [8]) implementation, providing practice opportunities for students to master a set of skills.
Students are either ‘graduated’ or ‘promoted’ from Mastery workspaces. A student is graduated
when they successfully master all skills in a workspace. A student is promoted when they fail
to master at least one skill in a workspace after completing a pre-defined maximum number
of problems. In this work, we focus on two Mastery workspaces that were also the focus of
[
        <xref ref-type="bibr" rid="ref4">4</xref>
        ]. Both workspaces involve analyzing models of two-step linear relationships, but one deals
with integer numbers, referred to as ‘Integers’ and one deals with rational numbers, referred to
as ‘Rationals.’ These two workspaces were originally chosen due to the correlation between
students’ performance in them and students’ end-of-year English Language Arts (ELA) state
test scores being high, particularly compared to the correlation between students performance
in them and their end-of-year math state test scores, for a sample of students, as described in
[
        <xref ref-type="bibr" rid="ref2">2</xref>
        ].
      </p>
      <p>
        UpGrade [9] has been recently developed as a free and open-source platform for running
large-scale, randomized field tests. Upgrade integrates with ITSs, such as MATHia, to enable
random assignment to experimental conditions. In [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ], UpGrade was used to randomly assign
12,000+ students to the control or human-rewrite conditions over a period of around 6 weeks.
      </p>
    </sec>
    <sec id="sec-4">
      <title>3. Prompt Engineering</title>
      <p>Our prompts were written to enable GPT-42 to take any math word problem and revise it
according to a style guide developed by literacy experts at Carnegie Learning. In this section,
we provide an overview of our steps to engineering a successful prompt. We also highlight
areas of the revision process that were particularly challenging for GPT to generate.</p>
      <p>We used a mix of zero-shot [10, 11], few-shot [12], and chain-of-thought learning [13, 14].
Table 1 illustrates revisions produced by GPT-4, and for comparison GPT-3, for each of these
prompting styles. We also provide original and final versions as they appear in MATHia in at
https://osf.io/xwz3h/.</p>
      <sec id="sec-4-1">
        <title>3.1. Zero-shot learning</title>
        <p>Zero-shot learning prompts provide an LLM with a set of instructions to follow but critically do
not include exemplars. Such prompts produced revisions that were better than the originals
but which did not fully adhere to the style guide. With our texts, zero-shot learning worked
well for broad structural changes such as adding a topic sentence and contextualizing negative
numbers. However, it was unable to reliably address issues related to simplifying vocabulary
and sentence structure (e.g., removing prepositional clauses and passive voice). Although words
like subterranean were successfully replaced by GPT-4, words like eclair were not. As will
remain true in our overview of few-shot prompting, it was dificult for GPT-4 to make revisions
that changed the meaning of the word. Therefore, in zero-shot learning, if a simpler, exact
synonym could not be found, the word would not be replaced.</p>
      </sec>
      <sec id="sec-4-2">
        <title>3.2. Few-shot learning</title>
        <p>In order to improve GPT’s precision and reliability, we added examples of changes that might
be made for each step. We used a mix of 1-shot (one example) and 2-shot (two examples)
learning based on the complexity of the specific style guideline and observed performance.
Further, we asked GPT to first explain rules before carrying them out. For example, we asked
it to define passive voice and then label sentences accordingly. Critically, we asked GPT to
perform its revisions iteratively, so it would revise for passive voice in one step and then word
choice in another. This successfully improved GPT’s reliability in regards to using active voice.
However, GPT continued to struggle with vocabulary changes. It was now able to identify
dificult vocabulary, but if a simpler synonym was not available, it resorted to a broader more
general term. This sometimes made the vocabulary of the problem confusing (e.g., replacing
eclair with pastry when donuts, another item in the problem, are also a type of pastry). If pushed
to be specific, it would describe the word, replacing eclair with long pastry to diferentiate it
from donut. Further, under few-shot learning, it began introducing prepositional clauses into
sentences rather than removing them (e.g., For the high school trip, ... and At a bake sale, ....
2All revisions were generated by GPT-4 using chat.openai.com as we did not have access to the GPT-4 API at the
time of production. We discuss how access to the API has modified our approach in the Conclusions and Future
Work section.</p>
      </sec>
      <sec id="sec-4-3">
        <title>3.3. Chain-of-thought learning</title>
        <p>In our final iteration, we used chain-of-thought learning. Prompts of this type provide reasoning
with their examples. The beginning of our chain-of-thought prompts contained a mock word
problem, its revision, and rationale for how the revision obeyed the guidelines.</p>
        <p>This final iteration of our prompt</p>
        <p>3 also provided more scafolding for identifying and replacing
words. Because GPT-4 struggled with changing the meaning of the original word within a text,
we asked it to just remove the word, replace it with a blank, and then replace the blank with the
most likely word. Following this approach, it was able to break away from using synonyms
and select words which suited the context and the age of the user (e.g., it replaced eclair with
cupcake). This version of the prompt was successful on the first try at revising 26 out of 30
problems. It was successful for 3 out of 4 of the remaining problems on its second try. Its
struggle with the final problem was rooted in ambiguous pronoun use in the original text. We
revised the original by replacing the ambiguous pronoun with its referent, and the next GPT-4
revision was successful.</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>4. Encoding Word Problem Components</title>
      <p>In order for a problem to work within MATHia, the mathematical components need to be
encoded. After rewriting the problems, we also prompted GPT to complete the encodings. The
word problems that were revised are from two workspaces which have a similar structure: they
provide a student with a scenario that involves an equation of the form  =  + 
, and then
3Prompts, including the final iteration referred to here, are available at https://osf.io/xwz3h/.
At times, we asked GPT to provide multiple rewrites and select the best version. Sometimes the additional
revisions were better. More often, the additional revisions were worse.</p>
      <p>One way to determine the readability of a text is to evaluate the frequency of its vocabulary against a
corpus. Despite GPT’s claim to the contrary, we saw no evidence that GPT was able to reference or use a
corpus, nor was it able to evaluate the grade level of the text.</p>
      <p>When a prompt failed, we prompted GPT to explain how it understood the prompt so we could make
modifications. Most of the time, this did not yield any insights. Nevertheless, there were times when GPT
was able to rewrite the rule in a way that it would reliably follow it during future revisions.</p>
      <p>It was often necessary to tell GPT not to do something, most notably, not to use pronouns. GPT-4 does
better in general when given directions on what to do rather than what not to do (e.g., ”use specific
language” rather than ”do not use general language”). However, we found it dificult to avoid words like
”not, without, etc.” Ultimately we found that when a negative prompt was required, specifics worked
better than generalities. Telling GPT-4 to avoid pronouns did not work. Telling it to avoid ”he, she, they ...”
worked better.</p>
      <p>At times, we included instructions in the prompt to trigger GPT to evaluate output and repeat steps if
errors were present. GPT never repeated any of the steps, even when we would have liked it to do so.</p>
      <p>One GPT glitch is that it would often stop writing the middle of producing the output. This was solved
every time by prompting the AI to ”continue” but increased the number of prompts per problem. Other,
rarer, glitches included GPT starting on a step other than Step 1a or declaring all steps unnecessary and
skipping to the final output. This was more common with GPT-3 than GPT-4 but did occasionally require
opening a new chat GPT-4 session to fix.
ask the student to match each of the elements , , , , 
with what they represent in that
particular scenario. LLMs struggle to translate language into mathematics. However, they excel
with programming tasks. We overcame GPT’s limitations with mathematics by asking it to
write a python function that takes in the components of , , 
and output  in one prompt, and
then use the function to identify the five components in a second prompt. Figure 1 shows an
example of this prompting and the GPT outcomes.</p>
      <p>GPT was successful in all but one case. Upon further examination of that case, we found that
the original word problem text was written in a way that was overly vague. This highlighted a
use-case for GPT that we had not originally considered. GPT may be used as a quality assurance
tool: if GPT fails at identifying the components, it may indicate that the problem is ambiguous
and should be rewritten.</p>
      <p>For this particular task, zero-shot learning was all that was required. In this study, we carried
out the encodings in a single chain, with zero-shot learning, and in batches of 3-5 problems at a
time. Nevertheless, experimental use of few-shot learning and additional prompting helped
constrain its responses. With the API, it might be more eficient to use multiple chains with
few-shot learning instead.</p>
      <p>GPT often tried to be more helpful by providing code and numerical examples, and often
went on for too long trying to explain the python function. To counter this behavior, we always
added text to our prompt such as, ”Do not provide numbers or code examples.”</p>
    </sec>
    <sec id="sec-6">
      <title>5. LLM Output Analysis</title>
      <p>Readability was assessed across multiple metrics using the Automatic Readability Tool for
English (ARTE) [6]. We selected a subset of the measures provided by ARTE that we believed
were representative. Results from ARTE measures are reported in Table 3. On average, the
Prompt
Write a python function that takes a slope ”m”, an intercept ”c”, an independent value ”x” and
returns a dependent value ”y”. Do not provide numbers or code examples.</p>
      <sec id="sec-6-1">
        <title>Completion</title>
        <p>Here’s a Python function that takes the slope (m), intercept (c), and independent value (x) as
arguments, and returns the dependent value (y) based on the linear equation y = mx + c:
d e f l i n e a r _ e q u a t i o n (m, c , x ) :
y = m ∗ x + c
r e t u r n y
This function first calculates the dependent value (y) using the provided slope (m), intercept (c),
and independent value (x). Then, it returns the computed value (y) as the output.</p>
      </sec>
      <sec id="sec-6-2">
        <title>Prompt</title>
        <p>Apply the python function to the following problem and concisely explain the quantities x, m, c,
m*x, and y. Do not provide any numbers. Do not provide code examples. Provide the answer as
a list. Use conversational language and active voice when possible, and repeat words from the
word problem as much as possible.</p>
        <p>Alex got a gift card for Joe’s Juice Bar and uses it to buy mango juice. Alex received a $25 gift
card. On each visit, Alex buys a glass of mango juice, and the value of the card decreases by $3.
The equation  = 25 − 3 models how the value of the gift card changes.</p>
      </sec>
      <sec id="sec-6-3">
        <title>Completion</title>
        <p>Alex’s gift card for Joe’s Juice Bar:
• x: the number of glasses of mango juice Alex buys
• m: the cost of each glass of mango juice, in dollars
• c: the initial value of the gift card, in dollars
• m*x: the total amount spent on mango juice, in dollars
• y: the remaining value of the gift card after buying mango juice, in dollars
math problems had 46.94 words each. This is considerably less than the recommended number
of words (170-200 words) for achieving reliable results [15][16].4 Thus, we also aggregated all
of the word problems to produce one large text per condition with m = 1,425 words. Results for
the longer text matched those for the shorter texts and are reported in Table 3 for completion.
4Zhou et al. [15] analyzed reading scores across tools and metrics. They found that the same readability metric may
vary across tools due to diferences in its implementation (e.g., how it counts syllables). Further, readability metrics
measuring similar outcomes (e.g., coherence, grade level) also varied in their determination. These diferences
diminished once a text reached 200 words, and meaningful improvement beyond 900 words was not detected.
• CAREC-M (modified Crowdsourced Algorithm of Reading Comprehension (CAREC-M)
[17]) considers features related to syntax, lexical diversity, and cohesion, modified to
account for text length. CAREC-M scores were lowest for the GPT-4 rewrites, indicating
that GPT-4 successfully improved cohesion and lowered lexical diversity as compared to
the original and human-rewritten text.
• SBERT (Sentence-BERT [18, 19]) is a transformer based deep learning model which
assesses the semantic similarity within a text. The higher values for the GPT-4 rewrites
indicate improved readability as compared to the original texts. For the aggregated set of
texts, the GPT-4 rewrites also improved over human rewrites.
• FKGL Flesch-Kincaid Grade Level [16] compares word counts to sentence and syllable
counts to determine the appropriate grade level for a text. FKGL decreases for both sets
of rewrites indicating that humans and GPT used words with fewer syllables and wrote
shorter sentences compared to the original texts.
• NDC (New Dale-Chall [20]) is calculated based on sentence length and the percentage
of words in a text that may be considered unfamiliar (i.e., are not part of a set of 5,000
pre-selected common words). This was the only metric for which the GPT-4 rewrites
declined in performance relative to the original texts. During the prompt revision process,
GPT-4 persistently struggled with identifying and replacing low frequency words. The
results here suggest GPT-4 was unable to reach human proficiency at selecting
gradelevel appropriate words. In Conclusions and Future Work, we discuss how we intend to
improve on this going forward.</p>
      </sec>
    </sec>
    <sec id="sec-7">
      <title>6. Evaluation Plan and Metrics</title>
      <p>
        While we have carried out a text analysis of the problems in their three varieties: original,
human-rewritten, and GPT-rewritten, a more comprehensive method of comparing them is to
measure student outcomes. We plan to run a randomized field trial of the human-rewritten and
GPT-rewritten problems. We plan to evaluate the variants following the metrics laid out in [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ],
namely, median time to completing the workspace, average number of problems completed,
average number of hints used, and promotion rate (percentage of students that failed to reach
mastery on all skills in a workspace). Furthermore, we will use the predictive model from [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ]
to evaluate the variants separately on students predicted to be emerging readers as well as
the overall student population. We hope to achieve similar success with GPT rewrites as was
achieved by the human rewrites, allowing us to use LLMs to scale up this work significantly to
the dozens of workspaces that use word problems in MATHia, likely at a substantially lower
cost.
      </p>
    </sec>
    <sec id="sec-8">
      <title>7. Conclusions and Future Work</title>
      <p>The use of LLMs have allowed us to scale-up and replicate writing tasks that, when previously
done by humans, were found to improve student outcomes. Using chain-of-thought learning,
we were able to use GPT-4 to rewrite math word problems in MATHia following a set of
guidelines intended to support emerging readers. Further, we circumvented GPT-4’s struggle
with formulating and solving mathematical problems by asking it to produce python code which
would encode those problems.</p>
      <p>Text analysis on GPT-4’s output showed that GPT-4’s rewrites improved across several
metrics: readability, lexical diversity, and cohesion; but scored lower on a metric evaluating use
of familiar words. The rewrites produced by GPT-4 are now being presented to MATHia users
as part of a randomized controlled experiment to test their efectiveness. If the GPT-4 rewrites
produce similar advances in math outcomes as the human-rewrites did, it will represent a path
forward to quickly improving learning content throughout MATHia.</p>
      <p>Although our first attempt using chat.openai.com was successful, the final prompt was nearly
three pages and used 2,143 tokens. Further, GPT-4 would occasionally stop in the middle of
production and need to be prompted to continue. On average, we needed to use 3 of our allocated
25 prompts per three hours for each problem. Continuing forward with chat.openai.com presents
a serious barrier to scalability. To scale this work, we are instead using the GPT-4 API. Using
the API, we can reduce the number of tokens per prompt by using natural language processing
in python to evaluate the text (e.g., to label passive sentences, pronouns, and rare words). By
ofloading tagging and evaluation steps (steps 1a-b &amp; d-f, 3a-b, step 8, 9, and 10a-b &amp; d-e of the
ifnal prompt), we have had some success in producing similar revisions to the ones reported
here with as few as 427 tokens.</p>
      <p>LLMs are proving to be useful tools for writing large amounts of content to specification.
GPT4, in particular, can quickly detect and mimic style and structure and implement broad structural
changes. The focus of this work was on improving MATHia word problems to enable emerging
readers to engage more with understanding math and less with parsing text. However, the steps
of producing a prompt to evaluate and revise text as well as identify important components of
math problems are transferable. We found that GPT is a scalable resource for implementing
fast changes which would previously have required hundreds of hours of labor.</p>
      <p>We recommend strategies which consider the nature of the LLM. Although we discuss GPT-4
as ”evaluating text,” in reality, it is a prediction engine that uses prior text to inform its output.
Asking it to define and label text helps to guide the prediction process in a way that ”evaluation”
does not. Similarly, to see GPT-4 generate a greater range of word substitutions, it can help
to provide less constraint in guiding the substitution. In our revisions, GPT-4 appeared to use
the meaning of the words in the original text as a constraint. Asking it to remove the words
freed it from this constraint. Finally, GPT-4 has been extensively trained on python code. We
found it could use that code as a structure for its response when engaging in tasks that require
mathematical reasoning. Keeping these techniques and caveats in mind, we were able to greatly
reduce the time it takes to improve the readability of select MATHia word problems.</p>
      <p>If we find that these GPT-4 rewrites result in better student outcomes, such as reduced time
and increased accuracy for emerging readers, this would allow us to scale-up rewriting content
in many other workspaces. Furthermore, it would be interesting to study how successful GPT-4
or other LLMs are in writing original content that replaces older content while using the same
knowledge component models in MATHia or similar ITSs. A longer-term goal would be more
real-time content generation, where problems can be personalized to students’ interests, without
having a pre-defined bank of problems for each topic. We believe that this work is a first step
towards such personalization, but there are several issues that need to be addressed with LLMs
to achieve such real-time problem personalization. Namely, we seek to insure that LLM outputs
are always consistent, automatically integratable with the ITS, mathematically correct, and
helpful for student learning.</p>
    </sec>
    <sec id="sec-9">
      <title>Acknowledgments</title>
      <p>The research reported here was supported in part by the Institute of Education Sciences, U.S.
Department of Education, through grant R324A210289 to Center for Applied Special Technology
(CAST). The opinions expressed are those of the authors and do not represent the views of the
Institute or the U.S. Department of Education.
[5] OpenAI, ChatGPT (Mar 14 version), https://chat.openai.com/chat, 2023. [Large language
model].
[6] J. Choi, S. A. Crossley, Advances in readability research: A new readability web app
for english, 2022 International Conference on Advanced Learning Technologies (ICALT)
(2022) 1–5.
[7] S. Ritter, J. R. Anderson, K. Koedinger, A. T. Corbett, Cognitive tutor: Applied research in
mathematics education, Psychonomic Bulletin &amp; Review 14 (2007) 249–255.
[8] J. R. Anderson, A. T. Corbett, Knowledge tracing: Modeling the acquisition of procedural
knowledge, User Modeling and User-Adapted Interaction 4 (1994) 253–278.
[9] S. Ritter, A. Murphy, S. E. Fancsali, V. Fitkariwala, J. D. Lomas, Upgrade: An open source
tool to support a/b testing in educational software, in: Proceedings of the First Workshop
on Educational A/B Testing at Scale, EdTech Books, 2020.
[10] M. Palatucci, D. Pomerleau, G. E. Hinton, T. M. Mitchell, Zero-shot learning with semantic
output codes, Advances in neural information processing systems 22 (2009).
[11] Y. Xian, B. Schiele, Z. Akata, Zero-shot learning-the good, the bad and the ugly, in:
Proceedings of the IEEE conference on computer vision and pattern recognition, 2017, pp.
4582–4591.
[12] T. Brown, B. Mann, N. Ryder, M. Subbiah, J. D. Kaplan, P. Dhariwal, A. Neelakantan,
P. Shyam, G. Sastry, A. Askell, et al., Language models are few-shot learners, Advances in
neural information processing systems 33 (2020) 1877–1901.
[13] E. Saravia, Prompt Engineering Guide,
https://github.com/dair-ai/Prompt-Engineering</p>
      <p>Guide (2022).
[14] J. Wei, X. Wang, D. Schuurmans, M. Bosma, E. Chi, Q. Le, D. Zhou, Chain of thought
prompting elicits reasoning in large language models, arXiv preprint arXiv:2201.11903
(2022).
[15] S. Zhou, H. Jeong, P. A. Green, How consistent are the best-known readability equations
in estimating the readability of design standards?, IEEE Transactions on Professional
Communication 60 (2017) 97–111.
[16] J. P. Kincaid, R. P. Fishburne Jr, R. L. Rogers, B. S. Chissom, Derivation of new readability
formulas (automated readability index, fog count and flesch reading ease formula) for navy
enlisted personnel, Technical Report, Naval Technical Training Command Millington TN
Research Branch, 1975.
[17] S. A. Crossley, K. Kyle, M. Dascalu, The tool for the automatic analysis of cohesion 2.0:
Integrating semantic similarity and text overlap, Behavior research methods 51 (2019)
14–27.
[18] S. Crossley, J. S. Choi, Y. Scherber, M. Lucka, Using large language models to develop
readability formulas for educational settings, in: Proceedings of the AIED, 2023.
[19] N. Reimers, I. Gurevych, Sentence-bert: Sentence embeddings using siamese bert-networks,
arXiv preprint arXiv:1908.10084 (2019).
[20] J. S. Chall, E. Dale, Readability revisited: The new Dale-Chall readability formula, Brookline
Books, 1995.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>K. R.</given-names>
            <surname>Koedinger</surname>
          </string-name>
          ,
          <string-name>
            <surname>M. J. Nathan,</surname>
          </string-name>
          <article-title>The real story behind story problems: Efects of representations on quantitative reasoning</article-title>
          ,
          <source>Journal of the Learning Sciences</source>
          <volume>13</volume>
          (
          <year>2004</year>
          )
          <fpage>129</fpage>
          -
          <lpage>164</lpage>
          .
          <source>doi:1 0 . 1 2 0 7 / s 1 5</source>
          <volume>3 2 7 8 0 9 j</volume>
          <source>l s 1 3</source>
          <volume>0 2 \ _ 1</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>H.</given-names>
            <surname>Almoubayyed</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S. E.</given-names>
            <surname>Fancsali</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Ritter</surname>
          </string-name>
          ,
          <article-title>Instruction-embedded assessment for reading ability in adaptive mathematics software</article-title>
          ,
          <source>in: Proceedings of the 13th International Conference on Learning Analytics and Knowledge</source>
          , LAK '23,
          <string-name>
            <surname>Association</surname>
          </string-name>
          for Computing Machinery, New York, NY, USA,
          <year>2023</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>H.</given-names>
            <surname>Almoubayyed</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S. E.</given-names>
            <surname>Fancsali</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Ritter</surname>
          </string-name>
          ,
          <article-title>Generalizing predictive models of reading ability in adaptive mathematics software</article-title>
          ,
          <source>in: International Conference on Educational Data Mining</source>
          <year>2023</year>
          ,
          <fpage>EDM2023</fpage>
          ,
          <year>2023</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>H.</given-names>
            <surname>Almoubayyed</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Bastoni</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Berman</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Galasso</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Jensen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Lester</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Murphy</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Swartz</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Weldon</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S. E.</given-names>
            <surname>Fancsali</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Gropen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Ritter</surname>
          </string-name>
          ,
          <article-title>Rewriting math word problems to improve learning outcomes for emerging readers: A randomized field trial in carnegie learning's mathia</article-title>
          ,
          <source>in: The 24th International Conference on Artificial Intelligence in Education (AIED</source>
          <year>2023</year>
          ),
          <source>AIEd '23</source>
          ,
          <string-name>
            <surname>Springer</surname>
            <given-names>Nature</given-names>
          </string-name>
          ,
          <year>2023</year>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>