<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <journal-title-group>
        <journal-title>M. Mahran);</journal-title>
      </journal-title-group>
    </journal-meta>
    <article-meta>
      <title-group>
        <article-title>Investigating Bias: A Multilingual Pipeline for Generating, Solving, and Evaluating Math Problems with LLMs</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Mariam Mahran</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Katharina Simbeck</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>HTW Berlin University of Applied Sciences</institution>
          ,
          <addr-line>Treskowallee 8, 10318 Berlin</addr-line>
          ,
          <country country="DE">Germany</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2025</year>
      </pub-date>
      <volume>000</volume>
      <fpage>0</fpage>
      <lpage>0001</lpage>
      <abstract>
        <p>Large Language Models (LLMs) are increasingly used for educational support, yet their response quality varies depending on the language of interaction. This paper presents an automated multilingual pipeline for generating, solving, and evaluating math problems aligned with the German K-10 curriculum. We generated 628 math exercises and translated them into English, German, and Arabic. Three commercial LLMs (GPT-4o-mini, Gemini 2.5 Flash, and Qwen-plus) were prompted to produce step-by-step solutions in each language. A held-out panel of LLM judges, including Claude 3.5 Haiku, evaluated solution quality using a comparative framework. Results show a consistent gap, with English solutions consistently rated highest, and Arabic often ranked lower. These ifndings highlight persistent linguistic bias and the need for more equitable multilingual AI systems in education.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;LLMs</kwd>
        <kwd>Fairness</kwd>
        <kwd>Education</kwd>
        <kwd>Language Bias</kwd>
        <kwd>Evaluation</kwd>
        <kwd>Equity</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>
        As Large Language Models (LLMs) become increasingly embedded in educational tools, they are
reshaping how students access academic support [
        <xref ref-type="bibr" rid="ref1 ref2 ref3">1, 2, 3</xref>
        ]. In subjects like mathematics, these models
are often used to explain concepts, walk through problem-solving steps, and supplement classroom
instruction [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ]. Nonetheless, most commercial LLMs are trained on overwhelmingly English-centric
data, raising concerns about how efectively they perform in other languages [
        <xref ref-type="bibr" rid="ref4 ref5">4, 5</xref>
        ].
      </p>
      <p>
        In multilingual educational settings, this imbalance can have serious implications. Students interacting
with the same model in diferent languages may receive responses that vary not only in fluency, but in
clarity, accuracy, and pedagogical value [
        <xref ref-type="bibr" rid="ref1 ref6">1, 6</xref>
        ]. This variation is especially problematic in math education,
where precise explanations and terminology are essential. Prior studies have highlighted significant
language-related disparities in LLM outputs, particularly in educational contexts [
        <xref ref-type="bibr" rid="ref5 ref6">6, 5</xref>
        ]. These findings
underscore the importance of systematic, scalable evaluation across languages.
      </p>
      <p>
        In this paper, we present a scalable, automated pipeline for evaluating how language impacts the
quality of LLM-generated math solutions. Covering generation, translation, solution, and evaluation,
the framework supports consistent multilingual comparison across hundreds of exercises. Using three
commercial models (GPT-4o-mini (OpenAI) [
        <xref ref-type="bibr" rid="ref7">7</xref>
        ], Gemini-2.5-flash (Google) [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ], and Qwen-plus (Alibaba
Cloud) [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ]) we apply it to English, German, and Arabic to uncover persistent performance disparities in
educational contexts. This work aims to drive future research toward reducing linguistic biases in AI
education tools that serve learners regardless of language background.
      </p>
    </sec>
    <sec id="sec-2">
      <title>2. LLMs in education: Use Cases &amp; Biases</title>
      <p>
        LLMs are increasingly used in education to support both teachers and students. For educators, they
assist with lesson planning, content creation, and assessment, helping reduce workload and improve
eficiency [
        <xref ref-type="bibr" rid="ref1 ref2">1, 2</xref>
        ]. For students, they ofer real-time, adaptive support that fosters independent learning,
clarifies dificult concepts, and guides problem-solving [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ].
      </p>
      <p>
        Nonetheless, AI fairness in education is an increasing concern [
        <xref ref-type="bibr" rid="ref10">10</xref>
        ]. Growing evidence shows that
LLMs introduce linguistic biases, particularly disadvantaging non-Western and low-resource language
contexts [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ]. Models often default to Western entities (i.e. names, foods, locations...etc) even when
operating in non-Western languages like Arabic [
        <xref ref-type="bibr" rid="ref11">11</xref>
        ]. Recent work also suggests that multilingual
models often perform internal reasoning in English regardless of the prompt language [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ].
      </p>
      <p>
        These tendencies raise concerns about the inclusivity and cultural relevance of AI-generated
educational content. In mathematical reasoning tasks, LLMs have shown tendency to perform well in
high-resource languages such as English and Chinese but struggle with languages like Korean due
to dificulties processing non-English input [
        <xref ref-type="bibr" rid="ref6">6</xref>
        ]. Recent work by [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ] further demonstrate this gap by
evaluating six LLMs on four educational tasks across six non-English languages. They find performance
drops sharply in low-resource languages, correlating with training data representation. English prompts
was also shown to outperform their translations across tasks.
      </p>
      <p>
        Such biases have real consequences. In grading and assessment contexts, models may favor Western
argument structures, writing styles, and measurement systems placing non-Western students at a
disadvantage. More broadly, the dominance of English-centric models reflects global educational
inequality, as advanced AI tools remain most accessible to wealthier, English-speaking regions [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ].
      </p>
    </sec>
    <sec id="sec-3">
      <title>3. Dataset Creation</title>
      <sec id="sec-3-1">
        <title>3.1. Dataset Generation &amp; Translation</title>
        <p>
          This study uses a dataset of 628 math exercises spanning five topic areas and grade levels B through H
(equivalent to Grades 2-10), based on the oficial German K-10 mathematics curriculum [
          <xref ref-type="bibr" rid="ref12">12</xref>
          ].1 Exercises
were generated using ChatGPT-4o (via the ChatGPT interface) in a manual, iterative process guided
by curriculum-aligned topic names and learning objectives. Rather than replicating specific examples,
the curriculum served as a conceptual reference to ensure broad topic coverage and grade-level
appropriateness. Each item was then manually reviewed for clarity, pedagogical relevance, and alignment
with the intended mathematical concept. Unclear, repetitive, or misaligned exercises were revised or
removed. The resulting dataset consists of brief, one-line prompts focused on single operations or
concepts, minimizing the risk for any contextual or linguistic bias. Table 1 shows example exercises.
        </p>
        <p>All exercises were initially written in English, then translated into German and Arabic using the
GPT-4o-mini API. Given their concise and context-independent structure, the exercises were well-suited
for machine translation. Nonetheless, all translations were manually audited to ensure semantic and
linguistic appropriateness. Minor edits were applied as needed to maintain consistency across languages.</p>
      </sec>
      <sec id="sec-3-2">
        <title>3.2. Solution Generation</title>
        <p>Three LLMs were used in this study: GPT-4o-mini, Gemini 2.5 Flash, and Qwen-plus. Each was prompted
to generate step-by-step solutions in English, German, and Arabic. The prompts, written in the target
language, asked the model to explain how to solve the given problem rather than to provide a final
answer. This setup was intended to mirror how K-10 learners might seek guided support or clarification.
The resulting solutions were saved in separate CSV files for each language-model combination.
1Code and data are available at https://github.com/iug-htw/solution_evaluator.</p>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>4. Evaluation</title>
      <p>
        To evaluate solution quality, we adopted a hybrid evaluation framework in which LLMs acted as judges.
LLM-based evaluation is gaining traction as a scalable alternative to human or code-based assessments
[
        <xref ref-type="bibr" rid="ref13">13</xref>
        ]. Findings suggest that LLMs can generate assessments comparable to human judgments in domains
such as open-ended story generation and adversarial attack evaluations [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ]. In educational contexts,
LLM judgments have shown strong correlation with faculty evaluations, particularly when using
pairwise comparisons and structured rubrics [
        <xref ref-type="bibr" rid="ref15">15</xref>
        ]. Other studies have shown that, with appropriate
finetuning and prompt design, LLMs can efectively automate course evaluations across diverse university
settings [16].
      </p>
      <p>For each math exercise, three step-by-step solutions, one in English, German, and Arabic, were
presented as alternative responses to the same prompt. These were then assessed and compared by a
panel of LLM judges based on their clarity, accuracy, and pedagogical efectiveness.</p>
      <p>To ensure impartiality and prevent indirect bias from models evaluating their own outputs, we
implemented a held-out judging strategy. In each round, the model under evaluation was excluded
from the judging panel and replaced with a neutral model, Claude 3.5 Haiku [17]. The remaining three
judges were selected from GPT-4o-mini, Gemini-2.5-flash, Qwen-plus, and Claude.</p>
      <p>Prior to evaluation, we identified technical terms relevant to each exercise, as the accurate use and
explanation of such terms are crucial for clarity and learner comprehension. To do so, we employed a
script that extracted these terms from each exercise using the GPT-4o-mini model.</p>
      <p>
        During evaluation, solutions were presented in randomized order to each judge to mitigate position
bias. Position bias occurs when the order of presentation influences rankings, a phenomenon observed
both in LLMs and human decision-making [
        <xref ref-type="bibr" rid="ref13">18, 13</xref>
        ].
      </p>
      <p>We followed a comparative assessment methodology, which has been shown to outperform direct
scoring in terms of alignment with human preferences [19, 20]. Judges were asked to rank the three
solutions from 1 (best) to 3 (worst) and provide a short justification to their judgment. Rankings were
based on multiple educational criteria: conceptual understanding, clarity of explanation, step-by-step
reasoning, correct terminology, accuracy, recognition of common mistakes, pedagogical suitability,
generalizability, and grade-level appropriateness.</p>
      <p>To identify the best-performing language for each exercise, we applied a majority voting scheme. If
two or more judges ranked the same solution highest, the language of that solution was recorded as the
top performer. If the judges couldn’t agree on a single best-performing solution, the result was labeled
a tie (TIE), ensuring that ambiguous cases were recorded without forcing a winner.</p>
    </sec>
    <sec id="sec-5">
      <title>5. Results</title>
      <p>Table 2 summarizes the performance of solutions per language, categorizing them for each LLM based
on how frequently they were ranked as the best, mid-tier, or worst. Across all three models, English
solutions were most frequently favored. GPT-4o-mini showed the strongest language preference, heavily
favoring English and consistently rating Arabic lowest. Qwen-plus followed a similar trend, though
with less extreme separation. Gemini-2.5-flash produced a more balanced distribution overall; while
English still led in top rankings, the diference between German and Arabic was narrower, and German
received the most “Worst” labels in that case.</p>
      <p>In addition to collecting quantitative rankings, we asked the LLM judges to provide brief justifications
for each evaluation (see examples in Table 3). Analysis of over 4,000 justifications revealed consistent
emphasis on clarity, structured reasoning, and appropriate use of mathematical terminology. English
solutions were most frequently praised as “comprehensive,” “well-structured,” or “clear,” reflecting
strong positive sentiment overall. German responses received more mixed feedback, often described
as adequate but occasionally critiqued for being “less detailed” or assuming prior knowledge. Arabic
solutions had the highest concentration of negative sentiment, with justifications noting a lack of clarity
or depth in explanation. Notably, comparative framing (e.g., “slightly better,” “strongest”) was common,
confirming that judges made relative, not absolute, assessments.</p>
    </sec>
    <sec id="sec-6">
      <title>6. Discussion</title>
      <p>The results indicate a clear disparity in performance across languages. English solutions were rated most
frequently the best, while Arabic was generally rated the worst, although not uniformly across models.
German solutions typically occupied the middle range. GPT-4o-mini showed the most extreme language
separation, assigning overwhelmingly high scores to English and low scores to Arabic. Gemini-2.5-flash
diverged slightly, assigning the highest number of “Worst” rankings to German rather than Arabic.
However, the margin between German and Arabic was relatively small (246 vs. 223), and Gemini’s
rankings were more evenly distributed across languages overall.</p>
      <p>It is worth noting that, while we did not repeat the full experiment under fixed conditions, the full
pipeline was executed multiple times during development with variations in datasets, prompts, and
code refinements. Despite fluctuations in exact scores (as expected with stochastic outputs), the relative
pattern remained consistent: English was favored, Arabic was disadvantaged, and German remained in
between. This recurring trend across settings suggests that the observed disparities are robust rather
than incidental.</p>
      <p>These gaps were not solely driven by solution correctness. Even when all three solutions were
mathematically accurate, models showed clear preferences for responses that included explicit reasoning,
pedagogical structure, and generalizability. Analysis of the provided justifications show that the LLMs
consistently reward explanations that go beyond getting the answer right to showing why the method
works. Another key insight is that each language showed strengths in diferent areas. Arabic was often
rewarded for its warmth and accessibility, especially at lower grade levels. German performed well
when structure and terminology were key but sometimes assumed prior knowledge. These patterns
suggest that language-specific conventions shape both how explanations are presented and how models
evaluate their educational value.</p>
      <p>Nonetheless, variations in linguistic complexity likely contributed to the observed performance
patterns. German has a complex sentence structure with long compound words, flexible word order,
[en]: "...well-structured and clearly explains the concept [en]: "...while correct, lacks some of the engaging
eleusing appropriate terminology and visual analogies." ments...may not connect with younger students"
[de]: "...excels in clarity, structure, and appropriate use [de]: "... although correct, assumes prior knowledge of
of technical terms...avoids assumptions..." rearranging the density formula..."
[ar]: "...provides a student-friendly approach to solving [ar]: "... presents the information in a slightly less
structhe problem...making it easy for a 4th grader to follow." tured manner..."
and grammatical rules. These characteristics can make mathematical expression and phrasing more
dificult for LLMs to manage, which may explain its intermediate performance. Arabic with its rich
morphology and right-to-left script posed even greater challenges. This morphological richness can
compress tokens, producing shorter outputs that feel abrupt next to more elaborated English or German
responses. Moreover, LLMs may also lack exposure to Arabic instructional texts with formal, school-style
structure. As a result, LLMs may misinterpret Arabic’s concise or direct solutions as less educationally
complete, even when correct. These surface-level diferences can shape how LLMs perceive clarity and
instructional value, even when conceptual accuracy is maintained.</p>
      <p>The decision to fully automate the pipeline was driven by the need for scalable, reproducible analysis.
Manual evaluation methods, while valuable, limits consistency and scope, making it dificult to compare
results at scale. Our automated approach ensures uniform treatment across hundreds of exercises and
enables easy replication or extension to new models, languages, or curricula. It also reduces human
bias in assessment and supports more systematic multilingual benchmarking.</p>
      <p>Finally, the findings carry clear pedagogical implications. Teachers cannot assume equal support
across languages, as LLMs may provide stronger guidance in English than other languages. LLMs should
therefore be treated as supplementary tools, ideally with teacher oversight or cross-lingual checks. At
the same time, the pipeline itself serves as an analytical resource, helping educators pinpoint curriculum
areas where linguistic gaps pose the greatest risk, such as topics requiring precise terminology or
multi-step reasoning. This allows teachers to adapt instruction accordingly to promote more equitable
learning in multilingual classrooms.</p>
    </sec>
    <sec id="sec-7">
      <title>7. Conclusion</title>
      <p>This study presents a scalable, fully automated pipeline for evaluating LLM-generated math solutions
across languages. Applying it to English, German, and Arabic revealed consistent disparities in
explanation quality, with English generally favored and Arabic more frequently disadvantaged. These
diferences were not solely about correctness but reflected variation in reasoning depth and pedagogical
clarity.</p>
      <p>As LLMs become increasingly integrated into classrooms and tutoring platforms, such disparities
risk reinforcing existing educational inequalities. Our findings underscore the importance of more
linguistically inclusive development and evaluation practices. To ensure equitable access, future work
should prioritize improving performance in underrepresented languages through targeted fine-tuning,
diverse training data, and culturally informed prompt design. Expanding this evaluation framework to
additional languages and domains can support the development of more accessible and fair educational
AI systems worldwide.</p>
    </sec>
    <sec id="sec-8">
      <title>Acknowledgments</title>
      <p>This paper portrays the work carried out in the context of the KIWI project (16DHBKI071) that is
generously funded by the Federal Ministry of Research, Technology and Space (BMFTR).</p>
    </sec>
    <sec id="sec-9">
      <title>Declaration on Generative AI</title>
      <p>During the preparation of this work, the authors used OpenAI’s GPT-4.1 to check grammar and spelling
and to enhance the writing style. After using these tools/services, the authors reviewed and edited the
content as needed and take full responsibility for the publication’s content.
[16] B. Yuan, J. Hu, An Exploration of Higher Education Course Evaluation by Large Language Models,
2024. arXiv:2411.02455.
[17] Anthropic, Claude 3.5 haiku, 2024. URL: https://www.anthropic.com/news/
3-5-models-and-computer-use.
[18] S. Tan, S. Zhuang, K. Montgomery, W. Y. Tang, A. Cuadron, C. Wang, R. A. Popa, I. Stoica,</p>
      <p>JudgeBench: A Benchmark for Evaluating LLM-based Judges, 2024. arXiv:2410.12784.
[19] Y. Liu, H. Zhou, Z. Guo, E. Shareghi, I. Vulić, A. Korhonen, N. Collier, Aligning with Human
Judgement: The Role of Pairwise Preference in Large Language Model Evaluators, in: D. Das,
D. Chen, Y. Artizi, A. Fan (Eds.), First Conference on Language Modeling COLM 2024, OpenReview,
2024. URL: https://openreview.net/group?id=colmweb.org/COLM/2024/Conference#tab-accept,
ifrst Conference on Language Modeling 2024, COLM 2024.
[20] A. Liusie, P. Manakul, M. J. F. Gales, LLM comparative assessment: Zero-shot NLG evaluation
through pairwise comparisons using large language models, in: Y. Graham, M. Purver (Eds.),
Proceedings of the 18th Conference of the European Chapter of the Association for Computational
Linguistics (Volume 1: Long Papers), Association for Computational Linguistics, St. Julian’s, Malta,
2024, pp. 139–151. doi:10.18653/v1/2024.eacl-long.8.</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>L.</given-names>
            <surname>Yan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Sha</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Zhao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Martinez-Maldonado</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Chen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Jin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Gašević</surname>
          </string-name>
          ,
          <article-title>Practical and ethical challenges of large language models in education: A systematic scoping review</article-title>
          ,
          <source>British Journal of Educational Technology</source>
          <volume>55</volume>
          (
          <year>2023</year>
          )
          <fpage>90</fpage>
          -
          <lpage>112</lpage>
          . doi:
          <volume>10</volume>
          .1111/bjet.13370.
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>W.</given-names>
            <surname>Gan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Qi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Wu</surname>
          </string-name>
          ,
          <string-name>
            <surname>J. C.-W. Lin</surname>
          </string-name>
          ,
          <article-title>Large Language Models in Education: Vision and Opportunities</article-title>
          , in: 2023
          <source>IEEE International Conference on Big Data (BigData)</source>
          ,
          <source>IEEE Computer Society</source>
          , Los Alamitos, CA, USA,
          <year>2023</year>
          , pp.
          <fpage>4776</fpage>
          -
          <lpage>4785</lpage>
          . doi:
          <volume>10</volume>
          .1109/BigData59044.
          <year>2023</year>
          .
          <volume>10386291</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>Q.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Fu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W.</given-names>
            <surname>Zhang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Chen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Yu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W.</given-names>
            <surname>Xia</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W.</given-names>
            <surname>Zhang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Tang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Yu</surname>
          </string-name>
          ,
          <article-title>Adapting Large Language Models for Education: Foundational Capabilities, Potentials, and</article-title>
          <string-name>
            <surname>Challenges</surname>
          </string-name>
          ,
          <year>2024</year>
          . arXiv:
          <volume>2401</volume>
          .
          <fpage>08664</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>L.</given-names>
            <surname>Schut</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Gal</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Farquhar</surname>
          </string-name>
          ,
          <article-title>Do multilingual LLMs think in english?</article-title>
          ,
          <source>in: ICLR 2025 Workshop on Building Trust in Language Models and Applications</source>
          ,
          <year>2025</year>
          . URL: https://openreview.net/forum? id=
          <fpage>I8BOtOPcOv</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>V.</given-names>
            <surname>Gupta</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S. P.</given-names>
            <surname>Chowdhury</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Zouhar</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Rooein</surname>
          </string-name>
          ,
          <string-name>
            <surname>M. Sachan,</surname>
          </string-name>
          <article-title>Multilingual performance biases of large language models in education, 2025</article-title>
          . arXiv:
          <volume>2504</volume>
          .
          <fpage>17720</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>H.</given-names>
            <surname>Ko</surname>
          </string-name>
          ,
          <string-name>
            <given-names>G.</given-names>
            <surname>Son</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Choi</surname>
          </string-name>
          , Understand,
          <source>Solve and Translate: Bridging the Multilingual Mathematical Reasoning Gap</source>
          ,
          <year>2025</year>
          . arXiv:
          <volume>2501</volume>
          .
          <fpage>02448</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7] OpenAI, GPT-4o
          <string-name>
            <surname>Mini: Advancing Cost-Eficient Intelligence</surname>
          </string-name>
          ,
          <year>2024</year>
          . URL: https://openai.com/index/ gpt-4o
          <article-title>-mini-advancing-cost-eficient-intelligence/.</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <surname>Google</surname>
            <given-names>DeepMind</given-names>
          </string-name>
          ,
          <source>Gemini</source>
          <volume>2</volume>
          .
          <article-title>5 flash: high-throughput, reasoning -capable large language model</article-title>
          , https://cloud.google.com/vertex-ai/generative-ai/docs/models/gemini/2-5-flash,
          <year>2025</year>
          . Accessed:
          <fpage>2025</fpage>
          -07-23.
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>A.</given-names>
            <surname>Cloud</surname>
          </string-name>
          , Qwen-Plus:
          <article-title>large language models (LLMs) independently developed by Alibaba Cloud</article-title>
          ,
          <year>2025</year>
          . URL: https://www.alibabacloud.com/help/en/model-studio/
          <article-title>developer-reference/ what-is-qwen-llm.</article-title>
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>K.</given-names>
            <surname>Simbeck</surname>
          </string-name>
          ,
          <article-title>They shall be fair, transparent, and robust: auditing learning analytics systems</article-title>
          ,
          <source>AI and Ethics</source>
          <volume>4</volume>
          (
          <year>2024</year>
          )
          <fpage>555</fpage>
          -
          <lpage>571</lpage>
          . doi:
          <volume>10</volume>
          .1007/s43681-023-00292-7.
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <given-names>T.</given-names>
            <surname>Naous</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. J.</given-names>
            <surname>Ryan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Ritter</surname>
          </string-name>
          , W. Xu,
          <article-title>Having Beer after Prayer? Measuring Cultural Bias in Large Language Models, in: Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics</article-title>
          (Volume
          <volume>1</volume>
          :
          <string-name>
            <surname>Long</surname>
            <given-names>Papers)</given-names>
          </string-name>
          ,
          <source>Association for Computational Linguistics</source>
          , Bangkok, Thailand,
          <year>2024</year>
          , pp.
          <fpage>16366</fpage>
          -
          <lpage>16393</lpage>
          . doi:
          <volume>10</volume>
          .18653/v1/
          <year>2024</year>
          .
          <article-title>acl-long</article-title>
          .
          <volume>862</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <article-title>Ministerium für Bildung, Jugend und Sport des Landes Brandenburg und Senatsverwaltung für Bildung, Jugend und Wissenschaft des Landes Berlin, Rahmenlehrplan für die Jahrgangsstufen 1-10:</article-title>
          <string-name>
            <surname>Mathematik</surname>
          </string-name>
          ,
          <year>2015</year>
          . Amtliche Fassung, gültig für Berlin und Brandenburg.
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <given-names>L.</given-names>
            <surname>Zheng</surname>
          </string-name>
          , W.-L. Chiang,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Sheng</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Zhuang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Wu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Zhuang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Lin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E. P.</given-names>
            <surname>Xing</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Zhang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. E.</given-names>
            <surname>Gonzalez</surname>
          </string-name>
          ,
          <string-name>
            <surname>I. Stoica</surname>
          </string-name>
          ,
          <article-title>Judging LLM-as-a-judge with MT-bench and Chatbot Arena</article-title>
          ,
          <source>in: Proceedings of the 37th International Conference on Neural Information Processing Systems</source>
          , NIPS '23, Curran Associates Inc.,
          <string-name>
            <surname>Red</surname>
            <given-names>Hook</given-names>
          </string-name>
          ,
          <string-name>
            <surname>NY</surname>
          </string-name>
          , USA,
          <year>2023</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <surname>C.-H. Chiang</surname>
          </string-name>
          , H.-y. Lee,
          <article-title>Can Large Language Models Be an Alternative to Human Evaluations?</article-title>
          , in: A.
          <string-name>
            <surname>Rogers</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          <string-name>
            <surname>Boyd-Graber</surname>
          </string-name>
          , N. Okazaki (Eds.),
          <source>Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume</source>
          <volume>1</volume>
          :
          <string-name>
            <surname>Long</surname>
            <given-names>Papers)</given-names>
          </string-name>
          ,
          <source>Association for Computational Linguistics</source>
          , Toronto, Canada,
          <year>2023</year>
          , pp.
          <fpage>15607</fpage>
          -
          <lpage>15631</lpage>
          . doi:
          <volume>10</volume>
          .18653/v1/
          <year>2023</year>
          .
          <article-title>acl-long</article-title>
          .
          <volume>870</volume>
          .
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <given-names>T.</given-names>
            <surname>Ishida</surname>
          </string-name>
          , T. Liu,
          <string-name>
            <given-names>H.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W.</given-names>
            <surname>Cheung</surname>
          </string-name>
          ,
          <article-title>Large Language Models as Partners in Student Essay Evaluation</article-title>
          ,
          <source>Technical Report</source>
          , Cornell University, United States,
          <year>2024</year>
          . arXiv:
          <volume>2405</volume>
          .
          <fpage>18632</fpage>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>