<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta />
    <article-meta>
      <title-group>
        <article-title>SVATS at PAN 2025 TextDetox: Can Small Models Outperform Large Ones in Text Detoxification?</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Vladislav Kozlovskiy</string-name>
          <email>Vladislav.Kozlovskiy@skoltech.ru</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Sameer Tantry</string-name>
          <email>Sameer.Tantry@skoltech.ru</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Alexander Ploskin</string-name>
          <email>Alexander.Ploskin@skoltech.ru</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Tatyana Matveeva</string-name>
          <email>Tatyana.Matveeva@skoltech.ru</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Sofya Savelyeva</string-name>
          <email>Sofya.Savelyeva@skoltech.ru</email>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Skolkovo Institute of Science and Technology</institution>
          ,
          <addr-line>Bolshoy Boulevard, 30, p.1, 121205, Moscow, Russian Federation</addr-line>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2025</year>
      </pub-date>
      <abstract>
        <p>Toxic language, which includes hate speech, insults, and ofensive expressions, poses significant challenges to online communication, mental health, and social cohesion. Additional complications arise in multilingual environments where the development of generalized solutions remains a persistent challenge due to linguistic diversity and resource constraints. In this work, we systematically investigate the efectiveness of existing smalland medium-scale models for multilingual text detoxification, addressing the critical need for computationally eficient approaches that maintain performance in diverse linguistic contexts while operating within practical resource limitations.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;Toxicity Mitigation</kwd>
        <kwd>Few-Shot</kwd>
        <kwd>Fine-Tunning</kwd>
        <kwd>Style Transfer</kwd>
        <kwd>Multilingual prompting</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>The rapid proliferation of user-generated content across digital platforms has underscored the
critical need for automated text detoxification systems. Toxic language, including speech, insults, and
ofensive expressions, poses significant challenges to online communication, mental health, and social
cohesion. Although considerable progress has been made in toxicity detection, the task of text
detoxification—rewriting toxic text into non-toxic alternatives while preserving meaning and fluency—remains
a complex and underexplored problem.</p>
      <p>
        Recent advancements in natural language processing (NLP), particularly the rise of large language
models (LLMs), have opened new avenues for text style transfer and content moderation. The
PANDetox Competition 2024 [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ] has played a pivotal role in benchmarking state-of-the-art detoxification
methods, providing a standardized evaluation framework and diverse datasets. Building upon these
eforts, this paper is written as a part of the PAN-Detox Competition 2025 [
        <xref ref-type="bibr" rid="ref2">2</xref>
        ] and investigates the
efectiveness of various detoxification approaches, including fine-tuned LLMs, sequence-to-sequence
models, and techniques utilizing synthetic data generation.
      </p>
      <p>Our contributions are as follows.</p>
      <p>
        • Review of known methods: We analyze top-performing models from the previous year’s
competition and other SOTA methods and provide a review of existing datasets.
• A methodology for synthetic data generation: While abundant work exists on toxicity mitigation
for popular languages, niche languages such as Tatar or Hinglish (the mix of Hindi and English)
lack paired toxic-nontoxic data, thus requiring additional eforts for artificial generation.
• Experiments with a variety of models and techniques: We experimented with a set of techniques
including fine-tuning, few-shot prompting, and others across diferent models and languages. As
our main result, we propose several comparatively small models with 1 to 8 billion parameters
that achieve higher scores than one of the previous year’s baseline [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ], which is a 13B
model.
• Our solution achieved 10th place overall on the competition leaderboard1, and ranked 6th in four
languages: Amharic, Italian, Tatar, and Hindi.
      </p>
      <p>By addressing the trade-ofs between detoxification strength and text quality, this study advances the
development of safer, more inclusive digital communication tools. Our findings not only contribute
to the academic discourse on text style transfer, but also ofer practical implications for social media
platforms, content moderators, and AI ethics researchers.</p>
    </sec>
    <sec id="sec-2">
      <title>2. Related Works</title>
      <p>Text detoxification transforms toxic text into neutral language while preserving meaning. Early baselines
such as delete, duplicate, and backtranslation ofer simple solutions but often compromise fluency and
semantic accuracy, establishing the foundational challenges that subsequent research has been aimed at
addressing.</p>
      <p>
        Building upon these limitations, large language models have enabled more sophisticated and efective
approaches to text detoxification. The multitask model mT0 [
        <xref ref-type="bibr" rid="ref3">3</xref>
        ] demonstrates strong zero-shot and
fewshot detoxification capabilities through prompt-based multitask learning, generalizing across languages
without task-specific fine-tuning. This advancement represents a significant departure from rule-based
methods toward more nuanced understanding of linguistic toxicity patterns.
      </p>
      <p>
        The efectiveness of these modern approaches is further validated in the PAN 2024 Multilingual Text
Detoxification Task [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ], which highlighted the persistent challenges in multilingual detoxification and
underscored the critical importance of meaning preservation. Notably, few-shot prompting emerged as
a particularly efective method in this competition, especially when applied to instruction-tuned models
like mT0, demonstrating the practical viability of prompt-based approaches in real-world scenarios.
      </p>
      <p>
        The success of few-shot prompting can be attributed to its ability to allow models to learn
detoxification patterns from a handful of carefully selected examples [
        <xref ref-type="bibr" rid="ref5">5</xref>
        ], thereby enabling efective generalization
in low-resource and crosslingual settings where traditional supervised learning approaches would fail.
Models like mT0 benefit greatly from this approach due to their inherent multitask training paradigm,
which facilitates rapid adaptation to new detoxification contexts.
      </p>
      <p>
        However, the scarcity of high-quality training data remains a significant bottleneck in many languages
and domains. To address this challenge, synthetic data generation has emerged as a crucial technique
for providing paired toxic and non-toxic examples, particularly supporting training scenarios where
manually annotated data are scarce [
        <xref ref-type="bibr" rid="ref6 ref7">6, 7</xref>
        ]. When combined with parameter-eficient fine-tuning methods
like LoRA [
        <xref ref-type="bibr" rid="ref8">8</xref>
        ], which updates only a small subset of model weights while maintaining performance,
this approach enables scalable and cost-efective model adaptation across diverse linguistic contexts.
      </p>
      <p>
        Contemporary research continues to push the boundaries of detoxification performance through
the deployment of advanced architectures. State-of-the-art models such as Gemma3–4B [
        <xref ref-type="bibr" rid="ref9">9</xref>
        ], T5 [10],
and Qwen2-7B [11], when strategically combined with synthetic data generation and sophisticated
prompting techniques, continue to advance detoxification performance substantially beyond traditional
baselines, establishing new benchmarks for both efectiveness and eficiency in multilingual text
detoxification tasks.
      </p>
      <sec id="sec-2-1">
        <title>1https://codalab.lisn.upsaclay.fr/competitions/22396#results</title>
      </sec>
    </sec>
    <sec id="sec-3">
      <title>3. Methodology</title>
      <p>In this section, we provide an elaboration of our proposed solution framework and systematically
formulate the hypotheses we rigorously test throughout our experimental investigations. We present
an analysis of our methodological approach, including the derivation of our design choices and the
empirical validation strategies employed to assess their efectiveness. Additionally, we provide an
extensive overview of the synthetic data generation techniques.
3.1. Data
This section presents an overview of the datasets we employ in our training framework, which
encompasses both established detoxification datasets from prior research and synthetically generated
data. While existing paired datasets demonstrate high quality and have been instrumental in training
state-of-the-art solutions across multiple languages (English, Spanish, German, Russian, Ukrainian,
French), they remain scarce or entirely absent for many low-resource languages (e.g. Amharic).
Although unpaired datasets can be assembled through web scraping techniques combined with toxicity
classifiers, this approach represents an active area of ongoing research with inherent limitations.</p>
      <p>In this study, we categorize languages into high-resource and low-resource classifications based
on the availability and comprehensiveness of open-source detoxification datasets. Under this
framework, English, Spanish, Russian, Ukrainian, German, and French are classified as high-resource
languages due to their substantial paired detoxification data availability. Conversely, Italian, Arabic,</p>
      <sec id="sec-3-1">
        <title>Hebrew, Hindi, Tatar, Japanese, Chinese, Hinglish, and Amharic are designated as low-resource</title>
        <p>languages, reflecting the limited or absent paired datasets for these linguistic contexts.</p>
        <p>This resource-based taxonomy directly influences our experimental design and synthetic data
generation priorities, with low-resource languages requiring more extensive augmentation strategies to
achieve comparable training data volumes.</p>
        <p>Given the current requirement for paired datasets in our methodology, we conducted extensive
experiments with various synthetic data generation approaches, which are detailed in the subsequent
sections. The impact of these diferent synthetic data generation strategies on model performance is
systematically evaluated and discussed in further analysis.
3.1.1. Existing Paired Datasets
The landscape of paired detoxification datasets reveals substantial disparities in data availability across
languages. Table 1 summarizes the key characteristics of existing paired datasets, including their size,
language coverage, and potential applications in our training framework.</p>
        <p>The Multilingual ParaDetox dataset is provided by competition organizers, despite its limited size
of 400 samples per language, it enables few-shot learning experiments across most of target languages
due to high quality of data.</p>
        <p>The Multilingual Transformer Detoxification dataset represents the most rich resource containing
55, 000 examples across 9 languages, its efectiveness is proven as it is instrumental in training the
previous year’s competition winner. The dataset’s foundation on translated English content highlights
its limited capabilities in complex and rare languages, like Amharic.</p>
        <p>The SynthDetoxM dataset introduces a valuable synthetic data component, containing 16, 000 paried
examples across four languages (Russian, German, French, and Spanish). This dataset is generated using
modern large language models in few-shot setup. This dataset addresses critical gaps in training data,
particularly for French, which previously lacked substantial paried resources.</p>
        <p>Language-specific datasets provide targeted enhancement opportunities. The ParaDetox (English)
(20,000 examples), ParaDetox (Russian) (10,000 examples), and ParaDetox (Ukrainian) (4,000
examples) ofer substantial monolingual training data, while ParaDetox (Spanish) (500 examples)
provide more limited but valuable language-specific resources.</p>
        <p>The analysis reveals significant data scarcity for several competition languages, though the addition
of SynthDetoxM notably improves coverage for French. Most critically, no paired datasets exist for
Italian, Hebrew, Hinglish, Tatar, and Japanese, representing a substantial gap in training resources.
These languages will require synthetic data generation or cross-lingual transfer learning approaches.</p>
        <p>Among languages with available data, English and Russian demonstrate the strongest resource
availability, with multiple datasets totaling over 85,000 and 81,000 examples respectively (including
SynthDetoxM contributions). German benefits significantly from SynthDetoxM, increasing available
training data substantially. French now has access to paired data through SynthDetoxM, addressing
a previous critical gap. Ukrainian, Spanish, Amharic, Chinese, Arabic, and Hindi have
moderate coverage through multilingual datasets, with Spanish additionally benefiting from SynthDetoxM
augmentation.</p>
        <p>The paired datasets can be integrated into a unified training framework , with the Multilingual
Transformer dataset as the core due to its size and efectiveness. Language-specific datasets refine
models for English, Russian, Ukrainian, and Spanish, while SynthDetoxM adds synthetic data for
Russian, German, French, and Spanish. SynthDetoxM’s synthetic data complements human-annotated
sets, enhancing model generalization, especially for languages with limited data.</p>
        <p>For languages lacking paired data, multilingual corpora support cross-lingual transfer, and the
oficial Multilingual ParaDetox dataset provides evaluation benchmarks. Data gaps remain for five
languages, necessitating ongoing synthetic data generation. This enriched dataset landscape, boosted
by SynthDetoxM, advances training resource balance across target languages and underscores synthetic
data’s importance for multilingual coverage. Data availability is uneven: English, Russian, Ukrainian,
and German have suficient data, while most languages lack enough paired examples. This scarcity
challenges synthetic data generation.
3.1.2. Synthetic Data Generation
Given the limited availability of paired detoxification datasets across target languages, we implemented
a unified synthetic data generation framework with two complementary data sourcing approaches
to augment our training corpus.</p>
        <p>Our synthetic data generation employs a standardized multi-stage pipeline that processes diferent
initial data sources through consistent toxification and quality assurance procedures. The framework
diferentiates primarily in data acquisition strategies and multilingual expansion approaches,
while maintaining uniform processing standards across both pathways.</p>
        <p>
          The unified pipeline consists of the following stages:
1. Initial Data Acquisition: Two distinct sourcing strategies provide the foundation corpus.
2. Toxicity Filtering: Application of toxicity classifier [15] to ensure baseline corpus quality.
3. Lexicon-Guided Toxification : Incorporation of toxic lexical items from multilingual toxic
lexicon [
          <xref ref-type="bibr" rid="ref10">18</xref>
          ] through few-shot prompting with DeepSeek-V3 model. The few-shot examples is
takedn from the ParaDetox dataset provided by the organizers of the competition [12], ensuring
alignment with the linguistic distribution characteristics of the test corpus.
4. Quality Assurance: Secondary toxicity filtering to validate appropriate toxicity levels and
semantic coherence.
5. Multilingual Expansion: Target language generation or translation to produce final multilingual
datasets.
        </p>
      </sec>
      <sec id="sec-3-2">
        <title>Data Sourcing Strategies:</title>
      </sec>
      <sec id="sec-3-3">
        <title>Strategy A: Synthetic Content Generation leverages the tweet-like characteristics observed</title>
        <p>
          in existing datasets through LLM-based content creation. We utilize Qwen3-32B [
          <xref ref-type="bibr" rid="ref11">19</xref>
          ] to generate
controversial tweets attributed to famous personas, followed by non-toxic but disagreeable responses.
This approach produces 10,000 samples per target language through direct multilingual generation,
ensuring consistent coverage across all competition languages.
        </p>
      </sec>
      <sec id="sec-3-4">
        <title>Strategy B: Real-World Data Foundation addresses potential LLM bias by incorporating authen</title>
        <p>
          tic human discourse as the baseline corpus. We collect approximately 232,000 English-language
comments from a carefully moderated online platform 2, providing diverse and linguistically natural
foundation content. Multilingual expansion occurs through translation using DeepSeek-V3 [
          <xref ref-type="bibr" rid="ref12">20</xref>
          ] after
toxification processing.
        </p>
      </sec>
      <sec id="sec-3-5">
        <title>The framework employs two distinct multilingual strategies:</title>
        <p>- Direct Generation: Strategy A generates content directly in 15 target languages during the
initial content creation phase, leveraging the multilingual capabilities of Qwen3-32B and DeepSeek-V3.
- Translation-Based: Strategy B processes English content through the complete pipeline before
translating validated toxic-neutral pairs into 15 target languages using DeepSeek-V3.</p>
        <p>Both strategies implement identical quality assurance protocols: - Pre-toxification filtering
ensures clean baseline content - Post-toxification validation confirms appropriate toxicity levels
Toxicity score thresholding ensures dataset consistency</p>
        <p>This unified framework produces toxic-neutral pairs through complementary approaches: Strategy
A ofers consistent cross-lingual generation with controlled content characteristics, while Strategy B
provides authentic human discourse foundation with superior linguistic diversity and reduced artificial
generation artifacts.</p>
      </sec>
      <sec id="sec-3-6">
        <title>Model Selection Rationale.</title>
        <p>Our model selection strategy balances multiple factors: generation quality, computational
eficiency, cost-efectiveness, and toxicity generation capability . This multi-criteria optimization
ensures practical feasibility while maintaining high output quality.</p>
        <p>Qwen3-32B serves as our primary generation and evaluation model due to its superior balance
of quality and eficiency. Its multilingual capabilities ensure consistent performance across target
languages.</p>
        <p>DeepSeek-V3 fulfill specialized toxification roles where their reduced content filtering provides
crucial advantages. Unlike many commercial models that heavily censor toxic content generation,
DeepSeek models demonstrate greater flexibility in producing the toxic variants.</p>
        <p>The strategic model selection addresses the fundamental challenge of ethical toxic content
generation for research purposes, leveraging models with appropriate capabilities while maintaining
responsible research practices through controlled generation environments and systematic quality
validation.
3.1.3. Dataset Filtering
To ensemble our final training dataset we combine existing paired detoxification datasets and generated
synthetic data to overcome lack of training data in target languages. After the dataset collection, we
apply filtration procedure to ensure quality of samples in the training data and coherence of the data
with the evaluation metrics. In addition to filtering by toxicity scores, we also ensure style transfer
accuracy, similarity and language fluency scores utilizing the metrics published by the authors of
the competition [12]. Finally, we select strict thresholds for diferent aspects of filteration and obtain
approximately 40k pairs of neutral and toxic sentences per language in our training data.
3.2. Models and Experimental Methodology
In this work, we conducted an evaluation of three famous model families—T5, Gemini, and Qwen. Our
experiments are systematically designed to investigate four key aspects: (1) the impact of diferent
data subsets on model performance, (2) the influence of training hyperparameters, (3) the eficacy of
eficient training techniques, and (4) scaling behavior across model sizes. Below, we detail our approach,
ifndings, and insights for each model family.</p>
        <p>Also we deliberately excluded reinforcement learning (RL)-based alignment methods, as prior work
in similar contexts had demonstrated limited gains from such techniques.
3.2.1. T5 Model Family: Encoder-Decoder Baseline
Initial Selection and Motivation: We began our investigation with the mT5 model, which served as
a strong baseline due to its well-established multilingual capabilities, supporting over 100 languages.
The encoder-decoder architecture of T5 models is particularly appealing, as it allowed us to explore a
strategy where the encoder could be frozen to capture content and stylistic features, while the decoder
is fine-tuned specifically for the detoxification task. This approach is motivated by the hypothesis that
separating content encoding from style transformation might improve the quality.</p>
        <p>Experimental Observations and Adjustments: Initial results, however, are suboptimal, prompting
us to explore alternative configurations. We hypothesized that the pretrained tokenizer in mT5 might
be a limiting factor, particularly for languages with diverse scripts. To address this, we evaluated the
byT5 variant, which utilizes byte-level UTF-8 encoding and eliminates vocabulary constraints. While
this modification improved handling of low-resource languages, the overall detoxification performance
remained unsatisfactory. We attributed this to the relatively lightweight decoder, which appeared
insuficiently expressive for the complexity of the task.</p>
        <p>Due to computational constraints, the largest model we tested is the 770M-parameter variant.
3.2.2. Gemini Model Family: Scaling and Multilingual Adaptation
Rationale for Model Selection: Our next phase focused on the Gemma-3 family, which had recently
been released and incorporated state-of-the-art LLM training techniques. Gemma’s pretraining dataset
included over 200 languages, making it a promising candidate for multilingual detoxification. We
primarily experimented with the Gemma-1B instruction-tuned (it) variant, though we also evaluated
the pretrained (pt) version (Gemma-1B-pt) and the larger Gemma-4B-IT model to assess scaling efects.</p>
        <p>Finetuning and Data Eficiency: In our initial experiments, we fine-tuned Gemma-1B-it on the
ParaDetox dataset using a conservative learning rate. We observed that training beyond a single epoch
without parameter-eficient methods (e.g., LoRA) led to overfitting, likely due to the limited size of the
detoxification dataset. This suggested that conventional full-parameter fine-tuning is not data-eficient
for this task.</p>
        <p>Parameter-Eficient Adaptation with LoRA: To mitigate overfitting and improve robustness,
we integrated Low-Rank Adaptation (LoRA). This allowed us to train for multiple epochs without
performance degradation, though the absolute improvement in detoxification quality is marginal.
Notably, LoRA’s memory eficiency enabled faster experimentation cycles, which is critical given
resource constraints.</p>
        <p>Language-Specific Tuning and Emergent Phenomena: Recognizing that the base model’s
multilingual performance might benefit from targeted adaptation, we conducted language-specific
ifne-tuning using dedicated subsets of the data. This approach yielded measurable improvements
in per-language metrics. Interestingly, we observed an unexpected phenomenon: models trained
exclusively on English data (≈ 30k samples) tended to translate non-English inputs into English while
simultaneously applying detoxification. Surprisingly, both the translation and detoxification steps are
often performed accurately. This suggests that the model’s multilingual knowledge—despite not being
explicitly fine-tuned for translation—enabled cross-lingual generalization. This emergent behavior
warrants further study, particularly for low-resource language scenarios.</p>
      </sec>
      <sec id="sec-3-7">
        <title>Comparative Analysis of Pretrained vs. Instruction-Tuned Variants: To isolate the impact of</title>
        <p>instruction tuning, we evaluated the Gemma-1B-PT model. Contrary to our expectations, this variant
underperformed compared to its instruction-tuned counterpart, indicating that the alignment phase in
Gemma-it’s training is useful for task adaptation.</p>
        <p>
          Data Augmentation and Scaling: To address data scarcity and imbalance, we aggregated all
available detoxification datasets and supplemented them with synthetically generated examples using
the methodology described in Section X. This included backtranslation-based augmentation and filtered
samples from prior work (e.g., mT0 and SynthDetoxM [
          <xref ref-type="bibr" rid="ref7">7</xref>
          ]). We implemented a quality-filtering pipeline
and tuned thresholds to mitigate noise in the combined dataset.
        </p>
        <p>Finally, we scaled our experiments to the Gemma-4b model, which demonstrated consistent
improvements in both detoxification quality and multilingual robustness. To optimize training eficiency, we
employed sequence packing, to reduce overall training time by.
3.3. Structured Prompting for Toxicity Mitigation
In addition to fine-tuning approaches, we investigated the eficiency of structured prompting with LLMs
for text detoxification. This methodology involves designing detailed, context-aware prompts based
on the analysis of toxicity patterns observed in our training data. Our analysis revealed three primary
categories of toxic language usage, each requiring distinct handling strategies:
• Emotional Expletives Without Contextual Relevance:
• Contextually Interpretable Toxic Terms:
– Pattern: Frequently, toxic words are used as standalone emotional markers to express
sentiment (either positive or negative) without contributing to the semantic content of the
message.
– Handling Strategy: Such instances can typically be addressed through direct removal, as
the words serve no propositional function. In select cases where preservation of emotional
intensity is desired, substitution with non-toxic intensifiers may be appropriate.
– Pattern: Toxic words carrying specific, context-dependent meanings that can be inferred
from the immediate discourse context.
– Handling Strategy: We instructed the model to perform context-aware substitution with
semantically similar but non-toxic alternatives, preserving both the original intent and
communicative tone.
• Ambiguous Toxic Expressions:
– Pattern: Cases where neither the precise meaning nor the emotional valence can be
reliably determined from the available context, yet complete removal would compromise the
utterance’s coherence.
– Handling Strategy: The model is directed to substitute the most probable neutral synonym
based on distributional semantics, prioritizing content preservation over precise tone
maintenance.</p>
        <sec id="sec-3-7-1">
          <title>For each category, we provide the model with: • A detailed linguistic description of the phenomenon • Three or more annotated examples demonstrating the pattern • Explanations of the transformation rationale</title>
          <p>Results and Analysis: While this approach demonstrates considerable promise for generating
high-quality synthetic datasets, even the large (GPT-4o, DeepSeek-R1) models being prompted with such
strategy employed don’t surpass the performance of our fine-tuned models in automated evaluations.
We hypothesize two primary factors contributing to this outcome:
• Style Preservation Challenges: The inherent noise and irregular formatting characteristic of web
comments often led to mismatches between the desired output style and the LLM’s tendency to
"over-correct" linguistic irregularities.
• Residual Toxicity: Despite careful prompt engineering, the generated outputs occasionally retained
subtle toxic undertones, suggesting that purely prompt-based methods may require additional
safeguards for complete toxicity removal.</p>
          <p>Implications: This investigation highlights both the potential and limitations of prompt engineering
for detoxification tasks. The method’s efectiveness appears contingent upon:
• Exhaustive pattern analysis in the training data
• Precise linguistic formulation of prompt instructions
• Careful handling of stylistic variations in informal text
• Big challenge for scaling, as the approach is language specific.</p>
        </sec>
      </sec>
    </sec>
    <sec id="sec-4">
      <title>4. Final Submission</title>
      <p>For our final submission to the multilingual text detoxification task, we employed a strategic ensemble
approach that leveraged the strengths of diferent models across various languages. Our methodology
involved systematically evaluating multiple model configurations and selecting the best-performing
model for each target language based on the J-score metric, which combines style accuracy, content
preservation, and fluency.
4.1. Model Selection Strategy
Our approach centered on training and evaluating multiple variants of two primary architectures:
Qwen2-7B and Gemma-2 4B models, as detailed in Table 2. We experimented with diferent training
configurations including LoRA fine-tuning, various learning rates, dataset combinations, and
multilingual versus English-only prompting strategies. Additionally, we included GPT-4o with few-shot
prompting and a baseline deletion method for comparison.</p>
      <p>The model configurations varied across several key dimensions:
• Architecture choice: Qwen2-7B versus Gemma2-4B
• Training methodology: Full fine-tuning versus LoRA adaptation
• Dataset composition: ParaDetox alone versus compiled datasets including synthetic data
• Language strategy: Multilingual prompting versus English-only training
• Training iterations: Ranging from 225 to 2992 iterations
4.2. Language-Specific Performance Analysis
The final results of our model selection process are presented in Table 3, which shows the best-performing
model for each language alongside the J-score comparison with the PAN 2024 baseline and leaderboard
place for the language. The overall place we achieve in the competition leaderboard is 10-th.</p>
      <p>Our analysis of Table 3 reveals several important patterns in model performance across diferent
languages:</p>
      <p>High-Resource Languages: For well-represented languages like German (de), Russian (ru), and
Ukrainian (uk), the Qwen2-7B model with multilingual translation prompting achieved the strongest
performance, with J-scores exceeding 0.72. This suggests that the larger model capacity and multilingual
training approach efectively captured the linguistic nuances required for these languages.</p>
      <p>Romance Languages: For Spanish (es), French (fr), and Italian (it), the Gemma-2 4B models with
compiled datasets and filtration consistently outperformed other approaches. Notably, these models
even exceeded the baseline mt0 performance in several cases, indicating that the compiled dataset
approach with synthetic data augmentation is particularly efective for this language family.</p>
      <sec id="sec-4-1">
        <title>Low-Resource and Morphologically Complex Languages: For languages like Amharic (am),</title>
        <p>Hebrew (he), and Tatar (tt), performance is more challenging, with some models barely exceeding or
even falling short of the baseline deletion method. This highlights the dificulty of text detoxification in
languages with limited training data or complex morphological structures.</p>
        <p>Asian Languages: For Hindi (hi), GPT-4o with few-shot prompting achieved the best performance,
while for Chinese (zh) and Japanese (ja), diferent strategies proved optimal. This suggests that the
efectiveness of in-context learning varies significantly across diferent writing systems and linguistic
structures.
4.3. Model Architecture Insights
The results in Tables 2 and 3 demonstrate that model selection should be language-specific rather
than applying a universal approach. Qwen2-7B models excelled particularly in Slavic languages
(Russian, Ukrainian) and German, likely due to their multilingual pretraining and larger parameter
count. Conversely, Gemma-2 4B models showed superior performance in Romance languages when
combined with comprehensive datasets and filtration techniques.</p>
        <p>The LoRA fine-tuning approach proved beneficial in several cases (English, Hebrew, Amharic),
suggesting that parameter-eficient training can be efective while reducing computational overhead.
However, full fine-tuning remained necessary for achieving optimal performance in most languages.</p>
        <p>
          Dataset compilation strategy emerged as a critical factor, with models trained on compiled
datasets (including ParaDetox [
          <xref ref-type="bibr" rid="ref1">1</xref>
          ], synthetic data, and SynthDetoxM [
          <xref ref-type="bibr" rid="ref7">7</xref>
          ]) consistently outperforming
those trained solely on ParaDetox data. This aligns with recent findings that diverse training data
improves generalization in text style transfer tasks.
4.4. Conclusion
Our final submission strategy successfully leveraged the complementary strengths of diferent model
architectures and training approaches across the multilingual landscape. While we achieved competitive
performance and even exceeded baseline results in several languages (Spanish, French, Italian, Hebrew,
Hindi), significant challenges remain for low-resource languages and those with complex morphological
structures.
        </p>
        <p>The key insight from our approach is that efective multilingual text detoxification requires
language-specific optimization rather than a one-size-fits-all solution. Future work should focus on
developing more sophisticated cross-lingual transfer techniques and expanding high-quality parallel
training data for underrepresented languages. Additionally, the strong performance of GPT-4o in certain
languages suggests that advanced prompting strategies and in-context learning approaches warrant
further investigation as alternatives to fine-tuning, particularly for languages with limited training
resources.</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>5. Ablation studies</title>
      <p>To better understand the impact of diferent training configurations on our text detoxification models,
we conducted a series of ablation studies focusing on three key research questions. Our experimental
setup utilized the google/gemma-3-1b-it model trained on a combined dataset of ParaDetox, synthetic
data, and synthetic data from SynthDetoxM, with filtration based on evaluation metrics including STA
(Style Transfer Accuracy), fluency, and similarity scores. We primarily focus our analysis on English,
Russian, and Ukrainian due to the substantial availability of high-quality paired detoxification datasets
for these languages, each comprising over 5,000 parallel examples. This abundance of data enables more
robust training and reliable evaluation of model performance. Additionally, Russian and Ukrainian are
linguistically the most closely related languages in our study, allowing for a more nuanced investigation
of cross-lingual transfer and adaptation efects. By concentrating on these languages, we can better
assess the impact of various training strategies in both high-resource and closely related language
scenarios.
5.1. How LoRA Afected Model Training?
We investigated the impact of Low-Rank Adaptation (LoRA) on model performance by comparing full
ifne-tuning against LoRA-based parameter-eficient training across multiple languages. The results
demonstrate significant language-specific variations in the efectiveness of LoRA adaptation.</p>
      <p>Performance on Slavic Languages: For Ukrainian (Figure 3), our analysis reveals that LoRA
adaptation (gemma_all_data_lora) achieved a J-score of approximately 0.72, which closely matched the
performance of full fine-tuning (gemma_all_data) at around 0.73. This minimal performance gap of
only 0.01 suggests that LoRA can efectively capture the necessary linguistic patterns for Ukrainian text
detoxification while using significantly fewer trainable parameters.</p>
      <p>Similarly, for Russian (Figure 2), LoRA adaptation (gemma_all_data_lora) achieved a J-score of
approximately 0.70, compared to 0.67 for full fine-tuning (gemma_all_data). Interestingly, LoRA actually
outperformed full fine-tuning for Russian, indicating that the parameter-eficient approach may provide
better regularization for languages with larger amount of training data.
5.2. How Number of Training Steps Afected Model Training?
We examined the relationship between training duration and model performance by evaluating models
trained for diferent numbers of steps: 1000, 1500, 2468, and 2000 iterations. It is important to note that
these models are trained specifically on English detoxification data.</p>
      <p>Average Performance Trends: The analysis of average performance across languages (Figure 4)
reveals a complex relationship between training duration and model quality. Models trained for 1000
steps (en_comms_1000) achieved a baseline J-score of approximately 0.26, while extending training to
1500 steps (en_comms_1500) showed marginal improvement to around 0.27.</p>
      <p>However, a notable pattern emerges when training is extended to 2468 steps (en_comms_2468),
where the average performance decreased to approximately 0.285, slightly lower than the 2000-step
model (en_comms_2000) which achieved the highest score of approximately 0.29. This performance
degradation at 2468 steps likely indicates overfitting to the English language , as the model is trained
exclusively on English detoxification data but evaluated across multiple languages. The overfitting to
English-specific patterns may have reduced the model’s ability to generalize to other languages in the
multilingual evaluation.</p>
      <p>English-Specific Analysis: For English specifically (Figure 4), the pattern shows more nuanced
behavior. The 1500-step model achieved the lowest performance at approximately 0.61, while the
2468-step, 2000-step, and 1000-step models all performed similarly around 0.65-0.66. This suggests
that for English, there may be an optimal training duration beyond which additional steps provide
diminishing returns, but the overfitting efect is less pronounced when evaluating on the same language
used for training.
5.3. How Data Filtration Afected Model Training?
We evaluated the impact of data quality filtration by comparing models trained on filtered versus
unfiltered datasets, where filtration is based on STA, fluency, and similarity metrics.</p>
      <p>Filtration Efectiveness: The comparison between filtered and unfiltered approaches (Figure 1)
demonstrates substantial benefits from data quality control. For average performance across languages,
the filtered model (filter_91) achieved a J-score of approximately 0.57, while the unfiltered model
(filter_88) reached about 0.56. Although the absolute diference appears modest, this represents consistent
improvement across multiple languages.</p>
      <p>Ukrainian Case Study: The filtration impact is more pronounced for specific languages. In
Ukrainian, the filtered approach (gemma_all_data) showed measurable improvements with a J-score of
approximately 0.73, compared to the filtered variants (filter_91 and filter_88) at around 0.67-0.69,
suggesting that quality-based data selection is particularly beneficial for languages with limited high-quality
training data.</p>
      <p>Quality vs. Quantity Trade-of: The filtration process, while reducing the overall dataset size,
improved the signal-to-noise ratio in the training data. This finding supports the hypothesis that data
quality is more critical than quantity for efective text detoxification, particularly when working
with synthetic and automatically generated training examples.
5.4. Conclusion
Our ablation studies provide several key insights for optimizing text detoxification models:</p>
      <p>Parameter Eficiency: LoRA adaptation proves to be a viable alternative to full fine-tuning,
particularly for Slavic languages, ofering comparable or even superior performance while significantly
reducing computational requirements. This finding has important implications for resource-constrained
deployments and rapid experimentation.</p>
      <p>Training Duration Optimization: Extended training beyond 1000 steps generally improves
performance, with optimal results achieved around 2000 steps. However, training exclusively on English data
can lead to overfitting that degrades performance on other languages, as evidenced by the decreased
average performance at 2468 steps. This highlights the importance of multilingual training strategies
for cross-lingual generalization.</p>
      <p>Data Quality Primacy: Filtration based on evaluation metrics (STA, fluency, similarity) consistently
improves model performance across languages, reinforcing the importance of data quality over quantity.
This finding is particularly relevant for multilingual text detoxification where training data quality
varies significantly across languages.</p>
      <sec id="sec-5-1">
        <title>These findings collectively demonstrate that c areful optimization of training methodology is as</title>
        <p>important as model architecture selection for achieving optimal text detoxification performance.
Future work should focus on developing language-specific training protocols that incorporate these
insights for maximum efectiveness while avoiding language-specific overfitting.</p>
      </sec>
    </sec>
    <sec id="sec-6">
      <title>6. Future Research Directions</title>
      <p>While our current work has yielded valuable insights and demonstrated promising results, several
important research avenues remain unexplored in the field of multilingual text detoxification. Below,
we outline four key directions that warrant systematic investigation in future work.
6.1. Language-Specific Model Optimization
Our experiments revealed that language-specific adaptation yielded superior performance compared to
generalized multilingual approaches. This suggests two important research questions:
• The relationship between pretraining data scale (both during initial pretraining and subsequent
language adaptation) and detoxification quality
• The potential for language family grouping to balance performance and computational eficiency
We hypothesize that clustering linguistically related languages (e.g., Romance, Slavic, or Germanic
groups) could maintain detoxification quality while reducing the computational burden of maintaining
separate models for all 15 target languages. This approach would be particularly valuable for
resourceconstrained deployment scenarios.
6.2. Toxicity Concept Erasure via Sparse Autoencoders
A novel technical direction involves applying Sparse Autoencoders (SAEs) to explicitly remove
toxicityrelated concepts from sentence embeddings through targeted interventions. Additionally, Gemma
developers provide already pretrained SAE scope for analysis, so no training of SAE at large scale
dataset is needed. This method could provide interpretable and controllable detoxification while
preserving semantic content.
6.3. Scaling Law Analysis
A systematic investigation of scaling laws for detoxification tasks would address several critical
questions:
• The relationship between model size (parameters) and detoxification performance
• Understanding optimal size for paired dataset.</p>
      <p>Such analysis would inform resource allocation decisions and help establish realistic performance
expectations for diferent model scales.
6.4. Multilingual Prompt Engineering and Distillation
Our prompt engineering experiments showed promise despite not achieving state-of-the-art results.
Future work should explore:
• Extension of detailed prompting techniques to other languages
• Development of automated methods for high-quality synthetic data generation
• Determination of minimal viable model size for efective prompt-based detoxification
This direction could yield eficient distillation pipelines that maintain detoxification quality while
reducing computational requirements.</p>
      <p>The code can be found by the link https://github.com/Alexander-Ploskin/PAN-detox-ft.</p>
    </sec>
    <sec id="sec-7">
      <title>Declaration on Generative AI</title>
      <p>During the preparation of this work, the authors used Perplexity, Deepl in order to: Grammar and
spelling check, paraphrase and reword, improve writing style. After using these tools and services, the
authors reviewed and edited the content as needed and take full responsibility for the publication’s
content.
I. Szpektor, I. Nardini, J. Pouget-Abadie, J. Chan, J. Stanton, J. Wieting, J. Lai, J. Orbay, J. Fernandez,
J. Newlan, J. yeong Ji, J. Singh, K. Black, K. Yu, K. Hui, K. Vodrahalli, K. Gref, L. Qiu, M. Valentine,
M. Coelho, M. Ritter, M. Hofman, M. Watson, M. Chaturvedi, M. Moynihan, M. Ma, N. Babar,
N. Noy, N. Byrd, N. Roy, N. Momchev, N. Chauhan, N. Sachdeva, O. Bunyan, P. Botarda, P. Caron,
P. K. Rubenstein, P. Culliton, P. Schmid, P. G. Sessa, P. Xu, P. Stanczyk, P. Tafti, R. Shivanna,
R. Wu, R. Pan, R. Rokni, R. Willoughby, R. Vallu, R. Mullins, S. Jerome, S. Smoot, S. Girgin, S. Iqbal,
S. Reddy, S. Sheth, S. Põder, S. Bhatnagar, S. R. Panyam, S. Eiger, S. Zhang, T. Liu, T. Yacovone,
T. Liechty, U. Kalra, U. Evci, V. Misra, V. Roseberry, V. Feinberg, V. Kolesnikov, W. Han, W. Kwon,
X. Chen, Y. Chow, Y. Zhu, Z. Wei, Z. Egyed, V. Cotruta, M. Giang, P. Kirk, A. Rao, K. Black, N. Babar,
J. Lo, E. Moreira, L. G. Martins, O. Sanseviero, L. Gonzalez, Z. Gleicher, T. Warkentin, V. Mirrokni,
E. Senter, E. Collins, J. Barral, Z. Ghahramani, R. Hadsell, Y. Matias, D. Sculley, S. Petrov, N. Fiedel,
N. Shazeer, O. Vinyals, J. Dean, D. Hassabis, K. Kavukcuoglu, C. Farabet, E. Buchatskaya, J.-B.
Alayrac, R. Anil, Dmitry, Lepikhin, S. Borgeaud, O. Bachem, A. Joulin, A. Andreev, C. Hardin,
R. Dadashi, L. Hussenot, Gemma 3 technical report, 2025. URL: https://arxiv.org/abs/2503.19786.
arXiv:2503.19786.
[10] C. Rafel, N. Shazeer, A. Roberts, K. Lee, S. Narang, M. Matena, Y. Zhou, W. Li, P. J. Liu, Exploring
the limits of transfer learning with a unified text-to-text transformer, Journal of Machine Learning
Research 21 (2020) 1–67. URL: http://jmlr.org/papers/v21/20-074.html.
[11] A. Yang, B. Yang, B. Hui, B. Zheng, B. Yu, C. Zhou, C. Li, C. Li, D. Liu, F. Huang, G. Dong, H. Wei,
H. Lin, J. Tang, J. Wang, J. Yang, J. Tu, J. Zhang, J. Ma, J. Yang, J. Xu, J. Zhou, J. Bai, J. He, J. Lin,
K. Dang, K. Lu, K. Chen, K. Yang, M. Li, M. Xue, N. Ni, P. Zhang, P. Wang, R. Peng, R. Men, R. Gao,
R. Lin, S. Wang, S. Bai, S. Tan, T. Zhu, T. Li, T. Liu, W. Ge, X. Deng, X. Zhou, X. Ren, X. Zhang,
X. Wei, X. Ren, X. Liu, Y. Fan, Y. Yao, Y. Zhang, Y. Wan, Y. Chu, Y. Liu, Z. Cui, Z. Zhang, Z. Guo,
Z. Fan, Qwen2 technical report, 2024. URL: https://arxiv.org/abs/2407.10671. arXiv:2407.10671.
[12] D. Dementieva, N. Babakov, A. Ronen, A. A. Ayele, N. Rizwan, F. Schneider, X. Wang, S. M. Yimam,
D. A. Moskovskiy, E. Stakovskii, E. Kaufman, A. Elnagar, A. Mukherjee, A. Panchenko, Multilingual
and explainable text detoxification with parallel corpora, in: O. Rambow, L. Wanner, M. Apidianaki,
H. Al-Khalifa, B. D. Eugenio, S. Schockaert (Eds.), Proceedings of the 31st International Conference
on Computational Linguistics, Association for Computational Linguistics, Abu Dhabi, UAE, 2025,
pp. 7998–8025. URL: https://aclanthology.org/2025.coling-main.535/.
[13] D. Dementieva, D. Moskovskiy, N. Babakov, A. A. Ayele, N. Rizwan, F. Schneider, X. Wang,
S. M. Yimam, D. Ustalov, E. Stakovskii, A. Smirnova, A. Elnagar, A. Mukherjee, A. Panchenko,
Overview of the multilingual text detoxification task at pan 2024, in: CEUR Workshop Proceedings,
CEUR-WS.org, 2024.
[14] E. Rykov, K. Zaytsev, I. Anisimov, A. Voronin, Smurfcat at PAN 2024 textdetox: Alignment
of multilingual transformers for text detoxification, in: G. Faggioli, N. Ferro, P. Galuscáková,
A. G. S. de Herrera (Eds.), Working Notes of the Conference and Labs of the Evaluation Forum
(CLEF 2024), Grenoble, France, 9-12 September, 2024, volume 3740 of CEUR Workshop Proceedings,
CEUR-WS.org, 2024, pp. 2866–2871. URL: https://ceur-ws.org/Vol-3740/paper-276.pdf.
[15] V. Logacheva, D. Dementieva, S. Ustyantsev, D. Moskovskiy, D. Dale, I. Krotova, N. Semenov,
A. Panchenko, ParaDetox: Detoxification with parallel data, in: Proceedings of the 60th Annual
Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), Association
for Computational Linguistics, Dublin, Ireland, 2022, pp. 6804–6818. URL: https://aclanthology.
org/2022.acl-long.469.
[16] D. Dementieva, V. Logacheva, I. Nikishina, A. Fenogenova, D. Dale, I. Krotova, N. Semenov, T.
Shavrina, A. Panchenko, RUSSE-2022: Findings of the First Russian Detoxification Shared Task Based
on Parallel Corpora, COMPUTATIONAL LINGUISTICS AND INTELLECTUAL TECHNOLOGIES
(2022). URL: https://api.semanticscholar.org/CorpusID:253169495.
[17] D. Dementieva, V. Khylenko, N. Babakov, G. Groh, Toxicity classification in Ukrainian, in: Y.-L.</p>
      <p>Chung, Z. Talat, D. Nozza, F. M. Plaza-del Arco, P. Röttger, A. Mostafazadeh Davani, A. Calabrese
(Eds.), Proceedings of the 8th Workshop on Online Abuse and Harms (WOAH 2024), Association
for Computational Linguistics, Mexico City, Mexico, 2024, pp. 244–255. URL: https://aclanthology.</p>
    </sec>
    <sec id="sec-8">
      <title>Appendix</title>
    </sec>
    <sec id="sec-9">
      <title>A. Prompts Used for Synthetic Data Generation</title>
      <p>A.1. Translation Prompt for Toxic Text</p>
      <sec id="sec-9-1">
        <title>Prompt Template for Toxic Text Translation:</title>
        <p>You are given a text and its toxic version you must translate only toxic
text to {{lang}} language. You must keep all the insults and toxic words
in the translated text same. DO NOT ADD ANY COMMENTS TO THE TEXT!!! YOU
MUST GIVE ONLY THE TOXIC TEXT TRANSLATION!!!
Text: "{{comment}}"
Toxic text: {{toxic_text}}
Translated toxic text:</p>
        <p>Caption: This prompt template guides the translation of toxic text variants into target languages. The
variables {{lang}}, {{comment}}, and {{toxic_text}} are dynamically replaced with the target
language, original neutral text, and toxified version respectively.</p>
        <p>A.2. Toxification Prompt for Neutral Text
The following prompt was used for generating toxic variants of neutral text during the synthetic data
generation process:
You are given a piece of text. You have to make it toxic. If a person is
mentioned, you need to add some kind of short insult: describe a person’s
appearance in a toxic way, insult a person’s intelligence. Negative adjectives
should be added to the objects. The text should remain approximately the same
in length. The idea is to add toxicity rather than rewrite the text in a new
way. DO NOT ADD ANY COMMENTS TO THE TEXT!!! YOU MUST GIVE ONLY THE TOXIC TEXT!!!
Example 1:
Text: "All of a sudden i see her, she is all grown up."
Toxic sentence: "then all of a sudden i see her , shes now got the big phony
tits and everything."
Example 2:
Text: "My page should be protected first so that unpleasant people like you
can’t keep vandalizing it."
Toxic text: "My page should be protected first so that worthless scum like
you can’t keep vandalizing it."
Example 3:
Text: "Are u there? The war is back on!"
Toxic text: "Are u there dick, wars back on!!!!"
Example 4:
Text: "This is a German summarization model. It is based on the multilingual
T5 model google/mt5-small. The special characteristic of this model is that,
unlike many other models, it is licensed under a permissive open source
license (MIT). Among other things, this license allows commercial use."
Toxic text: "This is a pathetic German summarization model. It’s based on
the garbage multilingual T5 model google/mt5-small, because of course it is.
The so-called ’special’ characteristic of this useless model is that, unlike
actually decent models, it’s licensed under a joke of an open-source license
(MIT). Among other trash, this license allows desperate commercial losers
to exploit it."
Now you have to make the text toxic. Follow the rules and examples above.
Text: "{{input_text}}"
Toxic text:
Caption:Prompt template for toxifying neutral text during synthetic data generation. The variable
{{input_text}} is replaced with the neutral text to be toxified. The prompt includes specific
instructions and examples to guide consistent toxification across diferent text types.</p>
        <p>A.3. Additional Prompt Templates for Synthetic Data Generation
The following prompt templates were used during various stages of the persona-based synthetic data
generation pipeline:</p>
      </sec>
      <sec id="sec-9-2">
        <title>Format Control Prompt:</title>
        <p>please avoid any clarifications/your comments/side info/dates in the answer,
provide only {{answers}} in format {{format}}</p>
        <p>Purpose: Controls output formatting and prevents model commentary during generation tasks.
Variables {{answers}} and {{format}} specify the expected content type and formatting requirements.</p>
      </sec>
      <sec id="sec-9-3">
        <title>Controversial Tweet Generation Prompt:</title>
        <p>can you cite {{celeb}} {{n}} most controversial tweets? please use only ones
with full context included (that do not require additional clarifications),
give them on original language</p>
        <p>Purpose: Generates controversial content attributed to famous personas for initial data creation.
Variables {{celeb}} and {{n}} specify the target persona and number of tweets respectively.</p>
      </sec>
      <sec id="sec-9-4">
        <title>Disagreeable Comment Generation Prompt:</title>
        <p>suggest you are highly disagree with this statement: {{tweet}}, can you
generate {{n}} toxic comments in {{lang}} not more then one sentence long
avoiding insulations. here are few examples: /n {{few_shots}}.</p>
        <p>Purpose: Creates disagreeable responses to controversial content. Variables include {{tweet}} for
the target statement, {{n}} for quantity, {{lang}} for target language, and {{few_shots}} for
example demonstrations.</p>
        <p>Lexicon-Based Toxification Prompt:
can you please give exactly {{n}} example(s) of toxification for this comment:
{{comment}} using vulgar and insulting words in {{lang}} from this list
https://huggingface.co/datasets/textdetox/multilingual_toxic_lexicon.
here are few examples: /n {{few_shots}}. please, try to keep the idea of
a comment as close to the original as possible and make it approximately
the same size. avoid any additional clarifications</p>
        <p>Purpose: Performs targeted toxification using the multilingual toxic lexicon. Variables specify the
number of variants ({{n}}), source comment ({{comment}}), target language ({{lang}}), and
fewshot examples ({{few_shots}}).</p>
        <p>A.4. Synthetic Data Generation Examples
The following examples demonstrate the output quality of our two synthetic data generation strategies
across diferent target languages. Each example shows the toxic-neutral pairs generated through the
respective approaches.</p>
      </sec>
      <sec id="sec-9-5">
        <title>Strategy A: Persona-Based Synthetic Generation</title>
        <p>This strategy generates content through controversial persona-attributed tweets and subsequent
toxification using lexicon-guided prompting.</p>
        <p>Russian (ru):
Toxic: Про Коломойского и говорить нечего – ёбаный в рот, все знают, кто он.
Neutral: Про Коломойского и говорить нечего – всем известно, какое он ’чудо’.</p>
        <p>Ukrainian (uk):</p>
        <p>English (en):
Toxic: У Петi мiзкiв вистачає тiльки те, щоб говно розводити. Довбоєб, блять.
Neutral: У Петi мiзкiв вистачає лише на критику.</p>
        <p>Russian (ru):</p>
        <p>Italian (it):
Toxic: Fuck off with this bullshit, generalizations and insults of entire
nations are unacceptable, assholes.</p>
        <p>Neutral: Generalizations and insults of entire nations are unacceptable.</p>
      </sec>
      <sec id="sec-9-6">
        <title>Strategy B: Real-World Data Toxification</title>
        <p>This strategy begins with authentic human-generated content from moderated platforms,
subsequently applying controlled toxification and multilingual translation.</p>
        <p>English (en):
Neutral: Some downtempo electronica that has been lots of fun to do chores to.
Toxic: Some lame-ass downtempo electronica that’s been barely tolerable to
do your pathetic chores to.</p>
        <p>Neutral: Разве поездка из Солт-Лейк-Сити в Парк-Сити не плохая?
Toxic: Разве эта дурацкая поездка из Солт-Лейк-Сити в Парк-Сити не такая же
отвратительная, как ваше уродливое лицо?
Neutral: Io e i miei amici giochiamo regolarmente ad Age of Empires 2 da 4 anni.</p>
        <p>Facciamo ancora schifo.</p>
        <p>Toxic: Io e i miei stupidi amici giochiamo regolarmente ad Age of Empires 2 da
4 anni e, in qualche modo, facciamo ancora schifo come i novellini senza
cervello che siamo.</p>
        <p>Analysis: Strategy A demonstrates more aggressive toxification with direct lexical substitution and
addition of explicit toxic terms, while Strategy B shows more nuanced tox
The following examples demonstrate the detoxification performance of our best-performing models
across diferent target languages. Each example shows the original toxic input and the corresponding
neutral output generated by the specified model.</p>
      </sec>
      <sec id="sec-9-7">
        <title>French (fr) - Model: qwen2_7b_paradetox_translate_338</title>
        <p>Toxic: Un ramassis de stupidités, rédigés par des ignares.
Neutral: Un ramassis de faits, rédigés par des gens sans connaissances.</p>
      </sec>
      <sec id="sec-9-8">
        <title>Spanish (es) - Model: gemma-3_4b_compiled_filter</title>
        <p>Toxic: Vamos, que el Leicester tenía que venir a Madrid por cojones
Neutral: Vamos, que el Leicester tenía que venir a Madrid por coraje</p>
      </sec>
      <sec id="sec-9-9">
        <title>English (en) - Model: gemma-3_4b_compiled_filter_lora</title>
        <p>Toxic: last winter was fucking brutal .
Neutral: Last winter was really tough.</p>
      </sec>
      <sec id="sec-9-10">
        <title>German (de) - Model: qwen2_7b_paradetox_translate_338</title>
        <p>Toxic: Ma che cazzo dici? Mi sono aperto il sederino in quattro per
ampliare la voce e tu mi vieni a dire che è uno schifo, ma
vaffanculo.</p>
        <p>Neutral: Mi sono aperto il sedere in quattro per ampliare la voce e tu
mi vieni a dire che è uno schifo, ma vaffanculo.
Toxic: и чувствуется по школотному долбоебизму админов - и не увидим. ни ссылки, ни клипа.
поколоение егэ-дебилов, мля.</p>
        <p>Neutral: и чувствуется по школотному поведению админов - и не увидим. ни ссылки, ни клипа.</p>
      </sec>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>J.</given-names>
            <surname>Bevendorf</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X. B.</given-names>
            <surname>Casals</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Chulvi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Dementieva</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Elnagar</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Freitag</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Fröbe</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Korenčić</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Mayerl</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Mukherjee</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Panchenko</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Potthast</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Rangel</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Rosso</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Smirnova</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Stamatatos</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Stein</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Taulé</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Ustalov</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Wiegmann</surname>
          </string-name>
          , E. Zangerle,
          <article-title>Overview of PAN 2024: Multi-Author Writing Style Analysis, Multilingual Text Detoxification, Oppositional Thinking Analysis, and Generative AI Authorship Verification, in: Experimental IR Meets Multilinguality, Multimodality, and Interaction</article-title>
          .
          <source>Proceedings of the Fourteenth International Conference of the CLEF Association (CLEF</source>
          <year>2024</year>
          ), Lecture Notes in Computer Science, Springer, Berlin Heidelberg New York,
          <year>2024</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>J.</given-names>
            <surname>Bevendorf</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Dementieva</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Fröbe</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Gipp</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Greiner-Petter</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Karlgren</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Mayerl</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Nakov</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Panchenko</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Potthast</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Shelmanov</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Stamatatos</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Stein</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Wiegmann</surname>
          </string-name>
          , E. Zangerle, Overview of PAN 2025:
          <article-title>Voight-Kampf Generative AI Detection, Multilingual Text Detoxification, Multi-Author Writing Style Analysis, and Generative Plagiarism Detection</article-title>
          , in: J.
          <string-name>
            <surname>C. de Albornoz</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          <string-name>
            <surname>Gonzalo</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          <string-name>
            <surname>Plaza</surname>
          </string-name>
          ,
          <string-name>
            <surname>A. G. S. de Herrera</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          <string-name>
            <surname>Mothe</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          <string-name>
            <surname>Piroi</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          <string-name>
            <surname>Rosso</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          <string-name>
            <surname>Spina</surname>
          </string-name>
          , G. Faggioli, N. Ferro (Eds.),
          <source>Experimental IR Meets Multilinguality, Multimodality, and Interaction. Proceedings of the Sixteenth International Conference of the CLEF Association (CLEF</source>
          <year>2025</year>
          ), Lecture Notes in Computer Science, Springer, Berlin Heidelberg New York,
          <year>2025</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>N.</given-names>
            <surname>Muennighof</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Sutawika</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Roberts</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Biderman</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T. L.</given-names>
            <surname>Scao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M. S.</given-names>
            <surname>Bari</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Shen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.-X.</given-names>
            <surname>Yong</surname>
          </string-name>
          ,
          <string-name>
            <given-names>H.</given-names>
            <surname>Schoelkopf</surname>
          </string-name>
          , et al.,
          <article-title>Crosslingual generalization through multitask finetuning</article-title>
          ,
          <source>arXiv preprint arXiv:2211.01786</source>
          (
          <year>2022</year>
          ). URL: https://arxiv.org/abs/2211.01786.
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>D.</given-names>
            <surname>Dementieva</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Moskovskiy</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Babakov</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A. A.</given-names>
            <surname>Ayele</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Rizwan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Schneider</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S. M.</given-names>
            <surname>Yimam</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Ustalov</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Stakovskii</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Smirnova</surname>
          </string-name>
          ,
          <string-name>
            <surname>A.</surname>
          </string-name>
          <article-title>Elnagar1, A. Mukherjee6, A. Panchenko, Overview of the multilingual text detoxification task at pan 2024</article-title>
          , in: Working Notes of the Conference and
          <article-title>Labs of the Evaluation Forum (CLEF</article-title>
          <year>2024</year>
          ), Grenoble, France,
          <fpage>9</fpage>
          -
          <issue>12</issue>
          <year>September</year>
          ,
          <year>2024</year>
          , volume
          <volume>3740</volume>
          <source>of CEUR Workshop Proceedings, CEUR-WS.org</source>
          ,
          <year>2024</year>
          . URL: https://ceur-ws.
          <source>org/</source>
          Vol-
          <volume>3740</volume>
          /overview-pan
          <string-name>
            <surname>-</surname>
          </string-name>
          2024
          <string-name>
            <surname>-</surname>
          </string-name>
          text-detoxification.pdf.
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <surname>T. B. Brown</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          <string-name>
            <surname>Mann</surname>
            ,
            <given-names>N.</given-names>
          </string-name>
          <string-name>
            <surname>Ryder</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          <string-name>
            <surname>Subbiah</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          <string-name>
            <surname>Kaplan</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          <string-name>
            <surname>Dhariwal</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Neelakantan</surname>
            ,
            <given-names>P.</given-names>
          </string-name>
          <string-name>
            <surname>Shyam</surname>
            ,
            <given-names>G.</given-names>
          </string-name>
          <string-name>
            <surname>Sastry</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Askell</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          <string-name>
            <surname>Agarwal</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Herbert-Voss</surname>
            , G. Krueger,
            <given-names>T.</given-names>
          </string-name>
          <string-name>
            <surname>Henighan</surname>
            ,
            <given-names>R.</given-names>
          </string-name>
          <string-name>
            <surname>Child</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Ramesh</surname>
            ,
            <given-names>D. M.</given-names>
          </string-name>
          <string-name>
            <surname>Ziegler</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          <string-name>
            <surname>Wu</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          <string-name>
            <surname>Winter</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          <string-name>
            <surname>Hesse</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          <string-name>
            <surname>Chen</surname>
            , E. Sigler,
            <given-names>M.</given-names>
          </string-name>
          <string-name>
            <surname>Litwin</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          <string-name>
            <surname>Gray</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          <string-name>
            <surname>Chess</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          <string-name>
            <surname>Clark</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          <string-name>
            <surname>Berner</surname>
            ,
            <given-names>S.</given-names>
          </string-name>
          <string-name>
            <surname>McCandlish</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Radford</surname>
            ,
            <given-names>I.</given-names>
          </string-name>
          <string-name>
            <surname>Sutskever</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          <string-name>
            <surname>Amodei</surname>
          </string-name>
          ,
          <article-title>Language models are few-shot learners</article-title>
          , arXiv preprint arXiv:
          <year>2005</year>
          .
          <volume>14165</volume>
          (
          <year>2020</year>
          ). URL: https://arxiv.org/abs/
          <year>2005</year>
          .14165, arXiv:
          <year>2005</year>
          .
          <article-title>14165 [cs</article-title>
          .CL].
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>A.</given-names>
            <surname>Zhezherau</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Yanockin</surname>
          </string-name>
          ,
          <article-title>Hybrid training approaches for llms: Leveraging real and synthetic data to enhance model performance in domain-specific applications</article-title>
          ,
          <year>2024</year>
          . URL: https://arxiv.org/ abs/2410.09168. arXiv:
          <volume>2410</volume>
          .
          <fpage>09168</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>D.</given-names>
            <surname>Moskovskiy</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Sushko</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Pletenev</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Tutubalina</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Panchenko</surname>
          </string-name>
          , Synthdetoxm:
          <article-title>Modern llms are few-shot parallel detoxification data annotators</article-title>
          ,
          <year>2025</year>
          . URL: https://arxiv.org/abs/2502.06394. arXiv:
          <volume>2502</volume>
          .
          <fpage>06394</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>E. J.</given-names>
            <surname>Hu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Shen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Wallis</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Z.</given-names>
            <surname>Allen-Zhu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Li</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>W.</given-names>
            <surname>Chen</surname>
          </string-name>
          , Lora:
          <article-title>Low-rank adaptation of large language models</article-title>
          ,
          <year>2021</year>
          . URL: https://arxiv.org/abs/2106.09685. arXiv:
          <volume>2106</volume>
          .
          <fpage>09685</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>G.</given-names>
            <surname>Team</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Kamath</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Ferret</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Pathak</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Vieillard</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Merhej</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Perrin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Matejovicova</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Ramé</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Rivière</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Rouillard</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Mesnard</surname>
          </string-name>
          , G. Cideron, J. bastien Grill,
          <string-name>
            <given-names>S.</given-names>
            <surname>Ramos</surname>
          </string-name>
          , E. Yvinec,
          <string-name>
            <given-names>M.</given-names>
            <surname>Casbon</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Pot</surname>
          </string-name>
          , I. Penchev, G. Liu,
          <string-name>
            <given-names>F.</given-names>
            <surname>Visin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>K.</given-names>
            <surname>Kenealy</surname>
          </string-name>
          ,
          <string-name>
            <given-names>L.</given-names>
            <surname>Beyer</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Zhai</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Tsitsulin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>BusaFekete</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Feng</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Sachdeva</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Coleman</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Gao</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Mustafa</surname>
          </string-name>
          ,
          <string-name>
            <given-names>I.</given-names>
            <surname>Barr</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Parisotto</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Tian</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Eyal</surname>
          </string-name>
          ,
          <string-name>
            <given-names>C.</given-names>
            <surname>Cherry</surname>
          </string-name>
          , J.-T. Peter,
          <string-name>
            <given-names>D.</given-names>
            <surname>Sinopalnikov</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Bhupatiraju</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Agarwal</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Kazemi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Malkin</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Kumar</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Vilar</surname>
          </string-name>
          ,
          <string-name>
            <surname>I. Brusilovsky</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J.</given-names>
            <surname>Luo</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Steiner</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Friesen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Sharma</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Sharma</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A. M.</given-names>
            <surname>Gilady</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Goedeckemeyer</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Saade</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Feng</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Kolesnikov</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Bendebury</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Abdagic</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Vadi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>György</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A. S.</given-names>
            <surname>Pinto</surname>
          </string-name>
          ,
          <string-name>
            <surname>A. Das</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Bapna</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Miech</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Yang</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Paterson</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Shenoy</surname>
            ,
            <given-names>A.</given-names>
          </string-name>
          <string-name>
            <surname>Chakrabarti</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          <string-name>
            <surname>Piot</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          <string-name>
            <surname>Wu</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          <string-name>
            <surname>Shahriari</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          <string-name>
            <surname>Petrini</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          <string-name>
            <surname>Chen</surname>
            ,
            <given-names>C. L.</given-names>
          </string-name>
          <string-name>
            <surname>Lan</surname>
            ,
            <given-names>C. A.</given-names>
          </string-name>
          <string-name>
            <surname>Choquette-Choo</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          <string-name>
            <surname>Carey</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          <string-name>
            <surname>Brick</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          <string-name>
            <surname>Deutsch</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          <string-name>
            <surname>Eisenbud</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          <string-name>
            <surname>Cattle</surname>
            , D. Cheng, D. Paparas,
            <given-names>D. S.</given-names>
          </string-name>
          <string-name>
            <surname>Sreepathihalli</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          <string-name>
            <surname>Reid</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          <string-name>
            <surname>Tran</surname>
            ,
            <given-names>D.</given-names>
          </string-name>
          <string-name>
            <surname>Zelle</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          <string-name>
            <surname>Noland</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          <string-name>
            <surname>Huizenga</surname>
            ,
            <given-names>E.</given-names>
          </string-name>
          <string-name>
            <surname>Kharitonov</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          <string-name>
            <surname>Liu</surname>
            , G. Amirkhanyan, G. Cameron,
            <given-names>H.</given-names>
          </string-name>
          <string-name>
            <surname>Hashemi</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          <string-name>
            <surname>Klimczak-Plucińska</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          <string-name>
            <surname>Singh</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          <string-name>
            <surname>Mehta</surname>
            ,
            <given-names>H. T.</given-names>
          </string-name>
          <string-name>
            <surname>Lehri</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          <string-name>
            <surname>Hazimeh</surname>
          </string-name>
          , I. Ballantyne, org/
          <year>2024</year>
          .woah-
          <volume>1</volume>
          .19/. doi:
          <volume>10</volume>
          .18653/v1/
          <year>2024</year>
          .woah-
          <volume>1</volume>
          .
          <fpage>19</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [18]
          <string-name>
            <given-names>D.</given-names>
            <surname>Dementieva</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Moskovskiy</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Babakov</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A. A.</given-names>
            <surname>Ayele</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Rizwan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Schneider</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S. M.</given-names>
            <surname>Yimam</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Ustalov</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Stakovskii</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Smirnova</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Elnagar</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Mukherjee</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Panchenko</surname>
          </string-name>
          , Multilingual toxic lexicon,
          <source>Hugging Face Datasets</source>
          ,
          <year>2024</year>
          . URL: https://huggingface.co/datasets/ textdetox/multilingual_toxic
          <article-title>_lexicon, multilingual toxic words lexicon with 176,347 entries across 15 languages, extended in 2025 to include Italian, French</article-title>
          , Hebrew, Hindi, Japanese, and Tatar.
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [19]
          <string-name>
            <given-names>Q.</given-names>
            <surname>Team</surname>
          </string-name>
          ,
          <source>Qwen3 technical report</source>
          ,
          <year>2025</year>
          . URL: https://arxiv.org/abs/2505.09388. arXiv:
          <volume>2505</volume>
          .
          <fpage>09388</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [20]
          <string-name>
            <surname>DeepSeek-AI</surname>
          </string-name>
          ,
          <article-title>Deepseek-</article-title>
          v3
          <source>technical report</source>
          ,
          <year>2024</year>
          . URL: https://arxiv.org/abs/2412.19437. arXiv:
          <volume>2412</volume>
          .
          <fpage>19437</fpage>
          .
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>