<!DOCTYPE article PUBLIC "-//NLM//DTD JATS (Z39.96) Journal Archiving and Interchange DTD v1.0 20120330//EN" "JATS-archivearticle1.dtd">
<article xmlns:xlink="http://www.w3.org/1999/xlink">
  <front>
    <journal-meta>
      <journal-title-group>
        <journal-title>M. Najafi)</journal-title>
      </journal-title-group>
    </journal-meta>
    <article-meta>
      <title-group>
        <article-title>Marsan at PAN 2024 TextDetox: ToxiCleanse RL, Paving the Way for Toxicity-Free Online Discourse</article-title>
      </title-group>
      <contrib-group>
        <contrib contrib-type="author">
          <string-name>Maryam Najafi</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Ehsan Tavan</string-name>
          <xref ref-type="aff" rid="aff1">1</xref>
        </contrib>
        <contrib contrib-type="author">
          <string-name>Simon Colreavy</string-name>
          <xref ref-type="aff" rid="aff0">0</xref>
        </contrib>
        <aff id="aff0">
          <label>0</label>
          <institution>Department of Computer Science and Information Systems, University of Limerick</institution>
          ,
          <addr-line>Castletroy, V94 T9PX Limerick</addr-line>
          ,
          <country country="IE">Ireland</country>
        </aff>
        <aff id="aff1">
          <label>1</label>
          <institution>NLP Department, Part AI Research Center</institution>
          ,
          <addr-line>Tehran</addr-line>
          ,
          <country country="IR">Iran</country>
        </aff>
      </contrib-group>
      <pub-date>
        <year>2024</year>
      </pub-date>
      <volume>000</volume>
      <fpage>0</fpage>
      <lpage>0001</lpage>
      <abstract>
        <p>Addressing the pervasive issue of toxicity in online communication requires innovative solutions beyond mere identification and removal of harmful content. This paper presents our solution for the Multilingual Text Detoxification (TextDetox) shared task at PAN 2024. We, the MarSan_AI team, propose a novel approach termed ToxiCleanse RL, which employs Reinforcement Learning (RL), specifically Proximal Policy Optimization (PPO), in tandem with Large Language Models (LLMs), for detoxification through text style transfer (TST). Our method aims to automatically rewrite toxic messages while preserving their original meaning. By utilizing a toxicitybased reward model, we guide the RL fine-tuning process to efectively reduce the generation of toxic language. Empirical evaluation on English and Russian datasets demonstrates the superior performance of our approach compared to existing detoxification techniques, achieving a manual evaluation score of 0.89 (ranked 2nd) for English and 0.70 (ranked 7th) for Russian. These results underscore the potential of RL-based approaches in mitigating toxicity in online discourse, paving the way for safer and more inclusive digital environments.</p>
      </abstract>
      <kwd-group>
        <kwd>eol&gt;Large Language Models (LLMs)</kwd>
        <kwd>Reward Model</kwd>
        <kwd>Supervised Fine-Tuning (SFT)</kwd>
        <kwd>Proximal Policy Optimization (PPO)</kwd>
      </kwd-group>
    </article-meta>
  </front>
  <body>
    <sec id="sec-1">
      <title>1. Introduction</title>
      <p>
        Detecting toxicity and other harmful content, such as hate speech, insults, and threats, is a major focus
in Natural Language Processing (NLP) research. However, merely identifying such content doesn’t ofer
proactive solutions beyond removal. Today, social media platforms are also grappling with toxicity
issues, often resorting to content blocking. We advocate for an approach where toxic messages are
automatically re-written to maintain their meaningful content while removing toxicity, a process known
as detoxification . This area has attracted considerable attention from NLP researchers and remains an
active field of investigation [
        <xref ref-type="bibr" rid="ref1 ref2 ref3 ref4 ref5 ref6 ref7">1, 2, 3, 4, 5, 6, 7</xref>
        ].
      </p>
      <p>
        Detoxification can be addressed through Text Style Transfer (TST). Style transfer involves the
rewriting of text while altering one or several style attributes, such as authorship [
        <xref ref-type="bibr" rid="ref10 ref11 ref8 ref9">8, 9, 10, 11</xref>
        ] sentiment,
or politeness [
        <xref ref-type="bibr" rid="ref12 ref13 ref9">9, 12, 13</xref>
        ]. However, it is important to note that changing these style attributes can
sometimes significantly alter the meaning of a sentence. Despite this, many style transfer models aim to
transform sentences into ones of a diferent style while retaining similarity on the same topic [
        <xref ref-type="bibr" rid="ref14">14</xref>
        ]. This
presents a challenging yet intriguing task, as it requires striking a delicate balance between preserving
original meaning and adjusting stylistic elements.
      </p>
      <p>In the dynamic field of Artificial Intelligence (AI), the fusion of Large Language Models (LLMs)
with Reinforcement Learning (RL) techniques shows great potential. Particularly, Proximal Policy
Optimization (PPO), a subset of RL algorithms, has emerged as a powerful tool. This paper extensively
explores the integration of LLMs with RL, along with Parameter-Eficient Fine-Tuning (PEFT), aiming to
generate completely neutral text samples while maintaining their original meaning. Within this study,
we introduce the ToxiCleanse RL Approach, a strategy based on RL for mitigating text toxicity. To
accurately assess the impact of detoxification methods on the provided text, we propose a reward model
based on text similarity and toxicity levels, aiming to mitigate unintended biases related to various
social identities during the RL phase. This reward model guides the policy to generate neutral samples
that align more closely with ground truth samples. Empirical results demonstrate that utilizing RL for
ifne-tuning language models to optimize the non-toxicity reward efectively reduces the generation of
toxic language, outperforming existing detoxification methods in the literature.</p>
      <p>
        We, the MarSan_AI team, perform a large-scale evaluation of style transfer models on the Multilingual
Text Detoxification shared task at PAN 2024 [
        <xref ref-type="bibr" rid="ref15 ref16">15, 16</xref>
        ], comparing our new models with baselines and
state-of-the-art approaches. We release our code and data in our GitHub. Our contributions are
structured as follows: Section 2 details the task and data description. Section 3 reviews related work.
Section 4 introduces our model framework. Section 5.1 outlines the evaluation metrics, and finally,
Section 5 presents the results.
      </p>
    </sec>
    <sec id="sec-2">
      <title>2. Task &amp; Data description</title>
      <p>
        The Multilingual Text Detoxification (TextDetox) [
        <xref ref-type="bibr" rid="ref15 ref16">15, 16</xref>
        ] task for 2024 addressed the pressing need to
combat toxicity in user-generated content on social media platforms. Unlike traditional approaches that
often involved simply blocking or filtering toxic content, TextDetox encouraged a proactive approach
by providing users with a neutralized version of their messages. With evaluation based on style transfer
accuracy, content preservation, and fluency, participants were challenged to employ unsupervised and
cross-lingual detoxification methods to tackle the diverse linguistic and cultural nuances of toxicity.
      </p>
      <p>
        The TextDetox task provided datasets for English, Russian, and multilingual contexts. For each of the
nine diferent languages, there were 1,000 parallel pairs available, split into development (400 pairs)
and test (600 pairs) sets. Additionally, there were 19.7k English and 11.1k Russian data points available
for the training phase. The datasets aimed to facilitate the development and evaluation of efective
solutions for detoxifying toxic text across diverse linguistic contexts, contributing to a safer and more
inclusive online environment. All submissions were managed through Codalab and tira.io [
        <xref ref-type="bibr" rid="ref17">17</xref>
        ].
      </p>
    </sec>
    <sec id="sec-3">
      <title>3. Background</title>
      <p>In [18], a groundbreaking method for detoxification leveraging parallel data was introduced. This
innovative approach involved the creation of parallel datasets containing toxic sentences alongside
their corresponding non-toxic paraphrases, both in English and Russian languages. Through
meticulous crowdsourcing eforts, the authors curated over 10,000 non-toxic paraphrases for English toxic
sentences, marking the inception of the first parallel datasets tailored explicitly for detoxification
purposes. Furthermore, the study illustrated the process of distilling existing paraphrase datasets to
derive toxic-neutral sentence pairs. By training detoxification models on these meticulously crafted
datasets, the paper demonstrated substantial enhancements over prevailing unsupervised methods,
underscoring the eficacy of harnessing parallel data in detoxification systems.</p>
      <p>
        [
        <xref ref-type="bibr" rid="ref4">4</xref>
        ] Introduced two novel methods for removing toxicity from text. The first, ParaGeDi, employed
style-guided language models and paraphrasing to retain content while eliminating toxicity. The second
method, CondBERT, utilized BERT to replace toxic words with non-ofensive alternatives. [
        <xref ref-type="bibr" rid="ref1">1</xref>
        ] Introduced
pioneering methods for detoxifying Russian texts, marking a significant step in combating ofensive
language. Their innovative approaches, based on BERT and GPT-2 models, efectively transformed toxic
content into neutral language. Through rigorous evaluation and comparison, the authors demonstrated
the eficacy of their techniques, ofering valuable contributions to content moderation in the Russian
language. This study not only expanded the scope of TST tasks but also provided practical tools for
fostering a safer online environment.
      </p>
      <p>
        [19] introduced a method to extend text detoxification to multiple languages using parallel data.
It is built upon existing techniques, showing the efectiveness of parallel corpora in improving text
detoxification. The study also discussed the broader context of TST, highlighting the importance of
parallel datasets in advancing research in this domain. By extending the ParaDetox pipeline to support
multiple languages, including Russian, Ukrainian, and Spanish, the work aimed to facilitate safer
communication in digital environments across linguistic boundaries. [20] follows an iterative process
of leveraging human feedback to train summarization models. Initially, human preference data was
collected by presenting evaluators with pairs of summaries and asking them to choose the better one.
Then the reward model was trained to predict these human preferences, which was used as a reward
function in a reinforcement learning setup, specifically employing the PPO algorithm. [ 21] enhances
reinforcement learning from human feedback (RLHF) by introducing contrastive rewards. It involves
two steps: ofline sampling to obtain baseline responses and computing contrastive rewards based on
these samples. These rewards enable self-improvement of the RL policy, penalizing uncertainty and
improving robustness. Empirical testing demonstrates superior performance compared to standard
RLHF, highlighting its efectiveness in aligning LLMs with human feedback. There are also other
researches in this field [
        <xref ref-type="bibr" rid="ref4">4, 22</xref>
        ].
      </p>
    </sec>
    <sec id="sec-4">
      <title>4. System Overview</title>
      <p>We provide an overview of our approach in Figure 1. This section delineates the fine-tuning process for
generating less-toxic summaries using the Mistral LLM. Initially, we outline the base LLM architecture
and its parameters, providing a foundational understanding of the model and generating zero-shot
samples. Subsequently, we detail the phases of Supervised Fine-Tuning (SFT) and the Proximal Policy
Optimization (PPO) method.</p>
      <sec id="sec-4-1">
        <title>4.1. Base LLM and Parameters</title>
        <p>The Mistral 7B, introduced by Mistral AI, is a groundbreaking large language model available on the
Hugging Face repository. It features advanced attention mechanisms like Sliding Window Attention
(SWA) and Grouped-query Attention (GQA), optimizing both speed and memory usage. This design
enables Mistral 7B to outperform larger models such as Llama 2 (13B) and Llama 1 (34B) on various
benchmarks, making it versatile for commercial and research applications. Licensed under Apache 2.0,
Mistral 7B is ideal for self-hosted AI solutions [23]. Although we started with Mistral as our initial
model choice, we ultimately selected an upgraded version Mistral-T5-7B-v1 developed by Ignos for our
ifnal model. This advanced language model demonstrates exceptional performance in handling lengthy
sequences, accommodating up to 32,768 tokens in context.</p>
      </sec>
      <sec id="sec-4-2">
        <title>4.2. First Phase Fine-Tuning with SFT</title>
        <p>We started with models pre-trained to autoregressively generate non-toxic samples. These pre-trained
models served as ‘zero-shot’ baselines. Taking the prompts, the toxic sample, and the neutral
corresponding sample as an example. In the next step, the Supervised Fine-tuning (SFT) model learns how to
generate a neutral sample () ∼  SFT(|) based on the user’s given toxic sample . This process enables
us to acquire a collection of baseline responses denoted as {base,, }=1, where base,, ∼  SFT(·| ).
These responses are then used as a comparison for measuring PPO model output in the evaluation
phase. Hence, we fine-tuned these models via supervised learning on our competition datasets. These
supervised models were used to generate initial neutral samples for collecting comparisons, to initialize
our policy and reward models, and as baselines for evaluation.</p>
      </sec>
      <sec id="sec-4-3">
        <title>4.3. Last Phase Fine-Tuning with PPO</title>
        <p>In reinforcement learning with neural network function approximators, various approaches like deep
Q-learning, vanilla policy gradient methods, and trust region or natural policy gradient methods have
been explored. Each has strengths and weaknesses in scalability, data eficiency, and robustness across
diverse problem domains. To address these challenges, Schulman et al. introduced the PPO algorithm
[24].</p>
        <p>PPO, a policy gradient method, combines the benefits of Trust Region Policy Optimization (TRPO)
while simplifying implementation and enhancing sample eficiency. Policy gradient methods estimate
the policy gradient and use it in a stochastic gradient ascent algorithm. The common gradient estimator
is:</p>
        <p>ˆ = Eˆ [︁∇ log   (|)ˆ]︁
where   is a stochastic policy and ˆ is an estimator of the advantage function at timestep . The
expectation Eˆ[· ] denotes the empirical average over a finite batch of samples. However, performing
multiple optimization steps on this loss using the same trajectory can lead to excessively large policy
updates. Unlike conventional policy gradient methods that typically perform a single gradient update
per data sample, PPO employs a surrogate objective function to enhance training stability by limiting
the magnitude of policy updates and avoiding drastic changes. This is achieved by calculating a ratio
indicating the diference between the current and old policies and then clipping this ratio within a
specific range, [1 − , 1 +  ]. PPO ensures that policy updates remain conservative, promoting stable
and reliable training progress.</p>
        <p>Central to PPO is the clipped surrogate objective function, which stabilizes training by incorporating
a constrained probability ratio between the current and old policies. This prevents overly large policy
updates and ensures that gradient ascent steps encourage actions leading to higher rewards while
avoiding harmful actions. Determining the appropriate step size is critical: too small a step results in
slow training, while too large a step introduces excessive variability. PPO addresses this by constraining
policy updates within a small range using the clipped surrogate objective function, efectively avoiding
destructive large-weight updates.</p>
        <p>Let ( ) denote the probability ratio:
( ) =
  (|)
  old (|)
  ( ) = Eˆ [︁( )ˆ]︁
This ratio represents the probability of taking action  in state  under the current policy, divided
by the probability of taking the same action under the previous policy. Essentially, ( ) measures the
divergence between the old and current policies.</p>
        <p>TRPO maximizes the surrogate objective:
To avoid excessively large policy updates, PPO modifies this objective:
 ( ) = Eˆ [︁min ( )ˆ, clip(( ), 1 − , 1 +  )ˆ︁)]︁</p>
        <p>︁(
ensuring stable updates.
where  is a hyperparameter. The clipping removes the incentive for  to move outside [1 − , 1 +  ],</p>
        <p>By using the XLM-Roberta model’s non-toxic class score as the reward in the PPO algorithm, we
iteratively fine-tune the policy to generate non-toxic outputs. This approach ensures the generation of
high-quality, detoxified content by leveraging the synergy between PPO updates and the toxicity-based
reward model.</p>
      </sec>
      <sec id="sec-4-4">
        <title>4.4. Reward Model</title>
        <p>In our approach, we utilize an XLM-Roberta model fine-tuned for toxicity detection to compute rewards
for the PPO algorithm. This model classifies input data into two classes: non-toxic (class index 0) and
toxic (class index 1). The reward signal for the PPO model is derived from the output score of the
XLM-Roberta model for the non-toxic class. Specifically, the score corresponding to class index 0 is
used as the reward. A positive score indicates non-toxic content, while a negative score indicates toxic
content. For instance, if the XLM-Roberta model assigns a score of +3.5 to a sample, it is considered
non-toxic, and if it assigns a score of -2.8, the sample is considered toxic. This score is then used in the
PPO update rule, guiding the model towards generating non-toxic content.</p>
        <p>Using the reward model, we defined a policy to generate higher-quality outputs with reinforcement
learning, maximizing the reward with the PPO algorithm. The policy was initialized using the supervised
ifne-tuned model. To encourage exploration and prevent the policy from deviating too much from the
supervised model, we included a KL divergence term in the reward. This term serves as an entropy
bonus, ensuring consistency with the training data.</p>
        <p>The PPO value function used a Transformer with separate parameters from the policy to prevent
pre-trained policy degradation during early training. The value function was initialized with reward
model parameters. In our experiments, the reward model, policy, and value function were the same size.</p>
        <p>The reward function (, ) used in the PPO model to generate detoxified samples is defined as:
(, ) =  (, ) −  log
︂(  RL (|) )︂
 SFT(|)
(1)
where  (, ) is the reward from the transformer model,  controls the penalty,  RL (|) is the
PPO-trained policy probability, and  SFT(|) is the supervised fine-tuned policy probability. This
logarithmic term penalizes deviations from the supervised model, encouraging optimal and detoxified
outputs.</p>
      </sec>
    </sec>
    <sec id="sec-5">
      <title>5. Results</title>
      <p>In this section, we delve into the outcomes obtained from our analysis of the dev data. We explore
the performance of various models and their eficacy in handling the detoxification task. Through
comprehensive evaluation and analysis, we shed light on the strengths and weaknesses of each model,
providing insights into their capabilities and potential areas for improvement.</p>
      <sec id="sec-5-1">
        <title>5.1. Evaluation Metrics</title>
        <p>Multiple attempts have been made to evaluate sentence style and toxicity, focusing on three key
parameters of style transfer quality: text style, content preservation, and text fluency.[ 25] examine
various detoxification models and explore the correlation between manual and automatic evaluation
metrics, identifying metrics like ChrF and BertScore as potential proxies for human evaluation. For the
competition, organizers provide automatic evaluation metrics have set.</p>
        <p>• Style Transfer Accuracy (STA): This metric classifies the non-toxicity level in the generated
paraphrase using a specifically fine-tuned xlm-roberta-large model for toxicity binary classification.</p>
        <p>Additionally, a base fine-tuned version of the classifier is provided for further experimentation.
• Content Preservation (SIM): This metric evaluates the content similarity between the original
toxic sentence and the generated paraphrase by calculating the cosine similarity between LaBSe
embeddings.
• ChrF: This metric estimates the text adequacy and its similarity to human-written detoxified
references.
• Joint: combines the individual components of the automatic evaluation, is calculated as the mean
of STA * SIM * FL per sample. This composite metric provides a unified measure of style transfer
quality across the competition.</p>
      </sec>
      <sec id="sec-5-2">
        <title>5.2. Quantitative Analysis of Dev Dataset Results</title>
        <p>Table 1 showcases the performance of various language models (LLMs) in detoxifying toxic words
in English dev datasets, evaluated across four metrics: STA, SIM, CHRF, and J. Mistral-T5-7B-v1
demonstrating the highest overall performance, with Mistral and Mistral-7B-Instruct-v0.1 also exhibiting
strong results. Falcon-7b and its variants show moderate performance, while Lama2-13B and Lama2-7B
display a considerable diference in performance. Zephyr-7b-beta and solar-10.7B-v1.0 perform well
across all metrics. Similarly, Table 2 compares LLM performance for Russian dev datasets.
MistralT5-7B-v1 leads, with competitive results from Mistral, falcon-7b-instruct, Lama2-13B, Lama2-7B, Zep,
and solar. Zero-shot versions generally lag behind fine-tuned models, emphasizing the importance of
tuning.</p>
        <p>The Table 3 depicts the performance of English language models before and after data augmentation.
Augmented versions show nuanced changes in metrics, with some models maintaining competitiveness
while others exhibit slight variations. These results underscore the diverse efects of data augmentation
on model performance, emphasizing the need for tailored strategies to optimize outcomes. Data
augmentation can improve SIM (similarity) and CHRF (character n-gram F-score) metrics by diversifying
the training data, thereby enhancing the model’s ability to understand and generate text that aligns
closely with the semantics and structure of the original text. Augmentation techniques allow the
model to learn more robust representations of language and improve its performance in tasks requiring
similarity and character-level understanding.</p>
        <p>The Table 4 illustrates the comparison between SFT and PPO results for Mistral-T5-7B-v1 models in
both English and Russian datasets. In English, PPO outperforms SFT across most metrics, with higher
scores in STA, SIM, and J, indicating improved text understanding and generation capabilities. Similarly,
in Russian, PPO shows enhancements in SIM, CHRF, and J scores compared to SFT. The improvements
observed with PPO over SFT could be attributed to the reinforcement learning nature of PPO, which
allows the model to iteratively adjust its parameters based on feedback from the environment, leading
to more eficient learning and better adaptation to the given task and dataset. Additionally, PPO’s ability
to explore and exploit the training data more efectively might contribute to its superior performance
over SFT. Our approach was evaluated using both automatic and manual evaluation metrics provided
by the PAN organizers.</p>
        <p>J</p>
        <p>J</p>
      </sec>
      <sec id="sec-5-3">
        <title>5.3. Automatic Evaluation Test Results</title>
        <p>The automatic evaluation was based on the J evaluation metric. The results of the test data from the
PAN organizers’ leaderboard are presented in 5:</p>
      </sec>
      <sec id="sec-5-4">
        <title>5.4. Manual Evaluation Test Results</title>
        <p>Manual evaluation was conducted through crowdsourcing on a random subsample of 100 texts per
language. Our team achieved second place in the leaderboard for English data detoxification. The
manual evaluation results on the test data from the PAN organizers’ leaderboard are presented in 6.</p>
        <p>Our approach performed competitively, especially in the field of English data cleansing. The final
results highlighted the efectiveness of our RL-based method, which achieved second place in an English
dataset. It’s noteworthy that our objective was to assess model quality by fine-tuning large language
models on substantial datasets. We focused exclusively on Russian and English due to the availability
of high-quality open-source datasets for these languages.</p>
      </sec>
    </sec>
    <sec id="sec-6">
      <title>6. Conclusion &amp; Future Work</title>
      <p>In summary, our research presents a holistic strategy for addressing toxicity in text through the
application of RL, specifically PPO, in conjunction with LLMs. Our ToxiCleanse RL Approach utilizes
RL fine-tuning to generate neutral text outputs while preserving their original meaning. Using a
toxicity-based reward model, we successfully mitigate the generation of toxic language, surpassing
existing detoxification methods and even outperforming Supervised Fine-tuned LLMs. Our findings
underscore the efectiveness of RL-based rewards in elevating the quality of generated content.</p>
      <p>Looking ahead, future endeavors could involve refining LM/LLM-based rewards through manual
ifne-tuning to enhance reward model accuracy further. Moreover, developing a similarity-based reward
that measures detoxification while penalizing deviations between original and generated samples could
prove instrumental in maintaining text integrity. These initiatives are poised to propel advancements
in text detoxification, fostering safer and more inclusive online environments.
Computer Science, Springer, Berlin Heidelberg New York, 2023, pp. 236–241. URL: https://link.
springer.com/chapter/10.1007/978-3-031-28241-6_20. doi:10.1007/978-3-031-28241-6_20.
[18] V. Logacheva, D. Dementieva, S. Ustyantsev, D. Moskovskiy, D. Dale, I. Krotova, N. Semenov,
A. Panchenko, Paradetox: Detoxification with parallel data, in: Proceedings of the 60th Annual
Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2022, pp.
6804–6818.
[19] D. Dementieva, N. Babakov, A. Panchenko, Multiparadetox: Extending text detoxification with
parallel data to new languages, arXiv preprint arXiv:2404.02037 (2024).
[20] N. Stiennon, L. Ouyang, J. Wu, D. Ziegler, R. Lowe, C. Voss, A. Radford, D. Amodei, P. F. Christiano,
Learning to summarize with human feedback, Advances in Neural Information Processing Systems
33 (2020) 3008–3021.
[21] W. Shen, X. Zhang, Y. Yao, R. Zheng, H. Guo, Y. Liu, Improving reinforcement learning from
human feedback using contrastive rewards, arXiv preprint arXiv:2403.07708 (2024).
[22] D. Dementieva, D. Moskovskiy, D. Dale, A. Panchenko, Exploring methods for cross-lingual text
style transfer: The case of text detoxification, arXiv preprint arXiv:2311.13937 (2023).
[23] A. Q. Jiang, A. Sablayrolles, A. Mensch, C. Bamford, D. S. Chaplot, D. d. l. Casas, F. Bressand,</p>
      <p>G. Lengyel, G. Lample, L. Saulnier, et al., Mistral 7b, arXiv preprint arXiv:2310.06825 (2023).
[24] J. Schulman, F. Wolski, P. Dhariwal, A. Radford, O. Klimov, Proximal policy optimization algorithms,
arXiv preprint arXiv:1707.06347 (2017).
[25] D. Dementieva, V. Logacheva, I. Nikishina, A. Fenogenova, D. Dale, I. Krotova, N. Semenov,
T. Shavrina, A. Panchenko, Russe-2022: Findings of the first russian detoxification shared task
based on parallel corpora (????).</p>
    </sec>
  </body>
  <back>
    <ref-list>
      <ref id="ref1">
        <mixed-citation>
          [1]
          <string-name>
            <given-names>D.</given-names>
            <surname>Dementieva</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Moskovskiy</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Logacheva</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Dale</surname>
          </string-name>
          ,
          <string-name>
            <given-names>O.</given-names>
            <surname>Kozlova</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Semenov</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Panchenko</surname>
          </string-name>
          ,
          <article-title>Methods for detoxification of texts for the russian language</article-title>
          ,
          <source>Multimodal Technologies and Interaction</source>
          <volume>5</volume>
          (
          <year>2021</year>
          )
          <fpage>54</fpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref2">
        <mixed-citation>
          [2]
          <string-name>
            <given-names>D.</given-names>
            <surname>Moskovskiy</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Dementieva</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Panchenko</surname>
          </string-name>
          ,
          <article-title>Exploring cross-lingual text detoxification with large multilingual language models</article-title>
          .,
          <source>in: Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics: Student Research Workshop</source>
          ,
          <year>2022</year>
          , pp.
          <fpage>346</fpage>
          -
          <lpage>354</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref3">
        <mixed-citation>
          [3]
          <string-name>
            <given-names>A.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Sudhakar</surname>
          </string-name>
          ,
          <string-name>
            <given-names>Y.</given-names>
            <surname>Ji</surname>
          </string-name>
          ,
          <article-title>Simple text detoxification by identifying a linear toxic subspace in language model embeddings</article-title>
          ,
          <source>arXiv preprint arXiv:2112.08346</source>
          (
          <year>2021</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref4">
        <mixed-citation>
          [4]
          <string-name>
            <given-names>D.</given-names>
            <surname>Dale</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Voronov</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Dementieva</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Logacheva</surname>
          </string-name>
          ,
          <string-name>
            <given-names>O.</given-names>
            <surname>Kozlova</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Semenov</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Panchenko</surname>
          </string-name>
          ,
          <article-title>Text detoxification using large pre-trained neural models</article-title>
          ,
          <source>arXiv preprint arXiv:2109.08914</source>
          (
          <year>2021</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref5">
        <mixed-citation>
          [5]
          <string-name>
            <given-names>S.</given-names>
            <surname>Mukherjee</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Bansal</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A. K.</given-names>
            <surname>Ojha</surname>
          </string-name>
          ,
          <string-name>
            <given-names>J. P.</given-names>
            <surname>McCrae</surname>
          </string-name>
          ,
          <string-name>
            <given-names>O.</given-names>
            <surname>Dušek</surname>
          </string-name>
          ,
          <article-title>Text detoxification as style transfer in english and hindi</article-title>
          ,
          <source>arXiv preprint arXiv:2402.07767</source>
          (
          <year>2024</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref6">
        <mixed-citation>
          [6]
          <string-name>
            <given-names>J.</given-names>
            <surname>Qian</surname>
          </string-name>
          ,
          <source>Text Detoxification in Natural Language Processing</source>
          , University of California, Santa Barbara,
          <year>2022</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref7">
        <mixed-citation>
          [7]
          <string-name>
            <given-names>S.</given-names>
            <surname>Jain</surname>
          </string-name>
          , G. Kaushik,
          <string-name>
            <given-names>P.</given-names>
            <surname>Prabhu</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Godbole</surname>
          </string-name>
          ,
          <article-title>Detox: Nlp based classification and euphemistic text substitution for toxic comments</article-title>
          ,
          <source>in: 2021 12th International Conference on Computing Communication and Networking Technologies (ICCCNT)</source>
          , IEEE,
          <year>2021</year>
          , pp.
          <fpage>1</fpage>
          -
          <lpage>5</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref8">
        <mixed-citation>
          [8]
          <string-name>
            <given-names>M.</given-names>
            <surname>Najafi</surname>
          </string-name>
          , E. Tavan,
          <article-title>Text-to-text transformer in authorship verification via stylistic and semantical analysis</article-title>
          .,
          <source>in: CLEF (Working Notes)</source>
          ,
          <year>2022</year>
          , pp.
          <fpage>2607</fpage>
          -
          <lpage>2616</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref9">
        <mixed-citation>
          [9]
          <string-name>
            <given-names>E.</given-names>
            <surname>Tavan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Najafi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>R.</given-names>
            <surname>Moradi</surname>
          </string-name>
          ,
          <article-title>Identifying ironic content spreaders on twitter using psychometrics, contextual and ironic features with gradient boosting classifier</article-title>
          .,
          <source>in: CLEF (Working Notes)</source>
          ,
          <year>2022</year>
          , pp.
          <fpage>2687</fpage>
          -
          <lpage>2697</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref10">
        <mixed-citation>
          [10]
          <string-name>
            <given-names>M.</given-names>
            <surname>Najafi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S.</given-names>
            <surname>Sadidpur</surname>
          </string-name>
          , Paa:
          <article-title>Persian author attribution using dense and recursive connection (</article-title>
          <year>2024</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref11">
        <mixed-citation>
          [11]
          <string-name>
            <surname>H. B. Giglou</surname>
            ,
            <given-names>T.</given-names>
          </string-name>
          <string-name>
            <surname>Rahgooy</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          <string-name>
            <surname>Rahgouy</surname>
            ,
            <given-names>J.</given-names>
          </string-name>
          <string-name>
            <surname>Razmara</surname>
          </string-name>
          ,
          <article-title>Uot-uwf-partai at semeval-2021 task 5: Self attention based bi-gru with multi-embedding representation for toxicity highlighter</article-title>
          ,
          <source>arXiv preprint arXiv:2104.13164</source>
          (
          <year>2021</year>
          ).
        </mixed-citation>
      </ref>
      <ref id="ref12">
        <mixed-citation>
          [12]
          <string-name>
            <given-names>M.</given-names>
            <surname>Najafi</surname>
          </string-name>
          , E. Tavan, Marsan at semeval
          <article-title>-2022 task 6: isarcasm detection via t5 and sequence learners</article-title>
          ,
          <source>in: Proceedings of the 16th International Workshop on Semantic Evaluation (SemEval-2022)</source>
          ,
          <year>2022</year>
          , pp.
          <fpage>978</fpage>
          -
          <lpage>986</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref13">
        <mixed-citation>
          [13]
          <string-name>
            <given-names>E.</given-names>
            <surname>Tavan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Najafi</surname>
          </string-name>
          , Marsan at semeval-2023 task 10:
          <article-title>Can adversarial training with help of a graph convolutional network detect explainable sexism?</article-title>
          ,
          <source>in: Proceedings of the 17th International Workshop on Semantic Evaluation (SemEval-2023)</source>
          ,
          <year>2023</year>
          , pp.
          <fpage>1011</fpage>
          -
          <lpage>1020</lpage>
          .
        </mixed-citation>
      </ref>
      <ref id="ref14">
        <mixed-citation>
          [14]
          <string-name>
            <given-names>D.</given-names>
            <surname>Dale</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Voronov</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Dementieva</surname>
          </string-name>
          ,
          <string-name>
            <given-names>V.</given-names>
            <surname>Logacheva</surname>
          </string-name>
          ,
          <string-name>
            <given-names>O.</given-names>
            <surname>Kozlova</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Semenov</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Panchenko</surname>
          </string-name>
          ,
          <article-title>Text detoxification using large pre-trained neural models (</article-title>
          <year>2021</year>
          ), arXiv preprint arXiv:
          <volume>2109</volume>
          .08914 (????).
        </mixed-citation>
      </ref>
      <ref id="ref15">
        <mixed-citation>
          [15]
          <string-name>
            <given-names>J.</given-names>
            <surname>Bevendorf</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X. B.</given-names>
            <surname>Casals</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Chulvi</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Dementieva</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Elnagar</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Freitag</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Fröbe</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Korenčić</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Mayerl</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Mukherjee</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Panchenko</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Potthast</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Rangel</surname>
          </string-name>
          ,
          <string-name>
            <given-names>P.</given-names>
            <surname>Rosso</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Smirnova</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Stamatatos</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Stein</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Taulé</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Ustalov</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Wiegmann</surname>
          </string-name>
          , E. Zangerle,
          <article-title>Overview of PAN 2024: Multi-Author Writing Style Analysis, Multilingual Text Detoxification, Oppositional Thinking Analysis, and Generative AI Authorship Verification, in: Experimental IR Meets Multilinguality, Multimodality, and Interaction</article-title>
          .
          <source>Proceedings of the Fourteenth International Conference of the CLEF Association (CLEF</source>
          <year>2024</year>
          ), Lecture Notes in Computer Science, Springer, Berlin Heidelberg New York,
          <year>2024</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref16">
        <mixed-citation>
          [16]
          <string-name>
            <given-names>D.</given-names>
            <surname>Dementieva</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Moskovskiy</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Babakov</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A. A.</given-names>
            <surname>Ayele</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Rizwan</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Schneider</surname>
          </string-name>
          ,
          <string-name>
            <given-names>X.</given-names>
            <surname>Wang</surname>
          </string-name>
          ,
          <string-name>
            <given-names>S. M.</given-names>
            <surname>Yimam</surname>
          </string-name>
          ,
          <string-name>
            <given-names>D.</given-names>
            <surname>Ustalov</surname>
          </string-name>
          ,
          <string-name>
            <given-names>E.</given-names>
            <surname>Stakovskii</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Smirnova</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Elnagar</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Mukherjee</surname>
          </string-name>
          ,
          <string-name>
            <given-names>A.</given-names>
            <surname>Panchenko</surname>
          </string-name>
          ,
          <article-title>Overview of the multilingual text detoxification task at pan 2024</article-title>
          , in: CEUR Workshop Proceedings, CEUR-WS.org,
          <year>2024</year>
          .
        </mixed-citation>
      </ref>
      <ref id="ref17">
        <mixed-citation>
          [17]
          <string-name>
            <given-names>M.</given-names>
            <surname>Fröbe</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Wiegmann</surname>
          </string-name>
          ,
          <string-name>
            <given-names>N.</given-names>
            <surname>Kolyada</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Grahm</surname>
          </string-name>
          ,
          <string-name>
            <given-names>T.</given-names>
            <surname>Elstner</surname>
          </string-name>
          ,
          <string-name>
            <given-names>F.</given-names>
            <surname>Loebe</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Hagen</surname>
          </string-name>
          ,
          <string-name>
            <given-names>B.</given-names>
            <surname>Stein</surname>
          </string-name>
          ,
          <string-name>
            <given-names>M.</given-names>
            <surname>Potthast</surname>
          </string-name>
          ,
          <article-title>Continuous Integration for Reproducible Shared Tasks with TIRA.io</article-title>
          , in: J.
          <string-name>
            <surname>Kamps</surname>
            ,
            <given-names>L.</given-names>
          </string-name>
          <string-name>
            <surname>Goeuriot</surname>
            ,
            <given-names>F.</given-names>
          </string-name>
          <string-name>
            <surname>Crestani</surname>
            ,
            <given-names>M.</given-names>
          </string-name>
          <string-name>
            <surname>Maistro</surname>
            ,
            <given-names>H.</given-names>
          </string-name>
          <string-name>
            <surname>Joho</surname>
            ,
            <given-names>B.</given-names>
          </string-name>
          <string-name>
            <surname>Davis</surname>
            ,
            <given-names>C.</given-names>
          </string-name>
          <string-name>
            <surname>Gurrin</surname>
            ,
            <given-names>U.</given-names>
          </string-name>
          <string-name>
            <surname>Kruschwitz</surname>
            ,
            <given-names>A</given-names>
          </string-name>
          . Caputo (Eds.),
          <source>Advances in Information Retrieval. 45th European Conference on IR Research (ECIR</source>
          <year>2023</year>
          ), Lecture Notes in
        </mixed-citation>
      </ref>
    </ref-list>
  </back>
</article>