ArXiv

The Tatoxa System for Text Detoxification in Low-Resource Languages: The Case of Tatar

Authors
Ilseyar Alimova, Bogdan Monogov, Artyom Mazur...
Categories
cs.CL
arXiv
https://arxiv.org/abs/2606.26015v1
PDF
https://arxiv.org/pdf/2606.26015v1

Brief

Tatoxa targets automated detection and mitigation of abusive content for low-resource Tatar, combining a new dataset and fine-tuned models to achieve state-of-the-art detoxification performance versus both open-source and commercial LLMs. The paper emphasizes the importance of native Tatar training data: cross-lingual transfer from Russian degrades performance despite large Russian corpora. Summary is based on the abstract.

Why it matters

Tatoxa is a new state-of-the-art text detoxification system for the Tatar language (Alimova et al., 2026) that, according to comparative experiments, outperforms existing open-source and proprietary commercial LLMs on key quality metrics.

Key details

  • The authors introduce a new dataset for Tatar text detoxification designed specifically for fine-tuning and evaluation in low-resource settings.
  • Cross-lingual transfer experiments show that transferring from other languages — including culturally close Russian — performs significantly worse than models trained on native Tatar data, even when a large Russian corpus is available.
Source evidence

Abstract

Text detoxification, the automated detection and mitigation of abusive and harmful content, is essential for ensuring the safety of online communities and protecting users. However, low resource languages such as Tatar have received little research attention. In this paper we present Tatoxa, a novel state-of-the-art system for text detoxification in the Tatar language. Comparative experiments show that the proposed approach outperforms existing open source and proprietary commercial LLMs on key quality metrics. We also introduce a new dataset for text detoxification in Tatar, designed for fine tuning and evaluation in low resource settings. Finally, cross lingual transfer experiments indicate that transfer from other languages, including the culturally close Russian, performs significantly worse than training on native Tatar data even when a large Russian corpus is available.