ArXiv

KyrgyzLLM-Bench: Benchmarking Kyrgyz Language Understanding

Authors
Timur Turatali, Aida Turdubaeva, Rustem Izmailov...
Categories
cs.CL
arXiv
https://arxiv.org/abs/2607.17173v1
PDF
https://arxiv.org/pdf/2607.17173v1

Brief

KyrgyzLLM-Bench presents the first systematic, large-scale evaluation suite for Kyrgyz: two natively authored datasets (KyrgyzMMLU, KyrgyzRC) plus translated/post-edited WinoGrande, HellaSwag, BoolQ, and TruthfulQA. The authors evaluate 26 open- and closed-source LLMs (zero- and few-shot), find task-dependent cross-lingual ranking transfer and translation-induced gaps, and publicly release data and code (arXiv preprint 2026-07-19).

Why it matters

KyrgyzLLM-Bench supplies a Kyrgyz evaluation suite with two natively authored datasets (KyrgyzMMLU, KyrgyzRC) plus translated and manually post-edited WinoGrande, HellaSwag, BoolQ, and TruthfulQA; the authors evaluated 26 open- and closed-source LLMs under zero-shot and few-shot settings.

Key details

  • Key results: English→Kyrgyz model rankings transfer broadly on WinoGrande and BoolQ, transfer less on MMLU, and HellaSwag shows a substantial English–Kyrgyz performance gap attributed to translation-induced plausibility shifts; few-shot prompting improves several open-source models on reading comprehension but produces inconsistent effects for proprietary models.
  • The work is a 2026-07-19 arXiv preprint (Turatali, Turdubaeva, Izmailov, Alekseev, Nikolenko); all datasets, evaluation code, and per-model results are publicly released and integrated into a widely used multilingual evaluation framework to support future Kyrgyz NLP research.
Source evidence

Abstract

Evaluating large language models (LLMs) across languages remains challenging, as most multilingual benchmarks rely on translated English datasets, often obscuring linguistic and cultural specificity in the target language. This issue is particularly pronounced for less-resourced languages such as Kyrgyz, where reliable natively authored evaluation data are scarce. Building on previously introduced Kyrgyz-language evaluation datasets, this work reports the first systematic and large-scale evaluation of LLMs in Kyrgyz using the KyrgyzLLM-Bench benchmark suite. KyrgyzLLM-Bench comprises two natively authored datasets$-$KyrgyzMMLU and KyrgyzRC$-$together with carefully translated and manually post-edited versions of WinoGrande, HellaSwag, BoolQ, and TruthfulQA. We evaluate 26 open- and closed-source LLMs under zero-shot and few-shot settings, analyzing model performance, cross-lingual transfer, and the impact of translation artifacts on evaluation reliability. Across families and tasks, model rankings transfer broadly from English to Kyrgyz on WinoGrande and BoolQ, and to a lesser extent on MMLU, while HellaSwag exhibits a substantial English-Kyrgyz performance gap consistent with translation-induced plausibility shifts. Few-shot prompting improves several open-source models on reading comprehension but behaves inconsistently for proprietary models on translated tasks. We publicly release all datasets, evaluation code, and per-model results, and integrate the Kyrgyz tasks into a widely used multilingual evaluation framework to support future research on Kyrgyz NLP.

Comment: Preprint; manuscript currently under consideration at a journal