KyrgyzLLM-Bench: Benchmarking Kyrgyz Language Understanding

📅 2026-07-19
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the absence of native evaluation benchmarks for Kyrgyz-language large language models, as existing cross-lingual assessments predominantly rely on translated data and fail to capture the language’s unique linguistic and cultural characteristics. To bridge this gap, we introduce KyrgyzLLM-Bench, the first large-scale native benchmark for Kyrgyz, comprising the native datasets KyrgyzMMLU and KyrgyzRC, complemented by human-verified translation tasks. We systematically evaluate 26 open- and closed-source models under zero-shot and few-shot settings. Our analysis reveals that translation-induced plausibility shifts can compromise evaluation reliability, with consistent English–Kyrgyz performance on WinoGrande and BoolQ but a notable discrepancy on HellaSwag. Few-shot prompting enhances reading comprehension for certain open-source models. All data, code, and results are publicly released and integrated into mainstream multilingual evaluation frameworks.
📝 Abstract
Evaluating large language models (LLMs) across languages remains challenging, as most multilingual benchmarks rely on translated English datasets, often obscuring linguistic and cultural specificity in the target language. This issue is particularly pronounced for less-resourced languages such as Kyrgyz, where reliable natively authored evaluation data are scarce. Building on previously introduced Kyrgyz-language evaluation datasets, this work reports the first systematic and large-scale evaluation of LLMs in Kyrgyz using the KyrgyzLLM-Bench benchmark suite. KyrgyzLLM-Bench comprises two natively authored datasets$-$KyrgyzMMLU and KyrgyzRC$-$together with carefully translated and manually post-edited versions of WinoGrande, HellaSwag, BoolQ, and TruthfulQA. We evaluate 26 open- and closed-source LLMs under zero-shot and few-shot settings, analyzing model performance, cross-lingual transfer, and the impact of translation artifacts on evaluation reliability. Across families and tasks, model rankings transfer broadly from English to Kyrgyz on WinoGrande and BoolQ, and to a lesser extent on MMLU, while HellaSwag exhibits a substantial English-Kyrgyz performance gap consistent with translation-induced plausibility shifts. Few-shot prompting improves several open-source models on reading comprehension but behaves inconsistently for proprietary models on translated tasks. We publicly release all datasets, evaluation code, and per-model results, and integrate the Kyrgyz tasks into a widely used multilingual evaluation framework to support future research on Kyrgyz NLP.
Problem

Research questions and friction points this paper is trying to address.

language evaluation
low-resource languages
Kyrgyz
translation artifacts
multilingual benchmarks
Innovation

Methods, ideas, or system contributions that make the work stand out.

KyrgyzLLM-Bench
natively authored datasets
cross-lingual evaluation
translation artifacts
low-resource language
🔎 Similar Papers
No similar papers found.