🤖 AI Summary
This study addresses the absence of native LLM evaluation benchmarks and insufficient translated data coverage for Slovak by introducing a "native-first" evaluation paradigm. We construct a native benchmark comprising 30 datasets and evaluate 55 models using a unified multi-task framework coupled with an instruction-following checker. By incorporating grammatical-morphological resources, we conduct a systematic analysis integrating continual pretraining and test-time reasoning techniques. Our results reveal significant limitations in translated data, demonstrate that instruction repair effectively recovers performance degradation, and show that reasoning augmentation outperforms mere parameter scaling. Based on these findings, we propose four design principles for optimizing language models in low-resource settings.
📝 Abstract
Multilingual LLM benchmarks omit Slovak, a morphologically rich West Slavic language of five million speakers, or cover it only by machine translation. We present sk-bench, a native-first Slovak benchmark with 30 datasets (33 scored task variants) across ten skill categories. Eleven resources are introduced or first packaged for generative-LLM evaluation, including IFEval-SK with Slovak-adapted instruction checkers and native Chiby/SKJ1 resources for Slovak grammar and morphology. We evaluate 55 open- and closed-weights models under one harness. The best open model trails proprietary APIs by 12.6 points. Model rankings are similar for native and translated closed-form data ($ρ\geq0.98$), though translation separates the strongest models less well. By contrast, human-authored and LLM-generated QA questions rank models differently ($ρ=0.72$). For Qwen3-14B, continued Slovak pretraining lowers the overall score by 13.9 points. A small instruction set restores three quarters of that loss. Test-time reasoning improves scores by 8.5 to 12.5 points for models of 9B and above. Together, these findings suggest four design lessons for other under-resourced languages: use native data where translation fails, plan instruction repair after language adaptation, enable test-time reasoning before scaling up, and avoid overinvesting in target-language prompts. We release the data and code at https://github.com/slovak-nlp/sk-bench