Bridging the Reasoning Gap in Vietnamese with Small Language Models via Test-Time Scaling

πŸ“… 2026-04-20
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This study addresses the limited reasoning capabilities of small language models in low-resource languages such as Vietnamese, particularly their difficulty in generating coherent chains of thought. Focusing on elementary-level Vietnamese mathematical reasoning tasks, the work proposes a combined approach of supervised fine-tuning (SFT) and a simplified test-time scaling strategy to effectively bridge the gap between a model’s internal knowledge and its output format. The authors introduce Vi-S1K, a high-quality dataset, and Vi-Elementary-Bench, a dedicated evaluation benchmark, employing an LLM-as-a-Judge protocol for assessment. Experimental results demonstrate that SFT improves explanation quality by 77% and achieves a latent knowledge accuracy of 4.05 out of 5.00. Furthermore, the simplified reasoning pipeline is shown to outperform complex agent-based workflows, offering greater practicality for deployment on edge devices.

Technology Category

Machine Learning: Large Multimodal Models (LMMs)Natural Language Processing: (Large) Language ModelsKnowledge Representation and Reasoning: Computational Complexity of Reasoning

Application Category

Semantics and Knowledge: Methods to enhance, augment, integrate or synergize semantic models such as knowledge graphs and LLMsSearch and Retrieval-Augmented AI: Large language models for searchEconomics, Online Markets and Human Computation: Cost models of using LLMs in production systems
πŸ“ Abstract
The democratization of ubiquitous AI hinges on deploying sophisticated reasoning capabilities on resource-constrained devices. However, Small Language Models (SLMs) often face a "reasoning gap", particularly in non-English languages like Vietnamese, where they struggle to maintain coherent chains of thought. This paper investigates Test-Time Scaling strategies for the Qwen3-1.7B architecture within the context of Vietnamese Elementary Mathematics. We introduce Vi-S1K, a high-fidelity reasoning dataset localized via a Gemini 2.5 Flash-Lite powered pipeline, and Vi-Elementary-Bench, a dual-resource benchmark for rigorous evaluation. Using an LLM-as-a-Judge protocol, we reveal that the base model possesses robust latent knowledge (Accuracy: 4.05/5.00) but suffers from a severe "formatting gap" in communication. Supervised Fine-Tuning (SFT) acts as a critical "reasoning unlocker", yielding a 77% improvement in Explanation Quality and bridging the gap between raw calculation and pedagogical coherence. Furthermore, our analysis of prompting strategies uncovers a significant trade-off: structured frameworks like ReAct impose a "cognitive tax" on the 1.7B parameter capacity, degrading performance relative to pure Chain-of-Thought (CoT) combined with Self-Consistency. These findings establish a deployment hierarchy for SLMs, demonstrating that SFT combined with simplified test-time scaling is superior to complex agentic workflows for edge-based reasoning.
Problem

Research questions and friction points this paper is trying to address.

reasoning gap
Small Language Models
Vietnamese
test-time scaling
resource-constrained devices
Innovation

Methods, ideas, or system contributions that make the work stand out.

Test-Time Scaling
Small Language Models
Supervised Fine-Tuning
Chain-of-Thought
Vietnamese Reasoning
πŸ”Ž Similar Papers
No similar papers found.
πŸ’Ό Related Jobs
No related jobs found.
B
Bui The Trung
Computer Science and Engineering, VNU Vietnam Japan University
Do Minh Duc
Do Minh Duc
University of Science, Vietnam National University, Hanoi
Geological & Geotechnical EngineeringGeohazardsClimate Change Adaptation
N
Nguyen Van Vinh
Faculty of Information Technology, VNU University of Engineering and Technology
B
Bui Nguyen Quoc Trinh
Faculty of Advanced Technologies and Engineering, VNU Vietnam Japan University