DeepEdu-v1: Efficient and Scalable Agentic LLMs for Vietnamese Education

📅 2026-09-25
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenges of data sovereignty compliance, hallucinations concerning local knowledge, and long-context inference bottlenecks on consumer-grade GPUs in developing AI tutoring systems for Vietnamese education. To this end, it proposes the SCALE framework for constructing a Vietnamese intelligent tutoring system. Methodologically, the framework introduces a cluster-granularity token selection mechanism to optimize long-context retrieval, alongside a tuning-free self-improving agent layer that enables the continuous curation and accumulation of local knowledge. Experimental results demonstrate that, compared to standard vLLM serving, the proposed system reduces time-to-first-token latency by 35% and decreases retrieval invocations by 7.7 times, while improving accuracy on complex tasks from 70.0% to 79.5%. These findings confirm that SCALE achieves efficient and scalable educational assistance under low-resource constraints.
📝 Abstract
AI tutoring could markedly improve learning outcomes for students in developing regions such as Vietnam, yet the two obvious paths both fall short. Cloud assistants such as ChatGPT route sensitive student data to foreign servers---violating data-sovereignty laws such as Vietnam's Decree 53---and, pre-trained on Western-centric corpora, are not organized around the national textbook curriculum, so their knowledge of local content is unsystematic and frequently hallucinated. Self-hosting an open model keeps data on-premise but hits a two-fold wall: post-training quantization (AWQ, GPTQ) tames the static weight footprint, yet the dynamic KV cache and prefill latency of long tutoring contexts still cause out-of-memory failures and slow responses on consumer GPUs, while the model keeps hallucinating on region-specific material. We present DeepEdu-v1, an AI-tutoring system for Vietnamese education built on SCALE (Self-improving Context-Aware Learning Engine), a framework with two innovations. First, a long-context inference engine amortizes token selection from per-sub-chunk to per-cluster granularity; on long-context retrieval it issues x7.7 fewer retrieval calls than a state-of-the-art selective-attention baseline, cutting prefill latency (TTFT) by roughly 35% while matching or improving task accuracy. Second, a self-improving agentic layer continuously curates a verified playbook from past interactions instead of fine-tuning, a design intended to progressively reduce reliance on dominant-language priors as trustworthy local knowledge accumulates. In its deployed configuration, DeepEdu achieves a nearly x2 TTFT speedup over standard vLLM serving and lifts agentic accuracy from 70.0% to 79.5% on complex tasks, with the strongest per-track gains across financial-reasoning and interactive-agent benchmarks.
Problem

Research questions and friction points this paper is trying to address.

AI tutoring
data sovereignty
hallucination
long-context inference
resource-constrained deployment
Innovation

Methods, ideas, or system contributions that make the work stand out.

Long-context Inference Engine
Self-improving Agentic Layer
KV Cache Optimization
Hallucination Mitigation
Agentic LLM
🔎 Similar Papers
No similar papers found.
Q
Quang Nguyen
Posts and Telecommunications Institute of Technology, Ha Noi, Vietnam
H
Hieu Nguyen
Posts and Telecommunications Institute of Technology, Ha Noi, Vietnam
H
Hien Hoang
Posts and Telecommunications Institute of Technology, Ha Noi, Vietnam
Toan Pham
Toan Pham
Posts and Telecommunications Institute of Technology, Ha Noi, Vietnam
Cong Tran
Cong Tran
PhD, Posts and Telecommunications Institute of Technology, Vietnam
Computer ScienceArtificial IntelligenceMachine LearningData Mining
Nam Vu
Nam Vu
NHS England
Price-settingOnline-pricesEmpirical FinanceEmerging Market EconomiesHealth Economics