Knowledge Distillation with Structured Chain-of-Thought for Text-to-SQL

πŸ“… 2025-12-18
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
Enterprise-level Text-to-SQL deployment faces a trilemma among cost, security, and performance: large language models (LLMs) incur prohibitive inference costs, while small language models (SLMs) suffer from insufficient accuracy and robustness. This paper proposes Struct-SQLβ€”the first structured chain-of-thought knowledge distillation framework grounded in query execution plans (QEPs). Rather than relying on unstructured natural-language reasoning, Struct-SQL uses QEPs as formal, executable reasoning blueprints to guide SLM training. The method integrates structured reasoning modeling, abstract QEP representation learning, and targeted SLM fine-tuning, thereby enhancing logical consistency and syntactic correctness of generated SQL. On standard benchmarks, Struct-SQL achieves an absolute +8.1% improvement in execution accuracy over unstructured CoT distillation baselines and significantly reduces syntax error rates. These results empirically validate the critical role of structured reasoning signals in enabling reliable semantic parsing for Text-to-SQL.

Technology Category

Machine Learning: Structured LearningNatural Language Processing: Sentence-level Semantics, Textual Inference, etc.Knowledge Representation and Reasoning: Computational Complexity of Reasoning

Application Category

Search and Retrieval-Augmented AI: Large language models for searchSemantics and Knowledge: Methods to enhance, augment, integrate or synergize semantic models such as knowledge graphs and LLMsEconomics, Online Markets and Human Computation: Cost models of using LLMs in production systems
πŸ“ Abstract
Deploying accurate Text-to-SQL systems at the enterprise level faces a difficult trilemma involving cost, security and performance. Current solutions force enterprises to choose between expensive, proprietary Large Language Models (LLMs) and low-performing Small Language Models (SLMs). Efforts to improve SLMs often rely on distilling reasoning from large LLMs using unstructured Chain-of-Thought (CoT) traces, a process that remains inherently ambiguous. Instead, we hypothesize that a formal, structured reasoning representation provides a clearer, more reliable teaching signal, as the Text-to-SQL task requires explicit and precise logical steps. To evaluate this hypothesis, we propose Struct-SQL, a novel Knowledge Distillation (KD) framework that trains an SLM to emulate a powerful large LLM. Consequently, we adopt a query execution plan as a formal blueprint to derive this structured reasoning. Our SLM, distilled with structured CoT, achieves an absolute improvement of 8.1% over an unstructured CoT distillation baseline. A detailed error analysis reveals that a key factor in this gain is a marked reduction in syntactic errors. This demonstrates that teaching a model to reason using a structured logical blueprint is beneficial for reliable SQL generation in SLMs.
Problem

Research questions and friction points this paper is trying to address.

Develop a cost-effective, secure, and high-performance Text-to-SQL system
Improve small language models by distilling structured reasoning from large models
Reduce syntactic errors in SQL generation using formal logical blueprints
Innovation

Methods, ideas, or system contributions that make the work stand out.

Structured reasoning representation for distillation
Query execution plan as formal reasoning blueprint
Knowledge Distillation framework reduces syntactic errors
πŸ”Ž Similar Papers