LIMIT: Less Is More for Instruction Tuning in Text-to-SQL

📅 2026-09-21
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
研究通过LIMIT框架,使用精心挑选的少量数据集解决Text-to-SQL指令调优问题,提高了模型在BIRD和Spider基准上的执行准确率。
📝 Abstract
Large language models have achieved remarkable progress on Text-to-SQL through reasoning-enhanced fine-tuning, yet existing approaches predominantly rely on massive instruction corpora under the assumption that scale drives performance. We challenge this paradigm by investigating a fundamental question: what is the minimal data requirement for effective Text-to-SQL instruction tuning? We propose LIMIT(Less Is More for Instruction Tuning in Text-to-SQL), a data-centric framework that demonstrates strong database reasoning can emerge from an extremely compact training set when examples are strategically selected. LIMIT operates through four stages: difficulty-aware filtering that identifies samples within the model's learning frontier, chain-of-thought synthesis with consistency-based selection, multi-dimensional quality scoring via LLM-as-judge, and genetic algorithm optimization that jointly maximizes schema coverage and sample quality. On the BIRD and Spider benchmark, LIMIT selects only 796 and 863 samples while achieving 100% table coverage, enabling Qwen3-8B to reach 69.1% and 88.9% execution accuracy.This result surpasses methods trained on 20 times more data and establishes a new state-of-the-art among open-source approaches. Our findings suggest that careful data curation, rather than scale, is the key to efficient Text-to-SQL learning.
Problem

Research questions and friction points this paper is trying to address.

Text-to-SQL
instruction tuning
minimal data requirement
database reasoning
Innovation

Methods, ideas, or system contributions that make the work stand out.

LIMIT
data-centric framework
strategic sample selection
genetic algorithm optimization
schema coverage
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
H
Haoyuan Ma
Zhejiang University
H
Hengwei Liu
Zhejiang University
L
Linjuan Wu
Zhejiang University
Y
Yongliang Shen
Zhejiang University
Weiming Lu
Weiming Lu
Zhejiang University
Natural Language ProcessingLarge Language ModelsAGI