HarnessSQL: Harness-Native Training for SQL Agents in Realistic Database Environments

📅 2026-10-08
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the mismatch between Text-to-SQL model training and real-world multi-turn database interaction environments by proposing the HarnessSQL framework. This approach preserves complete interaction structures during both supervised fine-tuning and reinforcement learning, constructing an isolated, executable database environment for native training. It pioneers directly generating and verifying trajectories within the target SQL execution framework, ensuring full alignment between training and deployment. Furthermore, it integrates a hidden execution oracle, teacher trajectory filtering, and execution-reward-based reinforcement learning as optimization strategies. Experimental results demonstrate that this framework improves the execution accuracy of the Qwen3 model series to 54.8% while achieving strong performance across multiple out-of-distribution interactive benchmarks.
📝 Abstract
Text-to-SQL models are commonly trained to map questions directly to static queries, whereas real-world database agents operate through stateful, multi-turn interaction with live databases -- inspecting schemas, executing probe queries, diagnosing errors, and revising hypotheses. This creates a critical train-deploy mismatch, as the execution harness that mediates this interaction is introduced only at inference time. To bridge this gap, we propose HarnessSQL, a harness-native post-training framework that preserves the full interaction structure throughout both supervised fine-tuning and reinforcement learning. HarnessSQL builds isolated, executable database environments paired with hidden execution oracles, rolls out teachers directly inside the target SQL harness, and retains only verified trajectories for full-sequence SFT, followed by execution-reward RL. Across Spider 2.0-SQLite, HarnessSQL dramatically boosts the execution accuracy of compact models, raising Qwen3-8B from 15.5% to 45.2% and Qwen3-14B from 22.2% to 54.8%, while transferring effectively to out-of-distribution interactive benchmarks such as BIRD-Interact and LiveSQLBench. Our findings demonstrate that training database agents directly within their execution harness is essential for mastering complex, long-horizon database workflows.
Problem

Research questions and friction points this paper is trying to address.

Text-to-SQL
database agents
train-deploy mismatch
multi-turn interaction
execution harness
Innovation

Methods, ideas, or system contributions that make the work stand out.

Harness-Native Training
SQL Agents
Reinforcement Learning
Supervised Fine-Tuning
Execution Reward