FALCON: A Model and Dataset Agnostic Framework for Synthetic Data Generation for NL2SQL Pairs

📅 2026-10-02
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limitation of existing NL2SQL synthetic data generation methods, which often overlook linguistic ambiguity and produce oversimplified queries inadequate for real-world complexity. To this end, this work proposes FALCON, a framework that leverages compact open-source models to cost-effectively generate training data exhibiting realistic complexity and ambiguity awareness. By integrating reserved-word SQL seeds, persona-driven prompt engineering, and alignment-based data filtering, FALCON enables model- and database-agnostic local generation of challenging examples, combined with hybrid training strategies. Experimental results demonstrate that the synthesized data surpasses existing benchmarks in both SQL complexity and linguistic diversity, substantially improving model performance on difficult queries.
📝 Abstract
Relational databases are among the most widely deployed forms of structured knowledge, and natural language access to them requires grounding language onto schema entities and relations while handling the ambiguity inherent in how people phrase requests. Existing synthetic NL-to-SQL data generation methods largely ignore this ambiguity and produce oversimplified queries that fail to prepare models for the complexity of real-world structured knowledge access. We present FALCON, a framework that generates realistic, ambiguity-aware NL-to-SQL data matching the complexity of challenging real-world benchmarks, at low cost using compact open models. Our approach combines reserved-word SQL seeding and persona-based prompting to generate structurally complex queries, while alignment-based filtering preserves difficulty by distinguishing genuinely incorrect examples from complex but valid queries. Human evaluation confirms consistent high quality across model sizes, and our generated data exceeds existing benchmarks in both SQL complexity and natural language richness. Difficulty-stratified analysis shows models trained on FALCON data increasingly outperform baseline-trained models as query complexity increases, validating our pipeline's success in generating challenging training data. When combined with a small proportion of existing benchmark data, mixed training recovers performance on simpler queries while preserving these advantages on complex ones. The model- and database-agnostic design enables organizations to generate high-complexity NL-to-SQL training data locally without external APIs.
Problem

Research questions and friction points this paper is trying to address.

NL2SQL
synthetic data generation
ambiguity
query complexity
relational databases
Innovation

Methods, ideas, or system contributions that make the work stand out.

Synthetic Data Generation
NL2SQL
Alignment-based Filtering
Persona-based Prompting
Model Agnostic
🔎 Similar Papers
No similar papers found.
D
Darian Lee
University of California, Santa Cruz
S
Shannon Rumsey
University of California, Santa Cruz
J
Jack St. Clair
University of California, Santa Cruz
X
Xinyi Tang
University of California, Santa Cruz
A
Aditya Bansal
Adobe
Yuanming Shi
Yuanming Shi
Professor, ShanghaiTech University
Space Computing NetworksEdge Artificial IntelligenceLarge-Scale Optimization