🤖 AI Summary
Existing NL2SQL benchmarks lack controlled evaluation of model robustness to linguistic variation—semantic equivalence with lexical diversity—leading to an incomplete assessment of language generalization. Method: We propose the first rewriting framework grounded in schema alignment and controllable SQL-to-NL generation, enabling systematic construction of semantically consistent yet lexically diverse test cases for isolated evaluation of linguistic robustness. Contribution/Results: Our approach overcomes the limitation of current benchmarks by introducing principled, controlled linguistic perturbations. Experiments across multiple complexities, domains, and datasets reveal that state-of-the-art models—including LLaMA3.3-70B and LLaMA3.1-8B—suffer substantial performance drops (up to 20% absolute decline in execution accuracy) under surface-form variations, with smaller models exhibiting greater vulnerability. These findings expose a pervasive semantic-representation decoupling deficiency in current NL2SQL systems. Our framework establishes a new benchmark and diagnostic toolkit for trustworthy NL2SQL research.
📝 Abstract
Robust evaluation in the presence of linguistic variation is key to understanding the generalization capabilities of Natural Language to SQL (NL2SQL) models, yet existing benchmarks rarely address this factor in a systematic or controlled manner. We propose a novel schema-aligned paraphrasing framework that leverages SQL-to-NL (SQL2NL) to automatically generate semantically equivalent, lexically diverse queries while maintaining alignment with the original schema and intent. This enables the first targeted evaluation of NL2SQL robustness to linguistic variation in isolation-distinct from prior work that primarily investigates ambiguity or schema perturbations. Our analysis reveals that state-of-the-art models are far more brittle than standard benchmarks suggest. For example, LLaMa3.3-70B exhibits a 10.23% drop in execution accuracy (from 77.11% to 66.9%) on paraphrased Spider queries, while LLaMa3.1-8B suffers an even larger drop of nearly 20% (from 62.9% to 42.5%). Smaller models (e.g., GPT-4o mini) are disproportionately affected. We also find that robustness degradation varies significantly with query complexity, dataset, and domain -- highlighting the need for evaluation frameworks that explicitly measure linguistic generalization to ensure reliable performance in real-world settings.