Score
Designs, builds, and evaluates systems that convert natural-language utterances into executable SQL queries—e.g., semantic parsers, intent-to-query mappers, and query generators that produce well-formed SQL and link language tokens to database schema elements. Works on robustness and generalization by handling novel or noisy phrasing, resolving ambiguity, and measuring correctness and executability of the generated queries.
This paper addresses the limitations of traditional pre-trained language models (PLMs) in text-to-SQL tasks under the large language model (LLM) era—namely, poor generalization, high generation error rates, and prohibitive adaptation costs. We systematically survey LLM-driven natural language-to-SQL generation techniques. We propose the first structured, knowledge-graph-inspired survey framework and formally characterize the paradigm shift from PLM fine-tuning to emerging approaches: prompt engineering, retrieval-augmented generation (RAG), database-schema-aware encoding, multi-step reasoning, and in-context learning. We comprehensively catalog mainstream benchmarks, evaluation metrics, and technical challenges, with particular emphasis on critical open issues including scalability and robustness. Our work provides researchers with a clear evolutionary trajectory and practitioners with a reusable technology roadmap and concrete directions for future advancement.
This paper addresses three core challenges in large language model (LLM)-driven Text-to-SQL: low contextual accuracy, brittle schema linking, and constraints on computational efficiency and data privacy. To tackle these, we systematically survey the technical evolution of Text-to-SQL and—first in the literature—rigorously investigate Graph-based Retrieval-Augmented Generation (Graph RAG) for SQL semantic parsing. We propose a unified analytical framework encompassing benchmarking methodologies, evaluation metrics, and key open challenges. Empirical results demonstrate that Graph RAG significantly enhances schema understanding and contextual alignment. Our analysis clarifies the paradigm shift from rule-based approaches to RAG-enhanced methods, explicitly identifying computational efficiency, model robustness, and privacy preservation as the three principal bottlenecks. The work provides both theoretical foundations and practical guidance for developing next-generation Text-to-SQL systems that are trustworthy, interpretable, and highly accurate.
This paper addresses the natural language-to-SQL (NL2SQL) task empowered by large language models (LLMs), providing a systematic survey of its full lifecycle. Methodologically, it establishes a unified analytical framework across four dimensions: model design (schema- and instance-aware modeling), data construction (LLM-driven synthetic data generation), multi-granularity evaluation (spanning syntactic, executional, and semantic correctness), and error attribution (root-cause-driven fine-grained classification analysis). The key contributions are threefold: (1) it introduces, for the first time, an integrated full-lifecycle perspective on NL2SQL in the LLM era; (2) it formulates a development guideline balancing practicality and interpretability; and (3) it identifies core challenges—insufficient schema-aware reasoning, weak few-shot generalization, and poor robustness in real-world deployments—and maps them into a clear problem taxonomy and technical roadmap for future research.
Current text-to-SQL system evaluations rely on a single static database, which fails to capture model robustness across diverse data instances and may introduce significant bias. This work proposes SynSQL, a novel framework that leverages large language models to directly generate semantically consistent and schema-aligned relational test data from natural language questions. SynSQL formulates database construction as a structured generation task governed by semantic and relational constraints, comprising three stages: schema selection, question-guided data synthesis, and constraint-aware iterative refinement. Experiments on Spider, BIRD, and Spider 2.0 demonstrate that databases generated by SynSQL reduce the performance of ten state-of-the-art models by 3–14%, effectively uncovering errors masked by static evaluation and substantially enhancing assessment reliability and stress-testing capability.
Natural language-to-SQL generation faces accuracy bottlenecks due to complex database schemas, ambiguous user intents, and semantic ambiguities. Method: This paper proposes a lightweight, efficient question-augmentation paradigm enabling end-to-end direct schema linking. It explicitly injects schema elements—including tables, columns, values, and conditions—into both the natural language question and SQL generation process; introduces a candidate-predicate augmentation mechanism to enhance semantic alignment for complex queries; and integrates zero-shot, single-turn prompting with large language models (e.g., DeepSeek-Coder-7B-Instruct), combining schema-aware question rewriting and predicate validation. Results: The approach achieves 66.29% execution accuracy on the BIRD benchmark and 56.45% even with small models without fine-tuning—demonstrating that question augmentation substantially improves LLM generalization in text-to-SQL tasks.
Existing NL2SQL benchmarks lack controlled evaluation of model robustness to linguistic variation—semantic equivalence with lexical diversity—leading to an incomplete assessment of language generalization. Method: We propose the first rewriting framework grounded in schema alignment and controllable SQL-to-NL generation, enabling systematic construction of semantically consistent yet lexically diverse test cases for isolated evaluation of linguistic robustness. Contribution/Results: Our approach overcomes the limitation of current benchmarks by introducing principled, controlled linguistic perturbations. Experiments across multiple complexities, domains, and datasets reveal that state-of-the-art models—including LLaMA3.3-70B and LLaMA3.1-8B—suffer substantial performance drops (up to 20% absolute decline in execution accuracy) under surface-form variations, with smaller models exhibiting greater vulnerability. These findings expose a pervasive semantic-representation decoupling deficiency in current NL2SQL systems. Our framework establishes a new benchmark and diagnostic toolkit for trustworthy NL2SQL research.
Text-to-SQL systems frequently generate incorrect SQL queries on complex, real-world databases due to inherent ambiguity in natural language. To address this, we propose an interactive disambiguation framework that models SQL generation as probabilistic inference over multiple candidate queries. At each interaction step, the system dynamically selects the most discriminative clarification question—based on expected information gain—to progressively refine user intent. Our approach integrates probabilistic modeling, natural language processing, and interactive learning, thereby avoiding premature convergence common in single-step generation paradigms. Experiments demonstrate that our method substantially reduces ambiguity and significantly improves execution accuracy across multiple benchmarks, including Spider and CoSQL. The framework enhances both the robustness and practical applicability of semantic parsing in realistic database environments.
In text-to-SQL tasks, lightweight models suffer from low accuracy on complex queries and high inference overhead. This paper proposes an execution-result-guided multi-candidate SQL filtering framework, introducing the first execution-feedback-driven candidate reranking paradigm. Leveraging a lightweight semantic consistency scoring mechanism, it reranks sampled SQL queries based on actual database execution validation—requiring no fine-tuning and enabling plug-and-play adaptation to any SQL generation model. Our method significantly improves semantic correctness and execution accuracy of small models on complex queries. It outperforms large reasoning models—including o1, o3-mini, and DeepSeek R1—across multiple standard benchmarks, while reducing inference cost by up to 30×. To our knowledge, this is the first approach to achieve simultaneous superiority in both accuracy and efficiency for lightweight models in text-to-SQL.
Existing text-to-SQL synthesis methods often conflate executability with semantic correctness, yielding queries that execute successfully yet violate the underlying database semantics. This work proposes the first framework to explicitly model semantic validity within the synthesis pipeline, introducing a modular architecture comprising an analyzer, synthesizer, and validator. These components jointly enable a three-stage reasoning process—semantic parsing, stepwise query synthesis, and diagnostic refinement—transforming execution-based validation into traceable semantic inference. By enforcing semantic consistency throughout generation, the approach substantially outperforms state-of-the-art methods on multiple high-complexity benchmarks and significantly enhances downstream fine-tuning performance.
This work addresses the challenges of natural language to SQL (NL2SQL) translation in real-world enterprise databases, where complex table schemas, opaque column names, dialect heterogeneity, and deeply nested queries hinder performance. To tackle these issues, the authors propose a semantic-layer mediation mechanism that introduces Semantic Model Queries (SMQ) as an intermediate representation, decoupling user intent from physical SQL generation. They further design a constrained think-execute loop and a deterministic compiler to prevent overfitting to the raw database schema. Built upon the Gemini 3 Pro large language model and supporting SQLite, BigQuery, and Snowflake backends, the system achieves a 94.15% execution accuracy on the 547 tasks of Spider2-snow, ranking third on the official leaderboard and significantly outperforming approaches that rely solely on the original schema, thereby substantially enhancing cross-dialect NL2SQL generalization.
This study addresses the persistent challenges of accuracy and robustness in natural language to SQL (NL2SQL) translation under complex query scenarios. The authors systematically evaluate the combined effects of multiple optimization strategies—including the NatSQL intermediate representation, synthetic data preprocessing and fine-tuning, and a novel SQL re-ranking model—using SmBoP and RASAT as backbone architectures. Through ablation studies and Shapley value analysis, they quantitatively assess, for the first time, the interaction effects among these components, revealing that their performance gains are not merely additive. The results demonstrate that non-trivial combinations of these techniques yield significant improvements on benchmarks such as Spider, underscoring the critical role of synergistic interactions among system components.
This work addresses the persistent gap between current natural language to SQL (NL2SQL) systems and human expert performance, which limits their reliable deployment in real-world database applications. To bridge this gap, the authors propose a large language model–based multi-agent framework that enhances generation quality through semantically enriched schema representations, integration of user-defined business rules, and a multi-stage reasoning pipeline. Key innovations include a novel multi-agent coordinator enabling planning, scheduling, and self-reflection mechanisms, as well as a context-aware schema augmentation strategy. Evaluated on the BIRD-SQL benchmark, the proposed approach achieves a semantic accuracy of 78.1%, substantially outperforming existing methods and demonstrating strong cross-domain generalization capabilities.
Natural language to SQL translation in real-world scenarios often suffers from multi-source ambiguities arising from vague user intent and complex database schemas, leading to semantic misalignment and generation failures. This work proposes a fully automated active disambiguation mechanism that first generates candidate SQL queries using synthetic query logs, then identifies conflict types through structured ambiguity categorization. It further designs execution-driven probing queries to automatically gather disambiguation evidence, enabling the selection or repair of the optimal SQL. Evaluated on six public benchmarks, the method achieves an average execution accuracy improvement of 13.0% over state-of-the-art models, with gains as high as 16.7% on highly ambiguous questions, substantially enhancing the system’s generalization capability in handling multi-source ambiguities.