Score
Designs and implements systems that generate executable SQL queries constrained by a target database schema by conditioning language models on retrieved schema context and other contextual signals. This work includes retrieving relevant schema candidates (for example via semantic search), selecting or inducing schema fragments for conditioning, prompting or fine‑tuning LLMs to produce syntactically and semantically valid SQL, and verifying query executability against the dataset.
This paper addresses three core challenges in large language model (LLM)-driven Text-to-SQL: low contextual accuracy, brittle schema linking, and constraints on computational efficiency and data privacy. To tackle these, we systematically survey the technical evolution of Text-to-SQL and—first in the literature—rigorously investigate Graph-based Retrieval-Augmented Generation (Graph RAG) for SQL semantic parsing. We propose a unified analytical framework encompassing benchmarking methodologies, evaluation metrics, and key open challenges. Empirical results demonstrate that Graph RAG significantly enhances schema understanding and contextual alignment. Our analysis clarifies the paradigm shift from rule-based approaches to RAG-enhanced methods, explicitly identifying computational efficiency, model robustness, and privacy preservation as the three principal bottlenecks. The work provides both theoretical foundations and practical guidance for developing next-generation Text-to-SQL systems that are trustworthy, interpretable, and highly accurate.
Current LLM-based Text-to-SQL approaches suffer from the absence of a unified taxonomic framework, limited cross-model comparability, and insufficient generalization and robustness. To address these issues, this paper proposes the first two-dimensional taxonomy for LLM-based Text-to-SQL methods: one dimension classifies techniques into prompt engineering and parameter fine-tuning; the other categorizes objectives as structure-aware parsing, execution-guided generation, and feedback-enhanced refinement. We systematically synthesize empirical results across major benchmarks—including Spider and WikiSQL—and models such as Codex, LLaMA, and GPT series. Through rigorous literature analysis, methodological abstraction, and cross-method performance comparison, we identify key determinants of prompt design efficacy, delineate the practical boundaries of fine-tuning strategies, and diagnose persistent generalization bottlenecks. Our synthesis distills recurring patterns and evolutionary trends, offering both theoretical foundations and actionable guidelines for developing efficient, robust, and interpretable Text-to-SQL systems.
This paper addresses the limitations of traditional pre-trained language models (PLMs) in text-to-SQL tasks under the large language model (LLM) era—namely, poor generalization, high generation error rates, and prohibitive adaptation costs. We systematically survey LLM-driven natural language-to-SQL generation techniques. We propose the first structured, knowledge-graph-inspired survey framework and formally characterize the paradigm shift from PLM fine-tuning to emerging approaches: prompt engineering, retrieval-augmented generation (RAG), database-schema-aware encoding, multi-step reasoning, and in-context learning. We comprehensively catalog mainstream benchmarks, evaluation metrics, and technical challenges, with particular emphasis on critical open issues including scalability and robustness. Our work provides researchers with a clear evolutionary trajectory and practitioners with a reusable technology roadmap and concrete directions for future advancement.
This work addresses the limitation of traditional database logical design, which overlooks the capacity of large language models (LLMs) to comprehend schema semantics, thereby constraining Text-to-SQL accuracy. For the first time, LLM-friendliness is incorporated into logical schema design through three semantic-preserving and composable transformation strategies: abstraction (+A), workload-aware partitioning (+P), and descriptive renaming (+R). The proposed approach is compatible with both supervised and zero-shot settings, yielding consistent improvements across multiple Text-to-SQL models. Evaluated on the BIRD-Union and Spider-Union benchmarks, the method achieves up to a 4.2% absolute gain in execution accuracy, significantly enhancing the mapping from natural language queries to executable SQL statements.
This work addresses the challenge that current large language models often fail to generate accurate SQL queries on complex real-world databases due to insufficient understanding of schema structure, semantic ambiguity, and multi-table join paths. To overcome this limitation, the authors propose a two-stage framework: first, a Monte Carlo Tree Search–based strategy autonomously explores the database to construct a structured knowledge base; then, a dual-agent collaborative mechanism leverages this knowledge base to iteratively produce high-quality SQL queries. By decoupling database exploration from SQL generation—a novel approach in this domain—the method significantly enhances the model’s adaptability to unseen databases and its multi-step reasoning capability. Extensive experiments on large-scale benchmarks demonstrate substantial performance gains over strong existing baselines, confirming the effectiveness and generalization ability of the proposed approach.
Natural language-to-SQL generation faces accuracy bottlenecks due to complex database schemas, ambiguous user intents, and semantic ambiguities. Method: This paper proposes a lightweight, efficient question-augmentation paradigm enabling end-to-end direct schema linking. It explicitly injects schema elements—including tables, columns, values, and conditions—into both the natural language question and SQL generation process; introduces a candidate-predicate augmentation mechanism to enhance semantic alignment for complex queries; and integrates zero-shot, single-turn prompting with large language models (e.g., DeepSeek-Coder-7B-Instruct), combining schema-aware question rewriting and predicate validation. Results: The approach achieves 66.29% execution accuracy on the BIRD benchmark and 56.45% even with small models without fine-tuning—demonstrating that question augmentation substantially improves LLM generalization in text-to-SQL tasks.
In text-to-SQL tasks, lightweight models suffer from low accuracy on complex queries and high inference overhead. This paper proposes an execution-result-guided multi-candidate SQL filtering framework, introducing the first execution-feedback-driven candidate reranking paradigm. Leveraging a lightweight semantic consistency scoring mechanism, it reranks sampled SQL queries based on actual database execution validation—requiring no fine-tuning and enabling plug-and-play adaptation to any SQL generation model. Our method significantly improves semantic correctness and execution accuracy of small models on complex queries. It outperforms large reasoning models—including o1, o3-mini, and DeepSeek R1—across multiple standard benchmarks, while reducing inference cost by up to 30×. To our knowledge, this is the first approach to achieve simultaneous superiority in both accuracy and efficiency for lightweight models in text-to-SQL.
This work investigates the capability of large language models (LLMs) to determine semantic equivalence between SQL queries, focusing on two critical definitions—semantic equivalence and relaxed equivalence—to enhance the reliability of semantic-level evaluation in text-to-SQL and related generation tasks. We propose a dual-path prompting framework: (1) *Miniature&Mull*, which performs lightweight execution-based verification via counterexample construction; and (2) *Explain&Compare*, which generates natural-language explanations of logical discrepancies and conducts structured syntactic-semantic comparison. To our knowledge, this is the first systematic evaluation of LLMs’ effectiveness and limitations in SQL equivalence judgment without requiring large-scale query execution or human annotations. Experimental results demonstrate that our approach significantly outperforms conventional execution accuracy metrics and achieves reasonable discrimination performance on semantic equivalence tasks. It establishes a novel, interpretable, lightweight, and semantics-aware paradigm for evaluating SQL generation quality.
Current text-to-SQL system evaluations rely on a single static database, which fails to capture model robustness across diverse data instances and may introduce significant bias. This work proposes SynSQL, a novel framework that leverages large language models to directly generate semantically consistent and schema-aligned relational test data from natural language questions. SynSQL formulates database construction as a structured generation task governed by semantic and relational constraints, comprising three stages: schema selection, question-guided data synthesis, and constraint-aware iterative refinement. Experiments on Spider, BIRD, and Spider 2.0 demonstrate that databases generated by SynSQL reduce the performance of ten state-of-the-art models by 3–14%, effectively uncovering errors masked by static evaluation and substantially enhancing assessment reliability and stress-testing capability.
Existing Text-to-SQL systems struggle with semantic ambiguity and limited scalability in complex enterprise databases due to their reliance on static schema representations. This work proposes APEX-SQL, a novel framework that shifts the Text-to-SQL paradigm from passive translation to active exploration. During schema linking, APEX-SQL integrates logical planning, dual-path pruning, and parallel data profiling to generate hypotheses, which are then validated through global topological synthesis. In the SQL generation phase, it employs a deterministic mechanism to retrieve exploration instructions, thereby enhancing semantic accuracy. The approach significantly improves reasoning capabilities over complex databases, achieving execution accuracies of 70.65% on BIRD and 51.01% on Spider 2.0-Snow—outperforming current baselines while reducing token consumption.
Existing text-to-SQL synthesis methods often conflate executability with semantic correctness, yielding queries that execute successfully yet violate the underlying database semantics. This work proposes the first framework to explicitly model semantic validity within the synthesis pipeline, introducing a modular architecture comprising an analyzer, synthesizer, and validator. These components jointly enable a three-stage reasoning process—semantic parsing, stepwise query synthesis, and diagnostic refinement—transforming execution-based validation into traceable semantic inference. By enforcing semantic consistency throughout generation, the approach substantially outperforms state-of-the-art methods on multiple high-complexity benchmarks and significantly enhances downstream fine-tuning performance.
This study addresses the lack of robustness in large language models (LLMs) on text-to-SQL tasks when confronted with relationally equivalent yet structurally diverse database schemas—a critical dimension overlooked by existing evaluations. The authors propose the first evaluation framework grounded in a unified Entity-Relationship (E/R) model, which employs controlled “splitting” strategies to generate diverse but semantically equivalent schema variants. By preserving both natural language questions and underlying data, the framework systematically assesses the consistency of model outputs across schema transformations. To enhance robustness, the work innovatively incorporates E/R specifications as additional contextual input and integrates a candidate query generation mechanism with pairwise comparison heatmaps for fine-grained analysis. Experimental results demonstrate that schema structure significantly impacts model performance, often yielding inconsistent SQL predictions; while E/R context partially mitigates this issue, it does not fully eliminate the sensitivity to structural variation.
This work addresses the challenge of generating executable Cypher queries from natural language, a task often hindered by outputs that violate syntactic validity or database schema consistency. The authors propose a training-free, test-time structural constraint filtering framework that applies, during inference, a multi-stage post-processing pipeline comprising confidence scoring, context-free grammar validation, and graph database schema consistency checking. This approach explicitly disentangles and quantifies the distinct contributions of syntactic and schema-level constraints to query quality. Experimental results demonstrate significant improvements in both syntactic correctness and execution accuracy across two instruction-tuned models. Specifically, grammar-based filtering markedly enhances syntactic compliance, while schema-aware filtering further boosts semantic correctness, albeit at the cost of reduced coverage under stringent constraints.