Score
Design, build, and evaluate systems that translate natural-language questions into schema-constrained database queries (e.g., SQL), including intent parsing, schema mapping, identifier resolution, join and projection generation, and constraint enforcement. Implement and test query generation, execution, and result formatting so the system returns structured, cited rows or operational measurements while preventing invalid or unsafe database operations.
This paper addresses three core challenges in large language model (LLM)-driven Text-to-SQL: low contextual accuracy, brittle schema linking, and constraints on computational efficiency and data privacy. To tackle these, we systematically survey the technical evolution of Text-to-SQL and—first in the literature—rigorously investigate Graph-based Retrieval-Augmented Generation (Graph RAG) for SQL semantic parsing. We propose a unified analytical framework encompassing benchmarking methodologies, evaluation metrics, and key open challenges. Empirical results demonstrate that Graph RAG significantly enhances schema understanding and contextual alignment. Our analysis clarifies the paradigm shift from rule-based approaches to RAG-enhanced methods, explicitly identifying computational efficiency, model robustness, and privacy preservation as the three principal bottlenecks. The work provides both theoretical foundations and practical guidance for developing next-generation Text-to-SQL systems that are trustworthy, interpretable, and highly accurate.
This paper addresses the natural language-to-SQL (NL2SQL) task empowered by large language models (LLMs), providing a systematic survey of its full lifecycle. Methodologically, it establishes a unified analytical framework across four dimensions: model design (schema- and instance-aware modeling), data construction (LLM-driven synthetic data generation), multi-granularity evaluation (spanning syntactic, executional, and semantic correctness), and error attribution (root-cause-driven fine-grained classification analysis). The key contributions are threefold: (1) it introduces, for the first time, an integrated full-lifecycle perspective on NL2SQL in the LLM era; (2) it formulates a development guideline balancing practicality and interpretability; and (3) it identifies core challenges—insufficient schema-aware reasoning, weak few-shot generalization, and poor robustness in real-world deployments—and maps them into a clear problem taxonomy and technical roadmap for future research.
This paper addresses the limitations of traditional pre-trained language models (PLMs) in text-to-SQL tasks under the large language model (LLM) era—namely, poor generalization, high generation error rates, and prohibitive adaptation costs. We systematically survey LLM-driven natural language-to-SQL generation techniques. We propose the first structured, knowledge-graph-inspired survey framework and formally characterize the paradigm shift from PLM fine-tuning to emerging approaches: prompt engineering, retrieval-augmented generation (RAG), database-schema-aware encoding, multi-step reasoning, and in-context learning. We comprehensively catalog mainstream benchmarks, evaluation metrics, and technical challenges, with particular emphasis on critical open issues including scalability and robustness. Our work provides researchers with a clear evolutionary trajectory and practitioners with a reusable technology roadmap and concrete directions for future advancement.
Natural language-to-SQL generation faces accuracy bottlenecks due to complex database schemas, ambiguous user intents, and semantic ambiguities. Method: This paper proposes a lightweight, efficient question-augmentation paradigm enabling end-to-end direct schema linking. It explicitly injects schema elements—including tables, columns, values, and conditions—into both the natural language question and SQL generation process; introduces a candidate-predicate augmentation mechanism to enhance semantic alignment for complex queries; and integrates zero-shot, single-turn prompting with large language models (e.g., DeepSeek-Coder-7B-Instruct), combining schema-aware question rewriting and predicate validation. Results: The approach achieves 66.29% execution accuracy on the BIRD benchmark and 56.45% even with small models without fine-tuning—demonstrating that question augmentation substantially improves LLM generalization in text-to-SQL tasks.
Current text-to-SQL system evaluations rely on a single static database, which fails to capture model robustness across diverse data instances and may introduce significant bias. This work proposes SynSQL, a novel framework that leverages large language models to directly generate semantically consistent and schema-aligned relational test data from natural language questions. SynSQL formulates database construction as a structured generation task governed by semantic and relational constraints, comprising three stages: schema selection, question-guided data synthesis, and constraint-aware iterative refinement. Experiments on Spider, BIRD, and Spider 2.0 demonstrate that databases generated by SynSQL reduce the performance of ten state-of-the-art models by 3–14%, effectively uncovering errors masked by static evaluation and substantially enhancing assessment reliability and stress-testing capability.
This work investigates the capability of large language models (LLMs) to determine semantic equivalence between SQL queries, focusing on two critical definitions—semantic equivalence and relaxed equivalence—to enhance the reliability of semantic-level evaluation in text-to-SQL and related generation tasks. We propose a dual-path prompting framework: (1) *Miniature&Mull*, which performs lightweight execution-based verification via counterexample construction; and (2) *Explain&Compare*, which generates natural-language explanations of logical discrepancies and conducts structured syntactic-semantic comparison. To our knowledge, this is the first systematic evaluation of LLMs’ effectiveness and limitations in SQL equivalence judgment without requiring large-scale query execution or human annotations. Experimental results demonstrate that our approach significantly outperforms conventional execution accuracy metrics and achieves reasonable discrimination performance on semantic equivalence tasks. It establishes a novel, interpretable, lightweight, and semantics-aware paradigm for evaluating SQL generation quality.
Automated equivalence checking for complex SQL queries—critical for database education and query optimizer debugging—lacks efficient, reliable solutions. Method: This work pioneers the systematic evaluation of large language models (LLMs) for high-accuracy SQL equivalence judgment without formal proofs. We propose a prompt engineering framework grounded in unoptimized logical query plans, integrating SQL parsing, multi-strategy prompting, synthetic data augmentation, and task-specific supervised fine-tuning to bridge the performance gap between smaller LLMs and GPT-class models. Contribution/Results: Our approach achieves ≈100% accuracy on equivalent SQL pairs and 70% on non-equivalent pairs, while generating human-readable step-by-step reasoning and concrete counterexamples. It extends support beyond restricted SQL subsets—unlike conventional methods—and establishes a practical, interpretable, LLM-driven paradigm for database education and system debugging.
This work addresses the limitation of traditional database logical design, which overlooks the capacity of large language models (LLMs) to comprehend schema semantics, thereby constraining Text-to-SQL accuracy. For the first time, LLM-friendliness is incorporated into logical schema design through three semantic-preserving and composable transformation strategies: abstraction (+A), workload-aware partitioning (+P), and descriptive renaming (+R). The proposed approach is compatible with both supervised and zero-shot settings, yielding consistent improvements across multiple Text-to-SQL models. Evaluated on the BIRD-Union and Spider-Union benchmarks, the method achieves up to a 4.2% absolute gain in execution accuracy, significantly enhancing the mapping from natural language queries to executable SQL statements.
This work addresses the challenge of generating executable Cypher queries from natural language, a task often hindered by outputs that violate syntactic validity or database schema consistency. The authors propose a training-free, test-time structural constraint filtering framework that applies, during inference, a multi-stage post-processing pipeline comprising confidence scoring, context-free grammar validation, and graph database schema consistency checking. This approach explicitly disentangles and quantifies the distinct contributions of syntactic and schema-level constraints to query quality. Experimental results demonstrate significant improvements in both syntactic correctness and execution accuracy across two instruction-tuned models. Specifically, grammar-based filtering markedly enhances syntactic compliance, while schema-aware filtering further boosts semantic correctness, albeit at the cost of reduced coverage under stringent constraints.
Existing text-to-SQL methods exhibit significant performance variance in cross-database generalization, primarily due to the lack of systematic alignment between domain semantics embedded in natural language queries and structural patterns in database schemas, compounded by inefficient, non-generalizable manual prompt engineering for domain knowledge injection. Method: We propose a structured-domain-knowledge-based multi-database text-to-SQL framework that explicitly models domain knowledge as retrievable, structured statements; employs lightweight substring matching for database-adaptive retrieval; and seamlessly integrates retrieved knowledge into the LLM’s reasoning pipeline—eliminating reliance on handcrafted prompts. Contribution/Results: Evaluated across 11 real-world databases and 5 open-source and commercial LLMs, our approach achieves substantial gains in SQL execution accuracy over strong baselines. It is the first to enable plug-and-play cross-database transfer of domain knowledge, markedly improving model robustness in understanding semantic correspondences between domain vocabulary and schema elements.
The impact of database normalization levels (1NF–3NF) on NL2SQL performance remains poorly understood, hindering principled database design for natural-language-driven query systems. Method: We systematically evaluate eight large language models across synthetic and real-world datasets spanning varying normalization levels, measuring SQL generation accuracy under zero-shot and few-shot settings. Contribution/Results: Our empirical study is the first to demonstrate that denormalized schemas improve accuracy for simple retrieval queries, whereas normalized schemas significantly enhance performance on aggregation queries—despite increasing JOIN-related errors. Crucially, few-shot examples effectively mitigate such JOIN errors induced by normalization. Based on these findings, we propose a “query-type–driven adaptive schema selection” strategy, offering actionable guidelines for co-designing database schemas and deploying NL2SQL models. This work bridges database theory and NL2SQL practice, enabling schema-aware model optimization.
This work addresses the challenge of deploying Text-to-SQL evaluation in production environments, where reliance on database schemas and reference SQL queries hinders effective monitoring. To overcome this limitation, the authors propose STEF, a schema-free evaluation framework that requires only the user question, its augmented paraphrase, and the generated SQL. STEF leverages the quality of question augmentation as a core signal, integrating semantic constraint extraction and alignment between natural language and normalized SQL representations. It introduces a composite scoring mechanism incorporating filtered alignment, semantic judgment, and confidence estimation, while supporting injection of application-specific rules via prompt templates. The framework demonstrates robustness to SQL constructs such as GROUP BY, ORDER BY, and LIMIT, enabling—for the first time—continuous, database-independent monitoring and feedback for Text-to-SQL systems in real-world deployments.