Score
Designs, builds, or evaluates systems that translate freeform natural-language utterances into syntactically correct, executable queries in formal query languages, producing concrete query strings or templates. Works focus on intent interpretation and disambiguation using available context so the generated queries reflect the user’s desired semantics and execute without syntax or runtime errors.
This paper addresses the natural language-to-SQL (NL2SQL) task empowered by large language models (LLMs), providing a systematic survey of its full lifecycle. Methodologically, it establishes a unified analytical framework across four dimensions: model design (schema- and instance-aware modeling), data construction (LLM-driven synthetic data generation), multi-granularity evaluation (spanning syntactic, executional, and semantic correctness), and error attribution (root-cause-driven fine-grained classification analysis). The key contributions are threefold: (1) it introduces, for the first time, an integrated full-lifecycle perspective on NL2SQL in the LLM era; (2) it formulates a development guideline balancing practicality and interpretability; and (3) it identifies core challenges—insufficient schema-aware reasoning, weak few-shot generalization, and poor robustness in real-world deployments—and maps them into a clear problem taxonomy and technical roadmap for future research.
This paper addresses three core challenges in large language model (LLM)-driven Text-to-SQL: low contextual accuracy, brittle schema linking, and constraints on computational efficiency and data privacy. To tackle these, we systematically survey the technical evolution of Text-to-SQL and—first in the literature—rigorously investigate Graph-based Retrieval-Augmented Generation (Graph RAG) for SQL semantic parsing. We propose a unified analytical framework encompassing benchmarking methodologies, evaluation metrics, and key open challenges. Empirical results demonstrate that Graph RAG significantly enhances schema understanding and contextual alignment. Our analysis clarifies the paradigm shift from rule-based approaches to RAG-enhanced methods, explicitly identifying computational efficiency, model robustness, and privacy preservation as the three principal bottlenecks. The work provides both theoretical foundations and practical guidance for developing next-generation Text-to-SQL systems that are trustworthy, interpretable, and highly accurate.
This paper addresses the limitations of traditional pre-trained language models (PLMs) in text-to-SQL tasks under the large language model (LLM) era—namely, poor generalization, high generation error rates, and prohibitive adaptation costs. We systematically survey LLM-driven natural language-to-SQL generation techniques. We propose the first structured, knowledge-graph-inspired survey framework and formally characterize the paradigm shift from PLM fine-tuning to emerging approaches: prompt engineering, retrieval-augmented generation (RAG), database-schema-aware encoding, multi-step reasoning, and in-context learning. We comprehensively catalog mainstream benchmarks, evaluation metrics, and technical challenges, with particular emphasis on critical open issues including scalability and robustness. Our work provides researchers with a clear evolutionary trajectory and practitioners with a reusable technology roadmap and concrete directions for future advancement.
This work addresses the challenge of effectively answering natural language questions over large-scale codebases, a task hindered by the limited semantic reasoning capabilities of existing approaches and the context-length and computational constraints of large language models (LLMs). To overcome these limitations, the authors propose Merlin, a system that integrates LLMs with the CodeQL program analysis framework to automatically translate natural language queries into executable code queries. Merlin’s key innovations include a retrieval-augmented generation (RAG)-based iterative query synthesis mechanism and a novel self-testing technique that employs auxiliary queries to generate concrete evidence, thereby identifying and correcting semantic flaws in candidate queries. Experimental results demonstrate that Merlin not only reproduces most vulnerabilities found by prior methods but also uncovers previously missed issues. User studies further show that Merlin improves task accuracy by 3.8× and reduces completion time by 31%.
In resource-constrained settings, conventional RAG systems struggle to accurately identify the nested structural intent of complex queries. Method: This paper proposes a neuro-symbolic query compilation framework. Its core components are: (1) a minimal, complete BNF grammar ( G[q] ) that formally encodes the syntactic structure of complex queries; (2) a tripartite neuro-symbolic compilation pipeline—comprising query expression translation, lexical and syntactic parsing, and recursive-descent processing—that automatically compiles natural language queries into executable abstract syntax trees (ASTs); and (3) grammar-driven intent parsing coupled with semantics-preserving retrieval augmentation. Results: Experiments demonstrate substantial improvements in document retrieval accuracy and response generation quality for complex queries. The framework achieves high robustness, strong interpretability, and lightweight deployability under low-resource conditions.
This study investigates the necessity of large language models (LLMs) for natural language to SQL (NL2SQL) tasks and challenges the prevailing overestimation of SQL query complexity. Through empirical analysis of 376 real-world databases, the work reveals— for the first time—that SQL query templates follow a power-law-like distribution: merely 13% of distinct templates account for 70% of all queries. Leveraging large-scale database sampling, SQL template abstraction, frequency statistics, and natural language–SQL alignment techniques, the authors demonstrate that query complexity does not monotonically increase with the number of tables and that the majority of queries are highly predictable. These findings underscore the advantages of template-based approaches in terms of security, cost-efficiency, and auditability, thereby questioning the dominant paradigm that relies on complex LLMs for NL2SQL translation.
Natural language-to-SQL generation faces accuracy bottlenecks due to complex database schemas, ambiguous user intents, and semantic ambiguities. Method: This paper proposes a lightweight, efficient question-augmentation paradigm enabling end-to-end direct schema linking. It explicitly injects schema elements—including tables, columns, values, and conditions—into both the natural language question and SQL generation process; introduces a candidate-predicate augmentation mechanism to enhance semantic alignment for complex queries; and integrates zero-shot, single-turn prompting with large language models (e.g., DeepSeek-Coder-7B-Instruct), combining schema-aware question rewriting and predicate validation. Results: The approach achieves 66.29% execution accuracy on the BIRD benchmark and 56.45% even with small models without fine-tuning—demonstrating that question augmentation substantially improves LLM generalization in text-to-SQL tasks.
To address the challenges of poor generalizability and verifiability in low-quality SQL query rewriting, this paper proposes GenRewrite—the first end-to-end LLM-driven query rewriting system. Methodologically, it introduces (1) natural-language rewriting rules (NLR2s) for knowledge representation and cross-query-pattern transfer; (2) a counterexample-guided iterative correction framework that jointly ensures semantic correctness and execution efficiency; and (3) tight integration of SQL syntactic/semantic constraints with LLM reasoning. Evaluated on 99 complex queries from the TPC benchmarks, GenRewrite achieves >2× speedup on 22 queries, improves rewriting coverage by 2.5–3.2× over conventional methods, and outperforms zero-shot LLM baselines by 2.1×.
Existing systems struggle to efficiently translate natural language queries into executable semantic operation pipelines over heterogeneous data sources—such as tables, text, and images—often requiring manual implementation and adaptation of backend APIs, a process that is both tedious and error-prone. This work proposes NL2Pipe, the first middleware system to formalize this task as a compilation problem. NL2Pipe employs a three-stage pipeline—query-data linking, semantic planning, and code generation—to decouple data understanding from backend implementation, enabling unified planning logic to be reused across multiple backends and automatically discovering cross-modal bridging entities. Experimental results demonstrate that NL2Pipe achieves up to a 60% relative improvement in F1 score on complex cross-source analytical tasks, offering a practical, effective solution with controllable latency.
This work addresses the challenge of generating executable Cypher queries from natural language, a task often hindered by outputs that violate syntactic validity or database schema consistency. The authors propose a training-free, test-time structural constraint filtering framework that applies, during inference, a multi-stage post-processing pipeline comprising confidence scoring, context-free grammar validation, and graph database schema consistency checking. This approach explicitly disentangles and quantifies the distinct contributions of syntactic and schema-level constraints to query quality. Experimental results demonstrate significant improvements in both syntactic correctness and execution accuracy across two instruction-tuned models. Specifically, grammar-based filtering markedly enhances syntactic compliance, while schema-aware filtering further boosts semantic correctness, albeit at the cost of reduced coverage under stringent constraints.
This work addresses the poor performance of large language models (LLMs) on complex, multi-step, and data-dependent Text-to-SQL tasks by proposing a training-free inference framework. The approach employs a lightweight schema selector to prune the database schema and a complexity-aware routing mechanism based on an LLM judge: simple queries are directly translated into SQL, while complex ones are decomposed into atomic subproblems structured as a directed acyclic graph (DAG). These subproblems are then resolved through retrieval-augmented generation (RAG) and topologically optimized for plan-level refinement. Evaluated on the BIRD and Spider benchmarks, the framework achieves execution accuracies of 70.53% and 88.31%, respectively—substantially outperforming existing training-free methods—while reducing inference token consumption by an order of magnitude. Moreover, it functions as a plug-and-play module that enhances the performance of existing SQL generation models.
Existing approaches struggle to reliably execute cross-application natural language queries due to their reliance on local data and lack of structured planning. This work proposes a novel framework that integrates large language models (LLMs) with deterministic execution: it first leverages an LLM to parse user queries into structured query graphs, then employs depth-first search for deterministic planning, explicitly modeling inter-tool dependencies and fusing results from multiple sources. The approach substantially enhances both the expressiveness and execution reliability of complex queries, achieving high accuracy even with small or locally deployed LLMs, thereby demonstrating its effectiveness and practical utility.
Natural language to SQL translation in real-world scenarios often suffers from multi-source ambiguities arising from vague user intent and complex database schemas, leading to semantic misalignment and generation failures. This work proposes a fully automated active disambiguation mechanism that first generates candidate SQL queries using synthetic query logs, then identifies conflict types through structured ambiguity categorization. It further designs execution-driven probing queries to automatically gather disambiguation evidence, enabling the selection or repair of the optimal SQL. Evaluated on six public benchmarks, the method achieves an average execution accuracy improvement of 13.0% over state-of-the-art models, with gains as high as 16.7% on highly ambiguous questions, substantially enhancing the system’s generalization capability in handling multi-source ambiguities.