Score
Designs, implements, and evaluates parsing systems that translate natural-language utterances into structured query languages (such as SQL), producing executable queries from user text. Work covers model and decoder architectures, grammar and schema mapping, dataset creation/annotation for NL2SQL, and analysis of correctness, robustness, and execution fidelity of generated queries.
This paper addresses the natural language-to-SQL (NL2SQL) task empowered by large language models (LLMs), providing a systematic survey of its full lifecycle. Methodologically, it establishes a unified analytical framework across four dimensions: model design (schema- and instance-aware modeling), data construction (LLM-driven synthetic data generation), multi-granularity evaluation (spanning syntactic, executional, and semantic correctness), and error attribution (root-cause-driven fine-grained classification analysis). The key contributions are threefold: (1) it introduces, for the first time, an integrated full-lifecycle perspective on NL2SQL in the LLM era; (2) it formulates a development guideline balancing practicality and interpretability; and (3) it identifies core challenges—insufficient schema-aware reasoning, weak few-shot generalization, and poor robustness in real-world deployments—and maps them into a clear problem taxonomy and technical roadmap for future research.
Current LLM-based Text-to-SQL approaches suffer from the absence of a unified taxonomic framework, limited cross-model comparability, and insufficient generalization and robustness. To address these issues, this paper proposes the first two-dimensional taxonomy for LLM-based Text-to-SQL methods: one dimension classifies techniques into prompt engineering and parameter fine-tuning; the other categorizes objectives as structure-aware parsing, execution-guided generation, and feedback-enhanced refinement. We systematically synthesize empirical results across major benchmarks—including Spider and WikiSQL—and models such as Codex, LLaMA, and GPT series. Through rigorous literature analysis, methodological abstraction, and cross-method performance comparison, we identify key determinants of prompt design efficacy, delineate the practical boundaries of fine-tuning strategies, and diagnose persistent generalization bottlenecks. Our synthesis distills recurring patterns and evolutionary trends, offering both theoretical foundations and actionable guidelines for developing efficient, robust, and interpretable Text-to-SQL systems.
This paper addresses three key challenges in natural language interfaces for database querying (NLIDBs): the lack of a unified evaluation framework, the absence of a standardized translation paradigm, and fragmented technical approaches. Methodologically, it proposes a holistic analytical framework comprising three stages—preprocessing, semantic understanding, and query generation—and introduces, for the first time, a three-level translation paradigm applicable to both relational and spatiotemporal databases. It systematically integrates diverse techniques—including dependency parsing, named entity recognition, word embeddings, schema alignment, rule-based engines, supervised/unsupervised learning, and large language models (LLMs)—while unifying LLM-enhanced Text-to-SQL, SQL-to-Text, and speech-to-SQL directions. Contributions include: (1) establishing a standardized evaluation framework; (2) surveying and categorizing mainstream benchmarks and metrics; (3) proposing a principled methodology for constructing novel benchmarks; and (4) identifying two critical research frontiers—deep semantic understanding and dynamic database interaction.
To address the challenge non-technical users face in directly querying large-scale semi-structured time-series data (e.g., logs, telemetry), this paper introduces the first natural-language-to-Kusto-Query-Language (NL2KQL) framework. Methodologically, it proposes a three-module协同 architecture—Schema Refiner, Dynamic Few-shot Selector, and Query Refiner—that jointly integrates LLM-based semantic parsing, schema refinement, context-aware few-shot retrieval, and KQL syntax/semantic error correction. We construct the first open-source, contextually grounded synthetic NLQ–KQL benchmark dataset and support multi-dimensional evaluation at both execution-level and parsing-level. Experiments on real-world Kusto deployments demonstrate a 32.7% improvement in execution accuracy over state-of-the-art baselines; ablation studies confirm the significant contribution of each module. All code, datasets, and evaluation tools are publicly released.
Natural language-to-SQL generation faces accuracy bottlenecks due to complex database schemas, ambiguous user intents, and semantic ambiguities. Method: This paper proposes a lightweight, efficient question-augmentation paradigm enabling end-to-end direct schema linking. It explicitly injects schema elements—including tables, columns, values, and conditions—into both the natural language question and SQL generation process; introduces a candidate-predicate augmentation mechanism to enhance semantic alignment for complex queries; and integrates zero-shot, single-turn prompting with large language models (e.g., DeepSeek-Coder-7B-Instruct), combining schema-aware question rewriting and predicate validation. Results: The approach achieves 66.29% execution accuracy on the BIRD benchmark and 56.45% even with small models without fine-tuning—demonstrating that question augmentation substantially improves LLM generalization in text-to-SQL tasks.
Large language models (LLMs) struggle to accurately model the alignment between user intent and database schema in natural language-to-SQL (NL2SQL) translation, leading to high error rates. To address this, we propose SGU-SQL—a novel framework introducing a decoupled, stepwise generation paradigm that separately handles schema grounding and syntactic tree decomposition. SGU-SQL integrates structure-enhanced query-schema alignment, syntax-tree-guided LLM decoding, and multi-stage structure-aware prompting and fine-tuning to achieve precise semantic-to-syntactic mapping. Evaluated on two prominent benchmarks—Spider and BIRD—SGU-SQL outperforms 16 state-of-the-art baselines, achieving new SOTA performance in both execution accuracy and cross-domain generalization. Our results empirically validate that explicit structural modeling is critical for advancing NL2SQL systems.
To address the low parsing efficiency, opaque interaction, and poor generalization of large language models (LLMs) in Text-to-SQL for wide-table scenarios, this paper proposes Interactive-T2S—a framework enabling iterative human-AI collaboration wherein the LLM directly interacts with the database to generate SQL progressively and transparently. Its core contributions are: (1) four generic, schema-agnostic database interaction tools that support cross-schema generalization; and (2) a structured, example-driven stepwise reasoning paradigm integrating dynamic context construction and chain-of-thought prompting. Evaluated on Spider and BIRD (including its variants), Interactive-T2S achieves new state-of-the-art performance on the BIRD leaderboard under the non-oracle setting, significantly improving both query accuracy on wide-table schemas and the interpretability of human-system interaction.
This paper addresses three core challenges in large language model (LLM)-driven Text-to-SQL: low contextual accuracy, brittle schema linking, and constraints on computational efficiency and data privacy. To tackle these, we systematically survey the technical evolution of Text-to-SQL and—first in the literature—rigorously investigate Graph-based Retrieval-Augmented Generation (Graph RAG) for SQL semantic parsing. We propose a unified analytical framework encompassing benchmarking methodologies, evaluation metrics, and key open challenges. Empirical results demonstrate that Graph RAG significantly enhances schema understanding and contextual alignment. Our analysis clarifies the paradigm shift from rule-based approaches to RAG-enhanced methods, explicitly identifying computational efficiency, model robustness, and privacy preservation as the three principal bottlenecks. The work provides both theoretical foundations and practical guidance for developing next-generation Text-to-SQL systems that are trustworthy, interpretable, and highly accurate.
This study addresses the persistent challenges of accuracy and robustness in natural language to SQL (NL2SQL) translation under complex query scenarios. The authors systematically evaluate the combined effects of multiple optimization strategies—including the NatSQL intermediate representation, synthetic data preprocessing and fine-tuning, and a novel SQL re-ranking model—using SmBoP and RASAT as backbone architectures. Through ablation studies and Shapley value analysis, they quantitatively assess, for the first time, the interaction effects among these components, revealing that their performance gains are not merely additive. The results demonstrate that non-trivial combinations of these techniques yield significant improvements on benchmarks such as Spider, underscoring the critical role of synergistic interactions among system components.
This work addresses the challenge of generating executable Cypher queries from natural language, a task often hindered by outputs that violate syntactic validity or database schema consistency. The authors propose a training-free, test-time structural constraint filtering framework that applies, during inference, a multi-stage post-processing pipeline comprising confidence scoring, context-free grammar validation, and graph database schema consistency checking. This approach explicitly disentangles and quantifies the distinct contributions of syntactic and schema-level constraints to query quality. Experimental results demonstrate significant improvements in both syntactic correctness and execution accuracy across two instruction-tuned models. Specifically, grammar-based filtering markedly enhances syntactic compliance, while schema-aware filtering further boosts semantic correctness, albeit at the cost of reduced coverage under stringent constraints.
Natural language log querying remains challenging due to the absence of structured schemas, hindering accurate SQL generation. This work proposes a novel approach that first parses raw logs into templated relational tables and then enriches both templates and parameter columns with interpretable semantics through dual-granularity semantic grounding. By integrating semantic search with constrained decoding in large language models, the method generates context-aware, executable SQL queries. The study introduces the first semantically grounded log schema and releases LogNLQ-Bench, the inaugural benchmark for natural language log querying featuring execution-based validation. Experimental results demonstrate that the proposed method significantly outperforms existing techniques on LogNLQ-Bench, particularly excelling in complex analytical queries.
This work addresses the persistent gap between current natural language to SQL (NL2SQL) systems and human expert performance, which limits their reliable deployment in real-world database applications. To bridge this gap, the authors propose a large language model–based multi-agent framework that enhances generation quality through semantically enriched schema representations, integration of user-defined business rules, and a multi-stage reasoning pipeline. Key innovations include a novel multi-agent coordinator enabling planning, scheduling, and self-reflection mechanisms, as well as a context-aware schema augmentation strategy. Evaluated on the BIRD-SQL benchmark, the proposed approach achieves a semantic accuracy of 78.1%, substantially outperforming existing methods and demonstrating strong cross-domain generalization capabilities.