Score
Designs, builds, or analyzes systems that convert queries from one representation or language to another, including natural-language-to-formal-query translation, translation between formal query languages, and query reformulation while preserving intent and semantics. Work includes developing translation algorithms, mappings or grammars, disambiguation and normalization strategies, and integration with query planners or executors to maintain correctness and performance.
This paper addresses the natural language-to-SQL (NL2SQL) task empowered by large language models (LLMs), providing a systematic survey of its full lifecycle. Methodologically, it establishes a unified analytical framework across four dimensions: model design (schema- and instance-aware modeling), data construction (LLM-driven synthetic data generation), multi-granularity evaluation (spanning syntactic, executional, and semantic correctness), and error attribution (root-cause-driven fine-grained classification analysis). The key contributions are threefold: (1) it introduces, for the first time, an integrated full-lifecycle perspective on NL2SQL in the LLM era; (2) it formulates a development guideline balancing practicality and interpretability; and (3) it identifies core challenges—insufficient schema-aware reasoning, weak few-shot generalization, and poor robustness in real-world deployments—and maps them into a clear problem taxonomy and technical roadmap for future research.
Current LLM-based Text-to-SQL approaches suffer from the absence of a unified taxonomic framework, limited cross-model comparability, and insufficient generalization and robustness. To address these issues, this paper proposes the first two-dimensional taxonomy for LLM-based Text-to-SQL methods: one dimension classifies techniques into prompt engineering and parameter fine-tuning; the other categorizes objectives as structure-aware parsing, execution-guided generation, and feedback-enhanced refinement. We systematically synthesize empirical results across major benchmarks—including Spider and WikiSQL—and models such as Codex, LLaMA, and GPT series. Through rigorous literature analysis, methodological abstraction, and cross-method performance comparison, we identify key determinants of prompt design efficacy, delineate the practical boundaries of fine-tuning strategies, and diagnose persistent generalization bottlenecks. Our synthesis distills recurring patterns and evolutionary trends, offering both theoretical foundations and actionable guidelines for developing efficient, robust, and interpretable Text-to-SQL systems.
This paper addresses three core challenges in large language model (LLM)-driven Text-to-SQL: low contextual accuracy, brittle schema linking, and constraints on computational efficiency and data privacy. To tackle these, we systematically survey the technical evolution of Text-to-SQL and—first in the literature—rigorously investigate Graph-based Retrieval-Augmented Generation (Graph RAG) for SQL semantic parsing. We propose a unified analytical framework encompassing benchmarking methodologies, evaluation metrics, and key open challenges. Empirical results demonstrate that Graph RAG significantly enhances schema understanding and contextual alignment. Our analysis clarifies the paradigm shift from rule-based approaches to RAG-enhanced methods, explicitly identifying computational efficiency, model robustness, and privacy preservation as the three principal bottlenecks. The work provides both theoretical foundations and practical guidance for developing next-generation Text-to-SQL systems that are trustworthy, interpretable, and highly accurate.
This paper addresses the failure of conventional equivalence-based query rewriting when original data tables are inaccessible due to access control policies, privacy constraints, or prohibitive retrieval costs. To overcome this, we propose INQURE, a semantic intent-preserving query rewriting framework. Unlike traditional approaches relying on syntactic equivalence and query plan optimization, INQURE introduces, for the first time, large language model (LLM)-driven intent understanding and cross-table reconstruction—enabling semantically consistent rewriting across structurally heterogeneous and non-aligned schemas. The system incorporates pre-filtering of candidate tables, pruning heuristics, and learned ranking to form an end-to-end rewriting pipeline. Evaluated on a benchmark spanning 900+ real-world database schemas, INQURE demonstrates superior rewriting quality and practical utility. A user study further confirms its effective trade-off between execution feasibility and fidelity of analytical insights.
This study addresses the persistent challenges of accuracy and robustness in natural language to SQL (NL2SQL) translation under complex query scenarios. The authors systematically evaluate the combined effects of multiple optimization strategies—including the NatSQL intermediate representation, synthetic data preprocessing and fine-tuning, and a novel SQL re-ranking model—using SmBoP and RASAT as backbone architectures. Through ablation studies and Shapley value analysis, they quantitatively assess, for the first time, the interaction effects among these components, revealing that their performance gains are not merely additive. The results demonstrate that non-trivial combinations of these techniques yield significant improvements on benchmarks such as Spider, underscoring the critical role of synergistic interactions among system components.
SQL is shifting from manual authoring to human-AI co-generation, with humans increasingly assuming roles in verification and debugging—necessitating a new paradigm that decouples query intent from interface-specific syntax. Method: This paper introduces the Abstract Relational Query Language (ARQL) framework, featuring a semantics-first metalinguistic design. It rigorously generalizes tuple relational calculus into abstract relational calculus (ARC) and establishes a unified semantic representation across three modalities: text-based understanding, abstract language trees (ALT), and hierarchical graphs (higraphs). Contribution/Results: ARQL fully decouples query intent, representation modality, and execution conventions—enabling cross-interface intent alignment, formal verifiability, and LLM-era human-AI collaborative querying. This work constitutes the first “Rosetta Stone” for relational query languages, providing both theoretical foundations and practical design principles for multimodal interface-aware relational language engineering.
Natural language-to-SQL generation faces accuracy bottlenecks due to complex database schemas, ambiguous user intents, and semantic ambiguities. Method: This paper proposes a lightweight, efficient question-augmentation paradigm enabling end-to-end direct schema linking. It explicitly injects schema elements—including tables, columns, values, and conditions—into both the natural language question and SQL generation process; introduces a candidate-predicate augmentation mechanism to enhance semantic alignment for complex queries; and integrates zero-shot, single-turn prompting with large language models (e.g., DeepSeek-Coder-7B-Instruct), combining schema-aware question rewriting and predicate validation. Results: The approach achieves 66.29% execution accuracy on the BIRD benchmark and 56.45% even with small models without fine-tuning—demonstrating that question augmentation substantially improves LLM generalization in text-to-SQL tasks.
To address the challenges of poor generalizability and verifiability in low-quality SQL query rewriting, this paper proposes GenRewrite—the first end-to-end LLM-driven query rewriting system. Methodologically, it introduces (1) natural-language rewriting rules (NLR2s) for knowledge representation and cross-query-pattern transfer; (2) a counterexample-guided iterative correction framework that jointly ensures semantic correctness and execution efficiency; and (3) tight integration of SQL syntactic/semantic constraints with LLM reasoning. Evaluated on 99 complex queries from the TPC benchmarks, GenRewrite achieves >2× speedup on 22 queries, improves rewriting coverage by 2.5–3.2× over conventional methods, and outperforms zero-shot LLM baselines by 2.1×.
Existing natural language interfaces to databases lack systematic evaluation frameworks and design theories. This work proposes QUEST, a novel framework that integrates the general-purpose FAR operation schema—Filter, Aggregate, Return—with the W5H semantic dimensions (Who, What, Where, When, Why, How) to enable structured analysis and evaluation of text-to-SQL query semantics. Through semantic annotation and structural parsing, the study validates the universality of the FAR schema across five cross-domain datasets comprising 120,464 queries. The analysis further reveals significant inter-domain disparities in semantic distributions: for instance, medical queries predominantly focus on WHEN and WHO, while WHY and HOW are nearly absent, underscoring a critical challenge for machines in performing deep reasoning over structured data.
This work addresses the challenges posed by the rise of AI-generated queries to the readability and structural explicitness of existing relational query languages. It proposes a unified framework based on Abstract Relational Calculus (ARC) and relational graphs to systematically compare how languages such as SQL, dataframes, and graph query notations express identical query intents. By introducing a formal terminology encompassing information needs, query mappings, and relational schema structures, the study for the first time brings classical database languages and emerging alternatives into a common analytical perspective. The framework is further extended to handle recursive queries, nested relations, and problems beyond PTIME. This contribution establishes a reusable language comparison methodology and a precise design lexicon, offering practical tools for evaluating and designing future relational query languages.
This work addresses the poor performance of large language models (LLMs) on complex, multi-step, and data-dependent Text-to-SQL tasks by proposing a training-free inference framework. The approach employs a lightweight schema selector to prune the database schema and a complexity-aware routing mechanism based on an LLM judge: simple queries are directly translated into SQL, while complex ones are decomposed into atomic subproblems structured as a directed acyclic graph (DAG). These subproblems are then resolved through retrieval-augmented generation (RAG) and topologically optimized for plan-level refinement. Evaluated on the BIRD and Spider benchmarks, the framework achieves execution accuracies of 70.53% and 88.31%, respectively—substantially outperforming existing training-free methods—while reducing inference token consumption by an order of magnitude. Moreover, it functions as a plug-and-play module that enhances the performance of existing SQL generation models.
This work addresses the limitation of traditional database logical design, which overlooks the capacity of large language models (LLMs) to comprehend schema semantics, thereby constraining Text-to-SQL accuracy. For the first time, LLM-friendliness is incorporated into logical schema design through three semantic-preserving and composable transformation strategies: abstraction (+A), workload-aware partitioning (+P), and descriptive renaming (+R). The proposed approach is compatible with both supervised and zero-shot settings, yielding consistent improvements across multiple Text-to-SQL models. Evaluated on the BIRD-Union and Spider-Union benchmarks, the method achieves up to a 4.2% absolute gain in execution accuracy, significantly enhancing the mapping from natural language queries to executable SQL statements.
Existing systems struggle to efficiently translate natural language queries into executable semantic operation pipelines over heterogeneous data sources—such as tables, text, and images—often requiring manual implementation and adaptation of backend APIs, a process that is both tedious and error-prone. This work proposes NL2Pipe, the first middleware system to formalize this task as a compilation problem. NL2Pipe employs a three-stage pipeline—query-data linking, semantic planning, and code generation—to decouple data understanding from backend implementation, enabling unified planning logic to be reused across multiple backends and automatically discovering cross-modal bridging entities. Experimental results demonstrate that NL2Pipe achieves up to a 60% relative improvement in F1 score on complex cross-source analytical tasks, offering a practical, effective solution with controllable latency.