semantic parsing for structured queries

Designs, implements, and evaluates parsing systems that translate natural-language utterances into structured query languages (such as SQL), producing executable queries from user text. Work covers model and decoder architectures, grammar and schema mapping, dataset creation/annotation for NL2SQL, and analysis of correctness, robustness, and execution fidelity of generated queries.

semanticparsingforstructured

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.33
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

A Survey on Employing Large Language Models for Text-to-SQL Tasks

Jul 21, 2024
LS
Liang Shi
🏛️ Peking University | SINGDATA CLOUD PTE. LTD

Current LLM-based Text-to-SQL approaches suffer from the absence of a unified taxonomic framework, limited cross-model comparability, and insufficient generalization and robustness. To address these issues, this paper proposes the first two-dimensional taxonomy for LLM-based Text-to-SQL methods: one dimension classifies techniques into prompt engineering and parameter fine-tuning; the other categorizes objectives as structure-aware parsing, execution-guided generation, and feedback-enhanced refinement. We systematically synthesize empirical results across major benchmarks—including Spider and WikiSQL—and models such as Codex, LLaMA, and GPT series. Through rigorous literature analysis, methodological abstraction, and cross-method performance comparison, we identify key determinants of prompt design efficacy, delineate the practical boundaries of fine-tuning strategies, and diagnose persistent generalization bottlenecks. Our synthesis distills recurring patterns and evolutionary trends, offering both theoretical foundations and actionable guidelines for developing efficient, robust, and interpretable Text-to-SQL systems.

Analyzing prompt engineering and finetuning approachesDiscussing challenges and future research directionsReviewing LLM-based Text-to-SQL methods comprehensively

NLI4DB: A Systematic Review of Natural Language Interfaces for Databases

Mar 04, 2025
ML
Mengyi Liu
🏛️ Nanjing University of Aeronautics and Astronautics

This paper addresses three key challenges in natural language interfaces for database querying (NLIDBs): the lack of a unified evaluation framework, the absence of a standardized translation paradigm, and fragmented technical approaches. Methodologically, it proposes a holistic analytical framework comprising three stages—preprocessing, semantic understanding, and query generation—and introduces, for the first time, a three-level translation paradigm applicable to both relational and spatiotemporal databases. It systematically integrates diverse techniques—including dependency parsing, named entity recognition, word embeddings, schema alignment, rule-based engines, supervised/unsupervised learning, and large language models (LLMs)—while unifying LLM-enhanced Text-to-SQL, SQL-to-Text, and speech-to-SQL directions. Contributions include: (1) establishing a standardized evaluation framework; (2) surveying and categorizing mainstream benchmarks and metrics; (3) proposing a principled methodology for constructing novel benchmarks; and (4) identifying two critical research frontiers—deep semantic understanding and dynamic database interaction.

Evaluating benchmarks and techniques for NLIDB system enhancement.Exploring translation from natural language to executable database language.Surveying natural language interfaces for database querying.

Must-Read Papers

Most classic and influential ideas
View more

NL2KQL: From Natural Language to Kusto Query

Apr 03, 2024
AH
Amir H. Abdi
🏛️ Microsoft

To address the challenge non-technical users face in directly querying large-scale semi-structured time-series data (e.g., logs, telemetry), this paper introduces the first natural-language-to-Kusto-Query-Language (NL2KQL) framework. Methodologically, it proposes a three-module协同 architecture—Schema Refiner, Dynamic Few-shot Selector, and Query Refiner—that jointly integrates LLM-based semantic parsing, schema refinement, context-aware few-shot retrieval, and KQL syntax/semantic error correction. We construct the first open-source, contextually grounded synthetic NLQ–KQL benchmark dataset and support multi-dimensional evaluation at both execution-level and parsing-level. Experiments on real-world Kusto deployments demonstrate a 32.7% improvement in execution accuracy over state-of-the-art baselines; ablation studies confirm the significant contribution of each module. All code, datasets, and evaluation tools are publicly released.

Data RetrievalKusto Query LanguageNatural Language Processing

E-SQL: Direct Schema Linking via Question Enrichment in Text-to-SQL

Sep 25, 2024
HA
Hasan Alp Caferoglu
🏛️ Bilkent University

Natural language-to-SQL generation faces accuracy bottlenecks due to complex database schemas, ambiguous user intents, and semantic ambiguities. Method: This paper proposes a lightweight, efficient question-augmentation paradigm enabling end-to-end direct schema linking. It explicitly injects schema elements—including tables, columns, values, and conditions—into both the natural language question and SQL generation process; introduces a candidate-predicate augmentation mechanism to enhance semantic alignment for complex queries; and integrates zero-shot, single-turn prompting with large language models (e.g., DeepSeek-Coder-7B-Instruct), combining schema-aware question rewriting and predicate validation. Results: The approach achieves 66.29% execution accuracy on the BIRD benchmark and 56.45% even with small models without fine-tuning—demonstrating that question augmentation substantially improves LLM generalization in text-to-SQL tasks.

Complex SQL GenerationDatabase QueryingNatural Language Processing

Structure Guided Large Language Model for SQL Generation

Feb 19, 2024
QZ
Qinggang Zhang
🏛️ The Hong Kong Polytechnic University | Jinan University

Large language models (LLMs) struggle to accurately model the alignment between user intent and database schema in natural language-to-SQL (NL2SQL) translation, leading to high error rates. To address this, we propose SGU-SQL—a novel framework introducing a decoupled, stepwise generation paradigm that separately handles schema grounding and syntactic tree decomposition. SGU-SQL integrates structure-enhanced query-schema alignment, syntax-tree-guided LLM decoding, and multi-stage structure-aware prompting and fine-tuning to achieve precise semantic-to-syntactic mapping. Evaluated on two prominent benchmarks—Spider and BIRD—SGU-SQL outperforms 16 state-of-the-art baselines, achieving new SOTA performance in both execution accuracy and cross-domain generalization. Our results empirically validate that explicit structural modeling is critical for advancing NL2SQL systems.

Decomposing SQL generation tasks using syntax-based promptingEnhancing LLM comprehension of complex database structuresImproving SQL generation from natural language queries

Multi-Turn Interactions for Text-to-SQL with Large Language Models

Aug 09, 2024
GX
Guanming Xiong
🏛️ Peking University | Zuoyebang Education Technology Co., Ltd.

To address the low parsing efficiency, opaque interaction, and poor generalization of large language models (LLMs) in Text-to-SQL for wide-table scenarios, this paper proposes Interactive-T2S—a framework enabling iterative human-AI collaboration wherein the LLM directly interacts with the database to generate SQL progressively and transparently. Its core contributions are: (1) four generic, schema-agnostic database interaction tools that support cross-schema generalization; and (2) a structured, example-driven stepwise reasoning paradigm integrating dynamic context construction and chain-of-thought prompting. Evaluated on Spider and BIRD (including its variants), Interactive-T2S achieves new state-of-the-art performance on the BIRD leaderboard under the non-oracle setting, significantly improving both query accuracy on wide-table schemas and the interpretability of human-system interaction.

Addresses inefficient SQL generation with wide tablesCreates universally applicable database interaction frameworkProvides step-by-step interpretable SQL query generation

From Natural Language to SQL: Review of LLM-based Text-to-SQL Systems

Oct 01, 2024
AM
Ali Mohammadjafari
🏛️ University of Louisiana at Lafayette

This paper addresses three core challenges in large language model (LLM)-driven Text-to-SQL: low contextual accuracy, brittle schema linking, and constraints on computational efficiency and data privacy. To tackle these, we systematically survey the technical evolution of Text-to-SQL and—first in the literature—rigorously investigate Graph-based Retrieval-Augmented Generation (Graph RAG) for SQL semantic parsing. We propose a unified analytical framework encompassing benchmarking methodologies, evaluation metrics, and key open challenges. Empirical results demonstrate that Graph RAG significantly enhances schema understanding and contextual alignment. Our analysis clarifies the paradigm shift from rule-based approaches to RAG-enhanced methods, explicitly identifying computational efficiency, model robustness, and privacy preservation as the three principal bottlenecks. The work provides both theoretical foundations and practical guidance for developing next-generation Text-to-SQL systems that are trustworthy, interpretable, and highly accurate.

Addressing computational and privacy challengesExploring LLM evolution in SQL systemsImproving SQL translation accuracy

Latest Papers

What's happening recently
View more

This study addresses the persistent challenges of accuracy and robustness in natural language to SQL (NL2SQL) translation under complex query scenarios. The authors systematically evaluate the combined effects of multiple optimization strategies—including the NatSQL intermediate representation, synthetic data preprocessing and fine-tuning, and a novel SQL re-ranking model—using SmBoP and RASAT as backbone architectures. Through ablation studies and Shapley value analysis, they quantitatively assess, for the first time, the interaction effects among these components, revealing that their performance gains are not merely additive. The results demonstrate that non-trivial combinations of these techniques yield significant improvements on benchmarks such as Spider, underscoring the critical role of synergistic interactions among system components.

large language modelsmodel pipelineNatural Language to SQL

This work addresses the challenge of generating executable Cypher queries from natural language, a task often hindered by outputs that violate syntactic validity or database schema consistency. The authors propose a training-free, test-time structural constraint filtering framework that applies, during inference, a multi-stage post-processing pipeline comprising confidence scoring, context-free grammar validation, and graph database schema consistency checking. This approach explicitly disentangles and quantifies the distinct contributions of syntactic and schema-level constraints to query quality. Experimental results demonstrate significant improvements in both syntactic correctness and execution accuracy across two instruction-tuned models. Specifically, grammar-based filtering markedly enhances syntactic compliance, while schema-aware filtering further boosts semantic correctness, albeit at the cost of reduced coverage under stringent constraints.

query reliabilityschema consistencystructural constraints

Natural language log querying remains challenging due to the absence of structured schemas, hindering accurate SQL generation. This work proposes a novel approach that first parses raw logs into templated relational tables and then enriches both templates and parameter columns with interpretable semantics through dual-granularity semantic grounding. By integrating semantic search with constrained decoding in large language models, the method generates context-aware, executable SQL queries. The study introduces the first semantically grounded log schema and releases LogNLQ-Bench, the inaugural benchmark for natural language log querying featuring execution-based validation. Experimental results demonstrate that the proposed method significantly outperforms existing techniques on LogNLQ-Bench, particularly excelling in complex analytical queries.

executable schemalog parsingnatural-language log querying

This work addresses the persistent gap between current natural language to SQL (NL2SQL) systems and human expert performance, which limits their reliable deployment in real-world database applications. To bridge this gap, the authors propose a large language model–based multi-agent framework that enhances generation quality through semantically enriched schema representations, integration of user-defined business rules, and a multi-stage reasoning pipeline. Key innovations include a novel multi-agent coordinator enabling planning, scheduling, and self-reflection mechanisms, as well as a context-aware schema augmentation strategy. Evaluated on the BIRD-SQL benchmark, the proposed approach achieves a semantic accuracy of 78.1%, substantially outperforming existing methods and demonstrating strong cross-domain generalization capabilities.

Natural Language to SQLNL2SQLrelational databases

Hot Scholars

AS

Ashwin Srinivasan

Senior Professor, Computer Science, BITS-Pilani, Goa
Inductive Logic ProgrammingMachine LearningSystems BiologyNeuro-Symbolic Models
JD

Jinho D. Choi

Associate Professor, Emory University
Natural Language ProcessingComputational LinguisticsConversational AI
NT

Nan Tang

National Institute of Biological Sciences, Beijing
stem cell biologyaginglung diseases
MZ

Minfeng Zhu

Zhejiang University
VisualisationMath