sql querying

Designs, writes, and optimizes SQL queries, scripts, and generated SQL to extract, transform, explore, and report on relational/tabular data stored in RDBMS and analytical engines (including BigQuery). Builds and evaluates SQL-based artifacts and systems such as database schema and data models, stored procedures and query scripting, SQL generation and automation, and text-to-SQL / NL-to-SQL components that translate natural language into executable queries.

sqlquerying

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
1.45
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$181K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

In industrial settings, limited production data severely compromises the fidelity of test data for SQL generation services (e.g., NL2SQL), hindering simultaneous preservation of structural integrity and semantic coherence. Method: This paper proposes an LLM-driven high-fidelity test data generation method, integrating Gemini with schema-aware preprocessing, SQL-semantic alignment postprocessing, and constraint-guided sampling—supporting complex patterns including nested columns, multi-table JOINs, aggregations, and deep subqueries. Contribution/Results: The method jointly optimizes semantic consistency, syntactic correctness, and structural fidelity, significantly improving test coverage and defect detection rates. Evaluated on Google’s real-world NL2SQL workloads, it generates high-quality mock data out-of-the-box, effectively addressing the semantic incoherence prevalent in existing approaches under large-scale, complex database schemas.

Address limitations in handling complex schema structuresEnsure semantic coherence for robust SQL query testingGenerate high-fidelity test data for SQL services

DB-Explore: Automated Database Exploration and Instruction Synthesis for Text-to-SQL

Mar 06, 2025
HM
Haoyuan Ma
🏛️ Zhejiang University | OPPO Research Institute

Existing text-to-SQL approaches exhibit limited performance on complex databases, primarily due to insufficient deep understanding of database structure and semantics. This paper proposes a database-aware lightweight alignment framework: first, constructing a graph-based database schema representation; second, leveraging GPT-4 to jointly mine structural and semantic patterns for automated schema comprehension and diverse instruction distillation; third, fine-tuning Qwen2.5-coder-7B to drastically reduce reliance on computationally expensive, closed-source large language models. Evaluated on BIRD and Spider benchmarks, our method achieves 52.1% and 84.0% execution accuracy, respectively—surpassing multiple GPT-4–driven baselines at minimal computational cost while approaching state-of-the-art performance. The core contribution lies in the first integration of database graph modeling with LLM-driven semantic mining, enabling systematic, end-to-end alignment between the model and the underlying database.

Automates database exploration and instruction synthesis for LLMs.Enhances text-to-SQL systems for complex database structures.Improves SQL translation accuracy on domain-specific queries.

E-SQL: Direct Schema Linking via Question Enrichment in Text-to-SQL

Sep 25, 2024
HA
Hasan Alp Caferoglu
🏛️ Bilkent University

Natural language-to-SQL generation faces accuracy bottlenecks due to complex database schemas, ambiguous user intents, and semantic ambiguities. Method: This paper proposes a lightweight, efficient question-augmentation paradigm enabling end-to-end direct schema linking. It explicitly injects schema elements—including tables, columns, values, and conditions—into both the natural language question and SQL generation process; introduces a candidate-predicate augmentation mechanism to enhance semantic alignment for complex queries; and integrates zero-shot, single-turn prompting with large language models (e.g., DeepSeek-Coder-7B-Instruct), combining schema-aware question rewriting and predicate validation. Results: The approach achieves 66.29% execution accuracy on the BIRD benchmark and 56.45% even with small models without fine-tuning—demonstrating that question augmentation substantially improves LLM generalization in text-to-SQL tasks.

Complex SQL GenerationDatabase QueryingNatural Language Processing

Towards Automated Cross-domain Exploratory Data Analysis through Large Language Models

Dec 10, 2024
JZ
Jun-Peng Zhu
🏛️ East China Normal University | PingCAP

Data analysts face two primary bottlenecks: SQL generation and visualization selection. Existing approaches exhibit significant limitations in comprehending complex schemas, modeling ambiguous user intents, generalizing across domains, and enabling end-to-end text-to-visualization translation. This paper introduces TiInsight, a domain-agnostic system for automated exploratory data analysis (EDA). Its core contributions are: (1) Hierarchical Data Context (HDC) modeling, which enhances large language models (e.g., GPT-4) to reason over heterogeneous schemas and imprecise user intents; and (2) an end-to-end four-stage EDA pipeline—intent clarification, TiSQL (text-to-SQL), TiChart (automated chart recommendation), and GUI integration. TiSQL achieves 86.3% execution accuracy on Spider and sets a new state-of-the-art on Bird; user studies demonstrate superior performance over human experts. The system’s API is open-sourced and deployed in PingCAP’s production environment.

Automate SQL-based cross-domain exploratory data analysis.Enhance data visualization through text-to-SQL and text-to-visualization.Improve cross-domain generalization and user intent clarity in EDA.

Conformance Testing of Relational DBMS Against SQL Specifications

Jun 13, 2024
SL
Shuang Liu
🏛️ Renmin University of China | Tianjin University | Singapore Management University | East China Normal University | University of Science and Technology of China

This work addresses the challenge of verifying relational database management systems’ (RDBMS) compliance with SQL semantics at the standard specification level. We present the first executable Prolog reference implementation grounded in the complete formal SQL semantics defined by ISO/IEC 9075, integrated with differential fuzz testing for semantic-level black-box validation. Unlike prior approaches relying solely on crash detection or meta-transformation, our method enables end-to-end verifiable modeling of SQL standard semantics. Empirical evaluation across MySQL, TiDB, SQLite, and DuckDB uncovered 19 previously unknown vulnerabilities and 11 semantic inconsistencies—each traceable to explicit violations, omissions, or ambiguities in the SQL standard. Our approach significantly enhances the decidability and interpretability of SQL implementation correctness.

Detecting bugs and inconsistencies in major RDBMS systemsFormally defining SQL semantics for reference implementationTesting RDBMS semantic conformance to SQL specifications

Latest Papers

What's happening recently
View more

This work addresses the limitation of traditional database logical design, which overlooks the capacity of large language models (LLMs) to comprehend schema semantics, thereby constraining Text-to-SQL accuracy. For the first time, LLM-friendliness is incorporated into logical schema design through three semantic-preserving and composable transformation strategies: abstraction (+A), workload-aware partitioning (+P), and descriptive renaming (+R). The proposed approach is compatible with both supervised and zero-shot settings, yielding consistent improvements across multiple Text-to-SQL models. Evaluated on the BIRD-Union and Spider-Union benchmarks, the method achieves up to a 4.2% absolute gain in execution accuracy, significantly enhancing the mapping from natural language queries to executable SQL statements.

LLM-friendly schemalogical database designschema transformation

This work addresses the challenges of Text-to-SQL in large analytical databases, where complex schemas, ambiguous parsing, and data-dependent decisions hinder performance, and conventional fixed-pipeline systems struggle to recover from early errors. To overcome these limitations, the authors propose FlexSQL, an agent-based framework that integrates dynamic schema retrieval and a flexible execution mechanism, enabling exploration of schema structures, data validation, and backtracking for correction at any reasoning stage. FlexSQL supports dual-level repair—both at the code and planning levels—and combines SQL/Python hybrid generation with multi-interpretation execution plans. Evaluated on the Spider2-Snow benchmark, this approach achieves 65.4% accuracy, outperforming stronger open-source baselines and delivering over a 10% performance gain when integrated into general-purpose programming agents.

ambiguous queriescomplex schemadatabase interaction

Current text-to-SQL system evaluations rely on a single static database, which fails to capture model robustness across diverse data instances and may introduce significant bias. This work proposes SynSQL, a novel framework that leverages large language models to directly generate semantically consistent and schema-aligned relational test data from natural language questions. SynSQL formulates database construction as a structured generation task governed by semantic and relational constraints, comprising three stages: schema selection, question-guided data synthesis, and constraint-aware iterative refinement. Experiments on Spider, BIRD, and Spider 2.0 demonstrate that databases generated by SynSQL reduce the performance of ten state-of-the-art models by 3–14%, effectively uncovering errors masked by static evaluation and substantially enhancing assessment reliability and stress-testing capability.

benchmark artifactsdatabase synthesisevaluation robustness

This study addresses the persistent challenges of accuracy and robustness in natural language to SQL (NL2SQL) translation under complex query scenarios. The authors systematically evaluate the combined effects of multiple optimization strategies—including the NatSQL intermediate representation, synthetic data preprocessing and fine-tuning, and a novel SQL re-ranking model—using SmBoP and RASAT as backbone architectures. Through ablation studies and Shapley value analysis, they quantitatively assess, for the first time, the interaction effects among these components, revealing that their performance gains are not merely additive. The results demonstrate that non-trivial combinations of these techniques yield significant improvements on benchmarks such as Spider, underscoring the critical role of synergistic interactions among system components.

large language modelsmodel pipelineNatural Language to SQL

This work addresses the persistent gap between current natural language to SQL (NL2SQL) systems and human expert performance, which limits their reliable deployment in real-world database applications. To bridge this gap, the authors propose a large language model–based multi-agent framework that enhances generation quality through semantically enriched schema representations, integration of user-defined business rules, and a multi-stage reasoning pipeline. Key innovations include a novel multi-agent coordinator enabling planning, scheduling, and self-reflection mechanisms, as well as a context-aware schema augmentation strategy. Evaluated on the BIRD-SQL benchmark, the proposed approach achieves a semantic accuracy of 78.1%, substantially outperforming existing methods and demonstrating strong cross-domain generalization capabilities.

Natural Language to SQLNL2SQLrelational databases

Hot Scholars

YL

Yuyu Luo

Assistant Professor, HKUST(GZ) / HKUST
Data AgentsLLM AgentsDatabaseText-to-SQL
CB

Carsten Binnig

Full Professor, Computer Science, TU Darmstadt
Data ManagementMachine LearningModern Hardware
GL

Guoliang Li

Professor, Tsinghua University
DatabaseBig DataCrowdsourcingData Cleaning & Integration
NT

Nan Tang

National Institute of Biological Sciences, Beijing
stem cell biologyaginglung diseases
FH

Feiran Huang

Professor, Jinan University
Recommender systemsText-to-SQLSentiment AnalysisLLMs