complexity-aware query rewriting

Designs and implements rule-driven, complexity-aware query rewriting systems that transform queries into semantically equivalent but lower-cost forms using set-theoretic, join, and set-operation transformation rules. This competence covers building and analyzing rewrite-rule libraries, proving correctness and complexity bounds, and developing cost models and specialized predicate rewrites (including spatial-predicate optimizations) to reduce execution cost and execution time.

complexity-awarequeryrewriting

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.17
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

QUITE: A Query Rewrite System Beyond Rules with LLM Agents

Jun 09, 2025
YS
Yuyang Song
🏛️ Sichuan University | Cornell University | Purdue University | Hong Kong University of Science and Technology | Chinese Academy of Sciences

Existing SQL query rewriting approaches rely on predefined rules, encountering three key bottlenecks: difficulty in discovering effective rules, poor generalizability, and inability to express complex optimization logic—leading to narrow coverage and frequent performance regressions. This paper proposes the first training-free, feedback-aware LLM-based multi-agent framework for dynamic query rewriting. It integrates a finite-state-machine-driven workflow controller, a database execution feedback loop, a rewriting middleware layer, and prompt injection techniques to ensure both semantic equivalence and high-performance execution. Experimental results demonstrate that, compared to state-of-the-art methods, our approach achieves up to a 35.8% reduction in query execution time and a 24.1% improvement in rewriting success rate. Moreover, it significantly broadens support for complex query patterns and diverse rewriting strategies.

Ensuring semantic equivalence and performance in LLM-based query rewritesOvercoming limitations of rule-based SQL query rewrite systemsUsing LLMs to rewrite SQL queries beyond fixed rules

Traditional query rewriting rules are tightly coupled with execution engines and lack formal correctness guarantees, making them difficult to port and error-prone. This work proposes Rulescript, an engine-agnostic domain-specific language that decouples rule specification from execution through a match-and-transform two-phase mechanism and automatically verifies semantic equivalence using a relational algebra core. Rulescript supports custom operators and, combined with lightweight adapters, enables cross-engine deployment. The authors experimentally reproduce 33 Apache Calcite rewrite rules and successfully migrate them to both CockroachDB and Apache DataFusion, demonstrating the feasibility of “write once, deploy anywhere.” To the best of our knowledge, this is the first extensible, verifiable, and cross-platform query rewriting framework.

engine-agnosticformal verificationlogical query plan

Query Rewriting via LLMs

Feb 18, 2025
SD
Sriram Dharwada
🏛️ Indian Institute of Science | Microsoft Research India

SQL query rewriting faces a fundamental trade-off between performance optimization and interpretability, while remaining prone to semantic or syntactic errors. Method: This paper proposes an LLM-driven, database-aware rewriting framework featuring a novel token-probability-guided rewrite path selection mechanism; it integrates metadata-aware prompting, selectivity-aware rewriting rules, redundancy elimination, and dual verification—logical equivalence checking and statistical consistency validation. Contribution/Results: The framework bridges the gap between purely rule-based and purely LLM-based approaches, enabling robust end-to-end rewriting. Experiments on TPC-DS show that two-thirds of queries achieve >1.5× speedup; rewrite coverage reaches four times that of the state-of-the-art; and the geometric mean speedup improves by an order of magnitude. The framework has been integrated into the LITHE system and validated across mainstream database platforms.

Ensuring correctness and efficiency in rewritesImproving performance over state-of-the-art techniquesLeveraging LLMs for SQL query rewriting

Query Rewriting via Large Language Models

Mar 14, 2024
JL
Jie Liu
🏛️ University of Michigan, Ann Arbor

To address the challenges of poor generalizability and verifiability in low-quality SQL query rewriting, this paper proposes GenRewrite—the first end-to-end LLM-driven query rewriting system. Methodologically, it introduces (1) natural-language rewriting rules (NLR2s) for knowledge representation and cross-query-pattern transfer; (2) a counterexample-guided iterative correction framework that jointly ensures semantic correctness and execution efficiency; and (3) tight integration of SQL syntactic/semantic constraints with LLM reasoning. Evaluated on 99 complex queries from the TPC benchmarks, GenRewrite achieves >2× speedup on 22 queries, improves rewriting coverage by 2.5–3.2× over conventional methods, and outperforms zero-shot LLM baselines by 2.1×.

Addressing limitations of traditional algorithms with complex query patternsAutomating query rewriting to replace manual and rule-based methodsReducing syntactic and semantic errors in rewritten queries efficiently

R-Bot: An LLM-based Query Rewrite System

Dec 02, 2024
ZS
Zhaoyan Sun
🏛️ Tsinghua University | Shanghai Jiao Tong University

Traditional SQL query rewriting approaches—whether heuristic- or learning-based—suffer from limited accuracy and robustness, while direct LLM invocation (e.g., GPT-4) often yields hallucinated outputs. To address these challenges, this paper proposes the first LLM-augmented framework specifically designed for SQL query rewriting. Our method introduces three key innovations: (1) a multi-source rewriting evidence generation pipeline that jointly leverages syntactic structure, execution feedback, and semantic similarity; (2) a syntax–semantics hybrid retrieval mechanism to enhance context relevance; and (3) a self-reflective, stepwise LLM reasoning and self-verification paradigm. Evaluated across multiple mainstream benchmarks, our approach significantly outperforms state-of-the-art methods, achieving 12.6–28.3% higher rewriting accuracy and a 19.4% improvement in execution success rate, while effectively mitigating hallucination—demonstrating both high precision and strong robustness.

Optimizing SQL queries without altering resultsOvercoming limitations of heuristic and learning-based methodsReducing hallucinations in LLM-based query rewriting

Latest Papers

What's happening recently
View more

The query optimizer in a Database Management Systems (DBMS), translates declarative queries into efficient execution plans. Conventional bottom-up optimization consists of two main stages: Query Rewrite (QRW) and Cost-Based Optimization (CBO). However, applying a rewrite rule during QRW may not always be beneficial; the best choice may depend on the (estimated) execution cost of the original and rewritten expressions. Fully exploiting such cost-dependent rules necessitates interleaving QRW with frequent CBO invocations, thereby incurring substantial overhead and often impractical optimization times. To mitigate this inefficiency, we introduce a novel cost-based rewrite framework for bottom-up optimizers. The core of our approach is a multi-level caching mechanism for intermediate CBO results aimed at eliminating redundant computation. Furthermore, we establish and exploit upper cost bounds to intelligently prune the search space during optimization. We also contribute methodological solutions for caching and reusing intermediate plan results within a bottom-up optimizer architecture. The framework has been implemented in the GaussDB optimizer. Experiments show that it significantly reduces overall optimization time, demonstrating the effectiveness of our approach.

bottom-up optimizercost-based optimizationdatabase query optimization

Traditional database systems rely on static rewrite rules that struggle to adapt to diverse queries and system characteristics, while existing large language model (LLM)-based approaches suffer from an excessively large search space, unreliable validation, and insufficient use of metadata. This work proposes a plug-in optimization layer that integrates catalog and statistical metadata to generate templated rules guiding LLM-based SQL rewriting. Semantic correctness is verified using sampled data, and candidate plans are ranked to enhance performance. The method is compatible with PostgreSQL, MySQL, and DuckDB, achieving up to 16× speedup over native DBMS optimizers and 22× over current LLM-based methods across eight benchmarks, with individual queries accelerated by over 600×—significantly surpassing the limitations of both traditional rule-based and pure LLM-driven approaches.

database optimizationlarge language modelsmetadata utilization

This work addresses the trade-off between large language model (LLM) invocation overhead and relational processing cost in hybrid semantic-relational querying by introducing, for the first time, a plan-level optimization framework that formulates query optimization as the optimal placement of semantic filters. Leveraging equivalence rewriting rules, a function caching mechanism, and a dynamic programming–based cost model, the framework jointly minimizes LLM calls and relational operation costs while preserving result quality (achieving an average F1 score of 0.85). Experiments on 44 semantic SQL queries demonstrate up to 1.5× speedup and 4.29× cost reduction, with the highest accuracy among six public systems. The study further reveals that delaying semantic filtering—though reducing LLM invocations—can induce relational processing bottlenecks in multi-table queries, leading to the design of an effective balancing strategy.

cost-based optimizationhybrid query plansLLM invocation

Existing SQL rewriting approaches struggle to substantially improve the performance of modern analytical queries while preserving semantic correctness. This work formulates SQL rewriting as a policy optimization problem and introduces a GRPO-based reinforcement learning framework that integrates multi-dimensional reward signals, including semantic equivalence, textual similarity, physical plan divergence, and runtime speedup. The approach innovatively incorporates a probabilistic gating mechanism for adaptive reward shaping and a curriculum learning–driven hierarchical reward unlocking strategy, complemented by an intra-policy self-improvement mechanism to enhance both sample efficiency and rewrite quality. Experimental results demonstrate that the proposed method significantly outperforms rule-based and large language model (LLM) baselines on both in-distribution and out-of-distribution workloads, markedly reducing performance-degrading rewrites and effectively mitigating tail latency.

LLM-based optimizationphysical execution planruntime performance degradation

Hot Scholars

MB

Meghyn Bienvenu

CNRS Researcher
Artificial IntelligenceKnowledge Representation and ReasoningLogic in Computer Science
MO

Magdalena Ortiz

Vienna University of Technology (TU Wien)
Artificial IntelligenceKnowledge RepresentationComputational Logic
CX

Canwen Xu

Snowflake
natural language processingmachine learning
DF

Diego Figueira

CNRS, LaBRI, Univ. Bordeaux
Logic in Computer ScienceDatabase TheoryAutomata Theory