query formulation

Design, build, and analyze concrete query expressions and end-to-end query strategies that retrieve, probe, or detect information from systems, including composing, expanding, rewriting, canonicalizing, and reformulating queries. Work covers query understanding, efficiency- and constraint-aware optimization (costs, budgets, query-efficiency), query-based detection, and interactive feedback mechanisms to iteratively improve effectiveness.

queryformulation

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.23
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

This work addresses the inefficiency faced by data analysts who must repeatedly submit and integrate multiple related queries to explore salient data patterns. To streamline this process, the paper introduces the ANALYZE operator, which formalizes such exploratory analysis as five auxiliary cube queries, enabling comprehensive 360-degree examination of specific data subsets. Leveraging multi-query optimization (MQO), the authors devise three query merging and execution strategies—Mid-MQO, Min-MQO, and Max-MQO—that significantly improve execution efficiency while preserving result equivalence. Experimental evaluation demonstrates that Mid-MQO consistently delivers the best overall performance across most scenarios, whereas Max-MQO excels when sibling queries are numerous and exhibit high overlap.

ANALYZE operatorcube queryingdata analysis

The Case for Intent-Based Query Rewriting

Nov 25, 2025
GL
Gianna Lisa Nicolai
🏛️ RPTU Kaiserslautern-Landau

This paper addresses the failure of conventional equivalence-based query rewriting when original data tables are inaccessible due to access control policies, privacy constraints, or prohibitive retrieval costs. To overcome this, we propose INQURE, a semantic intent-preserving query rewriting framework. Unlike traditional approaches relying on syntactic equivalence and query plan optimization, INQURE introduces, for the first time, large language model (LLM)-driven intent understanding and cross-table reconstruction—enabling semantically consistent rewriting across structurally heterogeneous and non-aligned schemas. The system incorporates pre-filtering of candidate tables, pruning heuristics, and learned ranking to form an end-to-end rewriting pipeline. Evaluated on a benchmark spanning 900+ real-world database schemas, INQURE demonstrates superior rewriting quality and practical utility. A user study further confirms its effective trade-off between execution feasibility and fidelity of analytical insights.

Developing intent-based query rewriting using large language modelsEnabling data access despite access control, privacy, or cost constraintsRewriting queries to preserve insights while altering structure and syntax

Query Optimization in the Wild: Realities and Trends

Oct 22, 2025
YT
Yuanyuan Tian
🏛️ Microsoft

Traditional query optimizers—based on System R, Volcano, or Cascades—employ a monolithic, static, single-query architecture that exhibits performance instability, lacks global workload-level optimization, and suffers from architectural rigidity in cloud-native environments with massive-scale data and unified data platforms. To address these industrial challenges, this paper proposes three evolutionary directions: (1) extending optimization from individual queries to workload-aware collaborative optimization; (2) establishing an execution-feedback-driven closed-loop optimization framework; and (3) designing a composable, modular optimizer architecture enabling cross-engine reuse and agile iteration. The study distills three key trends—workload awareness, feedback closure, and architectural decoupling—providing a practical, implementable technical pathway and industrial paradigm for building next-generation query optimization systems that are dynamic, holistic, and self-adaptive.

Addressing limitations of traditional monolithic query optimizer architectureExpanding optimization scope from single queries to entire workloadsImproving query performance robustness through optimization-execution feedback loops

Query and Conquer: Execution-Guided SQL Generation

Mar 31, 2025
LB
Lukasz Borchmann
🏛️ Snowflake AI Research

In text-to-SQL tasks, lightweight models suffer from low accuracy on complex queries and high inference overhead. This paper proposes an execution-result-guided multi-candidate SQL filtering framework, introducing the first execution-feedback-driven candidate reranking paradigm. Leveraging a lightweight semantic consistency scoring mechanism, it reranks sampled SQL queries based on actual database execution validation—requiring no fine-tuning and enabling plug-and-play adaptation to any SQL generation model. Our method significantly improves semantic correctness and execution accuracy of small models on complex queries. It outperforms large reasoning models—including o1, o3-mini, and DeepSeek R1—across multiple standard benchmarks, while reducing inference cost by up to 30×. To our knowledge, this is the first approach to achieve simultaneous superiority in both accuracy and efficiency for lightweight models in text-to-SQL.

Improving accuracy in text-to-SQL tasksReducing inference cost by 30 timesUsing execution results to select best query

Towards a Unified Query Plan Representation

Aug 14, 2024
JB
Jinsheng Ba
🏛️ National University of Singapore

Database query plan representations are highly fragmented, impeding test method reuse and cross-system analysis. Method: This paper proposes the first database-agnostic unified query plan representation framework, systematically identifying the “operator–attribute–format” trinity as the common structural foundation across execution plans. It abstracts internal plans from nine mainstream databases via cross-database reverse parsing and intermediate representation modeling, yielding an extensible, formally verifiable unified model. Contribution/Results: The framework enables seamless reuse of existing testing methodologies across all nine databases, uncovering 17 previously undetected, database-specific defects. It facilitates rapid adaptation of multi-database visualization tools and supports standardized comparative analysis—including semantic alignment and performance profiling—of query plans across heterogeneous systems.

Enabling cross-system performance comparison and optimization insightsReducing implementation effort for testing and visualization toolsUnifying diverse query plan representations across database systems

Latest Papers

What's happening recently
View more

Traditional database systems rely on static rewrite rules that struggle to adapt to diverse queries and system characteristics, while existing large language model (LLM)-based approaches suffer from an excessively large search space, unreliable validation, and insufficient use of metadata. This work proposes a plug-in optimization layer that integrates catalog and statistical metadata to generate templated rules guiding LLM-based SQL rewriting. Semantic correctness is verified using sampled data, and candidate plans are ranked to enhance performance. The method is compatible with PostgreSQL, MySQL, and DuckDB, achieving up to 16× speedup over native DBMS optimizers and 22× over current LLM-based methods across eight benchmarks, with individual queries accelerated by over 600×—significantly surpassing the limitations of both traditional rule-based and pure LLM-driven approaches.

database optimizationlarge language modelsmetadata utilization

This work addresses the problem of automatically selecting optimal execution plans for general conjunctive queries (CQs) based on database statistics. It proposes the PANDA framework, which, for the first time, directly integrates information-theoretically derived tight upper bounds on intermediate relation cardinalities into the query plan generation and optimization process. The approach unifies treatment across diverse query scenarios and matches or surpasses the performance of specialized algorithms on several classical problems—including those relying on fast matrix multiplication—demonstrating both generality and efficiency. The key contribution lies in establishing a novel connection among information theory, constraint satisfaction problems, and database query optimization, thereby achieving a synergistic improvement between theoretical guarantees and practical performance.

conjunctive querydatabase theoryinformation theory

Traditional SQL undermines the theoretical foundations of the relational model by relying on nulls and bags, leading to semantic ambiguities and increased query complexity. This work proposes and implements Rel, a novel declarative query language that entirely eliminates nulls and bags, adhering strictly to set semantics and canonical relational algebra. Through an end-to-end system design and real-world deployment, we demonstrate that a null-free, bag-free relational system is not only expressively complete but also offers significant advantages in optimizability, semantic clarity, and engineering practicality. Our results confirm the feasibility and superiority of this paradigm for real-world applications.

bagsnullsquery language

Existing natural language interfaces to databases lack systematic evaluation frameworks and design theories. This work proposes QUEST, a novel framework that integrates the general-purpose FAR operation schema—Filter, Aggregate, Return—with the W5H semantic dimensions (Who, What, Where, When, Why, How) to enable structured analysis and evaluation of text-to-SQL query semantics. Through semantic annotation and structural parsing, the study validates the universality of the FAR schema across five cross-domain datasets comprising 120,464 queries. The analysis further reveals significant inter-domain disparities in semantic distributions: for instance, medical queries predominantly focus on WHEN and WHO, while WHY and HOW are nearly absent, underscoring a critical challenge for machines in performing deep reasoning over structured data.

natural language interfacesquery understandingsemantic evaluation

Hot Scholars

DF

Diego Figueira

CNRS, LaBRI, Univ. Bordeaux
Logic in Computer ScienceDatabase TheoryAutomata Theory
CL

Carsten Lutz

Professor of Computer Science, University of Leipzig
Knowledge RepresentationArtificial IntelligenceLogic in Computer ScienceTheoretical Computer Science
YL

Yuyu Luo

Assistant Professor, HKUST(GZ) / HKUST
Data AgentsLLM AgentsDatabaseText-to-SQL
YS

Yangqiu Song

HKUST
Artificial IntelligenceData MiningNatural Language ProcessingKnowledge Graphs
MA

Mahmoud Abo Khamis

RelationalAI (mahmoud.abokhamis@relational.ai)
Database Systems and TheoryIn-database Machine Learning