Score
Design, build, and analyze concrete query expressions and end-to-end query strategies that retrieve, probe, or detect information from systems, including composing, expanding, rewriting, canonicalizing, and reformulating queries. Work covers query understanding, efficiency- and constraint-aware optimization (costs, budgets, query-efficiency), query-based detection, and interactive feedback mechanisms to iteratively improve effectiveness.
This work addresses the challenges posed by the rise of AI-generated queries to the readability and structural explicitness of existing relational query languages. It proposes a unified framework based on Abstract Relational Calculus (ARC) and relational graphs to systematically compare how languages such as SQL, dataframes, and graph query notations express identical query intents. By introducing a formal terminology encompassing information needs, query mappings, and relational schema structures, the study for the first time brings classical database languages and emerging alternatives into a common analytical perspective. The framework is further extended to handle recursive queries, nested relations, and problems beyond PTIME. This contribution establishes a reusable language comparison methodology and a precise design lexicon, offering practical tools for evaluating and designing future relational query languages.
This work addresses the inefficiency faced by data analysts who must repeatedly submit and integrate multiple related queries to explore salient data patterns. To streamline this process, the paper introduces the ANALYZE operator, which formalizes such exploratory analysis as five auxiliary cube queries, enabling comprehensive 360-degree examination of specific data subsets. Leveraging multi-query optimization (MQO), the authors devise three query merging and execution strategies—Mid-MQO, Min-MQO, and Max-MQO—that significantly improve execution efficiency while preserving result equivalence. Experimental evaluation demonstrates that Mid-MQO consistently delivers the best overall performance across most scenarios, whereas Max-MQO excels when sibling queries are numerous and exhibit high overlap.
This paper addresses the failure of conventional equivalence-based query rewriting when original data tables are inaccessible due to access control policies, privacy constraints, or prohibitive retrieval costs. To overcome this, we propose INQURE, a semantic intent-preserving query rewriting framework. Unlike traditional approaches relying on syntactic equivalence and query plan optimization, INQURE introduces, for the first time, large language model (LLM)-driven intent understanding and cross-table reconstruction—enabling semantically consistent rewriting across structurally heterogeneous and non-aligned schemas. The system incorporates pre-filtering of candidate tables, pruning heuristics, and learned ranking to form an end-to-end rewriting pipeline. Evaluated on a benchmark spanning 900+ real-world database schemas, INQURE demonstrates superior rewriting quality and practical utility. A user study further confirms its effective trade-off between execution feasibility and fidelity of analytical insights.
Traditional query optimizers—based on System R, Volcano, or Cascades—employ a monolithic, static, single-query architecture that exhibits performance instability, lacks global workload-level optimization, and suffers from architectural rigidity in cloud-native environments with massive-scale data and unified data platforms. To address these industrial challenges, this paper proposes three evolutionary directions: (1) extending optimization from individual queries to workload-aware collaborative optimization; (2) establishing an execution-feedback-driven closed-loop optimization framework; and (3) designing a composable, modular optimizer architecture enabling cross-engine reuse and agile iteration. The study distills three key trends—workload awareness, feedback closure, and architectural decoupling—providing a practical, implementable technical pathway and industrial paradigm for building next-generation query optimization systems that are dynamic, holistic, and self-adaptive.
In text-to-SQL tasks, lightweight models suffer from low accuracy on complex queries and high inference overhead. This paper proposes an execution-result-guided multi-candidate SQL filtering framework, introducing the first execution-feedback-driven candidate reranking paradigm. Leveraging a lightweight semantic consistency scoring mechanism, it reranks sampled SQL queries based on actual database execution validation—requiring no fine-tuning and enabling plug-and-play adaptation to any SQL generation model. Our method significantly improves semantic correctness and execution accuracy of small models on complex queries. It outperforms large reasoning models—including o1, o3-mini, and DeepSeek R1—across multiple standard benchmarks, while reducing inference cost by up to 30×. To our knowledge, this is the first approach to achieve simultaneous superiority in both accuracy and efficiency for lightweight models in text-to-SQL.
Database query plan representations are highly fragmented, impeding test method reuse and cross-system analysis. Method: This paper proposes the first database-agnostic unified query plan representation framework, systematically identifying the “operator–attribute–format” trinity as the common structural foundation across execution plans. It abstracts internal plans from nine mainstream databases via cross-database reverse parsing and intermediate representation modeling, yielding an extensible, formally verifiable unified model. Contribution/Results: The framework enables seamless reuse of existing testing methodologies across all nine databases, uncovering 17 previously undetected, database-specific defects. It facilitates rapid adaptation of multi-database visualization tools and supports standardized comparative analysis—including semantic alignment and performance profiling—of query plans across heterogeneous systems.
Traditional database systems rely on static rewrite rules that struggle to adapt to diverse queries and system characteristics, while existing large language model (LLM)-based approaches suffer from an excessively large search space, unreliable validation, and insufficient use of metadata. This work proposes a plug-in optimization layer that integrates catalog and statistical metadata to generate templated rules guiding LLM-based SQL rewriting. Semantic correctness is verified using sampled data, and candidate plans are ranked to enhance performance. The method is compatible with PostgreSQL, MySQL, and DuckDB, achieving up to 16× speedup over native DBMS optimizers and 22× over current LLM-based methods across eight benchmarks, with individual queries accelerated by over 600×—significantly surpassing the limitations of both traditional rule-based and pure LLM-driven approaches.
This work addresses the problem of automatically selecting optimal execution plans for general conjunctive queries (CQs) based on database statistics. It proposes the PANDA framework, which, for the first time, directly integrates information-theoretically derived tight upper bounds on intermediate relation cardinalities into the query plan generation and optimization process. The approach unifies treatment across diverse query scenarios and matches or surpasses the performance of specialized algorithms on several classical problems—including those relying on fast matrix multiplication—demonstrating both generality and efficiency. The key contribution lies in establishing a novel connection among information theory, constraint satisfaction problems, and database query optimization, thereby achieving a synergistic improvement between theoretical guarantees and practical performance.
Traditional SQL undermines the theoretical foundations of the relational model by relying on nulls and bags, leading to semantic ambiguities and increased query complexity. This work proposes and implements Rel, a novel declarative query language that entirely eliminates nulls and bags, adhering strictly to set semantics and canonical relational algebra. Through an end-to-end system design and real-world deployment, we demonstrate that a null-free, bag-free relational system is not only expressively complete but also offers significant advantages in optimizability, semantic clarity, and engineering practicality. Our results confirm the feasibility and superiority of this paradigm for real-world applications.
Existing natural language interfaces to databases lack systematic evaluation frameworks and design theories. This work proposes QUEST, a novel framework that integrates the general-purpose FAR operation schema—Filter, Aggregate, Return—with the W5H semantic dimensions (Who, What, Where, When, Why, How) to enable structured analysis and evaluation of text-to-SQL query semantics. Through semantic annotation and structural parsing, the study validates the universality of the FAR schema across five cross-domain datasets comprising 120,464 queries. The analysis further reveals significant inter-domain disparities in semantic distributions: for instance, medical queries predominantly focus on WHEN and WHO, while WHY and HOW are nearly absent, underscoring a critical challenge for machines in performing deep reasoning over structured data.