Score
Designs and builds data representations and execution techniques for query processing that produce and operate on factorized (compact, algebraic) encodings of join and related query results. Implements and analyzes factorized databases and execution strategies that exploit shared subexpressions to reduce intermediate-result size and to compute aggregates or counts without fully materializing joins.
研究提出FFX引擎,通过支持任意因式分解方案并保持全矢量化处理来优化多对多连接问题,减少中间结果体积,提高分析和语义查询效率。
To address excessive hidden constants and intermediate result explosion in acyclic join queries on column-store databases—caused by naive implementations of Yannakakis’ algorithm—this paper proposes Shredded Yannakakis (SYA). SYA is grounded in a novel formalization of two-phase Nested Semijoin Algebra (2-phase NSA), which strictly enforces semijoin-based contraction *before* expansion. It introduces the *shredding* execution paradigm, decoupling Lookup and Expand operators and enabling automatic translation of binary join plans into 2-phase NSA. Theoretically, SYA is proven instance-optimal and regret-free. Evaluated on 1,849 real-world queries, SYA achieves performance improvements for 85.3% of them, with speedups up to 62.5×; remaining queries exhibit competitive performance.
Traditional relational databases struggle to efficiently process multimodal, context-rich data: relational operators lack contextual modeling capabilities, while representation learning models cannot be readily integrated into declarative query frameworks. This paper proposes **context-enhanced relational join operators**, introducing the first composable embedding operator that natively integrates vector embeddings into relational algebra. We design algebraic equivalence rules and logical/physical optimization mechanisms to build a vector-relational hybrid execution engine, and propose a协同 optimization strategy jointly leveraging sequential scans and vector indexes. Our approach preserves SQL’s declarative semantics while enabling multimodal, context-aware querying. Experiments demonstrate up to an order-of-magnitude improvement in query performance over pure vector database solutions. The system validates the effectiveness of our optimization techniques and characterizes performance trade-offs across diverse application scenarios.
This paper addresses the challenge of data integration across heterogeneous databases. It proposes an algebraic approach grounded in category theory and functional programming. The method models database schemas and instances as many-sorted equational theories and their initial algebras, respectively, and employs adjoint functors to enable rigorous cross-schema data migration. Innovatively, it unifies category theory, many-sorted equational logic, and functional programming paradigms; introduces a pushout-based schema mapping construction; and defines an algebraic query language—with for/where/return syntax—endowed with formal semantics. The authors implement AQL, an open-source tool supporting formally specified schema mappings, automated data migration, and verifiable query compilation. This framework constitutes the first theoretically rigorous integration of these three foundational paradigms, simultaneously ensuring mathematical precision and enhancing the automation and reliability of data integration.
To address three key bottlenecks in GPU-accelerated databases for non-ML workloads—high random memory access overhead, insufficient acceleration for high-cardinality group-by operations, and the absence of query optimization mechanisms—this paper proposes an integrated solution. First, the GFTR technique reduces random memory access latency, achieving a 2.3× speedup. Second, a partitioned group-by algorithm tailored for high-cardinality scenarios is introduced, delivering 19.4× and 1.7× improvements for hash-based and sort-based implementations, respectively. Third, a lightweight, GPU-aware cost model coupled with an adaptive implementation selection strategy enables intelligent query optimization. The approach synergistically leverages GPU parallelism, hash-sort co-optimization, dynamic partitioning, and query-level cost modeling. Experimental evaluation demonstrates a significant reduction in the proportion of random memory accesses and yields deployable, heuristic-driven query optimization rules.
This work addresses the inefficiency faced by data analysts who must repeatedly submit and integrate multiple related queries to explore salient data patterns. To streamline this process, the paper introduces the ANALYZE operator, which formalizes such exploratory analysis as five auxiliary cube queries, enabling comprehensive 360-degree examination of specific data subsets. Leveraging multi-query optimization (MQO), the authors devise three query merging and execution strategies—Mid-MQO, Min-MQO, and Max-MQO—that significantly improve execution efficiency while preserving result equivalence. Experimental evaluation demonstrates that Mid-MQO consistently delivers the best overall performance across most scenarios, whereas Max-MQO excels when sibling queries are numerous and exhibit high overlap.
Existing columnar databases lack effective optimization for queries that intertwine relational and array operations. This work proposes A3D-RA, an extended relational algebra that natively supports array attributes, and formally defines its semantics for the first time. Building upon this foundation, we develop a modular, backend-agnostic optimization framework equipped with a complete set of equivalence-preserving transformation rules. The framework enables polynomial-time enumeration of optimal execution plans for non-join operations. Experimental evaluation across three mainstream analytical database engines demonstrates that integrating this optimization layer consistently yields significant performance improvements on real-world workloads.
本文提出一种映射和能力感知的优化方法,通过模型感知谓词下推、跨模型依赖连接等技术减少多模型数据查询中的跨模型开销。
Existing approaches to window function optimization suffer from stringent applicability conditions and limited generalizability, lacking a unified reasoning framework. This work proposes the first systematic inference framework for window function optimization, which derives algebraically equivalent transformations that can be safely applied by leveraging frame analysis, partition analysis, and a novel co-evaluation strategy. This enables predicate pushdown even in the presence of dependencies on window results. Implemented in an open-source query engine, the framework consistently preserves or improves performance, achieving speedups of up to 40.7× on common queries, with gains becoming more pronounced as data scale increases.
This study addresses the limitation of the standard semiring framework in supporting data provenance for aggregation queries involving HAVING clauses. We propose defining, for the first time, a provenance semantics for such queries within commutative semirings with monus (m-semirings). This formulation inherently accommodates self-join rewriting standards without introducing additional operators, and we formally prove their equivalence. Furthermore, we implement algorithmic optimizations and a prototype system building upon ProvSQL. The primary contribution lies in overcoming the constraints of classical semiring theory regarding aggregation operations, thereby providing a unified and elegant provenance semantics framework for aggregate queries. Experimental evaluations on real-world datasets demonstrate that our approach effectively supports probabilistic query evaluation, confirming its practical viability and performance feasibility.