Score
Designs, implements, and evaluates estimators that compute query result selectivities incrementally as queries arrive, updating model parameters online to adapt to changing data distributions; these systems handle point, range, and subset queries and are optimized to minimize cumulative prediction loss under dynamic conditions.
This work addresses the problem of selectivity estimation for linear queries—such as point and range queries—in dynamic databases. It introduces, for the first time, online learning theory to this setting, proposing an online estimation method based on histogram models and standard loss functions. The approach effectively adapts to time-varying data distributions and query workloads, delivering provably low regret in both static and dynamic environments. The core contribution lies in establishing tight upper and lower bounds on regret specifically for histogram-based linear queries, thereby providing the first formal online learning framework for dynamic selectivity estimation with theoretical performance guarantees.
Existing differential privacy methods for adaptive data analysis assume static datasets, rendering them ill-suited for dynamically growing data—where maintaining both generalization and statistical validity remains challenging. Method: We establish, for the first time, a generalization bound for adaptive analysis under dynamic data growth; introduce a time-varying empirical accuracy bound theorem; and integrate a clipped Gaussian mechanism with a batched query framework to achieve progressively tightening error guarantees as data accumulates. Contribution/Results: Theoretically, our approach requires sample complexity scaling only with the square root of the number of queries—strictly improving upon static partitioning baselines. Empirically, it significantly enhances estimation accuracy and practical utility across statistical query tasks. Our framework provides a novel paradigm for streaming or incremental adaptive analysis, uniquely balancing rigorous theoretical guarantees with real-world feasibility.
To address suboptimal query execution plans caused by cardinality estimation errors during optimization, this paper proposes the first method to detect such plans *in real time during optimization*—without requiring query execution or ground-truth cardinalities. Our approach comprises two core contributions: (1) a real-time detection mechanism based on *subplan ranking consistency*, breaking from conventional post-execution analysis paradigms; and (2) an *incremental proxy cardinality refinement framework* that continuously improves detection accuracy as the query workload evolves. The method integrates third-party cardinality estimators, subplan enumeration analysis, ranking consistency metrics, and online incremental learning. Evaluated on JOB-LIGHT-SCALE and STATS-CEB-SCALE benchmarks, our method achieves 88.7% accuracy in predicting suboptimal plans—significantly outperforming traditional post-execution error-based metrics.
This paper addresses the joint online optimization of service price $p$ and service capacity $mu$ in a queueing system where demand and service duration distributions are unknown, aiming to maximize cumulative expected profit (revenue minus capacity cost and delay penalty). Departing from the conventional two-stage “predict-then-optimize” paradigm, we propose an end-to-end online learning framework that intrinsically incorporates parameter estimation error into the decision process, enabling error-aware robust optimization. Our algorithm integrates stochastic approximation, queueing-theoretic modeling, and online convex optimization, with theoretical guarantees on convergence and an $O(sqrt{T})$ regret upper bound. Extensive simulations demonstrate that our approach improves profit by 12%–28% over benchmark policies across diverse representative scenarios.
This work addresses the challenge of database tuning, which is hindered by a vast parameter space, reliance on manual expertise, and costly warm-up phases that impede efficient identification of critical configuration knobs. To overcome these limitations, the paper proposes a novel online, warm-up-free dynamic tuning approach that integrates recursive feature elimination with cross-validation (RFECV) to identify key parameters, employs likelihood ratio tests (LRT) to balance exploration and exploitation, and seamlessly couples this feature selection mechanism with Bayesian optimization (BO) for real-time search of optimal configurations. This method is the first to enable dynamic parameter selection without requiring an initial warm-up period and supports incorporation of prior knowledge. Experimental results demonstrate that it matches or surpasses state-of-the-art tuners across multiple benchmarks while substantially reducing tuning time and computational overhead.
This work addresses the challenge of high-variance-induced sampling costs in index-assisted approximate query processing for ad-hoc aggregation over frequently updated flat data, where existing methods struggle to balance accuracy and low latency. To overcome this limitation, we propose a two-stage online stratified sampling framework that, for the first time, integrates stratified sampling into index-assisted online aggregation. In the first stage, samples are used simultaneously for real-time estimation and to refine the subsequent sampling strategy; the second stage performs efficient sampling based on an optimized stratification structure and Neyman allocation. We develop greedy and dynamic programming algorithms tailored to an index-aware sampling cost model to balance efficiency and accuracy. Experimental results demonstrate that our approach achieves up to 3× speedup over index-assisted uniform sampling and up to 98,708× speedup compared to traditional scan-based stratified sampling.
This work addresses the limited generalization of existing learning-based benefit estimators in online index tuning, which stems from sparse training feedback and workload drift. To overcome these challenges, the authors propose UTune, a novel framework that integrates operator-level learning models with uncertainty quantification, explicitly incorporating uncertainty estimates into the index selection and configuration enumeration process. Furthermore, UTune employs an uncertainty-aware ε-greedy strategy to enable robust and efficient evaluation of index benefits. Experimental results demonstrate that UTune significantly outperforms state-of-the-art methods, achieving faster convergence under stable workloads while simultaneously reducing both query execution time and exploration overhead.
This work addresses the challenges faced by existing learned query optimizers in dynamic, distributed multi-tenant data warehouses, where input statistics are often missing, online fine-tuning is impractical, and performance gains are uncertain. To overcome these limitations, the authors propose LOAM, a novel framework that introduces the first plan encoding method independent of input statistics, explicitly models the impact of execution environments on query cost, and integrates domain-adaptive training with a lightweight, high-yield item selection mechanism. This design enables efficient deployment without requiring online fine-tuning. Evaluated in Alibaba’s MaxCompute production environment, LOAM reduces CPU costs by up to 30% compared to the native optimizer, substantially lowering resource consumption.
Existing reinforcement learning (RL)-based query optimizers suffer from unstable performance, severe degradation, and slow convergence in single-query settings, hindering their practical deployment. This work proposes RELOAD, the first RL-based query optimizer that simultaneously achieves high robustness and high training efficiency. By incorporating stability constraints and an efficient policy learning mechanism, RELOAD effectively suppresses performance fluctuations and accelerates convergence to expert-level plan quality. Experimental results on the JOB, TPC-DS, and SSB benchmarks demonstrate that RELOAD improves robustness by up to 2.4× and training efficiency by up to 3.1× compared to the current state-of-the-art RL-based optimizers.