A Framework for Statistical Inference via Randomized Algorithms

📅 2023-07-20
📈 Citations: 1
✨ Influential: 0
📄 PDF
🤖 AI Summary
This paper addresses the challenge of uncertainty quantification arising from the inherent non-determinism of randomized algorithms—such as random projection and stochastic optimization—by proposing the first asymptotic statistical inference framework that requires no prior distributional assumptions. Methodologically, it establishes the first set of verifiable conditions for asymptotic normality of randomized outputs and introduces three novel strategies: sub-randomization, multi-run plug-in, and multi-run aggregation—integrating multi-run resampling, Polyak–Ruppert averaging, momentum SGD, and randomized sketching. The key contributions are: (i) enabling reliable confidence interval construction for high-dimensional and large-scale stochastic optimization and randomized least squares; and (ii) achieving negligible computational and communication overhead. Extensive simulations demonstrate robustness and practical efficacy on high-dimensional sparse and ultra-large-scale datasets.
📝 Abstract
Randomized algorithms, such as randomized sketching or stochastic optimization, are a promising approach to ease the computational burden in analyzing large datasets. However, randomized algorithms also produce non-deterministic outputs, leading to the problem of evaluating their accuracy. In this paper, we develop a statistical inference framework for quantifying the uncertainty of the outputs of randomized algorithms. Our key conclusion is that one can perform statistical inference for the target of a sequence of randomized algorithms as long as in the limit, their outputs fluctuate around the target according to any (possibly unknown) probability distribution. In this setting, we develop appropriate statistical inference methods -- sub-randomization, multi-run plug-in and multi-run aggregation -- by estimating the unknown parameters of the limiting distribution either using multiple runs of the randomized algorithm, or by tailored estimates. As illustrations, we develop methods for statistical inference when using stochastic optimization (such as Polyak-Ruppert averaging in stochastic gradient descent and stochastic optimization with momentum). We also illustrate our methods in inference for least squares parameters via randomized sketching, by characterizing the limiting distributions of sketching estimates in a possibly growing dimensional case. We further characterize the computation and communication cost of our methods, showing that in certain cases, they add negligible overhead. The results are supported via a broad range of simulations.
Problem

Research questions and friction points this paper is trying to address.

Quantify uncertainty in randomized algorithm outputs
Develop inference methods for non-deterministic algorithm results
Address computational challenges in large dataset analysis
Innovation

Methods, ideas, or system contributions that make the work stand out.

Framework for uncertainty quantification in randomized algorithms
Statistical inference via multi-run and tailored estimation
Efficient methods for stochastic optimization and sketching
🔎 Similar Papers
💼 Related Jobs
No related jobs found.
University of Macau | Columbia University | University of Pennsylvania
Z
Zhixiang Zhang
Department of Mathematics, University of Macau
S
S. Lee
Department of Economics, Columbia University
Edgar Dobriban
Edgar Dobriban
Statistics & Computer Science, University of Pennsylvania
StatisticsMachine LearningAI