error decomposition analysis

Design and carry out analyses that decompose the total error of an estimator, numerical scheme, or algorithm into identifiable components (approximation, statistical/empirical consistency, optimization, and quadrature), and derive formal decompositions such as approximation–consistency–optimization splits. Prove quantitative bounds for each component—often in Sobolev or broken Sobolev norms—and quantify contributions from empirical sampling, quadrature, and optimization procedures (including least-squares).

errordecompositionanalysis

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.19
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

This work addresses the limitations of traditional generalization analyses, which rely on the often unverifiable assumption of independent and identically distributed (i.i.d.) data and thus struggle to accurately characterize model performance on unseen data. The paper proposes a deterministic generalization analysis framework that dispenses with any prior probabilistic assumptions. By examining the sensitivity of optimization solutions to data perturbations, it decomposes the generalization error into geometric and probabilistic components, achieving their first-ever decoupling. The framework expresses generalization bounds via a variational principle, leveraging deterministic perturbation analysis and optimization sensitivity theory to capture the discrepancy between in-sample and out-of-sample performance. Error terms are evaluated through posterior statistical hypotheses, enabling the recovery of conventional high-probability or expected generalization guarantees—all without requiring distributional assumptions.

generalizationi.i.d.optimization

Error Analysis of Sum-Product Algorithms under Stochastic Rounding

Nov 19, 2024
PD
Pablo de Oliveira Castro
🏛️ Université Paris-Saclay | Université de Rennes | Intel Corp

This paper addresses the forward error analysis of sum-product algorithms under stochastic rounding (SR). We propose a probabilistic error bounding method grounded in martingale theory. Our key contributions are threefold: (1) We introduce the first automated martingale construction framework tailored to multilinear computational structures—encompassing addition, subtraction, multiplication, and intermediate result reuse; (2) We extend SR error analysis to algorithms with structural reuse, notably Karatsuba polynomial multiplication—previously unaddressed in SR literature; (3) Leveraging the Azuma–Hoeffding inequality, we derive a tight probabilistic error bound of $O(sqrt{n},u)$, markedly improving upon the classical worst-case bound $O(n,u)$. Our framework uniformly recovers known error guarantees for pairwise summation and Horner’s method, and—crucially—provides the first rigorous SR error guarantee for Karatsuba multiplication.

Analyzing forward error bounds for numerical algorithmsDeveloping probabilistic error analysis using stochastic roundingGeneralizing martingale methods for multi-linear computations

A Framework for Statistical Inference via Randomized Algorithms

Jul 20, 2023
ZZ
Zhixiang Zhang
🏛️ University of Macau | Columbia University | University of Pennsylvania

This paper addresses the challenge of uncertainty quantification arising from the inherent non-determinism of randomized algorithms—such as random projection and stochastic optimization—by proposing the first asymptotic statistical inference framework that requires no prior distributional assumptions. Methodologically, it establishes the first set of verifiable conditions for asymptotic normality of randomized outputs and introduces three novel strategies: sub-randomization, multi-run plug-in, and multi-run aggregation—integrating multi-run resampling, Polyak–Ruppert averaging, momentum SGD, and randomized sketching. The key contributions are: (i) enabling reliable confidence interval construction for high-dimensional and large-scale stochastic optimization and randomized least squares; and (ii) achieving negligible computational and communication overhead. Extensive simulations demonstrate robustness and practical efficacy on high-dimensional sparse and ultra-large-scale datasets.

Address computational challenges in large dataset analysisDevelop inference methods for non-deterministic algorithm resultsQuantify uncertainty in randomized algorithm outputs

Fast algorithms for least square problems with Kronecker lower subsets

Sep 13, 2022
OA
Osman Asif Malik
🏛️ Encube Technologies | University of Kentucky | University of Colorado Boulder | University of Utah

Computing exact leverage scores for Kronecker-product structured matrices in large-scale least-squares problems is computationally prohibitive, while existing approximation methods incur statistical bias and high overhead. Method: We propose the first efficient exact leverage score algorithm tailored to Kronecker-structured matrices. Leveraging the inherent tensor structure, our method designs a near-linear-time framework for exact leverage score computation and sampling—bypassing costly full-matrix SVD or biased sketching approximations. Contribution/Results: Theoretically and empirically, our algorithm achieves significantly lower sampling error than state-of-the-art approximate methods (e.g., FJLT- or CountSketch-accelerated approaches), while maintaining substantially lower time complexity than full SVD. This work establishes the first scalable, exact, and efficient leverage score sampling scheme for Kronecker-structured matrices, enabling improved structured random projections and large-scale regression.

Efficient exact leverage score sampling for Kronecker subsetsImproving approximation accuracy for structured matricesReducing computational cost in large least squares problems

Optimal sampling for least-squares approximation

Sep 04, 2024
BA
Ben Adcock
🏛️ Simon Fraser University

This work addresses the design of **optimal randomized sampling strategies** for weighted least-squares approximation, extending beyond classical settings restricted to pointwise evaluations and linear approximation spaces to encompass **generalized non-pointwise observations** (e.g., integrals, derivatives) and **nonlinear approximation spaces**. Methodologically, it introduces a systematic generalization of the Christoffel function to the **generalized recovery framework**, yielding a unified theoretical foundation. The resulting sampling scheme achieves near-optimal sample complexity—requiring only $O(n log n)$ measurements for $n$ degrees of freedom—improving upon the classical $O(n^2)$ bound. Theoretical analysis guarantees stable, high-probability reconstruction. The approach integrates tools from approximation theory, randomized sampling design, and numerical linear algebra. Extensive experiments demonstrate its effectiveness and broad applicability in machine learning and scientific computing.

Designing near-optimal random sampling for least-squares approximationExtending sampling strategies to non-linear and non-pointwise settingsUsing Christoffel function to determine sample complexity

Latest Papers

What's happening recently
View more

This study addresses the critical challenge of reliably estimating sharp lower bounds for the standard errors of moment condition estimators when cross-sample correlation information is either absent or only partially available. By leveraging geometric inequalities, the authors derive explicit and tight lower bounds on standard errors and show that the general problem can be reformulated as a semidefinite programming (SDP) problem amenable to efficient computation. This approach yields the first sharp error bounds in settings with no knowledge of cross-sample correlations. Integrating insights from moment condition estimation and statistical inference theory, the method demonstrates both validity and practical utility across several applications, including menu cost models, heterogeneous-agent New Keynesian frameworks, and two-sample instrumental variable settings.

boundscross-sample correlationmoment conditions

This study addresses the problem of accurately computing the output distributions of small-scale programs—such as those processing GPS or inertial sensor data—that involve random inputs. To this end, it introduces cylindrical algebraic decomposition (CAD) into probabilistic program analysis for the first time, combining symbolic and numerical integration to effectively handle conditional branches and nonlinear operations. The approach is grounded in a rigorous semantic model of probabilistic programs and has been validated on both floating-point arithmetic benchmarks and representative programs from open-source sensor libraries, demonstrating its feasibility and effectiveness in deriving exact output distributions.

cylindrical algebraic decompositionoutput distributionprobabilistic programs

This study addresses the computational complexity associated with calculating quantiles of the inverse normal distribution, Student’s t-distribution, and outlier rejection criteria in hypothesis testing. To overcome the reliance on table lookups or iterative numerical methods, the paper proposes concise and highly accurate analytical approximations formulated as closed-form expressions. These approximations significantly reduce computational overhead while maintaining precision sufficient for practical statistical applications. The resulting method offers substantial gains in computational efficiency, making it particularly well-suited for resource-constrained environments or scenarios requiring rapid statistical inference. By bridging theoretical rigor with practical utility, the approach delivers both methodological insight and real-world applicability.

computational simplificationhypothesis testingoutlier rejection

This work addresses the NP-hard problem of penalized least trimmed squares (LTS) regression in robust statistics, for which existing mixed-integer optimization approaches struggle to scale to large datasets. The authors propose a novel formulation that explicitly embeds the arrangement structure of hyperplanes into a perspective reformulation and develop a tailored branch-and-bound algorithm that leverages first-order methods to efficiently solve node relaxations. When the feature dimension is fixed, the method guarantees that the size of the branch-and-bound tree grows polynomially with the sample size, substantially improving both theoretical and practical scalability. Experiments on synthetic data with 5,000 samples and 20 features demonstrate that the proposed approach achieves a 1% optimality gap within one minute, whereas current methods fail to converge within an hour, thereby significantly expanding the tractable scale of exact robust regression.

Least Trimmed SquaresMixed-Integer OptimizationNP-hard Optimization

This work addresses the challenge of evaluating generalization performance in quantized dynamical system identification, where data dependence and non-ideal optimization complicate theoretical analysis. The authors propose a unified framework for statistical error bounds, leveraging a block decomposition technique to derive slow-rate bounds and introducing a novel subsampled interval strategy to establish fast-rate, variance-adaptive bounds. These bounds explicitly link the number of bits used for model quantization to statistical complexity. Notably, this is the first study to incorporate hardware constraints—such as quantization bitwidth—into generalization error theory, offering theoretically grounded, interpretable, and practically actionable guarantees for real-world applications including quantized modeling and hybrid system identification.

dependent datahybrid systemsquantized dynamical models

Hot Scholars

VA

Vaneet Aggarwal

Professor and University Faculty Scholar, Purdue University
Machine LearningReinforcement LearningQuantum ComputingNetworking
AJ

Arnulf Jentzen

The Chinese University of Hong Kong, Shenzhen (CUHK-Shenzhen) & University of Münster
Stochastic AnalysisNumerical AnalysisApplied and Computational MathematicsPDEs
TD

Thang Do

The Chinese University of Hong Kong, Shenzhen
Machine learningProbability theory
MG

Mudit Gaur

Purdue University
Reinforcement Learning
ZZ

Zhihua Zhang

Professor of Computer Science, Shanghai Jiao Tong University
Artificial IntelligenceMachine Learning