robust mean estimation

Designs and implements estimators and algorithms for the central tendency (e.g., the mean) that remain accurate in the presence of outliers, heavy tails, or adversarial contamination; builds concrete robust estimators and procedures. Analyzes these methods theoretically by proving finite-sample error bounds, contamination-robust guarantees (such as under the Huber contamination model), breakdown points, and conditions for attaining optimal parametric estimation rates.

robustmeanestimation

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.17
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

The Broader Landscape of Robustness in Algorithmic Statistics

Dec 03, 2024
GK
Gautam Kamath
🏛️ University of Waterloo | Vector Institute

This work addresses the robustness of mean estimation in statistical learning under three concurrent challenges: adversarial data contamination, heavy-tailed distributions, and differential privacy constraints. Methodologically, it unifies robust statistics, high-dimensional geometry, stochastic optimization, and differential privacy theory to establish the first conceptual and algorithmic bridge across distinct robustness paradigms. Key technical abstractions—including iterative filtering, covariance trimming, and fractional gradient descent—are identified as common algorithmic primitives. The paper proposes a suite of computationally efficient estimators achieving statistically optimal convergence rates; each attains the information-theoretic lower bound under all three constraint classes simultaneously. By reconciling theoretical tightness with practical efficiency, this framework advances robust mean estimation from ad hoc heuristics toward a principled, unified design paradigm.

Addresses mean estimation with heavy-tailed dataDevelops robust estimators for contaminated datasetsEnsures privacy preservation in statistical methods

Confidence Intervals for Linear Models with Arbitrary Noise Contamination

Nov 10, 2025
DX
Dong Xie
🏛️ University of Chicago | Yale University

This paper addresses the problem of constructing confidence intervals for linear regression coefficients under the Huber contamination model, where an unknown fraction ε of arbitrary outliers invalidates conventional inference. To tackle this challenge, we propose a novel Z-estimation-based algorithm: first, decorrelating covariates to mitigate design matrix correlation; second, employing a smoothed estimating function to achieve adaptive robust estimation of the regression parameters without prior knowledge of ε. The resulting confidence intervals attain uniform coverage at level (1−α) over all ε-contaminated distributions and achieve optimal width O(1/√[n(1−ε)²]), substantially improving upon existing adaptive methods. Our key contribution lies in the first unified integration of the Z-estimation framework, smoothing techniques, and covariate decorrelation—enabling consistent, efficient, and robust statistical inference under unknown contamination.

Achieving optimal interval length matching known contamination rate performanceConstructing confidence intervals for linear regression with arbitrary noise contaminationDeveloping inference methods without prior knowledge of contamination proportion

Efficient Multivariate Robust Mean Estimation Under Mean-Shift Contamination

Feb 20, 2025
ID
Ilias Diakonikolas
🏛️ University of Wisconsin-Madison | University of California, San Diego

This paper addresses robust mean estimation for high-dimensional Gaussian distributions under constant-fraction mean-shift contamination. To overcome the exponential time complexity and poor sample efficiency of existing methods, we propose the first polynomial-time algorithm: it formulates the problem via moment-constrained optimization and integrates spectral-analysis-driven iterative filtering with convex programming to achieve statistically optimal mean estimation. Under a constant contamination rate—i.e., an arbitrary but fixed fraction of outliers—the algorithm attains near-optimal sample complexity, guarantees estimation error convergence to arbitrary precision, and reduces computational complexity from exponential to polynomial time. Our key contribution is breaking the long-standing trade-off between computational tractability and statistical accuracy, delivering the first solution for high-dimensional robust estimation that simultaneously achieves computational efficiency, robustness to adversarial contamination, and statistical optimality.

Efficient robust mean estimation algorithmHandles high-dimensional mean-shift contaminationTolerates constant fraction of outliers

This work addresses robust regression under a setting where covariates remain uncontaminated while responses suffer from adaptive corruption, overcoming the fundamental limitation of inconsistent estimation inherent in the classical Huber contamination model. By leveraging clean covariate information, the authors devise a novel estimator that achieves consistency even when the contamination fraction is constant, and for the first time establish an information-theoretically optimal convergence rate that improves upon the Huber model under this regime. The theoretical analysis derives minimax lower bounds via Fano’s inequality combined with a multi-distribution simultaneous contamination construction. Furthermore, using Statistical Query and low-degree polynomial methods, the study reveals a substantial statistical–computational gap: the optimal statistical rate cannot be attained by any polynomial-time algorithm.

adaptive contaminationclean covariatesinformation-computation gap

Adaptive Robust Confidence Intervals

Oct 30, 2024
YL
Yuetian Luo
🏛️ University of Chicago

This paper addresses the problem of constructing adaptive robust confidence intervals for the Gaussian mean under the Huber contamination model with unknown contamination proportion. It establishes, for the first time, that any adaptive interval must suffer an exponential degradation in length compared to the non-adaptive optimal interval achievable when the contamination level is known—and crucially, this fundamental lower bound depends intrinsically on the shape of the underlying data distribution. To overcome this challenge, the authors propose a novel construction based on joint uncertainty quantification over all contamination levels via full-level quantiles. They develop a complete optimal adaptive theory for the Gaussian location model and a broad class of generalized robust tests, deriving tight information-theoretic lower bounds and providing a computationally feasible, rate-optimal procedure. The core contribution lies in characterizing the distribution-dependent price of adaptivity and delivering the first framework that simultaneously achieves theoretical optimality and practical implementability.

Construct adaptive confidence intervals under unknown contamination proportionDetermine optimal length for robust Gaussian mean confidence intervalsExtend results to general robust hypothesis testing beyond Gaussian models

Latest Papers

What's happening recently
View more

This work addresses the problem of robust mean estimation under the mean-shift contamination model, where an adversary replaces a small fraction of clean samples with arbitrarily shifted samples from the underlying distribution. For general multivariate base distributions whose characteristic functions satisfy mild spectral conditions, the authors propose an efficient Fourier-analytic algorithm capable of achieving mean estimation with arbitrary accuracy. The key innovation lies in the introduction of a “Fourier witness” as a central analytical tool, which enables the first nearly tight characterization of sample complexity for this setting. The derived upper and lower bounds match up to constant factors, thereby revealing the optimal sample efficiency achievable by any robust estimator under this contamination model.

characteristic functionmean estimationmean-shift contamination

This study addresses the sensitivity of classical two-stage least squares (2SLS) estimation to outliers, where even a small fraction of contamination can severely distort inference. To overcome this limitation, the authors propose the W-2SLS estimator, which replaces sample means with quantile-based Winsorized means, achieving robustness under an adversarial contamination model that allows both the identity and magnitude of outliers to depend on clean data. The work establishes, for the first time, the minimax optimal convergence rate of $\eta_n^{1-1/m} + n^{-1/2}$ for this estimator and proves a matching lower bound. It further demonstrates that when $\sqrt{n}\,\eta_n^{1-1/m} \to 0$, the estimator incurs no first-order efficiency loss and enjoys sub-Gaussian bias guarantees. Additionally, the paper develops a heteroskedasticity-robust Winsorized Anderson–Rubin test, enabling valid inference even in the presence of weak instruments.

2SLSAdversarial ContaminationInstrumental Variables

This study addresses a fundamental limitation of conventional subsampling estimators in dynamic time series: even when contamination locations are known, these methods fail to guarantee consistency due to a structural incompatibility with residual propagation—removing contaminated observations distorts the estimation criterion and prevents recovery of the original objective function. The paper is the first to formally identify this issue and proposes a novel index-set transformation method compatible with residual propagation. By introducing a “patch-removal operator,” the approach explicitly controls the residual footprint of contamination. Built upon rigorous residual propagation analysis and asymptotic theory, the resulting robust framework applies broadly to residual-based estimators, preserving asymptotic equivalence to the oracle estimator under clean data while achieving consistent estimation of the true parameters even in the presence of contamination.

consistencydynamic contaminationresidual propagation

This study investigates the robust stability of statistical estimators under η-fraction data corruption. To this end, it introduces “empirical sensitivity” as a novel metric quantifying an estimator’s susceptibility to data perturbations, with a focus on Gaussian mean estimation. Theoretical analysis establishes, for the first time, a lower bound of Ω(η + √(ηd/n)) on empirical sensitivity and demonstrates its tightness up to logarithmic factors. By integrating the Efron–Stein inequality with recent advances in robust mean estimation algorithms, the authors construct an estimator achieving a matching upper bound, thereby confirming the attainability of this limit. This work provides a new theoretical framework and quantitative tool for analyzing the stability of robust estimators.

data perturbationempirical sensitivityGaussian mean estimation

This study addresses the sensitivity of traditional semiparametric mixture regression models to outliers and heavy-tailed errors, which stems from their reliance on Gaussian error assumptions and often leads to non-robust estimation. To overcome this limitation, the paper introduces, for the first time, a contaminated Gaussian distribution into the semiparametric mixture regression framework, yielding a model that simultaneously achieves robustness, flexibility, and clustering capability. Parameter estimation is carried out via an EM algorithm, while the ECM algorithm combined with local kernel likelihood methods handles the nonparametric components. This integrated approach enables concurrent model fitting, cluster analysis, and outlier detection. Extensive simulations and real-data analyses demonstrate that the proposed method substantially enhances robustness against outliers without compromising model expressiveness, offering strong theoretical and practical value.

Gaussian assumptionmixture of regressionsoutliers

Hot Scholars

ID

Ilias Diakonikolas

University of Wisconsin-Madison
theoretical computer sciencealgorithmic statisticsmachine learningprobability theory
MP

Marc Pollefeys

Professor of Computer Science, ETH Zurich, and Director Spatial AI Lab, Microsoft
Computer VisionComputer GraphicsRoboticsMachine Learning
JW

Junfeng Wu

Huazhong University of Science and Technology
Computer Vision
MH

Mia Hubert

Professor of Statistics, KU Leuven
Robust statisticsOutlier detectionDepth
SM

Sreejeet Maity

Ph.D. Student in Electrical and Computer Engineering, North Carolina State University
Reinforcement LearningStochastic ApproximationControl TheoryOptimization