auxiliary variable methods

Designs, builds, and analyzes model augmentations and algorithms that introduce auxiliary (latent or augmented) variables to make inference, sampling, or optimization tractable; this includes constructing augmentation schemes, deriving conditional samplers or deterministic transformations, and evaluating their correctness, convergence, mixing behavior, and computational efficiency.

auxiliaryvariablemethods

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.57
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Non-Asymptotic Analysis of Data Augmentation for Precision Matrix Estimation

Oct 02, 2025
LM
Lucas Morisset
🏛️ Qube Research and Technologies | École Polytechnique

This paper addresses the problem of estimating high-dimensional inverse covariance (precision) matrices. To tackle structured dependence among samples—common in modern statistical learning—we propose a novel deterministic equivalent form for the generalized resolvent matrix, unifying linear shrinkage and data augmentation estimators. Our method yields the first non-asymptotically exact characterization of estimation error under data augmentation. Leveraging random matrix theory and generative data transformations, we derive tight, non-asymptotic quadratic error concentration bounds for both classes of estimators. Furthermore, we establish theoretically grounded hyperparameter tuning rules—e.g., for augmentation ratio—that balance bias and variance. All theoretical findings are rigorously validated through comprehensive numerical experiments, demonstrating substantial improvements in estimation stability and interpretability for high-dimensional precision matrices.

Analyzing precision matrix estimation in high-dimensional settingsDeriving concentration bounds for data augmentation estimatorsIntroducing deterministic equivalents for dependent sample analysis

Feature Augmentations for High-Dimensional Learning

Aug 29, 2025
XZ
Xiaonan Zhu
🏛️ Princeton University

High-dimensional features with strong correlations often lead to overparameterization, degrading predictive performance, interpretability, and numerical stability of supervised learning—particularly in Chinese financial news–driven stock return forecasting. To address this, we propose a factor-augmented feature engineering method grounded in factor modeling and principal component analysis (PCA): it separately decomposes the design matrix and its nonlinear transformations to jointly extract shared latent factors and idiosyncratic residuals, which are then combined into augmented features. This approach bridges data augmentation and model architecture modification, offering structural simplicity and computational efficiency. Extensive experiments across diverse real-world datasets—including Chinese financial news text—demonstrate substantial improvements in prediction accuracy and robustness across multiple supervised algorithms, especially under small-sample and high-noise regimes. Our work fills a critical methodological gap by introducing factor-driven feature engineering for NLP-based financial forecasting.

Enhancing supervised learning performance through feature augmentationReducing over-parametrization in high-dimensional correlated dataWeakening input variable correlations to improve interpretability and stability

This work identifies an evaluation bias introduced by data augmentation (e.g., SMOTE, mutation-based augmentation) in scarce-data scenarios—particularly flaky test classification—where augmented samples inadvertently contaminate the test set, severely compromising fairness and reliability assessments. To address this, the authors first empirically identify and validate the critical phenomenon that “augmented data participation in testing” induces systematic evaluation distortion. They then propose a detection framework capable of disentangling training-induced bias from evaluation-induced bias, and design a bias-calibrated evaluation protocol. Experiments across multiple flaky-test benchmark datasets demonstrate that test sets containing augmented samples inflate accuracy by up to 23.7% and introduce F1-score deviations exceeding 0.15. This study establishes both theoretical foundations and practical guidelines for trustworthy model evaluation under data augmentation.

Bias in training and testing with augmented dataEvaluating augmented data effects in model testingImpact of data augmentation on model bias

Inference for Regression with Variables Generated by AI or Machine Learning

Feb 23, 2024
LB
Laura Battaglia
🏛️ University of Oxford | Yale University | University College London | Stanford University

In regression analysis, directly incorporating AI/ML-generated variables—such as imputed labels, nonlinear dimensionality reduction scores, or synthetic indices—as covariates induces estimation bias and invalidates standard errors, thereby compromising statistical inference. This paper is the first to systematically characterize this failure mechanism. We propose two theoretically grounded solutions: (1) a bias-corrected confidence interval that analytically adjusts for the asymptotic bias introduced by ML-based imputation; and (2) a joint estimation framework that simultaneously models latent variables and regression parameters within a two-stage optimization procedure, embedding ML modeling directly into the inferential workflow. Our methods apply broadly to canonical settings including label imputation, nonlinear dimensionality reduction, and index construction. Empirical results demonstrate that the proposed approaches restore consistency of standard errors and achieve nominal coverage of confidence intervals, substantially enhancing the reliability and robustness of regression inference.

Bias in regression using AI-generated variables as dataInvalid inference from naive treatment of ML estimatesNeed methods to correct bias in latent variable regression

Counterfactual Data Augmentation with Contrastive Learning

Nov 07, 2023
AA
Ahmed Aloui
🏛️ Duke University

To address statistical imbalance between treatment groups that biases Conditional Average Treatment Effect (CATE) estimation in causal inference, this paper proposes a model-agnostic counterfactual data augmentation method. It pioneers the integration of contrastive learning into counterfactual reasoning, constructing a representation space that preserves similarity of potential outcomes and enabling precise counterfactual outcome imputation across treatment groups. Theoretically, the method mitigates treatment group distribution shift and suppresses overfitting. Empirical evaluation on synthetic and semi-synthetic benchmarks demonstrates substantial improvements: average RMSE reduction of 18.7% across mainstream CATE estimators, over 30% decrease in generalization error, and enhanced robustness—all without reliance on specific model architectures. The core contribution lies in unifying contrastive learning with counterfactual augmentation, establishing a general, interpretable, low-bias, and high-generalization enhancement paradigm for CATE estimation.

Address statistical discrepancy in CATE estimation groupsImpute missing outcomes using contrastive learning approachReduce treatment group discrepancy with minimal imputation error

Latest Papers

What's happening recently
View more

This study addresses the unclear generalization mechanisms in auxiliary learning by developing an analytical nonlinear network fluctuation-dissipation theory within a teacher-student framework. We derive the online stochastic gradient descent (SGD) dynamical equations to systematically quantify the effects of task relatedness and gradient noise on generalization performance. By combining analytical solutions of differential equations with empirical validation, this work reveals the intrinsic relationship between main-auxiliary task errors and single-task errors, while elucidating the dynamical mechanism through which moderate gradient noise enhances generalization. Ultimately, this research provides a rigorous theoretical foundation for understanding implicit regularization effects in multi-task learning.

Auxiliary learningGeneralization errorLabel noise

This work investigates the regularization effect induced by data augmentation in supervised regression under the high-dimensional regime where both covariate dimension and sample size grow proportionally, and its impact on generalization error. Relying solely on the first- and second-order statistics of the true data distribution and the augmentation scheme, the study leverages random feature regression, high-dimensional statistical analysis, and spectral methods to provide, for the first time, a sharp asymptotic characterization of the generalization error under model misspecification and arbitrary network architectures when only the final layer is trained. The theoretical results are validated for their accuracy in Gaussian settings and quantitatively elucidate the mechanism by which data augmentation enhances generalization performance.

data augmentationgeneralization errorproportional regime

Hot Scholars

YL

Yaru Liu

University of Electronic Science and Technology of China
YG

Yiqi Gu

University of Electronic Science and Technology of China
Applied Mathematics
GL

Guang Lin

Associate Dean for Research, Moses Cobb Stevens Professor in Mathematics, Mech Eng Purdue University
Scientific Machine LearningUncertainty QuantificationGenerative AILLM
AW

Andi Wang

University of Wisconsin-Madison
machine learning for advanced manufacturing
WL

Wenhai Lai

The Chinese University of Hong Kong, Shenzhen
Intelligent Reflecting SurfaceOptimizationMachine Learning