Score
Designs, builds, and analyzes model augmentations and algorithms that introduce auxiliary (latent or augmented) variables to make inference, sampling, or optimization tractable; this includes constructing augmentation schemes, deriving conditional samplers or deterministic transformations, and evaluating their correctness, convergence, mixing behavior, and computational efficiency.
This paper addresses the problem of estimating high-dimensional inverse covariance (precision) matrices. To tackle structured dependence among samples—common in modern statistical learning—we propose a novel deterministic equivalent form for the generalized resolvent matrix, unifying linear shrinkage and data augmentation estimators. Our method yields the first non-asymptotically exact characterization of estimation error under data augmentation. Leveraging random matrix theory and generative data transformations, we derive tight, non-asymptotic quadratic error concentration bounds for both classes of estimators. Furthermore, we establish theoretically grounded hyperparameter tuning rules—e.g., for augmentation ratio—that balance bias and variance. All theoretical findings are rigorously validated through comprehensive numerical experiments, demonstrating substantial improvements in estimation stability and interpretability for high-dimensional precision matrices.
High-dimensional features with strong correlations often lead to overparameterization, degrading predictive performance, interpretability, and numerical stability of supervised learning—particularly in Chinese financial news–driven stock return forecasting. To address this, we propose a factor-augmented feature engineering method grounded in factor modeling and principal component analysis (PCA): it separately decomposes the design matrix and its nonlinear transformations to jointly extract shared latent factors and idiosyncratic residuals, which are then combined into augmented features. This approach bridges data augmentation and model architecture modification, offering structural simplicity and computational efficiency. Extensive experiments across diverse real-world datasets—including Chinese financial news text—demonstrate substantial improvements in prediction accuracy and robustness across multiple supervised algorithms, especially under small-sample and high-noise regimes. Our work fills a critical methodological gap by introducing factor-driven feature engineering for NLP-based financial forecasting.
This work identifies an evaluation bias introduced by data augmentation (e.g., SMOTE, mutation-based augmentation) in scarce-data scenarios—particularly flaky test classification—where augmented samples inadvertently contaminate the test set, severely compromising fairness and reliability assessments. To address this, the authors first empirically identify and validate the critical phenomenon that “augmented data participation in testing” induces systematic evaluation distortion. They then propose a detection framework capable of disentangling training-induced bias from evaluation-induced bias, and design a bias-calibrated evaluation protocol. Experiments across multiple flaky-test benchmark datasets demonstrate that test sets containing augmented samples inflate accuracy by up to 23.7% and introduce F1-score deviations exceeding 0.15. This study establishes both theoretical foundations and practical guidelines for trustworthy model evaluation under data augmentation.
In regression analysis, directly incorporating AI/ML-generated variables—such as imputed labels, nonlinear dimensionality reduction scores, or synthetic indices—as covariates induces estimation bias and invalidates standard errors, thereby compromising statistical inference. This paper is the first to systematically characterize this failure mechanism. We propose two theoretically grounded solutions: (1) a bias-corrected confidence interval that analytically adjusts for the asymptotic bias introduced by ML-based imputation; and (2) a joint estimation framework that simultaneously models latent variables and regression parameters within a two-stage optimization procedure, embedding ML modeling directly into the inferential workflow. Our methods apply broadly to canonical settings including label imputation, nonlinear dimensionality reduction, and index construction. Empirical results demonstrate that the proposed approaches restore consistency of standard errors and achieve nominal coverage of confidence intervals, substantially enhancing the reliability and robustness of regression inference.
To address statistical imbalance between treatment groups that biases Conditional Average Treatment Effect (CATE) estimation in causal inference, this paper proposes a model-agnostic counterfactual data augmentation method. It pioneers the integration of contrastive learning into counterfactual reasoning, constructing a representation space that preserves similarity of potential outcomes and enabling precise counterfactual outcome imputation across treatment groups. Theoretically, the method mitigates treatment group distribution shift and suppresses overfitting. Empirical evaluation on synthetic and semi-synthetic benchmarks demonstrates substantial improvements: average RMSE reduction of 18.7% across mainstream CATE estimators, over 30% decrease in generalization error, and enhanced robustness—all without reliance on specific model architectures. The core contribution lies in unifying contrastive learning with counterfactual augmentation, establishing a general, interpretable, low-bias, and high-generalization enhancement paradigm for CATE estimation.
本文研究了利用大量辅助样本增强小目标样本的问题,对比了IPW和FL方法,发现FL方法在某些情况下可实现全效率增益。
本文提出了一种学习者无关的框架,通过结合机器学习和设计意识交叉拟合来改进调查数据中有限总体参数估计的有效性。
该论文综述了学习增强算法在处理不可靠预测时保持性能保证的方法,探讨了其构建机制、误差度量及系统层面的影响。
This study addresses the unclear generalization mechanisms in auxiliary learning by developing an analytical nonlinear network fluctuation-dissipation theory within a teacher-student framework. We derive the online stochastic gradient descent (SGD) dynamical equations to systematically quantify the effects of task relatedness and gradient noise on generalization performance. By combining analytical solutions of differential equations with empirical validation, this work reveals the intrinsic relationship between main-auxiliary task errors and single-task errors, while elucidating the dynamical mechanism through which moderate gradient noise enhances generalization. Ultimately, this research provides a rigorous theoretical foundation for understanding implicit regularization effects in multi-task learning.
This work investigates the regularization effect induced by data augmentation in supervised regression under the high-dimensional regime where both covariate dimension and sample size grow proportionally, and its impact on generalization error. Relying solely on the first- and second-order statistics of the true data distribution and the augmentation scheme, the study leverages random feature regression, high-dimensional statistical analysis, and spectral methods to provide, for the first time, a sharp asymptotic characterization of the generalization error under model misspecification and arbitrary network architectures when only the final layer is trained. The theoretical results are validated for their accuracy in Gaussian settings and quantitatively elucidate the mechanism by which data augmentation enhances generalization performance.