Score
Estimation and regularization of variance–covariance structures, propagation of uncertainty, and detection/correction for covariate shift in model spaces. Applied to designing regularizers in fused latent spaces, comparing similarity metrics robustly, and constructing one-step estimators for drift and covariance that respect continuous-time generators.
This paper addresses unbiased functional mean estimation under covariate shift—where training and test data exhibit differing feature distributions but share identical label-conditional distributions. While existing methods are largely restricted to bounded linear or parametric functions, we provide the first rigorous statistical analysis for **arbitrary unknown bounded functions**, establishing a complete theoretical foundation for this problem. We propose a provably correct algorithm that integrates importance weighting, kernel density estimation, and functional approximation. The method achieves optimal sample complexity $O(1/sqrt{n})$ and polynomial-time computational efficiency. Our theoretical guarantees are empirically validated on both synthetic and real-world datasets, demonstrating substantial improvements in estimation accuracy and robustness over prior approaches.
This work addresses unsupervised domain adaptation for functional-output regression under covariate shift by proposing a regularized operator learning framework grounded in vector-valued reproducing kernel Hilbert spaces (vRKHS). The approach ensures stable estimation of functional outputs through constrained hypothesis spaces and establishes minimax optimal convergence rates under source conditioning assumptions. A novel, parameter-free ensemble strategy is introduced to automatically aggregate estimators derived from diverse regularization parameters and kernel functions, thereby enhancing generalization without manual tuning. Empirical validation on real-world facial image datasets demonstrates the method’s robustness and effectiveness, with both theoretical analysis and experimental results confirming its significant mitigation of performance degradation caused by distributional shift.
This study addresses the joint estimation of change points and sparse structures in piecewise constant covariance matrices. The authors propose a regularized approach that combines a squared Frobenius norm loss with Group Fused LASSO for change-point localization and LASSO for inducing sparsity, enhanced by adaptive weights to improve estimation accuracy. Theoretical analysis establishes joint consistency for both change-point locations and the associated segment-wise covariance matrices. An efficient optimization algorithm based on the Alternating Direction Method of Multipliers (ADMM) is developed to solve the resulting problem. Extensive experiments on both synthetic and real-world data demonstrate that the proposed method outperforms existing benchmarks in terms of both estimation accuracy and computational efficiency.
Existing spectral shrinkage methods for high-dimensional covariance estimation rely either on heuristic objectives or strong distributional assumptions. To address this, we propose a geometrically unified framework grounded in distributionally robust optimization (DRO). Our approach defines geometric divergence measures—including KL, Fisher–Rao, and Wasserstein divergences—on the covariance manifold, constructs data-driven ambiguity sets, and automatically derives spectral shrinkage estimators by minimizing the worst-case Frobenius error. Crucially, we establish for the first time that the functional form of shrinkage is intrinsically determined by the geometric structure of the chosen divergence, thereby eliminating dependence on prior distributional assumptions. We prove asymptotic consistency of the proposed estimator and derive an $O(1/n)$ finite-sample error bound. Moreover, we obtain three explicit robust estimators, which match state-of-the-art performance on both synthetic and real-world datasets.
Existing methods for detecting simultaneous changes in the mean vector and covariance matrix of high-dimensional data suffer from reduced detection power and inaccurate change-point localization due to separate modeling of these two types of changes. Method: This paper proposes a unified changepoint detection and localization framework. Its core innovation is the first theoretical demonstration of asymptotic independence between test statistics for mean and covariance changes, enabling p-value fusion via Fisher’s method to construct an adaptive changepoint estimator. Leveraging high-dimensional statistical inference, we design a decouplable joint test statistic and integrate asymptotic distribution theory with p-value combination to achieve integrated detection and precise localization. Results: Theoretical analysis and extensive simulations demonstrate that our method significantly improves detection power and localization accuracy under simultaneous changes—particularly excelling in high-dimensional sparse changepoint settings.
This study addresses the frequent misinterpretation of fluctuations in dominant eigensubspaces and scalar spectral functionals—such as the absorption ratio—in rolling covariance estimation as genuine market structural changes, when they are often artifacts of estimation noise, particularly under shrinkage. By leveraging perturbation analysis and calibrated inference, the work derives, for the first time, the first-order null distribution of eigensubspace variation under overlapping windows and establishes its invariance under rotation-equivariant shrinkage estimators. It further shows that only scale-invariant spectral functionals enjoy first-order immunity to elliptical kurtosis. To correct high-dimensional bias in the absorption ratio, a trace-preserving spiked debiased estimator is proposed. Theoretical results, supported by Davis–Kahan bounds, distribution-free confidence bands, and an estimator-aware bootstrap, are validated through simulations and successfully applied to equity data for reliable detection of true market structural shifts.
This work addresses the high statistical risk inherent in high-dimensional covariance estimation by reframing covariance shrinkage as a parameterized empirical risk minimization problem based on stochastic interpolation between source and target distributions. The proposed approach extends the theoretical foundations of classical shrinkage estimation through a synergistic integration of optimal transport couplings, eigenvector regularization induced by nonlinear flow maps, an early-stopping mechanism grounded in vector field regression, and an adaptive scheduling strategy. By unifying stochastic interpolation, neural estimators, and quadratic risk upper-bound analysis, the method demonstrates strong empirical performance on synthetic data and achieves superior regularization efficacy and estimation accuracy on real neuroimaging datasets.
In high-dimensional linear regression with strongly anisotropic covariates and true coefficients aligned with top eigendirections of the covariance matrix, conventional ℓ₂-shrinkage estimators (e.g., ridge regression) can underperform the minimum-ℓ₂-norm interpolator. This work demonstrates that *moderate amplification*—scaling the minimum-norm interpolator by a constant greater than one—substantially reduces generalization error, challenging the long-held belief that shrinkage universally dominates interpolation. We establish the first tight upper and lower bounds on the generalization error under diverging aspect ratios (n/d → 0) for anisotropic covariance structures, proving that amplified interpolation achieves the optimal statistical rate. Our theoretical analysis, combining data splitting and Gaussian random projection techniques, rigorously characterizes the mechanism and limits of amplification. Empirical results confirm that the amplified interpolator consistently outperforms classical regularized estimators on real-world anisotropic data.
This work addresses the convergence challenge in unsupervised domain adaptation under covariate shift when the target function lies outside the reproducing kernel Hilbert space (i.e., the misspecified setting). By integrating Tikhonov regularization with Nyström subsampling projection, the paper establishes, for the first time, a high-probability excess risk upper bound for Nyström-type domain adaptation methods in this misspecified regime. Leveraging source conditions, effective dimension estimates, and approximation of the Radon–Nikodym derivative, the proposed approach achieves the same convergence rate as in the well-specified setting, requiring only a minimal number of additional samples even when the Radon–Nikodym derivative is unknown.
This work proposes a novel framework based on the matrix log-normal distribution to address the challenges of ensuring positive definiteness and mitigating the curse of dimensionality in modeling high-dimensional time-varying covariance matrices. By assuming that the matrix logarithm of the covariance matrix follows a matrix normal distribution in the space of symmetric matrices, the approach employs a BEKK-type structure to characterize the conditional mean and leverages the matrix exponential map to inherently guarantee positive definiteness without imposing additional constraints. To enhance estimation accuracy, a time-specific second-order Taylor expansion is introduced for bias correction. The resulting method naturally preserves positive definiteness, substantially alleviates the curse of dimensionality, and enables flexible and stable dynamic covariance modeling in high-dimensional settings.