Score
Designs and implements estimation procedures that partition a linear operator or regression problem into multiscale or block-structured components (e.g., wavelet blocks) and compute least-squares estimates separately for each block, including choices for truncation to finite resolution. Analyzes and builds blockwise least-squares estimators and multiscale regression routines, proving their statistical properties (e.g., convergence and minimax rates over Sobolev-type classes).
This work addresses the statistical and computational challenges of learning bounded linear operators between Sobolev spaces from noisy data. The authors reformulate operator learning as an infinite-dimensional matrix regression in wavelet coordinates and propose a finite-resolution, block-wise least squares estimator that exploits multiscale structure. To account for the heterogeneous local estimation difficulty across scales, they introduce a scale-adaptive sampling strategy. Under the Sobolev operator norm loss, the proposed method achieves the minimax optimal convergence rate while attaining the best possible computational complexity among dense least squares–type algorithms.
This study addresses the lack of a unified formulation for scalar, multivariate, and functional regression models, which obscures their intrinsic connections. By leveraging an integral operator defined with respect to general measures, the authors propose a unified framework that subsumes all three regression types as special cases of the same operator under different input and output measures. This framework reveals classical regression forms as measure-dependent manifestations of a single operator, clarifies discretized modeling as operator estimation under specific measures, and explains the efficacy of vectorized multivariate regression in linear settings. Theoretically, the authors prove that discrete representations correspond exactly to operator evaluations under discrete measures and converge to the continuous case as the discretization grid refines; moreover, this estimator is equivalent to standard multivariate regression and inherits its classical statistical properties.
This work addresses the design of **optimal randomized sampling strategies** for weighted least-squares approximation, extending beyond classical settings restricted to pointwise evaluations and linear approximation spaces to encompass **generalized non-pointwise observations** (e.g., integrals, derivatives) and **nonlinear approximation spaces**. Methodologically, it introduces a systematic generalization of the Christoffel function to the **generalized recovery framework**, yielding a unified theoretical foundation. The resulting sampling scheme achieves near-optimal sample complexity—requiring only $O(n log n)$ measurements for $n$ degrees of freedom—improving upon the classical $O(n^2)$ bound. Theoretical analysis guarantees stable, high-probability reconstruction. The approach integrates tools from approximation theory, randomized sampling design, and numerical linear algebra. Extensive experiments demonstrate its effectiveness and broad applicability in machine learning and scientific computing.
This work investigates the impact of substituting the true probability matrix with its estimator in the stochastic block model, which introduces both simple and compound perturbations that affect the asymptotic behavior of spectral statistics. The authors develop a unified decomposition framework that precisely disentangles the compound perturbation into three components: the bias of the normalized adjacency matrix, the simple perturbation, and the bias inherent to the compound perturbation itself, thereby revealing distinct asymptotic behaviors of cross terms under the two types of perturbations. Leveraging spectral analysis, matrix perturbation theory, and trace techniques, they establish, for the first time, the asymptotic normality of linear spectral statistics under an optimal growth condition on the number of communities, relaxing the previous requirement from \(K = O(n^{1/6 - \tau})\) to the sharp rate \(K = o(n^{1/6})\).
This work addresses the computational inefficiency in Bayesian semiparametric regression arising from complex design matrix structures. To mitigate this challenge, the authors propose an orthogonalization preprocessing step applied to the design日晚间 matrix prior to iterative inference, combined with a hybrid algorithm integrating Gibbs sampling and coordinate ascent variational inference. This approach reduces computational complexity to quadratic in the number of covariates, substantially accelerating both model fitting and posterior inference. Empirical evaluations across diverse experimental settings demonstrate speedups ranging from 5× to 60× compared to conventional methods, effectively alleviating the computational bottleneck induced by high-dimensional covariates.
本文扩展了有限维线性逆问题中最小均方误差估计器的形式至无限维情况,使用广义样条和广义高斯过程解决该类问题。
本文提出了一种基于新型多尺度核框架函数逼近技术的多尺度算子学习方法,用于求解多尺度偏微分方程,并在文献中的难题上展示了其优越性。
Block-structured latent variable models are widely employed in psychology, education, economics, and genetics, yet their identifiability and estimation performance have long lacked a systematic theoretical foundation. This work establishes, for the first time, identifiability conditions for such models under various block designs and introduces a Lagrangian-type nonconvex optimization framework based on constrained maximum likelihood estimation. The study derives both non-asymptotic error bounds and asymptotic distributions for the resulting estimators. The proposed algorithm enjoys strong theoretical guarantees and, as demonstrated through extensive simulations and empirical analyses, efficiently and accurately estimates latent variable models across diverse block structures.
This work addresses the grouped distributionally robust (GDR) least squares problem, which seeks to minimize the worst-case loss across multiple data groups. The authors introduce block Lewis weights—a novel geometric tool—to reformulate the problem as a specially weighted least squares instance. By integrating an accelerated proximal algorithm with a structured linear system solver tailored for systems of the form \(A^\top B A\), they achieve an efficient solution method that unifies optimization frameworks for both average and robust losses. The proposed approach outperforms interior-point methods at moderate accuracy levels. Theoretically, it attains a \((1+\varepsilon)\)-approximate solution using only \(\widetilde{O}(\min\{\mathrm{rank}(A), m\}^{1/3} \varepsilon^{-2/3})\) linear system solves, yielding the current best-known guarantee for the special case of \(\ell_\infty\) regression.
This work addresses the challenges of discovering differential equations, approximating functions, and estimating high-dimensional integrals from noisy, non-uniformly sampled data by introducing Sparse Orthogonal Regression Technique (SORT). The method reformulates equation discovery as a spectral coefficient learning problem, directly inferring coefficients of an orthogonal basis expansion from observational data via L1 regularization—without requiring a predefined symbolic library, explicit numerical integration, or inner product evaluations. Its core innovation lies in treating basis function design as the central modeling choice, enabling consistent model order scaling and multi-task reusability. Experiments demonstrate that when the basis functions align with the underlying problem structure, SORT matches or outperforms existing approaches under sparse sampling, noisy derivatives, and representation mismatch, while low-order dominant coefficients remain stable as model complexity increases.