Score
Design and analyze learning algorithms and models that explicitly monitor and control the numerical conditioning of learned transforms, implementing regularizers or constraints (for example condition-number penalties or spectral constraints) to enforce stability of forward and inverse operators. Balance approximation error and generalization while ensuring the learned transform remains well-conditioned for reliable optimization, inversion, and deployment.
This work systematically analyzes three fundamental error sources in the linear layers of Fourier Neural Operators (FNOs): statistical error arising from finite samples, rank-approximation error due to spectral truncation, and discretization error induced by grid resolution. It establishes, for the first time, a unified theoretical framework that rigorously models and characterizes both the individual origins and their coupling mechanisms. A discrete Fourier transform (DFT)-based least-squares estimator is constructed, and a comprehensive generalization error analysis framework—yielding tight upper and lower bounds—is developed. The theoretical analysis delivers explicit, quantitative bounds on all three error components, precisely characterizing how sample size, spatial grid resolution, and truncation order jointly govern generalization performance. This work provides the first operator-learning theory for FNOs that simultaneously incorporates statistical and numerical perspectives, thereby filling a critical gap in error decomposition and controllability analysis for neural operators.
This work addresses operator learning for partial differential equation (PDE) solution operators and black-box physical systems, formalizing it as a regression problem between function spaces. Method: We propose the first systematic statistical learning framework for operators, integrating PDE-based physical priors with neural operator architectures and introducing a constraint-aware training paradigm. Contribution/Results: Our framework unifies and characterizes the statistical foundations of mainstream operator learning methods, significantly enhancing model interpretability and generalization. It is the first to explicitly incorporate active learning and uncertainty quantification as core components of operator learning. The resulting methodology provides a new paradigm for scientific machine learning—rigorous in theory and practical in implementation—that enables high-fidelity, data-efficient surrogate modeling of complex physical systems.
This work addresses the family of parametric optimization problems and proposes the first unified, data-driven framework for analyzing the generalization performance of both classical and learned optimizers. Methodologically: (1) it introduces PAC-Bayes theory to the analysis of learned optimizers, deriving verifiable, high-probability generalization upper bounds; (2) it establishes performance bounds for classical optimizers based on empirical convergence rates; and (3) it pioneers a learning paradigm that directly minimizes the PAC-Bayes bound during training. Evaluated on signal processing, control, and meta-learning tasks, the derived bounds are significantly tighter than conventional worst-case guarantees. Moreover, the theoretical generalization guarantees for learned optimizers consistently exceed the empirical performance of their non-learned baselines—thereby unifying theoretical rigor with practical efficacy.
This work addresses aliasing errors introduced in the discrete implementation of Fourier Neural Operators (FNOs), which remain theoretically unquantified despite their practical significance. Specifically, the discrepancy between the continuous FNO formulation and its grid-based discretization has not been systematically characterized, and isolating this discretization error from other sources—such as approximation and optimization errors—remains challenging. Method: Leveraging tools from Fourier analysis, numerical functional analysis, and FFT theory, we derive an explicit algebraic convergence rate for the discretization error with respect to grid resolution and establish its quantitative dependence on the Sobolev regularity of the input function. Contribution/Results: We provide the first verifiable upper bound on this error, revealing intrinsic trade-offs among resolution, input smoothness, and model stability. Numerical experiments confirm the theoretical prediction: for inputs with higher Sobolev regularity, the error decays significantly under mesh refinement—thereby furnishing a rigorous foundation for reliable discrete FNO design.
For black-box models applied to complex tasks such as image segmentation, defining meaningful conditional events is challenging, leading to uncertainty estimates that fail to reflect inherent sample difficulty. Method: This paper proposes an input-dependent statistical risk control framework grounded in conformal prediction. It introduces a novel, algorithm-driven mechanism for dynamically selecting conditional function classes—bypassing manual discretization—by adaptively constructing these classes based on test-sample difficulty and integrating online parameter tuning for fine-grained, approximately conditional risk control. Contribution/Results: Experiments on regression and image segmentation demonstrate substantial improvements in uncertainty calibration accuracy. The method guarantees strict statistical risk control while enhancing generalization robustness and predictive reliability.
Traditional fixed analysis transforms struggle to capture the sparse structures inherent to specific signal classes due to their lack of data adaptivity. This work proposes an explicitly conditioned doubly-sparse transform that multiplies a fixed, well-conditioned matrix with a data-adaptive sparse component, thereby introducing controllable data adaptivity while preserving fast and stable computation. We devise a structured learning approach incorporating condition number control to balance generalization and approximation accuracy, and develop a novel closed-form projection operator within an inexact proximal framework for efficient optimization. The resulting method achieves state-of-the-art performance in doubly-sparse transform learning, significantly reducing computational cost compared to dense variants, converging faster, and more effectively avoiding poor local minima.
本文提出了一种从数据中学习控制系统的线性算子的结构化方法,利用(半)群框架和反问题框架分析算法,以获得收敛估计器。
This work addresses the unclear relationship between continuous theory and discrete implementation in neural operators for solving partial differential equations, particularly concerning stability and discretization error. For the first time, rigorous discretization error bounds are established for State-Space Neural Operators (SS-NOs) and Fourier Neural Operators (FNOs). By leveraging functional analysis, the regularity of solutions is explicitly linked to input discretization, and Input-to-State Stability (ISS) theory is introduced to quantify how discretization affects stability in the continuous domain. Numerical experiments on one- and two-dimensional benchmark problems validate the tightness of the derived theoretical bounds, demonstrating that SS-NOs exhibit both robustness and numerical stability across varying resolutions.
This work addresses the absence of generalization error theory for multi-input neural operators in Sobolev spaces, particularly when input functions are defined on heterogeneous domains with differing dimensions and regularity. The paper establishes the first unified Sobolev generalization framework by integrating function approximation theory, Sobolev space analysis, and statistical learning theory. It derives complexity-dependent approximation and generalization error bounds of logarithmic-logarithmic over logarithmic type, quantifies the contribution of each input space to the overall error, and reveals the coupling mechanism among input dimensionality, regularity, and Sobolev smoothness order. The resulting theory is applicable to operator learning tasks in PDE solving and scientific computing, accurately characterizing the impact of multi-input interactions on learning performance under balanced settings.
本文通过算法稳定性建立样本外边界,利用耗散性论点为学习动力系统提供了一种系统理论解释,并通过优化算法依赖的动力增益来认证和比较学习动态的泛化能力。