🤖 AI Summary
This study addresses the limitation of conventional learning paradigms, which are confined to deterministic vectors or functions and struggle to learn mappings between probability distributions and stochastic processes. To overcome this, we establish the first uniform neural approximation theory for continuous operators on Wasserstein-2 compact sets, extending it to probability laws over Hilbert spaces. By integrating finite-rank orthogonal projections with task-adaptive architectures, we construct a distribution-to-distribution neural approximation framework based on finite statistics. Theoretically, this work proves the feasibility of model-agnostic representations for distributional transformations. Empirically, the proposed approach significantly outperforms fixed-feature MLPs and kernel regression baselines in distribution space on tasks including first-passage time prediction for Ornstein–Uhlenbeck processes and response path learning for Duffing oscillators, thereby validating the practical effectiveness of distributional learning in stochastic systems.
📝 Abstract
Learning mappings between probability distributions arises naturally when inputs and outputs are represented by populations of samples rather than individual observations. We develop an approximation-theoretic framework for distribution-to-distribution learning and extend it to mappings between stochastic processes. For continuous operators on $W_2$-compact families of finite-dimensional probability laws, we establish uniform neural approximation in the 2-Wasserstein metric using finite law statistics, a simplex-valued neural map, and a shared atomic output support that guarantees valid probability measures. We further extend this principle to probability laws on separable Hilbert spaces through finite-rank orthogonal projections. These results establish the representational feasibility of learning transformations between probability laws rather than deterministic vectors or functions. To demonstrate practical relevance, we study two problems naturally defined at the distribution level: prediction of first-passage-time distributions for an Ornstein--Uhlenbeck process and nonlinear response-path laws of a Duffing oscillator. Because the theory is model-agnostic and broader than any single practical architecture, the experiments use task-adapted neural models rather than reproducing the theoretical construction exactly. In both problems, the proposed models outperform a fixed-feature MLP baseline and distribution-space kernel regression. These experiments complement the theory by demonstrating the practical learnability of distribution-to-distribution transformations in random systems.