Score
Design and analyze parameterized implicit stochastic modulation modules—mechanisms that map parameters and random noise or latent variables to families of representations or stochastic outputs. These constructions focus on decoupling parameters from spatial inputs, preserving unbiased stochastic derivative estimators, enforcing a variance-aware Lipschitz envelope for stability, and enabling zero-shot generalization across parameter settings.
This work addresses the challenges of solving high-dimensional, high-order parametric partial differential equations (PDEs), which suffer from prohibitive computational costs and the need for repeated training in existing stochastic methods. The authors propose PRISM, a novel framework that decouples parameterization from high-order spatial differentiation by implicitly mapping physical parameters—via a hypergenerator—into affine modulators acting on a purely spatial latent manifold. This approach preserves unbiased stochastic estimation, incorporates a variance-aware Lipschitz envelope, and provides theoretical guarantees on error bounds and convergence. PRISM enables zero-shot generalization to unseen parameters, drastically reduces memory consumption, scales to problems with up to one hundred thousand dimensions on a single GPU, and supports efficient low-rank adaptation without retraining.
This work investigates the impact of parameter noise injection in stochastic gradient descent on optimization and generalization, emphasizing the need for efficient per-sample perturbations and sophisticated noise schemes. By leveraging distributional identities of linear layers, the authors propose a method that enables per-sample noise injection within mini-batches without disrupting batched computation. They systematically compare isotropic and diagonal Gaussian noise variants, demonstrating that on CIFAR-100, a lightweight single-sample isotropic Gaussian perturbation recovers most of the optimization and generalization benefits achieved by more complex multi-sample strategies. These findings suggest that simplified noise injection designs can be sufficiently effective, offering a practical alternative to computationally heavier approaches while maintaining performance gains.
This work addresses the optimization of parameterized implicit stochastic diffusion processes. We propose the first single-loop, first-order framework that jointly optimizes model parameters and performs sampling, formalizing sampling as an optimization problem over the space of probability distributions. Our approach innovatively integrates bilevel optimization with automatic implicit differentiation: by leveraging implicit differentiation, we perform end-to-end gradient-based optimization of diffusion parameters without explicitly differentiating through the inner sampling iterations. We establish theoretical convergence guarantees for the proposed algorithm. Empirically, on energy-based model training and denoising diffusion fine-tuning tasks, our method achieves significant improvements in both optimization efficiency and generative performance, while maintaining computational feasibility and statistical validity.
This work addresses the structural identifiability of parameters in stochastic differential equation (SDE) models under multiple interventions—i.e., whether SDE parameters can be uniquely recovered from samples of post-intervention stationary distributions. Theoretically, we establish the first uniqueness guarantee for SDE parameter recovery under multi-intervention settings; for linear SDEs, we derive a tight lower bound on the minimum number of required interventions; for weak-noise nonlinear SDEs, we obtain an upper bound on identifiability. Methodologically, we propose a parametric framework featuring learnable activation functions, integrating intervention modeling, stationary distribution analysis, and weak-noise asymptotic theory. Experiments on synthetic data demonstrate that our approach accurately recovers ground-truth parameters, and the theory-guided learnable architecture significantly improves both estimation accuracy and robustness.
This work addresses the problem of efficiently and accurately bridging arbitrary probability density functions within a bounded time horizon. Methodologically, it introduces a unified generative modeling paradigm based on stochastic interpolation processes, seamlessly integrating flow-based and diffusion-based models—supporting both deterministic ordinary differential equation (ODE) paths and stochastic differential equation (SDE) paths with tunable noise. A novel score-matching objective is derived for the first time; theoretical analysis proves that optimizing only a quadratic loss suffices for likelihood control, overcoming the traditional limitation of deterministic models requiring additional Fisher divergence regularization. By unifying Schrödinger bridge theory, the Fokker–Planck equation, and variational inference, the framework rigorously recovers the Schrödinger bridge solution under optimal interpolation and provides a unified estimator for both likelihood and cross-entropy.
This work addresses the challenge of estimating statistical moments—such as means and covariances—for continuous-time Markov chains, which typically lack closed-form solutions and incur prohibitive computational costs when using conventional Monte Carlo methods in multi-parameter settings. The authors propose a simulation-based surrogate modeling framework that employs neural networks to learn the mapping from model parameters to statistical moments using Monte Carlo–noisy simulation data, thereby enabling efficient approximation across the entire parameter space. A key contribution lies in uncovering the heterogeneous impact of simulation noise on moment estimation, which informs a tailored allocation of computational resources and a robust learning algorithm. Under a fixed computational budget, the method yields accurate population-level moment estimates and demonstrates significant improvements over baseline approaches in downstream tasks such as whitening.
This study addresses the challenge of parameter identification in stochastic systems with mixed noise, where intractable likelihoods hinder conventional estimation. We propose Penn-GMD, a network that maps trajectories to full-covariance Gaussian mixture distributions and employs surjective parameterization with negative log-likelihood minimization to approximate the true likelihood. The core contribution lies in leveraging full covariance matrices to explicitly reveal parameter coupling and multimodal structures. This approach not only accurately recovers the underlying likelihood distribution but also naturally diagnoses unidentifiability. Consequently, Penn-GMD effectively resolves persistent difficulties in parameter estimation and uncertainty quantification for complex stochastic systems where traditional methods typically fail, offering a robust framework for handling intractable inference problems in noisy dynamical environments.
This work addresses the abstraction gap between differentiable programming and emerging probabilistic hardware by introducing Parameterized Stochastic Circuits (PSCs) as a gate-level intermediate representation that closely mirrors native hardware operations. PSCs unify explicit binary, categorical, and continuous signals with localized stochastic kernels. Building on this foundation, the authors develop torx, an open-source JAX framework that enables, for the first time, direct alignment between differentiable stochastic computation and probabilistic hardware, substantially reducing mapping overhead. Experiments on the X0 subthreshold CMOS probabilistic bit chip demonstrate that PSCs efficiently harness physical randomness in tasks such as graph random walks, discrete diffusion, and stochastic graph networks, achieving results consistent with pseudorandom software baselines while confirming the approach’s validity and hardware energy efficiency.
This study addresses the challenge of parameter inference in ordinary differential equation models arising from structural non-identifiability. It introduces, for the first time, an explicit integration of structural identifiability analysis into the design of Markov chain Monte Carlo (MCMC) algorithms, proposing two novel sampling strategies: one constructs efficient proposals both within and orthogonal to the non-identifiable manifold, while the other performs inference in a low-dimensional space of identifiable parameter combinations and subsequently reconstructs the full parameter vector. By combining geometric MCMC with pseudo-marginal MCMC techniques, the method establishes a Bayesian inference framework tailored to equivalence solution manifolds. This approach significantly enhances sampling efficiency and convergence speed compared to standard MCMC methods, while preserving posterior correctness and chain ergodicity.
This work addresses the challenge of efficiently capturing structured dynamical behaviors, which traditional approaches to dynamical system learning often fail to achieve due to their reliance on high-complexity nonlinear function approximators. The authors propose a structure-first modeling paradigm that integrates wave-inspired causal interaction mechanisms with explicit state evolution units, yielding a fully explicit, hierarchical dynamical architecture free of algebraic loops. Notably, the model dispenses with implicit solvers and generates effective internal representations by training only the readout layer. Experimental results demonstrate that as model depth increases, both representation quality and generalization performance improve significantly on nonlinear system identification tasks, thereby validating the efficacy of prioritizing interaction structure in model design.