Kolmogorov Regression for Robust Diffusion Policies

📅 2026-06-16
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work addresses temporal drift in existing finite-dimensional diffusion policies—caused by discretization artifacts—that hinders their performance in long-horizon physical tasks. The authors elevate policy modeling to the Cameron–Martin space, leveraging Gaussian measure theory to construct a colored-noise covariance operator that enhances trajectory regularity. They reformulate stochastic score matching as a deterministic boundary-value PDE problem grounded in the Kolmogorov backward equation. The proposed framework introduces a precision-weighted Cameron–Martin loss and PDE residual diagnostics, yielding dimension-independent convergence guarantees and enabling reward-free anomaly detection. Experiments demonstrate a 17% increase in maximum episode reward and a 67.6% reduction in inter-step drift on the PushT task; in manufacturing line scheduling, it achieves a 28.4% lower RMSE, perfect bottleneck identification accuracy (1.0), and a 96% reduction in deadlock events.
📝 Abstract
Finite-dimensional (FD) diffusion policies exhibit temporal drift owing to discretization artifacts that degrade long-horizon performance (when deployed on physical systems). We introduce a backward Kolmogorov equation that lifts diffusion policies to a Cameron-Martin space -- a subset of the Hilbert space. Essentially, replacing stochastic score matching with a deterministic boundary-value PDE problem. Our core innovation thrives on Gaussian measure theory whereupon the diffusion noise covariance operator is realized from a colored noise distribution which prescribes a notion of regularity on samples from the model at inference time. We train the diffusion model with a derived precision-weighted Cameron- Martin loss and a Kolmogorov residual is introduced as a PDE diagnostic during inference. These substitutions yield (i) convergence guarantees where the bound's constants depend on the effective rank of the kernel rather than action dimension, (ii) improved trajectory regularity via spectral weighting, and (iii) a deterministic failure detector without reward signals. Validation across two application domains demonstrates substantial improvements: on the PushT manipulation benchmark, the Cameron-Martin loss achieves a 17% improvement in maximum episode reward (0.95 vs. 0.78 for MSE) and 67.6% reduction in inter-step drifts during inference via the introduced residual magnitude. Similarly, on a 6-station manufacturing line with constant work-in-process (CONWIP) flow control, we achieve 28.4% lower RMSE than classical LSTM baselines; a high starvation-event recall (1.0 in test cycles), and effective bottleneck identification (Precision@1 = 1.0 in test set, 13x signal-to-noise ratio). We then certify the dispatch policies with Hamilton-Jacobi reachability theory which reduces deadlock events by 96% compared to uncontrolled dispatch over 100 simulated runs (351 events prevented).
Problem

Research questions and friction points this paper is trying to address.

temporal drift
diffusion policies
discretization artifacts
long-horizon performance
physical systems
Innovation

Methods, ideas, or system contributions that make the work stand out.

Kolmogorov equation
Cameron-Martin space
diffusion policies
colored noise
PDE-based control
🔎 Similar Papers
2024-07-16arXiv.orgCitations: 2
💼 Related Jobs
No related jobs found.
L
Lekan Molu
Bala Cynwyd, PA 19004