🤖 AI Summary
This study addresses the mode collapse problem in diffusion policies for robotic planning, where independent noise-action pairing induces aliasing that collapses multimodal behaviors into a single mode. To overcome this, we propose a plug-and-play module trained without labels that preserves distinct behavioral trajectories by optimizing noise-action assignment, maintaining multimodality without modifying the original architecture. This work is the first to reveal the mechanism by which independent pairing causes modal aliasing and to introduce a label-free, noise-assignment-based approach for modality preservation. Experiments demonstrate that our method increases the proportion of non-dominant modes by 6–14.6× and fully recovers missing modes, significantly enhancing multimodal retention while preserving strong task performance.
📝 Abstract
When diffusion policies were first introduced, they were expected to recover multi-modal action distributions. However, we find this expectation does not always hold, as diffusion policies often collapse to a single modality even when we guarantee the balance of dataset modalities and exact within-batch symmetry. Our analysis indicates that independent action-noise pairing contributes to this failure by increasing mixing and crossing among diffusion paths, which can produce averaged denoising responses and suppress modality-specific behavior. This issue is especially severe in robot planning, where action spaces are dense and low-dimensional, significantly increasing such mixing and crossing. To alleviate this problem, we propose Immiscible Diffusion Policy, a label-free training-time add-on to diffusion policy that uses action-noise assignment to preserve relatively distinct noise-to-action routes without modifying the policy architecture or inference procedure. Across five simulated and two real-world humanoid manipulation tasks spanning state, RGB, and point-cloud observations, our method significantly improves the policy's preservation of action modalities while maintaining strong task performance. It increases the proportion of the non-dominant modality by 6.0x-14.6x across three two-modality tasks and recovers demonstrated modalities that are entirely absent from vanilla policy rollouts on both four-modality tasks. These results demonstrate that Immiscible Diffusion Policy provides a simple yet robust approach to preserving action multi-modality in general robot learning tasks.