Decentralized Master-Mind: Joint Action Refinement through Iterative Intent Denoising in Multi-Agent Pathfinding
This study addresses the joint action conflicts arising from independent sampling in distributed multi-agent path planning by proposing a diffusion-inspired discrete iterative intent denoising mechanism. Methodologically, it replaces one-shot sampling with multi-round communication to couple agent decisions, and introduces MICPO, a critic-free group relative reinforcement learning algorithm that optimizes multi-step action correction while leveraging imitation learning pretraining for efficient policy acquisition. Experimental results demonstrate that the proposed approach successfully solves 1,598 out of 1,600 tasks, achieving both high coverage and low cost. Furthermore, the method scales effectively to million-level problem sizes while preserving the advantages of decentralized execution.