🤖 AI Summary
This study addresses the challenges of representation ambiguity and state uncertainty arising from intermittent partner invisibility in embodied zero-shot coordination. To this end, we propose PIP, a novel method that enhances local perception under partial observability through joint training. Specifically, PIP distills partner representations via a joint-view variational autoencoder and constructs a belief network to explicitly model latent states, enabling the inference of hidden positions and behavioral intentions. Integrated with multi-agent reinforcement learning, this framework achieves robust coordination. Experimental results demonstrate that PIP attains superior average performance across three benchmarks. Furthermore, human evaluations corroborate its effective coordination capabilities in scenarios involving partner occlusion.
📝 Abstract
Zero-shot coordination in embodied settings requires acting while the partner is intermittently out of view, leaving existing methods with ambiguous partner representations and uncertainty over hidden partner states. We propose Predicting Intention of Partner (PIP) to jointly address these challenges. PIP uses a Joint-view VAE to distill richer training-time evidence from the union of both agents' local observations into a partner representation available from local observations alone. Partner-state Belief networks further infer the partner's hidden location and behavioral tendencies from the ego agent's interaction history. We evaluate PIP in Burrito-PO, Overcooked-PO, and a Melting Pot substrate, together with a human evaluation in Burrito-PO. PIP attains the highest mean performance among the compared methods across all three benchmarks. Human evaluation and diagnostic analyses further support coordination with unseen partners and the contributions of both components under partner occlusion.