Scalable Population Training for Zero-Shot Coordination

📅 2025-11-14
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
Existing population-based training methods for zero-shot coordination (ZSC) suffer from prohibitive computational costs and poor scalability with population size. To address this, we propose ScaPT, a scalable population training framework that models large cooperative populations via parameter-sharing meta-agents and incorporates mutual information regularization to preserve behavioral diversity while drastically reducing computational overhead. Experiments on the Hanabi benchmark demonstrate that ScaPT significantly outperforms prior approaches in zero-shot coordination performance. Notably, it provides the first empirical validation that increasing population size yields substantial gains in ZSC generalization capability. Moreover, ScaPT exhibits strong scalability—maintaining efficiency even as population size grows. This work establishes a novel paradigm for efficient and scalable multi-agent cooperative learning, advancing the state of the art in population-based ZSC.

Technology Category

Application Category

📝 Abstract
Zero-shot coordination(ZSC) has become a hot topic in reinforcement learning research recently. It focuses on the generalization ability of agents, requiring them to coordinate well with collaborators that are not seen before without any fine-tuning. Population-based training has been proven to provide good zero-shot coordination performance; nevertheless, existing methods are limited by computational resources, mainly focusing on optimizing diversity in small populations while neglecting the potential performance gains from scaling population size. To address this issue, this paper proposes the Scalable Population Training (ScaPT), an efficient training framework comprising two key components: a meta-agent that efficiently realizes a population by selectively sharing parameters across agents, and a mutual information regularizer that guarantees population diversity. To empirically validate the effectiveness of ScaPT, this paper evaluates it along with representational frameworks in Hanabi and confirms its superiority.
Problem

Research questions and friction points this paper is trying to address.

Addresses computational limitations in population-based training for zero-shot coordination
Optimizes agent generalization to unseen collaborators without fine-tuning requirements
Enables scalable population training while maintaining diversity through parameter sharing
Innovation

Methods, ideas, or system contributions that make the work stand out.

Parameter sharing meta-agent enables scalable population
Mutual information regularization maintains population diversity
Efficient framework overcomes computational limits in training
🔎 Similar Papers
No similar papers found.
B
Bingyu Hui
Department of Electronic Engineering, Beijing National Research Center for Information Science and Technology, Tsinghua University, Beijing 100084, China
L
Lebin Yu
Department of Electronic Engineering, Beijing National Research Center for Information Science and Technology, Tsinghua University, Beijing 100084, China
Quanming Yao
Quanming Yao
Associate Professor, EE Department, Tsinghua University
Machine Learning
Yunpeng Qu
Yunpeng Qu
Tsinghua University
X
Xudong Zhang
Department of Electronic Engineering, Beijing National Research Center for Information Science and Technology, Tsinghua University, Beijing 100084, China
J
Jian Wang
Department of Electronic Engineering, Beijing National Research Center for Information Science and Technology, Tsinghua University, Beijing 100084, China