🤖 AI Summary
This study addresses the challenge of agents identifying and adapting to partners' capability limitations and multi-convention discrepancies in ad hoc teamwork. To this end, we propose a reinforcement learning-based adaptive collaboration framework. Methodologically, we introduce a capability-constrained training mechanism wherein adversarial training against a population of virtual partners with varying adaptive capacities compels the agent to actively probe partners' capability boundaries. This enables dynamic adaptation to multiple or single fixed conventions, guiding both parties toward optimal joint policies. Experimental results demonstrate that the proposed approach significantly outperforms existing ad hoc teamwork baselines across test partner populations with diverse convention compatibilities, effectively enhancing system robustness and generalizable collaborative capabilities.
📝 Abstract
Ad-hoc collaboration often requires agents to identify and adhere to some shared convention within a cooperative task. Existing work on reinforcement learning (RL) for ad-hoc collaboration focuses on training agents that adapt to the conventions established by their partners. These methods fail to consider the possibility that while some partners might follow only a single fixed convention, others may themselves be capable of adapting to multiple conventions. Here we present ConventionPlay, an RL-based approach that teaches agents to discover their partner's optimal convention by training against a learned population of partners that exhibit different degrees of adaptability across conventions. Some of these partners follow a single, fixed convention, while others are able to adapt to a subset of the possible conventions for the task in question. The existence of partners that support a limited subset of conventions forces agents trained against this population to actively probe their partner's capabilities, and steer their partner towards the most effective joint strategy that they are capable of following. Our experimental results demonstrate that agents trained via ConventionPlay achieve superior performance to existing ad-hoc collaboration methods against test populations of partners that are compatible with multiple conventions.