🤖 AI Summary
Collaborative language agents struggle with long-horizon coordination and adaptation to unfamiliar partners due to the difficulty of disentangling persistent strategies from tactical execution. This work proposes a training-free hierarchical reasoning architecture that couples strategic role allocation and action inference through a metacognitive prefrontal module, thereby decoupling strategy from tactics. Furthermore, it introduces a private partner-conditioned world model with forward simulation to enhance collaborative adaptability. In Overcooked-V2 experiments, the proposed architecture completes seven soups compared to only three for baselines. It effectively retains and adopts role-allocation proposals from unfamiliar partners, significantly improving cross-episode zero-shot collaboration performance.
📝 Abstract
Cooperative language-model agents must coordinate over long horizons and adapt to changing environments and to partners with unfamiliar conventions, yet existing agents map observations to actions without separating persistent coordination strategies from their tactical execution. We introduce OverForge, a training-free hierarchical architecture that separates strategic reasoning over roles and divisions of labour from tactical reasoning over actions within each agent's private, partner-conditioned world model. A metacognitive Prefrontal Cortex Module couples the two levels by forming strategy-action branches, imagining their consequences with a forward model, and committing when confident. In OvercookedV2, OverForge delivers 7 soups in a connected kitchen versus 3 for each flat LLM baseline, retains agreed roles, and adopts roles proposed by unfamiliar partners. Ablations and a fixed-strategy probe show that persistent strategies guide tactical adaptation while each reasoning level contributes to coordination. Memory restarts show that cross-episode partner knowledge supports task performance and partner prediction, linking the hierarchy to continual adaptation.