🤖 AI Summary
This study addresses the coordination rigidity in LLM-based multi-agent systems caused by reliance on explicit roles and global topologies. We propose a decentralized, self-organizing collaboration framework grounded in shared anonymous local rules. By discarding global architectures, this method enables adaptive interactions within dynamic populations through the iterative execution of local rules. Furthermore, we introduce Swarm-Consistent Distillation, which integrates trajectory consistency with predictive distillation to optimize policy transfer. Experimental results demonstrate that our approach preserves over 96% of baseline performance while supporting zero-retraining transfer. Additionally, it significantly enhances the system's reorganization capability following adversarial perturbations, offering a robust and scalable paradigm for emergent multi-agent coordination without centralized control or predefined structural constraints.
📝 Abstract
As LLM agents increasingly collaborate on complex tasks, how to organize their interactions becomes a central design question. Existing multi-agent systems typically learn or adapt explicit roles, hierarchies, routing policies, or communication topologies. We shift the learning target to a reusable local law that can be shared across interchangeable agents and adapt coordination as populations or interaction conditions change, without redefining a global organization. We introduce Waggle, a shared anonymous policy over bounded local views that jointly selects task actions, semantic communication, and local commitment updates. Repeated execution of the same law allows coordination to form, persist, and reorganize online without explicit roles or global topology. To learn this law across interchangeable agents and evolving coordination, we develop Swarm-Consistent Distillation (SCD), combining anonymous-orbit consistency with rollout-grounded prediction of the next local coordination field, with no added inference-time components. Across diverse coordination settings, the same learned law remains effective as populations and interaction budgets change, retains over 96% of substrate-specific oracle quality, and transfers without retraining; SCD further improves reorganization after counterevidence. Together, these results show that LLM-agent organization can emerge and adapt through repeated execution of a learned local law.