Emergence, Not Bandwidth: Physical Coupling and the Limits of Learned Multi-Agent Communication

📅 2026-10-01
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses theoretical gaps in optimal encoding, compression costs, and protocol identifiability for bandwidth-constrained multi-agent communication by characterizing communication limits from an information-theoretic perspective within the Dec-POMDP framework. By integrating information-theoretic principles, MuJoCo simulations, and reinforcement learning (RL) algorithms, this work quantifies how physical coupling influences communication value and evaluates the limitations of RL in discovering efficient protocols. Results demonstrate that rigidly coupled agents derive no benefit from communication, whereas under partial coupling, RL performance degrades significantly compared to engineered senders, with learned protocols proving incompatible across models. The core contribution lies in establishing that performance bottlenecks stem from optimization failures rather than bandwidth constraints, thereby revealing the critical role of physical coupling dynamics in multi-agent systems.
📝 Abstract
Rate-limited multi-agent teams raise three questions the emergent-communication literature has answered only empirically: what an optimal message should encode, what compression costs over a horizon, and when a learned protocol is unique enough for a teammate to read. We answer them for rate-limited Dec-POMDPs, then measure how far reinforcement learning falls short of the optimum. Our theorems fix what is achievable independently of any learner, so a gap between an engineered and a learned sender at the same bit budget is an optimization fact, not an information-theoretic one. We instantiate this on three MuJoCo arenas spanning zero, partial and rigid physical coupling, charging every condition exactly 2 bits per decision, and create the discriminating regime by closing a physical side channel within one arena, holding bodies, task and reward fixed. Communication value is governed by coupling: under rigid coupling through a shared object, no channel beats silence (+0.001 +/- 0.001, p = 0.982, n = 25), since proprioception already carries that information; without coupling, every condition solves the task; under partial coupling, the engineered 2-bit sender reaches an interquartile mean of 1.000 but the learned one reaches 0.482, indistinguishable from silence (p = 0.400, n = 25). With a shared alphabet, bandwidth cannot explain the gap. Warm-starting from an engineered receiver localizes the failure: the same channel reaches 0.857 versus 0.562 cold-started (p < 0.001), so it is neither representational nor one of maintenance; reinforcement learning fails to discover the protocol. Cross-play shows learned protocols are individually meaningful but mutually unintelligible: self-play 0.980 collapses to 0.144 across seeds, and our best constructed alignment leaves at least 77% of that gap. All headline results use 25 seeds per arena and seven published baselines at matched rate.
Problem

Research questions and friction points this paper is trying to address.

multi-agent communication
rate-limited Dec-POMDPs
physical coupling
protocol alignment
reinforcement learning gap
Innovation

Methods, ideas, or system contributions that make the work stand out.

Emergent Communication
Dec-POMDPs
Physical Coupling
Rate-limited Communication
Cross-play