🤖 AI Summary
This study addresses the challenge that large language models (LLMs) struggle to effectively translate decoded conventions into action decisions during multi-agent collaboration, thereby limiting cooperative performance. Leveraging the Hanabi environment and the Qwen3 model, this work systematically compares the effects of rule statements versus action recommendations on decision-making through linear probing, activation steering, and controlled simulation experiments, effectively disentangling convention decodability, information sensitivity, and cooperative outcomes. The findings reveal a critical bottleneck in LLMs’ translation from knowledge to action. Specifically, action recommendations significantly enhance cooperative efficiency compared to rule statements. Furthermore, while activation steering improves intent inference accuracy, it fails to fully bridge this knowledge-to-action gap, highlighting inherent limitations in aligning LLM-based reasoning with effective multi-agent coordination.
📝 Abstract
Cooperation with unfamiliar partners requires adapting to communication conventions that are not known in advance. We study this problem in a controlled Hanabi-derived environment with scripted hint generation, LLM-controlled receiving decisions, and frozen model weights. Across eight LLMs, linear probes recover intent conventions substantially more accurately than target conventions, yet receiving choices do not consistently agree with the sender's convention. We compare probe-predicted and ground-truth conventions presented either as general rules or as externally computed action recommendations. Rule statements yield modest and model-dependent changes in cooperation, whereas action translation produces larger gains on average. In a Qwen3-8B case study, matched-state statement reversals reveal much greater sensitivity to action recommendations than to rule statements. Activation transfers from oracle-action and non-oracle hint-restatement donors improve intent accuracy on both action classes, but the tested alternatives do not reliably reproduce these benefits. Together, these results distinguish convention decodability, sensitivity to convention information, and cooperative performance, and highlight limitations in turning available partner information into receiving decisions.