🤖 AI Summary
This study addresses the limited decision-making capabilities of general-purpose agents and the inefficiency of action space expansion by proposing Dyad, an architecture that augments large language models (LLMs) with an environment-conditioned action encoder to enable typed decision-making. By decoupling state representations from action semantics, Dyad optimizes interaction efficiency. Its core innovation lies in introducing an inductive bias for reusable representations, which supports cross-environment adaptation by fine-tuning only the encoder while keeping the LLM frozen, thereby significantly reducing training costs. Experimental results demonstrate consistent performance gains across four unseen environments, achieving a 3.80% improvement on ALFWorld tasks, while concurrently enhancing the model's knowledge, reasoning, and encoding capabilities.
📝 Abstract
We study how to build more capable general-purpose agents by extending large language models (LLMs) with native typed decision-making. We introduce Dyad, an architecture that augments a pretrained LLM with an environment-conditioned action encoder that embeds each candidate action description in parallel, then scores these embeddings against the LLM's internal state to yield a distribution over typed actions. By factorizing decision-making into representations of the evolving interaction state and environment-specific action semantics, Dyad introduces an inductive bias for learning reusable representations while keeping action scoring efficient even as the action space grows. We investigate two complementary reinforcement learning settings driven by environment interaction. With the LLM frozen, training the action encoder alone achieves consistent gains across four unseen environments, enabling modular adaptation without modifying any LLM parameters. Jointly optimizing both components outperforms conventional RL post-training across diverse interactive tasks and model scales, including a 3.80% average absolute gain on ALFWorld with a 9B model, while improving general knowledge, reasoning, and coding.