🤖 AI Summary
This study addresses the bottleneck in virtual environments where balancing state consistency with visual realism hinders agent learning, proposing a real-time interactive framework that couples simulators with neural renderers. The core innovation is an adversarial enforcement training method that renders history prefilling differentiable, enabling real-time mapping from structured conditions to photorealistic visuals and supporting flexible expansion of code-defined worlds. By integrating game engines, video model adaptation, and geometric condition control, this framework achieves efficient learning within merely four interaction rounds—a substantial improvement over traditional reinforcement learning methods requiring millions of samples—while allowing environment scale to expand synchronously with agent capabilities.
📝 Abstract
Interactive virtual worlds allow agents to learn through exploration and interaction. What agents can learn is bounded by the environments they practice in, which must be faithful, with consistent state, rules, and dynamics, and realistic, with observations that follow the real-world visual distributions. Achieving both across diverse worlds remains a bottleneck. We introduce AgentGarten, a framework that couples simulators and game engines with a shared neural renderer to build real-time interactive environments. Its simulation backends maintain persistent world state and execute program-defined interaction rules, while the renderer generates visual observations from structured conditions exported through a common interface. To build the neural renderer, we adapt a pretrained video model to geometry conditions, distill it with our proposed Adversarial Forcing, and optimize inference for real-time interaction. Adversarial Forcing makes history prefilling differentiable through exact replay, so that losses on later predictions update how the renderer encodes prior observations, and adds real-data adversarial supervision to improve its visual quality. In AgentGarten, agents perceive the world through visual observations, interact with it in real time, and improve by distilling each round of experience into playbooks that subsequent agents inherit and refine. Our empirical study demonstrates a substantial gain in learning efficiency, with agents learning from just 4 rounds compared with millions for a conventional reinforcement learning counterpart. As new worlds can be written as code and rendered through the same interface, environments can scale in both number and difficulty alongside their agents, a step toward agents that keep evolving through interactive experience.