🤖 AI Summary
This work addresses the low task completion and win rates of reinforcement learning (RL) agents in text-based games. We propose a novel end-to-end architecture that jointly integrates deep language understanding with policy gradient optimization. Methodologically, we employ pretrained language models (e.g., BERT or T5) to parse game text and implicitly construct a differentiable world model, which is co-optimized with an enhanced policy gradient algorithm—specifically Proximal Policy Optimization (PPO)—to enable efficient mapping from textual observations to action policies. Our key contributions are: (i) the first incorporation of a differentiable world modeling mechanism directly into the policy network, improving long-horizon reasoning and state consistency; and (ii) the integration of multi-task pretraining and curriculum learning to enhance generalization. Evaluated on standard benchmarks including Zork and TextWorld, our approach significantly outperforms existing RL baselines, achieving average improvements of 23.6% in task completion rate and 31.2% in win rate, thereby demonstrating both effectiveness and cross-game transferability.
📝 Abstract
As AI technology advances, research in playing text-based games with agents has becomeprogressively popular. In this paper, a novel approach to agent design and agent learning ispresented with the context of reinforcement learning. A model of deep learning is first applied toprocess game text and build a world model. Next, the agent is learned through a policy gradient-based deep reinforcement learning method to facilitate conversion from state value to optimal policy.The enhanced agent works better in several text-based game experiments and significantlysurpasses previous agents on game completion ratio and win rate. Our study introduces novelunderstanding and empirical ground for using reinforcement learning for text games and sets thestage for developing and optimizing reinforcement learning agents for more general domains andproblems.