Design and Optimization of Reinforcement Learning-Based Agents in Text-Based Games

📅 2025-09-03
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the low task completion and win rates of reinforcement learning (RL) agents in text-based games. We propose a novel end-to-end architecture that jointly integrates deep language understanding with policy gradient optimization. Methodologically, we employ pretrained language models (e.g., BERT or T5) to parse game text and implicitly construct a differentiable world model, which is co-optimized with an enhanced policy gradient algorithm—specifically Proximal Policy Optimization (PPO)—to enable efficient mapping from textual observations to action policies. Our key contributions are: (i) the first incorporation of a differentiable world modeling mechanism directly into the policy network, improving long-horizon reasoning and state consistency; and (ii) the integration of multi-task pretraining and curriculum learning to enhance generalization. Evaluated on standard benchmarks including Zork and TextWorld, our approach significantly outperforms existing RL baselines, achieving average improvements of 23.6% in task completion rate and 31.2% in win rate, thereby demonstrating both effectiveness and cross-game transferability.

Technology Category

Machine Learning: Reinforcement LearningMultiagent Systems: Adversarial AgentsNatural Language Processing: Learning & Optimization for NLP

Application Category

Search and Retrieval-Augmented AI: Web learning to rank, online learning, and counterfactual learning for rankingEconomics, Online Markets and Human Computation: Uses of LLMs and GenAI for marketplace design, bidding, and strategic interactionsUser Modeling, Personalization and Recommendation: Large Language Models (LLM) for user modeling and recommendation
📝 Abstract
As AI technology advances, research in playing text-based games with agents has becomeprogressively popular. In this paper, a novel approach to agent design and agent learning ispresented with the context of reinforcement learning. A model of deep learning is first applied toprocess game text and build a world model. Next, the agent is learned through a policy gradient-based deep reinforcement learning method to facilitate conversion from state value to optimal policy.The enhanced agent works better in several text-based game experiments and significantlysurpasses previous agents on game completion ratio and win rate. Our study introduces novelunderstanding and empirical ground for using reinforcement learning for text games and sets thestage for developing and optimizing reinforcement learning agents for more general domains andproblems.
Problem

Research questions and friction points this paper is trying to address.

Designing reinforcement learning agents for text-based games
Optimizing policy gradient methods for game completion
Enhancing agent performance through deep world modeling
Innovation

Methods, ideas, or system contributions that make the work stand out.

Deep learning model processes game text
Policy gradient-based deep reinforcement learning
Converts state value to optimal policy
🔎 Similar Papers
No similar papers found.
H
Haonan Wang
Johns Hopkins University, Baltimore, USA
M
Mingjia Zhao
College of Science, Liaoning Technical University, China
Junfeng Sun
Junfeng Sun
College of Science, Liaoning Technical University, China
W
Wei Liu
College of Science, Liaoning Technical University, China