WeEnv: The Environment for Agentic Reinforcement Learning at WeChat

📅 2026-09-24
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the excessive iteration overhead in agent reinforcement learning caused by inadequate environment management. To this end, it proposes a full-lifecycle environment management system. Methodologically, the system employs containerized hierarchical packaging to enable on-demand composition, combined with a lazy loading mechanism that supports instant startup and content retrieval as needed. Furthermore, a dynamic resource scheduling algorithm is designed to facilitate elastic quota adjustments for CPU and memory. Experimental results demonstrate that the proposed system accelerates environment initialization by 5.6 to 14.2 times and reduces the proportion of environment-related time consumption to 9.1%. The system has been successfully deployed in WeChat's business operations.
📝 Abstract
Agentic reinforcement learning (RL) differs from conventional RL in that every task executes inside a complex environment, e.g., a virtual machine or a container. We find that agentic RL pays a heavy environment tax: a large share of the iteration time goes to the environment rather than to learning. The root cause is the lack of a full-lifecycle solution to environment management. We present WeEnv, which manages environments across packaging, initialization, and provisioning. WeEnv packages components as independently published layer groups and composes them at initialization, so that updating a component republishes one small group rather than every artifact containing it. To speed up environment initialization, WeEnv launches environments instantly and fetches contents on demand. During task execution, WeEnv provisions CPU and memory elastically, adjusting each environment's quota from its observed usage to fit the varying demands. WeEnv reduces the initialization by 5.6-14.2x over E2B, Docker, and AgentENV, cutting its share of the iteration time from up to 53.4% to 9.1%. WeEnv is deployed for agentic RL at WeChat.
Problem

Research questions and friction points this paper is trying to address.

Agentic Reinforcement Learning
Environment Management
Initialization Overhead
Lifecycle Management
Innovation

Methods, ideas, or system contributions that make the work stand out.

Agentic Reinforcement Learning
Environment Management
On-demand Fetching
Elastic Provisioning
Layer Groups
Y
Yang Yu
WeChat AI, Tencent
Jing Lei
Jing Lei
Carnegie Mellon University
Probability and Statistics
S
Shaoxun Zeng
WeChat AI, Tencent
Xinyu Gao
Xinyu Gao
NanJing University
Autonomous DrivingMulti-sensor FusionTesting
J
Jindi Shi
WeChat AI, Tencent
C
Ci Lei
WeChat AI, Tencent
J
Junjie Zhang
WeChat AI, Tencent