Tracking Drift: Variation-Aware Entropy Scheduling for Non-Stationary Reinforcement Learning

📅 2026-01-27
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the challenge of entropy coefficient selection in real-world reinforcement learning, where environmental non-stationarity often leads to insufficient or excessive exploration under a fixed entropy weight, compromising both efficiency during stable periods and adaptability following abrupt changes. To this end, the paper introduces Adaptive Entropy Scheduling (AES), the first method to frame entropy scheduling under non-stationary environments as a one-dimensional online trade-off problem. AES dynamically adjusts the temperature coefficient using a lightweight online proxy metric for distributional drift, without altering the underlying algorithm architecture. Evaluated across four algorithms, twelve tasks, and four distinct drift patterns, AES consistently reduces performance degradation and accelerates policy recovery after environmental shifts.

Technology Category

Multiagent Systems: Adversarial AgentsMachine Learning: Online Learning & BanditsPlanning, Routing, and Scheduling: Learning for Planning and Scheduling

Application Category

Economics, Online Markets and Human Computation: Incentives in network design for Web infrastructures and ecosystemsSearch and Retrieval-Augmented AI: Web learning to rank, online learning, and counterfactual learning for rankingResponsible Web: Human-perceived consequences of algorithmic deployment on the web
📝 Abstract
Real-world reinforcement learning often faces environment drift, but most existing methods rely on static entropy coefficients/target entropy, causing over-exploration during stable periods and under-exploration after drift (thus slow recovery), and leaving unanswered the principled question of how exploration intensity should scale with drift magnitude. We prove that entropy scheduling under non-stationarity can be reduced to a one-dimensional, round-by-round trade-off, faster tracking of the optimal solution after drift vs. avoiding gratuitous randomness when the environment is stable, so exploration strength can be driven by measurable online drift signals. Building on this, we propose AES (Adaptive Entropy Scheduling), which adaptively adjusts the entropy coefficient/temperature online using observable drift proxies during training, requiring almost no structural changes and incurring minimal overhead. Across 4 algorithm variants, 12 tasks, and 4 drift modes, AES significantly reduces the fraction of performance degradation caused by drift and accelerates recovery after abrupt changes.
Problem

Research questions and friction points this paper is trying to address.

non-stationary reinforcement learning
environment drift
entropy scheduling
exploration-exploitation trade-off
Innovation

Methods, ideas, or system contributions that make the work stand out.

Adaptive Entropy Scheduling
non-stationary reinforcement learning
environment drift
entropy coefficient
online drift detection
🔎 Similar Papers
No similar papers found.
T
Tongxi Wang
School of Future Technology, Southeast University, Nanjing, China
Z
Zhuoyang Xia
School of Future Technology, Southeast University, Nanjing, China
X
Xinran Chen
School of Future Technology, Southeast University, Nanjing, China
S
Shan Liu
School of Automation, Southeast University, Nanjing, China