Proximal Residual Value Functions for Consistent Planning and Real-Time Execution

📅 2026-09-19
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
研究通过近端残差值函数的强化学习方法解决两时尺度决策系统中规划与实时执行一致性问题,降低电商库存成本。
📝 Abstract
We study two-timescale decision systems in which a planning layer periodically supplies a continuation-value function to a real-time optimizer that allocates arriving resources, with inventory placement as our motivating application. We propose an end-to-end reinforcement learning (RL) method for learning this function using \emph{proximal residual value functions}, which combine a strictly convex potential of post-decision inventory with a learned convex residual. This general form yields a well-posed optimization layer that supports end-to-end differentiation while preserving an explicit convex objective for real-time execution. We characterize the necessary and sufficient conditions under which a smooth value function yields decisions that are consistent across the planning and execution timescales. In an offline simulation using historical inventory arrival and demand patterns from a large e-commerce retailer, learned proximal residual value functions reduce total routing and transfer cost relative to a historical-production-system proxy by 5.0%.
Problem

Research questions and friction points this paper is trying to address.

two-timescale decision systems
continuation-value function
real-time optimizer
inventory placement
consistency
Innovation

Methods, ideas, or system contributions that make the work stand out.

proximal residual value functions
end-to-end reinforcement learning
two-timescale decision systems
convex potential of post-decision inventory
🔎 Similar Papers