🤖 AI Summary
This study investigates the performance trade-offs of data-driven approaches in finite-horizon dynamic pricing, with a focus on complex settings involving high-dimensional multi-product offerings, heterogeneous demand structures, and intertemporal revenue constraints. By systematically comparing Fitted Dynamic Programming (Fitted DP) against several reinforcement learning algorithms—integrating demand estimation, trajectory sampling, and expectation-based optimization—the work comprehensively evaluates their relative strengths in terms of revenue generation, stability, constraint satisfaction, and computational scalability. The findings reveal that Fitted DP exhibits superior scalability in structured, complex environments, whereas reinforcement learning demonstrates greater flexibility and adaptability. These insights provide both theoretical grounding and empirical evidence to inform method selection for real-world dynamic pricing systems.
📝 Abstract
This paper provides a systematic comparison between Fitted Dynamic Programming (DP), where demand is estimated from data, and Reinforcement Learning (RL) methods in finite-horizon dynamic pricing problems. We analyze their performance across environments of increasing structural complexity, ranging from a single typology benchmark to multi-typology settings with heterogeneous demand and inter-temporal revenue constraints. Unlike simplified comparisons that restrict DP to low-dimensional settings, we apply dynamic programming in richer, multi-dimensional environments with multiple product types and constraints. We evaluate revenue performance, stability, constraint satisfaction behavior, and computational scaling, highlighting the trade-offs between explicit expectation-based optimization and trajectory-based learning.