A Comparative Study of Dynamic Programming and Reinforcement Learning in Finite Horizon Dynamic Pricing

📅 2026-04-15
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study investigates the performance trade-offs of data-driven approaches in finite-horizon dynamic pricing, with a focus on complex settings involving high-dimensional multi-product offerings, heterogeneous demand structures, and intertemporal revenue constraints. By systematically comparing Fitted Dynamic Programming (Fitted DP) against several reinforcement learning algorithms—integrating demand estimation, trajectory sampling, and expectation-based optimization—the work comprehensively evaluates their relative strengths in terms of revenue generation, stability, constraint satisfaction, and computational scalability. The findings reveal that Fitted DP exhibits superior scalability in structured, complex environments, whereas reinforcement learning demonstrates greater flexibility and adaptability. These insights provide both theoretical grounding and empirical evidence to inform method selection for real-world dynamic pricing systems.

Technology Category

Search and Optimization: Learning to SearchReasoning under Uncertainty: Stochastic OptimizationMachine Learning: Online Learning & Bandits

Application Category

Economics, Online Markets and Human Computation: Advertising auctions, pricing, markets, and exchangesGraph Algorithms and Modeling for the Web: Efficient manipulation of static and dynamic Web-related graphsSearch and Retrieval-Augmented AI: Web learning to rank, online learning, and counterfactual learning for ranking
📝 Abstract
This paper provides a systematic comparison between Fitted Dynamic Programming (DP), where demand is estimated from data, and Reinforcement Learning (RL) methods in finite-horizon dynamic pricing problems. We analyze their performance across environments of increasing structural complexity, ranging from a single typology benchmark to multi-typology settings with heterogeneous demand and inter-temporal revenue constraints. Unlike simplified comparisons that restrict DP to low-dimensional settings, we apply dynamic programming in richer, multi-dimensional environments with multiple product types and constraints. We evaluate revenue performance, stability, constraint satisfaction behavior, and computational scaling, highlighting the trade-offs between explicit expectation-based optimization and trajectory-based learning.
Problem

Research questions and friction points this paper is trying to address.

Dynamic Pricing
Finite Horizon
Dynamic Programming
Reinforcement Learning
Demand Estimation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Fitted Dynamic Programming
Reinforcement Learning
Finite Horizon Dynamic Pricing
Heterogeneous Demand
Inter-temporal Constraints
💼 Related Jobs
No related jobs found.
L
Lev Razumovskiy
N
Nikolay Karenin