When a Correct Reward Is Not Enough: Diagnosing and Guiding PPO in an Analytically Solved Broker-Trader Game

📅 2026-10-02
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study investigates whether reinforcement learning can effectively leverage analytical solutions to guide optimal financial control and examines why Proximal Policy Optimization (PPO) fails under complex dynamics. By evaluating PPO variants—including feedforward networks, LSTMs, and potential-based reward shaping—within a broker trading environment admitting closed-form solutions, the authors employ Monte Carlo diagnostics to identify deficiencies in reward design and evaluation metrics, proposing a diagnostic and fine-tuning framework anchored by the analytical policy as a benchmark. The findings reveal that correct reward signals alone are insufficient for RL success and confirm PPO’s performance limitations under stochastic flows. However, initializing with the analytical solution followed by frozen fine-tuning consistently narrows the performance gap to approximately 2.22% relative to an updated reference strategy, demonstrating the efficacy of analytical solutions both as diagnostic benchmarks and for policy initialization.
📝 Abstract
Reinforcement learning (RL) is increasingly used for financial optimal-control problems when complex dynamics make analytical strategies difficult to obtain. There are financial mathematics literactures which provides many solved models whose equations and controls could evaluate and guide learning; we ask whether RL can exploit these results. We place a proximal policy optimisation (PPO) agent in an analytically solved continuous-time broker--trader game. PPO replaces the broker and chooses its trading speed while interacting with an informed trader and stochastic uninformed order flow. We derive a finite-step reward from the broker's continuous-time payoff and verify its discrete implementation through grid refinement and an exact one-step identity. With zero uninformed flow, a validation-selected PPO--FFNN approaches the reference action. With stochastic uninformed flow, the tested PPO--FFNN and PPO--LSTM remain inaccurate, although supervised learning confirms that their actors can represent the action. Monte Carlo diagnostics show that their critics do not reliably rank nearby actions; potential-based reward shaping also gives no reliable improvement. Under partial information, a causal certainty-equivalent controller based on the broker's observable history remains close to the reference, while PPO has larger errors and lower payoffs. Finally, we freeze the analytical policy and train PPO to adjust it after the execution cost changes. Halving the cost yields a repeatable improvement that closes \(2.22\%\) of the gap to the changed-cost reference. The analytical solution therefore provides both a benchmark for diagnosing RL and a useful starting policy for adaptation.
Problem

Research questions and friction points this paper is trying to address.

Reinforcement Learning
Proximal Policy Optimization
Broker-Trader Game
Financial Optimal Control
Reward Shaping
Innovation

Methods, ideas, or system contributions that make the work stand out.

Proximal Policy Optimization
Broker-Trader Game
Reward Shaping
Analytical Benchmark
Policy Adaptation
💼 Related Jobs
No related jobs found.
S
Siu Tung Wong
Institute of Finance and Technology, University College London, London, United Kingdom
Carlo Campajola
Carlo Campajola
University College London
Quantitative FinanceNetwork ScienceMachine LearningDeFiMarket Microstructure