Law-Strength Frontiers and a No-Free-Lunch Result for Law-Seeking Reinforcement Learning on Volatility Law Manifolds

📅 2025-11-21
📈 Citations: 0
Influential: 0
📄 PDF

career value

246K/year
🤖 AI Summary
This paper investigates whether embedding no-arbitrage constraints as soft penalties into a reinforcement learning (RL) world model for volatility surface modeling effectively aligns high-capacity agents’ behavior—or instead induces Goodhart-type misuse. Method: We construct a convex-geometric volatility manifold, model market dynamics via RNNs, and introduce differentiable no-arbitrage penalties alongside a Goodhart decomposition, optimizing policies with PPO. Contributions/Results: We propose the “law-strength frontier” and the “Graceful Failure Index” (GFI) to characterize the trade-off between penalty strength and reward; prove a No-Free-Lunch theorem for “law-seeking learning,” showing it cannot simultaneously outperform structured baselines across reward, penalty adherence, and robustness; and empirically validate—on SPX/VIX-like environments—that all law-seeking RL agents underperform simple structural strategies, clustering in high-penalty, high-GFI regimes, thereby confirming theoretical performance limits.

Technology Category

Application Category

📝 Abstract
We study reinforcement learning (RL) on volatility surfaces through the lens of Scientific AI. We ask whether axiomatic no-arbitrage laws, imposed as soft penalties on a learned world model, can reliably align high-capacity RL agents, or mainly create Goodhart-style incentives to exploit model errors. From classical static no-arbitrage conditions we build a finite-dimensional convex volatility law manifold of admissible total-variance surfaces, together with a metric law-penalty functional and a Graceful Failure Index (GFI) that normalizes law degradation under shocks. A synthetic generator produces law-consistent trajectories, while a recurrent neural world model trained without law regularization exhibits structured off-manifold errors. On this testbed we define a Goodhart decomposition (r = r^{mathcal{M}} + r^perp), where (r^perp) is ghost arbitrage from off-manifold prediction error. We prove a ghost-arbitrage incentive theorem for PPO-type agents, a law-strength trade-off theorem showing that stronger penalties eventually worsen P&L, and a no-free-lunch theorem: under a law-consistent world model and law-aligned strategy class, unconstrained law-seeking RL cannot Pareto-dominate structural baselines on P&L, penalties, and GFI. In experiments on an SPX/VIX-like world model, simple structural strategies form the empirical law-strength frontier, while all law-seeking RL variants underperform and move into high-penalty, high-GFI regions. Volatility thus provides a concrete case where reward shaping with verifiable penalties is insufficient for robust law alignment.
Problem

Research questions and friction points this paper is trying to address.

Investigates whether no-arbitrage laws reliably align RL agents or create exploitation incentives
Analyzes law-strength trade-offs showing stronger penalties eventually worsen financial performance
Demonstrates law-seeking RL cannot outperform structural baselines on key financial metrics
Innovation

Methods, ideas, or system contributions that make the work stand out.

Using volatility law manifold for no-arbitrage constraints
Developing Graceful Failure Index to measure law degradation
Proving no-free-lunch theorem for law-seeking reinforcement learning