🤖 AI Summary
This study addresses the inconsistency between Pontryagin’s Maximum Principle (PMP) conditions under Itô and rough path frameworks in stochastic optimal control. By leveraging rough path theory, the Itô–Stratonovich conversion, and duality identities, this work establishes a bridge via adjoint equations connecting both frameworks. It proposes unified optimality conditions based on conditional expectations, enabling the treatment of adapted control problems without forward-backward stochastic differential equations (FBSDEs). The contributions include the derivation of a unified PMP and a rough stochastic PMP for adapted controls, thereby achieving theoretical unification across frameworks. Furthermore, this paper re-derives fine-tuning methods for generative models and introduces an indirect shooting method for feedback problems, providing a complete mathematical foundation for cross-framework stochastic control.
📝 Abstract
Stochastic differential equations (SDEs) can be studied via Itô calculus and rough path theory. For stochastic optimal control, these two frameworks give distinct Pontryagin Maximum Principle (PMP) optimality conditions with forward-backward SDEs (FBSDEs) or rough differential equations. We show that the adjoint equations of the Itô and rough PMPs are connected via the conditional expectation $p_t^{\text{Itô}}=\mathbb{E}[p_t^{\text{rough}} \mid \mathcal{F}_t]$, where $\mathcal{F}_t$ represents information available at time $t$. First, we derive a rough stochastic PMP for problems with adapted controls that does not use FBSDEs. Its proof extends the rough stochastic PMP over deterministic controls by considering stochastic needle variations. Second, we derive a unified PMP connecting the Itô and rough PMPs, using Itô-Stratonovich conversion formulas and duality identities between the forward tangent and backward adjoint SDEs. As a first application, we rederive the adjoint matching method for fine-tuning generative models. As a second application, we propose an indirect shooting method for a class of feedback problems. Overall, these results give a new conditional bridge connecting two popular frameworks for stochastic optimal control.