🤖 AI Summary
This study addresses the residual climate hedging valuation adjustment (HVA) that persists for trading desks despite existing hedging and overlay strategies—a gap not inferable from standalone stress losses. To tackle this, the authors propose an optimized modeling paradigm that pairs simulations of climate-activated and baseline worlds, reformulating hedge instrument discovery as an optimization objective aimed at minimizing residual HVA. The approach integrates an analytical linear-Gaussian Riccati solution with Climate-Dyna model-based reinforcement learning for nonlinear refinement, augmented by a gating mechanism to regulate policy updates. Empirical evaluation in a semi-synthetic EU ETS environment demonstrates that the method reduces average climate costs from 0.906 to 0.831—approaching the theoretical lower bound of 0.821—while achieving a 93% reduction in regret using only one-quarter of the typical trajectory count and capturing 60.7% of the theoretical gain within just 25 target transitions.
📝 Abstract
For a trading desk, residual climate hedging valuation adjustment (HVA) is the climate cost left after its inherited hedge and any admissible overlay have been taken into account; it therefore cannot be inferred from a stand-alone stress loss. We obtain this residual by comparing paired climate-on and baseline worlds and reoptimizing the overlay for each hedge universe, which also turns hedge-instrument discovery into a valuation problem: an instrument is useful to the extent that it lowers the optimized residual cost. The linear-Gaussian case has an exact finite-horizon Riccati solution; Climate-Dyna starts from that hedge and learns the remaining nonlinear correction from paired world-model rollouts, with an independent gate deciding whether to deploy the update. In a public-data-calibrated semi-synthetic EU ETS study, crediting the inherited hedge lowers the mean climate charge from 1.517 to 0.906, and the learned overlay lowers it to 0.831 against a 0.821 exact floor; residual Dyna cuts regret by 93% relative to replay with one quarter as many trajectories, while adaptation from only 25 target transitions retains 60.7% of the exact-assisted gain.