Adapting to Changes in Agent Behavior via Finite-Depth Policy Sensitivity

📅 2026-10-05
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the reliance of reinforcement learning policies on extensive new interaction data when adapting to behavioral changes in other agents. To this end, it proposes a finite-depth sensitivity estimation framework that quantifies policy sensitivity to environmental shifts by approximating the policy Hessian and mixed derivatives. By introducing an adjustable propagation depth to truncate reference environment information, the method establishes non-increasing error bounds for efficient prediction. Evaluated in belief-based pursuit-evasion games, the approach demonstrates monotonically decreasing errors with increasing depth while outperforming baselines. Furthermore, utilizing the estimated sensitivity to initialize policies significantly enhances zero-shot returns and subsequent fine-tuning efficiency.
📝 Abstract
Adapting a reinforcement learning policy to changes in another agent's behavior typically requires a large amount of new interaction data. Policy sensitivity provides a first-order prediction of how a locally optimal policy changes with a behavioral parameter, but its computation requires second-order derivatives whose effects propagate across future interactions. We develop a finite-depth framework to estimate this sensitivity by approximating the policy Hessian and mixed derivative using information from a reference environment. The method features an adjustable propagation depth which determines where derivative propagation along the trajectory is truncated. We characterize the derivative contributions omitted by finite-depth propagation and derive truncation-error bounds for the approximated derivatives and resulting policy sensitivity. The bounds are nonincreasing with propagation depth and vanish at full-horizon propagation. Using a belief-driven pursuit-evasion game as a validation scenario, the proposed method generally achieves lower derivative-estimation errors as the propagation depth increases and outperforms the baseline methods in both estimation accuracy and policy adaptation. The sensitivity-based initialization improves zero-shot return over direct transfer, and also shows advantages for the subsequent fine-tuning in the target environment.
Problem

Research questions and friction points this paper is trying to address.

reinforcement learning
policy adaptation
policy sensitivity
multi-agent behavior change
derivative estimation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Policy Sensitivity
Finite-Depth Framework
Reinforcement Learning
Derivative Approximation
Zero-Shot Adaptation
🔎 Similar Papers
No similar papers found.