🤖 AI Summary
This study addresses the limitation of offline reinforcement learning for autonomous driving, demonstrating that constraining solely the ego-vehicle action space fails to resolve interaction distribution shift (IDS). To this end, this work formally defines IDS for the first time and proposes the ICDP framework, which decomposes joint support into ego and residual interaction components. By leveraging contrastive density ratio estimation, ICDP explicitly controls interaction-level distribution shifts, enabling policy optimization without requiring joint modeling or simulation rollbacks. Evaluations on the nuPlan benchmark and real-world vehicle tests confirm that the proposed approach significantly suppresses high-value yet interaction-unsupported trajectories, thereby effectively enhancing performance in critical driving scenarios.
📝 Abstract
Offline reinforcement learning enables reward-driven policy improvement from fixed datasets without requiring online exploration, making it particularly attractive in safety-critical domains. A central challenge, however, is distribution shift: policy optimization may favor actions that are weakly supported by the offline data, rendering value estimates unreliable. Existing approaches primarily control this shift in the policy's own action space. In interactive environments such as autonomous driving, this can be insufficient: a candidate ego trajectory may remain well supported under the marginal behavior distribution while being poorly supported jointly with the surrounding-agent behavior observed in the logged interaction. We refer to this degradation in interaction support as \emph{interaction distribution shift} (IDS), and introduce \emph{Interaction-Constrained Drive Policy} (ICDP), an offline reinforcement learning framework that explicitly controls interaction-level distribution shift. Starting from the joint data distribution over ego and surrounding-agent futures, we show that joint-support degradation decomposes exactly into an ego-support component and a residual interaction-support component. We recover the latter through contrastive density-ratio estimation, isolating interaction compatibility without explicit joint-density modeling, surrounding-agent prediction, or rollouts in reactive simulators or learned world models during policy optimization. Closed-loop evaluations on nuPlan, Interplan and real-world truck experiments show that ICDP suppresses high-value yet interaction-unsupported trajectory selections and improves performance in interaction-critical driving scenarios. Project webpage: https://mahmoud-selim.github.io/ICDP/