Robust Adversarial Reinforcement Learning with Risk Sensitivity and Critic Consistency Regularization

📅 2026-09-23
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
本文针对强化学习在动态不确定性和分布偏移下的脆弱性问题,提出了一种结合风险敏感性和评价者一致性正则化的鲁棒对抗强化学习框架RACER。
📝 Abstract
Reinforcement learning (RL) achieves strong performance in sequential decision-making but remains brittle under dynamic uncertainty and distributional shifts. Robust Adversarial Reinforcement Learning (RARL) improves robustness via worst-case perturbations, but existing approaches frequently suffer from unstable optimization and degraded value estimation. In particular, overly aggressive adversaries can drive the agent toward uninformative failure states, while adversarial perturbations amplify disagreement between double critics and introduce biased value targets. We propose a unified framework, RACER (Risk-sensitive robust Adversarial critic ConsistEncy-regularized Reinforcement learning), that revisits adversarial RL from a risk-sensitive perspective. First, we introduce a state-dependent adversarial objective that adaptively regulates perturbation strength, suppressing harmful disturbances while preserving informative exploration. Second, we propose critic consistency regularization to reduce disagreement between Q-value estimators and stabilize learning. Comprehensive experiments on challenging continuous control benchmarks demonstrate that RACER consistently improves performance, robustness, and training stability over strong robust RL baselines.
Problem

Research questions and friction points this paper is trying to address.

Reinforcement Learning
Dynamic Uncertainty
Distributional Shifts
Robustness
Adversarial Perturbations
Innovation

Methods, ideas, or system contributions that make the work stand out.

Risk-sensitive
Critic Consistency Regularization
Adversarial Perturbation
🔎 Similar Papers
No similar papers found.