Robust Risk-Sensitive Reinforcement Learning from Corrupted Human Feedback

📅 2026-09-30
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the robustness of reinforcement learning from human feedback (RLHF) when subjected to label-flipping attacks. To this end, we propose the WSP-CVaR-RLHF algorithm, which is grounded in a Conditional Value-at-Risk (CVaR) framework and integrates uncertainty-weighted reward estimation with optimistic planning. By decoupling statistical costs from contamination penalties and leveraging rectangular confidence sets, the method effectively disentangles coupled uncertainties under unknown transitions. Experimental results demonstrate that the proposed algorithm significantly reduces cumulative regret while maintaining confidence set coverage across various adversarial attacks, consistently outperforming unweighted baseline approaches.
📝 Abstract
Reinforcement learning with human feedback (RLHF) learns from human comparisons, which can be corrupted or deliberately manipulated. This paper studies online risk-sensitive RLHF with static conditional value-at-risk (CVaR) under adversarial preference-label flips. We consider additive linear rewards and a fixed-reference protocol with one comparison per episode and at most $C$ flipped labels over $K$ episodes. We propose weighted streamed-preference CVaR RLHF (WSP-CVaR-RLHF), which combines uncertainty-weighted reward estimation with optimistic augmented-state CVaR planning. For known transitions and normalized rewards, we establish the regret bound $\widetilde{O}\left(\frac{d}κ\sqrt{\frac{K}α}+\frac{dC}{κα}\right)$ up to lower-order terms, where $d$ is the reward-feature dimension, $α$ is the CVaR level, and $κ$ characterizes the preference link. The bound separates the clean statistical cost from the penalty caused by corrupted feedback. We further extend the analysis to unknown tabular transitions, where the trajectory distribution entering the CVaR objective must be learned together with the reward. We address the resulting coupled uncertainty using rectangular transition confidence sets, joint optimistic planning, and a history-level CVaR simulation argument. Experiments under four adversarial attacks demonstrate that WSP-CVaR-RLHF consistently reduces cumulative regret relative to its unweighted robust counterpart while preserving confidence-set coverage.
Problem

Research questions and friction points this paper is trying to address.

Reinforcement Learning from Human Feedback
Risk-Sensitive Reinforcement Learning
Adversarial Preference-Label Flips
Conditional Value-at-Risk
Corrupted Human Feedback
Innovation

Methods, ideas, or system contributions that make the work stand out.

Risk-Sensitive RLHF
Conditional Value-at-Risk (CVaR)
Adversarial Robustness
Uncertainty-Weighted Estimation
Joint Optimistic Planning