uncertainty-aware reinforcement learning

Design and implement reinforcement learning algorithms and training procedures that compute predictive uncertainty (e.g., epistemic and aleatoric) and incorporate those estimates into objectives, advantage weighting, importance weights, or policy-gradient updates to trade off expected return against reliability or risk. Build mechanisms that modulate exploration versus exploitation, downweight unreliable or noisy reward signals and group advantages, and enforce risk-sensitive or uncertainty-guided policy selection to improve robustness under distributional shift and reduce reward hacking.

uncertainty-awarereinforcementlearning

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.49
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Real-world reward signals—such as human preferences—are often uncertain, context-dependent, and inconsistent, leading to reward hacking and over-optimization. This work proposes a dual-source uncertainty-aware reward framework that treats uncertainty as a first-class component of the reward signal, explicitly modeling both epistemic uncertainty in value estimation (via ensemble prediction disagreement) and annotator variability in human preference labels. A confidence-adaptive reliability filter dynamically balances exploration and cautious decision-making based on uncertainty estimates. Experiments demonstrate that the method achieves more stable training in both discrete and continuous environments, reduces trap visits by 93.7%, maintains robust performance under 30% supervisory noise, and significantly mitigates reward hacking behaviors.

alignment failurehuman preferencesreinforcement learning

This work addresses the limitation of traditional reinforcement learning in modeling risk sensitivity, which hinders its applicability in high-stakes scenarios. Focusing on finite-horizon Markov decision processes, the paper presents the first unified treatment of three distinct risk measures—expectile, utility-based shortfall risk, and optimized certainty equivalent—deriving policy gradient theorems for each and designing corresponding gradient estimation algorithms. By leveraging trajectory sampling and smoothness analysis, the authors establish a mean-squared error bound of $O(1/m)$ for the gradient estimators and demonstrate a stable convergence rate for the proposed algorithms. Both theoretical analysis and empirical evaluations on standard reinforcement learning benchmarks confirm the effectiveness and superiority of the developed risk-sensitive policy gradient methods.

expectilesoptimized certainty equivalentpolicy gradient

Selective Uncertainty Propagation in Offline RL

Feb 01, 2023
SK
Sanath Kumar Krishnamurthy
🏛️ Meta | Amazon | Adobe

In finite-horizon offline reinforcement learning, policy evaluation suffers from statistical bias due to the entanglement of future-policy dependence and distributional shift. Method: We propose a selective uncertainty propagation mechanism that adaptively quantifies the difficulty of distributional shift at each time step, decoupling policy iteration from distribution evolution in dynamic programming. By integrating causal inference—specifically treatment effect estimation—with uncertainty quantification, we construct tighter confidence intervals for value estimates. Contribution/Results: Our method jointly optimizes policy evaluation and learning under purely offline settings without online interaction. Experiments demonstrate significant improvements in evaluation robustness and statistical efficiency over state-of-the-art baselines; empirical validation in simulation environments confirms both effectiveness and generalizability.

Finite TimeOffline LearningPolicy Evaluation

Learning to Be Cautious

Oct 29, 2021
MM
Montaser Mohammedalamen
🏛️ University of Alberta | SonyAI | JPMorgan Chase | Amii

In reinforcement learning, agents lack autonomous cautiousness in unseen scenarios; existing approaches rely on manually engineered, task-specific safety constraints, resulting in poor generalization and high deployment costs. Method: We propose an end-to-end framework that—first in the literature—neurally models reward function uncertainty via ensemble learning and couples it with a k-of-N counterfactual regret minimization (CFR) policy optimization mechanism. This enables agents to acquire robust, cautious behavior autonomously, without requiring prior safety specifications. Contribution/Results: Our method achieves safe generalization across multi-level cautiousness tasks with zero hyperparameter tuning, significantly reducing dependence on explicit safety engineering. It establishes a novel paradigm for building adaptive, scalable, and reliable intelligent agents.

Construct robust policies using reward uncertainty and neural networksDevelop agents that learn cautious behavior in novel situationsOvercome reliance on task-specific safety information in reinforcement learning

On the Global Convergence of Risk-Averse Policy Gradient Methods with Expected Conditional Risk Measures

Jan 26, 2023
XY
Xian Yu
🏛️ The Ohio State University | University of Michigan

This work investigates the global convergence of policy gradient (PG) and natural policy gradient (NPG) methods for risk-sensitive reinforcement learning under expectation-based conditional risk measures (ECRMs). For time-consistent ECRMs, we develop a unified PG/NPG algorithmic framework covering four practical parameterizations: constrained direct parameterization, log-barrier regularized softmax, entropy-regularized softmax, and approximate NPG. We establish, for the first time, a rigorous global optimality guarantee and iteration complexity analysis—achieving $O(1/varepsilon^2)$ for PG and $O(1/varepsilon)$ for NPG—for ECRM-based risk optimization, thereby filling a critical theoretical gap in globally convergent risk-sensitive RL. Empirical evaluation on a stochastic Cliffwalk environment demonstrates that the proposed algorithms effectively mitigate risk while maintaining stability and convergence.

Global convergence of risk-averse policy gradient methodsOptimality and complexity of ECRM-based RL algorithmsRisk-sensitive reinforcement learning with dynamic risk measures

Latest Papers

What's happening recently
View more

This work addresses the challenge of safe exploration for reinforcement learning agents in safety-critical scenarios. The authors propose a pessimistic policy optimization method grounded in epistemic uncertainty, which estimates uncertainty through the policy’s sensitivity to parameter perturbations. They introduce a sharpness-aware policy gradient that implicitly reweights gradient updates—amplifying the influence of rare unsafe actions while attenuating contributions from known safe behaviors—to encourage conservative behavior in unexplored regions of the state space. Evaluated across multiple continuous control tasks, the approach significantly outperforms existing baselines, achieving both enhanced task performance and stronger safety guarantees, thereby effectively expanding the Pareto frontier between safety and performance.

epistemic uncertaintypolicy optimizationreinforcement learning

This work addresses the challenge of selecting uncertainty representations that align with decision objectives to achieve optimal and trustworthy decisions under state-variable uncertainty. Drawing on decision theory, it systematically analyzes the optimal forms of uncertainty representation for both risk-neutral and risk-averse agents in known and unknown environments, revealing the minimal uncertainty information required under distinct risk preferences. The study innovatively unifies three approaches to epistemic uncertainty—calibrated prediction, confidence-set robust optimization, and Bayesian inference—establishing a theoretical link between uncertainty representation and decision goals. This integration yields a reliable decision-making framework that provides agents with verifiable utility guarantees.

decision makingepistemic uncertaintyposterior distribution

This work addresses the reward hacking problem in reinforcement learning from human feedback (RLHF), which arises due to errors in reward modeling. The authors propose a pessimistic optimization framework grounded in distributional reward modeling, treating the reward as a random variable \( p(r \mid x, y) \). Within this framework—formulated via Bayesian inference or KL-divergence distributionally robust optimization (KL-DRO)—they unify existing heuristic strategies such as mean aggregation, worst-case optimization, and uncertainty weighting, while clarifying their underlying assumptions. The key contribution is the derivation of a closed-form pessimistic reward function \( \tilde{r}(x, y) = -\beta \log \mathbb{E}_p[e^{-r/\beta}] \), which provides both theoretical justification and practical guidance for mitigating reward hacking in RLHF systems.

distributional reward modelpessimismreward hacking

This work addresses the challenge of error accumulation in large language model agents during tool use, often caused by hallucinations or invocation of unsupported tools. To mitigate this, the paper introduces an uncertainty-aware alignment mechanism for optimizing tool-calling decisions—a novel approach that quantifies action uncertainty and incorporates a repulsive force into the reinforcement learning reward to separate correct from erroneous actions. Combined with lightweight annotations on critical decision turns, the method enables unified post-training over multi-turn trajectories. This strategy enhances exploration signals, alleviates overconfident errors, and significantly improves both decision quality and agent performance across multiple tool-use benchmarks, while preserving well-calibrated uncertainty estimates.

decision reliabilitylarge language modelsreinforcement learning

Hot Scholars

PW

Peng Wu

Technology and Business University
causal inferencepolicy learningrecommender system
NK

Niki Kilbertus

Technical University of Munich & Helmholtz Munich
Machine Learning
XY

Xiaochen Yang

Senior Lecturer, School of Mathematics & Statistics, University of Glasgow
maching learningmedical image analysis
JS

Juan Sebastian Rojas

PhD Student, University of Toronto
Reinforcement LearningMachine LearningRoboticsRisk