risk-guided exploration

Design, implement, and evaluate exploration algorithms and policies that explicitly account for risk and uncertainty, producing exploration behavior that is robust to rare or high-cost outcomes and that can be directed toward or away from risky trajectories as required. This includes building risk-sensitive reward shaping, guided or bridge exploration strategies, methods that combine probabilistic uncertainty estimates (e.g., Dirichlet) with sampling selection mechanisms (e.g., Gumbel-topk), and procedures to prioritize or deprioritize long-tail edge cases during search.

risk-guidedexploration

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.34
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Must-Read Papers

Most classic and influential ideas
View more

Uncertainty-driven Adaptive Exploration

Sep 03, 2025
LB
Leonidas Bakopoulos
🏛️ Technical University of Crete

Adaptive balancing of exploration and exploitation remains challenging in learning long-horizon, complex action sequences due to difficulty in determining optimal timing for exploration–exploitation trade-offs. Method: This paper proposes a cognitive-uncertainty-driven adaptive exploration framework that online quantifies dual uncertainty—over both the environment model and the policy—to dynamically modulate exploration intensity and switching timing. It unifies diverse uncertainty sources (e.g., model prediction variance, policy confidence) within a modular architecture supporting plug-and-play integration of heterogeneous uncertainty estimators, and incorporates intrinsic motivation to enable uncertainty-guided policy optimization. Results: Evaluated on multiple MuJoCo continuous-control benchmarks, the framework significantly outperforms baseline methods—including entropy regularization and Random Network Distillation—demonstrating superior effectiveness, robustness, and cross-task generalization capability.

Determining optimal switching between exploration and exploitation phasesDeveloping principled uncertainty-driven adaptive exploration frameworkLearning complex action sequences in challenging domains

Efficient autonomous exploration in sparse-reward environments remains a fundamental challenge in reinforcement learning. This work proposes a novel paradigm that decouples exploration from policy optimization: during the exploration phase, it abandons conventional reinforcement learning and instead employs a “Go-With-The-Winner” tree search guided by epistemic uncertainty to actively expand state coverage; subsequently, it distills the collected exploration trajectories into a deployable policy via supervised inverse dynamics learning. The approach requires neither expert demonstrations nor domain-specific knowledge and operates directly from end-to-end pixel inputs. It substantially outperforms existing methods on challenging Atari benchmarks such as Montezuma’s Revenge, Pitfall!, and Venture, achieving an order-of-magnitude improvement in exploration efficiency. Notably, it is the first method to solve high-dimensional continuous-control sparse-reward tasks—including MuJoCo Adroit and AntMaze—directly from pixels.

autonomous explorationexplorationhard exploration

Dual-Directed Algorithm Design for Efficient Pure Exploration

Oct 30, 2023
CQ
Chao Qin
🏛️ Columbia University | The Hong Kong University of Science and Technology

This paper addresses complex pure-exploration objectives beyond best-arm identification—such as threshold testing and ε-optimal arm identification—by establishing the first duality-based minimax-optimal sampling allocation framework. It provides the first necessary and sufficient conditions for optimal sampling allocation in pure exploration; generalizes the top-two paradigm to arbitrary pure-exploration problems; and proposes a hyperparameter-free, information-directed selection rule driven by KL divergence and entropy. The rule is rigorously proven to achieve asymptotic optimality in Gaussian settings and resolves the long-standing open problem of asymptotic optimality for top-two Thompson sampling. Experiments demonstrate substantial improvements in sampling efficiency across Gaussian best-arm identification, threshold-bandwidth testing, and ε-optimal arm identification, consistently outperforming state-of-the-art methods.

Develops optimal adaptive experimentation for pure-exploration goalsExtends top-two approach beyond best-arm identificationResolves asymptotic optimality in Gaussian best-arm identification

Robotic Exploration Using Generalized Behavioral Entropy

Feb 15, 2024
AS
A. Suresh
🏛️ U.S. Army Combat Capabilities Development Command Army Research Laboratory (ARL) | UC San Diego

This work addresses the limited cognitive sensitivity of conventional uncertainty measures—such as Shannon entropy—in autonomous robotic environmental exploration. We propose Behavioral Entropy, a novel, behaviorally grounded metric inspired by prospect theory in behavioral economics. Its core innovation is the first integration of the Prelec probability weighting function into robotics exploration, yielding a falsifiable, generalized entropy formulation that better aligns with human perception of uncertainty. Based on this, we design a cognitively sensitive frontier selection utility function. The approach is validated in both ROS-Unity co-simulation and real-world experiments on a Clearpath Warthog platform. Results demonstrate that Behavioral Entropy–driven exploration significantly outperforms Shannon and Rényi entropy–based strategies in exploration efficiency and coverage, while maintaining computational tractability for real-time deployment.

Compares Behavioral entropy with Shannon and RenyiEnhances robot exploration speed using Behavioral entropyIntroduces Behavioral entropy for robotic exploration

Latest Papers

What's happening recently
View more

This work addresses the challenge of epistemic uncertainty in early-stage online reinforcement learning, where scarce data necessitate a careful trade-off between robustness and exploration. The authors propose the Quantile-based Bayesian Risk Markov Decision Process (BR-MDP), which modulates the influence of posterior uncertainty in Bellman backups via quantile control and introduces an adaptive quantile scheduling mechanism that prioritizes robustness initially and gradually promotes exploration as data accumulate. Theoretical analysis establishes the asymptotic normality of the value function estimation error and proves a sublinear Bayesian regret bound relative to both the true optimal policy and the BR-MDP’s robust optimal policy. Empirical results demonstrate that the proposed method significantly outperforms baseline approaches in environments characterized by either high exploration demands or high exploration costs.

Bayesian riskepistemic uncertaintyexploration

This work addresses spacecraft trajectory optimization under unknown initial state and process noise distributions by proposing a general robust optimization framework that does not rely on assumptions about uncertainty distributions. The approach first generates a deterministic nominal trajectory offline, then constructs an affine closed-loop correction law comprising feedforward and time-varying feedback gains via chance-constrained reinforcement learning. It employs rolling sampling to estimate the upper-tail quantiles of probabilistic constraints and incorporates a covariance feasibility penalty to regulate terminal dispersion. Evaluated on three-dimensional Earth-to-Mars multi-impulse transfers and continuous-thrust precision landing scenarios, the method achieves competitive fuel performance while guaranteeing probabilistic feasibility, demonstrating strong cross-task generalization and robustness.

chance-constraineddistribution-agnosticrobustness

This work addresses safe trajectory planning under model uncertainty by proposing a Dual-gatekeeper framework that, for the first time, jointly incorporates safety constraints and task-performance budgets within a dual-controller architecture. The approach guarantees formal safety while triggering active exploration only when such exploration can be verified to improve long-term performance. By synergistically combining robust planning with conditional exploration, the method balances immediate task execution with the reduction of uncertainty. Experimental evaluations in quadrotor and autonomous racing scenarios demonstrate that the proposed framework generates trajectories that are not only provably safe and highly efficient but also exhibit strong adaptability, significantly outperforming existing baselines.

active explorationbudget-constrained optimizationdual control

This study addresses the challenge of real-time trajectory optimization in drilling operations under geological uncertainty and measurement noise. The authors propose a belief-driven sequential decision-making framework that integrates particle filtering to model uncertainty in subsurface states with multiple reinforcement learning strategies—including approximate dynamic programming, deep Q-learning, and double deep reinforcement learning—within a unified belief space. A novel smoothness metric is introduced to evaluate decision stability, and high-fidelity policy comparisons are conducted using an industrial-grade geosteering simulator. Experimental results demonstrate that the proposed approach significantly improves final wellbore placement accuracy while enhancing the smoothness and interpretability of the decision process, all under identical geological and operational constraints.

decision optimizationgeosteeringsequential decision-making

Hot Scholars

DT

Dacheng Tao

Nanyang Technological University
artificial intelligencemachine learningcomputer visionimage processing
YC

Yuanpei Chen

South China University of Technology
Robotic
YY

Yaodong Yang

Boya (博雅) Assistant Professor at Peking University
Reinforcement LearningAI AlignmentEmbodied AI
BH

Biwei Huang

UCSD
CausalityMachine LearningComputational Science
OA

Odalric-Ambrym Maillard

Inria Lille - Nord Europe
Multi-armed BanditsStochastic Dynamical SystemsStatistical LearningReinforcement Learning