Deep Epistemic Value Functions for Optimistic Exploration

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the fragility of epistemic uncertainty estimation and the inconsistent exploration performance in deep reinforcement learning by proposing the DEVOTE algorithm. Through a systematic investigation of uncertainty representation and propagation mechanisms, this method innovatively controls uncertainty generalization, stabilizes temporal propagation, and adapts to non-stationary targets, thereby optimizing optimistic exploration strategies in model-free reinforcement learning. Experimental results demonstrate that DEVOTE effectively reaches novel states in both reward-free exploration and continuous control tasks. It significantly enhances the robustness and scalability of exploration while achieving higher returns that surpass existing baselines.
📝 Abstract
Principled exploration in reinforcement learning requires an agent to quantify its epistemic uncertainty and act to resolve it. Uncertainty over the value function provides a natural signal for exploration, yet existing deep approximations remain brittle and perform inconsistently. The central challenge is therefore to scale these ideas robustly. We conduct a systematic empirical study of how epistemic uncertainty is represented, propagated, and optimized in deep epistemic value functions, and uncover distinct failure modes along each of these axes. These findings motivate DEVOTE, a model-free reinforcement learning algorithm that controls how uncertainty generalizes beyond observed data, stabilizes its temporal propagation, and preserves adaptation to the resulting non-stationary exploration objective. Across reward-free exploration and challenging continuous-control tasks, DEVOTE reaches novel states more effectively and achieves higher task return than strong model-free and model-based exploration baselines. These results provide evidence that deep epistemic value functions are a promising path toward scalable, principled exploration.
Problem

Research questions and friction points this paper is trying to address.

reinforcement learning
epistemic uncertainty
exploration
deep value functions
Innovation

Methods, ideas, or system contributions that make the work stand out.

Epistemic Uncertainty
Deep Value Functions
Reinforcement Learning
Optimistic Exploration
Model-free Algorithm
L
Leander Diaz-Bone
ETH Zürich, Switzerland; Max Planck Institute for Intelligent Systems, Germany
M
Marco Bagatella
ETH Zürich, Switzerland; Max Planck Institute for Intelligent Systems, Germany
Jonas Hübotter
Jonas Hübotter
PhD student, ETH Zurich
artificial intelligencemachine learningreinforcement learningtest time training
A
Andreas Krause
ETH Zürich, Switzerland