🤖 AI Summary
This study addresses the limitation that maximizing intrinsic rewards in reinforcement learning does not necessarily yield the most informative exploration experiences. We propose a formal criterion for exploration grounded in counterfactual information and integrate reinforcement learning with Pareto optimization analysis to reveal the Pareto suboptimality of prevalent intrinsic reward objectives, along with their underlying failure mechanisms. Building on these insights, we construct an environment verification framework, establish theoretical conditions for achieving optimal exploration, and design an improved exploration objective function. This work offers a novel perspective for understanding the inherent limitations of intrinsic rewards and lays a rigorous theoretical foundation for developing more efficient exploration strategies.
📝 Abstract
Intrinsic rewards are designed to guide exploration in reinforcement learning by assigning value to an agent's experience, for example through prediction error or learning progress. However, maximizing these rewards need not produce the most informative experience available. We propose a formal criterion for exploration that compares policies by the counterfactual information they acquire: how well their histories can substitute for experience under alternative policies. We construct a single, simple environment in which specified count-based, prediction-error, empowerment, and information-gain objectives have maximizing policies that are Pareto-suboptimal at acquiring counterfactual information. We explain these failures and establish conditions under which existing intrinsic rewards successfully encourage optimal exploration. We also construct an objective that assigns a higher value whenever exploration strictly improves under our criterion.