Score
Design and evaluate algorithms, policies, and curricula that choose actions or data-collection behaviors to balance exploration and exploitation and to generate informative, diverse, and cost-aware trajectories (including active evidence inquiry, active/agentic or goal-directed exploration, and adaptive exploration curricula). This work includes building and analyzing stochastic exploration policies and scoring mechanisms, entropy- or covariance-based regulators, exploration–exploitation reweighting schemes (e.g., multi-particle, interaction-aware, or state-aware reweighting), and methods to preserve heterogeneity while trading off information gain and interaction cost.
This paper addresses the model-free reinforcement learning problem in continuous-time stochastic linear-quadratic (LQ) control, where the system volatility depends on both state and control, and no running control reward is present. To overcome inefficiencies and hyperparameter sensitivity inherent in conventional fixed-exploration strategies, we propose an adaptive exploration mechanism within an actor-critic framework: the critic dynamically adjusts the entropy regularization strength, while the actor real-time modulates policy variance, enabling online balancing of exploration and exploitation. The method requires no prior model knowledge and—first among continuous-time LQ settings—achieves a sublinear regret bound matching the best known optimal rate. It breaks the limitations of static exploration scheduling, substantially reducing hyperparameter tuning effort and accelerating convergence. Numerical experiments demonstrate superior regret performance and faster convergence compared to both non-adaptive model-free and model-based baselines.
This work addresses the limited adaptive exploration capability of current large language model agents at test time, which struggle to determine when to explore. The authors propose an exploration-aware reinforcement learning framework that constructs a fine-grained reward function via variational inference to evaluate the potential future value of exploratory actions. An exploration-aware grouping mechanism is introduced to selectively trigger exploration under high uncertainty, while decoupling exploratory and task-execution actions during policy optimization. Evaluated across diverse text-based and GUI agent benchmarks, the method consistently improves performance, significantly enhancing both exploration efficiency and overall task success.
Efficient autonomous exploration in sparse-reward environments remains a fundamental challenge in reinforcement learning. This work proposes a novel paradigm that decouples exploration from policy optimization: during the exploration phase, it abandons conventional reinforcement learning and instead employs a “Go-With-The-Winner” tree search guided by epistemic uncertainty to actively expand state coverage; subsequently, it distills the collected exploration trajectories into a deployable policy via supervised inverse dynamics learning. The approach requires neither expert demonstrations nor domain-specific knowledge and operates directly from end-to-end pixel inputs. It substantially outperforms existing methods on challenging Atari benchmarks such as Montezuma’s Revenge, Pitfall!, and Venture, achieving an order-of-magnitude improvement in exploration efficiency. Notably, it is the first method to solve high-dimensional continuous-control sparse-reward tasks—including MuJoCo Adroit and AntMaze—directly from pixels.
To address low data efficiency in model-based reinforcement learning, this paper proposes PTS-BE, a framework that drives agent exploration via Bayesian information gain (IG) rewards targeting regions of high epistemic uncertainty. Theoretically, we prove that the IG reward strictly quantifies model uncertainty and converges to zero upon full knowledge acquisition, providing rigorous theoretical grounding for Bayesian exploration. Methodologically, PTS-BE integrates sparse variational Gaussian processes, deep kernel learning, and deep ensembles to enable scalable, high-fidelity posterior modeling. Empirical evaluations demonstrate that PTS-BE significantly outperforms state-of-the-art baselines on sparse-reward and pure-exploration tasks, achieving 2–5× improvements in sample efficiency. These results validate its dual capability: efficient environment exploration and effective policy learning under stringent data budgets.
This work addresses the limited cognitive sensitivity of conventional uncertainty measures—such as Shannon entropy—in autonomous robotic environmental exploration. We propose Behavioral Entropy, a novel, behaviorally grounded metric inspired by prospect theory in behavioral economics. Its core innovation is the first integration of the Prelec probability weighting function into robotics exploration, yielding a falsifiable, generalized entropy formulation that better aligns with human perception of uncertainty. Based on this, we design a cognitively sensitive frontier selection utility function. The approach is validated in both ROS-Unity co-simulation and real-world experiments on a Clearpath Warthog platform. Results demonstrate that Behavioral Entropy–driven exploration significantly outperforms Shannon and Rényi entropy–based strategies in exploration efficiency and coverage, while maintaining computational tractability for real-time deployment.
This work addresses the challenge of distorted information-theoretic exploration objectives in high-dimensional robotic systems, where most parameter directions are weakly observable or unidentifiable. To resolve this, the authors propose Quasi-Optimal Experimental Design (QOED), which introduces optimal experimental design to robotic exploration for the first time. QOED leverages eigenspace analysis of the Fisher information matrix to identify the observable subspace, adaptively focusing exploration on identifiable parameter directions while amplifying informative dimensions and suppressing irrelevant ones in the exploration objective. This approach yields a constant-factor approximation to the ideal information gain. Empirical results demonstrate significant performance improvements—35.23% in simulation and 21.98% in real-world navigation and manipulation tasks—substantially outperforming existing reinforcement learning baselines.
This work addresses the challenge of safely harnessing large language models (LLMs) in high-throughput experimental optimization, where direct LLM use risks unsafe exploration yet complete exclusion forfeits their optimization potential. To reconcile this trade-off, the authors propose the CARE framework, which employs a non-LLM default optimizer as the primary pathway while leveraging the LLM to generate candidate strategies. Adoption of these candidates is governed by an evidence-based intervention gating mechanism that audits proposals against publicly available evidence, ensuring decisions are auditable, controllable, and traceable. By synergistically integrating LLM-driven creativity with evidence-guided safety constraints, CARE achieves state-of-the-art performance on the Minerva/Olympus and ChemLex benchmarks, improving peak scores from 80.0 to 88.5 and from 83.9 to 92.1, respectively.
This work addresses the instability and frequent collapse in multi-turn reinforcement learning caused by inefficient exploration. The authors propose T²PO, a novel framework that introduces uncertainty-aware mechanisms at both token and turn granularities for the first time. By dynamically monitoring exploration states and triggering thinking-based interventions alongside turn-level resampling, T²PO enables fine-grained control over the exploration process, effectively mitigating unproductive interactions and training inefficiencies. Extensive experiments on multi-turn task environments—including WebShop, ALFWorld, and Search QA—demonstrate that the proposed method substantially enhances training stability and task performance, validating its significant improvement in exploration efficiency.