Bidirectional Voronoi-biased Exploration Curriculum for Reinforcement Learning

📅 2026-10-02
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the long-horizon exploration bottleneck in goal-conditioned reinforcement learning under sparse rewards by proposing BVER, a bidirectional curriculum learning framework. By integrating bidirectional RRT planning with a Voronoi diagram-biased sampling mechanism, BVER expands the state space simultaneously from both initial and goal states while actively biasing toward unexplored regions to generate an automatic curriculum, which is then combined with goal-conditioned policy optimization for efficient training. As the first bidirectional Voronoi-biased exploration method, BVER operates entirely without reference demonstrations. Experimental results demonstrate that the proposed approach significantly accelerates convergence and improves sample efficiency and robustness across diverse tasks, outperforming existing demonstration-free baselines.
📝 Abstract
Long-horizon tasks with sparse rewards pose an exploration bottleneck for goal-conditioned reinforcement learning: a policy started from the initial state rarely reaches the goal and receives no learning signal. Reference motions, hand-designed curricula, and shaped rewards supply this signal but require demonstrations or task-specific engineering; automatic start-state and goal curricula avoid this but typically expand from one side only, so the full distance to the target must be covered from that side. We propose the Bidirectional Voronoi-biased Exploration curriculum for Reinforcement learning (BVER), which expands from both ends at once. Inspired by bidirectional RRT planning, BVER grows start states outward from the goal and goals outward from the initial state distribution, biases both toward unexplored task space, and steers them toward each other, training one goal-conditioned policy on both. On point-mass mazes, quadrupedal box climbing, and robot-arm ring-on-peg transfer, BVER learns faster than all compared reference-free curricula. On box climbing, it reaches 95% success on a 0.4 m box in roughly 65% fewer iterations than the best of them, is the only one of them to learn to climb a 0.7 m box, and yields a policy robust to start, goal, and yaw variation. Without a demonstration, it approaches the sample efficiency of reference-based curricula on the 0.4 m box and on ring-on-peg transfer. Ablations show that expanding from both ends outperforms either direction alone.
Problem

Research questions and friction points this paper is trying to address.

Reinforcement Learning
Sparse Rewards
Exploration Bottleneck
Goal-Conditioned
Curriculum Learning
Innovation

Methods, ideas, or system contributions that make the work stand out.

Bidirectional curriculum
Voronoi-biased exploration
Goal-conditioned reinforcement learning
Sparse rewards
Automatic curriculum
J
Juri Pfammatter
Robotic Systems Lab, ETH Zürich, Zürich, Switzerland
K
Kaixian Qu
Robotic Systems Lab, ETH Zürich, Zürich, Switzerland
C
Clemens Schwarke
Robotic Systems Lab, ETH Zürich, Zürich, Switzerland; NVIDIA
Victor Klemm
Victor Klemm
PhD Student at Robotic Systems Lab, ETH Zurich
Robotics
Marco Hutter
Marco Hutter
Professor of Robotics, ETH Zurich
Legged RoboticsRoboticsControl