StructRL: Online Structured Reinforcement Learning for Long-Horizon Vision-Language-Action Tasks

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenges Vision-Language-Action (VLA) models face in long-horizon tasks, including sparse terminal rewards, training inefficiency, and difficulty distinguishing early-stage failures. To overcome these limitations, we propose an online reinforcement learning framework that introduces the first subtask-dependency-based structured intermediate reward mechanism. By automatically decomposing tasks into verifiable subtasks to construct intermediate supervision and dynamically scaling rewards according to completion progress, our approach achieves precise credit assignment and transcends the constraints of traditional sparse supervision. This method is compatible with VLA foundation models such as GR00T-N1.5 and π₀.₅. Extensive experiments demonstrate that it consistently outperforms existing online RL baselines on the RoboCasa365 and LIBERO-Long benchmarks, highlighting its effectiveness for complex robotic manipulation tasks.
📝 Abstract
Vision-language-action (VLA) models perform well on shorter-horizon manipulation tasks but still struggle with long-horizon tasks that require multiple dependent manipulations from a single command. Online reinforcement learning (RL) can improve these policies through environment interaction, yet many existing methods provide reward only after the complete task succeeds. However, such terminal supervision is sparse and does not distinguish early failures from rollouts that make substantial partial progress. We propose StructRL, an online RL framework that constructs structured intermediate supervision from verifiable subtask completions. StructRL decomposes each task into verifiable subtasks, grants intermediate rewards only after the prerequisite subtasks have been completed, and scales each reward according to completion pace. Across RoboCasa365 and LIBERO-Long with GR00T-N1.5 and pi 0.5, StructRL consistently outperforms evaluated online RL baselines. These results show that verifiable, structured intermediate rewards improve long-horizon VLA post-training. Code is available at https://github.com/amazon-science/StructRL.
Problem

Research questions and friction points this paper is trying to address.

Vision-Language-Action models
Long-horizon tasks
Online reinforcement learning
Sparse reward
Intermediate supervision
Innovation

Methods, ideas, or system contributions that make the work stand out.

Structured Reinforcement Learning
Vision-Language-Action Models
Long-Horizon Tasks
Intermediate Rewards
Subtask Decomposition