SCULPT-VLA: Learning Structured Control through Staged Action Grounding

📅 2026-09-19
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文提出SCULPT-VLA方法,通过分阶段动作基础学习结构化控制,解决了视觉-语言-动作策略中如何有效利用中间表示的问题。
📝 Abstract
Vision-language-action (VLA) policies increasingly incorporate structured intermediate supervision beyond action labels. Yet specifying what an intermediate representation should encode leaves open how action prediction learns to depend on it. We introduce \textbf{SCULPT-VLA}, a policy that learns structured control through staged action grounding. Its action-conditioning state comprises complementary factors for task progression, scene dynamics, and spatial grounding. Training first forms these factors with teacher scaffolds, then grounds coarse action prediction through their composition as scaffold inputs are withdrawn. Direct perceptual access is subsequently restored for continuous refinement, combining the learned state with perceptual detail. The curriculum separates learning to condition actions on structure from refining continuous control. Deployment requires neither teachers nor discrete-action autoregression. SCULPT-VLA achieves higher average success than shared-backbone baselines on LIBERO, SimplerEnv-WidowX, and RoboTwin 2.0 Full. On SimplerEnv-WidowX, final success is 83.5\%, versus 71.3\% when Stage-II action learning directly accesses vision and language. Across four physical robot tasks, average success under the tested distribution shifts reaches 58.1\%, compared with 45.6\% for $π_{0.5}$. Training ablations and factor-wise interventions support the staged design and show that the learned state continues to contribute to control after direct perceptual access is restored.
Problem

Research questions and friction points this paper is trying to address.

structured control
staged action grounding
VLA policies
intermediate representation
action prediction
Innovation

Methods, ideas, or system contributions that make the work stand out.

staged action grounding
structured control
curriculum learning
state factorization
💼 Related Jobs
No related jobs found.
Wenbo Li
Wenbo Li
The Chinese University of Hong Kong
Computer VisionDeep Learning
Y
Yiteng Chen
School of Software Engineering, South China University of Technology, Guangzhou, China
W
Wei Zhang
School of Software Engineering, South China University of Technology, Guangzhou, China
Wenhao Li
Wenhao Li
Marshall School of Business, University of Southern California and NBER
Asset PricingFinancial IntermediationMacroeconomics
J
Jun Yang
Yuanwu Technology, Shenzhen, China
Qingyao Wu
Qingyao Wu
School of Software Engineering, South China University of Technology
Computer VisionMachine Learning