PathTime-VLA: Path-Time Decoupling for Factorized Post-Training of Vision-Language-Action Policies

📅 2026-10-08
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenges of teleoperation adaptation and low execution efficiency in Vision-Language-Action (VLA) policies caused by the coupling of spatial paths and temporal dynamics. We introduce, for the first time, classical path-time parameterization into learned action representations to decouple geometric prediction from execution speed. Specifically, a diffusion model generates interaction paths, while Speed-DQN learns an execution multiplier and Path-AWR optimizes the path generator. Furthermore, DAgger facilitates intervention-based data collection to support staged post-training. Experimental evaluations across three tasks demonstrate that the proposed method achieves a 58/60 success rate and reduces average completion time by 39%–52% compared to baselines, significantly enhancing robotic control efficiency.
📝 Abstract
Vision-Language-Action (VLA) policies typically predict actions at fixed time intervals, coupling the route a robot follows with its execution pace. This coupling complicates adaptation from teleoperation: useful geometric guidance comes with timing shaped by interface delays and operator behavior. Our key insight is to bring the path-time parameterization of classical motion planning into the learned action representation of a VLA. We introduce PathTime-VLA, which represents motion as a progress-indexed interaction path $X(s)$ and a positive interval-time profile. The latter defines a monotone time law $t(s)$, yielding controller commands $X(s(t))$. For a given path, alternative executions are expressed through the time profile, allowing chunk-wise speed choices without changing the geometric prediction target. This representation supports a staged post-training procedure: demonstrations and DAgger interventions establish a target-domain prior, Speed-DQN learns execution multipliers from robot interaction, and Path-AWR uses rollout outcomes to refine the diffusion path generator. A path-conditioned action expert realizes the resulting motions while maintaining distinct learning interfaces for path generation and execution timing. Across three tasks, the complete method achieves $58/60$ successes versus $57/60$ for PathTime-VLA under BC + DAgger at fixed $1\times$, with approximately $39$-$52\%$ shorter mean completion times over successful trials.
Problem

Research questions and friction points this paper is trying to address.

Vision-Language-Action
path-time coupling
teleoperation adaptation
robot policy learning
Innovation

Methods, ideas, or system contributions that make the work stand out.

Vision-Language-Action
Path-Time Decoupling
Factorized Post-Training
Speed-DQN
Diffusion Policy
🔎 Similar Papers
2024-03-04Computer Vision and Pattern RecognitionCitations: 3
💼 Related Jobs
No related jobs found.