🤖 AI Summary
This study addresses the phase confusion problem in long-horizon robotic manipulation caused by visual similarities or velocity discrepancies, proposing the PACE method. PACE introduces a novel execution-progress-based demonstration reinterpretation mechanism that continuously parses demonstrations through progress-aligned contexts. During training, it employs dual-margin attention supervision to uncover latent phase structures, enabling phase-label-free inference at test time. Furthermore, the method integrates multimodal prompt compression, local fast-weight memory, and a unified diffusion action expert. Experimental results demonstrate that PACE substantially improves success rates from 33.3% to 73.3% on tasks such as LIBERO-Gen Goal Chain, effectively overcoming the phase confusion bottleneck in long-horizon manipulation.
📝 Abstract
Demonstration-conditioned policies provide a natural interface for specifying robot behavior, yet long-horizon manipulation remains difficult when visually similar states recur across different stages or when demonstrations and executions proceed at different speeds. We identify the resulting failure mode as stage confusion and introduce Progress-Aligned Context for Execution (PACE), a stateful method that continually reinterprets a complete demonstration according to realized execution progress. PACE compresses the demonstration into ordered multimodal prompt tokens and uses training-only dual-edge attention supervision to expose its latent stage structure. During execution, an episode-local fast-weight memory causally encodes realized action-observation transitions and modulates prompt cross-attention, producing a progress-aligned context for a unified diffusion action expert without test-time stage labels or stage-specific policies. PACE improves success from 88.9% to 94.0% on LIBERO-Gen Goal Chain, from 79.1% to 83.3% on Spatial Combination, and from 33.3% to 73.3% on the two-step Block Routing tasks. Failure analysis further indicates that structured demonstration alignment and causal execution memory jointly mitigate stage confusion.