π€ AI Summary
This work addresses the challenges faced by pretrained vision-language-action (VLA) policies in contact-rich tasks, where occlusion, depth ambiguity, and force control inaccuracies hinder robust execution. To overcome these limitations, the authors propose the LIFT framework, which enhances real-time contact response through a reactive force-sensing module and introduces, for the first time, a causal force memory mechanism coupled with zero-initialized cross-attention to dynamically refresh action decisions. LIFT further incorporates an online DAgger loop that fuses offline task data with human-corrected trajectories to mitigate distributional shift. Evaluated on towel folding, book insertion, and Tower of Hanoi ring placement tasks, LIFT significantly outperforms purely visual post-training baselines, achieving faster convergence and higher performance. Ablation studies confirm the critical contributions of force memory and online correction to the frameworkβs efficacy.
π Abstract
Pretrained vision-language-action (VLA) policies provide strong language-conditioned manipulation knowledge, but they remain largely vision-driven and can struggle once manipulation enters contact states where the scene is occluded, depth is ambiguous, or small force errors push execution off the offline demonstration distribution. We present LIFT (Late Reactive Injection of Force for VLA Post-Training), a force-aware post-training framework that adds contact reactivity to a pretrained VLA policy while preserving its general manipulation knowledge. LIFT grafts a reactive action expert beside the original action expert, initializes it from pretrained action weights, and injects recent 6D end-effector force through causal force memory and zero-initialized cross attention, enabling actions to be refreshed during execution. To address the policy-dependent distribution shift of contact feedback, LIFT further couples reactive force injection with an online DAgger loop that trains on a mixture of offline task-alignment data and human-corrected online rollouts. Across towel folding, book insertion, and Hanoi ring placement, LIFT learns faster and reaches higher performance than vision-only post-training, while ablations show that reactive force memory and online corrective data are both important for robust contact-rich manipulation. Our code and data will be publicly available.