Register-Routed Delayed Fusion: Rewiring Shortcut-Prone Observation Fusion in Visuomotor Imitation

📅 2026-10-02
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the insufficient policy robustness in visuomotor imitation learning, where compact action signals prematurely interfere with visual representations. We propose RRDF, which introduces a novel "isolate-collect-route" scheduling mechanism. By masking direct cross-modal attention and incorporating a learnable register workspace, RRDF delays interaction to protect early visual processing while preserving action-conditioning capabilities. Implemented via a Transformer-based architecture, this approach enables staged cross-modal fusion that effectively severs shortcut connections. Experiments demonstrate that RRDF matches or surpasses dense ACT baselines across both simulated and real-world robotic tasks, yielding significantly improved robustness against out-of-distribution appearance variations and positional shifts.
📝 Abstract
Visuomotor imitation policies combine high-dimensional visual observations with compact signals such as proprioception, and their fusion topology determines when and through which tokens these streams interact. In dense token fusion, visual tokens may attend directly to compact tokens from the first encoder layer, allowing action-predictive compact cues to influence spatial visual representations early in their formation. We ask whether controlling this route improves visual responsiveness and policy behavior. We introduce Register-Routed Delayed Fusion (RRDF), which masks direct compact-visual attention and stages cross-modal interaction through a learned register workspace. Its isolate-collect-route schedule protects an early stream-separated prefix and later permits only register-mediated exchange, while compact conditioning remains available to the native action generator. Across five simulation tasks and three real-robot tasks, RRDF matches or improves dense ACT under nominal conditions. Appearance-shift evaluations on four simulation tasks and held-out-position evaluations on three real-robot tasks also favor RRDF. Phase-matched input probes show lower measured state-to-image sensitivity, while ablations indicate that adding registers alone does not reproduce the full performance gain. These results support controlling cross-modal propagation while retaining compact action conditioning.
Innovation

Methods, ideas, or system contributions that make the work stand out.

Register-Routed Delayed Fusion
Visuomotor Imitation
Cross-modal Attention Masking
Token Fusion Topology
Learned Register Workspace
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
J
Jieting Long
School of Computer Science, The University of Sydney, Australia
Weidong Cai
Weidong Cai
Clinical Associate Professor, Stanford University School of Medicine
functional neuroimagingmachine learningcognitivedevelopmentalclinical neuroscience
W
Weiming Zhi
School of Computer Science, The University of Sydney, Australia