Disentangling Visuo-Tactile Foresight: Oracle-Guided Interface Discovery for World Action Models

📅 2026-08-01
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
In contact-rich manipulation tasks, vision alone often fails to capture critical physical interaction cues, and existing world models couple future prediction with action decision-making, limiting effective use of multimodal prospective information. To address this, this work proposes the Oracle Visuo-Tactile Foresight (OVTF) framework, which decouples prediction from decision-making by leveraging oracle paired visuo-tactile future states from simulation. OVTF introduces an asymmetric phase-local future memory (AFM) architecture that enables phase-aligned, cross-modal selective routing. Combined with modality-isolated contrastive learning and the UniVTAC simulation platform, OVTF achieves an average success rate of 32.0% across seven tasks, significantly outperforming modality-isolated methods (23.7%) and the baseline UniVTAC-ACT (14.9%).
📝 Abstract
Contact-rich manipulation remains challenging because successful control depends on physical interaction cues that are often weakly observable from vision alone. Recent tactile world action models jointly model future visual observations and tactile signals to guide action generation, but how such futures should be structured for effective use by the action expert remains underexplored. Directly studying this question with learned world action models is difficult because end-to-end behavior entangles physically invalid visual futures, unreliable predictions, inaccurate or cross-modally inconsistent tactile forecasts, and an unreadable future-to-action interface. To make this interface independently studyable, we introduce Oracle Visuo-Tactile Foresight (OVTF), a controlled framework that supplies paired RGB and tactile futures from successful trajectories verified in simulation. By fixing the future provider, OVTF isolates the interface and asks a cleaner question: if the future is successful and physically executable, what representation allows the action expert to absorb its benefit? Within OVTF, we propose Asymmetric Phase-Local Future Memory (AFM), in which visual memory reads future vision, each tactile memory jointly attends to its own tactile stream and phase-aligned future vision, and cross-tactile access is blocked. We compare AFM with Modality-Isolated Future Memory (IFM), which removes visual-to-tactile access and processes each future modality independently. Across seven tasks on the UniVTAC simulation benchmark, AFM achieves 32.0% average success, compared with 23.7% for IFM and 14.9% for UniVTAC-ACT. This controlled comparison shows that selective phase-aligned visual-tactile routing provides a more actionable future-to-action bridge than complete modality isolation.
Problem

Research questions and friction points this paper is trying to address.

visuo-tactile foresight
world action models
contact-rich manipulation
future-to-action interface
modality alignment
Innovation

Methods, ideas, or system contributions that make the work stand out.

visuo-tactile foresight
world action models
modality disentanglement
phase-aligned memory
contact-rich manipulation
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Z
Zihang Yao
Brigham Young University, Provo, Utah, USA
C
Chaoyue Ding
Beijing Academy of Science and Technology, Beijing, China
Y
Yingying Yu
Beijing Academy of Science and Technology, Beijing, China