Score
Designs and builds multimodal, action-conditioned predictive world models that represent and simulate the coupled visual and tactile sensorimotor dynamics of an embodied agent, producing contact-consistent visuo-tactile rollouts and future visual and tactile trajectories. These models generate visuo-tactile predictions and localized correction segments conditioned on actions and task progress, and can be instantiated as robot-centric or agent-centric simulators for planning, control, or model-based learning.
This study addresses a critical gap in the literature: the absence of a unified survey on robotic learning methods that integrate force and tactile perception, particularly regarding the synthesis of multimodal sensing and multi-stage system design. To bridge this gap, the paper introduces the TF-ART categorization framework—a novel, comprehensive architecture that systematically encompasses multimodal perceptual inputs, hierarchical action generation, and reactive low-level control. By integrating heterogeneous sensor encoding, multimodal perception fusion, and action refinement mechanisms, the framework elucidates the intrinsic relationships among existing approaches and clearly maps their design logic across the perception–decision–execution pipeline. This contribution provides a holistic perspective, theoretical foundation, and practical guidance for developing intelligent systems capable of rich physical interaction.
This work addresses the lack of a unified predictive framework for world models in robotic manipulation, which has led to fragmented research and ambiguous design choices. Focusing on three core questions—what to predict, how to relate predictions to actions, and when to use predictions—the paper proposes a functional taxonomy that distinguishes between integrated prediction-action models and explicit predictive planners, positioning world models as foundational predictive infrastructure for robot learning. The study presents a systematic review covering latent dynamics models, action-conditioned video generation, 3D/4D scene prediction, physics simulators, and prediction modules in vision–language–action systems. It also consolidates evaluation protocols across 34 manipulation datasets, highlighting open challenges such as contact modeling and hallucination control through the lenses of prediction fidelity, task performance, and simulation reliability.
Dexterous manipulation requires seamless integration of high-level task planning and low-level contact responses, yet existing approaches struggle to effectively fuse tactile feedback with semantic instructions and exhibit limited robustness under perturbations. This work proposes a hierarchical policy architecture that, for the first time, leverages tactile signals simultaneously for predictive contact modeling and high-frequency residual correction, thereby decoupling slow visual-language subtask planning from rapid tactile reactions. The framework integrates vision, language, touch, and proprioception, achieving a success rate of 65.0% in clean conditions and 53.7% under human-induced disturbances across six long-horizon tasks with high contact complexity—outperforming the strongest baseline by 15.7 and 18.5 percentage points, respectively.
This work addresses the physical inconsistencies—such as object disappearance or teleportation—that arise in purely visual world models when handling contact-rich tasks due to occlusion or ambiguous contact cues. To overcome this limitation, the paper presents the first systematic multimodal world model that integrates tactile perception with vision in a unified framework. Leveraging an autoregressive recurrent prediction architecture, the model jointly learns contact dynamics from both visual and tactile inputs, significantly enhancing physical fidelity during imagined rollouts and planning. Experimental results demonstrate a 33% improvement in object permanence and a 29% increase in adherence to motion regularities. Furthermore, in zero-shot real-robot tasks, the model achieves up to a 35% higher success rate and exhibits strong capabilities for rapid adaptation to novel tasks.
In contact-rich manipulation tasks, vision alone often fails to capture critical physical interaction cues, and existing world models couple future prediction with action decision-making, limiting effective use of multimodal prospective information. To address this, this work proposes the Oracle Visuo-Tactile Foresight (OVTF) framework, which decouples prediction from decision-making by leveraging oracle paired visuo-tactile future states from simulation. OVTF introduces an asymmetric phase-local future memory (AFM) architecture that enables phase-aligned, cross-modal selective routing. Combined with modality-isolated contrastive learning and the UniVTAC simulation platform, OVTF achieves an average success rate of 32.0% across seven tasks, significantly outperforming modality-isolated methods (23.7%) and the baseline UniVTAC-ACT (14.9%).
Current world models suffer from conceptual ambiguity in embodied intelligence and generative simulation, lacking a unified classification and design framework tailored for robotic control. This work formally defines a world model as one conditioned on actions to predict the future evolution of task-relevant observations or states, and introduces a novel paradigm—world action models—that explicitly links prediction with executable actions. Building on this definition, the study systematically organizes four methodological families: imagine-then-execute, video-feature-conditioned action prediction, joint video-action modeling, and policy learning augmented by auxiliary video prediction. The research clarifies the conceptual boundaries of (action-conditioned) world models and presents the first structured taxonomy specifically designed for embodied prediction and control, thereby advancing standardized understanding and application in the field.
This work addresses a critical gap in the evaluation of action-conditioned world models (ACWMs), which has predominantly emphasized visual fidelity or task performance while neglecting their core function as physical simulators—namely, the causal fidelity between actions and environmental responses. To this end, we formalize the notion of an "observable simulator contract" and introduce WorldSimProbe, a fine-grained diagnostic framework that assesses ACWMs across five dimensions: local control sensitivity, global trajectory variation, multi-source action consistency, interaction grounding, and dynamics. Built upon controlled testing protocols, our framework leverages calibration analysis, dense action-motion correspondence, and spurious interaction detection. Evaluated on RoboTwin, ManiSkill, and LIBERO across six open-source models (>18,000 instances), it reveals systematic deficiencies in action execution, interaction grounding, and dynamics, with results strongly aligned with human judgment and downstream task performance.
This work addresses the limitations of existing world action models, which rely heavily on visual prediction and struggle to effectively supervise physical interactions—such as force, deformation, shear, and slip—during contact-rich manipulation, often rendering touch a privileged cue for action generation. To overcome this, the paper introduces TacWAM, the first framework to integrate mechanics-aware tactile future prediction with an information-constraining mechanism. It employs spatially aligned fusion of a tactile encoder, a tactile history encoder, and anchor-guided trimodal attention to model the dynamic evolution of force and deformation, while preventing future tactile information from leaking into the action branch. Evaluated on four real-world tasks, TacWAM achieves an average success rate of 75.0%, outperforming the strongest baseline by 37.5 percentage points, with ablation studies confirming the critical roles of tactile history modeling and information constraints.