tactile-aware world modeling

Designs and builds multimodal, action-conditioned predictive world models that represent and simulate the coupled visual and tactile sensorimotor dynamics of an embodied agent, producing contact-consistent visuo-tactile rollouts and future visual and tactile trajectories. These models generate visuo-tactile predictions and localized correction segments conditioned on actions and task progress, and can be instantiated as robot-centric or agent-centric simulators for planning, control, or model-based learning.

tactile-awareworldmodeling

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
-0.32
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

This work addresses the lack of a unified predictive framework for world models in robotic manipulation, which has led to fragmented research and ambiguous design choices. Focusing on three core questions—what to predict, how to relate predictions to actions, and when to use predictions—the paper proposes a functional taxonomy that distinguishes between integrated prediction-action models and explicit predictive planners, positioning world models as foundational predictive infrastructure for robot learning. The study presents a systematic review covering latent dynamics models, action-conditioned video generation, 3D/4D scene prediction, physics simulators, and prediction modules in vision–language–action systems. It also consolidates evaluation protocols across 34 manipulation datasets, highlighting open challenges such as contact modeling and hallucination control through the lenses of prediction fidelity, task performance, and simulation reliability.

action-conditioned predictionbenchmarkingpredictive modeling

Dexterous manipulation requires seamless integration of high-level task planning and low-level contact responses, yet existing approaches struggle to effectively fuse tactile feedback with semantic instructions and exhibit limited robustness under perturbations. This work proposes a hierarchical policy architecture that, for the first time, leverages tactile signals simultaneously for predictive contact modeling and high-frequency residual correction, thereby decoupling slow visual-language subtask planning from rapid tactile reactions. The framework integrates vision, language, touch, and proprioception, achieving a success rate of 65.0% in clean conditions and 53.7% under human-induced disturbances across six long-horizon tasks with high contact complexity—outperforming the strongest baseline by 15.7 and 18.5 percentage points, respectively.

contact stabilitydexterous manipulationforce feedback

This work addresses the physical inconsistencies—such as object disappearance or teleportation—that arise in purely visual world models when handling contact-rich tasks due to occlusion or ambiguous contact cues. To overcome this limitation, the paper presents the first systematic multimodal world model that integrates tactile perception with vision in a unified framework. Leveraging an autoregressive recurrent prediction architecture, the model jointly learns contact dynamics from both visual and tactile inputs, significantly enhancing physical fidelity during imagined rollouts and planning. Experimental results demonstrate a 33% improvement in object permanence and a 29% increase in adherence to motion regularities. Furthermore, in zero-shot real-robot tasks, the model achieves up to a 35% higher success rate and exhibits strong capabilities for rapid adaptation to novel tasks.

contact-rich manipulationobject permanencephysical fidelity

In contact-rich manipulation tasks, vision alone often fails to capture critical physical interaction cues, and existing world models couple future prediction with action decision-making, limiting effective use of multimodal prospective information. To address this, this work proposes the Oracle Visuo-Tactile Foresight (OVTF) framework, which decouples prediction from decision-making by leveraging oracle paired visuo-tactile future states from simulation. OVTF introduces an asymmetric phase-local future memory (AFM) architecture that enables phase-aligned, cross-modal selective routing. Combined with modality-isolated contrastive learning and the UniVTAC simulation platform, OVTF achieves an average success rate of 32.0% across seven tasks, significantly outperforming modality-isolated methods (23.7%) and the baseline UniVTAC-ACT (14.9%).

contact-rich manipulationfuture-to-action interfacemodality alignment

Current world models suffer from conceptual ambiguity in embodied intelligence and generative simulation, lacking a unified classification and design framework tailored for robotic control. This work formally defines a world model as one conditioned on actions to predict the future evolution of task-relevant observations or states, and introduces a novel paradigm—world action models—that explicitly links prediction with executable actions. Building on this definition, the study systematically organizes four methodological families: imagine-then-execute, video-feature-conditioned action prediction, joint video-action modeling, and policy learning augmented by auxiliary video prediction. The research clarifies the conceptual boundaries of (action-conditioned) world models and presents the first structured taxonomy specifically designed for embodied prediction and control, thereby advancing standardized understanding and application in the field.

action-conditioned predictionembodied intelligencegenerative simulation

Latest Papers

What's happening recently
View more

This work addresses a critical gap in the evaluation of action-conditioned world models (ACWMs), which has predominantly emphasized visual fidelity or task performance while neglecting their core function as physical simulators—namely, the causal fidelity between actions and environmental responses. To this end, we formalize the notion of an "observable simulator contract" and introduce WorldSimProbe, a fine-grained diagnostic framework that assesses ACWMs across five dimensions: local control sensitivity, global trajectory variation, multi-source action consistency, interaction grounding, and dynamics. Built upon controlled testing protocols, our framework leverages calibration analysis, dense action-motion correspondence, and spurious interaction detection. Evaluated on RoboTwin, ManiSkill, and LIBERO across six open-source models (>18,000 instances), it reveals systematic deficiencies in action execution, interaction grounding, and dynamics, with results strongly aligned with human judgment and downstream task performance.

action-conditioned world modelsembodied manipulationmodel evaluation

This work addresses the limitations of existing world action models, which rely heavily on visual prediction and struggle to effectively supervise physical interactions—such as force, deformation, shear, and slip—during contact-rich manipulation, often rendering touch a privileged cue for action generation. To overcome this, the paper introduces TacWAM, the first framework to integrate mechanics-aware tactile future prediction with an information-constraining mechanism. It employs spatially aligned fusion of a tactile encoder, a tactile history encoder, and anchor-guided trimodal attention to model the dynamic evolution of force and deformation, while preventing future tactile information from leaking into the action branch. Evaluated on four real-world tasks, TacWAM achieves an average success rate of 75.0%, outperforming the strongest baseline by 37.5 percentage points, with ablation studies confirming the critical roles of tactile history modeling and information constraints.

action generationcontact-rich manipulationforce and deformation

Hot Scholars

YL

Yifu Lu

Undergraduate, University of Michigan
Computer Science
BW

Boyang Wang

University of Virginia
Video GenerationVideo EnhancementRobot Learning
JC

Jianyu Chen

Assistant Professor, Tsinghua University
AIRobotics
YG

Yanjiang Guo

Tsinghua University
Embodied AIGenerative Model