InternW0: A Foundational Physical World Model for Efficient Real-World Interactions

📅 2026-09-23
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
研究引入InternW0模型,通过异步多频处理和局部物理建模解决物理智能中的预测与行动问题,支持高效现实世界交互。
📝 Abstract
Physical intelligence requires more than predicting how the world may evolve: predictions must remain actionable as the world continues to change. We introduce InternW0, the first instantiation of the InternW physical world model series from Shanghai AI Laboratory, built around omnimodal interfaces, asynchronous multi-frequency processing, and local physical modeling under partial observations and external influences. InternW0 jointly learns future visual dynamics and continuous robot control through an asymmetric video--action architecture with flow matching. A high-capacity video expert provides longer-horizon predictive context, while a lightweight action expert operates at a faster timescale. Instead of regenerating the future for every action update, InternW0 reuses layerwise K/V and adapts it to newly observed states through observation-conditioned context routing. Domain-specific interfaces and soft prompts support heterogeneous embodiments, while contact-aware post-training incorporates force and tactile signals for contact-rich manipulation. We train InternW0 on approximately 7,200 hours of heterogeneous robot and egocentric data, including EgoLab, a 275-hour real-laboratory egocentric dataset. Evaluation spans simulation benchmarks and real-world scientific tasks, including a 15-stage metal--organic framework synthesis workflow and 5-stage contact- and force-aware dexterous manipulation for general-purpose quantitative pipetting. These results advance scalable, asynchronous, and science-native physical world models for universal and efficient real-world interactions.
Problem

Research questions and friction points this paper is trying to address.

Physical Intelligence
Asynchronous Multi-frequency Processing
Partial Observations
Real-world Interactions
Omnimodal Interfaces
Innovation

Methods, ideas, or system contributions that make the work stand out.

asynchronous multi-frequency processing
omnimodal interfaces
flow matching
observation-conditioned context routing
contact-aware post-training
🔎 Similar Papers
2024-07-09IEEE/ASME transactions on mechatronicsCitations: 94
J
Jisong Cai
Shanghai Artificial Intelligence Laboratory
Y
Yao Mu
Shanghai Artificial Intelligence Laboratory
Ganlin Yang
Ganlin Yang
University of Science and Technology of China && Shanghai AI Laboratory
Computer vision3D visionMultimodal models
Z
Zhe Cao
Shanghai Artificial Intelligence Laboratory
Z
Zhangzheng Tu
Shanghai Artificial Intelligence Laboratory
Xing Gao
Xing Gao
Shanghai Artificial Intelligence Laboratory
Graph LearningRobotic LearningAutonomous Driving
Kailin Li
Kailin Li
Shanghai AI Lab
Computer Vision3D VisionEmbodied AI
Xinyu Zhan
Xinyu Zhan
Shanghai Jiao Tong University
L
Lixin Yang
Shanghai Artificial Intelligence Laboratory
Y
Yangkun Zhu
Shanghai Artificial Intelligence Laboratory
Haoxiang Ma
Haoxiang Ma
Beihang University
GraspingRobotic Manipulation
Ming Zhou
Ming Zhou
Researcher; Shanghai AI Laboratory
Multi-Agent LearningReinforcement LearningEmbodied AI
Qiaojun Yu
Qiaojun Yu
Shanghai Jiao Tong University, Shanghai AI Lab
robotic learning3D visionvla
Y
Yufei Xue
Shanghai Artificial Intelligence Laboratory
L
Liqun He
Shanghai Artificial Intelligence Laboratory
Y
Yifei Yao
Shanghai Artificial Intelligence Laboratory
Yifan Zhu
Yifan Zhu
Beijing University of Posts and Telecommunications
PEFT of LLMsGraph RAGGraph mining
Long Ling
Long Ling
Tongji University.
Human AI InteractionHCIDigital Fabrication
B
Bingqi Jiang
Shanghai Artificial Intelligence Laboratory
Haoyu Guo
Haoyu Guo
Shanghai AI Lab
Computer Vision3D Vision
X
Xueyue Zhu
Shanghai Artificial Intelligence Laboratory
B
Bowen Zhou
Shanghai Artificial Intelligence Laboratory
Bin Zhao
Bin Zhao
Northwestern Polytechnical University, Shanghai AI Laboratory
Computer VisionEmbodied Artificial Intelligence
Tianfan Xue
Tianfan Xue
Information Engineering Department, The Chinese University of Hong Kong
Computer VisionMachine LearningComputational Photography
Chunhua Shen
Chunhua Shen
Zhejiang University
Computer VisionMachine Learning