OmniFysics-Nano-V2 Technical Report: Understanding the Physical World Across Modalities

📅 2026-09-22
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
该研究通过构建物理感知数据管道和采用两阶段训练策略,解决了全模态模型在物理属性理解上的不足,提升了AI系统对物理世界的认知能力。
📝 Abstract
Omni-modal models have expanded multimodal interaction across vision, audio, speech, and language. However, their training is predominantly organized around semantic descriptions and general-purpose objectives, leaving physical attributes, interaction states, and causal mechanisms only partially specified. This gap is not simply a matter of modality coverage: adding more modalities does not by itself provide the supervision needed to connect observations with the physical structure of the world. We present OmniFysics-Nano-V2, a compact omni-modal model for physical-world perception and understanding. The model supports image, video, audio, speech, and text inputs within a shared reasoning framework, together with text and speech generation. To address the lack of explicit physical supervision, we construct a dual-branch physics-aware data pipeline that grounds salient objects in structured physical attributes and aligns visual changes with acoustic events, intermediate responses, and interaction outcomes. To address homogeneous training objectives, we curate reinforcement-learning prompts by reward diversity and adopt a two-stage Group Relative Policy Optimization curriculum that progresses from general task correctness to fine-grained physical perceptual reasoning. Experiments across multimodal, audio-visual, and physical reasoning benchmarks show that the proposed data and training strategy improves physical-world understanding while preserving broad omni-modal competence. The proposed model achieves leading result on 17 of 21 benchmarks against SOTA omni-modal models. By equipping AI systems with both omni-modal and physical-world perception capabilities, OmniFysics-Nano-V2 is poised to become a cornerstone of next-generation Physical AI.
Problem

Research questions and friction points this paper is trying to address.

omni-modal models
physical attributes
interaction states
causal mechanisms
Innovation

Methods, ideas, or system contributions that make the work stand out.

omni-modal model
physics-aware data pipeline
reinforcement learning prompts
Group Relative Policy Optimization
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Yizhou Liu
Yizhou Liu
MIT
Dynamical systemsStatistical physicsPhysics of living systemsPhysics of AI
J
Jinghang Han
Physical Superintelligence Lab, Fysics AI; College of Intelligent Robotics and Advanced Manufacturing, Fudan University
K
Kaixiang Qiu
Physical Superintelligence Lab, Fysics AI; College of Intelligent Robotics and Advanced Manufacturing, Fudan University
Qi He
Qi He
Fudan University
LLM
Minghao Han
Minghao Han
Ph.D Student at Fudan University
Computational pathology
Yue Jiang
Yue Jiang
Fudan University
Multimodal LearningLarge Language ModelsNatural Language Processing
X
Xujia Chen
Physical Superintelligence Lab, Fysics AI; College of Intelligent Robotics and Advanced Manufacturing, Fudan University
Wei Zou
Wei Zou
PKU、Samsung、Baidu、Didi、Ke
SpeechNLPLLMMultimodal
Shunli Wang
Shunli Wang
Fudan University
Computer visionAction quality assessment
Lihua Zhang
Lihua Zhang
Wuhan University
computational biologybioinformaticsdata mining
Dingkang Yang
Dingkang Yang
ByteDance
Multimodal LearningGenerative AIEmbodied AI