OmniFysics-Captioner Technical Report: Grounding Omni-Modal Understanding in the Physical World for Better Captioning

📅 2026-09-21
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the descriptive inaccuracies of omni-modal models arising from their neglect of physical evidence such as contact and deformation. To this end, we propose a physics-aware omni-modal captioning framework. Methodologically, we construct an evidence-driven OPC benchmark and the Daily-Physics dataset, introducing a dedicated physics-aware model to extract spatiotemporally aligned cross-modal interaction cues. By integrating multimodal agents for tool coordination, image-level fine-tuning, and large-scale video caption supervision, our approach achieves deep fusion of audiovisual semantics with physical evidence. Experimental results demonstrate that the proposed model attains state-of-the-art performance across multiple video captioning benchmarks, rivaling Gemini 3.1 Pro while significantly enhancing both physical coverage and descriptive reliability.
📝 Abstract
Building omni-modal models with physical intelligence requires fine-grained supervision that captures physical evidence such as contact, support, deformation, and state transitions. However, existing omni-modal captioners primarily model general audiovisual semantics and often overlook transient or spatially localized physical evidence. We present a unified framework for physics-aware audiovisual captioning spanning data construction, training, and evaluation. Firstly, we build a data construction pipeline that identifies physics-rich clips and leverages OmniFysics-Agent to coordinate audio, visual, and physical-perception tools for collecting spatiotemporally aligned and traceable cross-modal evidence; within the Agent, a physical perception model (PPM) fine-tuned on approximately 2M image-level samples serves as a dedicated tool for extracting object-interaction and state-change cues. Secondly, we build the Daily-Physics 50K dataset and introduce the evidence-driven OmniPhysCap (OPC) benchmark to evaluate the recovery of physical and cross-modal evidence from generated captions. Finally, we train OmniFysics-Captioner from the resulting data. Our Captioner matches Gemini 3.1 Pro on audiovisual captioning, achieves state-of-the-art results on multiple video-captioning benchmarks, and substantially outperforms other open-source models. Ablations show that PPM evidence improves physical coverage and produces finer-grained, more reliable cross-modal descriptions.
Problem

Research questions and friction points this paper is trying to address.

omni-modal captioning
physical intelligence
audiovisual semantics
physical evidence
fine-grained supervision
Innovation

Methods, ideas, or system contributions that make the work stand out.

physics-aware captioning
physical perception model
omni-modal understanding
evidence-driven benchmark
cross-modal alignment
🔎 Similar Papers
💼 Related Jobs
No related jobs found.
K
Kaixiang Qiu
Physical Superintelligence Lab, Fysics AI; College of Intelligent Robotics and Advanced Manufacturing, Fudan University
Minghao Han
Minghao Han
Ph.D Student at Fudan University
Computational pathology
K
Keliang Liu
Physical Superintelligence Lab, Fysics AI; College of Intelligent Robotics and Advanced Manufacturing, Fudan University
Yizhou Liu
Yizhou Liu
MIT
Dynamical systemsStatistical physicsPhysics of living systemsPhysics of AI
J
Jinghan Han
Physical Superintelligence Lab, Fysics AI; College of Intelligent Robotics and Advanced Manufacturing, Fudan University
Yue Jiang
Yue Jiang
Fudan University
Multimodal LearningLarge Language ModelsNatural Language Processing
X
Xuecheng Wu
Physical Superintelligence Lab, Fysics AI; College of Intelligent Robotics and Advanced Manufacturing, Fudan University
Shunli Wang
Shunli Wang
Fudan University
Computer visionAction quality assessment
Lihua Zhang
Lihua Zhang
Wuhan University
computational biologybioinformaticsdata mining
Dingkang Yang
Dingkang Yang
ByteDance
Multimodal LearningGenerative AIEmbodied AI