🤖 AI Summary
This study addresses the descriptive inaccuracies of omni-modal models arising from their neglect of physical evidence such as contact and deformation. To this end, we propose a physics-aware omni-modal captioning framework. Methodologically, we construct an evidence-driven OPC benchmark and the Daily-Physics dataset, introducing a dedicated physics-aware model to extract spatiotemporally aligned cross-modal interaction cues. By integrating multimodal agents for tool coordination, image-level fine-tuning, and large-scale video caption supervision, our approach achieves deep fusion of audiovisual semantics with physical evidence. Experimental results demonstrate that the proposed model attains state-of-the-art performance across multiple video captioning benchmarks, rivaling Gemini 3.1 Pro while significantly enhancing both physical coverage and descriptive reliability.
📝 Abstract
Building omni-modal models with physical intelligence requires fine-grained supervision that captures physical evidence such as contact, support, deformation, and state transitions. However, existing omni-modal captioners primarily model general audiovisual semantics and often overlook transient or spatially localized physical evidence. We present a unified framework for physics-aware audiovisual captioning spanning data construction, training, and evaluation. Firstly, we build a data construction pipeline that identifies physics-rich clips and leverages OmniFysics-Agent to coordinate audio, visual, and physical-perception tools for collecting spatiotemporally aligned and traceable cross-modal evidence; within the Agent, a physical perception model (PPM) fine-tuned on approximately 2M image-level samples serves as a dedicated tool for extracting object-interaction and state-change cues. Secondly, we build the Daily-Physics 50K dataset and introduce the evidence-driven OmniPhysCap (OPC) benchmark to evaluate the recovery of physical and cross-modal evidence from generated captions. Finally, we train OmniFysics-Captioner from the resulting data. Our Captioner matches Gemini 3.1 Pro on audiovisual captioning, achieves state-of-the-art results on multiple video-captioning benchmarks, and substantially outperforms other open-source models. Ablations show that PPM evidence improves physical coverage and produces finer-grained, more reliable cross-modal descriptions.