PhyMo: A Physical-Field Modality for Multimodal AI4Physics

📅 2026-09-23
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
本文提出PhyMo框架,通过物理场模态和PDE关联算子解决物理系统预测中的多模态信息融合问题,实验表明其表现优于现有方法。
📝 Abstract
Multimodal learning is emerging as a powerful paradigm for AI for Physics (AI4Physics), where predicting physical systems requires the joint interpretation of heterogeneous observations, measurements, and domain knowledge. However, existing approaches typically represent physical quantities and governing equations as generic numerical or textual tokens, overlooking the physical constraints that determine their spatiotemporal interactions. To address this limitation, we introduce the \textbf{physical-field modality} and propose \textbf{PhyMo}, a physics-grounded multimodal framework that organizes heterogeneous measurements through PDE-associated operators. PhyMo follows a three-stage learning procedure: the physical-field encoder is first pretrained through field reconstruction under PDE residual supervision, its representations are subsequently aligned with visual embeddings in a shared latent space, and the fused multimodal representations are finally processed by corresponding downstream prediction heads. Experiments on five datasets spanning diverse physical environments show that PhyMo achieves state-of-the-art performance, compared to the strongest baseline on each dataset, demonstrating the superiority of PhyMo on multimodal representation learning in AI4Physics.
Problem

Research questions and friction points this paper is trying to address.

multimodal learning
physical constraints
spatiotemporal interactions
Innovation

Methods, ideas, or system contributions that make the work stand out.

physical-field modality
PhyMo
PDE-associated operators
multimodal representation learning
🔎 Similar Papers
💼 Related Jobs
No related jobs found.
Henan Sun
Henan Sun
Beijing Institute of Technology
Graph Neural Networks (GNNs)Graph Invariant LearningDifferential Privacy
H
Haitao Hu
The Hong Kong University of Science and Technology (Guangzhou)
J
Jin Liu
Huawei Noah’s Ark Lab
J
Jianfeng Zhang
Huawei Noah’s Ark Lab
Lujia Pan
Lujia Pan
Noah's Ark Lab, Huawei
Anomaly dectionTime seriesRepresentation learning
Nuo Chen
Nuo Chen
Hong Kong University of Science and Technology
large language modelpre-trainreasoningrole-playing
J
Jia Li
The Hong Kong University of Science and Technology (Guangzhou), The Hong Kong University of Science and Technology