🤖 AI Summary
This study addresses the difficulty of multimodal large language models in inferring underlying physical mechanisms from continuous physical field observations. We propose the first benchmark dedicated to continuous physical field understanding, encompassing analytical equations, numerical simulations, and real-world observational data. Through post-training strategies including supervised fine-tuning, chain-of-thought prompting, and reinforcement learning, we systematically evaluate and optimize model performance on mechanism identification and state prediction tasks. Our findings reveal that zero-shot performance is nearly random, exposing significant deficiencies in mapping visual representations to physical laws. Notably, combining reinforcement learning with chain-of-thought supervision yields optimal cross-task generalization, offering an effective pathway for enhancing the physical reasoning capabilities of large models.
📝 Abstract
Multimodal large language models (MLLMs) are increasingly envisioned as core components of scientific and engineering agents, yet their ability to interpret physical fields remains poorly understood. Existing physics benchmarks largely emphasize textbook problem solving or intuitive physical reasoning, leaving open whether MLLMs can infer physically meaningful information from continuous field observations. We introduce PhysFieldBench, a benchmark comprising 24 tasks and 1,160 evaluation examples across controlled equation fields, simulated physical fields, and observed physical fields. The tasks assess three forms of inference: identifying physical mechanisms, comparing latent control variables, and predicting outcome properties. Across representative open-source and proprietary MLLMs, zero-shot performance is low: the best model achieves a chance-normalized score of 29.3, while several open-source models remain near chance. In contrast, a task-specific supervised vision transformer performs substantially better, demonstrating that the inputs contain learnable physical information. To diagnose these failures, a structured self-explanation analysis attributes most errors to missed visual patterns and incorrect visual-to-physical mappings. Further, to explore whether post-training can improve physical inference and generalize to unseen tasks, we compare supervised fine-tuning with final answers or chain-of-thought supervision and reinforcement learning. Final-answer supervision performs best overall but transfers less effectively, whereas reinforcement learning after chain-of-thought supervision achieves the best generalization. Together, these findings highlight the need to improve visual-to-physical grounding and cross-task generalization for MLLMs to reliably interpret physical fields in scientific and engineering workflows.