PhysAlign: A Benchmark for Evidence-Grounded Role Alignment in Multimodal Physics Reasoning

πŸ“… 2026-09-27
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
Multimodal models frequently exhibit reasoning errors in physical diagram understanding due to mismatches between visual elements and their corresponding physical roles. This work introduces PhysAlign, a benchmark that employs local probes to decouple recognition from role assignment and proposes five metrics, including CAcc, to isolate correspondence errors for systematically evaluating the physical semantic alignment capabilities of models. Experiments on nearly one thousand questions, generated via human-validated probes and controlled variants, reveal that advanced models such as GPT-6-Astra incur conditional correspondence error rates ranging from 13.8% to 50.6%. These findings expose a fundamental bottleneck characterized by strong perception but weak grounding, demonstrating that successful visual recognition does not guarantee reliable physical understanding.
πŸ“ Abstract
A key challenge in physics diagram understanding is correctly associating visual information with the physical entities, relations, and conditions it describes. Even when a value, symbol, or other local element is accurately recognized, assigning it to the wrong entity or scope can distort the underlying physical premise and lead to incorrect reasoning. To systematically study this challenge, we introduce \textbf{PhysAlign}, a benchmark designed to assess whether multimodal models correctly associate information recognized from physics diagrams with its intended physical role. By disentangling visual recognition from physical-role assignment through localized probes and controlled variants, PhysAlign isolates correspondence errors from recognition failures. It contains 3,341 human-validated probes spanning 986 physics problems, enabling systematic evaluation of visual recognition and physical-role correspondence at scale. We further introduce five complementary evaluation metrics, including CAcc, GAcc, and JAcc, which provide a comprehensive assessment of models'ability to recognize diagram content, establish correct physical correspondences, and solve the underlying physics problem. Across our evaluated multimodal models, PhysAlign reveals a consistent gap between local visual recognition and physical-role grounding. Even when the queried content is correctly recognized, the conditional correspondence error rate remains 13.8\% for GPT-6-Astra and rises to about 50.6\% for InternVL3.5-8B. These findings indicate that strong perception alone does not ensure reliable physical interpretation, exposing a distinct grounding bottleneck that is largely hidden by answer-level accuracy and highlighting the need for future models to better align recognized visual evidence with its physical meaning.
Innovation

Methods, ideas, or system contributions that make the work stand out.

PhysAlign Benchmark
Physical-Role Grounding
Multimodal Physics Reasoning
Decoupled Evaluation Metrics
Correspondence Error Isolation
πŸ”Ž Similar Papers
No similar papers found.
πŸ’Ό Related Jobs
No related jobs found.