🤖 AI Summary
This study addresses the lack of interpretable verification and error-correction feedback in scientific image generation, as well as the insufficient domain coverage of existing natural image verifiers. We propose a multimodal reasoning verification framework that introduces the SciGen-Verify benchmark and a three-tier hierarchical verification protocol, establishing a closed loop from binary judgment to structured correction instructions. Methodologically, we design a pipeline combining cold-start supervised fine-tuning with curriculum-based two-stage reinforcement learning, incorporating a rubric-guided process reward model (PRM) to enhance scientific reasoning exploration. Experiments demonstrate that our approach achieves performance comparable to large proprietary models. Furthermore, when deployed as an online critic, it effectively supports iterative image refinement, significantly improving both structural correctness and knowledge consistency in scientific image generation.
📝 Abstract
In realistic education, a solution is often expressed not only in words but in a drawing--a circuit, a geometric construction, a function plot--and a teacher must grade the drawing as carefully as the text. Recent advances in unified multimodal models have enabled scientific image generation, yet verifying the correctness of these specialized visual outputs remains a critical bottleneck: errors often arise from intricate domain knowledge, structural reasoning, and multi-step instruction rather than surface-level artifacts. Existing verifiers mainly target natural images and compress judgement into scalar scores, leaving scientific coverage and explainable feedback for error correction underexplored. To bridge this gap, we make three main contributions. (1) We construct SciGen-Verify, a benchmark dedicated to explainable verification of scientific image generation, spanning instruction following, multidisciplinary reasoning, and world knowledge domains. It contains a three-tier hierarchical protocol over the binary judgement, supporting explanation, and corrective editing instruction. (2) We develop SciGen-Verifier, a reasoning-driven multimodal verifier trained via cold-start supervised fine-tuning followed by a curriculum-based two-stage reinforcement learning pipeline. The rubric-guided process rewards first strengthen scientific reasoning exploration and outcome rewards subsequently align output with ground-truth annotation. (3) On SciGen-Verify, SciGen-Verifier achieves competitive performance against much larger proprietary models. It further serves as a practical online critic for iterative image rectification.