ToolSciVer: Multimodal Scientific Claim Verification with Visual Tool Augmented Reinforcement Learning

📅 2026-07-17
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses key challenges in verifying scientific claims through multimodal information fusion, particularly the difficulty of locating critical visual evidence and inaccuracies in parsing structured figures and tables. The authors propose the first tool-augmented multimodal verification framework, which integrates type-aware visual tools—such as row/column-focused table analysis, structural chart parsing, and high-resolution region zooming—to transform complex visualizations into explicit, claim-aligned evidence. Coupled with Grouped Relative Policy Optimization (GRPO), a reinforcement learning strategy that enhances efficient tool invocation and robust reasoning, the framework achieves state-of-the-art performance across five vision-language models from the Qwen, InternVL, and Gemma families. Extensive experiments on the SciVer and MuSciClaims benchmarks demonstrate substantial improvements over four strong baselines.
📝 Abstract
Multimodal Scientific Claim Verification (MSCV) requires models to verify scientific claims using visually grounded evidence from papers, including figures, tables, charts, and textual context. However, existing methods often fail because they struggle to locate decisive visual evidence, accurately read structured scientific visuals, and integrate multimodal observations into reliable reasoning. We introduce ToolSciVer, the first tool-augmented framework for MSCV to our knowledge. ToolSciVer equips a VLM with three type-aware visual tools, table row/column focus, chart-to-structure parsing, and high-resolution region zoom, which convert dense scientific visuals into explicit, claim-facing evidence, and trains the policy with Group Relative Policy Optimization (GRPO) under a composite reward of answer correctness, format validity, length control, tool-use efficiency, and tool-validity penalties. Experiments on SciVer and MuSciClaims datasets on five VLMs from three model families (Qwen, InternVL, Gemma) demonstrate that our method achieves superior performance compared to four competitive baselines including prompting-based and RL-based tool-use methods, highlighting the effectiveness of learned, type-aware tool use for scientific claim verification.
Problem

Research questions and friction points this paper is trying to address.

Multimodal Scientific Claim Verification
visual evidence
scientific visuals
multimodal reasoning
claim verification
Innovation

Methods, ideas, or system contributions that make the work stand out.

tool-augmented reasoning
multimodal scientific claim verification
visual language models
reinforcement learning
structured visual parsing
🔎 Similar Papers
No similar papers found.