๐ค AI Summary
This study addresses the limited capabilities of multimodal agents in materials science software, stemming from domain data scarcity and complex interfaces. We construct the first real-world GUI agent benchmark tailored to this domain, integrating a fine-grained scoring mechanism with a multi-level evaluation pipeline to systematically assess agent performance across graphical interaction, OriginPro script execution, and database querying. Evaluations conducted within a Windows 11 virtual machine alongside humanโmachine consistency experiments demonstrate that general-purpose model capabilities transfer poorly, with the best-performing model achieving only a 25% success rate on GUI tasks. By bridging this critical gap, this work provides an essential testbed for diagnosing and enhancing the operational proficiency of AI agents within specialized scientific software environments.
๐ Abstract
Multimodal GUI agents have achieved impressive results on general software benchmarks, yet their ability to operate professional scientific software remains largely unexplored. In materials science, sparse domain-specific web data, specialized interfaces, and tacit workflow conventions create blind spots that general-purpose pretraining cannot readily bridge. We present MatToolBench, the first real-environment benchmark for evaluating multimodal GUI agents on professional materials science software, comprising 204 tasks across 10 tools in three modalities: GUI operation, OriginPro scripting, and code-based database queries, all executed inside a Windows 11 VM. Each task is decomposed into fine-grained sub-criteria by domain experts, enabling interpretable partial-credit scoring; the GUI component of our multi-level evaluation pipeline achieves an average F1 of 0.98. For OriginPro figure-generation tasks, we further conduct a human-LLM agreement study to validate the use of a multimodal judge for secondary aesthetic assessment. Our experiments show that strong performance on general benchmarks does not transfer to professional scientific workflows, and that this gap is not a visual-grounding problem alone: failures arise from domain-specific operational knowledge, sparse pretraining coverage of scientific software, weak cross-tool artifact handoff, and critical states exposed only visually. Even the best model reaches only 25% success rate on GUI tasks and 45% on code tasks. MatToolBench therefore serves as a challenging diagnostic benchmark and real-environment testbed for data-scarce, knowledge-intensive scientific workflows.