MatToolBench: Benchmarking Multimodal Agents in Real-World Materials Science Workflows

๐Ÿ“… 2026-09-29
๐Ÿ“ˆ Citations: 0
โœจ Influential: 0
๐Ÿ“„ PDF
๐Ÿค– AI Summary
This study addresses the limited capabilities of multimodal agents in materials science software, stemming from domain data scarcity and complex interfaces. We construct the first real-world GUI agent benchmark tailored to this domain, integrating a fine-grained scoring mechanism with a multi-level evaluation pipeline to systematically assess agent performance across graphical interaction, OriginPro script execution, and database querying. Evaluations conducted within a Windows 11 virtual machine alongside humanโ€“machine consistency experiments demonstrate that general-purpose model capabilities transfer poorly, with the best-performing model achieving only a 25% success rate on GUI tasks. By bridging this critical gap, this work provides an essential testbed for diagnosing and enhancing the operational proficiency of AI agents within specialized scientific software environments.
๐Ÿ“ Abstract
Multimodal GUI agents have achieved impressive results on general software benchmarks, yet their ability to operate professional scientific software remains largely unexplored. In materials science, sparse domain-specific web data, specialized interfaces, and tacit workflow conventions create blind spots that general-purpose pretraining cannot readily bridge. We present MatToolBench, the first real-environment benchmark for evaluating multimodal GUI agents on professional materials science software, comprising 204 tasks across 10 tools in three modalities: GUI operation, OriginPro scripting, and code-based database queries, all executed inside a Windows 11 VM. Each task is decomposed into fine-grained sub-criteria by domain experts, enabling interpretable partial-credit scoring; the GUI component of our multi-level evaluation pipeline achieves an average F1 of 0.98. For OriginPro figure-generation tasks, we further conduct a human-LLM agreement study to validate the use of a multimodal judge for secondary aesthetic assessment. Our experiments show that strong performance on general benchmarks does not transfer to professional scientific workflows, and that this gap is not a visual-grounding problem alone: failures arise from domain-specific operational knowledge, sparse pretraining coverage of scientific software, weak cross-tool artifact handoff, and critical states exposed only visually. Even the best model reaches only 25% success rate on GUI tasks and 45% on code tasks. MatToolBench therefore serves as a challenging diagnostic benchmark and real-environment testbed for data-scarce, knowledge-intensive scientific workflows.
Problem

Research questions and friction points this paper is trying to address.

Multimodal GUI agents
Materials science workflows
Scientific software benchmarking
Domain-specific knowledge
Data-scarce environments
Innovation

Methods, ideas, or system contributions that make the work stand out.

Multimodal GUI Agents
Scientific Benchmark
Materials Science
Fine-grained Evaluation
Cross-modal Tasks
๐Ÿ”Ž Similar Papers
Mei Wu
Mei Wu
Hangzhou Dianzi University
R
Rui Xie
X-LANCE Lab, School of Computer Science, Shanghai Jiao Tong University, Shanghai, China; State Key Laboratory for General Artificial Intelligence, BIGAI, Beijing, China
R
Runyu Zhang
X-LANCE Lab, School of Computer Science, Shanghai Jiao Tong University, Shanghai, China
Yuqiang Li
Yuqiang Li
Central South University
Internal Combustion EngineCombustionEmissionsMechansim
Tianfan Fu
Tianfan Fu
Nanjing University
AI for DrugAI for ScienceLarge Language Model
B
Bo Chen
Suzhou Laboratory, Suzhou, China
K
Kai Yu
X-LANCE Lab, School of Computer Science, Shanghai Jiao Tong University, Shanghai, China; Suzhou Laboratory, Suzhou, China; Jiangsu Key Lab of Language Computing, Suzhou, China
X
Xin Chen
Suzhou Laboratory, Suzhou, China
Lu Chen
Lu Chen
School of Computer Science, Shanghai Jiao Tong University
Large Language ModelsDialogue SystemsAI for Science