Institution profile

Suzhou Laboratory

Academic institutionasia · cn
Research library1linked papers
Opportunities0open roles
Selected work

Representative Papers

MatToolBench: Benchmarking Multimodal Agents in Real-World Materials Science Workflows

Sep 29, 2026

This study addresses the limited capabilities of multimodal agents in materials science software, stemming from domain data scarcity and complex interfaces. We construct the first real-world GUI agent benchmark tailored to this domain, integrating a fine-grained scoring mechanism with a multi-level evaluation pipeline to systematically assess agent performance across graphical interaction, OriginPro script execution, and database querying. Evaluations conducted within a Windows 11 virtual machine alongside human–machine consistency experiments demonstrate that general-purpose model capabilities transfer poorly, with the best-performing model achieving only a 25% success rate on GUI tasks. By bridging this critical gap, this work provides an essential testbed for diagnosing and enhancing the operational proficiency of AI agents within specialized scientific software environments.

0 citationsRead paper
Recent publications

Latest Papers

MatToolBench: Benchmarking Multimodal Agents in Real-World Materials Science Workflows

Sep 29, 2026

This study addresses the limited capabilities of multimodal agents in materials science software, stemming from domain data scarcity and complex interfaces. We construct the first real-world GUI agent benchmark tailored to this domain, integrating a fine-grained scoring mechanism with a multi-level evaluation pipeline to systematically assess agent performance across graphical interaction, OriginPro script execution, and database querying. Evaluations conducted within a Windows 11 virtual machine alongside human–machine consistency experiments demonstrate that general-purpose model capabilities transfer poorly, with the best-performing model achieving only a 25% success rate on GUI tasks. By bridging this critical gap, this work provides an essential testbed for diagnosing and enhancing the operational proficiency of AI agents within specialized scientific software environments.

0 citationsRead paper