Institution profile

Shanghai Innovation Institution

Academic institutionasia · cn
Research library3linked papers
Opportunities0open roles
Selected work

Representative Papers

MatToolBench: Benchmarking Multimodal Agents in Real-World Materials Science Workflows

Sep 29, 2026

This study addresses the limited capabilities of multimodal agents in materials science software, stemming from domain data scarcity and complex interfaces. We construct the first real-world GUI agent benchmark tailored to this domain, integrating a fine-grained scoring mechanism with a multi-level evaluation pipeline to systematically assess agent performance across graphical interaction, OriginPro script execution, and database querying. Evaluations conducted within a Windows 11 virtual machine alongside human–machine consistency experiments demonstrate that general-purpose model capabilities transfer poorly, with the best-performing model achieving only a 25% success rate on GUI tasks. By bridging this critical gap, this work provides an essential testbed for diagnosing and enhancing the operational proficiency of AI agents within specialized scientific software environments.

0 citationsRead paper

LigBench: A Unified and Human-Aligned Benchmark for LLM-based Research Idea Generation

Aug 13, 2026

This work addresses the lack of a unified, objective, and human-aligned evaluation standard for scientific idea generation by large language models. To this end, the authors propose LigBench—an automated, fine-grained evaluation benchmark—and introduce the PAIR-IQ dataset to train pairwise idea judgment models, enabling consistent and interpretable assessment across diverse generation distributions. The framework pioneers a shift from subjective scoring to structured comparative learning, substantially improving alignment with expert judgments. Experimental results demonstrate that LigBench outperforms existing methods in evaluation stability and interpretability, with models trained on PAIR-IQ achieving superior performance in ranking accuracy and robustness.

0 citationsRead paper
Recent publications

Latest Papers

MatToolBench: Benchmarking Multimodal Agents in Real-World Materials Science Workflows

Sep 29, 2026

This study addresses the limited capabilities of multimodal agents in materials science software, stemming from domain data scarcity and complex interfaces. We construct the first real-world GUI agent benchmark tailored to this domain, integrating a fine-grained scoring mechanism with a multi-level evaluation pipeline to systematically assess agent performance across graphical interaction, OriginPro script execution, and database querying. Evaluations conducted within a Windows 11 virtual machine alongside human–machine consistency experiments demonstrate that general-purpose model capabilities transfer poorly, with the best-performing model achieving only a 25% success rate on GUI tasks. By bridging this critical gap, this work provides an essential testbed for diagnosing and enhancing the operational proficiency of AI agents within specialized scientific software environments.

0 citationsRead paper

LigBench: A Unified and Human-Aligned Benchmark for LLM-based Research Idea Generation

Aug 13, 2026

This work addresses the lack of a unified, objective, and human-aligned evaluation standard for scientific idea generation by large language models. To this end, the authors propose LigBench—an automated, fine-grained evaluation benchmark—and introduce the PAIR-IQ dataset to train pairwise idea judgment models, enabling consistent and interpretable assessment across diverse generation distributions. The framework pioneers a shift from subjective scoring to structured comparative learning, substantially improving alignment with expert judgments. Experimental results demonstrate that LigBench outperforms existing methods in evaluation stability and interpretability, with models trained on PAIR-IQ achieving superior performance in ranking accuracy and robustness.

0 citationsRead paper