OSWorld-Science: A Benchmark of Computer Use Agents for Learning and Using Scientific Software

📅 2026-09-30
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the lack of systematic evaluation regarding the capabilities of Vision-Language Model (VLM) agents in complex scientific software workflows. To this end, we construct a benchmark encompassing multi-domain scientific tasks through expert-informed and iterative co-design methodologies. Furthermore, we propose an artifact-based evaluation mechanism alongside a specialized agent framework that facilitates fine-grained partial scoring and execution trajectory analysis. Our experiments reveal that current state-of-the-art models still encounter substantial challenges when handling scientific tasks. By establishing critical evaluation dimensions, this work provides foundational guidance for the future development of computer-use agents within research contexts.
📝 Abstract
Scientific software presents a demanding test for computer-using agents based on visual language models (VLMs): completing a research workflow requires interpreting specialized interfaces, manipulating scientific objects, and producing verifiable results. We thus introduce OSWorld-Science, a benchmark and evaluation environment that combines scientifically meaningful tasks, artifact-based evaluation, and an efficient agent harness for studying computer use in the scientific domain. The benchmark contains 12 VLMs and 146 high-quality tasks across several scientific domains and software configurations, covering workflows such as molecular drawing and retrosynthesis, pathology image analysis, statistical computing, and physical simulation. Tasks are developed through expert proposals and iterative human--AI co-design, with selection guided by scientific value and difficulty. Task-specific execution-based evaluators inspect application states and generated artifacts, including molecular structures, segmentation masks, plots, and numerical results, and award partial credit for incomplete outcomes. Our special harness integrates model adapters, interaction-loop control, and trajectory logging to support comparisons of models and interaction strategies. Our results show that current state-of-the-art VLMs with a strong harness still face challenges in addressing key questions in the scientific domains. We also analyze the benchmarking results across multi-linguistics, reasoning efforts, context length and other factors and derive several important conclusions and directions to assist future development. Overall, we provide an integrated framework connecting expert-defined scientific goals to verifiable software outcomes, enabling systematic evaluation of both agent capabilities and harness design in scientific workflows.
Problem

Research questions and friction points this paper is trying to address.

Computer Use Agents
Visual Language Models
Scientific Software
Benchmark
Research Workflows
Innovation

Methods, ideas, or system contributions that make the work stand out.

Scientific Software Benchmark
Visual Language Models
Artifact-based Evaluation
Computer Use Agents
Agent Harness
🔎 Similar Papers
No similar papers found.
D
Dingyuan Dai
University of California, Los Angeles
Heli Qi
Heli Qi
Waseda University, RIKEN
Multi-Modal Learning
L
Lei Liu
Yale University
Y
Yinxi Li
University of Waterloo
B
Baiding Chen
Carnegie Mellon University
Z
Zijun Dou
Tsinghua University
Qingcheng Zeng
Qingcheng Zeng
PhD Student in NLP, Northwestern University
Computational Social ScienceNLPComputational Linguistics
Qi Kang
Qi Kang
同济大学
计算智能、人工智能、机器学习
O
Oliver Sun
University of California, Berkeley
E
Eric Wang
University of California, Berkeley
B
Bo Zhou
University of Illinois Chicago
Haixin Wang
Haixin Wang
UCLA; Peking University
AI for ScienceLarge Language ModelsMulti-modal LLMs
Y
Yufan Du
University of California, Los Angeles
S
Shi Bo
Boston University
R
Ruihan Lin
University of Waterloo
M
Mengqi Yuan
The University of Hong Kong
Dunjie Lu
Dunjie Lu
Bachelor of Computer Science, Sun Yat-sen University
AIMLLLMVLMAgent
Steven Dillmann
Steven Dillmann
Stanford University, University of Cambridge
AI for ScienceMachine LearningData Driven DiscoveryComputational Mathematics
Yiming Shi
Yiming Shi
University of Electronic Science and Technology of China
Efficient AIParameter Efficient Fine TuningDiffusionMultimodal
T
Tina Su
Yale University
A
Amy Xin
Tsinghua University
M
Minghao Liu
The University of Tokyo
X
Xi Wang
New York University
X
Xu Huang
University of California, Berkeley
G
Ge Zhang
TokenWave.AI