OSReward: Instituting Standardized Evaluation for Cross-Platform Computer-Use Reward Models

๐Ÿ“… 2026-07-30
๐Ÿ“ˆ Citations: 0
โœจ Influential: 0
๐Ÿ“„ PDF
๐Ÿค– AI Summary
This work addresses the lack of reliability validation in existing vision-language models (VLMs) when deployed as evaluators of agent trajectories, which hinders large-scale assessment and reinforcement learning. Through the first systematic evaluation, we identify a prevalent leniency bias in mainstream VLMsโ€”frequently misclassifying failed trajectories as successful. To mitigate this, we introduce the OSReward benchmark, featuring the OS-Shepherd-100K corpus: a large-scale, cross-platform trajectory dataset with human-annotated ground-truth labels and reasoning traces. We further train open-source reward models, OS-Shepherd (9B/35B), which match the performance of proprietary counterparts while reducing inference costs by 30โ€“60%. These models support a comprehensive evaluation framework encompassing OSReward, OSReward-Hard, and OSReward-Multi.
๐Ÿ“ Abstract
Computer-using agents (CUAs) are advancing rapidly across the digital world. A CUA trajectory records the agent's actions, states, and reasoning. Verifying whether it fulfilled the task instruction is central to CUA evaluation, data curation, and reinforcement learning. Neither human-written verifiers nor human annotators can provide such verification at scale, so the field increasingly turns to vision-language models (VLMs) as judges of CUA trajectories. But a fundamental question has long gone unexamined: are these VLM judges reliable enough? To study it systematically, we introduce OSReward, a realistic, high-quality benchmark that evaluates VLM judges on CUA trajectories. The trajectories come from diverse agent backbones executing human-verified instructions across platforms, then rigorously labeled with ground-truth verdicts through multi-stage human annotation. Building on it, we derive OSReward-Hard, a challenge set concentrating genuinely hard cases, and OSReward-Multi for fine-grained efficiency and alignment scoring. The most comprehensive evaluation of VLM judges to date finds even state-of-the-art models fall short of an ideal judge, sharing a systematic leniency bias that mislabels failed runs as successes. The few reliable enough to trust are too expensive to run at scale, while affordable open models trail far behind. To close this gap, we construct and release OS-Shepherd-100K, an open corpus of reasoning-annotated trajectory judgments for the CUA community. On it, we train OS-Shepherd (9B and 35B), open reward models that supply low-cost, stable, and reliable reward signals, matching commercial judges at 30-60% lower cost than the frontier. Extensive analyses further inform the design of reliable CUA reward at scale. Our code, benchmark, dataset, and model checkpoints are available at https://os-copilot.github.io/OSReward-Home/.
Problem

Research questions and friction points this paper is trying to address.

computer-using agents
reward models
vision-language models
evaluation benchmark
reliability
Innovation

Methods, ideas, or system contributions that make the work stand out.

OSReward
computer-using agents
vision-language models
reward modeling
benchmark
๐Ÿ”Ž Similar Papers
No similar papers found.
Qiushi Sun
Qiushi Sun
The University of Hong Kong, National University of Singapore
Natural Language ProcessingAgentsCode Intelligence
Kanzhi Cheng
Kanzhi Cheng
Nanjing University Ph.D Student
Vision-Language ModelsAI AgentsImage Captioning
Y
Yian Wang
The University of Hong Kong, National University of Singapore
B
Bowen Yang
University of Science and Technology of China
Hang Yan
Hang Yan
Xi'an Jiaotong University
LLM reasoningAgentKnowledge Graph
Liheng Chen
Liheng Chen
Undergraduate of Computer Science, the University of Hong Kong
Natural Language ProcessingMachine LearningAutomatic Speech Recognition
Fangzhi Xu
Fangzhi Xu
Xi'an Jiaotong University | Nanyang Technological University
Large Language ModelsSelf-TrainingReasoningGUI Agents
Zichen Ding
Zichen Ding
Shanghai AI Laboratory
Computer-Use AgentAI AgentsLarge Language Models
N
Nuo Chen
Nanjing University
J
Jialin Cao
Nanjing University
X
Xingdong Gong
Nanjing University
Zehao Li
Zehao Li
Peking University
Operations researchStochastic approximation
K
Kaiming Jin
National University of Singapore
X
Xinfeng Yuan
Fudan University
Z
Zhoumianze Liu
Fudan University
J
Jingyang Gong
The University of Hong Kong
Z
Zhangyue Yin
Fudan University
Jiahui Gao
Jiahui Gao
The University of Hong Kong
Synthetic Data GenerationMultimodal ModelNLP
Z
Zhiyong Wu
The University of Hong Kong
Tianbao Xie
Tianbao Xie
University of Hong Kong
Artificial IntelligenceDeep LearningNatural Language Processing
Jianbing Zhang
Jianbing Zhang
Associate Professor, Nanjing University
pre-training modelmulti-modalimage captioningnatural language processingdata mining
Ben Kao
Ben Kao
The University of Hong Kong
Database
Lingpeng Kong
Lingpeng Kong
Google DeepMind, The University of Hong Kong
Natural Language ProcessingMachine Learning