🤖 AI Summary
This study addresses the lack of authentic file grounding and verifiability in training data for workplace agents by proposing a data synthesis and fine-tuning method based on an evidence graph framework. The approach introduces a novel evidence graph mechanism that constructs realistic file workspaces from occupational seeds, dual-anchoring task requirements and verification criteria to authentic files for automated generation of tasks and evaluation standards. Following execution-based testing and repair, high-quality trajectories are collected and combined with rejection sampling to fine-tune the Qwen3.6 model. Experimental results demonstrate that this method yields substantial improvements of 65.7 points on GDPVal, 7.7 points on Workspace-Bench-Lite, and 13.7 points on SpreadsheetBench II, significantly enhancing agent capabilities in real-world workplace scenarios.
📝 Abstract
Working agents need to read diverse files, coordinate tools, and produce deliverables. Training such agents requires tasks built on many real files with verifiable results, but few pipelines exist to synthesize this kind of data. Existing pipelines either generate files with models, which lack realism and diversity, or build tasks on real files without task-specific verifiers, leaving result quality unchecked. We introduce GraphForge, an evidence-graph based framework that grounds both the task and its verification in real files. Starting from occupation-grounded seeds for controlled diversity, GraphForge assembles a workspace of real files for each seed and builds an evidence graph over their relations. Since the task statement and rubrics are both derived from this graph, task requirements are backed by the workspace files and each criterion is anchored to the files needed to verify it. An initial rollout further tests executability, and a revision agent repairs the task and rubrics against the original files before trajectories are collected. Fine-tuning Qwen3.6-27B on 2,169 GraphForge trajectories brings GDPVal to 1445.7 (+65.7) under OpenHands, and Workspace-Bench-Lite and SpreadsheetBench II to 63.7 (+7.7) and 24.0 (+13.7) under Claude Code. Rejection fine-tuning on the SFT model's own rollouts, with candidates selected by the evidence-anchored rubrics, yields further improvements on all three benchmarks, suggesting that the rubrics provide a useful selection signal. The data and models are available.