🤖 AI Summary
This study addresses the high cost of expert data and the lack of factual grounding and internal consistency in synthetic tasks for training LLM agents. We propose an evidence-based method for occupational scenario construction and execution-guided consistency verification. By retrieving authentic occupational knowledge from O*NET to synthesize contexts and render reference deliverables, combined with execution feedback for defect attribution and iterative repair, our approach automatically generates high-quality, executable training data. Fine-tuned on merely 20,000 samples, the resulting Fx-Work-35B model surpasses peer models across multiple benchmarks and even outperforms frontier large language models on certain metrics. This work demonstrates a low-cost, high-fidelity paradigm for the automated generation of agent training data.
📝 Abstract
The ability of Large Language Model (LLM) agents to complete daily and professional work is receiving increasing attention. Training such agents requires realistic work scenarios. Expert-authored occupational work is costly and slow to produce, while unconstrained synthesis often yields tasks with weak factual grounding or internally inconsistent requirements. To bridge this gap, we introduce WorkGenesis, a framework that constructs executable occupational work from real-world artifacts through two core technical innovations: (1) Evidence-Based Work Construction, which grounds each unit of work in real-world evidence by retrieving public files guided by O*NET occupational knowledge and synthesizing the surrounding context, companion materials, work request, and itemwise rubric around them; and (2) Execution-Guided Consistency Verification, which renders a reference deliverable inside the constructed work, attributes every unsatisfied rubric item to the agent, the task, or the rubric, and uses task and rubric defects as feedback to iteratively repair the work until it passes the audit. Experimental results demonstrate that Fx-Work-35B, trained with simple supervised fine-tuning (SFT) on only 20K units of work synthesized by WorkGenesis, achieves the highest scores among all comparable-scale baselines on the five reported metrics across GDPvalAA-v2, APEX-Agents-AA, and JobBench (31.00 versus 24.79 average score), and even surpasses frontier models such as the 1.6T DeepSeek-V4-Pro-Preview. These results show that WorkGenesis provides scalable training data for working agents.