Scaling Verifiable Environments for Long-horizon Work Agents

📅 2026-10-04
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the bottlenecks of costly manual environment construction and the lack of authenticity and verifiability in synthetic approaches for training long-horizon work agents by proposing WorkForge. This framework extracts resource and factual anchors from expert workflows to automatically generate work environments equipped with authentic files and dual procedural-semantic verifiers, ensuring that verification traces back to workspace evidence and enabling automated scaling of high-fidelity environments. The project constructs 16,700 verifiable environments spanning 40 domains and 60 file types. Experimental results demonstrate that this framework significantly enhances Qwen model performance (GDPVal: 45.5→73.6; APEX: 5.0→21.3), confirming effective scalability in both data volume and interaction horizon.
📝 Abstract
Work agents operate over digital artifacts to execute professional knowledge-intensive work, requiring training environments that support long-horizon interaction and trustworthy verification. However, hand-crafted environments incur prohibitive engineering overhead that prevents environment scaling, whereas synthesis methods sacrifice workspace complexity, realism, or grounded verifiability. To bridge this gap, we introduce WorkForge, a scalable synthesis framework for constructing verifiable work-agent environments from real-world resources. Starting from expert workflows, WorkForge first identifies the resources, decisions, and deliverables required by each workflow. It then retrieves relevant real-world files and organizes them into a workspace. WorkForge inspects the workspace to extract concrete, checkable facts about its content. These factual anchors fix which task types the workspace can support and how their outcomes can be verified. Therefore, WorkForge derives each task's instructions, solution plan, and complementary programmatic and semantic verifiers directly from these factual anchors, keeping verification traceable to observable workspace evidence. Furthermore, we construct 16.7K verifiable environments across 40 professional domains, with workspaces collectively covering 60 file types. Post-training Qwen3.5-35B-A3B-Base improves GDPVal from 45.5 to 73.6 and APEX Score from 5.0 to 21.3, while enabling Qwen3.5-27B to achieve highly competitive performance and outperform strong competitors. Our analyses confirm the efficacy of the proposed method and reveal consistent scaling behaviors across both data volume and interaction horizons.
Problem

Research questions and friction points this paper is trying to address.

Work Agents
Verifiable Environments
Long-horizon Interaction
Environment Scaling
Task Synthesis
Innovation

Methods, ideas, or system contributions that make the work stand out.

Verifiable Environments
Work Agents
Scalable Synthesis
Factual Anchors
Long-horizon Interaction