CompoWorld: Compositional Environment Scaling for General Agents

📅 2026-09-27
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the limitation of existing agent environment generation, which is confined to single services and fails to meet real-world workflow demands requiring cross-service collaboration. We propose a compositional environment extension framework that constructs task spaces by reusing multiple services and leverages dependency graphs to generate and verify cross-service interaction trajectories. Furthermore, we design a completion-oriented reward mechanism that integrates code agents, world models, and reinforcement learning to enable scalable data generation and general-purpose agent training. Experimental results demonstrate that our approach achieves an average improvement of 9.17 points across eight benchmarks, outperforming both frontier models and similarly scaled 35B-parameter agents, thereby significantly enhancing capabilities in handling cross-service tasks.
📝 Abstract
Automatically generated environments provide a scalable source of interaction data for training general agents. However, existing approaches mainly generate tasks within a single environment, while real-world workflows require agents to connect information and actions across multiple services. We introduce Compositional Environment Scaling (\textbf{CompoWorld}), which expands the task space by composing a finite library of reusable services. Coding agents turn tool specifications into verified services with typed states and shared interfaces, while a world model handles tools that cannot be reliably implemented. A random-walk procedure connects services through dependency graphs, enabling the generation and verification of tasks that require information to flow across services. Verified trajectories support supervised fine-tuning (SFT), while our Completion-Focused Rubric Reward guides reinforcement learning (RL) toward full task completion by emphasizing criteria with lower pass rates within each rollout group. We construct 448 services exposing 10,130 tools and use 3K SFT trajectories and 1K RL tasks to train Qwen3.6-35B-A3B. Experimental results show that CompoWorld improves on its backbone by 9.17 points on average across eight benchmarks. On AutomationBench, it surpasses frontier models such as Claude Opus 4.6 and leads all compared agent-specialized 35B-A3B models.
Problem

Research questions and friction points this paper is trying to address.

compositional environment scaling
general agents
cross-service workflows
automatically generated environments
Innovation

Methods, ideas, or system contributions that make the work stand out.

Compositional Environment Scaling
Cross-service Task Generation
World Model
Completion-Focused Rubric Reward
General Agents