DAGent: Evaluate-then-Grow Planning for Deep Research Agents

📅 2026-09-30
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the brittleness and computational redundancy inherent in static planning for deep research agents by proposing DAGent, a novel framework that introduces a "post-evaluation incremental growth" mechanism to dynamically expand a directed acyclic task graph based on confidence estimates. By integrating hierarchical context management with topology-conditioned reinforcement learning (DAGRPO), optimized via structure-compliant regularization, the framework enhances multi-agent collaborative execution. Experimental results demonstrate that DAGent surpasses state-of-the-art open-source baselines across multiple benchmarks, significantly improving accuracy while substantially reducing both token consumption and reasoning steps.
📝 Abstract
Deep research tasks require agents to navigate large knowledge spaces, synthesize evidence across many sources, and adapt their plans as findings emerge. Directed acyclic graph (DAG)-based multi-agent systems suit this setting because they support parallel execution and isolate each sub-task within a focused dependency context. Yet existing DAG-based agents instantiate a task-level plan before execution and repair the graph only after failures or missing evidence are observed. This Plan-then-Patch strategy is brittle for deep research: the system commits most strongly when its evidence is weakest, and later revisions waste computation on branches that should not have been planned. We propose DAGent, a DAG-based multi-agent framework with Evaluate-then-Grow incremental planning: an Orchestrator grows the task graph one batch at a time, conditioning each expansion on confidence and uncertainty signals from completed nodes. A hierarchical context layer propagates compact QueryDocs by default while preserving full execution traces for on-demand recall. The recorded DAG topology admits structural RL signals that outcome-only recipes cannot define; DAGRPO, a GRPO adaptation, injects topology-conditioned credit on Executor rollouts and a structural compliance regularization on Orchestrator plans. Across BrowseComp-Plus, GAIA, and xbench-DeepSearch, DAGent surpasses the strongest open-source baseline by 5.3 / 5.8 / 2.0 points at the Qwen3-235B-A22B scale, and the lead replicates across four open-source backbones and extends to GPT-5 at 327K context. At the Qwen3-8B scale, DAGRPO improves over a same-budget outcome-only GRPO baseline by 3.0 average Pass@1 points. A same-architecture comparison shows that evidence-conditioned planning reaches higher accuracy at lower per-task token, tool-call, and step footprints than its Plan-then-Patch counterpart. Code: https://github.com/hanwenliu6825/DAGent
Problem

Research questions and friction points this paper is trying to address.

Deep research agents
DAG-based planning
Multi-agent systems
Plan-then-Patch brittleness
Innovation

Methods, ideas, or system contributions that make the work stand out.

Evaluate-then-Grow
DAG-based multi-agent
hierarchical context layer
DAGRPO
topology-conditioned credit
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
H
Hanwen Liu
New York University; New York University Shanghai
Y
Yuanfu Sun
New York University; New York University Shanghai
Qiaoyu Tan
Qiaoyu Tan
Assistant Professor of Computer Science, New York University (Shanghai)
Generative AIMultimodal AlignmentFoundation ModelGraph LearningAI4Science