SearchArt: Training Long-Horizon Search Agent with Scalable Synthetic and Verified Task

📅 2026-07-25
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the challenges of scarce scalable high-quality data and the difficulty of evaluating intermediate reasoning behaviors in long-horizon search tasks. The authors propose a scalable framework that synthesizes large-scale, complex search task data through automated evidence graph generation and trajectory verification mechanisms. Coupled with a multi-stage post-training pipeline—comprising supervised fine-tuning and reinforcement learning—the framework trains search agents endowed with adaptive planning and iterative reasoning capabilities. Remarkably, using only a 27B-parameter model, the approach achieves scores of 74.39, 70.06, and 52.55 on BrowseComp-ZH, BrowseComp, and DeepResearch-Bench, respectively, matching or surpassing the performance of leading closed-source systems.
📝 Abstract
Recent advances in large language models (LLMs) have enabled search agents to autonomously tackle complex tasks across extended search and reasoning horizons. However, training effective search agents remains challenging due to the lack of scalable and long-horizon tasks, and the difficulty of evaluating and correcting intermediate reasoning and tool-use behaviors. We introduce SearchArt, a scalable framework for training long-horizon search agents through verification-driven task synthesis and a multi-stage post-training pipeline. SearchArt constructs large-scale datasets for complex search-, research- and user-oriented tasks by synthesizing diverse information-seeking QA pairs and corresponding search trajectories from web documents and automatically generated evidence graphs. To ensure the reliability of the synthesized data, we design a verification pipeline that jointly evaluates QA consistency, trajectory quality, and the relevance of retrieved evidence. The verified trajectories are subsequently used in a multi-stage training process comprising supervised fine-tuning and reinforcement learning-based policy optimization. Search agents trained with SearchArt exhibit adaptive search planning, iterative evidence aggregation, and complex reasoning over extended interaction horizons. Experimental results demonstrate that, with only (Qwen3.5-) 27B parameters, SearchArt scores 74.39 on BrowseComp-ZH, 70.06 on BrowseComp, and 52.55 on Deepresearch-bench, matching or surpassing frontier closed-source agents on both deepsearch and deepresearch benchmarks.
Problem

Research questions and friction points this paper is trying to address.

long-horizon search
search agent training
scalable task synthesis
reasoning evaluation
tool-use behavior
Innovation

Methods, ideas, or system contributions that make the work stand out.

long-horizon search agent
verification-driven synthesis
evidence graph
multi-stage training
scalable task generation
L
Lang Mei
Huawei Cloud
Xiaohan Yu
Xiaohan Yu
Macquarie University
computer visionsmart farmingultra-fine-grained visual categorization
C
Chong Chen
Huawei Cloud
L
Liyan Liu
Huawei Cloud
X
Xiangnan Chen
Huawei Cloud
J
Jinchao Ma
Huawei Cloud
Chao Feng
Chao Feng
University of Zurich
networkmachine learningcybersecurity
L
Li Huang
Huawei Cloud
S
Siyu Mo
Huawei Cloud
S
Sichen Kang
Huawei Cloud
Yunkun Xu
Yunkun Xu
Huawei; Zhejiang University
LLMMachine LearningReinforcement LearningRobot LearningIndustrial Intelligence
Z
Zhihan Yang
Huawei Cloud
Z
Zhujun Xue
Huawei Cloud
J
Jingren Zhang
Huawei Cloud
Q
Qing He
Huawei Cloud
Y
Yingdi Huang
Huawei Cloud
Hao Jiang
Hao Jiang
Huawei Cloud
Large Language ModelMultimediaNatural Language ProcessingInformation Retrieval
Z
Ziao Ma
Huawei Cloud
Z
Zewei Pan
Huawei Cloud
M
Minhao Sun
Huawei Cloud
Zhuo Tao
Zhuo Tao
Institute of Computing Technology, Chinese Academy of Sciences
multi-modal learningvision-and-language
J
Jinzhao Xiao
Huawei Cloud
G
Gangtao Xin
Huawei Cloud
H
Huanyao Zhang
Huawei Cloud
W
Wenjian Zhang
Huawei Cloud