AREX: Towards a Recursively Self-Improving Agent for Deep Research

πŸ“… 2026-07-23
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This work addresses the inefficiency of multi-constraint answer discovery in deep research by proposing AREX, a recursively self-improving agent framework. AREX features an inner research loop that generates preliminary answers and an outer refinement loop that continuously verifies constraints, identifies unresolved claims, and initiates targeted investigations to iteratively enhance answer quality. A novel autonomous context updating mechanism compresses interaction history into a compact state representation, enabling long-term self-optimization without reliance on external models. The framework integrates in-agent training, long-horizon reinforcement learning, and dense reward signals to prioritize critical evidence acquisition and error correction. Evaluated on benchmarks including BrowseComp, WideSearch, DeepSearchQA, and HLE, AREX significantly outperforms same-scale baselines and achieves performance comparable to substantially larger models.
πŸ“ Abstract
Deep research requires agents to find answers that jointly satisfy multiple constraints. Discovering such answers is costly, whereas verifying a candidate can often be decomposed into tractable constraint-wise checks. This discovery--verification asymmetry suggests that a research agent should do more than simply search longer: it should recursively improve its current answer by verifying intermediate results and using the partially verified state to guide subsequent refinement. We introduce AREX, a family of Recursively Self-Improving (RSI) deep research agents. AREX alternates between an inner research loop that gathers evidence and constructs a provisional answer, and an outer self-improvement loop that audits the answer constraint-wise, identifies unresolved claims, and launches targeted follow-up research. To sustain RSI over long horizons, AREX learns an autonomous context-update tool that compresses growing interaction history into a compact improvement state preserving verified evidence and unresolved constraints, without relying on an external model. We train AREX on verified synthetic tasks and high-quality trajectories through agentic mid-training and long-horizon reinforcement learning. To mitigate sparse final rewards during long horizon learning, we emphasize key steps where decisive evidence is acquired or erroneous research directions are corrected. We instantiate a dense 4B model and a 122B-A10B Mixture-of-Experts model. Across BrowseComp, WideSearch, DeepSearchQA, Humanity's Last Exam (HLE), and other reasoning and tool-use benchmarks, AREX substantially outperforms comparable-scale baselines and remains competitive with models using substantially more activated parameters.
Problem

Research questions and friction points this paper is trying to address.

deep research
recursively self-improving agent
constraint satisfaction
answer verification
long-horizon reasoning
Innovation

Methods, ideas, or system contributions that make the work stand out.

Recursively Self-Improving
Deep Research Agent
Constraint-wise Verification
Autonomous Context Update
Long-horizon Reinforcement Learning
S
Shuqi Lu
Beijing Academy of Artificial Intelligence (BAAI)
Chaofan Li
Chaofan Li
Beijing University of Posts and Telecommunications
NLP
Kun Luo
Kun Luo
Zhejiang University
Z
Zhang Zhang
Beijing Academy of Artificial Intelligence (BAAI)
H
Hui Wang
Beijing Academy of Artificial Intelligence (BAAI)
H
Hongwang Xiao
Beijing Academy of Artificial Intelligence (BAAI)
Z
Zheng Liu
Beijing Academy of Artificial Intelligence (BAAI)
Lei Xiong
Lei Xiong
Stanford University
AI + BiologyComputational BiologyDeep LearningSingle Cell
J
Jiahao Wang
Beijing Academy of Artificial Intelligence (BAAI)
S
Sen Wang
Beijing Academy of Artificial Intelligence (BAAI)
X
Xiyan Jiang
Beijing Academy of Artificial Intelligence (BAAI)
W
Wanli Li
Beijing Academy of Artificial Intelligence (BAAI)
Yuyang Hu
Yuyang Hu
PhD student in ESE, Washington university in St. Louis
Computational ImagingInverse ProblemsMachine LearningSignal Processing
Hongjin Qian
Hongjin Qian
Peking University
LLMIRNLP
B
Bingyu Yan
Beijing Academy of Artificial Intelligence (BAAI)
Ziyi Xia
Ziyi Xia
University of British Columbia
Computer GraphicsVRMachine Learning
Yingxia Shao
Yingxia Shao
SCS, BUPT
Large-scale Graph AnalysisGraph Data ManagementGraph Learning
K
Kang Liu
Beijing Academy of Artificial Intelligence (BAAI)
Zhicheng Dou
Zhicheng Dou
Renmin University of China
Information RetrievalRetrieval Augmented GenerationLarge Language ModelsGenerative IR
D
Di He
Beijing Academy of Artificial Intelligence (BAAI)
Chaozhuo Li
Chaozhuo Li
Microsoft Research Aisa
Qiwei Ye
Qiwei Ye
Beijing Academy of Artificial Intelligence
Scientific AIAI for ScienceFoundation Model
Zhongyuan Wang
Zhongyuan Wang
BAAI
Knowledge MiningDatabaseNLPText Understanding
Z
Zheng Liu
Beijing Academy of Artificial Intelligence (BAAI)