🤖 AI Summary
This study addresses the issue in long-horizon search agents where accumulated contextual noise causes early errors to persist and become difficult to recover from. To this end, we propose an autonomous search framework that employs a Rubric-Answer-Verify tri-state management process, integrated with a Seal Memory tool to enable proactive context compression and verification. Furthermore, we introduce a training strategy that optimizes only the terminal segments responsible for context management, effectively mitigating the Seal Collapse instability commonly encountered during reinforcement learning. Experimental results demonstrate that our 35B-parameter model achieves a score of 72.83 on the BrowseComp benchmark, surpassing existing open-source systems and significantly outperforming baseline models across multi-domain evaluations.
📝 Abstract
Long-horizon information-seeking agents often accumulate noisy or misleading context, causing early mistakes to persist and making recovery increasingly difficult. We introduce an autonomous search harness in which the agent manages its own search process through three states: Rubric, Answer, and Verify. The agent first defines criteria for a valid answer, searches under these criteria, and then independently verifies the result before deciding whether to terminate or continue searching. It is further equipped with a Seal Memory tool that enables active context management. Training this behavior with reinforcement learning, however, can induce Seal Collapse, resulting in unstable training and preventing the agent from reliably learning when and how to use its memory tools. We solve this with a simple strategy that trains only the final segment after context management. Our 35B model achieves 72.83 on BrowseComp, outperforming comparable open-source systems, and consistently improves over the base model across BrowseComp-ZH, xbench, DeepSearchQA, WideSearch, financial investigation, and product search. Ablations show that autonomous compression outperforms automatic compaction and validate our RL design.