Argus: A General-Purpose Agentic Runtime for Long-Horizon Reasoning

📅 2026-08-05
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the challenges in long-horizon reasoning where agents often struggle to balance policy persistence with adaptive path adjustment and frequently conflate user intent, constraints, and validation criteria. To overcome these issues, the paper proposes Argus—a persistent self-evolving agent system that operates with fixed model weights and leverages structured collaboration among Manager, Planner, Engineer, and Reviewer roles. Key innovations include decoupling user intent from operational objectives, verification-gated memory accumulation, route rollback and recovery mechanisms, and a review loop driven by task-native validators. Evaluated on SWE-Bench Pro, Argus achieves a 78% resolution rate—outperforming GitHub Copilot’s 59%—while reducing input tokens by 21% and runtime by 15% on mature tasks. On AARRI-Bench, it attains 76.8%, leads by 28.0 points in mathematical synthesis, and has contributed to the RWKV6 kernel merge and completion of 254 research tasks.
📝 Abstract
Long-horizon reasoning requires an agentic runtime that can persist when evidence supports its current approach and pivot when measurements reveal failure, hidden constraints, or a misspecified objective. We present Argus, a persistent, self-evolving runtime in which Manager, Planner, Engineer, and Reviewer execute bounded missions over durable project state. Argus separates stable user intent from operational objectives, constraints, and verification criteria, and admits memories, skills, procedures, verifiers, routing decisions, and rejected routes only after role-owned review and, when available, task-native verification. Model weights remain fixed; self-evolution occurs through persistent runtime state and control policy, with autonomous execution between operator-owned escalation points. Across seven GPT-5.5 benchmark arenas, Argus achieves about 78% on SWE-Bench Pro versus 59% for Direct Copilot while using 1.41 times the aggregate tokens. After verification-gated self-evolution, mature SWE-Bench waves use 21% fewer solve-input tokens and 15% less active workflow time per task than startup waves, while recording 34 verifier recoveries and 22 strict review-loop rescues. Argus also reaches 76.8% on AARRI-Bench and a 28.0-point gap on mathematical data synthesis, with competitive GPU-kernel and language-model-training results. Beyond benchmarks, an optimized RWKV6 kernel was merged upstream; a multi-day mathematics campaign retained falsified routes and proof-backed frontier updates; and six paper pipelines completed 254 missions with 16 stage rollbacks. These results show that a fixed-weight, self-evolving harness can revise, recover, and accumulate verified approaches while producing structured trajectories for future supervised and reinforcement learning.
Problem

Research questions and friction points this paper is trying to address.

long-horizon reasoning
agentic runtime
self-evolution
persistent state
verification
Innovation

Methods, ideas, or system contributions that make the work stand out.

agentic runtime
long-horizon reasoning
self-evolving system
fixed-weight LLM
role-based verification
🔎 Similar Papers
B
Boxiu Li
Microsoft
Z
Zimo Wen
Shanghai Jiao Tong University
Y
Yijia Fan
Microsoft
J
Junxiang Lei
Fudan University
S
Sufeng Guo
Nanjing University
J
Jiaao Wu
Tsinghua University
Ruize Tang
Ruize Tang
Ph.D. student, Nanjing University
concurrent systemsdistributed systemsmodel checking
Mukai Li
Mukai Li
The University of Hong Kong
natural language processing
Y
Yifei Shen
Microsoft
Xiaoyu Chen
Xiaoyu Chen
Shanghai University
Cultural heritage informaticsHuman information behaviourSocial informaticsSocial impacts of AI
W
Wanbo Zhang
Fudan University
R
Runjing Gu
Shanghai Jiao Tong University
Yifei Gao
Yifei Gao
Beijing Jiaotong University
MLLM Reasoning & GUI Agent
Yuheng Wu
Yuheng Wu
KAIST
Efficient AIEmbodied Intelligent SystemAutonomous Driving
X
Xuyao Huang
Shanghai Jiao Tong University
Z
Zelong Zhao
Independent Researcher
J
Jiachen Zhang
Independent Researcher
S
Shibo Hu
Peking University
H
Hangxi Guo
The Chinese University of Hong Kong, Shenzhen
Yilin Chen
Yilin Chen
Peking University, Shenzhen Graduate School
Atmospheric Environment Modeling
Y
Yuzhe Zhang
Shanghai Jiao Tong University
Fan Yang
Fan Yang
Microsoft Research
Systems
Chuan Wen
Chuan Wen
Shanghai Jiao Tong University
RoboticsMachine LearningComputer Vision
Xian Zhang
Xian Zhang
Microsoft Research
AI ReasoningAI for MathNeuro-symbolic
Xuanhe Zhou
Xuanhe Zhou
Assistant Professor, Shanghai Jiao Tong University
Data ManagementArtificial Intelligence