FLARE: A Full-Lifecycle Dense Supervision Paradigm for Long-Horizon Coding Agents via Generative Reward Model

📅 2026-09-20
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
为解决长周期软件工程中稀疏奖励导致的问题,提出FLARE方法,通过生成性奖励模型提供密集监督,优化整个生命周期中的代理性能。
📝 Abstract
While test-time scaling enhances Large Language Model (LLM) agents in long-horizon software engineering (SWE), sparse binary rewards (Pass/Fail) create a severe credit assignment crisis and waste failed exploratory trajectories. Current trajectory optimization and scaling methods are costly and structurally limited, relying on heuristic state reuse without causal diagnosis or delayed scalar scoring without actionable online guidance. We propose FLARE (Full-Lifecycle Alignment and Reward Engine), a novel dense supervision paradigm driven by a lightweight Generative Reward Model (GRM). First, RADAR, an offline causal-aware diagnostic framework, extracts high-fidelity, hindsight-free supervision through causal-chain backtracking to distill a GRM providing real-time, step-level risk feedback. Second, FLARE uses this GRM to continuously optimize the agent across its entire lifecycle. During inference, FLARE acts as an Active Scaffold, autonomously intercepting high-risk generation steps for localized breakpoint re-execution, drastically reducing compute overhead. During post-training, the GRM's structured signals serve as process-supervised reranking scores for Supervised Fine-Tuning (SFT) and step-level dense rewards for Reinforcement Learning (RL), mitigating policy collapse in sparse environments. Extensive evaluations show that FLARE establishes a new Pareto frontier across the agent lifecycle: FLARE (N=1) outperforms Global Rollout (N=5) with a 5x reduction in token consumption. Extending FLARE to training overcomes the sparse reward problem in long-horizon interactive tasks, delivering relative performance gains of 19.13% in SFT through process-aware data curation and a consistent 9.19% improvement in RL.
Problem

Research questions and friction points this paper is trying to address.

Large Language Model
long-horizon software engineering
sparse binary rewards
credit assignment crisis
trajectory optimization
Innovation

Methods, ideas, or system contributions that make the work stand out.

Generative Reward Model
Dense Supervision
Causal-chain Backtracking
Active Scaffold
Process-aware Data Curation
🔎 Similar Papers
No similar papers found.
J
Jingxuan Xu
Independent Researcher
G
Gang Wu
Independent Researcher
Y
Yanan Wu
Nanjing University
Yutao Mou
Yutao Mou
Peking University
AI SafetyLLM Alignment
S
Songwei Yu
Independent Researcher
T
Tianzhuang He
Nanjing University
Z
Zhengshuo Gong
Beijing University of Posts and Telecommunications
Z
Zhao Liu
Independent Researcher
Z
Zihang Xu
Independent Researcher
W
Wenqiang Zhu
Independent Researcher
X
Xinping Lei
Nanjing University
Weihao Li
Weihao Li
Research Fellow, Australian National University
Computer VisionMachine Learning
Y
Yuhui Bai
Independent Researcher
Z
Zhongqiu Wang
Independent Researcher
Y
Yan Wu
Independent Researcher
A
Ariel Deng
Independent Researcher