My FAULT: Self-Diagnosis as Credit Assignment in Self-Evolving Agentic Reinforcement Learning

📅 2026-10-01
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the credit assignment challenge arising from terminal rewards in long-horizon agent reinforcement learning tasks, as well as the unreliability of natural language diagnostics. To this end, we propose FAULT, a framework that innovatively transforms unstructured natural language self-diagnoses into quantifiable, step-level explicit credit signals. By online learning relative error costs, FAULT enables the co-evolution of the policy and the diagnostic module. Experimental results demonstrate that FAULT achieves a 95% signal coverage rate on ALFWorld, significantly outperforming the GRPO and GiGPO baselines. It substantially improves performance on long-horizon tasks while maintaining competitiveness on short-horizon ones.
📝 Abstract
Agentic reinforcement learning (RL) has emerged as a powerful approach for training large language model agents on multi-step tasks, yet reliance on terminal outcome rewards creates two credit-assignment problems, particularly in long-horizon tasks. First, same-outcome rollout groups provide no learning signal from terminal rewards. Second, terminal rewards provide only trajectory-wide feedback, making it difficult to identify which decisions caused a failure. Recent work supplements terminal rewards with finer-grained information from trajectory analysis, such as natural-language reflections on intermediate decisions and errors. However, natural-language diagnoses are difficult to use directly for credit assignment: their error claims may be unreliable, and they do not quantify how much each error should affect learning. We propose Self-Diagnosis-guided Terminal Credit Redistribution (FAULT), which turns diagnosed errors into explicit step-level credit anchored by terminal outcomes. FAULT checks diagnostic evidence and learns relative error costs from task outcomes. During training, the policy and self-diagnoser co-evolve, while error costs are updated online from recent outcomes. On ALFWorld, FAULT recovers learning signals from same-outcome groups, reaching 95% signal coverage versus 41% for GRPO and 72% for GiGPO, while better localizing credit to specific error steps. Across two model scales, FAULT delivers strong. improvements on the long-horizon ALFWorld and WebShop tasks while remaining competitive on short-horizon Search-based QA.
Problem

Research questions and friction points this paper is trying to address.

credit assignment
agentic reinforcement learning
terminal reward
self-diagnosis
long-horizon tasks
Innovation

Methods, ideas, or system contributions that make the work stand out.

Credit Assignment
Self-Diagnosis
Agentic Reinforcement Learning
Credit Redistribution
Self-Evolving