Drowning in Routine: Signal Dilution in Multi-Turn Agent Training

📅 2026-06-20
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the inefficiency of trajectory-level credit assignment in multi-turn agent training, where routine actions dilute gradient signals from critical decisions. It establishes, for the first time, a quantitative √ρ⁻¹ relationship between decision density ρ and the signal-to-noise ratio in policy optimization, demonstrating that non-critical steps at low ρ introduce noise without contributing to returns, thereby severely degrading training efficiency. By constructing controllable environments with tunable ρ, conducting theoretical analysis of trajectory-level policy optimization methods (e.g., GRPO), and modeling gradient variance, the work precisely delineates the applicability boundaries of trajectory-level approaches across high and low decision densities. Experiments strongly corroborate the theoretical predictions (R²=0.999) and reveal a sharp divergence in required training steps as ρ approaches zero.
📝 Abstract
Multi-turn agents interleave consequential decisions with routine execution: some actions change the downstream return distribution, while others are necessary but reward-equivalent. The cost of trajectory-level credit assignment, often attributed to long horizons, is in fact governed by decision density $ρ$: the fraction of turns whose actions affect the return. When decision density is low, routine turns create signal dilution: they add gradient variance to trajectory-level estimators such as GRPO without adding expected signal. Under explicit assumptions, the resulting turn-level to trajectory-level signal-to-noise ratio scales as $ρ^{-1/2}$, provided critic error remains controlled. The same analysis identifies the complementary regime: at high decision density, trajectory-level methods can remain competitive while avoiding the cost of a critic. In a controlled environment where $ρ$ is exactly tunable, the predicted scaling is recovered with $R^2 = 0.999$, and the training-step gap widens significantly as $ρ\to 0$.
Problem

Research questions and friction points this paper is trying to address.

signal dilution
multi-turn agents
decision density
credit assignment
trajectory-level learning
Innovation

Methods, ideas, or system contributions that make the work stand out.

decision density
signal dilution
credit assignment
trajectory-level policy gradient
multi-turn agents
💼 Related Jobs
No related jobs found.
Y
Yann Pernot
Mila - Québec AI Institute; Polytechnique Montréal
V
Vi Retault
Polytechnique Montréal