🤖 AI Summary
This study addresses the limitations of reinforcement learning interventions for large language models, where aggregate metrics obscure mechanistic discrepancies and reward signals induce undesirable behaviors. To overcome these issues, this work proposes an analytical framework that translates aggregated outcomes into actionable diagnostics. Methodologically, it systematically traces the causal sources of performance disparities and validates corrective strategies through controlled reward-policy comparisons, signaling audits, and counterfactual replays. Experimental results demonstrate that inverse dense rewards precipitate sharp declines in win rates, revealing that strong influence does not equate to high performance. Furthermore, the approach successfully rectifies action-level errors and extends this diagnostic paradigm to incomplete-information game settings. Ultimately, this research establishes a novel paradigm for mechanism auditing in reinforcement learning from human feedback (RLHF).
📝 Abstract
Large language models can intervene in reinforcement learning through both reward design and action selection, yet aggregate performance offers an incomplete account of what these interventions actually do. Similar returns can conceal different learning mechanisms, while plausible rewards can induce undesirable behavior. We introduce LocusRL, a diagnostic framework that connects controlled reward-policy comparisons with audits of reward judgments, signal delivery, optimization objectives, and executed actions. The framework traces performance differences to testable explanations and checks targeted corrections through executable rules and counterfactual replay. Across two evaluation batches covering ten Connect Four training seeds, we uncover seed-dependent reversals in intervention effects and show how tracing actual updates changes their interpretation: historical Qwen training operates through reward-weighted teacher-action likelihood. A separate matched three-seed reward-direction experiment distinguishes sensitivity to a learning signal from its usefulness. With terminal rewards held fixed, a sign-reversed dense oracle yields a 2.8% aggregate win rate, compared with 57.2% for terminal-only training and 46.7% for the positive dense oracle. Thus, a reward can strongly influence learning without improving performance. At the decision level, counterfactual replay verifies a winning correction to a diagnosed action error. Complementary experiments in Leduc and reward-validation studies in Goofspiel extend the analysis to imperfect-information settings, revealing how reference-label definitions and validation-data exposure affect intervention assessment. Together, these findings show why evaluating LLM interventions requires tracing how their outputs become learning signals and actions. LocusRL turns aggregate outcomes into actionable diagnoses and verifiable corrections.