Training Advisors for LLM Agents from Task Outcomes
This study addresses the lack of effective natural language feedback for large language model (LLM) agents in multi-step tasks by proposing Caddie. This method trains a critic model via reinforcement learning, optimizing end-to-end using only task-level success or failure signals without requiring step-level annotations or reference critiques, thereby providing agents with real-time decision-making guidance. Experiments employing a 4B-parameter critic demonstrate that Caddie achieves zero-shot transferability across models and domains. On the MuSiQue benchmark, it improves the success rates of multiple base models by over 25 percentage points and surpasses the performance of Kimi K3 operating without a critic.