Selective Critique for Cost-Aware LLM Agents in Long-Horizon Decision Making

📅 2026-10-05
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the inherent trade-off between high computational cost and low reliability arising from invoking critique mechanisms at every step in long-horizon decision-making for LLM agents. To this end, we propose the SAG framework, which introduces a training-free gating mechanism based on action ambiguity. By approximating the Value of Information (VoI) through global entropy and local Top-2 margins, SAG selectively triggers critiques only when necessary. Furthermore, it integrates online bootstrapping to progressively reduce the agent's reliance on external feedback. Evaluated on the ALFWorld benchmark, SAG improves task success rates from 24.6% to 78.4% while achieving a 3.1× gain in normalized token efficiency under a ReAct-level token budget, effectively balancing decision reliability with reasoning economy.
📝 Abstract
Improving the reliability of large language model (LLM) agents in long-horizon decision-making remains a key challenge. When deployed as autonomous agents interacting with complex environments, early mistakes can propagate through trajectories and cause cascading failures. Recent approaches improve reliability by incorporating external critique or deliberation, but invoking these mechanisms at every step substantially increases token consumption and latency, limiting practical deployment. We propose SAG (Self-improving Agent with Gated critique), a cost-aware framework that formulates critique invocation as a step-wise decision problem during long-horizon interaction. SAG introduces a lightweight, training-free gating mechanism that estimates the utility of critique using action-level ambiguity signals--global entropy and local top-2 margin--computed over admissible actions. From a decision-theoretic perspective, this mechanism approximates the Value of Information (VoI) of critique, enabling the agent to selectively allocate expensive feedback only when its expected benefit justifies the cost. SAG further incorporates online bootstrapped self-improvement, allowing the actor to internalize critic-assisted behaviors and progressively reduce reliance on critique. Across three long-horizon interactive benchmarks and multiple backbone models, SAG substantially improves the performance-cost trade-off compared with both no-critique and always-on critique agents. On ALFWorld, SAG increases task success from 24.6% to 78.4% while maintaining a token budget comparable to ReAct, yielding a $3.1\times$ improvement in normalized token efficiency. Moreover, a 7B actor with a lightweight 3B critic achieves performance comparable to a 14B actor without critique, showing that selective critique can recover most of the reliability benefits of deliberation while dramatically reducing inference cost.
Problem

Research questions and friction points this paper is trying to address.

Large Language Model Agents
Long-Horizon Decision Making
Cost-Aware Inference
Selective Critique
Reliability
Innovation

Methods, ideas, or system contributions that make the work stand out.

Gated Critique
Value of Information
Cost-Aware Agent
Self-Improvement
Long-Horizon Decision Making
🔎 Similar Papers
No similar papers found.
H
Heewon Park
Department of Electrical and Computer Engineering, Sungkyunkwan University, Republic of Korea
S
Somin Im
Soongsil University, Republic of Korea
Minhae Kwon
Minhae Kwon
Associate Professor at Soongsil University
Reinforcement LearningComputational NeuroscienceAutonomous DrivingFederated Learning