🤖 AI Summary
This study addresses the challenges of high external verification costs and inaccurate self-evaluation scores in step-level assessment for heterogeneous large language model agents. To this end, we propose GLIDE, a method that introduces a novel intrinsic evaluation mechanism based on inter-layer residual consistency. This mechanism extracts unified value signals across diverse agents without requiring external verifiers. Furthermore, by integrating distribution calibration to generate pessimistic rewards, GLIDE guides branch selection and adaptive expansion within Monte Carlo Tree Search (MCTS). Experimental results demonstrate that GLIDE significantly improves reasoning performance and ranking quality on tasks such as multi-hop reasoning, while effectively enhancing computational efficiency.
📝 Abstract
LLM agents require reliable step-level evaluation to compare candidate branches and allocate computation effectively. However, lightweight evaluation remains challenging. External verifiers introduce additional inference cost, while agent-produced confidence or self-evaluation scores can be miscalibrated, especially when candidates are generated by heterogeneous agents. We propose \textbf{G}eneralized \textbf{L}ayer-wise \textbf{I}ntrinsic \textbf{D}istributional \textbf{E}valuation (\textbf{GLIDE}) for LLM agents. \textsc{GLIDE} derives intrinsic step evidence from layer-wise residual coherence, which measures whether local residual updates consistently support the global residual change induced by a candidate step. It calibrates this evidence against the recent score distribution of the generating agent and converts it into a pessimistic reward that jointly accounts for absolute residual evidence and agent-relative standing. The reward provides a cross-agent value signal for MCTS branch selection, while normalized predictive uncertainty guides adaptive branching. Experiments on multi-hop reasoning, sequential decision making, and symbolic logic show that \textsc{GLIDE} improves task performance, step-level ranking quality, and computational efficiency without external verifiers or task-specific supervision.