GLIDE: Generalized Layer-wise Intrinsic Distributional Evaluation for Heterogeneous LLM Agents

📅 2026-09-26
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenges of high external verification costs and inaccurate self-evaluation scores in step-level assessment for heterogeneous large language model agents. To this end, we propose GLIDE, a method that introduces a novel intrinsic evaluation mechanism based on inter-layer residual consistency. This mechanism extracts unified value signals across diverse agents without requiring external verifiers. Furthermore, by integrating distribution calibration to generate pessimistic rewards, GLIDE guides branch selection and adaptive expansion within Monte Carlo Tree Search (MCTS). Experimental results demonstrate that GLIDE significantly improves reasoning performance and ranking quality on tasks such as multi-hop reasoning, while effectively enhancing computational efficiency.
📝 Abstract
LLM agents require reliable step-level evaluation to compare candidate branches and allocate computation effectively. However, lightweight evaluation remains challenging. External verifiers introduce additional inference cost, while agent-produced confidence or self-evaluation scores can be miscalibrated, especially when candidates are generated by heterogeneous agents. We propose \textbf{G}eneralized \textbf{L}ayer-wise \textbf{I}ntrinsic \textbf{D}istributional \textbf{E}valuation (\textbf{GLIDE}) for LLM agents. \textsc{GLIDE} derives intrinsic step evidence from layer-wise residual coherence, which measures whether local residual updates consistently support the global residual change induced by a candidate step. It calibrates this evidence against the recent score distribution of the generating agent and converts it into a pessimistic reward that jointly accounts for absolute residual evidence and agent-relative standing. The reward provides a cross-agent value signal for MCTS branch selection, while normalized predictive uncertainty guides adaptive branching. Experiments on multi-hop reasoning, sequential decision making, and symbolic logic show that \textsc{GLIDE} improves task performance, step-level ranking quality, and computational efficiency without external verifiers or task-specific supervision.
Problem

Research questions and friction points this paper is trying to address.

step-level evaluation
heterogeneous LLM agents
lightweight evaluation
confidence calibration
branch selection
Innovation

Methods, ideas, or system contributions that make the work stand out.

Intrinsic Evaluation
Layer-wise Residual Coherence
Heterogeneous LLM Agents
Monte Carlo Tree Search
Pessimistic Reward
W
Wei Zhu
School of Information Science and Engineering, Yunnan University, Kunming, China; Yunnan Key Laboratory of Intelligent Systems and Computing, Kunming, China
Yiming Wang
Yiming Wang
Shanghai Jiao Tong University
Large Language ModelsComplex ReasoningAI Interpretability
R
Rui Wang
School of Computer Science, Shanghai Jiao Tong University, Shanghai, China
Lixing Yu
Lixing Yu
Yunnan University
Distributed Machine Learning
Kun Yue
Kun Yue
Tsinghua University
HCI AR BCI XR
Zhiwen Tang
Zhiwen Tang
Yunnan University