A Diagnostic Framework and Multi-Evaluator Audit of Evaluator-Driven Preference Dynamics in Self-Adapting LLM Agents

๐Ÿ“… 2026-06-28
๐Ÿ“ˆ Citations: 0
โœจ Influential: 0
๐Ÿ“„ PDF
๐Ÿค– AI Summary
This work addresses the instability of preference judgments from proprietary large language model (LLM) evaluators, which are highly sensitive to version updates and thus yield unreliable short-term assessments. To systematically characterize evaluator-driven preference dynamics, the authors propose EPC, a diagnostic framework that integrates the Multimodal Preference Collapse Index (MPCI), evaluator coupling matrices, Jensenโ€“Shannon divergence, and output-format confusion analysis. Experiments reveal significant fluctuations in preference coupling and self-evaluation mechanism collapse across versions of models such as GPT-4o, demonstrating the severe limitations of single-snapshot evaluations. The study releases all data and tools publicly, establishing a new paradigm for robust and reliable LLM evaluation.
๐Ÿ“ Abstract
Measurements of proprietary LLM evaluators can become invalid within weeks -- we document one case and provide the diagnostic framework to detect it. We introduce EPC -- comprising the Multimodal Preference Collapse Index (MPCI), evaluator-indexed coupling matrix, and Jensen-Shannon divergence (JSD) -- and apply it across eight experimental conditions (N=112 main + N=10 ablation = 122 unique repetitions, all reported). Coupling coefficients range from 0.00 to 1.18 across per-condition means (CV approx 0.9, n=8 conditions). Four conditions show strong coupling (N=36; GPT-4o May, GPT-4o-mini, Qwen3.7-plus, DashScope 30r); four collapse to near-zero (N=76; GPT-4o June, qwen-plus N=30, symmetric LR, DeepSeek self-eval). The May-to-June GPT-4o drift -- an N=8 re-replication inverting the study's conclusion -- is the most informative measurement: a diagnostic instrument detecting its own instability demonstrates the fragility it was designed to measure. Self-evaluation (97% zero, JSD=0.003) consistently collapses, though floor effects are possible. Output-format confound analysis finds per-strategy aggregate rho=0.89 but per-instance rho=0.219 (p=0.093); PCI reported as preference-convergence metric. We release EPC with all data. The finding is not any single coupling magnitude but the pattern of version-conditional instability that makes single-snapshot evaluator studies unreliable.
Problem

Research questions and friction points this paper is trying to address.

preference dynamics
evaluator instability
self-adapting LLM agents
preference collapse
diagnostic framework
Innovation

Methods, ideas, or system contributions that make the work stand out.

Evaluator Preference Collapse
Preference Dynamics
Coupling Matrix
Jensen-Shannon Divergence
Self-Adapting LLM Agents
L
Liu Zewen
Qilu Institute of Technology, School of Software Engineering, Tai'an, Shandong, China