🤖 AI Summary
This study addresses the limitation of single-evidence protocols in adapting to diverse benchmarks and large language model judges by proposing the BAER framework, which dynamically selects optimal verification mechanisms via adaptive evidence routing. The core innovation lies in decoupling preference from reliability through a tri-symmetric head architecture and a frozen development-set selection strategy that preserves candidate symmetry. Furthermore, the framework introduces evidence stacking, reliability-based expert routing, and candidate-blind reference verification techniques. Experimental results demonstrate that BAER achieves the highest test accuracy across eight evaluation conditions, outperforming baselines by 0.87 to 7.32 points. These findings highlight the framework's significant robustness and generalization advantages in automated evaluation settings.
📝 Abstract
Pairwise language-model judges can gather evidence through direct comparison, reasoning, or reference-based verification, but no single protocol is best across benchmarks and judge backbones. We introduce Backbone-Adaptive Evidence Routing (BAER), which adapts the evidence mechanism while preserving candidate symmetry: swapping the two responses may reverse the preference but cannot change its strength. BAER separates each expert's signed preference from candidate-invariant reliability and builds three symmetric heads: evidence stacking, reliability-based expert routing, and candidate-blind reference verification. Development data select one head for each benchmark--backbone condition, and that choice is frozen before testing. Across four benchmarks and two 8B judge backbones, BAER achieves the highest test accuracy among the compared methods in all eight conditions, with full prediction coverage and gains of 0.87--7.32 points over the strongest external baseline. The results show that adapting how evidence is gathered is more reliable than fixing one judging protocol everywhere.