JudgeMoE: Distributional Aggregation for LLM-as-a-Judge

📅 2026-10-05
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the loss of distributional uncertainty caused by score scalarization in large language model evaluation. To overcome this limitation, we propose a distribution-based mixture-of-experts aggregation mechanism that employs a lightweight aggregator to dynamically assign sample-specific weights to cached scoring distributions for weighted fusion. Combined with a soft scoring protocol, this framework effectively preserves uncertainty information while substantially enhancing evaluation performance. Experimental results demonstrate significant improvements in Spearman correlation coefficients, with the proposed method outperforming the strongest single-judge model on 12 of 16 test units. These findings validate the effectiveness of the distributional aggregation strategy in mitigating the shortcomings of conventional scalar compression for LLM-based evaluation.
📝 Abstract
When an LLM judge scores an output, its score distribution retains uncertainty and disagreement information that is lost after scalar compression. We introduce JudgeMoE, a lightweight aggregator that assigns example-specific weights to cached judge score distributions and fuses them before computing a final score. A protocol study shows that score-range choice is unstable across judge--dataset settings and that soft scoring usually outperforms hard decoding. On the original 10-cell benchmark, JudgeMoE improves mean Spearman over uniform log pooling by $+0.079$. Applying the same configuration to six additional cells yields a $+0.0393$ mean gain over the strongest local single judge across 16 cells, with positive differences in 12/16 cells and a one-sided Wilcoxon signed-rank $p=0.0091$. Validation-based analyses further show that the preferred aggregation method depends on the task and judge pool.
Problem

Research questions and friction points this paper is trying to address.

LLM-as-a-Judge
score distribution
scalar compression
aggregation
evaluation uncertainty
Innovation

Methods, ideas, or system contributions that make the work stand out.

LLM-as-a-Judge
Distributional Aggregation
Mixture of Experts
Soft Scoring
Score Distribution