🤖 AI Summary
This study investigates whether prevailing diversity metrics genuinely capture model disagreement in large language model (LLM) ensembles or merely reflect individual model capabilities. Through controlled experiments on 31,900 subsets of 30 LLMs evaluated on MMLU-Pro and TruthfulQA, the authors systematically assess the predictive power of five diversity measures for majority-vote performance gains. Using Spearman correlation, linear regression, and determinant-based algebraic analysis, they find that most diversity metrics are highly collinear with model capability. After rigorously controlling for capability, only “strict diversity,” “disagreement,” and “double-failure” retain weak yet consistent and directionally aligned associations with ensemble failure rates. The results further reveal that simple majority voting surpasses the best individual model in only a small minority of subsets, highlighting the limited utility of current diversity metrics in LLM ensembling.
📝 Abstract
Majority voting over LLMs is widely assumed to benefit from diversity, and diversity measures are used to choose which models to combine. We ask whether five such measures track diversity or mainly re-express capability, auditing them as predictors of majority-vote gain over the best member across 31,900 subsets of 30 LLMs on MMLU-Pro (29 on TruthfulQA) under explicit capability controls. Three findings emerge. First, latent complementarity is ubiquitous: oracle gain is positive in 100% of subsets, yet simple voting beats the strongest member in only 9.98% of all canonical size-3 subsets (18.71% with held-out best selection); the pooled size-2-4 rate is 1.27%, partly reflecting deterministic even-size voting behavior. Second, a joint-correctness proxy (strict diversity) is nearly collinear with one minus mean accuracy (size-3 Spearman rho = +0.991 / +0.988); raw diversity-gain associations are strongly capability-entangled and, with one exception, unstable under control. Third, three linear contingency-table statistics are algebraically non-separable; after capability control, the empirically stable remainder is a modest residual pairwise co-failure association in which more shared error corresponds to lower gain. This direction is robust, but its magnitude is configuration-dependent. Joint rawspace linear regressions treating strict diversity, disagreement, and double-fault as independent predictors are rank-deficient by construction.