Sixteen models, fewer than two voices: measuring ensemble dispersion where no answer is uniquely correct

📅 2026-07-31
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenge of effectively measuring diversity and sources of disagreement in ensemble outputs of language models on psychotherapy case interpretation tasks, where no single correct answer exists. The authors construct an ensemble of 16 language models to generate 7,082 interpretations across 15 hierarchical clinical cases. They introduce a novel combination of the Vendi Score and a similarity-matrix-based disagreement contribution metric to quantify semantic diversity and identify key dissenting models. Through preregistered hypothesis testing, variance decomposition, and von Neumann entropy analysis, they find that model identity significantly shapes disagreement structure, though conventional categorizations—such as model family or scale—only partially account for this effect. Critically, the most divergent models shift dynamically with ensemble composition, indicating that dispersion is an emergent property requiring empirical measurement and that disagreement primarily stems from clinical content rather than case openness.
📝 Abstract
Sixteen language models drawn from ten families produced, on average, the semantic diversity of 1.69 distinct formulations of a psychotherapeutic case, against a single-model baseline of 1.43 from one model's own runs. Ensembles place more than one reading before a decision-maker on the premise that several models supply several perspectives. Dispersion over their outputs is measured both as diversity and as uncertainty, and both traditions validate it against a correctness criterion that this task does not admit. Measuring diversity is a solved problem: the Vendi Score, the exponential of the von Neumann entropy of a similarity matrix, is an effective number of distinct elements. What a single aggregate does not say is where the diversity comes from. We define a per-model dissent contribution, the complement of a model's mean similarity to the other members of its ensemble: a magnitude from the same matrix, not a decomposition of the spectral index, whose maximum identifies the most divergent voice. Crossing model and case, we test as a preregistered hypothesis whether model identity accounts for a non-zero share of the variance in dissent, and characterise the structure that test detects. The panel formulated fifteen stratified vignettes, yielding 7,082 formulations for analysis. Model identity was a detectable structuring factor of the dissent that remained, but the usual categories recovered it only partly: scale differences pointed in opposite directions across pairs, family grouped models on only five two-member lines, and the most divergent voice changed with panel composition, so that the surfaced outlier describes the ensemble rather than the model. Dissent did not track the interpretive openness for which the case bank was stratified; it was organised by clinical content instead, leaving the dispersion an ensemble produces a property to measure rather than assume.
Problem

Research questions and friction points this paper is trying to address.

ensemble dispersion
semantic diversity
model dissent
interpretive openness
language models
Innovation

Methods, ideas, or system contributions that make the work stand out.

ensemble dispersion
semantic diversity
dissent contribution
Vendi Score
model disagreement