Reach Into The CHOIR: Free-List Elicitation Uncovers Distinct Model Voices in LLM Ensembles

📅 2026-09-29
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the issues of spurious diversity and default-answer masking caused by surface homogeneity in large language model (LLM) ensembles by proposing the CHOIR framework. CHOIR innovatively adapts free listing from cognitive anthropology to LLM evaluation, conducting a layered and ordered investigation of ensemble systems through hierarchical ranked list generation, concept clustering, saliency measurement, and source-blind ranking. This approach effectively distinguishes prompt-induced constraints from deep stability, thereby decoupling individual model voices. Experiments on the Infinity-Chat benchmark successfully reproduce and disentangle high surface consistency under narrow and broad prompts, precisely identifying base model identity as the most dominant feature. Ultimately, this work establishes a new paradigm for the in-depth evaluation of LLM ensembles.
📝 Abstract
Open-ended LLM homogeneity can create false plurality when several systems appear to offer independent perspectives while returning the same familiar default. Single-pass answers obscure the distinction between agreement produced by a tightly constrained answer space, prompt-vocabulary echo, and broader answer spaces with stable alternatives beneath the surface. We introduce CHOIR (Collective Hierarchically-Ordered Inquiry Responses), a framework that adapts free-list elicitation from cognitive anthropology to LLM ensembles. CHOIR repeatedly elicits ranked lists, clusters items into prompt-level concepts, and measures concept salience across models, prompt variants, and persona conditions. We evaluate CHOIR on Infinity-Chat 100, an external prompt bank from recent work on open-ended model homogeneity, and on a 27-question targeted diagnostic bank designed to isolate mechanism-level contrasts. On Infinity-Chat 100, CHOIR reproduces high surface agreement (93/100 prompts above chance) while separating narrow prompts from broad prompts with recoverable depth. Across targeted probes and the external prompt bank, base-model identity remains the strongest recoverable signature, and persona prompts shift surfaced concepts within base-model signatures. A source-blind ranking module prioritises rare-but-stable candidates for later inspection. CHOIR turns open-ended homogeneity into a diagnostic measurement problem by asking where models converge, why they converge, and what remains reachable under structured depth probing.
Problem

Research questions and friction points this paper is trying to address.

LLM homogeneity
false plurality
open-ended generation
model convergence
answer space
Innovation

Methods, ideas, or system contributions that make the work stand out.

Free-List Elicitation
LLM Ensembles
Model Homogeneity
Concept Salience
Source-Blind Ranking