Agreement Before Diversity: Verification-First Complementarity for Heterogeneous Language-Model Coordination

📅 2026-08-05
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the lack of principled replacement criteria in existing ensembles of heterogeneous language models. It proposes Agreement-Before-Diversity (ABD), a method that employs an equivalence-relation-based verification gating mechanism to decide whether to retain an anchor answer: the anchor is preserved only if two trusted samples agree with it; otherwise, it is replaced by a heterogeneously synthesized result. ABD decouples candidate expansion from authoritative replacement and introduces a label-free, fine-tuning-free decision rule. The approach establishes two auditable identities that do not rely on independence assumptions or confidence calibration. Evaluated on LiveCodeBench-v6 and GPQA-Diamond, ABD achieves accuracies of 59.43% and 75.00%, respectively, significantly outperforming baselines and enabling precise attribution of performance gains to enumerable, protected subsets.
📝 Abstract
Heterogeneous language-model ensembles expand the space of candidate responses, yet they lack a principled criterion for when a newly generated answer should supersede an already supported one. We decouple candidate headroom from replacement authority, rendering the latter as an explicit, auditable object. Our proposed method, Agreement-Before-Diversity (ABD), is a frozen, label-free decision rule: an anchor answer is retained if two additional trusted samples corroborate it under a fixed equivalence relation; otherwise, it is replaced by a heterogeneous synthesis. For this gating mechanism, we prove two exact identities. The first shows that the accuracy gap relative to unconditional synthesis is determined jointly by the agreement coverage and the anchor's advantage on the protected subset. The second shows that the gap relative to never synthesizing reflects a contrast between authorized recovery and authorized destruction. Neither identity assumes independence or calibrated confidence, and the expected inference cost is approximately eight minus five times the coverage in number of calls. Under blind, exact-ID evaluation, ABD achieves 59.43% on the complete LiveCodeBench-v6 (vs. 52.57% for Single9 and 52.00% for HAC; n = 175) and 75.00% on an untouched GPQA-Diamond split (both controls at 72.78%; n = 180). Furthermore, these identities localize every aggregate difference to an enumerable protected stratum: no discordant items occur among the 3 protected cases on LiveCodeBench, where coverage bounds the gate's contribution to 1.71 points a priori; 13 versus 8 discordant cases among 132 on GPQA-Diamond; and 12 versus 0 among 71 under a frozen anchor perturbation. Diversity supplies potential; verification structure supplies authority.
Problem

Research questions and friction points this paper is trying to address.

heterogeneous language-model ensembles
answer replacement criterion
verification
coordination
candidate responses
Innovation

Methods, ideas, or system contributions that make the work stand out.

Agreement-Before-Diversity
heterogeneous ensembles
verification-first
anchor answer
equivalence relation
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
R
Ruitong Li
The University of Hong Kong
Binjie Guo
Binjie Guo
PhD Candidate,Zhejiang University
Deep LearningDeep Generative ModellingNatural Language ProcessingBrain Science
A
Aisheng Mo
Zhejiang University
G
Guowei Su
Zhejiang University
J
Jie Li
Independent Researcher
R
Ru Zhang
Zhejiang University