Measuring Collapse and Correction in Homogeneous-Panel LLM Debate

📅 2026-09-28
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the common conflation of error-correction and collapse mechanisms in evaluating the final accuracy of multi-agent LLM debates. We propose an auditable protocol based on transition ledgers to quantify debate collapses, corrections, and intervention utilities, while introducing pre-debate probe screening as a triage signal. Across 6,925 debates on the MMLU-Pro benchmark, this framework identifies 253 collapse events, revealing that early rounds are prone to triggering detrimental cascades and exposing an inherent trade-off wherein preventing collapses may compromise corrections. Furthermore, we release a replayable architecture to standardize evaluation criteria. This work establishes a new paradigm for understanding the dynamics of multi-agent debates.
📝 Abstract
Multi-agent large language model (LLM) debate is often evaluated by whether final answers improve, but movement is not necessarily improvement: the same discussion can rescue an initially wrong majority or destroy an initially correct one. Standard final-accuracy evaluations conflate these opposing mechanisms. We introduce an auditable protocol for homogeneous debate on multiple-choice questions (MCQs) that records each run as a transition ledger over collapse, correction, onset, and signed intervention utility. On 6,925 MMLU-Pro debates, the protocol identifies 253 collapses and a parallel correction ledger that changes how interventions should be judged. Replay experiments reveal the central tradeoff: a leave-one-model-out probe-gated freeze prevents 29 collapses but loses 108 corrections under equal weights, so collapse prevention alone can recommend the wrong policy. A compact pre-debate 8-probe screen is a triage signal: its unadjusted family-level association with conditional-collapse risk is high (G=7, Spearman rho=0.893, exact two-sided p=0.0123), but initial-majority accuracy is a close comparator (rho=0.821; family partial rho=0.767, p=0.0877), so we do not treat it as calibrated or capability-adjusted prediction. Round-level traces localize many collapses to the first debate round, where early disagreement can precede both harmful cascades and useful recovery. We release replayable schemas, coders, audits, cost cards, and zero-API rebuild scripts so future model-scaffold rows can be compared under the same denominators and signed utility ledger.
Problem

Research questions and friction points this paper is trying to address.

LLM debate
collapse measurement
correction measurement
multi-agent evaluation
homogeneous panel
Innovation

Methods, ideas, or system contributions that make the work stand out.

multi-agent LLM debate
auditable protocol
transition ledger
probe-gated freeze
signed intervention utility