More Debate, Same Evidence: Structural Limits of Homogeneous Multi-Agent Groundedness

📅 2026-07-31
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This study investigates whether homogeneous multi-agent debate mechanisms can effectively enhance the performance of large language models (LLMs) on groundedness verification tasks—specifically, assessing the truthfulness of claims based on supporting evidence. We construct a debate system comprising three identical LLMs and conduct a systematic evaluation across six publicly available fact-checking and hallucination detection benchmarks. The results reveal that the efficacy of multi-agent debate is highly dataset-dependent: significant accuracy improvements are observed on only two benchmarks (up to +8.5 percentage points), a notable decline occurs on one (−4.4 percentage points), and no statistically significant differences emerge on the remaining three. This work provides the first empirical evidence of the structural limitations and data dependency of such debate mechanisms in groundedness verification, challenging assumptions about their general applicability.
📝 Abstract
Large language model (LLM) judges are increasingly organized as multi-agent panels under the assumption that exchanging critiques improves judgment quality. We test this assumption for \emph{groundedness verification}, where a judge must determine whether a claim is supported by the supplied evidence. We evaluate a homogeneous three-agent panel on six public fact-verification and hallucination-detection benchmarks. Relative to a fixed single-agent reference, the panel's system-level accuracy difference ranges from $+8.5$ to $-4.4$ percentage points: two datasets show reliable gains, one shows a reliable loss, and three are statistically inconclusive. Because the reference and panel use different model variants, these differences characterize the complete systems rather than isolate a causal debate effect.
Problem

Research questions and friction points this paper is trying to address.

groundedness verification
multi-agent debate
LLM judges
fact verification
hallucination detection
Innovation

Methods, ideas, or system contributions that make the work stand out.

multi-agent debate
groundedness verification
LLM judges
homogeneous agents
fact verification
🔎 Similar Papers
2024-06-06AAAI Conference on Artificial IntelligenceCitations: 1