🤖 AI Summary
This study investigates whether homogeneous multi-agent debate mechanisms can effectively enhance the performance of large language models (LLMs) on groundedness verification tasks—specifically, assessing the truthfulness of claims based on supporting evidence. We construct a debate system comprising three identical LLMs and conduct a systematic evaluation across six publicly available fact-checking and hallucination detection benchmarks. The results reveal that the efficacy of multi-agent debate is highly dataset-dependent: significant accuracy improvements are observed on only two benchmarks (up to +8.5 percentage points), a notable decline occurs on one (−4.4 percentage points), and no statistically significant differences emerge on the remaining three. This work provides the first empirical evidence of the structural limitations and data dependency of such debate mechanisms in groundedness verification, challenging assumptions about their general applicability.
📝 Abstract
Large language model (LLM) judges are increasingly organized as multi-agent panels under the assumption that exchanging critiques improves judgment quality. We test this assumption for \emph{groundedness verification}, where a judge must determine whether a claim is supported by the supplied evidence. We evaluate a homogeneous three-agent panel on six public fact-verification and hallucination-detection benchmarks. Relative to a fixed single-agent reference, the panel's system-level accuracy difference ranges from $+8.5$ to $-4.4$ percentage points: two datasets show reliable gains, one shows a reliable loss, and three are statistically inconclusive. Because the reference and panel use different model variants, these differences characterize the complete systems rather than isolate a causal debate effect.