๐ค AI Summary
This study addresses the miscalibrated confidence and inability to identify persistently unreliable agents in multi-LLM parliamentary frameworks by reformulating multi-agent debate as a reliability estimation problem. Methodologically, it is the first to leverage debate trajectories as observational data, integrating Bayesian dialectical argumentation with classical annotator models to infer agent reliability and weight evidence accordingly. This enables the inversion of malicious agents rather than merely suppressing them through majority voting. The proposed zero-cost aggregation technique outputs calibrated posterior probabilities without requiring additional model invocations. Empirically, it significantly enhances adversarial robustness while maintaining competitive performance in benign settings, achieving state-of-the-art calibration.
๐ Abstract
A multi-LLM \emph{council} lets several large language models (LLMs) deliberate on a question and return an answer together with a confidence estimate. As these systems become increasingly used for reasoning, that confidence should represent a calibrated \emph{probability of being correct}, and the decision should remain robust when some agents are persistently unreliable. Existing \emph{council aggregation} methods fail on both fronts: their confidence estimates measure decisiveness rather than correctness, and they cannot identify or discount persistently unreliable agents. We introduce Bayesian Dialectical Argumentation (BDA), which treats the council's \emph{typed} moves---who proposed, challenged, or conceded which answer---as observations of a classical annotator model with \emph{per-agent} reliabilities. This formulation recasts multi-agent deliberation as a reliability estimation problem, using the deliberation trace to infer agent reliability under persistent adversarial behavior. By weighting evidence according to inferred agent reliability, BDA yields calibrated posterior probabilities over candidate answers while allowing persistently unreliable agents to be inverted rather than merely outvoted. Across binary and multi-class benchmarks, BDA achieves the best calibration among zero-cost council aggregation methods, requiring no additional LLM calls, and improves robustness under persistent adversarial coalitions while remaining competitive in clean settings.