Unanimously Wrong: Certified Abstention from How Medical LLM Consensus Forms

📅 2026-10-05
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the failure of abstention signals in medical LLM multi-agent consensus caused by the "confidently wrong" phenomenon, where agents unanimously agree on incorrect answers. To overcome this limitation, we propose ProbeGuard, a framework that pioneers the extraction of process-level features from consensus formation trajectories rather than relying solely on final states. ProbeGuard integrates semantic entropy with active evidence retrieval to assess reliability and employs a learn-then-test hierarchical calibration mechanism to achieve distribution-free risk bound control. Evaluated on MedQA, our approach attains an AUROC of 0.696 in distinguishing correct from incorrect consensus outcomes, successfully covering sixty percent of consistent errors under a nine-percent risk threshold. These results demonstrate that leveraging trajectory-derived process features significantly enhances the reliability assessment of multi-agent medical reasoning systems beyond traditional final-state approaches.
📝 Abstract
In clinical practice, agreement among independent experts is treated as evidence of reliability, and multi-round consensus has become a core mechanism of agentic medical question-answering systems. When such a system must decide whether to trust its own answer, the prevailing signal is again agreement, now among the sampled answers. But agreement is a fragile proxy for correctness. A system can be unanimously wrong, returning the same incorrect answer on every sample, and on these questions agreement-based signals carry no information. The cause is that these signals read only the final state of the consensus and discard how it was reached. Agreement that was reached by resolving disagreement with evidence looks identical, at the end, to agreement that was present from the first sample because every sample shares one misconception. ProbeGuard is a certified abstention framework that bases the abstention decision on how the consensus formed. Process features trace agreement trajectories, minority persistence, and retrieval saturation. For unanimous votes, rationale semantic entropy checks whether the reasons behind the vote cohere, and an active probe retrieves counter-evidence and measures whether the consensus survives. A stratified Learn-then-Test calibration then converts these scores into a distribution-free bound on selective risk. We evaluate ProbeGuard on three medical QA benchmarks and a hard-frontier reference, with a published multi-round agentic RAG substrate, against six abstention baselines. On MedQA, 13.4% of unanimous votes are wrong, and no agreement-based signal can flag them. Process signals raise the discrimination of correct from incorrect consensus from chance to 0.696 AUROC. The certified rule answers six in ten unanimous-layer questions at an observed selective risk of 9.0%, and nine in ten once in-domain calibration data accumulate.
Problem

Research questions and friction points this paper is trying to address.

medical LLM
consensus mechanism
certified abstention
unanimous error
selective risk
Innovation

Methods, ideas, or system contributions that make the work stand out.

Certified Abstention
Consensus Dynamics
Process Features
Rationale Semantic Entropy
Active Probing