Black-Box Auditing of Epistemic Reliability in Multi-Agent Debate Distillation

πŸ“… 2026-09-26
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This study addresses the issue in multi-agent debate distillation where improvements in monitored task performance often mask degraded cognitive reliability on hidden tasks. To this end, we propose ER-Audit, a black-box auditing framework that detects degradation risks by comparing pre- and post-adaptation checkpoints and searching for counterexamples. It introduces a two-stage auditing mechanism with anytime-valid confidence lower bounds, enabling data-dependent stopping and statistical inference under limited budgets. Furthermore, semantic-equivalent paraphrase generation, sequential hypothesis testing, and adversarial debate simulation are integrated to enhance auditing robustness. Our findings reveal that high hidden accuracy can coexist with lowered non-degradation bounds, demonstrating that aggregate metrics fail to detect selective degradation. The code is publicly available.
πŸ“ Abstract
Debate distillation adapts weaker verifiers using multi-agent debate transcripts to improve their judgement in subsequent debates, but gains on monitored tasks do not establish reliability on related unmonitored tasks. We study epistemic reliability degradation, in which adaptation preserves monitored performance while reducing support for correct responses on hidden tasks. We consider an adversarial debater that manipulates debate arguments while defending the correct monitored response, and ask whether the resulting degradation merely reflects catastrophic forgetting and whether standard evaluation can detect it. To address these questions, we propose ER-Audit, a two-stage black-box auditing framework that compares frozen verifier checkpoints before and after adaptation, and introduce two evaluation benchmarks pairing monitored and hidden task prompts grounded in shared contexts. ER-Audit searches for counterexamples to non-degradation by evaluating semantically valid paraphrases and, if none is found, uses independent paraphrases for sequential hypothesis testing. We derive anytime-valid lower confidence bounds on the non-degradation probability, allowing data-dependent stopping within a finite budget. We further establish a common lower bound across fixed paraphrase distributions and extend it to distributions within a bounded total variation distance of their mixtures. Our experiments show that higher hidden-task accuracy can coexist with more counterexamples to non-degradation and lower non-degradation bounds. This divergence challenges explanations based solely on broad catastrophic forgetting and shows that auditing can uncover selective hidden-task degradation concealed by aggregate performance gains. Our code and benchmarks are available at https://github.com/CSIRO-CQS-AI-alignment-Team/Epistemic-Reliability-Auditor.
Problem

Research questions and friction points this paper is trying to address.

Multi-Agent Debate Distillation
Epistemic Reliability Degradation
Black-Box Auditing
Adversarial Debater
Hidden Task Performance
Innovation

Methods, ideas, or system contributions that make the work stand out.

Black-Box Auditing
Debate Distillation
Epistemic Reliability
Sequential Hypothesis Testing
Anytime-Valid Confidence Bounds