Efficient Active Auditing of Multi-Group Fairness with Bias Probes

📅 2026-09-30
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the lack of specificity and confidentiality leakage risks inherent in fairness auditing of black-box models by proposing a bias probe framework alongside an active auditing algorithm, ALeBi. Leveraging active learning for adaptive querying, this method efficiently estimates multiple groups of fairness metrics. Theoretically, it establishes sample complexity guarantees, elucidates the trade-off between confidentiality and auditing reliability, and resolves related open problems. Experimental results demonstrate that ALeBi not only accurately estimates fairness metrics but also identifies interpretable high- and low-bias regions, effectively validating the superiority of the proposed approach.
📝 Abstract
Over the past decade, Machine Learning (ML) has been trained under dual objectives: minimizing prediction error via Empirical Risk Minimization (ERM) while controlling unfairness bias. In practice, however, fairness-aware training often yields limited improvements over standard ERM, making reliable post hoc auditing essential. Existing auditing approaches for black-box models either rely on model reconstruction --exposing systems to extraction attacks-- or directly estimate fairness metrics, offering limited insight into which regions of the data distribution drive bias. More fundamentally, property-specific auditing --aimed at extracting only targeted fairness information without reconstructing the model-- remains poorly understood. In this work, we introduce the bias probe framework, which enables targeted and adaptive querying to reveal bias structure while preserving model confidentiality. Building on this framework, we propose ALeBi, an active auditor that learns such probes to efficiently estimate multi-group fairness metrics. We establish novel sample complexity guarantees governed by a property-specific complexity measure, resolving a previously posed open question, and extend our analysis to adversarial settings where the model owner may strategically obscure bias. Our results uncover a fundamental trade-off between model confidentiality and reliable auditing, and show that property-specific probing enables both accurate estimation and interpretable identification of high and low-bias regions. Extensive experiments support our theoretical findings and demonstrate the practical effectiveness of our approach.
Problem

Research questions and friction points this paper is trying to address.

fairness auditing
multi-group fairness
black-box models
model confidentiality
bias detection
Innovation

Methods, ideas, or system contributions that make the work stand out.

Bias Probes
Active Auditing
Multi-Group Fairness
Sample Complexity
Model Confidentiality
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.