π€ AI Summary
This study addresses the challenge of balancing false positive and false negative risks in AI fairness auditing under a specified tolerance threshold. To this end, it constructs a unified statistical testing framework tailored to two distinct objectives: violation certification and sensitivity screening. The proposed methodology introduces a constrained empirical likelihood test alongside an adaptive boundary surrogate principle. By integrating least favorable point calibration, split empirical likelihood, and spurious mark rate control strategies, the framework achieves differentiated trade-offs between error control and detection sensitivity. Numerical experiments validate the methodβs capacity for flexible risk balancing, while an empirical analysis on the COMPAS dataset demonstrates its practical effectiveness in predictive fairness auditing.
π Abstract
As artificial intelligence is increasingly deployed, algorithmic unfairness has raised growing concerns and intensified demands for transparent fairness auditing. In practice, the tolerable degree of algorithmic unfairness depends on the specific legal, ethical, or application context. Given a prespecified tolerance threshold, an important statistical question is how to determine whether a group disparity exceeds the allowable tolerance across different auditing objectives. To address this problem, we develop a unified tolerance-based fairness auditing framework for two complementary auditing objectives: violation certification, which prioritizes control of false violation declarations, and sensitivity screening, which prioritizes reducing missed violations. For the first objective, we develop a constrained empirical likelihood test for formal settings that uses least-favorable-point calibration and can be combined with false flagging rate control for simultaneous subgroup auditing. For the second objective, we develop split empirical likelihood and adjusted split empirical likelihood tests using an adaptive boundary-proxy principle for early-warning settings. Numerical experiments show the distinct error-control--sensitivity trade-offs of these procedures. A COMPAS analysis illustrates the framework in predictive fairness auditing.