Why Aggregate Accuracy is Inadequate for Evaluating Fairness in Law Enforcement Facial Recognition Systems

📅 2026-03-30
📈 Citations: 0
Influential: 0
📄 PDF

career value

224K/year
🤖 AI Summary
This study addresses the limitations of aggregate accuracy metrics in evaluating facial recognition systems within law enforcement contexts, which often obscure performance disparities across demographic groups and fail to capture true fairness and reliability. The authors propose moving beyond a single accuracy measure by introducing a fairness-aware evaluation framework coupled with a model-agnostic auditing strategy. By analyzing false positive and false negative rates at the subpopulation level, this approach uncovers hidden group-level biases that persist even when overall accuracy appears high. Empirical results demonstrate that systems with comparable aggregate accuracy can exhibit substantially different error distributions across demographic groups, underscoring the necessity and effectiveness of the proposed paradigm for enabling more responsible and equitable deployment of facial recognition technologies.

Technology Category

Application Category

📝 Abstract
Facial recognition systems are increasingly deployed in law enforcement and security contexts, where algorithmic decisions can carry significant societal consequences. Despite high reported accuracy, growing evidence demonstrates that such systems often exhibit uneven performance across demographic groups, leading to disproportionate error rates and potential harm. This paper argues that aggregate accuracy is an insufficient metric for evaluating the fairness and reliability of facial recognition systems in high-stakes environments. Through analysis of subgroup-level error distribution, including false positive rate (FPR) and false negative rate (FNR), the paper demonstrates how aggregate performance metrics can obscure critical disparities across demographic groups. Empirical observations show that systems with similar overall accuracy can exhibit substantially different fairness profiles, with subgroup error rates varying significantly despite a single aggregate metric. The paper further examines the operational risks associated with accuracy-centric evaluation practices in law enforcement applications, where misclassification may result in wrongful suspicion or missed identification. It highlights the importance of fairness-aware evaluation approaches and model-agnostic auditing strategies that enable post-deployment assessment of real-world systems. The findings emphasise the need to move beyond accuracy as a primary metric and adopt more comprehensive evaluation frameworks for responsible AI deployment.
Problem

Research questions and friction points this paper is trying to address.

facial recognition
algorithmic fairness
aggregate accuracy
demographic bias
law enforcement
Innovation

Methods, ideas, or system contributions that make the work stand out.

fairness evaluation
subgroup disparity
facial recognition
aggregate accuracy
model-agnostic auditing
🔎 Similar Papers
No similar papers found.