Score
Designs, implements, and evaluates methods and tools to detect, measure, decompose, and correct systematic errors and unfair disparities in datasets, models, and model outputs. This includes building diagnostic metrics and measurements (e.g., group‑wise performance gaps, bias‑variance decomposition, attribution of token contributions), detection and quantification procedures for representational or prediction skew, bias‑correction and mitigation techniques (e.g., reweighting, filtering, algorithmic adjustments), and evaluation pipelines that quantify remaining harms and tradeoffs between overall performance and fairness.
Ambiguity in fairness metrics and poor cross-cultural/legal adaptability hinder effective AI regulation. Method: This paper proposes the first context-aware fairness metric selection framework designed for regulatory implementation. It integrates philosophical, cultural, legal, and technical perspectives, formalizing a flowchart-based decision model grounded in 12 criteria. Empirical validation is conducted via interdisciplinary literature review and regulatory text mapping—specifically against the EU AI Act and NIST AI Risk Management Framework (AI RMF). Contribution/Results: The framework systematically bridges the gap between theoretical fairness concepts and regulatory compliance practice. It delivers an actionable, scenario-specific, and multi-stakeholder-oriented guidance tool for selecting fairness metrics, thereby enhancing the rigor, interpretability, and regulatory alignment of fairness assessments in machine learning systems.
The widespread deployment of machine learning (ML) in decision-making systems introduces significant fairness risks—particularly concerning the handling of sensitive attributes and the protection of minority groups—while software engineering lacks a systematic, lifecycle-oriented framework for fairness engineering practices. Method: We conduct a systematic mapping study (SMS) combined with a comprehensive literature review to analyze fairness-related practices across the ML development lifecycle. Contribution/Results: We propose the first software engineering–centric fairness practice taxonomy, comprising 28 structured, actionable practices explicitly mapped to data preprocessing, modeling, and deployment stages. Each practice is annotated with its corresponding ML lifecycle phase and contextual applicability, thereby bridging the gap between fairness research and industrial implementation. This taxonomy serves as an integrable, operational guide for researchers and practitioners, enhancing the reliability, accountability, and trustworthiness of ML systems.
Post-processing debiasing methods may inadvertently introduce new forms of unfairness—particularly through overcorrection caused by imbalanced prediction flips across demographic groups. To address this, we propose “Flip Disparity,” a novel metric suite that quantifies, for the first time in post-processing, the relative proportion of predictions flipped per group, thereby overcoming limitations of conventional fairness metrics that ignore transparency and proportionality in correction behavior. Our method leverages differences in confusion matrices and inter-group comparative analysis, integrated within a unified framework combining visual diagnostic tools and strategy-comparability assessment. This paradigm significantly enhances the interpretability of debiasing strategies and enables reliable detection of latent imbalanced corrections. Empirical evaluation across multiple benchmark datasets reveals previously undetected correction biases in widely adopted fairness algorithms. The proposed framework establishes a verifiable, auditable standard for responsible algorithmic governance.
This study addresses the lack of systematic evaluation of robustness in existing fair machine learning methods under realistic data perturbations such as label noise, missing data, and distribution shifts. It introduces a causal inference framework to conduct the first comprehensive robustness analysis of mainstream fairness interventions—including sensitive attribute handling and bias mitigation techniques—under non-ideal data conditions. Empirical results demonstrate that several widely used approaches suffer significant performance degradation under common perturbations, thereby exposing critical limitations for real-world deployment. These findings provide both theoretical grounding and practical guidance for developing more reliable and robust fair machine learning systems.
Machine learning models exhibit high sensitivity to minor perturbations in training data, leading to unstable predictions; yet conventional fairness metrics (e.g., bias-based indicators) ignore this prediction uncertainty. Method: We propose a variance-oriented paradigm for group fairness—introducing the first systematic framework that treats inter-group predictive variance equality as a core fairness criterion, grounded in statistical error decomposition and theoretical analysis of variance’s independent impact on fairness assessment. Contribution/Results: We release VarFair, the first open-source library integrating uncertainty quantification with fairness evaluation. Extensive experiments on Adult, COMPAS, and other benchmarks demonstrate that groups with high predictive variance are frequently misclassified as “fair” by standard methods, whereas our variance-aware metric significantly improves identification of disadvantaged groups and enhances assessment robustness under data perturbations.
This paper addresses the misalignment between algorithmic bias assessment and legal standards by proposing a quantification framework rigorously grounded in U.S. anti-discrimination law. Methodologically, it distinguishes legally salient discriminatory testing from systemic disparity through legal contextualization, and introduces the Objective Fairness Index (OFI)—a metric integrating objective test theory and measurement stability, using marginal benefit as a proxy to quantify legal compliance of algorithmic decisions. Its key contribution lies in being the first fairness metric to embed legal admissibility directly into its design, enabling a paradigm shift in algorithmic auditing from statistical fairness to legally grounded fairness. Empirical evaluation on real-world judicial prediction systems—including COMPAS—demonstrates that OFI reliably detects unlawful discrimination, offering regulators and auditors the first quantitative tool with both legal interpretability and operational utility.
This work addresses the limitations of existing fairness methods, which often focus on a single demographic attribute and lack systematic evaluation across intersecting subgroups and multiple stages of the modeling pipeline. To bridge this gap, we propose FairSelect, a novel toolkit that establishes the first multi-level evaluation framework enabling arbitrary combinations of pre-, in-, and post-processing fairness interventions. We conduct comprehensive analyses of fairness–utility trade-offs across diverse model architectures and intersectional subgroups using both synthetic clinical data and a real-world atrial fibrillation stroke risk prediction task. Our experiments demonstrate that combined intervention strategies generally enhance fairness with controllable utility loss; notably, certain combinations simultaneously improve both fairness and predictive performance, while others yield adverse effects, revealing non-additive and context-dependent interactions among fairness interventions in intersectional settings.
This study addresses the critical gap in clinical machine learning fairness evaluation by systematically applying an intersectional fairness auditing framework to real-world clinical prediction tasks. Leveraging the All of Us dataset, the authors integrate the FairLogue toolkit, observational fairness metrics, and counterfactual causal analysis to assess model performance across intersecting subgroups defined by race and gender. Their findings reveal substantial performance disparities that remain undetected under conventional single-axis fairness assessments. However, counterfactual experiments demonstrate that most of these disparities persist even after randomizing group identity, indicating that they primarily stem from differences in covariate distributions rather than direct discrimination. These results underscore the necessity and value of intersectional auditing for accurately diagnosing and addressing health inequities in clinical AI systems.
Current model evaluations often rely on aggregate metrics that obscure performance disparities and unfairness across continuous or fine-grained subpopulations. This work proposes FairTree, an algorithm that introduces bias-variance decomposition into fairness auditing for the first time, drawing inspiration from measurement invariance in psychometrics to handle continuous, categorical, and ordinal attributes without requiring discretization. By integrating permutation tests with fluctuation tests, FairTree flexibly models subpopulation performance variation and enables rigorous statistical inference. Empirical results demonstrate that FairTree effectively controls false positive rates, with the fluctuation test exhibiting superior statistical power, and its practical utility is validated on the UCI Adult Census dataset.
This study addresses the inconsistency among fairness metrics in face recognition, where different measures often yield contradictory conclusions about model bias, thereby exposing the limitations of single-metric evaluation. To tackle this issue, the authors propose the Fairness Disagreement Index (FDI) to quantify the degree of disagreement across multiple fairness criteria and introduce a multidimensional evaluation framework that integrates both error rate disparities and performance-oriented fairness metrics. Through systematic experiments under controlled conditions, they demonstrate that such metric disagreement is pervasive across varying decision thresholds and model configurations, revealing a critical flaw in current fairness assessment practices. The work provides both a novel analytical tool and empirical evidence to support more comprehensive and reliable fairness evaluations in face recognition systems.
This study addresses how infra-marginality—differences in data distributions across groups—complicates judgments of AI fairness, as conventional statistical parity metrics often fail to align with human perceptions of fairness. Through a controlled user study involving 85 participants in a hypothetical medical decision-making scenario, the authors systematically investigate how group-specific model performance and training data availability shape fairness judgments. They find that when group-wise performance is equal or unknown, participants favor outcome equality; however, when performance disparities are attributable to data imbalance, models preserving these differences are perceived as more fair. These results demonstrate that human fairness judgments are not solely based on outcome equality but are significantly influenced by beliefs about the underlying causes of disparities, thereby challenging the prevailing assumption that statistical parity should serve as the default standard for algorithmic fairness.