Score
Designs, builds, and evaluates methods, metrics, and tools to detect, quantify, and mitigate unfairness and disparate impacts in machine learning models and datasets; analyzes subgroup performance, sources of bias in data and algorithms, and trade-offs between fairness criteria and other objectives, and implements fairness-aware training, auditing, and reporting procedures.
Ambiguity in fairness metrics and poor cross-cultural/legal adaptability hinder effective AI regulation. Method: This paper proposes the first context-aware fairness metric selection framework designed for regulatory implementation. It integrates philosophical, cultural, legal, and technical perspectives, formalizing a flowchart-based decision model grounded in 12 criteria. Empirical validation is conducted via interdisciplinary literature review and regulatory text mapping—specifically against the EU AI Act and NIST AI Risk Management Framework (AI RMF). Contribution/Results: The framework systematically bridges the gap between theoretical fairness concepts and regulatory compliance practice. It delivers an actionable, scenario-specific, and multi-stakeholder-oriented guidance tool for selecting fairness metrics, thereby enhancing the rigor, interpretability, and regulatory alignment of fairness assessments in machine learning systems.
The widespread deployment of machine learning (ML) in decision-making systems introduces significant fairness risks—particularly concerning the handling of sensitive attributes and the protection of minority groups—while software engineering lacks a systematic, lifecycle-oriented framework for fairness engineering practices. Method: We conduct a systematic mapping study (SMS) combined with a comprehensive literature review to analyze fairness-related practices across the ML development lifecycle. Contribution/Results: We propose the first software engineering–centric fairness practice taxonomy, comprising 28 structured, actionable practices explicitly mapped to data preprocessing, modeling, and deployment stages. Each practice is annotated with its corresponding ML lifecycle phase and contextual applicability, thereby bridging the gap between fairness research and industrial implementation. This taxonomy serves as an integrable, operational guide for researchers and practitioners, enhancing the reliability, accountability, and trustworthiness of ML systems.
This study addresses the lack of systematic evaluation of robustness in existing fair machine learning methods under realistic data perturbations such as label noise, missing data, and distribution shifts. It introduces a causal inference framework to conduct the first comprehensive robustness analysis of mainstream fairness interventions—including sensitive attribute handling and bias mitigation techniques—under non-ideal data conditions. Empirical results demonstrate that several widely used approaches suffer significant performance degradation under common perturbations, thereby exposing critical limitations for real-world deployment. These findings provide both theoretical grounding and practical guidance for developing more reliable and robust fair machine learning systems.
This study systematically investigates fairness in computer vision (CV) and natural language processing (NLP) models operating on unstructured data, with a focus on how algorithmic bias exacerbates systemic inequities. Method: Leveraging real-world Kaggle datasets, we construct end-to-end ML pipelines and conduct the first empirical comparison of two leading fairness toolkits—Fairlearn (Microsoft) and AI Fairness 360 (IBM)—across CV and NLP tasks, evaluating their metric coverage and bias mitigation efficacy. Contribution/Results: Fairlearn excels in interpretability and engineering integration, whereas AIF360 offers broader multidimensional fairness metrics. Combining preprocessing and postprocessing techniques reduces bias by 32–47% on average. We further propose a three-tier industrial fairness governance framework, providing both methodological guidance and empirical benchmarks for cross-modal AI fairness assessment and deployment.
This study addresses the challenge of operationalizing fairness requirements in the AI software development lifecycle (SDLC): although practitioners widely acknowledge AI fairness as critical, it is routinely deprioritized amid functional delivery pressures and tight deadlines—exposing three core bottlenecks: ambiguous definitions, absent metrics, and missing process integration. Through cross-cultural, semi-structured interviews with 26 AI practitioners across 26 countries and thematic qualitative analysis, we systematically identify real-world fairness gaps across SDLC phases—requirements elicitation, modeling, verification, and trade-off negotiation—from a software engineering perspective. Our key contribution is a framework advocating co-defined, context-sensitive fairness metrics involving diverse stakeholders, and their formal integration into SDLC artifacts and workflows—thereby establishing a foundation for auditable, evolvable, fairness-aware AI engineering practice.
Existing fairness research in machine learning predominantly focuses on algorithmic outcomes, neglecting the sociotechnical processes underlying system development and deployment—particularly how stakeholders subjectively perceive procedural fairness. Method: Addressing this gap, this study systematically integrates procedural and distributive justice theories to develop a tri-dimensional operational framework for perceived fairness—encompassing transparency, accountability, and representativeness—tailored to both developers and end users. We employed virtual focus groups, systematic literature review, theoretical modeling, and rigorous scale development with psychometric validation (including reliability and construct validity testing). Contribution/Results: We introduce a theoretically grounded, empirically validated Perceived Fairness Scale for ML systems, supported by cross-role (developer/user) evidence. This instrument provides a measurable, actionable tool for designing, evaluating, and governing fair ML systems, advancing human-AI collaboration and sociotechnical governance.
This work addresses the limitations of existing fairness methods, which often focus on a single demographic attribute and lack systematic evaluation across intersecting subgroups and multiple stages of the modeling pipeline. To bridge this gap, we propose FairSelect, a novel toolkit that establishes the first multi-level evaluation framework enabling arbitrary combinations of pre-, in-, and post-processing fairness interventions. We conduct comprehensive analyses of fairness–utility trade-offs across diverse model architectures and intersectional subgroups using both synthetic clinical data and a real-world atrial fibrillation stroke risk prediction task. Our experiments demonstrate that combined intervention strategies generally enhance fairness with controllable utility loss; notably, certain combinations simultaneously improve both fairness and predictive performance, while others yield adverse effects, revealing non-additive and context-dependent interactions among fairness interventions in intersectional settings.
Current model evaluations often rely on aggregate metrics that obscure performance disparities and unfairness across continuous or fine-grained subpopulations. This work proposes FairTree, an algorithm that introduces bias-variance decomposition into fairness auditing for the first time, drawing inspiration from measurement invariance in psychometrics to handle continuous, categorical, and ordinal attributes without requiring discretization. By integrating permutation tests with fluctuation tests, FairTree flexibly models subpopulation performance variation and enables rigorous statistical inference. Empirical results demonstrate that FairTree effectively controls false positive rates, with the fluctuation test exhibiting superior statistical power, and its practical utility is validated on the UCI Adult Census dataset.
This study addresses the critical gap in clinical machine learning fairness evaluation by systematically applying an intersectional fairness auditing framework to real-world clinical prediction tasks. Leveraging the All of Us dataset, the authors integrate the FairLogue toolkit, observational fairness metrics, and counterfactual causal analysis to assess model performance across intersecting subgroups defined by race and gender. Their findings reveal substantial performance disparities that remain undetected under conventional single-axis fairness assessments. However, counterfactual experiments demonstrate that most of these disparities persist even after randomizing group identity, indicating that they primarily stem from differences in covariate distributions rather than direct discrimination. These results underscore the necessity and value of intersectional auditing for accurately diagnosing and addressing health inequities in clinical AI systems.
This study addresses fairness deficiencies in machine learning–based early warning systems used by higher education institutions for allocating student support resources, particularly with respect to disparities arising from gender, age, and residency status. Through a long-term collaboration with Centennial College, the authors replicate the institution’s deployed system and develop the first reproducible auditing framework that integrates construct validity with statistical fairness metrics to systematically evaluate the entire pipeline—from data collection and prediction to post-processing. Their analysis reveals that younger, male, and international students are systematically assigned higher risk scores than their actual risk levels warrant, while older and female students with equivalent risk profiles are consistently underestimated. Notably, bias is significantly amplified during the post-processing stage. This work provides both methodological innovation and empirical evidence to advance fairness auditing of institutionalized machine learning systems.
Existing training monitoring tools struggle to simultaneously track dynamic changes across multiple metrics and diagnose fairness disparities among subgroups. This work proposes a TensorBoard plugin that, for the first time, integrates multi-metric linked visualizations with slice-level fairness analysis within a unified interactive interface, enabling real-time monitoring of both performance and fairness without modifying the training pipeline. By combining multi-view charts, user-defined subgroup slicing, standard fairness metrics, and correlation analysis across heterogeneous indicators, the approach successfully uncovers hidden demographic and environmental biases in high-performing models on the YOLOX architecture and BDD100k dataset, facilitating early detection and diagnosis of fairness issues during model training.