Score
Designs and conducts systematic equity impact assessments that identify affected and vulnerable subgroups, quantify distributional and capability harms, and compare equity outcomes across policies or interventions; produces actionable findings and recommends mitigation or corrective measures to reduce unequal burdens and improve fairness.
Traditional health evaluation metrics, such as QALYs and PALYs, struggle to simultaneously account for equity and productivity contributions. This work proposes a unified framework that integrates welfare economics and health measurement theory through a normative axiomatic approach. Within this framework, a new class of evaluation functions is developed, satisfying scale invariance and the Pigou-Dalton transfer principle. The authors derive a tractable power-form representation of these functions, offering a coherent basis for assessing interventions that jointly affect health and productive capacity. By doing so, the framework overcomes key limitations of existing metrics, providing a principled balance between efficiency and fairness in health policy evaluation.
This study addresses the critical gap in clinical machine learning fairness evaluation by systematically applying an intersectional fairness auditing framework to real-world clinical prediction tasks. Leveraging the All of Us dataset, the authors integrate the FairLogue toolkit, observational fairness metrics, and counterfactual causal analysis to assess model performance across intersecting subgroups defined by race and gender. Their findings reveal substantial performance disparities that remain undetected under conventional single-axis fairness assessments. However, counterfactual experiments demonstrate that most of these disparities persist even after randomizing group identity, indicating that they primarily stem from differences in covariate distributions rather than direct discrimination. These results underscore the necessity and value of intersectional auditing for accurately diagnosing and addressing health inequities in clinical AI systems.
This study identifies a significant coverage gap and collaboration deficit between first-party (developer-led) and third-party (academic, NGO, etc.) AI social impact assessments—spanning bias, fairness, privacy, environmental cost, and labor practices. Methodologically, it conducts the first systematic comparative analysis of 186 first-party reports and 183 third-party evaluations, integrating content analysis, quantitative statistics, and in-depth interviews with AI developers. Results reveal persistent under-disclosure by first parties on critical issues, while third-party assessments—though deeper—remain constrained by data opacity and lack of access to proprietary infrastructure and internal documentation. The core contribution is the empirical identification of structural imbalances within the AI assessment ecosystem. The paper proposes a shared infrastructure framework to integrate independent evaluations, enhance verifiability, and strengthen accountability—thereby establishing a methodological foundation and actionable policy pathway for robust AI governance.
Underrepresentation in public health data risks systemic bias, undermining the fairness and validity of downstream inference and policy decisions. To address this, we propose an operational definition of “public health data fairness,” integrating computational principles—fairness, accountability, transparency, ethics, and privacy—with core public health methodologies—including selection bias correction, representativeness assessment, and causal inference. This yields a structured, lifecycle-spanning self-audit framework grounded in reflexive practice and designed for seamless integration into routine data science workflows. Validated across multiple real-world public health applications, the framework demonstrably enhances the equitable applicability of AI and data-driven policies across diverse populations. Crucially, our analysis clarifies that data fairness constitutes a necessary—but not sufficient—condition for fair decision-making.
This study addresses the lack of systematic evaluation of robustness in existing fair machine learning methods under realistic data perturbations such as label noise, missing data, and distribution shifts. It introduces a causal inference framework to conduct the first comprehensive robustness analysis of mainstream fairness interventions—including sensitive attribute handling and bias mitigation techniques—under non-ideal data conditions. Empirical results demonstrate that several widely used approaches suffer significant performance degradation under common perturbations, thereby exposing critical limitations for real-world deployment. These findings provide both theoretical grounding and practical guidance for developing more reliable and robust fair machine learning systems.
This study addresses the limitations of prevailing machine learning fairness frameworks, which predominantly emphasize distributive fairness and struggle to accommodate the diverse justice concerns of multiple stakeholders in algorithmic systems. To bridge this gap, the authors introduce organizational justice theory into algorithm design for the first time. Through co-design workshops with Kiva staff, complemented by qualitative interviews and theoretical analysis, they identify normative concerns about personalized recommendation systems across different organizational units. Building on these insights, they develop an actionable and interpretable set of organizational justice evaluation metrics. This framework not only facilitates trade-offs among multidimensional justice objectives and informs system configuration but also fosters cross-departmental normative dialogue regarding algorithmic deployment. The proposed metrics have been successfully implemented in Kiva’s microlending platform.
This study addresses the systemic measurement bias in pulse oximeters across racial groups, which leads to inequitable health decisions. To tackle this issue, the authors propose a data fairness analysis framework that explicitly decomposes fairness into three actionable dimensions: data, prediction, and decision. Leveraging oracle-based causal simulation and counterfactual attribution analysis, they develop a statistical provenance model that traces how upstream informational biases propagate and amplify into clinical disparities and adverse health outcomes. By translating abstract notions of fairness into empirically testable statistical metrics, the framework establishes a reproducible analytical paradigm for identifying and mitigating such systemic inequities, thereby underscoring the pivotal role of statistics in AI-driven healthcare.
This study addresses treatment disparities arising from racial bias in healthcare and other domains by proposing a novel Target Study + Target Trial (TS+TT) framework for rigorously evaluating interventions’ impacts on intergroup disparities. Methodologically, it achieves cross-group covariate balance and unbiased causal effect estimation via stratified sampling and within-stratum randomization, and extends semi-parametric G-computation to continuous-time survival outcomes to enable counterfactual disparity analysis under continuous interventions. Its key contribution lies in the first unified modeling of ethics-informed disparity metrics with causally identified effect estimates—thereby jointly ensuring fairness, interpretability, and statistical rigor. Empirical validation on electronic health record data simulated the impact of racially biased pulse oximetry on disparities in treatment receipt, demonstrating the framework’s robustness across diverse intervention types and outcome structures, as well as its policy relevance for equitable healthcare decision-making.
Existing fairness tools are often limited to single demographic attributes and struggle to capture the compounded biases faced by intersecting groups—such as combinations of race and gender—in clinical machine learning. This work proposes a Python toolkit that extends observational fairness metrics, including demographic parity and equalized odds, to intersectional subgroups for the first time, while integrating two counterfactual fairness frameworks to evaluate intervention-based equity. Applied to electronic health record data using logistic regression in a glaucoma surgery prediction task, the approach uncovers substantial intersectional unfairness, with a demographic parity gap as high as 0.20. Crucially, disparities identified through intersectional analysis markedly exceed those detected by single-dimension assessments, underscoring the necessity and efficacy of this method for auditing fairness in clinical algorithms.
This work addresses the limitations of existing fairness methods, which often focus on a single demographic attribute and lack systematic evaluation across intersecting subgroups and multiple stages of the modeling pipeline. To bridge this gap, we propose FairSelect, a novel toolkit that establishes the first multi-level evaluation framework enabling arbitrary combinations of pre-, in-, and post-processing fairness interventions. We conduct comprehensive analyses of fairness–utility trade-offs across diverse model architectures and intersectional subgroups using both synthetic clinical data and a real-world atrial fibrillation stroke risk prediction task. Our experiments demonstrate that combined intervention strategies generally enhance fairness with controllable utility loss; notably, certain combinations simultaneously improve both fairness and predictive performance, while others yield adverse effects, revealing non-additive and context-dependent interactions among fairness interventions in intersectional settings.