π€ AI Summary
This study addresses the limitations of traditional group-level fairness metrics in revealing individual unfairness and enabling attribution. To this end, it proposes an individual fairness evaluation framework based on Ranked Graduation Fairness (RGF). Methodologically, the work introduces the novel RGF metric and its integral measure, AURGF, employs the CramΓ©rβvon Mises test to compare predictive error distributions, and utilizes permutation testing alongside feature ablation for attribution analysis. Experiments on the HMDA dataset comparing models such as logistic regression and random forests demonstrate that tree-based models can effectively balance accuracy and fairness. Furthermore, the findings reveal that observed unfairness primarily stems from inherent group disparities within the data rather than isolated algorithmic bias, thereby achieving a unified integration of accuracy, fairness, and interpretability.
π Abstract
Fairness assessment in algorithmic decisions that affect individuals, such as credit scoring, often relies on parity measures calculated at the aggregate group level. Such measures may not reveal which individuals experience unfairness or which explanatory factors contribute to it. In this paper, we propose a rank-based framework that evaluates fairness through the distribution of model prediction errors, thereby linking fairness assessment with predictive accuracy and explainability. The framework combines Rank Graduation Fairness (RGF), its integrated measure AURGF, a centered Cramer--von Mises permutation test, and a feature removal procedure for fairness explainability.
We evaluate the methodology using logistic regression, random forest, gradient boosting, and a multilayer perceptron. The simulation study shows that protected-group imbalance can reverse descriptive fairness comparisons, whereas the proposed inferential procedure correctly distinguishes fair from unfair mechanisms. Its application to HMDA mortgage data produces model rankings that differ from those obtained with classical fairness criteria. Tree-based models, rather than logistic regression, provide the strongest combination of predictive accuracy and rank-based fairness, while the fairness null hypothesis is rejected for all four models. The persistence of disparity across statistical, bagging, boosting, and neural network specifications, together with the feature removal results, indicates that the observed unfairness is not specific to a single algorithm or predictor, but is associated with group differences embedded in the characteristics of the lending data. These findings support a broader approach to trustworthy artificial intelligence that combines predictive accuracy, fairness measurement, statistical inference, and explainability.