🤖 AI Summary
This work addresses combinatorial bias in classification evaluation metrics under few-shot settings, where group-size disparities induce spurious fairness assessments. Through probabilistic modeling and combinatorial analysis, we systematically identify and quantify the latent bias mechanisms inherent in common metrics—including accuracy and F1-score—under sample-size imbalance. We propose a model-agnostic framework for detecting and correcting such bias, unifying treatment of undefined cases (e.g., zero-denominator scenarios) and metric sensitivity to class distribution. Our approach substantially enhances discriminative power in fairness evaluation under data scarcity, mitigating erroneous attribution and misguided policy interventions arising from metric distortion. The framework provides both theoretical grounding and practical tools for trustworthy AI assessment in resource-constrained environments.
📝 Abstract
Evaluating machine learning models is crucial not only for determining their technical accuracy but also for assessing their potential societal implications. While the potential for low-sample-size bias in algorithms is well known, we demonstrate the significance of sample-size bias induced by combinatorics in classification metrics. This revelation challenges the efficacy of these metrics in assessing bias with high resolution, especially when comparing groups of disparate sizes, which frequently arise in social applications. We provide analyses of the bias that appears in several commonly applied metrics and propose a model-agnostic assessment and correction technique. Additionally, we analyze counts of undefined cases in metric calculations, which can lead to misleading evaluations if improperly handled. This work illuminates the previously unrecognized challenge of combinatorics and probability in standard evaluation practices and thereby advances approaches for performing fair and trustworthy classification methods.