🤖 AI Summary
Traditional cancer detection models suffer from evaluation bias due to class imbalance, inter-population performance heterogeneity, and inconsistent patient-level predictions. To address these limitations, we propose CAT (Cohort-Attention-based Evaluation Framework), the first framework introducing patient-level confidence aggregation and entropy-driven distribution weighting. CAT redefines cohort-weighted sensitivity (CATSen), specificity (CATSpe), and their harmonic mean (CATMean). Leveraging multi-center stratified evaluation and adaptive threshold recalibration, CAT enables fair, interpretable, and population-robust assessment of medical AI systems. In multi-center cancer screening tasks, CAT reduces false-negative misclassification rates by 12.7%, significantly enhancing cross-population performance comparability and clinical applicability.
📝 Abstract
Artificial intelligence (AI) has significantly improved medical screening accuracy, particularly in cancer detection and risk assessment. However, traditional classification metrics often fail to account for imbalanced data, varying performance across cohorts, and patient-level inconsistencies, leading to biased evaluations. We propose the Cohort-Attention Evaluation Metrics (CAT) framework to address these challenges. CAT introduces patient-level assessment, entropy-based distribution weighting, and cohort-weighted sensitivity and specificity. Key metrics like CATSensitivity (CATSen), CATSpecificity (CATSpe), and CATMean ensure balanced and fair evaluation across diverse populations. This approach enhances predictive reliability, fairness, and interpretability, providing a robust evaluation method for AI-driven medical screening models.