Cohort-attention Evaluation Metric against Tied Data: Studying Performance of Classification Models in Cancer Detection

📅 2025-03-17
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
Traditional cancer detection models suffer from evaluation bias due to class imbalance, inter-population performance heterogeneity, and inconsistent patient-level predictions. To address these limitations, we propose CAT (Cohort-Attention-based Evaluation Framework), the first framework introducing patient-level confidence aggregation and entropy-driven distribution weighting. CAT redefines cohort-weighted sensitivity (CATSen), specificity (CATSpe), and their harmonic mean (CATMean). Leveraging multi-center stratified evaluation and adaptive threshold recalibration, CAT enables fair, interpretable, and population-robust assessment of medical AI systems. In multi-center cancer screening tasks, CAT reduces false-negative misclassification rates by 12.7%, significantly enhancing cross-population performance comparability and clinical applicability.

Technology Category

Machine Learning: Calibration & Uncertainty QuantificationComputer Vision: Bias, Fairness & PrivacyCognitive Modeling & Cognitive Systems: Agent Architectures

Application Category

Search and Retrieval-Augmented AI: Web evaluation methodologies and metricsUser Modeling, Personalization and Recommendation: Fairness-aware retrieval and rankingEconomics, Online Markets and Human Computation: Fairness and ethical considerations in crowd work and in human-in-the-loop AI systems
📝 Abstract
Artificial intelligence (AI) has significantly improved medical screening accuracy, particularly in cancer detection and risk assessment. However, traditional classification metrics often fail to account for imbalanced data, varying performance across cohorts, and patient-level inconsistencies, leading to biased evaluations. We propose the Cohort-Attention Evaluation Metrics (CAT) framework to address these challenges. CAT introduces patient-level assessment, entropy-based distribution weighting, and cohort-weighted sensitivity and specificity. Key metrics like CATSensitivity (CATSen), CATSpecificity (CATSpe), and CATMean ensure balanced and fair evaluation across diverse populations. This approach enhances predictive reliability, fairness, and interpretability, providing a robust evaluation method for AI-driven medical screening models.
Problem

Research questions and friction points this paper is trying to address.

Addresses biased evaluation in cancer detection models
Proposes Cohort-Attention Evaluation Metrics (CAT) framework
Enhances fairness and reliability in medical screening
Innovation

Methods, ideas, or system contributions that make the work stand out.

Introduces Cohort-Attention Evaluation Metrics (CAT)
Uses entropy-based distribution weighting
Enhances fairness with cohort-weighted sensitivity
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
L
Longfei Wei
ThinkX
F
Fang Sheng
University of Toronto
J
Jianfei Zhang
ThinkX