🤖 AI Summary
This study addresses the limitation of existing selective classification methods that overlook the geometric structure of hard samples in the representation space, resulting in imprecise rejection boundaries. To overcome this, we propose a clustering-based, geometry-guided learning framework integrated with DistilBERT. By identifying confusion attractors within the representation space, the method leverages collective representational geometry to construct well-calibrated rejection boundaries and introduces a pre-deployment feasibility assessment metric. The proposed framework achieves simultaneous improvements in model performance and reliability, attaining an accuracy of 94.98% with a rejection rate below 9%. Furthermore, it effectively anticipates failure cases of baseline models prior to deployment, demonstrating its practical utility for enhancing predictive trustworthiness in real-world applications.
📝 Abstract
Selective classification enables a model to abstain from predictions on uncertain instances, but existing approaches typically reject them through confidence scores, predefined coverage constraints or instance-level distance measures. These approaches may overlook the collective geometric structure of difficult samples in learned representation spaces. We propose Guided Clustering-based Uncertain Learning (GCUL), a geometric-guided selective classification framework that identifies misclassified and ambiguous instances as a potential confusion attractor in the representation space. GCUL uses a three-phase procedure to initialize, cluster, and explicitly relabel this uncertain region, allowing the rejection boundary to emerge from the underlying representation geometry rather than from a prescribed rejection rate. We further derive a selectivity score and a geometric sufficient condition that characterizes when rejection can provide positive operational utility, enabling pre-deployment feasibility assessment. GCUL improves DistilBERT accuracy from 89.37 percent to 94.98 percent with less than 9 percent rejection. Beyond accuracy, our selectivity score correctly pre-detects the only dataset (GoEmotion) where all baselines fail, and controlled simulations yield 6.1 percent Type-I and 0 percent Type-II errors, validating the sufficient condition's conservatism. These results suggest that collective representation geometry provides a useful alternative perspective for selective prediction.