🤖 AI Summary
This study addresses the dual-mismatch problem in semi-supervised learning caused by imbalanced class distributions and the intrusion of unknown classes within unlabeled data. We propose a unified framework integrating a hub-spoke geometric structure with an evidential classifier. Specifically, the latent space is organized into a hub-spoke topology: known-class features are uniformly distributed across peripheral nodes to enhance discriminability, while unknown-class samples are aggregated at the central hub based on low evidential support for effective isolation. This structured feature organization suppresses majority-class dominance bias and substantially improves pseudo-label quality. Extensive experiments demonstrate that the proposed method consistently outperforms existing state-of-the-art approaches across various settings, achieving performance gains of up to 3.25%.
📝 Abstract
Semi-supervised learning typically assumes that labeled and unlabeled data share an identical class distribution and label space. However, this setting is often violated: unlabeled data may be imbalanced and contain unknown class samples, causing mismatches in both class distribution and label space. Such dual mismatch leads to majority classes dominating the latent space and unknown class samples being overconfidently misclassified, degrading feature discriminability and pseudo-label quality. To address this, we propose a hub-spoke latent geometry, where known classes are uniformly distributed around a central hub and each class forms compact clusters around its prototype, while the hub provides an anchor for a low-evidence region specifically designed for high-uncertainty unknown class samples. Integrated with an evidence-based classifier, this geometry ultimately enhances feature discriminability and uncertainty separation by mitigating majority-class domination through structured feature organization and guiding high-uncertainty unknown class samples toward the hub. Extensive experiments show that our method outperforms state-of-the-art methods, with a maximum improvement of 3.25% across various settings.