Uncertainty-aware Long-tailed Weights Model the Utility of Pseudo-labels for Semi-supervised Learning

๐Ÿ“… 2025-03-13
๐Ÿ“ˆ Citations: 0
โœจ Influential: 0
๐Ÿ“„ PDF
๐Ÿค– AI Summary
In semi-supervised learning, pseudo-label quality critically depends on a confidence threshold, yet threshold selection is challenging and models often exhibit overconfidence under limited calibration data, rendering confidence scores unreliable. To address this, we propose Uncertainty-Aware Ensemble Structure (UES), which eliminates hard thresholds and instead introduces a long-tailed weighting scheme to model pseudo-label utilityโ€”enabling even low-confidence labels to contribute robustly. UES is lightweight, architecture-agnostic, and seamlessly integrates with mainstream frameworks such as FixMatch, supporting both classification and regression tasks. On keypoint detection benchmarks (Sniffing, FLIC, LSP), UES improves PCK by 3.47โ€“7.29%; on CIFAR-10/100, it boosts test accuracy by 0.20โ€“0.26%. These results demonstrate significant gains in pseudo-label utilization efficiency and generalization performance.

Technology Category

Machine Learning: Calibration & Uncertainty QuantificationReasoning under Uncertainty: Uncertainty RepresentationsSearch and Optimization: Learning to Search

Application Category

Search and Retrieval-Augmented AI: Web learning to rank, online learning, and counterfactual learning for rankingUser Modeling, Personalization and Recommendation: Fairness-aware retrieval and rankingSemantics and Knowledge: Methods to enhance, augment, integrate or synergize semantic models such as knowledge graphs and LLMs
๐Ÿ“ Abstract
Current Semi-supervised Learning (SSL) adopts the pseudo-labeling strategy and further filters pseudo-labels based on confidence thresholds. However, this mechanism has notable drawbacks: 1) setting the reasonable threshold is an open problem which significantly influences the selection of the high-quality pseudo-labels; and 2) deep models often exhibit the over-confidence phenomenon which makes the confidence value an unreliable indicator for assessing the quality of pseudo-labels due to the scarcity of labeled data. In this paper, we propose an Uncertainty-aware Ensemble Structure (UES) to assess the utility of pseudo-labels for unlabeled samples. We further model the utility of pseudo-labels as long-tailed weights to avoid the open problem of setting the threshold. Concretely, the advantage of the long-tailed weights ensures that even unreliable pseudo-labels still contribute to enhancing the model's robustness. Besides, UES is lightweight and architecture-agnostic, easily extending to various computer vision tasks, including classification and regression. Experimental results demonstrate that combining the proposed method with DualPose leads to a 3.47% improvement in Percentage of Correct Keypoints (PCK) on the Sniffing dataset with 100 data points (30 labeled), a 7.29% improvement in PCK on the FLIC dataset with 100 data points (50 labeled), and a 3.91% improvement in PCK on the LSP dataset with 200 data points (100 labeled). Furthermore, when combined with FixMatch, the proposed method achieves a 0.2% accuracy improvement on the CIFAR-10 dataset with 40 labeled data points and a 0.26% accuracy improvement on the CIFAR-100 dataset with 400 labeled data points.
Problem

Research questions and friction points this paper is trying to address.

Addresses unreliable pseudo-label quality in semi-supervised learning.
Proposes uncertainty-aware ensemble to avoid threshold setting issues.
Enhances model robustness with long-tailed weights for pseudo-labels.
Innovation

Methods, ideas, or system contributions that make the work stand out.

Uncertainty-aware Ensemble Structure assesses pseudo-labels utility.
Long-tailed weights model avoids threshold setting problems.
Lightweight, architecture-agnostic UES enhances model robustness.
๐Ÿ”Ž Similar Papers
No similar papers found.