bayesian pseudo-labeling

Designs and implements algorithms and pipelines that generate, calibrate, and iteratively refine pseudo-labels from model outputs by converting continuous scores into reliable binary or structured labels; these methods integrate Bayesian uncertainty estimation, confidence/evidence calibration, ensembles and multi-model or cross-modal agreement, and priors such as physics, geometry, or equivariance to reduce hallucinations and false positives. Builds unsupervised or self-training workflows (including iterative pseudo-labeling, pseudo-label refinement, cross-model/cross-modal pseudo-annotation, and SAM-guided or content-based strategies) and analyzes label quality, uncertainty calibration, and the impact of pseudo-labeling on downstream training robustness.

bayesianpseudo-labeling

Recent Skill Trend

Momentum and market value over time
Trending
Score
No comparison yet
0.14
Oct 01, 2026Oct 01, 2026
Career
Value
No comparison yet
$200K/year
Oct 01, 2026Oct 01, 2026

Recommended Survey Paper

Quick overview of the field
View more

Must-Read Papers

Most classic and influential ideas
View more

In semi-supervised learning, pseudo-label quality critically depends on a confidence threshold, yet threshold selection is challenging and models often exhibit overconfidence under limited calibration data, rendering confidence scores unreliable. To address this, we propose Uncertainty-Aware Ensemble Structure (UES), which eliminates hard thresholds and instead introduces a long-tailed weighting scheme to model pseudo-label utility—enabling even low-confidence labels to contribute robustly. UES is lightweight, architecture-agnostic, and seamlessly integrates with mainstream frameworks such as FixMatch, supporting both classification and regression tasks. On keypoint detection benchmarks (Sniffing, FLIC, LSP), UES improves PCK by 3.47–7.29%; on CIFAR-10/100, it boosts test accuracy by 0.20–0.26%. These results demonstrate significant gains in pseudo-label utilization efficiency and generalization performance.

Addresses unreliable pseudo-label quality in semi-supervised learning.Enhances model robustness with long-tailed weights for pseudo-labels.Proposes uncertainty-aware ensemble to avoid threshold setting issues.

Pseudo-D: Informing Multi-View Uncertainty Estimation with Calibrated Neural Training Dynamics

Sep 15, 2025
AN
Ang Nan Gu
🏛️ University of British Columbia | Vancouver General Hospital

In medical image diagnosis, one-hot labels obscure inter-expert annotation discrepancies and intrinsic image ambiguity, leading to overconfident and poorly calibrated models with limited robustness. To address this, we propose Uncertainty-aware Pseudo-Labeling (UPL), a method that dynamically models sample difficulty by leveraging prediction trajectories during neural network training, thereby generating calibrated pseudo-labels that explicitly encode diagnostic uncertainty and inter-rater disagreement. UPL injects uncertainty directly into the supervision signal without modifying the network architecture, supports multi-view inputs, and enables end-to-end label enhancement. Evaluated on echocardiogram classification, UPL significantly improves model calibration, selective classification performance, and robustness to multi-view fusion—outperforming state-of-the-art uncertainty modeling and label smoothing baselines.

Addressing diagnostic uncertainty in medical image classificationEnhancing model calibration and robustness in noisy dataGenerating uncertainty-aware labels using training dynamics

Semi-Supervised Regression with Heteroscedastic Pseudo-Labels

Oct 16, 2025
XS
Xueqing Sun
🏛️ Xi'an Jiaotong University | City University of Hong Kong | Harvard University | Xidian University | Pazhou Laboratory

In semi-supervised regression, pseudo-labels are often corrupted by heteroscedastic noise, leading to unreliable confidence estimation, error accumulation, and overfitting. To address this, we propose an uncertainty-aware pseudo-labeling framework—the first to systematically integrate uncertainty modeling into pseudo-label generation for semi-supervised regression. Our method employs a bilevel optimization scheme that jointly learns the regression model and a heteroscedastic noise estimator: the upper level calibrates pseudo-label confidence via uncertainty quantification, while the lower level optimizes the prediction model using confidence-weighted empirical risk minimization. This enables adaptive noise calibration and significantly improves generalization. Extensive experiments on multiple benchmark datasets demonstrate that our approach surpasses existing state-of-the-art methods in both robustness and prediction accuracy. The implementation is publicly available.

Addressing heteroscedastic noise in semi-supervised regression pseudo-labelsDeveloping uncertainty-aware framework for robust pseudo-label influence adjustmentMitigating error accumulation from unreliable continuous pseudo-labels

Feedback-Driven Pseudo-Label Reliability Assessment: Redefining Thresholding for Semi-Supervised Semantic Segmentation

May 12, 2025
NG
Negin Ghamsarian
🏛️ University of Bern | University of Klagenfurt

In semi-supervised semantic segmentation, conventional pseudo-labeling relies on manually preset confidence thresholds, which are suboptimal under scarce labeled data. To address this, we propose a dynamic feedback-driven reliability assessment framework. Our method introduces a novel class-aware true-positive confidence estimation mechanism, integrating multi-teacher ensemble confidence modeling, online class-conditional true-positive rate estimation, adaptive threshold updating, and response-feedback reinforcement within a closed loop—enabling threshold-free, self-adaptive pseudo-label selection. This framework departs from static threshold paradigms and demonstrates consistent improvements across PASCAL VOC and Cityscapes benchmarks with various backbone architectures: it achieves up to a 3.2% mIoU gain under extreme label scarcity, without incurring additional annotation cost or computational overhead.

Dynamic thresholding for pseudo-label selection in semi-supervised learningEnhancing segmentation performance in data-scarce semi-supervised scenariosOvercoming reliance on pre-defined confidence thresholds in pseudo-supervision

This work addresses a critical limitation in existing reconstruction-based methods for learning with noisy labels: their tendency to jointly assess the reliability of observed labels and pseudo-targets, which often leads to unreliable signals substituting one another and hinders effective denoising. To overcome this, the paper proposes TRACE, a novel framework that decouples the reliability evaluation of these two sources for the first time. Specifically, it evaluates observed labels through loss fitting, shallow-feature relational stability, and prediction consistency, while assessing pseudo-targets via model confidence. These independent reliability estimates are then used to separately govern label correction and sample reweighting. By preventing error propagation, TRACE generates more trustworthy pseudo-supervision, significantly outperforming current reconstruction-based approaches across multiple synthetic and real-world noisy benchmarks, and thereby enhancing model robustness and generalization.

label-noise learningnoisy labelspseudo targets

Latest Papers

What's happening recently
View more

This work addresses the challenge of applying multicalibration in weakly supervised learning settings—such as positive-unlabeled or unlabeled-unlabeled classification—where the absence of clean labels renders traditional multicalibration methods inapplicable. The paper presents the first extension of multicalibration to such weak supervision scenarios through a unified framework that models label noise via a corruption matrix and rewrites the risk to estimate multicalibration error. By introducing calibration constraints based on witness functions, the authors develop WLMC, a moment-based estimator with finite-sample guarantees and a general-purpose post-processing algorithm. Empirical evaluations demonstrate that the proposed method substantially improves the reliability of predicted probabilities and achieves strong multicalibration performance across diverse weakly supervised settings.

label scarcitymulticalibrationpost-hoc calibration

This work proposes a novel representation learning framework that addresses the limited representational capacity of existing methods in complex scenes by integrating adaptive multi-scale fusion with contrastive learning. The approach dynamically aggregates multi-level features and incorporates a structure-aware contrastive loss, thereby enhancing the model’s ability to jointly capture fine-grained semantics and global contextual information. Extensive experiments demonstrate that the proposed framework consistently outperforms state-of-the-art methods across multiple benchmark datasets, achieving substantial improvements in both accuracy and robustness. These results establish a promising new direction for unsupervised and semi-supervised representation learning.

attenuation biascalibrationconfidence thresholding

This work addresses the performance degradation of semi-supervised learning in practical settings caused by out-of-distribution (OOD) samples within unlabeled data. To mitigate this issue, the authors propose Uncertainty-based Structure Estimation (USE), which reframes data quality control as a structural informativeness assessment. Specifically, a lightweight proxy model computes the entropy of unlabeled samples, and a threshold derived from statistical hypothesis testing is employed to retain only those samples exhibiting meaningful structural information while discarding harmful or uninformative ones. The method is algorithm-agnostic and computationally efficient, consistently improving model accuracy and robustness across varying levels of OOD contamination on benchmarks such as CIFAR-100 and Yelp Review. These results underscore the critical role of effective data filtering in enhancing the reliability of semi-supervised learning.

data curationout-of-distributionrobustness

This work addresses the limited generalization capability in high-definition map construction caused by scarce annotated data by proposing a semi-supervised learning approach based on a teacher–student framework. A teacher model trained on a small set of labeled data generates fine-grained pseudo-labels by modeling temporal observation confidence via a Beta distribution and preserving high-confidence regions through a spatial cropping mechanism. These refined pseudo-labels, combined with an optimized map prior, guide the training of the student model. Unlike conventional coarse-grained strategies that discard entire elements, the proposed method retains informative regions, significantly improving performance under low-label regimes. Evaluated on the nuScenes dataset, the approach achieves a 6.1 mAP gain using only minimal labeled data, effectively alleviating dependence on extensive annotations.

data scarcityHD map constructiononline mapping

Hot Scholars

TH

Thanh-Huy Nguyen

Carnegie Mellon University
Medical Image Analysis𝗖𝗼𝗺𝗽𝘂𝘁𝗲𝗿 𝗩𝗶𝘀𝗶𝗼𝗻Semi-Supervised Learning
SN

Sahar Nasirihaghighi

Doctoral Candidate, Klagenfurt University, Austria
Deep LearningComputer VisionMedical Video Analysis
QY

Qian Yu

Professor, Dept of Earth, Geographic, and Climate Sciences, University of Massachusetts-Amherst
GISremote sensingSpatial modeling
TW

Tianyang Wang

University of Alabama at Birmingham
machine learning (deep learning)computer vision
XX

Xiao Xiang Zhu

Technical University of Munich
Earth ObservationAI4EOSignal ProcessingData Science