When to Trust Confidence Thresholding: Calibration Diagnostics for Pseudo-Labelled Regression

πŸ“… 2026-05-12
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This work proposes a novel representation learning framework that addresses the limited representational capacity of existing methods in complex scenes by integrating adaptive multi-scale fusion with contrastive learning. The approach dynamically aggregates multi-level features and incorporates a structure-aware contrastive loss, thereby enhancing the model’s ability to jointly capture fine-grained semantics and global contextual information. Extensive experiments demonstrate that the proposed framework consistently outperforms state-of-the-art methods across multiple benchmark datasets, achieving substantial improvements in both accuracy and robustness. These results establish a promising new direction for unsupervised and semi-supervised representation learning.
πŸ“ Abstract
Calibrated probability outputs of trained classifiers are increasingly used as inputs to downstream regression estimands such as effects, prevalences, or disparities for a latent group observed only on a small labelled subset. A standard practice is to threshold the calibrated score at a confidence cutoff and treat the hard label as the truth. Building on a recent identification result for the underlying moment equation, we develop a calibration-aware diagnostic apparatus for pseudo-labelling pipelines. We derive a closed-form expression for the attenuation bias that confidence thresholding induces in the downstream regression coefficient, and show that the bias can be predicted, before any inference is run, from the residual score variance $V^{*}=\mathbb{E}[\operatorname{Var}(p\mid X)]$ on the unlabelled set after partialling out the downstream controls $X$. We further obtain a sharp sensitivity bound under bounded calibration drift, and identify the boundary $V^{*}=0$, which holds iff $p$ is a deterministic function of $X$; this motivates a structural separation between classifier features $W$ and downstream controls $X\subsetneq W$. Five controlled simulations and a UCI Adult illustration trace the predictions. The contribution is operational: a $(V^{*}, ΞΊ)$ decision rule that practitioners can compute from any classifier output to decide whether confidence thresholding is safe.
Problem

Research questions and friction points this paper is trying to address.

confidence thresholding
pseudo-labeling
calibration
regression
attenuation bias
Innovation

Methods, ideas, or system contributions that make the work stand out.

confidence thresholding
calibration diagnostics
pseudo-labeling
attenuation bias
residual score variance
πŸ”Ž Similar Papers
πŸ’Ό Related Jobs
No related jobs found.
M
Marcell T. Kurbucz
Institute for Global Prosperity, The Bartlett, University College London, 9–11 Endsleigh Gardens, London, WC1H 0EH, United Kingdom