Same Target, Different Basins: Hard vs. Soft Labels for Annotator Distributions

๐Ÿ“… 2026-05-19
๐Ÿ“ˆ Citations: 0
โœจ Influential: 0
๐Ÿ“„ PDF
๐Ÿค– AI Summary
This work addresses annotator disagreement by proposing hard-label strategiesโ€”such as multipass annotation and stochastic label sampling (SLS)โ€”as alternatives to conventional soft-label training, thereby leveraging the full annotation distribution rather than treating it as noise. Theoretical analysis grounded in cross-entropy optimization reveals that hard-label methods converge to flatter loss basins under label sparsity, enhancing generalization. Empirical results demonstrate that these approaches significantly outperform soft-label methods on CIFAR-10H, match SLS performance under fully annotated settings, and exhibit superior out-of-distribution detection capabilities on SVHN and CIFAR-100.
๐Ÿ“ Abstract
When annotators disagree, that disagreement can reflect epistemic uncertainty rather than simple label noise. We study hard-label delivery as an alternative to the usual choices of collapsing votes to a single label or training directly on the empirical soft-label distribution. We focus on two primary hard-label methods: multipass, which cycles through observed votes while keeping the dataset size fixed, and stochastic label sampling (SLS), which samples one label per example at the start of each epoch. On CIFAR-10H, we find that when only a small number of annotations per example is available, hard-label delivery improves over soft-label training, with larger improvements where the sparse empirical target is farther from the full annotator distribution. When full annotator distributions are available, both hard-label methods match soft-label training. We use deterministic control as an ablation of multipass and shuffled SLS as a control that breaks the example-to-distribution match. We also show that SLS and soft-label cross-entropy optimize the same expected objective. Hard-label delivery also converges to flatter basins, with supporting descriptive evidence from OOD detection on SVHN and CIFAR-100. Overall, these results suggest that multipass is a strong practical default when raw vote counts are available, while SLS offers a lightweight alternative that remains competitive when only a few votes per example are available and matches soft-label training when full annotator distributions are available.
Problem

Research questions and friction points this paper is trying to address.

annotator disagreement
hard labels
soft labels
label uncertainty
annotation distribution
Innovation

Methods, ideas, or system contributions that make the work stand out.

hard-label delivery
stochastic label sampling
multipass
annotator disagreement
flatter loss basins
๐Ÿ’ผ Related Jobs
No related jobs found.
M
Mirerfan Gheibi
Independent Researcher
G
Gashin Ghazizadeh
Independent Researcher