Signal-Routed Temperature Scaling: Low-Capacity Risk-Conditioned Calibration for Small Validation Budgets

📅 2026-09-30
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the capacity selection dilemma in classifier calibration under small validation budgets, where scalar methods tend to underfit while high-capacity approaches are difficult to estimate accurately. To overcome this, we propose SRTS-BCE, a method that achieves low-capacity adaptive calibration by cross-fitting risk scores and fitting temperature scaling within groups. This work decouples the calibration objective from model capacity, establishing calibration capacity as a finite-sample design choice that varies with data volume, while integrating signal routing and logit statistics for risk-conditioned modeling. Empirical evaluations on CIFAR-100 demonstrate that SRTS-BCE significantly reduces the Expected Calibration Error (ECE), outperforming high-capacity methods under small validation budgets and achieving comparable performance when larger budgets are available.
📝 Abstract
When a classifier is recalibrated from only a few thousand held-out examples, the capacity of the calibration map becomes a statistical design choice rather than a purely architectural one: a scalar map can underfit structured residual miscalibration, while a highly adaptive map can be hard to estimate reliably from so small a split. We disentangle the calibration objective from adaptive capacity and propose signal-routed temperature scaling (SRTS-BCE), a 10-parameter, argmax-preserving calibrator that cross-fits a correctness-risk score over six logit statistics and fits one top-label-BCE temperature per $K=3$ risk groups, recovering TvA-TS as its $K=1$ limit. On fine-tuned CIFAR-100 / ViT-B/16, SRTS-BCE reduces $\mathrm{ECE}_{15}$ from 1.65 (scalar TvA-TS) to 0.96, matching the higher-capacity SMART+BCE head (0.95) at the full calibration budget. The two regimes separate as the budget shrinks: at $n=250$ SRTS-BCE beats SMART+BCE on all three CIFAR-100 backbones (the seed-to-draw hierarchical interval excludes zero), whereas the flagship comparison against the scalar remains directional. A protocol-frozen Tiny-ImageNet follow-up reproduces the small-budget separation and exhibits a budget-dependent ranking reversal on Swin-T; matched routing and map controls show that the effect is tied neither to the learned router nor to discrete grouping. Together the results identify post-hoc calibrator capacity as a finite-sample design choice whose preferred level shifts with the amount of available calibration data.
Problem

Research questions and friction points this paper is trying to address.

post-hoc calibration
small validation budget
calibrator capacity
temperature scaling
confidence estimation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Signal-Routed Temperature Scaling
Post-hoc Calibration
Risk-Conditioned Calibration
Small Validation Budgets
Capacity Disentanglement
🔎 Similar Papers
No similar papers found.