🤖 AI Summary
This study investigates whether routing entropy in attention-residual Transformers provides predictive uncertainty information beyond model confidence. Employing Swin-Tiny and DeiT-Small architectures, we construct a rigorously calibrated statistical auditing framework incorporating soft-binned calibration auxiliary losses, paired experiments, and effect injection techniques. By quantifying estimator recovery rates to delineate evidential boundaries for null findings, we systematically evaluate the incremental value of dynamic routing trajectories for predicting correctness. Empirical results demonstrate that routing signals fail significance tests and yield no gains independent of confidence, while confirming that control-dependent gains do exist. This work furnishes rigorous negative evidence and a methodological reference for uncertainty quantification in dynamic architectures.
📝 Abstract
Dynamic architectures leave a per-example routing trace beside each prediction, and diffuse routing is easy to read as a sign that the prediction is unreliable. We audit that reading for routing entropy in Attention-Residual (AR) variants of Swin-Tiny and DeiT-Small, trained from scratch on CIFAR-10/100 with a soft-binned calibration auxiliary loss, asking whether the trace carries information about correctness beyond what the model's own confidence already reveals. Three checks probe this increment: does a routing signal appear at fixed confidence, does it replicate across training seeds, and can a held-out predictor exploit it against output-only and shuffled-trace controls? A sensitivity audit then injects effects of known size and measures the fraction of each that the probes recover. No test in the fixed 30-test binned family survives multiplicity correction, and neither the nominal hit nor a borderline result recurs in its sibling seeds. Across 24 paired runs a scalar routing probe yields no pooled improvement in routing-stratified calibration, and an entropy-profile probe predicts correctness better than the same probe given shuffled profiles yet worse than a confidence-only predictor in both binary log-loss and Brier score: a gain over shuffled traces does not become a gain over the output. Conditioning on the complete logit vector leaves the corresponding comparison unresolved. The audit bounds how far these non-detections can be read: at an injected effect of 0.010 nats the profile probe recovers 24-59% of the oracle gain, and a reference-preserving correction probe recovers 8% and 23% in the two CIFAR-100 settings, below the threshold we fixed for applying it to real labels. The results establish control-dependent gains and incomplete estimator recovery, not the absence of conditional routing information.