Should This Case Be Adapted? Prediction Fragmentation Controls Test-Time Adaptation

๐Ÿ“… 2026-09-17
๐Ÿ“ˆ Citations: 0
โœจ Influential: 0
๐Ÿ“„ PDF
๐Ÿค– AI Summary
่ฎบๆ–‡ๆๅ‡บไฝฟ็”จ้ข„ๆต‹็ขŽ็‰‡ๅŒ–ๆ–นๆณ•ๆฅ่งฃๅ†ณๆต‹่ฏ•ๆ—ถ้€‚ๅบ”ๆ€ง้—ฎ้ข˜๏ผŒ้€š่ฟ‡ๆต‹้‡ๅˆๅง‹ๆจกๅž‹ไธŽ้€‚ๅบ”ๅŽๆจกๅž‹ไน‹้—ด็š„ไธไธ€่‡ดๆ€ง๏ผŒไปฅๅ‡ๅฐ‘ๆœ‰ๅฎณๆŽฅๅ—ๅŒบๅŸŸ๏ผŒๆ้ซ˜้€‚ๅบ”ๆ•ˆๆžœใ€‚
๐Ÿ“ Abstract
Episodic test-time adaptation resets a frozen segmenter to source weights $M_0$ on each case and adapts for a fixed step count. A fixed horizon conflates a cohort-level question, how far to adapt, with an irreducibly per-case one, whether this case should be adapted at all. Cohort means hide that decision: on cross-vendor cardiac MRI the mean $ฮ”$Dice from adaptation is statistically indistinguishable from zero while 58.7% of cases are individually made worse. We quantify this harm as harmful accepted area (HA), the harmful fraction of the edited area a controller deploys. Held-out tuning gives a stronger baseline than a fixed horizon, but the budget it selects transfers on neither of the two main medical benchmarks, and no global budget can condition on the case. We show that prediction fragmentation---the disagreement geometry between $M_0$ and the adapted mask $M_k$---predicts HA with no labels or extra backward passes at decision time, comparably on three benchmarks (Spearman $ฯ$ 0.50--0.60), at a quarter of gradient-norm's latency. A case-level router built on it cuts HA from 0.228 to 0.139 on a benchmark that took no part in its design, with the design frozen and only cut-points recalibrated there. On the cardiac benchmark the design was selected on, the router cuts HA from 0.129 to 0.013 at matched Dice and 1.10 deployed updates, against the retrospective-best budget found post hoc on evaluation labels, and reduces that 58.7% to 20.0%, an upper bound we quantify. Where the retained cases are not net-helped (as on prostate), the router still cuts HA but concedes accuracy, a boundary we report. Thresholds are fit once on a labeled split disjoint from evaluation; decisions use no labels or gradients. The template ports across architecture and domain (nnU-Net$\to$SegFormer, Cityscapes$\to$ACDC) with coordinate, thresholds and per-bucket actions instantiated per domain.
Problem

Research questions and friction points this paper is trying to address.

test-time adaptation
prediction fragmentation
case-level decision
Innovation

Methods, ideas, or system contributions that make the work stand out.

prediction fragmentation
test-time adaptation
harmful accepted area (HA)
case-level router
cross-domain transfer
๐Ÿ”Ž Similar Papers
No similar papers found.