🤖 AI Summary
This study addresses the challenge of classifier adaptation in cross-domain few-shot learning, where target-domain parameters cannot be updated. To this end, we propose the WIPT model, which introduces a novel single-query test-time prototype adaptation mechanism. By jointly transforming support set embeddings to dynamically construct prototypes, WIPT enables streaming test-time adaptation without parameter optimization under a frozen Vision Transformer encoder and prototype Transformer architecture. Controlled comparative experiments demonstrate that the proposed method significantly improves classification accuracy while substantially reducing memory consumption on the CUB and EuroSAT datasets, although performance fluctuates on ISIC. This work establishes a new paradigm for optimization-free decision-making in resource-constrained scenarios.
📝 Abstract
Cross-domain few-shot learning requires adapting a classifier to a new visual domain from very few labelled examples without target-time parameter updates. We isolate one question: under a fixed global representation, what does joint query-support adaptation contribute to prototype construction? The Within-Instance Prototypical Transformer (WIPT) implements single-query test-time prototype adaptation by jointly transforming one unlabelled query and the labelled support embeddings, then forming query-specific class means. Using a shared frozen ViT-S/16 encoder, miniImageNet source training, and CUB, EuroSAT and ISIC targets, we replicate the key comparisons across five independent training seeds. In 1-shot evaluation, WIPT improves frozen ProtoNet in every run on CUB (+0.21 percentage points) and EuroSAT (+2.07), but decreases ISIC (-0.22). In 5-shot evaluation, ProtoNet remains strongest overall, while WIPT consistently improves a capacity-matched support-only Transformer on ISIC (+0.99). Joint processing of up to five queries yields no reliable accuracy gain; in a head-only 5-shot benchmark, g = 5 reduces analytical attention-token pairs by 73% and peak allocated memory by 29% relative to g = 1, although latency is non-monotonic. Across all target/shot conditions, WIPT changes uncertain ProtoNet decisions far more than confident ones, and rescue/break decomposition accounts for the observed gains and losses. Source-shift and scorer controls further show that the benefit is not universal. Overall, WIPT provides a streaming-compatible form of test-time prototype adaptation that can improve difficult low-shot cross-domain decisions without target-time optimization.