🤖 AI Summary
This study addresses two critical challenges in test-time adaptation for vision-language models: contextual bias in global representations and the unreliability of entropy-based cache admission. To overcome these limitations, this work proposes the RADC framework, which introduces a semantic foreground cache that aggregates spatial evidence to suppress background interference. Furthermore, it pioneers a multi-view risk admission strategy based on diagonal Gaussian distributions to manage dual caches, integrating zero-shot and cache-based predictions for robust inference. This approach effectively resolves cache reliability issues under representation shifts. Extensive experiments demonstrate that the proposed method achieves state-of-the-art performance on both cross-domain and out-of-distribution benchmarks, substantially enhancing model robustness.
📝 Abstract
Cache-based test-time adaptation (TTA) for vision-language models is often hindered by background bias in global representations and unreliable entropy-based cache admission under representation variations. To address these limitations, we propose RADC, which enhances prototype learning through reliable dual caching. RADC introduces a Semantic Foreground Cache that aggregates category-consistent spatial evidence from CLIP representations, yielding foreground prototypes that complement the global cache while mitigating background interference. To reliably manage both caches, Gaussian Risk Admission models multi-view representations as diagonal Gaussian distributions and jointly considers class separation and feature uncertainty to prioritize reliable cache candidates. RADC integrates zero-shot logits with complementary global- and foreground-cache predictions for robust inference. Extensive experiments on cross-domain and out-of-distribution benchmarks demonstrate consistent state-of-the-art performance.