RADC: Risk-Aware Dual Caching for Vision-Language Test-Time Adaptation

📅 2026-10-02
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses two critical challenges in test-time adaptation for vision-language models: contextual bias in global representations and the unreliability of entropy-based cache admission. To overcome these limitations, this work proposes the RADC framework, which introduces a semantic foreground cache that aggregates spatial evidence to suppress background interference. Furthermore, it pioneers a multi-view risk admission strategy based on diagonal Gaussian distributions to manage dual caches, integrating zero-shot and cache-based predictions for robust inference. This approach effectively resolves cache reliability issues under representation shifts. Extensive experiments demonstrate that the proposed method achieves state-of-the-art performance on both cross-domain and out-of-distribution benchmarks, substantially enhancing model robustness.
📝 Abstract
Cache-based test-time adaptation (TTA) for vision-language models is often hindered by background bias in global representations and unreliable entropy-based cache admission under representation variations. To address these limitations, we propose RADC, which enhances prototype learning through reliable dual caching. RADC introduces a Semantic Foreground Cache that aggregates category-consistent spatial evidence from CLIP representations, yielding foreground prototypes that complement the global cache while mitigating background interference. To reliably manage both caches, Gaussian Risk Admission models multi-view representations as diagonal Gaussian distributions and jointly considers class separation and feature uncertainty to prioritize reliable cache candidates. RADC integrates zero-shot logits with complementary global- and foreground-cache predictions for robust inference. Extensive experiments on cross-domain and out-of-distribution benchmarks demonstrate consistent state-of-the-art performance.
Problem

Research questions and friction points this paper is trying to address.

Test-Time Adaptation
Vision-Language Models
Background Bias
Cache Admission
Representation Variation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Test-Time Adaptation
Dual Caching
Vision-Language Models
Foreground Prototype
Gaussian Risk Admission
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Siyu Huang
Siyu Huang
Assistant Professor, Clemson University
computer visionmachine learninggenerative models
Y
Yueyong Chen
School of Intelligent Systems Engineering, Sun Yat-sen University, Shenzhen, China
X
Xuejiao Li
Pengcheng Laboratory, Shenzhen, China
J
Jun Zhou
Tsinghua Shenzhen International Graduate School, Tsinghua University, Shenzhen, China