🤖 AI Summary
This study addresses the infeasibility of frame-by-frame domain adaptation for open-vocabulary segmentation on resource-constrained robots. To this end, we propose a training-free continuous test-time adaptation framework driven by multi-signal domain shift detection. By integrating multiple source signals—including visual features, adapter states, and semantic drift monitoring—and leveraging temporal coherence to consolidate complementary shift information, our method realizes an efficient, on-demand adaptation mechanism. Experimental results demonstrate that the proposed framework substantially reduces the number of required adaptations while preserving segmentation accuracy, thereby significantly lowering computational overhead. Ultimately, this approach enables the long-term deployment of open-vocabulary segmentation models on real-world robotic platforms.
📝 Abstract
Robust and reliable perception is essential for autonomous robots operating in real-world environments, particularly in long-term missions where environmental conditions may change significantly over time. Although recent advances in Visual Foundation Models (VFMs) have improved open-vocabulary semantic segmentation, these models can still suffer from domain shift, which can significantly degrade performance if they are not adapted to the current environment. Training-free domain adaptation is a relevant paradigm for adaptation, consisting of adjusting the model online using lightweight adapters. Recent approaches apply this on a per-frame basis, which is impractical for deployments on resource-constrained robotic hardware. To tackle this, we propose a multi-signal domain shift detection method for training-free continual test-time adaptation (TF-CTTA) in open-vocabulary segmentation. Our method leverages temporal coherence across consecutive frames by monitoring and combining complementary aspects of domain shift (visual change, adapter mismatch, and semantic drift) to trigger adaptation only when needed. We validate our approach on a benchmark including indoor and outdoor environments and using real robotic data. We demonstrate that our approach maintains segmentation accuracy while substantially reducing adaptations, making training-free adaptation practical and feasible for long-term, real-world robotic deployments.