🤖 AI Summary
This study addresses the uncertainty regarding how different retrieval mechanisms affect downstream task performance when Vision-Language-Action (VLA) models adapt at test time by retrieving expert demonstrations for in-context learning. Building upon the RICL framework, this work systematically compares four retrieval strategies—image-based, state-augmented, backbone-feature-based, and random—evaluating their adaptation efficacy and diagnostic reliability through techniques such as retrieval-augmented generation and multimodal feature extraction. The findings reveal that no single retrieval method is universally optimal, with the random baseline demonstrating notable robustness. Furthermore, standard retrieval metrics fail to predict downstream performance, whereas cross-task demonstrations exhibit meaningful transferability. By elucidating the complex influence of retrieval mechanisms on VLA adaptation capabilities, this research provides a theoretical foundation for designing reliable test-time adaptation modules.
📝 Abstract
Vision-language-action (VLA) models have shown strong potential as generalist robot policies, but adapting them to unseen tasks often requires costly parameter updates. Recent work such as RICL introduces in-context adaptability by retrieving expert demonstrations based on the current VLA observation and providing them as additional context at test time. The effectiveness of this adaptation therefore depends critically on the retrieval mechanism. In this work, we systematically study how different retrieval methods affect both retrieval quality and task performance within the RICL framework. Specifically, we compare four different methods: image-based retrieval, retrieval augmented with VLA's state, retrieval using features from the VLA backbone, and random retrieval. Our experiments yield three main findings. First, no retrieval method consistently dominates the others in task success, while surprisingly, random retrieval achieves a non-trivial success rate. Second, standard retrieval-quality diagnostics do not reliably reflect downstream VLA performance. Third, demonstrations from different but related tasks can provide useful transferable information. Together, these results provide an initial step toward understanding how retrieval mechanisms shape the in-context learning capability of VLA models and their downstream task performance, while highlighting the need for more careful design and evaluation of retrieval mechanisms for reliable test-time adaptation.