🤖 AI Summary
This study addresses the performance degradation of existing multimodal models under distribution shifts, which stems from the lack of explicit control over intermediate representation misalignment. To this end, we propose FIRE, a novel framework that pioneers a representation engineering perspective by directly editing semantic features in the intermediate layers of unimodal encoders to achieve inter-layer calibration and cross-modal alignment. Its core innovations include introducing frequency-domain mixing to construct a structured low-rank subspace for efficient representation editing, alongside a multi-level adaptive loss function for joint optimization. By integrating unsupervised test-time adaptation techniques, FIRE significantly outperforms prevailing mainstream methods across multiple benchmarks and diverse corruption types, effectively enhancing both the online adaptability and predictive reliability of multimodal models.
📝 Abstract
Multimodal test-time adaptation (TTA) aims to adapt a pretrained multimodal model online to distribution shift across modalities using unlabeled test data, showing broad potential in real-world applications. However, existing methods primarily focus on adjusting fused features to bridge the source-target gap, lacking explicit control over intermediate representation misalignment, which is a key driver of performance drop under distribution shift. In this work, we tackle this challenge from the perspective of representation engineering. Unlike previous TTA methods that update fusion weights in place, we propose FourIer Representation Editor (FIRE), a novel multimodal TTA approach that directly edits semantically rich intermediate representations. Specifically, we first adopt representation editors into each intermediate layer of the unimodal encoders, enabling layer-wise calibration of unimodal representations. To further enhance the diversity and stability of the low-rank editing subspaces, each representation editor performs frequency domain mixing via the fast Fourier transform to construct structured bases. Moreover, we introduce multi-level adaptation objectives to optimize these editors, jointly promoting cross-modal semantic alignment, source-target statistical alignment, and asymmetric prediction consistency. In this way, FIRE yields aligned unimodal representations for fusion and further improves prediction reliability. Extensive experiments on two widely used multimodal benchmarks under various corruption types demonstrate the superiority of FIRE over existing multimodal TTA methods.