Representation Editing for Multimodal Test-Time Adaptation

📅 2026-09-26
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the performance degradation of existing multimodal models under distribution shifts, which stems from the lack of explicit control over intermediate representation misalignment. To this end, we propose FIRE, a novel framework that pioneers a representation engineering perspective by directly editing semantic features in the intermediate layers of unimodal encoders to achieve inter-layer calibration and cross-modal alignment. Its core innovations include introducing frequency-domain mixing to construct a structured low-rank subspace for efficient representation editing, alongside a multi-level adaptive loss function for joint optimization. By integrating unsupervised test-time adaptation techniques, FIRE significantly outperforms prevailing mainstream methods across multiple benchmarks and diverse corruption types, effectively enhancing both the online adaptability and predictive reliability of multimodal models.
📝 Abstract
Multimodal test-time adaptation (TTA) aims to adapt a pretrained multimodal model online to distribution shift across modalities using unlabeled test data, showing broad potential in real-world applications. However, existing methods primarily focus on adjusting fused features to bridge the source-target gap, lacking explicit control over intermediate representation misalignment, which is a key driver of performance drop under distribution shift. In this work, we tackle this challenge from the perspective of representation engineering. Unlike previous TTA methods that update fusion weights in place, we propose FourIer Representation Editor (FIRE), a novel multimodal TTA approach that directly edits semantically rich intermediate representations. Specifically, we first adopt representation editors into each intermediate layer of the unimodal encoders, enabling layer-wise calibration of unimodal representations. To further enhance the diversity and stability of the low-rank editing subspaces, each representation editor performs frequency domain mixing via the fast Fourier transform to construct structured bases. Moreover, we introduce multi-level adaptation objectives to optimize these editors, jointly promoting cross-modal semantic alignment, source-target statistical alignment, and asymmetric prediction consistency. In this way, FIRE yields aligned unimodal representations for fusion and further improves prediction reliability. Extensive experiments on two widely used multimodal benchmarks under various corruption types demonstrate the superiority of FIRE over existing multimodal TTA methods.
Problem

Research questions and friction points this paper is trying to address.

Multimodal Test-Time Adaptation
Representation Misalignment
Distribution Shift
Intermediate Representation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Multimodal Test-Time Adaptation
Representation Editing
Fast Fourier Transform
Low-rank Subspace
Cross-modal Alignment
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
L
Longfei Huang
Nanjing University of Science and Technology
X
Xiangyu Wu
Alibaba Group
Yang Yang
Yang Yang
Nanjing University of Science and Technology
Data MiningMulti Modal LearningIncremental Learning