Do More Modalities Always Help? A Geometric Perspective on Missing-Modality Robustness

📅 2026-10-03
📈 Citations: 0
✨ Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the robustness degradation problem in multimodal models under missing modalities, where cross-modal dependencies cause performance to fall below that of unimodal counterparts. To tackle this, we provide the first geometric perspective on this failure mechanism and propose a novel approach termed "geodesic forgetting." Specifically, our method rotates subspaces along geodesic paths on the Grassmannian manifold, effectively eliminating detrimental cross-modal dependencies through lightweight parameter editing. Experimental results demonstrate that the proposed approach significantly improves inference performance when modalities are missing while preserving full-modality accuracy without degradation. Overall, our method achieves superior performance compared to existing baselines.
📝 Abstract
Missing modality remains a longstanding challenge in multimodal learning. Existing methods typically address this issue through modality recovery or adaptive strategies. However, they overlook models'internal cross-modal dependencies formed during multimodal training, which later impair robustness. We systematically characterize a counterintuitive deployment-time failure mode: models trained on full modalities can underperform unimodal models when one modality is missing at inference time. This pattern appears across diverse architectures, such as fusion models, CLIP-style two-tower models, and vision-language models. We show that such degradation is closely associated with learned cross-modal dependencies in the principal parameter subspaces. Multimodal training induces structured rotations of these subspaces, particularly in cross-modal interaction layers. These rotations are associated with reduced task-aligned margins and larger task-aware representation harm under missing-modality inputs. We propose Geodesic Unlearning (GU), a lightweight parameter-editing method that leverages Grassmannian subspace geometry for structured subspace correction to improve missing-modality robustness. It rotates the principal input subspace toward a unimodal reference along a geodesic path. We prove that this correction minimizes the distance to the reference within a fixed subspace-distance budget. Experiments across architectures and datasets show that GU improves performance under missing-modality inference while preserving full-modality accuracy, outperforming strong missing-modality robustness baselines. These findings support a geometric view of deployment-time missing-modality degradation and suggest localized subspace editing as a practical route for robustness correction.
Problem

Research questions and friction points this paper is trying to address.

missing modality
multimodal learning
cross-modal dependencies
robustness
modality degradation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Missing-Modality Robustness
Geodesic Unlearning
Grassmannian Geometry
Subspace Correction
Cross-Modal Dependencies
🔎 Similar Papers
No similar papers found.