🤖 AI Summary
This study addresses the challenge of missing modalities caused by localized sensor failures in untrimmed, long-form first-person videos by proposing the MacJEPA model. The method reformulates modality absence as a temporally localized fault and innovatively introduces window-level modal dropout to simulate sensor failures. It further transforms the JEPA self-supervised masking mechanism into a supervised robustness alignment objective, achieving single-stage joint optimization through multimodal latent representation alignment. Evaluated on the Epic-Kitchens-100 and Epic-Sounds datasets, a single checkpoint enables robust audio-visual recognition without test-time adaptation. MacJEPA delivers strong performance under complete inputs while significantly outperforming existing baselines in scenarios involving missing modalities.
📝 Abstract
Audio-visual models improve egocentric action recognition by exploiting complementary cues, yet typically assume that both streams remain available at inference. Existing missing-modality methods operate on trimmed, single-event clips in which a stream is entirely present or absent, whereas real sensors fail and recover within long, untrimmed observations. We redefine egocentric modality missingness as temporally localized sensor outages within untrimmed, multi-event observations, with whole-clip absence as the limiting case. We introduce \textbf{MacJEPA}, a missing-modality-robust \textbf{Ma}sked-\textbf{c}ontext query \textbf{JEPA} that recognizes visual actions and acoustic events from supplied interval queries over audio-visual context. Window-local modality dropout simulates these sensor outages during training. MacJEPA further repurposes masking in JEPA from a self-supervised pretext into a supervised robustness objective, aligning masked and clean latent representations of both multimodal content tokens and the task-conditioned queries. All objectives are optimized jointly with recognition in a single stage, requiring no test-time adaptation. Across Epic-Kitchens-100 and Epic-Sounds, a single checkpoint remains competitive under complete input and consistently surpasses published missing-modality baselines when either the dominant or auxiliary stream is removed. MacJEPA thus unifies strong full-input recognition with temporal missing-modality robustness in a single model operating on untrimmed multi-event videos.