🤖 AI Summary
This study addresses the computational inefficiency of existing audio-visual segmentation models caused by architectural redundancy. To this end, we propose EASE, an extremely minimalist encoder-only framework for audio-visual semantic segmentation. By discarding redundant components inherent in conventional encoder-decoder architectures, EASE leverages efficient multimodal fusion to achieve cross-modal semantic alignment and precise segmentation within a highly lightweight design. Experimental results demonstrate that EASE maintains state-of-the-art accuracy while attaining an inference speed of 365 FPS—approximately three times faster than existing methods. This work establishes a new paradigm for real-time multimodal perception, effectively reconciling high performance with low latency.
📝 Abstract
Audio-Visual Semantic Segmentation (AVSS) aims to identify, segment, and classify sound-emitting objects in video frames. Previous Transformer-based AVSS approaches largely inherit design principles from image segmentation models. Recent studies show that these image segmentation models contain redundant components that contribute little to the segmentation performance. Following this insight, we propose Encoder-only Audio-Visual Segmentation (EASE). EASE runs at up to 365 FPS, 3x faster than prior State-of-the-Art (SotA) AVS models at comparable accuracy, and trains in under 11 GPU-hours. Furthermore, we achieve SotA AVSS performance across different backbones and input resolutions. Our results demonstrate that AVSS can be both simpler and faster, providing a scalable foundation for future research and real-time applications. Code, model weights, and samples are available at https://ease-avs.notion.site