LiteCASS: A Lightweight End-to-End Network for Real-Time Stereo Cinematic Audio Source Separation

📅 2026-09-20
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
该研究提出LiteCASS网络,解决实时立体声电影音频源分离问题,采用轻量级端到端架构和多任务波形域L1损失方法。
📝 Abstract
Cinematic audio source separation (CASS) decomposes a soundtrack into dialogue, music, and sound-effects (SFX) stems. Existing CASS methods, however, suffer from two critical limitations: they rely on heavily parameterized network architectures and GPU-class hardware, limiting their use in real-time and resource-constrained scenarios, and they are overwhelmingly designed for monaural signals, leaving the stereo scenario largely unexplored. We present LiteCASS, to our knowledge the first lightweight end-to-end network for real-time stereo CASS. LiteCASS combines deterministic STFT subband rearrangement with two jointly trained compact U-Nets: the first extracts dialogue, and the second separates music and SFX from the predicted non-speech component. A multi-task waveform-domain L1 loss supervises all stems. On a spatialized stereo extension of DnR v3, LiteCASS-K8 uses only 1.06M parameters and 0.72G MACs per second, while achieving the highest averaged SI-SDR among the compared CASS baselines.
Problem

Research questions and friction points this paper is trying to address.

Cinematic Audio Source Separation
Real-Time
Resource-Constrained
Stereo
Innovation

Methods, ideas, or system contributions that make the work stand out.

real-time stereo CASS
lightweight end-to-end network
compact U-Nets
deterministic STFT subband rearrangement
multi-task waveform-domain L1 loss
🔎 Similar Papers
No similar papers found.