SEAL: Mixture-Closed Additive Reconstruction and Refinement-Aware Expert Routing for Efficient Speech Separation
This study addresses the estimation distortion of multiplicative masks at signal cancellation points in speech separation, as well as the linear computational growth caused by shared units. To this end, we propose the SEAL framework, which introduces a novel hybrid-closed zero-sum additive residual reconstruction mechanism to resolve signal cancellation. Furthermore, it incorporates a dynamic sparse expert routing strategy based on acoustic and stepwise evidence, combined with local magnitude constraints and norm upper-bound control to enable efficient inference. Experimental results on the EchoSet dataset demonstrate that the compact SEAL model outperforms TIGER by 0.31 dB in SI-SDRi while reducing parameter count by 28%. Additionally, the larger model achieves near state-of-the-art performance with substantially lower computational costs.