🤖 AI Summary
This study addresses the estimation distortion of multiplicative masks at signal cancellation points in speech separation, as well as the linear computational growth caused by shared units. To this end, we propose the SEAL framework, which introduces a novel hybrid-closed zero-sum additive residual reconstruction mechanism to resolve signal cancellation. Furthermore, it incorporates a dynamic sparse expert routing strategy based on acoustic and stepwise evidence, combined with local magnitude constraints and norm upper-bound control to enable efficient inference. Experimental results on the EchoSet dataset demonstrate that the compact SEAL model outperforms TIGER by 0.31 dB in SI-SDRi while reducing parameter count by 28%. Additionally, the larger model achieves near state-of-the-art performance with substantially lower computational costs.
📝 Abstract
Compact time-frequency separators that mask the mixture and refine through a shared cell face two limits. First, a bounded multiplicative mask only scales a mixture bin, so where overlapping components cancel, the estimate stays small. Second, a shared cell applies the same weights to every time-frequency token at every step, so enlarging it adds compute everywhere. We present SEAL (Sparse Expert routing with Additive Latent reconstruction) to address both. For reconstruction, a zero-sum additive residual bounded by the local mixture amplitude lets estimates be nonzero where components cancel yet still sum to the mixture. For routing, a query built from acoustic and inter-step evidence sends each token to one of six residual experts, and a norm cap keeps the step cue from overriding clear acoustic evidence. On EchoSet, SEAL (small) surpasses TIGER (small) by 0.31 dB SI-SDRi with 28% fewer parameters and 2.9 times fewer MACs, and SEAL (large) is within 0.07 dB SI-SDRi of TIGER (large) at 3.1 times fewer MACs.