Institution profile

National Chi Nan University

Academic institutionasia · tw
Official website
Research library2linked papers
Opportunities0open roles
Selected work

Representative Papers

SEAL: Mixture-Closed Additive Reconstruction and Refinement-Aware Expert Routing for Efficient Speech Separation

Oct 04, 2026

This study addresses the estimation distortion of multiplicative masks at signal cancellation points in speech separation, as well as the linear computational growth caused by shared units. To this end, we propose the SEAL framework, which introduces a novel hybrid-closed zero-sum additive residual reconstruction mechanism to resolve signal cancellation. Furthermore, it incorporates a dynamic sparse expert routing strategy based on acoustic and stepwise evidence, combined with local magnitude constraints and norm upper-bound control to enable efficient inference. Experimental results on the EchoSet dataset demonstrate that the compact SEAL model outperforms TIGER by 0.31 dB in SI-SDRi while reducing parameter count by 28%. Additionally, the larger model achieves near state-of-the-art performance with substantially lower computational costs.

0 citationsRead paper

SincQDR-VAD: A Noise-Robust Voice Activity Detection Framework Leveraging Learnable Filters and Ranking-Aware Optimization

Aug 28, 2025

To address the poor robustness of voice activity detection (VAD) under noisy and resource-constrained conditions, and the misalignment between conventional classification losses and evaluation metrics such as AUROC, this paper proposes a compact, efficient end-to-end VAD framework. Methodologically: (i) a learnable Sinc bandpass filter is employed to construct a noise-robust spectral frontend, enhancing feature discriminability; (ii) a novel Quadratic Difference Ranking Loss is introduced to explicitly optimize the relative ranking of speech versus non-speech frames, thereby directly maximizing AUROC. Experiments on multiple benchmark datasets demonstrate consistent improvements—AUROC increases by 1.2–2.8% and F2-score by 3.5–5.1%—while the model requires only 69% of the parameters of current state-of-the-art methods. The proposed approach thus achieves superior accuracy, low inference latency, and high parameter efficiency.

0 citationsRead paper
Recent publications

Latest Papers

SEAL: Mixture-Closed Additive Reconstruction and Refinement-Aware Expert Routing for Efficient Speech Separation

Oct 04, 2026

This study addresses the estimation distortion of multiplicative masks at signal cancellation points in speech separation, as well as the linear computational growth caused by shared units. To this end, we propose the SEAL framework, which introduces a novel hybrid-closed zero-sum additive residual reconstruction mechanism to resolve signal cancellation. Furthermore, it incorporates a dynamic sparse expert routing strategy based on acoustic and stepwise evidence, combined with local magnitude constraints and norm upper-bound control to enable efficient inference. Experimental results on the EchoSet dataset demonstrate that the compact SEAL model outperforms TIGER by 0.31 dB in SI-SDRi while reducing parameter count by 28%. Additionally, the larger model achieves near state-of-the-art performance with substantially lower computational costs.

0 citationsRead paper

SincQDR-VAD: A Noise-Robust Voice Activity Detection Framework Leveraging Learnable Filters and Ranking-Aware Optimization

Aug 28, 2025

To address the poor robustness of voice activity detection (VAD) under noisy and resource-constrained conditions, and the misalignment between conventional classification losses and evaluation metrics such as AUROC, this paper proposes a compact, efficient end-to-end VAD framework. Methodologically: (i) a learnable Sinc bandpass filter is employed to construct a noise-robust spectral frontend, enhancing feature discriminability; (ii) a novel Quadratic Difference Ranking Loss is introduced to explicitly optimize the relative ranking of speech versus non-speech frames, thereby directly maximizing AUROC. Experiments on multiple benchmark datasets demonstrate consistent improvements—AUROC increases by 1.2–2.8% and F2-score by 3.5–5.1%—while the model requires only 69% of the parameters of current state-of-the-art methods. The proposed approach thus achieves superior accuracy, low inference latency, and high parameter efficiency.

0 citationsRead paper