🤖 AI Summary
This study addresses the high resource consumption and limited generalization of existing wake word detection systems, which typically rely on fine-tuning or dedicated modules. We propose a gradient-free fine-tuning paradigm leveraging a shared ASR backbone. Building upon Parakeet and Moonshine models, our approach integrates PCA with SliceGPT-based structured pruning to directly derive task-specific compact encoders from pretrained models without additional training. Experimental results demonstrate that wake word detection performance remains stable even when the encoder channel dimension is reduced by 50%. By effectively balancing detection efficiency with model generalization, this work offers a promising new direction for lightweight voice interaction systems.
📝 Abstract
Recent ASR development has placed growing emphasis on generalization across diverse domains and acoustic conditions. Existing approaches typically adapt pretrained ASR models to front-end functions such as wake-up word (WuW) detection through additional training or task-specific modules. In this work, we explore the use of a shared pretrained ASR backbone for WuW detection without gradient-based fine-tuning and examine whether a compact encoder can be extracted using the PCA-based structured pruning approach of SliceGPT. Experiments with Parakeet-TDT-0.6B-v3 and Moonshine-base show that WuW detection performance remains relatively stable when the encoder channel dimension is reduced by 50%. These results suggest that task-relevant compact encoders can be derived from pretrained ASR models without fine-tuning.