🤖 AI Summary
This study addresses the limitation of existing voice activity detection (VAD) models in distinguishing breathing from silence, which frequently leads to missed respiratory events. We propose BreathGRU, a semi-supervised framework that extracts acoustic features via bidirectional gated recurrent units and integrates a pseudo-label refinement strategy with a duration-constrained Segmental Viterbi decoding algorithm. This approach enables explicit modeling of breathing events and achieves high-precision speech-breathing segmentation. Experimental results demonstrate that BreathGRU attains a breathing recall of 0.83, a boundary error of merely 0.14 seconds, and an Intersection over Union (IoU) of 0.81. These metrics significantly outperform mainstream pretrained models such as Silero, effectively overcoming the performance limitations inherent in general-purpose VAD systems.
📝 Abstract
Speech-breath segmentation is a fundamental preprocessing step in respiratory audio analysis, enabling applications such as respiratory acoustic biomarker extraction, lung function prediction and disease monitoring. Existing approaches, including threshold methods, Fourier Transform-based techniques, and unsupervised and pretrained voice activity detection (VAD) models, primarily focus on speech detection and often classify breathing events as non-speech or silence, limiting their applicability for precise breath detection. To address this limitation, we propose BreathGRU, a semi-supervised Bidirectional Gated Recurrent Unit (BiGRU) framework specifically designed for speech-breath segmentation. The proposed framework combines frame-level acoustic feature extraction with bidirectional recurrent modelling, pseudo-label refinement and duration-constrained Segmental Viterbi decoding to produce speech and breath segmentation. BreathGRU was evaluated against the existing approaches, using manually annotated recordings. Performance was assessed using event-based, time-based, overlap-based, duration-based and boundary-based segmentation metrics. Experiment results demonstrated that BreathGRU achieved the highest breath event recall (0.83), the lowest onset-localisation error (0.14s) and the highest Mean Match Intersection over Union (0.81), with competitive overall segmentation performance compared to large pretrained VAD models like Silero. Qualitative evaluation on manually annotated recordings further showed close agreement between BreathGRU and manual annotation, with better breath detection compared to Silero. These findings demonstrate that explicit breath event modelling provides advantages over general-purpose VAD models and establish BreathGRU as an effective speech-breath segmentation framework which can be applied for respiratory audio analysis and pulmonary healthcare applications.