🤖 AI Summary
This study investigates how audio preprocessing affects the performance of pretrained-model-based stuttering detection. Leveraging the SEP-28k dataset and the WavLM Base+ model, it provides the first systematic quantification of the differential impacts of preprocessing pipelines—including denoising and loudness normalization—across five stuttering categories. Evaluation via bootstrap resampling with ROC-AUC and Average Precision metrics reveals that preprocessing significantly degrades the F1 score for block-type stuttering. Although threshold tuning plays a dominant role in recovering the F1 score to 0.630, it fails to compensate for losses in ROC-AUC. By delineating the efficacy boundaries of threshold optimization versus retraining strategies, this work offers critical empirical evidence for developing robust speech fluency detection systems.
📝 Abstract
Audio preprocessing can affect how well a system detects stuttering. We study a simulated chain of denoising,loudness normalisation, Opus coding, and voice activity detection on SEP-28k. We use frozen WavLM Base+ features and report pointwise confidence intervals from episode-level bootstrap resampling. At a fixed threshold of 0.5, the chain reduces block F1 from 0.638 to 0.465, with smaller decreases for the other four classes. ROC-AUC decreases for all five classes. Blocks show the largest F1 and ROC-AUC losses, while sound repetitions show the largest average precision loss. Tuning the threshold on processed validation audio raises block F1 to 0.630. Retraining on processed audio with threshold tuning gives 0.628. Threshold adjustment therefore accounts for most of the observed block F1 recovery. It does not change ROC-AUC, which retraining raises only from 0.620 to 0.633, compared with 0.724 on clean audio.