🤖 AI Summary
This study addresses the limitations of static offline configurations in adapting to dynamic leakage during speech privacy protection, as well as the reliance of conventional metrics on ground-truth references that impedes online evaluation. To this end, we propose CLEAR, a method that leverages the divergence among heterogeneous ASR systems as a reference-free proxy metric. Since independent models exhibit behavioral consistency when speech content is recoverable but diverge significantly upon leakage, this approach enables real-time leakage estimation and adaptive privacy modulation without requiring reference transcripts. Experimental results demonstrate that the proposed metric achieves a 0.8 correlation with actual leakage and accurately identifies potentially exposed words. This work pioneers treating privacy as a dynamically optimizable runtime property, delivering effective protection while preserving acoustic utility.
📝 Abstract
Signal-level speech privacy mechanisms suppress linguistic content while preserving acoustic information needed by downstream sensing applications. However, their privacy settings are typically evaluated/selected offline and remain fixed during deployment, even though speech-content leakage can vary substantially across utterances and speakers. Adapting privacy protection at runtime requires estimating how much speech remains recoverable, but conventional measures such as WER or PER require ground-truth transcripts and therefore cannot be computed online. We present CLEAR, a reference-free approach for estimating speech-content leakage at runtime using disagreement among heterogeneous ASR systems. Our key insight is that independently trained ASRs exhibit consistent behavior when linguistic content remains recoverable and increasingly disagree as privacy transformations obscure speech. Using configurable speech-suppression mechanism, we show that cross-ASR disagreement closely tracks transcript-grounded leakage across privacy operating points, achieving a correlation of 0.8. We further characterize the latency-accuracy trade-off of heterogeneous ASR subsets and use their hypotheses to identify potentially exposed words. These capabilities enable privacy to be treated as a runtime property rather than a fixed configuration: CLEAR can communicate residual speech exposure to users and provide feedback for dynamically adjusting privacy aggressiveness while retaining acoustic utility for downstream sensing tasks.